<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.anunna.wur.nl/index.php?action=history&amp;feed=atom&amp;title=Ollama</id>
	<title>Ollama - Revision history</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.anunna.wur.nl/index.php?action=history&amp;feed=atom&amp;title=Ollama"/>
	<link rel="alternate" type="text/html" href="https://wiki.anunna.wur.nl/index.php?title=Ollama&amp;action=history"/>
	<updated>2026-10-09T03:55:21Z</updated>
	<subtitle>Revision history for this page on the wiki</subtitle>
	<generator>MediaWiki 1.45.3</generator>
	<entry>
		<id>https://wiki.anunna.wur.nl/index.php?title=Ollama&amp;diff=3213&amp;oldid=prev</id>
		<title>Haars001: Created page with &quot;== How to use ollama on the HPC == On the HPC we have &lt;code&gt;ollama&lt;/code&gt; available through modules, in the GPU bucket.  To load the module, do this:&lt;syntaxhighlight lang=&quot;bash&quot;&gt; ml GPU ollama &lt;/syntaxhighlight&gt;Once the module is loaded, &lt;code&gt;ollama&lt;/code&gt; is available on your path:&lt;syntaxhighlight lang=&quot;bash&quot;&gt; ollama --version Warning: could not connect to a running Ollama instance Warning: client version is 0.35.0 &lt;/syntaxhighlight&gt;  === Intera...&quot;</title>
		<link rel="alternate" type="text/html" href="https://wiki.anunna.wur.nl/index.php?title=Ollama&amp;diff=3213&amp;oldid=prev"/>
		<updated>2026-10-02T08:47:38Z</updated>

		<summary type="html">&lt;p&gt;Created page with &amp;quot;== How to use ollama on the HPC == On the HPC we have &amp;lt;code&amp;gt;ollama&amp;lt;/code&amp;gt; available through &lt;a href=&quot;/Environment_Modules&quot; title=&quot;Environment Modules&quot;&gt;modules&lt;/a&gt;, in the GPU bucket.  To load the module, do this:&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt; ml GPU ollama &amp;lt;/syntaxhighlight&amp;gt;Once the module is loaded, &amp;lt;code&amp;gt;ollama&amp;lt;/code&amp;gt; is available on your path:&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt; ollama --version Warning: could not connect to a running Ollama instance Warning: client version is 0.35.0 &amp;lt;/syntaxhighlight&amp;gt;  === Intera...&amp;quot;&lt;/p&gt;
&lt;p&gt;&lt;b&gt;New page&lt;/b&gt;&lt;/p&gt;&lt;div&gt;== How to use ollama on the HPC ==&lt;br /&gt;
On the HPC we have &amp;lt;code&amp;gt;ollama&amp;lt;/code&amp;gt; available through [[Environment Modules|modules]], in the GPU bucket.&lt;br /&gt;
&lt;br /&gt;
To load the module, do this:&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
ml GPU ollama&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;Once the module is loaded, &amp;lt;code&amp;gt;ollama&amp;lt;/code&amp;gt; is available on your path:&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
ollama --version&lt;br /&gt;
Warning: could not connect to a running Ollama instance&lt;br /&gt;
Warning: client version is 0.35.0&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Interactive session ===&lt;br /&gt;
To make working with ollama easier, we have created a small helper script, this will help you select a GPU, and set it up with the values you entered.&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
menu_ollama.sh&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;Once you have entered all the values, the script will start an interactive job (so with output and input in your shell), which will look something like this:&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
Now running:&lt;br /&gt;
srun --immediate=60 --nodes=1 --ntasks=1 --partition=gpu_amd --gres=gpu:1 --time=1:0:0 --mem-per-cpu=8G --cpus-per-gpu=8 /shared/apps/ollama/0.35.0/start_ollama.sh&lt;br /&gt;
&lt;br /&gt;
srun: job 43006080 queued and waiting for resources&lt;br /&gt;
srun: job 43006080 has been allocated resources&lt;br /&gt;
remove ollama 0.35.0 binaries.&lt;br /&gt;
load ollama 0.35.0 binaries.&lt;br /&gt;
------------------------------------------------------------------------------&lt;br /&gt;
Use this on your host to setup a tunnel to the running instance:&lt;br /&gt;
ssh -L 11434:gpua201.internal.anunna.wur.nl:11435 haars001@login.anunna.wur.nl&lt;br /&gt;
Then connect to http://localhost:11434&lt;br /&gt;
For applications on other hosts within Anunna, you can use:&lt;br /&gt;
OLLAMA_HOST=gpua201.internal.anunna.wur.nl:11435&lt;br /&gt;
(Stop instance by pressing CTRL-C twice)&lt;br /&gt;
------------------------------------------------------------------------------&lt;br /&gt;
time=2026-10-02T10:32:43.672+02:00 level=INFO source=routes.go:2099 msg=&amp;quot;server config&amp;quot; env=&amp;quot;map[CUDA_VISIBLE_DEVICES:0 GGML_VK_VISIBLE_DEVICES: GPU_DEVICE_ORDINAL:0 HIP_VISIBLE_DEVICES: HSA_OVERRIDE_GFX_VERSION: HTTPS_PROXY: HTTP_PROXY: LLAMA_ARG_FIT: LLAMA_ARG_FIT_TARGET: NO_PROXY: OLLAMA_CONTEXT_LENGTH:0 OLLAMA_CREATE_REMOTE:false OLLAMA_DEBUG:INFO OLLAMA_DEBUG_LOG_REQUESTS:false OLLAMA_EDITOR: OLLAMA_FLASH_ATTENTION:false OLLAMA_GO_TEMPLATE:true OLLAMA_GPU_OVERHEAD:0 OLLAMA_HOST:http://0.0.0.0:11435 OLLAMA_IGPU_ENABLE: OLLAMA_KEEP_ALIVE:5m0s OLLAMA_KV_CACHE_TYPE: OLLAMA_LLM_LIBRARY: OLLAMA_LOAD_TIMEOUT:5m0s OLLAMA_MAX_LOADED_MODELS:0 OLLAMA_MAX_QUEUE:512 OLLAMA_MAX_TRANSFER_STREAMS:4 OLLAMA_MODELS:/lustre/scratch/GUESTS/haars001/ollama-models OLLAMA_NOHISTORY:false OLLAMA_NOPRUNE:false OLLAMA_NO_CLOUD:false OLLAMA_NUM_PARALLEL:1 OLLAMA_ORIGINS:[http://localhost https://localhost http://localhost:* https://localhost:* http://127.0.0.1 https://127.0.0.1 http://127.0.0.1:* https://127.0.0.1:* http://0.0.0.0 https://0.0.0.0 http://0.0.0.0:* https://0.0.0.0:* app://* file://* tauri://* vscode-webview://* vscode-file://*] OLLAMA_REMOTES:[ollama.com] OLLAMA_SCHED_SPREAD:false OLLAMA_VULKAN:true ROCR_VISIBLE_DEVICES:0 http_proxy: https_proxy: no_proxy:]&amp;quot;&lt;br /&gt;
time=2026-10-02T10:32:43.674+02:00 level=INFO source=routes.go:2101 msg=&amp;quot;Ollama cloud disabled: false&amp;quot;&lt;br /&gt;
time=2026-10-02T10:32:44.055+02:00 level=INFO source=images.go:922 msg=&amp;quot;total blobs: 0&amp;quot;&lt;br /&gt;
time=2026-10-02T10:32:44.055+02:00 level=INFO source=images.go:930 msg=&amp;quot;total unused blobs removed: 0&amp;quot;&lt;br /&gt;
time=2026-10-02T10:32:44.063+02:00 level=INFO source=routes.go:2159 msg=&amp;quot;Listening on 0.0.0.0:11435 (version 0.35.0)&amp;quot;&lt;br /&gt;
time=2026-10-02T10:32:44.074+02:00 level=INFO source=runner.go:60 msg=&amp;quot;discovering available GPUs...&amp;quot;&lt;br /&gt;
time=2026-10-02T10:32:44.075+02:00 level=WARN source=runner.go:722 msg=&amp;quot;user overrode visible devices&amp;quot; CUDA_VISIBLE_DEVICES=0&lt;br /&gt;
time=2026-10-02T10:32:44.075+02:00 level=WARN source=runner.go:722 msg=&amp;quot;user overrode visible devices&amp;quot; ROCR_VISIBLE_DEVICES=0&lt;br /&gt;
time=2026-10-02T10:32:44.075+02:00 level=WARN source=runner.go:722 msg=&amp;quot;user overrode visible devices&amp;quot; GPU_DEVICE_ORDINAL=0&lt;br /&gt;
time=2026-10-02T10:32:44.075+02:00 level=WARN source=runner.go:726 msg=&amp;quot;if GPUs are not correctly discovered, unset and try again&amp;quot;&lt;br /&gt;
time=2026-10-02T10:32:44.325+02:00 level=INFO source=model_recommendations.go:177 msg=&amp;quot;model recommendations cache sleep scheduled&amp;quot; wait=3h44m59.585365476s consecutive_failures=0&lt;br /&gt;
time=2026-10-02T10:32:49.707+02:00 level=INFO source=types.go:32 msg=&amp;quot;inference compute&amp;quot; id=0 filter_id=0 library=ROCm compute=gfx90a name=ROCm0 description=&amp;quot;AMD Instinct MI210&amp;quot; libdirs=ollama,rocm_v7_2 driver=0.0 pci_id=0000:67:00.0 type=discrete total=&amp;quot;64.0 GiB&amp;quot; available=&amp;quot;63.9 GiB&amp;quot;&lt;br /&gt;
time=2026-10-02T10:32:49.707+02:00 level=INFO source=routes.go:2209 msg=&amp;quot;vram-based default context&amp;quot; total_vram=&amp;quot;64.0 GiB&amp;quot; default_num_ctx=262144&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;As you can see, it will tell you how to connect to it, either from your computer, or within the cluster.&lt;br /&gt;
&lt;br /&gt;
=== Batch wise usage ===&lt;br /&gt;
If you don&amp;#039;t want or need to have &amp;lt;code&amp;gt;ollama&amp;lt;/code&amp;gt; in an interactive session, you can bypass the menu, and just use the &amp;lt;code&amp;gt;start_ollama.sh&amp;lt;/code&amp;gt; script directly:&amp;lt;syntaxhighlight lang=&amp;quot;bash&amp;quot;&amp;gt;&lt;br /&gt;
sbatch --ntasks=1 --partition=gpu_amd --gres=gpu:1 --time=1:0:0 --mem-per-cpu=8G --cpus-per-gpu=8 start_ollama.sh&lt;br /&gt;
&amp;lt;/syntaxhighlight&amp;gt;Here we chose an AMD GPU, but one can of course also select Nvidia GPUs, that partition is called &amp;lt;code&amp;gt;gpu&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
After that you will have to look up the information on how to connect in the job output.&lt;/div&gt;</summary>
		<author><name>Haars001</name></author>
	</entry>
</feed>