Jump to content

Ollama

From HPCwiki
Revision as of 08:47, 2 October 2026 by Haars001 (talk | contribs) (Created page with "== How to use ollama on the HPC == On the HPC we have <code>ollama</code> available through modules, in the GPU bucket. To load the module, do this:<syntaxhighlight lang="bash"> ml GPU ollama </syntaxhighlight>Once the module is loaded, <code>ollama</code> is available on your path:<syntaxhighlight lang="bash"> ollama --version Warning: could not connect to a running Ollama instance Warning: client version is 0.35.0 </syntaxhighlight> === Intera...")
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)

How to use ollama on the HPC

On the HPC we have ollama available through modules, in the GPU bucket.

To load the module, do this:

ml GPU ollama

Once the module is loaded, ollama is available on your path:

ollama --version
Warning: could not connect to a running Ollama instance
Warning: client version is 0.35.0

Interactive session

To make working with ollama easier, we have created a small helper script, this will help you select a GPU, and set it up with the values you entered.

menu_ollama.sh

Once you have entered all the values, the script will start an interactive job (so with output and input in your shell), which will look something like this:

Now running:
srun --immediate=60 --nodes=1 --ntasks=1 --partition=gpu_amd --gres=gpu:1 --time=1:0:0 --mem-per-cpu=8G --cpus-per-gpu=8 /shared/apps/ollama/0.35.0/start_ollama.sh

srun: job 43006080 queued and waiting for resources
srun: job 43006080 has been allocated resources
remove ollama 0.35.0 binaries.
load ollama 0.35.0 binaries.
------------------------------------------------------------------------------
Use this on your host to setup a tunnel to the running instance:
ssh -L 11434:gpua201.internal.anunna.wur.nl:11435 haars001@login.anunna.wur.nl
Then connect to http://localhost:11434
For applications on other hosts within Anunna, you can use:
OLLAMA_HOST=gpua201.internal.anunna.wur.nl:11435
(Stop instance by pressing CTRL-C twice)
------------------------------------------------------------------------------
time=2026-10-02T10:32:43.672+02:00 level=INFO source=routes.go:2099 msg="server config" env="map[CUDA_VISIBLE_DEVICES:0 GGML_VK_VISIBLE_DEVICES: GPU_DEVICE_ORDINAL:0 HIP_VISIBLE_DEVICES: HSA_OVERRIDE_GFX_VERSION: HTTPS_PROXY: HTTP_PROXY: LLAMA_ARG_FIT: LLAMA_ARG_FIT_TARGET: NO_PROXY: OLLAMA_CONTEXT_LENGTH:0 OLLAMA_CREATE_REMOTE:false OLLAMA_DEBUG:INFO OLLAMA_DEBUG_LOG_REQUESTS:false OLLAMA_EDITOR: OLLAMA_FLASH_ATTENTION:false OLLAMA_GO_TEMPLATE:true OLLAMA_GPU_OVERHEAD:0 OLLAMA_HOST:http://0.0.0.0:11435 OLLAMA_IGPU_ENABLE: OLLAMA_KEEP_ALIVE:5m0s OLLAMA_KV_CACHE_TYPE: OLLAMA_LLM_LIBRARY: OLLAMA_LOAD_TIMEOUT:5m0s OLLAMA_MAX_LOADED_MODELS:0 OLLAMA_MAX_QUEUE:512 OLLAMA_MAX_TRANSFER_STREAMS:4 OLLAMA_MODELS:/lustre/scratch/GUESTS/haars001/ollama-models OLLAMA_NOHISTORY:false OLLAMA_NOPRUNE:false OLLAMA_NO_CLOUD:false OLLAMA_NUM_PARALLEL:1 OLLAMA_ORIGINS:[http://localhost https://localhost http://localhost:* https://localhost:* http://127.0.0.1 https://127.0.0.1 http://127.0.0.1:* https://127.0.0.1:* http://0.0.0.0 https://0.0.0.0 http://0.0.0.0:* https://0.0.0.0:* app://* file://* tauri://* vscode-webview://* vscode-file://*] OLLAMA_REMOTES:[ollama.com] OLLAMA_SCHED_SPREAD:false OLLAMA_VULKAN:true ROCR_VISIBLE_DEVICES:0 http_proxy: https_proxy: no_proxy:]"
time=2026-10-02T10:32:43.674+02:00 level=INFO source=routes.go:2101 msg="Ollama cloud disabled: false"
time=2026-10-02T10:32:44.055+02:00 level=INFO source=images.go:922 msg="total blobs: 0"
time=2026-10-02T10:32:44.055+02:00 level=INFO source=images.go:930 msg="total unused blobs removed: 0"
time=2026-10-02T10:32:44.063+02:00 level=INFO source=routes.go:2159 msg="Listening on 0.0.0.0:11435 (version 0.35.0)"
time=2026-10-02T10:32:44.074+02:00 level=INFO source=runner.go:60 msg="discovering available GPUs..."
time=2026-10-02T10:32:44.075+02:00 level=WARN source=runner.go:722 msg="user overrode visible devices" CUDA_VISIBLE_DEVICES=0
time=2026-10-02T10:32:44.075+02:00 level=WARN source=runner.go:722 msg="user overrode visible devices" ROCR_VISIBLE_DEVICES=0
time=2026-10-02T10:32:44.075+02:00 level=WARN source=runner.go:722 msg="user overrode visible devices" GPU_DEVICE_ORDINAL=0
time=2026-10-02T10:32:44.075+02:00 level=WARN source=runner.go:726 msg="if GPUs are not correctly discovered, unset and try again"
time=2026-10-02T10:32:44.325+02:00 level=INFO source=model_recommendations.go:177 msg="model recommendations cache sleep scheduled" wait=3h44m59.585365476s consecutive_failures=0
time=2026-10-02T10:32:49.707+02:00 level=INFO source=types.go:32 msg="inference compute" id=0 filter_id=0 library=ROCm compute=gfx90a name=ROCm0 description="AMD Instinct MI210" libdirs=ollama,rocm_v7_2 driver=0.0 pci_id=0000:67:00.0 type=discrete total="64.0 GiB" available="63.9 GiB"
time=2026-10-02T10:32:49.707+02:00 level=INFO source=routes.go:2209 msg="vram-based default context" total_vram="64.0 GiB" default_num_ctx=262144

As you can see, it will tell you how to connect to it, either from your computer, or within the cluster.

Batch wise usage

If you don't want or need to have ollama in an interactive session, you can bypass the menu, and just use the start_ollama.sh script directly:

sbatch --ntasks=1 --partition=gpu_amd --gres=gpu:1 --time=1:0:0 --mem-per-cpu=8G --cpus-per-gpu=8 start_ollama.sh

Here we chose an AMD GPU, but one can of course also select Nvidia GPUs, that partition is called gpu.

After that you will have to look up the information on how to connect in the job output.