LLM Fine-tuning

Fine-tuning gpt-oss-20b with Unsloth and running it in Ollama

By Khadim Hussain · · 5 min read

In short

gpt-oss-20b fine-tunes with QLoRA on a single consumer GPU (Unsloth puts it at about 14 GB of VRAM). Shipping it is where the chat template matters: gpt-oss uses OpenAI's Harmony format, so the GGUF needs a Harmony template with <|return|> as a stop token, or Ollama gives gibberish or never stops.

On this page

I fine-tuned gpt-oss-20b for STEM reasoning and published it three ways: a LoRA adapter, a merged 16-bit model and GGUF files. Below are the settings and versions behind it, the export code, and the Modelfile that makes the GGUF behave in Ollama.

QLoRA settings for gpt-oss-20b

The published run, as recorded on the model card:

SettingValue
Base modelopenai/gpt-oss-20b (21B parameters, mixture of experts)
MethodQLoRA, 4-bit base weights
LoRA rank / alpha32 / 64
Learning rate1e-4
Batch size1, with gradient accumulation 16
OptimizerAdamW 8-bit
Epochs1
Data4,260 training and 474 evaluation examples of STEM Q&A with chain-of-thought
Resulttrain loss 1.087, eval loss 0.837
FrameworkUnsloth + TRL

The full training config, with target modules and scheduler, is in configs/gpt_oss_20b.yaml in kllm, the small toolkit I use for these runs.

How much VRAM does fine-tuning gpt-oss-20b need?

QLoRA keeps the base weights in 4-bit and trains only a small adapter, so a 20B model fits on one consumer card. Unsloth's gpt-oss guide lists about 14 GB of VRAM for gpt-oss-20b with QLoRA and recommends at least 16 GB for stable runs. LoRA on unquantized BF16 weights needs about 44 GB, which is why QLoRA is the practical choice on one card.

The environment kllm is verified on:

ComponentVersion
GPURTX 5090 (32 GB)
OSUbuntu 24.04
CUDA12.8 toolkit (13.0 driver)
Python3.12.3
PyTorch2.9.1+cu128
Unsloth2026.1.4
transformers4.57.3
vLLM0.15.0
uv pip install -U vllm --torch-backend=cu128   # cu121 for RTX 40 series
uv pip install unsloth unsloth_zoo bitsandbytes
uv pip install --force-reinstall "transformers==4.57.3"

Export to LoRA, merged and GGUF with Unsloth

After training, the same model object gives all three formats. These are the calls kllm makes:

# 1. LoRA adapter (~61 MB): for people who already run the base model
model.save_pretrained(output_dir)
tokenizer.save_pretrained(output_dir)
 
# 2. Merged 16-bit model (~41 GB): for Transformers or vLLM, no adapter handling
model.save_pretrained_merged(f"{output_dir}-merged", tokenizer, save_method="merged_16bit")
 
# 3. GGUF for llama.cpp, Ollama and LM Studio, one call per quantization
model.save_pretrained_gguf("gguf/gpt-oss-20b-stem", tokenizer, quantization_method="q4_k_m")

quantization_method takes f16, q8_0, q4_k_m and the other llama.cpp types listed in Unsloth's GGUF docs. The GGUF files for this model came out at:

FileSize
f1641.9 GB
Q8_022.3 GB
Q4_K_M15.8 GB

If the GGUF step runs out of memory, Unsloth's docs suggest lowering maximum_memory_usage from its 0.75 default to around 0.5.

The Harmony chat template is what breaks

gpt-oss doesn't use ChatML. It was trained on OpenAI's Harmony format, where every message is wrapped in special tokens and the assistant writes on separate channels:

  • analysis for its chain of thought, which end users shouldn't see
  • final for the answer
  • commentary for tool calls

A final answer ends with <|return|>, which OpenAI's docs call "a valid stop token indicating that you should stop inference." Messages kept in the conversation history end with <|end|> instead, while supervised training targets should end with <|return|>. OpenAI's model card is blunt about it: the models "should only be used with the harmony format as it will not work correctly otherwise."

So the template you ship has to match the one used in training. Unsloth's docs name a mismatch as the most common reason a model gives gibberish or endless output on other platforms. And <|return|> has to be a stop token, or generation runs past the answer.

This exact question, how to template a fine-tuned gpt-oss for Ollama, is still open on OpenAI's forum.

A working Ollama Modelfile for fine-tuned gpt-oss

This is the Modelfile that ships with the GGUF repo:

FROM ./gpt-oss-20b-finetuned-q8_0.gguf
 
TEMPLATE """<|start|>system<|message|>You are a helpful assistant trained by OpenAI.
 
Reasoning: medium
 
# Valid channels: analysis, commentary, final. Channel must be included for every message.<|end|><|start|>user<|message|>{{ .Prompt }}<|end|><|start|>assistant"""
 
PARAMETER stop "<|return|>"
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_ctx 4096
PARAMETER num_predict 2048

Setting reasoning effort

Reasoning: medium in the system message sets how much the model thinks in its analysis channel before it answers. Harmony accepts low, medium and high, and medium is the default. To change it, edit that line in the TEMPLATE and run ollama create again. Ollama shows the analysis channel as "Thinking...", so low is the setting for short, fast answers.

ollama run hf.co or a Modelfile?

With the Modelfile, the template is the one you wrote:

ollama create gpt-oss-20b-stem -f Modelfile
ollama run gpt-oss-20b-stem

Ollama can also pull straight from Hugging Face:

ollama run hf.co/khadim-hussain/gpt-oss-20b-stem-reasoning-GGUF:Q4_K_M

That route doesn't read the Modelfile in the repo. Per Hugging Face's Ollama docs, it uses the chat template stored in the GGUF's tokenizer.chat_template metadata, or a file named template in the repo. The quantization tag is case-insensitive, and without a tag Ollama picks Q4_K_M. If you publish a Harmony model, put a template file in the repo or point people to the Modelfile.

Which format people actually download

All-time Hugging Face downloads as of October 2026:

Formatgpt-oss-20bQwen3-14B
GGUF513274
Merged 16-bit7042
LoRA adapter4651

Most people want a file they can run on their own machine. If you only publish one format, make it a GGUF with a working template. The adapter is worth adding for anyone who wants to keep training.

Troubleshooting

The failures documented in kllm's troubleshooting guide:

SymptomFix
Training process dies with no errorOut of memory. Check dmesg | grep oom and lower the batch size
UnboundLocalError mentioning dropoutSet lora_dropout: 0.0
ImportError: cannot import name ... from 'transformers'uv pip install --force-reinstall "transformers==4.57.3"
Python.h: No such file or directorysudo apt-get install python3.12-dev

The full project, with the Qwen3-14B run and the download numbers, is in the STEM reasoning LLMs case study.

Frequently asked questions

How much VRAM do you need to fine-tune gpt-oss-20b?

With QLoRA, Unsloth's gpt-oss guide puts gpt-oss-20b at about 14 GB of VRAM and recommends at least 16 GB for stable runs. LoRA on BF16 weights needs about 44 GB. kllm, the toolkit I trained with, is verified on an RTX 5090 with 32 GB.

Can I run a fine-tuned gpt-oss-20b in Ollama?

Yes. Export it to GGUF, then either run ollama create with a Modelfile that carries the Harmony template, or pull it with ollama run hf.co/<user>/<repo>:Q4_K_M. The second route uses the chat template embedded in the GGUF file, not the Modelfile in the repo.

Why does my fine-tuned gpt-oss model output gibberish or never stop?

Usually the chat template. The template at inference has to match the one used in training, and Harmony ends a final answer with <|return|>. If that token isn't a stop token, generation runs past the answer.

Which GGUF quantization should I use for gpt-oss-20b?

Q4_K_M (15.8 GB for this model) is the smallest file and the one Ollama picks by default. Use Q8_0 (22.3 GB) if you have the memory and want output closer to the 16-bit model.

How do I set reasoning effort for gpt-oss in Ollama?

Harmony reads it from the system message as Reasoning: low, medium or high, and medium is the default. In a Modelfile, change that line inside the TEMPLATE and run ollama create again.

Want a second pair of eyes on this?

Book a 15-minute intro call. Bring the problem, and you leave with a concrete next step.