LLM Fine-tuning

STEM reasoning LLMs: fine-tuning gpt-oss-20b and Qwen3-14B

By Khadim Hussain · · Updated · 2 min read

In short

QLoRA fine-tunes of gpt-oss-20b and Qwen3-14B for STEM reasoning and Q&A, trained with Unsloth. Each model is published three ways: a LoRA adapter, a merged bf16 model, and GGUF quantizations that run locally in Ollama, llama.cpp and LM Studio.

4,260
training examples
780+
GGUF downloads
3
formats
  • Unsloth
  • QLoRA
  • PEFT
  • GGUF
  • Ollama
  • llama.cpp
  • Hugging Face
  • Transformers
STEM reasoning model training results
On this page

What was trained

Two open-weight models, fine-tuned with QLoRA through Unsloth:

Base modelTaskPublished as
openai/gpt-oss-20bSTEM reasoningLoRA adapter, merged bf16 model, GGUF (f16, Q8_0, Q4_K_M)
Qwen3-14BSTEM Q&ALoRA adapter, merged bf16 model, GGUF (f16, Q4_K_M)

QLoRA trains a small low-rank adapter on top of a 4-bit quantized copy of the base model, which cuts the GPU memory a 20B fine-tune needs to a fraction of a full fine-tune.

Why three formats

Each format is for a different user:

  • LoRA adapter for people who already run the base model and want to load the fine-tune on top
  • Merged bf16 model for serving with Transformers or vLLM without handling adapters
  • GGUF for running on a laptop or a single consumer GPU

The GGUF sizes decide what hardware works:

FileSize
gpt-oss-20b Q4_K_M15.8 GB
gpt-oss-20b Q8_022.3 GB
gpt-oss-20b f1641.9 GB
Qwen3-14B Q4_K_M9.0 GB
Qwen3-14B f1629.5 GB

Running it locally

Each GGUF repo ships a Modelfile with the right chat template. gpt-oss needs OpenAI's Harmony format, so this is the reliable route:

# after downloading a .gguf and the Modelfile from the repo
ollama create gpt-oss-20b-stem -f Modelfile
ollama run gpt-oss-20b-stem

ollama run hf.co/khadim-hussain/gpt-oss-20b-stem-reasoning-GGUF:Q4_K_M also works, but it uses the template embedded in the GGUF file rather than the Modelfile. The full walkthrough, from the QLoRA settings to the Harmony template, is in Fine-tuning gpt-oss-20b with Unsloth and running it in Ollama.

Training results

ModelTrain lossEval lossExamples (train / eval)
gpt-oss-20b1.0870.8374,260 / 474
Qwen3-14B0.4610.6924,260 / 474

Both runs used LoRA rank 32, alpha 64, learning rate 1e-4, batch size 1 with gradient accumulation 16, and AdamW 8-bit, as recorded on the model cards.

What got used

All-time downloads on Hugging Face as of October 2026: 513 for the gpt-oss-20b GGUF and 274 for the Qwen3-14B GGUF. The adapters and merged models have far fewer, so most people who use a fine-tune want it in a format that runs locally.

Want a second pair of eyes on this?

Book a 15-minute intro call. Bring the problem, and you leave with a concrete next step.