What was trained
Two open-weight models, fine-tuned with QLoRA through Unsloth:
| Base model | Task | Published as |
|---|---|---|
| openai/gpt-oss-20b | STEM reasoning | LoRA adapter, merged bf16 model, GGUF (f16, Q8_0, Q4_K_M) |
| Qwen3-14B | STEM Q&A | LoRA adapter, merged bf16 model, GGUF (f16, Q4_K_M) |
QLoRA trains a small low-rank adapter on top of a 4-bit quantized copy of the base model, which cuts the GPU memory a 20B fine-tune needs to a fraction of a full fine-tune.
Why three formats
Each format is for a different user:
- LoRA adapter for people who already run the base model and want to load the fine-tune on top
- Merged bf16 model for serving with Transformers or vLLM without handling adapters
- GGUF for running on a laptop or a single consumer GPU
The GGUF sizes decide what hardware works:
| File | Size |
|---|---|
| gpt-oss-20b Q4_K_M | 15.8 GB |
| gpt-oss-20b Q8_0 | 22.3 GB |
| gpt-oss-20b f16 | 41.9 GB |
| Qwen3-14B Q4_K_M | 9.0 GB |
| Qwen3-14B f16 | 29.5 GB |
Running it locally
Each GGUF repo ships a Modelfile with the right chat template. gpt-oss needs OpenAI's Harmony format, so this is the reliable route:
# after downloading a .gguf and the Modelfile from the repo
ollama create gpt-oss-20b-stem -f Modelfile
ollama run gpt-oss-20b-stemollama run hf.co/khadim-hussain/gpt-oss-20b-stem-reasoning-GGUF:Q4_K_M also works, but it uses the template embedded in the GGUF file rather than the Modelfile. The full walkthrough, from the QLoRA settings to the Harmony template, is in Fine-tuning gpt-oss-20b with Unsloth and running it in Ollama.
Training results
| Model | Train loss | Eval loss | Examples (train / eval) |
|---|---|---|---|
| gpt-oss-20b | 1.087 | 0.837 | 4,260 / 474 |
| Qwen3-14B | 0.461 | 0.692 | 4,260 / 474 |
Both runs used LoRA rank 32, alpha 64, learning rate 1e-4, batch size 1 with gradient accumulation 16, and AdamW 8-bit, as recorded on the model cards.
What got used
All-time downloads on Hugging Face as of October 2026: 513 for the gpt-oss-20b GGUF and 274 for the Qwen3-14B GGUF. The adapters and merged models have far fewer, so most people who use a fine-tune want it in a format that runs locally.
