Notes from production ML work: computer vision on the factory floor, inference optimization, and fine-tuning open LLMs.
LLM Fine-tuning
QLoRA settings, VRAM, GGUF export and a working Ollama Modelfile for a fine-tuned gpt-oss-20b, including the Harmony template and stop token it needs.
12 Oct 2026 · 5 min read