Writing
Blog
Current writing lives on TokenCost; earlier machine learning articles live on Medium. Everything below links out to the original.
Build log
Fine-Tuning Qwen3-4B on Agent Traces: the Thinking-Budget Failure Benchmarks Don't Catch
A QLoRA fine-tune of Qwen3-4B on real agent traces: completion masking, replay mix, HumanEval+ results against the base model, and a 21% silent-failure mode in hybrid-thinking models that standard benchmarks miss.
2026-07-22