Distillation Fine-Tuning for One Narrow Task

Bridgewater tuned Qwen3-235B on private financial data and beat GPT on six tasks at 13.8× lower inference cost. Here's what distillation fine-tuning actually involves - and where the real cost hides.

Cover art for Distillation Fine-Tuning for One Narrow Task

Bridgewater Associates trained on Thinking Machines Lab's Tinker platform, built on top of the open model Qwen3-235B. In their own evaluation, the fine-tuned model hit 84.7% accuracy versus 78.2% for the best frontier model tested - and cost nearly 14 times less to run. That is the pitch for distillation fine-tuning in one sentence. But the gap between what that result sounds like and what it actually required is wide enough to matter.

This post covers what distillation fine-tuning of open-weight models genuinely is, where it beats prompting and RAG, where it does not, and what the Bridgewater case reveals about the real cost structure that most write-ups gloss over.

What distillation fine-tuning actually means for teams using closed APIs

Distillation fine-tuning means training a smaller student model to approximate the behavior of a larger teacher on a specific task. In academic settings the teacher hands the student its full probability distribution. In practice, if your teacher is GPT or Claude behind an API, the textbook version falls apart on contact with reality: classic distillation needs the teacher's probability distribution, and if your teacher lives behind somebody's API, you are not getting a probability distribution - you are getting text.

The version that actually ships: collect real inputs from your product, run them through the expensive teacher once, keep every output, and fine-tune a small open model on those input/output pairs. That is supervised fine-tuning with AI-generated labels. Calling it distillation is not wrong, but it is a looser sense of the term than the literature uses. The practical consequence: your student inherits the teacher's outputs but not its reasoning process, so tasks that require multi-step chain-of-thought benefit less than tasks that are essentially classification or extraction at the output level.

Distillation fine-tuning is most effective when the teacher model is substantially more capable than the student on the target task - recommended specifically for transferring complex multi-step reasoning capabilities, including scientific and domain-specific question answering that requires step-by-step reasoning.

It provides smaller gains on tasks where the student model already performs close to the teacher, or on short-form retrieval tasks where the teacher's reasoning trace does not add value.

When it beats prompting, and when it does not

Prompt engineering with in-context examples handles most narrow tasks if you can fit 5-10 examples in the prompt. Frontier models in 2026 have 128K-1M context windows, so you can include a small "training set" inline and skip the fine-tune entirely. That is an important first filter. Distillation fine-tuning earns its cost in a narrower band of situations than its advocates usually admit.

Situation Prompt + RAG Distillation fine-tune
Task is new, volume is low ✓ Fast and cheap ✗ Overkill
Answers live in changing external docs ✓ Retrieval handles updates ✗ Weights go stale
Same narrow task, >100K calls/month Expensive at scale ✓ Inference cost collapses
Right answers are in private data, not public web ✗ Model never saw them ✓ Core use case
Output format must be exactly consistent Drifts under prompting pressure ✓ Enforced in weights
Style and brand voice Fragile in long prompts ✓
500-2,000 example QLoRA reliably teaches house tone, faster than a 10K-token system prompt

Fine-tuning wins when the problem is behavior: you need a reliable output format every time, a specific brand voice, adherence to a niche taxonomy, or a narrow task (classification, extraction, structured generation) repeated at high volume where you want a smaller, cheaper model to punch above its weight.

RAG wins when the problem is knowledge: the model needs your current policies, product catalog, contracts, or support history. Facts change, so you want them retrieved at query time - not baked into weights that go stale the day you train.

Huge pools of proprietary corporate data and untrained human expertise still exist and hold real room for improvement - especially where companies deliberately keep their most valuable data private. Anyone who hands that data to a frontier lab risks competing against a product built on top of it.

What the Bridgewater case actually shows - and what it does not

Bridgewater Associates and Thinking Machines Lab say the tuned Qwen3-235B model outperformed GPT, Claude, and Gemini variants in an internal finance-task evaluation, where expert labels, prompt rules, and fine-tuning helped encode private workflow judgments that public web knowledge lacked.

The reported figures - 84.7% accuracy and a 13.8× inference-cost reduction - are company-run measurements, not independent benchmarks. That caveat matters. Both companies have a product to sell.

What the case does confirm, independent of the exact numbers, is a structural point: the real work in finance is a constant stream of small, repeated judgment calls about what actually matters. The researchers defined six tasks drawn from an investor's daily routine - one example being whether a financial article is relevant to a specific executive. Highly repeated, deeply private, low tolerance for inconsistency. That profile is where fine-tuning has always been strongest.

On the training side, the team did not just QLoRA and ship. For their multi-task training recipe, they compared three batching strategies and found that interleaving one batch per task in round-robin order improved accuracy by 12.1% over fully mixed batches.

They also used on-policy distillation, where the student model was regularized to stay near a teacher distribution that was updated only when validation accuracy hit new highs. That is meaningfully more sophisticated than the standard "run teacher outputs through student" recipe most teams start with.

84.7%accuracy on six finance taskstuned Qwen3-235B, Bridgewater internal eval
13.8×inference cost reductionvs. the frontier model it replaced
$10-16QLoRA run on a single H100for a 7B-class model, 8-12 hours
28×data annotation cost over computefor frontier-model-level quality data

The cost structure nobody leads with

The compute bill for distillation fine-tuning has collapsed. A QLoRA run on a single H100 takes 8-12 hours and costs $10-16, producing a LoRA adapter file of 50-200 MB that you merge with the base model.

Fine-tuning is dramatically cheaper than pre-training; fine-tuning a 7B model costs approximately $50-$500 on a GPU marketplace.

What is not cheap is data. Human expertise and data annotation costs now exceed compute expenses by up to 28× for frontier models, making data quality and team expertise the primary cost drivers alongside infrastructure efficiency. For most teams, the hours spent collecting real inputs, curating edge cases, and evaluating whether the tuned model actually got better - not the GPU bill - are where the budget goes.

No data, no useful fine-tune, and no way to measure success. Without evaluation, you cannot tell whether the trained model is better - or quietly worse. That last part is underemphasized. A fine-tuned model that regresses on cases outside your curated slice will fail silently in production. The eval harness is not optional scaffolding; it is the thing that prevents a narrow win from masking a broad regression.

Triaging 200 analyst queries a day
Without Beagle
every query hits GPT at frontier pricing; inconsistent format means downstream parsing breaks on 15% of responses
With Beagle
a Beagle-triggered workflow routes the query to a tuned open model; the adapter enforces schema, inference cost drops, and a human approves before anything sensitive is logged

A practical note on tooling: Unsloth handles single-GPU QLoRA on a free tier for 7B models; a rented H100 handles 70B in under three hours.

The Bridgewater team trained on Tinker from Thinking Machines Lab, which let them iterate quickly without worrying about GPU infrastructure. For teams that want control without managing cluster orchestration, Tinker and Together AI's fine-tuning API are the two platforms most cited in primary sources right now.

Beagle in action#data-team, 2:47pm
The ask
'can we run the doc-relevance classifier on today's filings before EOD?'
Beagle drafts
checks the linked workflow, sees the tuned model endpoint is configured, drafts a reply confirming it can queue the batch and post results to the shared sheet
You approve
you approve; the job kicks off without anyone touching a notebook
Do this in your workspace →

Distillation fine-tuning: common questions

What is the difference between distillation and standard fine-tuning?

Standard fine-tuning updates a model on labeled examples from humans. Distillation uses a larger, stronger model to generate those labels. In practice, when the teacher is a closed API, distillation is supervised fine-tuning on AI-generated outputs - you gain labeled data at scale but lose access to the teacher's internal probability distribution, which limits certain theoretical benefits.

When does fine-tuning an open-weight model beat prompting a frontier model?

Fine-tuning wins when the task is narrow, repeated at high volume, demands a consistent output format, or depends on private data the frontier model was never trained on. The strongest commercial case in 2026 is distillation: tune a small open-source model on frontier model outputs to match near-frontier quality on a narrow task at roughly one-tenth the inference cost.

How much does distillation fine-tuning actually cost?

Compute is now the smaller line item. A QLoRA run on a single H100 takes 8-12 hours at $10-16, producing a LoRA adapter file of 50-200 MB. The real cost is data curation and evaluation - collecting clean examples, reviewing edge cases, and building the eval suite that tells you whether the tuned model is genuinely better on held-out data.

Does fine-tuning an open model pose a data security risk?

The reverse: it reduces one. Fine-tuning open-weights models lets you host the result yourself after training. That lets regulated teams keep data and models in an environment their rules require, which closed fine-tuning services do not allow. The training run itself still touches a cloud GPU, so where that GPU sits matters for compliance.

What open-weight models are most commonly fine-tuned right now?

QLoRA with rank 16 on top of Llama 3, Qwen 3, Gemma 4, or Mistral is the current default starting point.

NVIDIA's Nemotron 3 (June 2026) comes in Nano (30B), Super (120B), and Ultra (~550B) tiers, and is often fine-tuned specifically for tool-use agents, long-horizon planning, and throughput-sensitive production pipelines.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle