Fine-Tuning Is Not the Answer to Cheaper Tokens

As frontier API costs drop 10x annually, teams keep reaching for fine-tuning to cut spend. That's the wrong reason. Here's what fine-tuning is actually for.

Cover art for Fine-Tuning Is Not the Answer to Cheaper Tokens

A May 2025 paper comparing fine-tuned small models to prompted frontier ones found that fine-tuning improved quality by about 10% on a structured-output task - while a "good prompt" got most of the way there on its own. That 10% gap is real. But it's not a cost argument anymore. It's a behavior argument. Teams that miss that distinction are about to fine-tune themselves into a maintenance hole they didn't need to dig.

The debate - fine-tune a small model or prompt a big one? - is getting murkier as the economics shift beneath it. AI infrastructure costs have dropped 280-fold, and inference costs decline roughly 10x annually. Frontier model API calls that cost a team real money in 2023 are rounding errors by 2026. Which means the dominant reason most teams give for fine-tuning - it's cheaper - is eroding faster than their training pipelines can keep up.

The position here is specific: fine-tuning for cost savings is a weak bet in most cases today. Fine-tuning for output behavior is still worth doing, when you know what you're asking it to fix.

When "just prompt it better" actually wins

Most teams asking about fine-tuning should not fine-tune. They should fix their prompts, build a real RAG pipeline, and write evals - in that order.

That sounds harsh, but the evidence backs it. Even without any fine-tuning, a well-designed prompt can enable an LLM to perform well on a narrow task. The failure mode is teams reaching for fine-tuning before they've established a proper evaluation set - which means they can't measure whether the tuned model is actually better, or just different.

If you're under 100 requests a day across varied tasks, just call an API - you'll never recoup the time spent training and maintaining your own model. If you're somewhere in the middle, start with prompting plus RAG, and only reach for fine-tuning once your evaluation set stops improving.

The practical sequence matters. The right sequence is: prompt → RAG → fine-tune → distill. Skipping steps because fine-tuning feels more rigorous is a common and expensive mistake. A fine-tuned model needs version control, regression evals, retraining schedules, and documentation. A well-crafted system prompt costs none of that.

What fine-tuning is actually for

There are four places where fine-tuning genuinely moves the needle: structured output reliability when prompt-only solutions still hallucinate fields, domain vocabulary or jargon that base models hedge on, refusal and tone control where prompt instructions get overridden, and cost compression through small-model distillation from a working large-model pipeline.

Notice the shape of that list. Three of the four are about behavior, not knowledge. The fourth - distillation - is really a cost play at scale, not a quality play.

Notice what's missing: knowledge injection. Research has consistently shown that RAG outperforms fine-tuning for factual recall. Baking facts into weights produces stale, unverifiable answers and can erode the model's general capability through catastrophic forgetting. If your problem is "the model doesn't know our docs," fine-tuning is the wrong tool.

This is where most teams go wrong. They encounter a model that gives a wrong answer about their product, and they assume the fix is to train the right answer into the weights. It isn't. The answer lives in a document somewhere - retrieve it.

Problem Right tool Fine-tuning helps?
Model doesn't know our internal docs RAG No - facts go stale
Output format keeps breaking JSON schema Fine-tune Yes - reliably
Model hedges on our industry jargon Fine-tune (SFT or LoRA) Yes - tone and vocab
We're paying too much on 50 req/day Better prompt No - overhead exceeds savings
We're paying too much on 50,000 req/day Distillation Yes - at that volume
Model ignores our refusal instructions Fine-tune Yes - prompt gets overridden
10%quality gain from fine-tuningover a strong prompt, on structured-output tasks
10xannual decline in inference costsmaking pure cost-savings arguments weaker each year
~1/10inference cost of a distilled small modelvs. a frontier API call, at near-parity quality on narrow tasks

The one cost argument that still holds: distillation at volume

Here's the steelman for the cost case, and it's a real one.

Use a frontier model to generate high-quality outputs on your specific task, then fine-tune a small open-source model on those outputs. The tuned smaller model delivers near-frontier quality on the narrow task at roughly one-tenth the inference cost and a fraction of the latency.

This is distillation, and it's different from fine-tuning on your own task examples directly. You're using the frontier model as a labeler, not as a production inference path. The small model inherits its output quality without inheriting its per-token price.

If you're running more than 10 requests per second on a single narrow task, fine-tuning a 3-to-8B model and self-hosting it is defensible - the volume justifies the upfront effort and the cost savings compound.

Below that threshold, the math rarely closes. A fine-tune run, the infrastructure to serve it, the eval suite to validate it, and the retraining cadence to keep it current represent real ongoing cost that most teams undercount. A Towards Data Science analysis put the break-even threshold bluntly: under roughly 100 requests per day on varied tasks, the overhead never pays off.

Beagle in action#ml-platform, 11:02am
The ask
engineer asks whether they should fine-tune Llama-3-8B to reduce support ticket classification costs
Beagle drafts
pulls current ticket volume from Linear (340/week), runs the inference cost estimate, drafts a reply: 'At 340 tickets/week, the API cost is ~$28/month - fine-tune overhead likely exceeds that for 12+ months. Recommend: write evals first, then decide.'
You approve
the team skips the fine-tune, builds evals instead, finds three prompt edge cases that were causing misclassification
Do this in your workspace →

Reinforcement fine-tuning: the narrow case that's actually new

There is one genuinely new development worth tracking. OpenAI's Reinforcement Fine-Tuning (RFT) is now generally available on o-series reasoning models, currently scoped to o4-mini. RFT trains a model against a custom grader rather than labeled outputs.

This is different in kind from standard supervised fine-tuning. You're not teaching the model to mimic your examples - you're giving it a reward signal and letting it discover the right reasoning path. For tasks where correctness is verifiable (code that runs, answers that match a schema, citations that resolve), this is a sharper tool than SFT. The tradeoff: you need a reliable grader, which is its own engineering project.

GRPO has largely replaced pure supervised fine-tuning as the technique of interest. In 2023, SFT dominated - you gave the model input-output pairs, it learned to mimic your data. That still works. But the frontier has moved to GRPO and reinforcement learning from human feedback.

For teams not running verifiable tasks, this doesn't change the calculus. For teams with a clear correctness signal - a compiler, a schema validator, a unit test suite - it's worth a look before assuming you need a massive labeled dataset.

Cutting classification errors in a support workflow
Without Beagle
team decides to fine-tune, spends three weeks collecting training examples, discovers mid-run that they have no eval set to validate against
With Beagle
team writes 60 labeled eval examples first, finds that two prompt changes fix 80% of errors, reserves fine-tuning for the residual structured-output failures that prompting can't close

Fine-tuning vs prompting: common questions

When does fine-tuning beat prompting for narrow tasks?

Fine-tuning reliably wins on four problems: structured output reliability when prompting still produces malformed JSON, domain vocabulary where the base model hedges or paraphrases, refusal and tone behavior that prompt instructions cannot reliably enforce, and latency-sensitive tasks where a smaller model's faster inference matters more than marginal quality differences.

Is fine-tuning still worth it as API costs drop?

For pure cost reduction at modest volumes, less so each year. Inference costs are dropping roughly 10x annually, which shortens the payback period for any fine-tune done primarily to save money. The exception is distillation at high throughput: if you're running a narrow task at more than 10 requests per second, a distilled small model still saves meaningfully at that scale.

What's the right sequence before reaching for fine-tuning?

Fix your prompt first, then add retrieval for factual grounding, then write an evaluation set that covers your failure modes. Fine-tune only once the eval set stops improving with prompt changes. Most teams who follow this sequence discover they don't need to fine-tune - or they fine-tune with much cleaner data and clearer goals.

What is distillation and how is it different from fine-tuning?

Distillation uses a large frontier model to generate labeled outputs for your task, then trains a smaller model on those outputs. The small model inherits near-frontier quality on the specific task at a fraction of the inference cost. It's a fine-tuning technique, but the key difference is that your training data comes from a stronger model, not from hand-labeled examples.

What is reinforcement fine-tuning and when should I use it?

Reinforcement fine-tuning (OpenAI calls theirs RFT) trains a model against a reward signal - a grader that scores outputs - rather than mimicking labeled examples. It's worth considering when your task has a clear, automatable correctness signal: code that compiles, answers that match a schema, or claims that can be verified against a database. If you can't define a reliable grader, stick with supervised fine-tuning or prompting.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle