A May 2025 paper comparing fine-tuned small models to prompted frontier ones found that fine-tuning improved quality by about 10% on a structured-output task - while a "good prompt" got most of the way there on its own. That 10% gap is real. But it's not a cost argument anymore. It's a behavior argument. Teams that miss that distinction are about to fine-tune themselves into a maintenance hole they didn't need to dig.
The debate - fine-tune a small model or prompt a big one? - is getting murkier as the economics shift beneath it. AI infrastructure costs have dropped 280-fold, and inference costs decline roughly 10x annually. Frontier model API calls that cost a team real money in 2023 are rounding errors by 2026. Which means the dominant reason most teams give for fine-tuning - it's cheaper - is eroding faster than their training pipelines can keep up.
The position here is specific: fine-tuning for cost savings is a weak bet in most cases today. Fine-tuning for output behavior is still worth doing, when you know what you're asking it to fix.
When "just prompt it better" actually wins
Most teams asking about fine-tuning should not fine-tune. They should fix their prompts, build a real RAG pipeline, and write evals - in that order.
That sounds harsh, but the evidence backs it. Even without any fine-tuning, a well-designed prompt can enable an LLM to perform well on a narrow task. The failure mode is teams reaching for fine-tuning before they've established a proper evaluation set - which means they can't measure whether the tuned model is actually better, or just different.
If you're under 100 requests a day across varied tasks, just call an API - you'll never recoup the time spent training and maintaining your own model. If you're somewhere in the middle, start with prompting plus RAG, and only reach for fine-tuning once your evaluation set stops improving.
The practical sequence matters. The right sequence is: prompt → RAG → fine-tune → distill. Skipping steps because fine-tuning feels more rigorous is a common and expensive mistake. A fine-tuned model needs version control, regression evals, retraining schedules, and documentation. A well-crafted system prompt costs none of that.
What fine-tuning is actually for
There are four places where fine-tuning genuinely moves the needle: structured output reliability when prompt-only solutions still hallucinate fields, domain vocabulary or jargon that base models hedge on, refusal and tone control where prompt instructions get overridden, and cost compression through small-model distillation from a working large-model pipeline.
Notice the shape of that list. Three of the four are about behavior, not knowledge. The fourth - distillation - is really a cost play at scale, not a quality play.
Notice what's missing: knowledge injection. Research has consistently shown that RAG outperforms fine-tuning for factual recall. Baking facts into weights produces stale, unverifiable answers and can erode the model's general capability through catastrophic forgetting. If your problem is "the model doesn't know our docs," fine-tuning is the wrong tool.
This is where most teams go wrong. They encounter a model that gives a wrong answer about their product, and they assume the fix is to train the right answer into the weights. It isn't. The answer lives in a document somewhere - retrieve it.
| Problem | Right tool | Fine-tuning helps? |
|---|---|---|
| Model doesn't know our internal docs | RAG | No - facts go stale |
| Output format keeps breaking JSON schema | Fine-tune | Yes - reliably |
| Model hedges on our industry jargon | Fine-tune (SFT or LoRA) | Yes - tone and vocab |
| We're paying too much on 50 req/day | Better prompt | No - overhead exceeds savings |
| We're paying too much on 50,000 req/day | Distillation | Yes - at that volume |
| Model ignores our refusal instructions | Fine-tune | Yes - prompt gets overridden |
The one cost argument that still holds: distillation at volume
Here's the steelman for the cost case, and it's a real one.
Use a frontier model to generate high-quality outputs on your specific task, then fine-tune a small open-source model on those outputs. The tuned smaller model delivers near-frontier quality on the narrow task at roughly one-tenth the inference cost and a fraction of the latency.
This is distillation, and it's different from fine-tuning on your own task examples directly. You're using the frontier model as a labeler, not as a production inference path. The small model inherits its output quality without inheriting its per-token price.
If you're running more than 10 requests per second on a single narrow task, fine-tuning a 3-to-8B model and self-hosting it is defensible - the volume justifies the upfront effort and the cost savings compound.
Below that threshold, the math rarely closes. A fine-tune run, the infrastructure to serve it, the eval suite to validate it, and the retraining cadence to keep it current represent real ongoing cost that most teams undercount. A Towards Data Science analysis put the break-even threshold bluntly: under roughly 100 requests per day on varied tasks, the overhead never pays off.
Reinforcement fine-tuning: the narrow case that's actually new
There is one genuinely new development worth tracking. OpenAI's Reinforcement Fine-Tuning (RFT) is now generally available on o-series reasoning models, currently scoped to o4-mini. RFT trains a model against a custom grader rather than labeled outputs.
This is different in kind from standard supervised fine-tuning. You're not teaching the model to mimic your examples - you're giving it a reward signal and letting it discover the right reasoning path. For tasks where correctness is verifiable (code that runs, answers that match a schema, citations that resolve), this is a sharper tool than SFT. The tradeoff: you need a reliable grader, which is its own engineering project.
GRPO has largely replaced pure supervised fine-tuning as the technique of interest. In 2023, SFT dominated - you gave the model input-output pairs, it learned to mimic your data. That still works. But the frontier has moved to GRPO and reinforcement learning from human feedback.
For teams not running verifiable tasks, this doesn't change the calculus. For teams with a clear correctness signal - a compiler, a schema validator, a unit test suite - it's worth a look before assuming you need a massive labeled dataset.
Fine-tuning vs prompting: common questions
When does fine-tuning beat prompting for narrow tasks?
Fine-tuning reliably wins on four problems: structured output reliability when prompting still produces malformed JSON, domain vocabulary where the base model hedges or paraphrases, refusal and tone behavior that prompt instructions cannot reliably enforce, and latency-sensitive tasks where a smaller model's faster inference matters more than marginal quality differences.
Is fine-tuning still worth it as API costs drop?
For pure cost reduction at modest volumes, less so each year. Inference costs are dropping roughly 10x annually, which shortens the payback period for any fine-tune done primarily to save money. The exception is distillation at high throughput: if you're running a narrow task at more than 10 requests per second, a distilled small model still saves meaningfully at that scale.
What's the right sequence before reaching for fine-tuning?
Fix your prompt first, then add retrieval for factual grounding, then write an evaluation set that covers your failure modes. Fine-tune only once the eval set stops improving with prompt changes. Most teams who follow this sequence discover they don't need to fine-tune - or they fine-tune with much cleaner data and clearer goals.
What is distillation and how is it different from fine-tuning?
Distillation uses a large frontier model to generate labeled outputs for your task, then trains a smaller model on those outputs. The small model inherits near-frontier quality on the specific task at a fraction of the inference cost. It's a fine-tuning technique, but the key difference is that your training data comes from a stronger model, not from hand-labeled examples.
What is reinforcement fine-tuning and when should I use it?
Reinforcement fine-tuning (OpenAI calls theirs RFT) trains a model against a reward signal - a grader that scores outputs - rather than mimicking labeled examples. It's worth considering when your task has a clear, automatable correctness signal: code that compiles, answers that match a schema, or claims that can be verified against a database. If you can't define a reliable grader, stick with supervised fine-tuning or prompting.