A Friday Prompt Change That Broke Monday Morning

LLM evals are the quality layer most teams skip until a regression finds them. Here's how evaluation pipelines actually work - and why they're nothing like unit tests.

Cover art for A Friday Prompt Change That Broke Monday Morning

A team ships a new chatbot prompt on Friday. By Monday, customer-support escalations are up 14 percent. Nobody knows which slice of users is affected. There's no offline regression suite that ran on the prompt change, no faithfulness score on production traces, no quality signal in the dashboard - just token usage and latency. Every change was a coin flip; every regression got found by users instead of CI.

That is the state most teams are in. LLM evals exist to fix it, but they work nothing like the tests you already know.

What LLM evals actually are

LLM evaluation is the process of verifying whether your AI application is performing as intended - quality control for AI outputs. You provide your system with test questions, check the answers it produces, and assess whether they meet your standards for accuracy, relevance, and usefulness. The key difference from testing conventional software is that AI doesn't always produce the same output twice, and a "good answer" often depends on context rather than a single right answer.

That last point is the one that trips people up. LLM pipelines don't produce deterministic outputs; their responses are often subjective and context-dependent. A response might be factually accurate but have the wrong tone, or sound persuasive while being completely wrong. This ambiguity makes evaluation fundamentally different from conventional software testing.

LLM evaluation has three distinct components: the test dataset - a curated collection of representative inputs, typical cases, edge cases, and adversarial inputs; the success criterion - an explicit definition of what correct output looks like for each case; and the evaluator - the mechanism that scores outputs against those criteria, which can be a script, a human reviewer, or another AI model.

The third option - another AI model as evaluator - is now the dominant approach for anything open-ended.

The evaluator types, from fast to fuzzy

Not every output can be scored with a regex. Teams layer four types of evaluation, roughly in this order:

  • Deterministic checks. Deterministic metrics compare model output to a ground truth using a fixed algorithm. They're fast, cheap, and reproducible, but fail on open-ended generation where many surface forms are correct. They're the right tool for math, classification, code, and any task with a single right answer.

  • Heuristic scoring. Regex, keyword presence, format validation. Low cost, brittle at the edges.

  • LLM-as-a-Judge. Combines deterministic tests, reference-based scoring, LLM-as-a-Judge, human review, and production signals to measure behaviors that do not have one exact correct output. A stronger model (often GPT-4o or Claude Opus) reads the output and applies a rubric. The non-obvious failure mode: the judge can inherit the same biases as the model it's evaluating - it tends to prefer verbose answers and fluent-sounding nonsense.

  • Human review. Human-in-the-loop evaluation is a step where a person scores the output of an AI system. Your automated scorers and LLM judges handle the bulk of evaluation. Human review handles the cases they can't.

Evaluator type Cost Speed Works on open-ended? Catches tone/bias?
Deterministic Very low Fast No No
Heuristic Low Fast Rarely No
LLM-as-a-Judge Medium Medium Yes Partially
Human review High Slow Yes Yes

The strongest production setups layer all four. Deterministic checks for format, heuristic scoring for quality, LLM-as-Judge for nuance, humans for calibration.

Offline evals vs. online evals - and why you need both

Offline evaluation checks quality before changes are released. Teams test updates using a fixed set of example questions to make sure nothing breaks when prompts, models, or retrieval logic change. Online evaluation checks the quality after the system is live. It analyzes real user interactions to detect issues that don't appear in test cases, such as new question patterns, edge cases, or gradual performance drops.

Most teams start with offline evals and stop there. That's a mistake. The same evaluator may be used before release on a curated dataset and after release on sampled production traces, although production evaluation introduces additional requirements around latency, cost, privacy, monitoring, and incident response.

Integrating LLM evaluation into CI/CD pipelines fundamentally changes how teams develop AI systems. Instead of treating evaluation as a one-time checkpoint, it becomes a continuous process that runs alongside every code change, prompt modification, or model update - allowing teams to catch potential issues before they reach production.

The 2026 practice is to fail the pipeline if any of the headline metrics - faithfulness, answer correctness, task adherence, safety - drops more than a defined threshold, and to flag warnings on smaller drops.

You usually look at benchmark results once, when you're choosing a base model. You run evals continuously, as you change the agent around it. Benchmarks tell you what the model can do. Evals tell you whether your system does what you need, today.

Beagle in action#eng-ai, 3:47pm Tuesday
The ask
'the new summarisation prompt is in - anyone run evals on it yet?'
Beagle drafts
pulls the eval dataset from the linked Notion doc, drafts a summary of which test cases changed score and by how much
You approve
you review and post the diff in 30 seconds; the PR reviewer has signal before they click merge
Do this in your workspace →

What a working eval pipeline looks like in practice

Here is the workflow a team running a support-ticket classifier might actually use, from first test case to production monitoring:

  1. Build the dataset from production failures first. Start with 50 production failure cases. Grow it every time you find a bug. Cold-start test cases that nobody ever typed are nearly useless.

  2. Write success criteria before the prompt. What does a correct classification look like? A routed ticket with the right label, under 200ms, no hallucinated fields. Make it checkable.

  3. Run offline evals on every prompt change in CI. OpenAI's evaluation documentation makes this operational: run evals on every change, not just at launch. This practice - eval-driven development - treats evaluation as infrastructure rather than a final quality check.

  4. Sample live traffic and score it. A 5% sample scored by an LLM judge catches the drift that test cases miss.

  5. Close the loop. Add failures and boundary cases to the eval dataset. Update the prompt, model, retrieval, tools, or orchestration. Run an experiment against the baseline. Release the change when it meets the quality bar. Continue monitoring for new failures.

GitLab shares how they build Duo, their suite of AI-powered features. They created an evaluation framework with thousands of ground truth answers, which they test daily. They also have smaller proxy datasets for quick iterations. That shape - a large slow suite and a small fast suite - is the right model. The large one finds regressions; the small one unblocks developers.

85%GenAI projects that faildue to bad data or inadequate testing, per Gartner
95%cache hit rate neededfor ~76% cost reduction in a support app with long system prompt
5%live traffic sampletypical starting point for online eval coverage

Anthropic's engineering team identifies three distinct infrastructure components: an eval suite (specific task capabilities), an agent harness (tool calls and orchestration state), and an evaluation harness (end-to-end tests run concurrently, aggregating results at the trajectory level). If you're evaluating an agent, you're grading the whole sequence, not just the last response.

Catching a prompt regression
Without Beagle
the new prompt ships on Friday; escalations spike Monday; someone spends half a day correlating the deploy with the support queue
With Beagle
the eval suite runs on the PR, flags a 12-point drop in answer-correctness on the refund-policy test cases, and blocks merge before it ships

LLM evals: common questions

What is LLM evaluation?

LLM evaluation is the process of assessing the performance of large language models by using tasks, data and metrics to gauge their effectiveness. In practice it means running a set of representative inputs through your system, scoring the outputs against defined criteria, and treating a drop in score as a blocker - the same way a failing unit test blocks a deploy.

How is LLM evaluation different from normal software testing?

LLMs are non-deterministic, meaning there's no guarantee they'll provide the same answer to the same question twice. This makes it more complicated to verify that things work according to spec than it does with other software, for which automated tests are available. You measure pass rates across a test set, not the presence or absence of a single right answer.

What is LLM-as-a-Judge?

LLM-as-a-Judge uses a second, typically stronger language model to score the output of your application against a rubric. It handles open-ended outputs - tone, factual coherence, policy compliance - that a regex or exact-match check cannot. The main risk is that the judge can share the biases of the model it evaluates, so calibrate it against human-labeled examples before trusting it in CI.

How do you build an eval dataset?

You need an LLM evaluation dataset: a collection of sample inputs paired with their approved outputs. You can generate such a dataset or curate it from historical logs, like using past responses from human support agents. The closer these cases reflect real-world scenarios, the more reliable your evaluations will be. Start with your actual failure cases - not hypothetical ones.

When should evals run in a CI/CD pipeline?

Evals are becoming part of CI/CD pipelines. The right trigger is any change that touches the prompt, the model version, the retrieval logic, or the tool definitions - anything that could shift output quality. LLMs are non-deterministic by nature. The same prompt can produce different outputs across runs, and subtle changes in retrieval pipelines, model versions, or prompt templates can quietly degrade quality without triggering traditional error alerts.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle