Stop Measuring Agentic Coding Tools by How Fast They Write

Developers now spend more time reviewing AI-generated code than writing new code. The productivity story for agentic coding tools has a measurement problem - and it changes how you should deploy them.

Cover art for Stop Measuring Agentic Coding Tools by How Fast They Write

Developers now report spending 11.4 hours a week reviewing AI-generated code versus 9.8 hours writing new code

  • a reversal that happened quietly inside a single year. Nobody put out a press release. The vendor dashboards still show PRs merged and suggestions accepted. But the actual shape of an engineering day shifted, and most teams are still measuring the wrong thing.

This is the argument: agentic coding tools are genuinely useful, the productivity number is probably positive now, and the debate about them is almost entirely aimed at the wrong variable.

The productivity number you keep hearing is real - and beside the point

METR's original controlled trial found AI caused tasks to take 19% longer, with a confidence interval between +2% and +39%; their early 2026 follow-up with returning developers estimated a speedup of 18%, with a confidence interval between -38% and +9%. The study flipped sign in a year. Both newer confidence intervals include zero - neither slowdown nor speedup is firmly established.

Vendor figures are less equivocal, and more suspect. GitHub reports average task completion time of 1 hour 11 minutes versus 2 hours 41 minutes for the control group - a 55% reduction.

METR's slowdown was measured with Cursor running Claude 3.5 and 3.7 Sonnet on mature codebases; GitHub's 55% figure came from Copilot on a greenfield task with no existing conventions to violate. Both numbers are real. They describe different jobs.

The honest summary: routine task savings run around 46% according to McKinsey's 4,500-developer study, but complex work shows under 10% gains. That gap is the actual signal. Agentic coding tools are fast at what is already well-defined, and slow at what is not. Teams that have figured this out route accordingly. Teams that haven't are measuring total PR count and wondering why velocity feels flat.

Why the review bottleneck is harder to fix than the generation problem

The explosion in AI-generated code hasn't yet led to the productivity gains most expected. Instead, a verification bottleneck has emerged, creating a new set of challenges.

While 70% of developers say AI has positively impacted their time-to-market, fewer than half - 47% - say it's had a positive impact on end-user experience or on reducing technical debt. That gap describes a team shipping code that looks done but isn't reliable.

Here's what makes this structurally hard. Generation is fast because the model operates without ambiguity about its own output. Review is slow because a human has to reconstruct ambiguity: does this implementation match the intent, does it break something upstream, does it introduce a subtle vulnerability? Review fatigue is an underreported productivity drag; when the AI produces more code than a developer can meaningfully review, teams either merge under-reviewed work or queue PRs indefinitely - both failure modes surfaced in 37% of long-form survey responses.

Productivity gains also plateau after about 180 days, with self-reported speed jumping 34% in the first 60 days then flattening, and gains concentrating in specific task types rather than across-the-board velocity. The early wins are real. The problem is that teams extrapolate from them.

11.4 hrsreviewing AI-generated codeper developer per week, Q1 2026
9.8 hrswriting new code with AIdown from a 4-hr lead in Q4 2024
37%of survey respondentsreport merge-or-queue failure modes from review overload
46%routine-task savingsvs under 10% on complex tasks (McKinsey, 4,500 devs)

What teams that actually gained ground did differently

The teams with measurable, durable gains did two things that most coverage ignores.

First, they routed tasks explicitly. Agentic tools go to: writing tests for existing logic, drafting boilerplate to a known spec, generating a first-pass migration script that a human will review before running. They do not go to: the hairy distributed systems problem with six implicit constraints that aren't written down anywhere. A follow-up study revisiting technical workers whose AI-assisted workflows had matured found output ran 1.4 to 2 times higher on comparable tasks. The tool didn't change. The routing did.

Second, they stopped measuring what the tool produces and started measuring what it costs to verify. Healthy ROI on AI coding tools runs 2.5-3.5x on average and 4-6x in the top quartile - but only when the cost denominator includes actual token and usage-based costs, not just seat licenses. The seat license is the visible line item. The 11 hours of review time per developer per week is not. Organisations that track quality alongside velocity consistently outperform those chasing speed alone.

There is a narrower operational point here that rarely surfaces. One enterprise team found an 18% productivity lift tied to AI usage, but only after identifying that rapid context switching between multiple AI tools was creating spiky commit patterns that disrupted team-wide flow; once they standardised on fewer, more deeply integrated tools, rework rates dropped 3x while gains held. Tool sprawl is a real cost. A team running Cursor, Claude Code, and Copilot in parallel for different workflows is not tripling its surface area of help - it's tripling the cognitive switching load on every reviewer.

Beagle in action#engineering, 2:47pm
The ask
'anyone reviewed the auth migration PR? been sitting three days'
Beagle drafts
checks open PRs, identifies the stale review, drafts a summary of what changed and tags the right reviewer based on prior authorship
You approve
you approve the nudge; the PR gets a human set of eyes the same afternoon - the queue shortens, not the review
Do this in your workspace →

Steelmanning the other side

The counter-argument deserves a fair run: maybe the review bottleneck is temporary, an adoption curve artifact. The tool didn't fundamentally change across a ten-month period - what changed was the habit: knowing when to trust a suggestion, how to structure a task so the AI has enough context, and how to review output like a colleague's work rather than rubber-stamping it. If that is true, then the right move is patience, not pessimism.

That case is plausible. 69% of developers in the METR study kept using AI tools after the experiment ended, despite being measurably slower

  • a behavioral signal that they were sensing value that the task timer wasn't capturing. Habits that persist against measured friction are usually habits that pay off eventually.

But "eventually" is not a team plan. And the 180-day plateau suggests the habit does not keep compounding on its own.

Treating all code tasks the same vs. routing by type
Without Beagle
agent runs on a complex cross-service refactor; developer spends two days reviewing output that solves the wrong problem
With Beagle
agent handles the test suite and migration scaffolding; developer writes the architectural logic directly, reviews the rest in 90 minutes

The right framing isn't "do agentic coding tools work?" They do, for certain tasks, and the evidence is now positive enough to act on. The right question is: which tasks, and at what review cost? Teams that answer that question concretely - not with a vibe, but with an actual tally of where review hours are going - are the ones where the productivity gains stick.

Agentic coding tools: common questions

Do agentic coding tools actually increase developer productivity?

The honest answer is: yes, on routine and well-defined tasks, probably not on complex work with implicit constraints. METR's 2026 update estimates an 18% speedup for experienced developers, flipped from a 19% slowdown a year earlier - both with wide confidence intervals. Vendor figures average 55%. Task type is the key variable.

Why are developers spending more time reviewing code than writing it?

Agentic tools generate pull requests and review-ready diffs at a pace that outstrips human review capacity. In Q1 2026, the median developer spent 11.4 hours reviewing AI-generated code versus 9.8 hours writing new code - a reversal from 2024. The output rate went up; the verification rate did not.

Which tasks should teams route to agentic coding agents?

Route agents to well-specified, bounded work: test generation for existing logic, boilerplate against a known schema, first-pass migration scripts. Keep agents away from tasks with implicit constraints, complex multi-service interactions, or work where the spec lives only in someone's head. Productivity gains concentrate in the first category and mostly disappear in the second.

Does it matter which agentic coding tool you use?

Yes, but less than how you use it. Teams standardised on fewer, more integrated tools report lower rework rates than those running multiple agents in parallel. GitHub Copilot's deep integration with issues, PRs, and CI/CD creates a tighter review feedback loop than most third-party tools

  • the value is in the closed loop, not the raw generation speed.

What should engineering leaders actually measure?

Measure review hours per developer per week alongside PR cycle time and rework rate. If review hours are climbing faster than velocity, the tool is producing faster than the team can absorb. That is a routing problem, not a tooling problem. Organisations that track quality alongside velocity consistently outperform those chasing speed alone.

Or just watch me work

Point me at your website.

I will read up on your business and come back with what I would run for you. No account, no card, about a minute.

I only read what is public. Nothing is saved to your name until you say so.

Keep reading

Beagle does this work for you, in your Slack.1,000 free credits. No card.Hire Beagle