Senior engineers are spending 4.3 minutes reviewing AI-generated pull requests - versus 1.2 minutes for human-written code - while PR volume has roughly doubled. That math does not close. Faster generation was supposed to be the unlock. It turns out it just moved the constraint one stage downstream.
Most of the conversation about agentic coding tools is still about which model writes the cleanest diff. That was the right question in 2024. It is not the right question now.
The review queue is where agentic coding tools stall
When a coding agent finishes a task, it opens a pull request. The engineer's job used to start at the keyboard; now it starts at the diff. Coding agents are genuinely good and getting better fast. The hard part of engineering moved from writing code to deciding whether to trust it, which makes review the most leveraged skill in software right now.
That shift has a cost most teams are not tracking. The most visible change is the volume of code and the size of the PRs carrying it. Across a sample of 1,450+ engineering organizations in Swarmia, median batch size - lines changed per PR - roughly doubled between Q1 2025 and Q1 2026. More lines per PR means more context to reconstruct. A 2025 study found that senior engineers spend an average of 4.3 minutes reviewing AI-generated suggestions, compared to 1.2 minutes for human-written code. Meanwhile, the volume of PRs has roughly doubled.
The paradox is real: the tools that made developers more productive created a downstream review problem that those same developers now have to absorb manually.
The problem compounds at the cost level. A 2026 study of eight frontier models on SWE-bench Verified found that agentic coding tasks consumed about 1,000 times more tokens than code reasoning and code chat. Runs on the same task varied by as much as 30 times, and higher token use did not consistently improve accuracy. That variance is almost entirely invisible inside most teams' dashboards - they see seat licenses, not run-level token spend.
What the numbers actually say about AI coding productivity
The productivity gains from AI are real, but raw output overstates them: about four times the code for a tenth more delivered value. The gap between those numbers is review work, which is exactly why review is where the leverage now sits.
AI tool costs in 2026 are no longer trivial seat licenses. Inline completion tools cost $20-60 per engineer per month. Agentic tools - Claude Code, high-autonomy Cursor agents, custom LLM pipelines - introduce usage-based token costs of $200-$2,000+ per engineer per month. Most engineering teams now use tools from multiple tiers, making the total cost per engineer $200-$600 per month on average.
At that spend level, a 1.6x productivity return is not obviously a win. If token costs run higher because engineers use agentic tools for low-value tasks, ROI can drop below breakeven. Quality gates prevent rework from eating the gains, and complexity-adjusted throughput ensures AI tools are applied to high-value work.
The honest version: teams that are still measuring success by PR count or lines of code are running the old playbook with new tooling. Those lousy metrics are returning with a vengeance. Organizations have bragged about their percentage of new code written by AI. Engineers might be ranked by token usage. That is Goodhart's Law in a new jacket.
The steelman: more output is still progress
Here is the case for the other side, stated honestly. Engineers who have leaned into agentic workflows are more optimistic about agentic engineering than ever before. The agents are genuinely good, get better every month, and enable shipping things that would not have been attempted a year ago.
Anthropic's internal Code Review agent reports under 1% of its findings marked incorrect by their engineers, and raised their internal rate of PRs receiving a substantive review from 16% to 54%. That second number matters: the long tail of changes that previously got a glance and a rubber stamp is now actually read. If AI is the reviewer as well as the writer, the review capacity problem becomes solvable.
The right question in 2026 is not "which model is best?" It is "which agent workflow fits my codebase, budget, and risk boundary?" That framing is correct. The problem is most teams answer it by evaluating generation quality in isolation and stopping there.
DeepSeek Harness, released on August 13 under an MIT license, makes the architectural stakes concrete. Built on the Cordis meta-framework, it is an agent harness built around the principle that everything is a plugin. Models, tools, skills, sessions, sandboxes, execution loops - even the graphical interface - are all swappable components a developer can mount or replace from configuration without touching the project's source code.
The repository had more than 210,000 GitHub stars by September 2. That adoption is a signal: developers are not just picking a model, they are picking an orchestration layer. The model is increasingly interchangeable. The harness - and what it does to your review workflow - is not.
Stop watching the wrong part of the pipeline
The engineer is the final approval gate. Never skip this. That principle is not just about safety - it is where most of the productivity is being lost.
The frame to apply: treat your agent's output the way a good engineering manager treats a strong junior hire. High throughput, real capability, but the review burden is yours, and it scales with their output. The question is not whether to hire the junior; it is whether you have the review capacity to make use of them.
Right now, most teams are adding agents faster than they are building review infrastructure. A 2026 paper analyzing developer discussions about AI-generated code included one line that has stayed with engineers reading it: reviewing an agent's PR made them "the first human being to ever lay eyes on this code." That is not a failure of the tool. It is a workflow that was never designed for this production rate.
The places where agentic coding is actually compounding - rather than just increasing raw output - share a pattern: they use the agent for well-scoped tasks, they have fast test suites that give the agent signal, and they treat the review queue as a first-class engineering constraint. Fast feedback loops enable productive agentic workflows: fast compilation, fast tests, fast tool responses. Hanging tools break agent flow; if your toolchain is slow, agents will struggle.
A teammate like Beagle can help surface which agent-opened PRs are sitting unreviewed and flag them before standups - not as a code review tool, but as the connective layer between what the agent shipped and who on the team needs to know about it.
Agentic coding tools and code review: common questions
Does using a coding agent increase my team's total review burden?
Yes, measurably. Median PR batch size has roughly doubled since Q1 2025, and AI-generated code takes around 3.5x longer per PR to review than human-written code. Net effect: total review time across a team rises faster than output does, which is why review capacity is now the binding constraint for most agentic workflows.
Should I use an AI code review tool alongside a coding agent?
Probably. Dedicated review agents like CodeRabbit can raise the proportion of PRs getting substantive review substantially - Anthropic's internal tooling moved that figure from 16% to 54%. The key is using separate agents for generation and review: the model that wrote the code is too close to it to catch its own errors reliably.
How do I measure whether my agentic coding tool is actually helping?
Track cost per durable outcome - a change that merged, reached production, and did not require rework. PR count and token usage are activity signals, not productivity signals. Agentic runs on the same task can vary 30x in token consumption without improving accuracy, so run-level cost and outcome data matter more than headline seat costs.
What is DeepSeek Harness and why is it relevant to this debate?
DeepSeek Harness is an open-source, MIT-licensed agent harness released in August 2026. Its "everything is a plugin" architecture makes the model, tools, session store, sandbox, and review checkpoints all swappable. It crossed 210,000 GitHub stars in under three weeks. Its significance: it makes visible that the real competition in agentic coding is now at the harness layer, not the model layer - and the harness is where review workflow gets designed in or omitted.
What is a realistic ROI for agentic coding tools?
Somewhere between 1.6x and 3.5x for most teams, against a total cost of $200-$600 per engineer per month across tool tiers. Top-quartile organizations reach 4-6x, not by spending less but by applying expensive agentic tools to genuinely hard work and running quality gates that prevent rework from consuming the gains.