Something counterintuitive is happening at companies that successfully deploy AI coding tools: Output goes up, and velocity doesn't follow it.
The team is producing more code, more PRs, and more deliverables, but the most experienced engineers are spending their days reviewing instead of building. The net effect is that the team's most capable people have stopped producing, and the extra output from everyone else isn't making up the difference.
We've been calling this the Sisyphus Predicament, borrowing the name from a May 2026 study of simulated AI teams. The study measured what its authors call the short-board effect, named for the bucket that holds only as much water as its shortest board: how much one underperforming agent drags down a whole team. Among the results, the paper names a regime it calls the Sisyphus predicament, where a team at a critical capability threshold is spinning in place while showing pseudo-high efficiency. The version I keep seeing in real teams is a little different: the generating side is fine (often very good actually) and the checking side can't keep up. We borrowed the name because in both cases the boulder is back at the bottom of the hill every morning.
Why is this happening? AI made generating work cheap and left checking it about where it was. Every team has some capacity to check and approve work, and it's concentrated in the people who know the work the best. When the rate of generation passes the rate of checking, a queue forms in front of those people, and it refills faster than they can clear it. Even valiant efforts to get back in control of the backlog aren't enough as agents take over more of the generating, because an agent can refill the queue overnight while the reviewers sleep. This can be fixed.
I'm on both ends of this. I use AI to extend my capacity in a variety of areas, and I ensure that any AI output with my name on it has been checked and confirmed accurate. For many projects that means I'm the person filling the queue and the person clearing it.
Most of the examples here are AI-assisted pull requests: one person prompting, another reviewing. That's still where most teams live. As more of the generating shifts to agentic systems, the shape holds and the mismatch widens. The question stops being how good any one engineer's judgment is and becomes how much senior attention each unit of net value costs, whoever or whatever produced it.
How it works
AI coding tools are an output multiplier, but "faster at producing output" and "faster at producing value" aren't the same thing. A change still has to be read by someone who understands the codebase well enough to know whether it belongs there, and that person's time was already scarce before any of this started. Generation got easier while checking stayed expensive and demand for it increased.
What makes AI-assisted output especially expensive to check compared to conventional tools is that the failures don't repeat consistently. Tools that fail the same way every time are cheap to review, because the reviewer learns what to look for and a checklist catches it. AI-assisted output fails wherever the prompt leaves something unsaid or the AI works from incomplete information, so the errors change with the context rather than with the tool. Every piece needs full attention because any piece could be the one with the subtle problem.
The team enters the Sisyphus Predicament when the people who can check the work spend more time reviewing and correcting AI-assisted output than they would doing the work themselves. In the myth, the boulder rolls back down every time it nears the top. Here the boulder is the review queue, the reviewers are the ones pushing it, and generation is what rolls it back.
The five signals
These patterns show up reliably in teams experiencing the Predicament:
1. Review time exceeds generation time. Someone spends 45 minutes reviewing what AI generated in 10. The review-to-generation ratio is the clearest diagnostic: if it's consistently above 3:1, the team is in the Predicament.
2. Rework loops. Generate, review, reject, re-prompt, review, reject, manual fallback. Multiple AI-assisted rounds that end in manual work consumed more total resources than starting manual would have.
3. The most experienced people become review bottlenecks. They stop producing because they're clearing the queue full-time. Their velocity drops toward zero while the team's raw output climbs. This is the most visible signal from outside the team, and it's easy to misread from there, because from outside it just looks like the senior people slowed down.
4. "Faster to do it myself." When experienced team members start saying this, the Predicament has arrived. The statement is usually accurate: it often IS faster, because the review cost exceeds the production cost.
5. Architecture degradation. Each AI-generated PR looks fine on its own, and the codebase still gets worse over time: inconsistent patterns, unnecessary abstractions, duplicated logic with slight variations. Nobody reviewing changes one at a time can see the trajectory, because the trajectory isn't in any one change.
Seeing these signals?
Talk to us about where your review time is going
We've rebuilt our own engineering practice around agentic workflows, including how work gets checked. If your senior people are spending their days in review, we're happy to walk through what's filling the queue and what would bring it down.
Talk to us about your review queueThe metric that matters
Output volume and merge rate both miss it, because neither one counts the senior time a change consumes on the way in. The metric that tracks that is senior attention consumed per unit of net value produced.
Ten changes where seven merge clean cost the team very little. Ten changes where three merge and seven consume senior review time before they're rejected cost more than if the seven had never been generated. Same output volume, very different value, and the hidden cost (senior time diverted from production to review) makes the total picture worse than nothing at all.
Underneath the number is an overall direction. Senior attention either flows into the work or into checking the work, and the Predicament is what it feels like when the net flow reverses.
What the research says
This isn't theoretical, and the most direct evidence for it now comes from a tool vendor's own data. Two kinds of evidence follow, and they're worth keeping distinct: some of it is about people reviewing AI-assisted work, and some is about agents working under automated review.
On August 7, 2026, Anthropic announced that automated permission checks would become the default in Claude Code for Pro, Max, and Team plans on August 14. The study behind the change: across 1,053 paid testers, human reviewers caught a deliberately dangerous command 13.6% of the time, while an automated check caught 89%. By the time a tester had been through 50 prompts, the human catch rate was closer to 5%. Anthropic also reports that users approve 97% of permission prompts, which they read as clicking through rather than reviewing. That is the Predicament stated as an instrument reading: human review degrades under volume, and it degrades fastest under exactly the load AI throughput creates. The result says nothing about how experienced the reviewers were, only about what happens to anyone's attention when the volume of things to approve goes up. The vendor's answer was an automated gate rather than more human attention.
Anthropic's published measurements of AI development inside frontier labs show what that answer looks like at scale. As of August 2026, about 30,000 agents were doing research and engineering work at Anthropic at any one time on its most-used internal platform. An automated monitor checks 100% of their actions before they're executed, and people review any action it blocks. An offline monitor flags roughly 100,000 transcripts a week. Several classifiers resolve most of those flags, and the highest-priority ones, about 50 a week, go to people for review. Those are agents working on one company's own platform (so the numbers describe one place rather than an industry) but it's the clearest picture I've seen of what checking looks like once generation runs at a scale that no review team could cover by hand: the machine checks everything, and people's attention goes to what the machine stopped.
When a team has more than one gap, the short-board study finds their effects compound: the authors write that "team performance is shaped by the aggregated impact of all weaknesses" and that mitigation "should extend beyond the remediation of individual weak links." The WORC study lands in the same place from the other direction: individual errors amplify through collaboration rather than averaging out, and the fix that worked in the paper was spending extra reasoning budget on the weakest agent rather than the strongest to raise the floor.
Two studies fill in the mechanism on the people side. Microsoft's CHI 2025 survey of 319 knowledge workers found that higher confidence in AI correlates with less critical thinking, and that AI shifts the nature of critical thinking toward "information verification, response integration, and task stewardship," which is the review cost this piece is about, measured independently. And BCG's read on what separates the companies getting value from AI isn't a technology gap. They put the split at 10% algorithms, 20% data and technology, and 70% people, processes, and cultural transformation, and their own phrasing is that the soft stuff turns out to be the hard stuff.
What to do about it
The Predicament isn't an argument against AI tools, and the fix isn't to adopt fewer of them. Teams that get real value are usually running the same tools as everyone else, so what separates them is how those tools get deployed and used across the team.
The tempting move is to find whoever is generating the most review load and start there, but it's treating the symptom rather than the cause. The person generating review load is using the same tools as everyone else, and without changes the processes and systems supporting those tools will hand the next hire (and the next agent), the same problem.
Diagnose first. Measure senior review time against what ships, for the team as a whole and then by source: which kinds of change, which parts of the codebase, which workflows are consuming the most attention per unit of value. The metric is "senior hours consumed per shipped unit of value," not "changes generated." If the ratio is high across the board the system is the problem, and that's the usual case.
Raise the floor. This is the heart of the fix, and it comes before any conversation about individuals. The short-board study says as much: mitigation should reach past fixing individuals one at a time. In practice it means three things, and they're the same three things whether the generating is being done by a person with an assistant or by an agent on its own:
- Context the tool reads on every run. Shared conventions, codebase documentation and architectural decisions, written down where the tool loads them (CLAUDE.md files, style guides, decision records). Every prompt is different, and this is the part that stays the same across all of them, which makes it the cheapest way to make the output consistent.
- Automated checks on every change. Tests, linters, content validators, whatever catches the class of error the reviewers keep catching by hand. Many of these can be built once with AI and then run as ordinary code on every change, the same move that keeps an AI budget from spiraling.
- People reviewing what the checks stop. Senior attention goes to the changes the automated layer flagged and to the design questions no check can answer, instead of to every change in the queue. One alarm is useful, but a hundred at once are just noise.
These raise the minimum quality of AI-assisted output regardless of who or what is producing it, and when the floor rises, the review cost drops. It scales, too, because individual coaching has a linear cost while infrastructure compounds across every team member, every agent and every future hire. The vendor data points the same way: the durable answer to review load is an automated gate rather than more standing human attention, and Anthropic's 30,000-agent example is that answer running in production.
Then look at what's left. Once the floor is up some review load usually remains, but it's now small enough to look at closely. At that point the question about a member of a team is what support they need, and it's usually one of three:
- Harness. The context isn't reaching the work. The conventions exist, but this workflow doesn't load them, or the prompts aren't pointing at them. Fixable with better prompting practices and project-specific context. This is the most common cause and the easiest to fix.
- Judgment. The domain knowledge isn't there yet. Evaluating AI output in an unfamiliar part of the system needs knowledge of that part, and that takes time to build. Pairing with an experienced reviewer bridges it, with a defined window and explicit hand-off criteria. Standing human review is the thing the vendor data says stops scaling under volume, so it belongs as a temporary bridge rather than the permanent fix.
- Role. Sometimes a role is structured so that AI amplifies its hardest parts rather than its most valuable ones. That's a question about how the work is deployed rather than a verdict on the person: the fix is to match people to work where their judgment compounds.
The number to watch is the same one throughout: senior attention consumed per unit of net value. Raise the floor and it drops for the whole team at once. Leave it where it is and it climbs with every new tool, and increasingly with every agent pointed at the codebase, because each one adds to the generating side and nothing to the checking side.
Raising the floor
Want help raising the floor?
Our own delivery runs on this layer: context the agents read on every run and automated checks on every change. We can help your team build the same layer around the tools you already use.
Talk to us about raising the floorBritton Russell is Director of Agentic Operations at Rangle, where he helps organizations build AI systems that compound in value over time.





