Skip to main content
Blue code blocks flowing fast into a delivery pipeline and piling up at an amber review gate
AI engineering

Agentic AI is moving the bottleneck in software delivery

AI agents write code faster than teams can review, test and release it. Where work now waits in the delivery lifecycle, how to measure it, and what to fix first.

Idan Sofer
Founder, iSofer · · 5 min read
Key takeaways
  • AI made writing code faster; review, testing and release still run at the old speed.
  • Measure where work waits before adding reviewers or testers.
  • Capture the intent of every change and keep pull requests small.

AI coding assistants and agents have changed how fast code gets written. A developer can now describe a feature, let an agent draft it, review the result and open a pull request in a fraction of the time it used to take. For engineering leaders, the promise is obvious: more output from the same team.

Code that has been written is not code that has shipped.

In many teams we work with, the extra output is not reaching production any faster. It is piling up in review queues, waiting for testing, or sitting behind a manual release process. The bottleneck has not gone away. It has moved downstream.

The code got faster. The delivery didn’t.

Software delivery is a chain: understand the problem, write the code, review it, test it, release it and learn from how it behaves in production. For years, writing code was the slowest link, so most of the tooling, hiring and process was built around it.

Agentic AI speeds up that one link dramatically. Every other link still runs at the old speed. The result is predictable: more pull requests opened than reviewed, more features waiting for testing than being tested, and more changes queued for release than released. Developers feel productive, while the business sees the same delivery dates as before.

Where the work waits now

When we look at a team’s delivery flow, the new queues usually appear in four places:

  • Code review. Agents produce larger changes, and more of them. Reviewers who once handled a few pull requests a day now face a steady stream, and big diffs get either skimmed or left waiting.
  • Testing and quality assurance. When testing is a separate stage, or a separate team, it receives finished work in batches. Its capacity has not grown with development output, so the queue grows instead.
  • Release. Manual approvals, release windows and fragile deployment pipelines were tolerable when changes arrived slowly. With more changes arriving, they become the limit.
  • Decisions. Agents can build whatever they are asked to build, quickly. When requirements are vague, the team builds the wrong thing faster, and the wait moves to product decisions and rework.

Why handoffs hurt more with agents

Before anyone can review or test a change properly, they need to understand its intent. What problem does it solve? Which scenarios are in scope? What are the edge cases, and what was deliberately left out?

In a team where developers, testers and product people work closely together, that context builds up naturally as the feature evolves. In a team where work moves through formal handoffs, from a development column to a testing column days later, context leaks at every step. The tester has to go back and ask what was meant, and the developer has to remember a change made three features ago.

Agents make this worse in a subtle way. When a developer writes code by hand, the reasoning lives in their head and can be recovered with a conversation. When an agent writes most of the code, much of the reasoning lived in a prompt and a chat session that nobody else saw. If that intent is not captured with the change, it is simply gone.

Measure before you add people

The tempting response is to add more reviewers or testers. Sometimes more capacity is needed, but it is rarely the whole answer. If the process still depends on late handoffs and missing context, adding people scales an inefficient model, and every new person also needs time to learn the system.

Start by measuring where work actually waits. Break your cycle time into stages: time to first commit, time a pull request waits for its first review, time in review, time waiting for testing, time to deploy. Look at pull request size and at how long items sit in each column of your board. These numbers usually point to one or two queues that matter far more than the rest.

This is why our AI-first engineering work includes an engineering insights dashboard. It is hard to remove a bottleneck you cannot see, and it is easy to optimise the wrong stage.

What to fix this quarter

You do not need to redesign your organisation to get relief. These changes work inside most existing team structures:

  • Ship intent with every change. Ask the agent to write the pull request description from the ticket and its own plan: the problem, the expected behaviour, the affected areas, the edge cases and the known risks. Reviewers and testers start with context instead of reconstructing it.
  • Keep changes small. Set limits on pull request size and ask agents to split work into reviewable steps. Small changes get reviewed quickly and properly; large ones do not.
  • Put the tests in the same change. Agents are good at writing tests alongside the code. Reviewers should read the tests first: they show what the author believes the change does.
  • Use AI for impact analysis. Let AI summarise what changed, which components are affected and where similar defects appeared before, so human testers spend their time on the risky parts.
  • Make quality gates explicit. Decide where automated checks are enough, where human review is required and where extra validation is needed, especially for security, data and payments.
  • Automate the path to production. Reliable CI/CD pipelines and preview environments turn release from a scheduled event into a routine step.

The longer-term fix: quality in the flow, not after it

The lasting answer is to stop treating testing as a stage that receives finished work. In an AI-native team, developers and quality engineers agree on the intent of a change before it is built: the acceptance criteria, the risky scenarios and the edge cases. Agents then build against that shared definition, and automated checks verify it.

The tester’s role shifts from writing test cases for work they have just seen, to shaping intent up front, challenging assumptions, exploring what automation misses and supervising automated verification. Human judgement stays central. AI handles more of the creation and checking around it.

This is a change in how teams work, not just in which tools they use, and that is usually the harder part.

From faster code to faster delivery

The first wave of AI adoption in software engineering focused on writing code. The teams that benefit most from the next wave will be the ones that speed up everything after it: review, testing, release and the decisions that come before the code. That means capturing intent early, measuring where work waits, and building quality into the flow instead of inspecting it at the end.

Free · 30 minutes

Not sure where work is waiting in your delivery pipeline?

Book a free 30-minute consultation and we will walk through your flow from ticket to production with you, find the biggest queue and suggest what to fix first.

Idan Sofer
Founder, iSofer

Senior engineer and founder of iSofer, which embeds hands-on AI and cloud engineers in startup and enterprise teams.

LinkedIn →

Keep reading