Coding agents let a small team start far more work than it can finish. Many engineering leaders respond by adding review capacity, often by pointing a second model at the AI-generated code. Adding reviewers moves the queue along without changing how much work enters it. Teams that keep agents productive write down what the software may change, where a person has to sign off, and what has to pass the same way on every run.
Stephen McGowan is Founder of TrolleyRelay, a multi-tenant SaaS product that syncs products, inventory and orders between an in-store point-of-sale system and an online storefront. His client work includes a global insurer, where he serves as AI solutions architect and helped establish the company's AI governance framework. He has also served as CTO or technical founder for several companies, including a regulatory technology startup and a telehealth provider he helped scale through the pandemic. McGowan writes and speaks about AI-enabled development, and published an eight-part benchmark series on LLM code review.
"You start building with AI and it's like, I can do 10 things at once, it's amazing. But what that actually leads to is more things you haven't quite finished, and that's where the real problem is," says McGowan. He traces the pattern to research on delivery performance that predates AI by decades. From his perspective, the number that governs a team's output is how quickly work closes.
The queue behind the review
Work opened with an agent often stalls in several places at once. One branch waits on a test run, another waits on a decision, and a third sits nearly finished. Each new task joins a queue that was already full, and the oldest items in it keep aging while newer ones arrive. "I don't see reviews necessarily being the bottleneck," notes McGowan. "It's the same problem we've had forever. It goes back to Little's Law in 1961, and then to the Toyota production system. Work in progress is the enemy of progress. Review is often the first thing that breaks, but review isn't the cause."
Pull requests on TrolleyRelay took several hours to close when McGowan started. He says they now average three to 15 minutes, and automation accounts for most of that change. He treats the figure as a reading on how much work the platform is carrying at any moment. "I want to go do this new thing," McGowan explains. "The problem is I didn't quite finish that thing yesterday, and I didn't finish the one before it either. All these things stack up, and it gets to the point where you're making no progress at all."
Continuous integration is one of the jobs McGowan now hands to agents outright. Work like that sits further upstream than code generation, and it's the part of the pipeline he automated first. Building it takes longer than most developers expect, and every run waits on a GitHub runner before it even starts. "It's very awkward for context switching," adds McGowan. "They go off and work on something else for another 15 minutes, and then they realize it's a day later and they haven't fixed the problem. Being able to pass it off to AI and say, 'Could you just go make this work for me please?' is hugely powerful."
Making the checks deterministic
Running the same review ten times against the same code can return a slightly different answer each time. A test passes or fails the same way every time it runs against that code, so that's where McGowan puts his guarantees. He also argues that each team has to build its own process, since the rules worth keeping come out of the failures a team runs into. "I don't trust the code I write, and I don't trust the code that anybody else who has worked for me writes," McGowan notes. "At the end of the day we're all fallible, and so are the LLMs. That's why there are so many different gates even in the standard development cycle."
The ticket a piece of work starts from gets the same treatment as the code it produces. One of McGowan's scripts reads a ticket and reports whether it has been checked against the latest architectural decision on record. A ticket that fails goes back for review before an agent touches it. "Everything should be testable," says McGowan. "Even my documentation is testable. When documentation is submitted alongside the code, it has to be in a particular format, with particular headings and files, and I have linting run across it."
The firmest checks on TrolleyRelay run in the pipeline, where a failure blocks a merge outright. Writing checks used to cost enough that many teams skipped the ones they could live without. Agents have now made that work cheap, so McGowan writes one for every rule he wants to hold. "There's a sliding scale of how to enforce these things," he explains. "You can say in a chat, 'Don't do that.' You can say, 'You've done that twice, commit it to memory.' You can add it to your project instructions. These are all very soft ways of doing it. Then at a certain point you give it a script or a tool saying, 'This is how you evaluate it.'"
Deciding what's off limits
TrolleyRelay is middleware, so it isn't the source of truth for anything it handles. A point-of-sale system holds the real stock count, and an online store holds the real record of orders. A wrong number in the middle can be fixed by running the transfer again. Billing and pricing behave differently, because a charge that goes out wrong has already left the system. "As soon as it touches billing or pricing, it has to go to me for human review," says McGowan. "The same goes for anything to do with an architectural decision. At first I was reviewing every commit, then I was reviewing every pull request. Then I asked, 'What do I actually need to review?'"
GitHub enforces that boundary directly, blocking agents from merging any pull request that reaches the folders holding billing and pricing code, and putting the human sign-off in the path itself. McGowan built a second set of rules that run before an agent writes anything, catching work that would otherwise be built on the wrong premise. "If I tell it, 'Can you quickly change this file?', it'll tell me there's no ticket for that," McGowan notes. "If there is a ticket, it'll tell me the ticket hasn't been baselined against the most recent architectural decision. I'm not going to start working on it until we've done that. Assumptions are the place that will really bite you."
McGowan has agents maintain high-level solution design documents for the platform, and he reads through them about once a month. A passage that doesn't match how he thought the system worked can mean the build has a real problem. It can also mean the tickets underneath it are wrong, or that the code and its documentation have drifted apart. "You can't test for good architecture," McGowan concludes. "This is how I'm solving for validating it, making sure the shared assumptions we have about how it should be architected are true."




