The views expressed in this article reflect those of Joseph Feld alone, and do not represent the position of any organization.
When a specification leaves a decision unstated, a coding agent makes that decision itself. The decision is an architectural one, and it lands in the codebase as a database call or an error path the specification never mentioned. Nothing in the code marks that line as a guess, and a reviewer sees it alongside everything else in the diff. Catching it earlier means changing how completely the specification gets written and how many checks stand between it and the commit.
Joseph Feld works as a Senior Full-Stack Engineer in AI Systems Architecture at Ford Motor Company through TEKsystems, where he designed the standard its manufacturing IT organization now uses for AI-assisted software development. He's spent more than thirty years in full-stack engineering, with work spanning Fortran, .NET, Java, Angular, and cloud-native delivery. He also writes on spec-driven development and the divide he expects to open between AI-native and traditional developers.
"Ninety-nine times out of a hundred, you could just take the AI's recommendations and rubber-stamp them all the way through, but that hundredth time is going to burn you. That hundredth time is where the headlines get written," says Feld. His response is a development process built to catch that hundredth case before it ships. His teams have run it both on new applications and on a legacy mainframe system.
Specification before implementation
Feld's teams settle the business intent and the full architecture in writing before a coding agent sees anything. The order reverses how most developers have worked, and Feld resisted it himself when a manager first raised it with him. "All your heavy load is done at the beginning," he says. "We are trying to write a specification so detailed that by the time it goes to a coding agent, there are no architectural decisions left to be made. It is simply implementing at this point."
A design agent produces the first version from a stated business intent, and it's the one model in the sequence permitted to infer anything. The draft goes back to the business teams to confirm the software it describes is the software they asked for. Feld has yet to see one go through his engineers unchanged. "This stuff goes in front of the actual human designers, and they have to approve it and adjust it," he notes. "And they do."
The sequence has no debugging step. Feld objects to long debugging passes because the changes they make rarely get written down anywhere. "We don't go in and change the code after we're done," he explains. "If something comes out wrong, you go back and amend the spec." Regeneration takes about half an hour, and the finished specification serves as the documentation.
The chain of instruction
Feld's workflow assigns separate models to separate roles. A person directs an architect model, and the architect passes instructions down to the builder that writes the code. Feld found that a model starts drifting when the language it receives stops matching the role it was assigned. "Nobody's talking to it but the architect," Feld notes. "And the only thing the architect ever gives it is instructions."
Three bootstrap layers load before a project starts. One carries the corporate coding rules, one carries the rules for the project, and one carries the rules for the tech stack in use. Work is broken into implementation chunks to keep the context window small. Feld explains that specification stays out of the conversation entirely, held in files the models re-read as needed. "All of these bootstraps have restrictions against inference," he says. "They have a long list of don'ts, things they cannot do in their code."
One check runs before implementation and one after. The builder first repeats back what it understands the instruction to be, and the architect confirms that the description matches the intent before any code is written. "Before the code gets checked in, we go back and we do a zero drift audit on the completed code," Feld explains. "Basically, what did you just build? Describe it." Drift still occurs, but it takes predictable forms, with the architect jumping a step or the builder writing code before anyone asks it to.
Business intent in schema design
Schema design is the one place Feld found the standard sequence breaking down. The enterprise doesn't map its data models one to one onto data transfer objects, and the model gets that judgment wrong consistently. "We need to have a separate workflow specifically for schema design," he says. "We cannot mix in schema design with the rest, because what ends up happening is it starts hallucinating DTOs and we have to correct a lot."
Inside the separate workflow, the sequence matches the one governing application code. "We give it the business intent for what we need to store there, and we let it assign an overall schema," Feld observes. "Then we put it in front of our data engineers to pick it apart or add suggestions." The approved specification then goes to a coding assistant that builds the schema, and the same drift audit runs on the result before anything is committed.
Legacy systems arrive with no written intent at all. Feld's teams reverse engineer those codebases to reconstruct what the business originally wanted, separating it from architectural decisions no longer worth carrying forward. One mainframe application they modernized was roughly forty years old, and even its comments were unreliable. "We actually found a comment in the code that said 'remove when permanent fix supplied,'" he recalls. Everyone who could have explained it had long since left the company, so the team built scenarios from the code and put them to the current business owners.
Owning an unsupervised mistake
Vendors have been encouraging teams to let long agent runs proceed without interruption. Feld treats that as an operations question before an AI question, and he says the accountability problem is older than the technology. "If the AI gets something wrong at two in the morning, who owns that?" Feld says. "You've got no human supervision, so who owns fixing that?"
The failures Feld has watched began with a process that was already weak to start with. An agent working at speed accelerates whatever the workflow permits, and the output arrives faster than reviewers can read it. "You can't leave this thing unbound yet," he adds. Frequent human checks catch a small divergence early, before it grows into a run the team has to abandon.
Oversight of that kind depends on the reviewer being able to do the work themselves. Feld tells the developers on his team that the skills are what keep them relevant. He gives the same message to the ones straight out of college, who worry that what they learned no longer applies. "If you can't build it, you can't run a properly human-governed workflow," Feld concludes. "You're not going to understand what you're seeing."




