AI companies measure new models with benchmarks, standardized tests that score each model and rank it on public leaderboards. Some engineers believe models are now tuned too closely to those scores, a pattern they call "benchmaxxing." The tests often favor eye-catching tasks, like rendering a 3D scene. A model can rank first on them and still fall short on the everyday work enterprise teams depend on. Those teams need models that follow instructions and ask questions when something is unclear, a need that grows as companies give AI agents larger tasks with less oversight. Engineers who build production systems at scale see the difference most clearly.

Sairam Krishnan is a Senior AI/ML Engineering Leader at Apple, where he leads a multi-site team building the media knowledge graph that supports search, recommendations, and discovery for more than 100 million users. He previously led an Alexa data team at Amazon as a Senior Software Engineer, and he teaches AI strategy to executives as a guest instructor at Praxtera AI Institute. Krishnan uses frontier models daily for everything from design plans to code reviews, and he sees a clear trade-off between what scores well and what production systems need.

"In the enterprise, the work is supposed to be reliable and stable. For that, what you need to do is slow down," says Krishnan. Newer models often try to finish entire projects in one pass and rarely pause for feedback, so engineers have a harder time catching mistakes early.

In the enterprise, the work is supposed to be reliable and stable. For that, what you need to do is slow down.

Sairam KrishnanSairam KrishnanSenior AI/ML Engineering Leader, Apple

Questions before code

Krishnan traces part of this to how newer models are trained. Earlier models learned mostly from people reviewing their work, while newer ones learn more from automated feedback, including scores from other AI models. The result behaves like an eager junior intern who wants to prove they can work alone. How much independence a model should have depends on the task. A front-end developer may want it to stop for sign-off at every step, while a back-end engineer with a finished plan may want it to follow the plan without interruption.

Before writing any code, a model should question the user the way a good product manager would. It should ask why a feature is needed and whether the codebase already has something that does the job. "You need a handbrake that says, 'Don't jump into implementation. Let's think about what you're trying to do,'" Krishnan notes.

Next, the model should list its assumptions so the user can fix any wrong ones before work starts. Teams could also give each step its own agent. One agent challenges the goal, a second designs the solution, and a third plans the tests and breaks the work into weekly tasks. Krishnan gets better results when he asks a model to teach him how to solve a problem than when he asks it for a plan, since explaining is something these models do especially well.

Grounded in the database

Good questions only help if the model knows what's already in the codebase. Coding agents often skim files for keywords without learning how the system fits together, so they end up rebuilding code that already exists. They also forget. Once a session ends, earlier decisions are lost unless someone writes them down.

Knowledge graphs, databases that map how pieces of information connect, can give agents a longer-lasting memory, with a separate model for each kind of knowledge. "One local model could be your expert on the codebase. It only knows the codebase. The other one only cares about the business requirements. The bridge could be a more powerful model that understands how to make these two talk to each other," Krishnan adds.

That way, an agent building a new feature can find related work and reuse existing code. The setup also protects private information. Proprietary code and user data stay on the company's own servers, and anything sent to an outside model gets scrubbed of sensitive details first.

Spreading knowledge across several models means they all have to stay in sync, a problem cloud computing faced when it first took off. "You really need that data consistency. I think you need the database to be the source of truth, that common, agreed-upon contract between all the agents," Krishnan warns.

Teams often guide agents with written instruction files, but agents can treat those instructions as optional. Krishnan favors a set of firm rules that every agent must follow. A planning agent could gather those rules at the start of a project, and a shared system would hold every agent to them. Model developers could build that discipline in during training, test for it afterward, or both.

Pick one pain point

Every agent costs money to run and time to manage. Teams can start with one pain point, for customers or inside the business, and try writing out the steps to solve it. "If you can, then you can automate it and make it a workflow. If you can make it a workflow but you also need some customizability, then make it an agent," Krishnan says. Strict guardrails only matter once a team runs several agents, and most companies aren't there yet. Plenty still need to get their data into a database before any agent can put it to work.

AI still has a lot to offer enterprise teams, and many will find its best uses once the initial buzz fades. The Gartner hype cycle, which tracks how new technologies move from inflated expectations to a slump and then to practical use, calls that slump the trough of disillusionment. "We need to get into that trough of disillusionment for people to start thinking, let's not jump into the deep end. Let's think about what we need and address one problem at a time," says Krishnan.

The views and opinions expressed are those of Sairam Krishnan and do not represent the official policy or position of any organization.