Not Every Problem Deserves an Agent
Most agentic AI failures happen before a line of code is written. The use case was wrong. The cost model was missing. Two frameworks most teams skip, and pay for later.
Most agentic AI failures happen before a line of code is written. The use case was wrong. The cost model was missing. Two frameworks most teams skip, and pay for later.
Designing agentic systems is like teaching a class of toddlers to behave. You set the rules. You arrange the room. But you cannot test every move. The design problem changed. Most enterprise teams haven't noticed yet.
Your team ran evals. Every test passed. The results look like the QA matrix you've signed off on a hundred times. They're not. Here's what AI evals actually catch, and where they go silent.
After launch, trust builds naturally and supervision erodes naturally. If the product wasn't designed to hold oversight stable, the agent ends up at an autonomy level nobody authorized. You can't measure what you didn't design.
Most change management programs treat agentic AI as an adoption problem. The diagnosis is wrong. Change management is being treated as adoption when the real work is supervision. That is a different skill, and most teams are not designing for it.
Your agent shipped. Your team moved on. Silent degradation is what happens next, and nobody is watching for it. The most-deployed clinical AI in American hospitals ran at half its advertised accuracy for six years before anyone audited it. Three disciplines before the next ship.
Your AI worked in the lab. Your users bypassed it in the real world. The gap isn’t the model, it’s the environment. Without the right incentives, workflows, and accountability, even the best-designed AI gets reduced to output generation. The system around it determines everything.
Agentic AI is not a smarter chatbot. It’s a system that can hold the full clinical picture across fragmented data, coordinate actions, and escalate when needed. But its value depends on data, interfaces, and accountability designed right from the start.
Twenty years of customer interviews, workshops, and journey maps. Then agentic AI arrived, and every framework I trusted turned out to share one assumption I had stopped noticing: that the human is always smarter than the tool. Here's what breaks when that stops being true.
I still read NEJM and BMJ cover to cover. Lately they are filled with elegant AI breakthroughs: models that promise to transform diagnosis, prediction, resource allocation. The deeper I build enterprise software, the more critical I get. A paper is a demo. A scaled deployment is a different problem.
For 20 years, I translated between customers and engineers. Now I architect decision systems with AI. But here's the catch: you're designing for two customers: humans and their AI agents. Design for one, forget the other, and it fails.
The most important AI benchmark you have never heard of measures one thing: how long can an AI system work on a task before it loses the thread? In healthcare, that matters more than single-answer accuracy. A sepsis protocol is a multi-hour trajectory, not a prompt. Attention span is the real gate.