The Quiet Erosion
AI raises your confidence whether it's right or wrong. Two preprints from MIT and Wharton show it also degrades the skill you need to catch it when it fails. Aviation solved this problem decades ago. Medicine and software haven't.
AI raises your confidence whether it's right or wrong. Two preprints from MIT and Wharton show it also degrades the skill you need to catch it when it fails. Aviation solved this problem decades ago. Medicine and software haven't.
A new BMJ paper asks why humans are still in the loop now that AI outperforms physicians on reasoning tasks. The answer is a framework. The problem is that framework assumes a physician whose independent competence AI is quietly eroding.
We are running an experiment on human oversight without a control group, without outcome tracking, and without a policy framework designed for the result. The current generation of experts may be the last one capable of catching what AI gets wrong.
Most AI clinical studies measured the wrong thing. They constrained the reasoning mechanism and then evaluated what was left. The Brodeur Science paper finally tests the model the way medicine actually works. The results are hard to dismiss.
Utah just became the first US state to let an AI autonomously renew prescriptions. The legal debate is real. The design question nobody is asking is more important: who designed the rungs on the autonomy ladder before the climb began?
Amazon owns the patient at the moment of health decision intent. OpenEvidence owns the physician at the moment of prescribing intent. Both are monetized by the same pharmaceutical industry. The prescription is the handshake between them.
AI was trained on what physicians write. Not on how they reason. The note comes hours after the decision, shaped by billing codes and fatigue. What actually saved the patient was never recorded.
Most PMs raising agentic AI concerns are on the wrong side of the room. The question is not whether to speak up. It is what to put on the table. Four moves that trade gut-feeling objections for artifacts, unit economics, and the question that aligns the whole room.
Most CEOs are optimizing for when the agent ships. The question that matters is what the customer experiences during the rollback. Five things the current agentic AI plan is missing, and the four questions the CEO should be asking before any agent ships.
Microsoft just analyzed 500,000 Copilot health conversations. The paper is honest: they observed questions but not outcomes. The physician in me cannot stop thinking about it. In medicine we audit what happened to the patient, not just what we considered. Why should AI be different?
A patient record can be comprehensive and clinically wrong at the same time. The AI reasons from what it receives. The failure is upstream, in the data, not in the model. That is the hallucination type nobody is naming.
Healthcare regulation still assumes a clinician is watching the AI before anything happens. That assumption worked when AI meant a score on a screen. It breaks down when AI runs continuously in a patient's pocket, shaping decisions no one reviews.