Five Hundred Thousand Questions. What Did the Answers Look Like?
Microsoft just analyzed 500,000 Copilot health conversations. The paper is honest: they observed questions but not outcomes. The physician in me cannot stop thinking about it. In medicine we audit what happened to the patient, not just what we considered. Why should AI be different?