The Exercise That Teaches What a Model Is
Ten runs of one boring question about fruit teach more about AI than any slide deck. Probabilistic is not a definition; it is something you watch happen to your own keyboard.
I run the same exercise in every AI session I teach, and it still works every time.
I ask everyone in the room to open whatever model they have in front of them and type one prompt: in 20 words or less, what is the difference between apples and oranges. Then I ask them to run it ten times. Not once, not three times. Ten.
Thirty seconds later the room gets quiet in a specific way. Ten answers come back. Every one of them is reasonable. One leads with taste, another with botany, another with color, another with how each fruit is peeled. None of them are identical. A few people check their screens twice, certain they mistyped something. They did not. The model simply answered the same question ten different, defensible ways, and nothing in the interface told them that was coming.
That thirty seconds is the most useful thing I teach. Not because the answers are interesting, they are not, apples and oranges is a deliberately boring question. It is useful because of what happens in the room while people are looking at ten screens at once. "Probabilistic" stops being a word from a slide and becomes something they just watched happen to their own keyboard. You cannot unsee it once you have seen it. And you cannot really teach it any other way. I have tried explaining probabilistic output with diagrams, with the dice analogy, with the weather forecast analogy. All of it lands as information. The exercise lands as experience, and experience is what changes how a person actually behaves the next time a model answers something that matters.
This is the gap my last article pointed at: judgment about a probabilistic system is built by living with one, not by reading about one. Most AI training I have sat through, and most I have built, teaches people to operate a tool. Click here, prompt like this, avoid these words. That is operating instructions for a calculator. A calculator gives you the same answer every time you press the same buttons. The tool in front of your product team does not, and no amount of operating instructions will teach that difference. Only watching it happen will. Fluency with AI is not knowing which buttons to press. It is knowing what kind of system you are talking to, and knowing what your own job becomes once you are standing next to it instead of in front of it alone.
Once that shift happens in a room, three things tend to follow in a predictable order, and I have started teaching them as three pillars, because they build on each other in the order people actually discover them.
Probabilistic as experience, not definition. Once you have watched ten different answers come out of one prompt, a few habits stop being optional. You stop trusting a single run of anything that matters. You start asking what an eval actually checks before you believe a benchmark number, because a benchmark is one sample from a distribution, and a distribution is exactly what that exercise just showed you. And you retire the sentence "it worked yesterday" as evidence of anything. It worked yesterday, on that input, on that sample. The apples and oranges exercise proves that sentence wrong in thirty seconds, for free, before a single dollar has been spent on an incident.
The autonomy ladder. Nearly every practical AI decision I have made in the last two years has actually been one decision wearing different clothes: where on the spectrum from suggestion to full delegation does this particular task belong right now. A model that drafts a summary for a human to edit sits low on that ladder. A model that files the summary and moves to the next task without anyone reading it sits at the top. Most teams pick a rung once, at launch, and leave it there out of habit or nerves. The actual skill is moving a task up the ladder deliberately as trust is earned through evidence, not vibes, and moving it back down the moment the evidence turns. The ladder is not a one-time design decision. It is a dial you keep your hand on.
The agent replaces the person at the screen, not the screen. This is the one that takes longest to land, and it is the one I think matters most. An agent does not take over a piece of software. It takes over the seat of the person who used to operate that software, and that person's job does not disappear, it splits into two jobs that used to be one. Some of their time is still spent doing the work directly. The rest is now spent supervising the agent that does the work in their place, catching what it gets wrong, knowing when to step back in. Supervision of a probabilistic system is a real, learnable skill, and almost nobody is teaching it, because most AI curricula still assume the human is the one typing, not the one watching.
I think about this through medicine more than I think people expect. Teaching hospitals taught me something before I ever touched a product roadmap: explaining something clearly to the person who has to act on it next is not a soft skill bolted onto the real job, it is the job. A resident who cannot explain a finding to the nurse who will act on it overnight has not finished the work, no matter how correct the finding was. Supervising an agent is the same discipline pointed at a different kind of colleague. You are not done when the model produces an answer. You are done when the person relying on that answer, human or agent, knows what to trust in it and what to check.
None of this requires a platform change or a new tool. It requires thirty seconds and one boring question about fruit, and a willingness to actually look at what ten honest runs of the same prompt hand you back. I put the fuller version of this argument into a book, "The Agentic AI Practitioner: Keeping the Judgment the Machine Cannot Hold," out this week, because the exercise is a door and the book is what is on the other side of it. But you do not need the book to run the exercise tomorrow morning. You need a prompt, a room, and ten runs.
Try it before your next AI meeting. Ask your team the same twenty-word question ten times and read the answers out loud together. Watch what happens in the room. That is the whole lesson.