A fluent answer can feel like a confident answer. Those are different things.
Most AI systems are built to return something: a prediction, a recommendation, a ranked list, or the next word in a sentence. That is useful when the task is routine and the system has been tested in conditions like the ones it will face. It is risky when the system is outside those conditions, when the stakes are high, or when a plausible answer can hide uncertainty.
The better question is not whether an AI system can answer. It is whether it can recognize the circumstances in which it should pause.
“I don't know” is a design choice
In machine learning, there is a long-standing idea called selective classification or a reject option. Instead of forcing a prediction for every input, a system can decline to answer some cases and send them for review. That tradeoff is not magical: the system needs a way to estimate uncertainty, a threshold for acting, and a real process for handling the cases it rejects. But it changes the goal from “answer everything” to “answer the cases we can support.”
For a chatbot, the equivalent might be asking a clarifying question, pointing to a source, refusing to make a high-stakes call, or escalating to a qualified person. For a medical image system, it might mean flagging an unfamiliar scan for specialist review. For an automated factory, it might mean stopping a workflow when the signals no longer resemble conditions it has been tested on.
That is not a sign the system failed. It can be a sign that its limits were made visible.
What uncertainty can and cannot tell us
A model's confidence score is not a guarantee of correctness. A system can be confidently wrong, especially when its training data did not represent a new situation well. It can also be underconfident on a task where it is useful.
Researchers therefore study several related questions: whether a model is calibrated, whether it recognizes inputs unlike its training data, and whether it can abstain in cases where the expected cost of an error is high. A recent survey describes this family of approaches as uncertainty estimation with a “reject option.” Read the survey
The important practical point is modest: a score alone does not create safety. Someone must decide what the score means, where the threshold sits, and what happens next.
The human handoff has to be real
“Human in the loop” is often used as a reassurance. It only helps when the person has authority, time, context, and a clear reason to question the system.
NIST's AI Risk Management Framework says human roles and responsibilities in decision-making and oversight need to be clearly defined. It also notes that AI systems may defer to a human expert, serve as an additional opinion, or operate autonomously depending on the use case. Read NIST's human-AI interaction guidance
That makes the handoff a workflow question, not a staffing label. A useful design specifies:
- what the system is allowed to decide on its own;
- what signals trigger a pause or escalation;
- who receives the case and what information they see;
- whether that person can override the system;
- and how the organization learns from overrides and mistakes.
Without those details, a human reviewer can become a rubber stamp or a convenient place to assign blame after a failure.
A promising example: teaching systems to notice surprise
Georgia Tech researchers recently described Mutual Information Surprise, a framework meant to help autonomous systems tell the difference between a merely unusual event and one that should change their understanding. The work is early-stage, but its aim is sensible: an agent, robot, or automated factory should be able to recognize when its current model of a situation is inadequate and reassess. Read Georgia Tech's report
That is a long way from proving a system will behave safely in the real world. The researchers have not shown a finished safety mechanism, and every setting has different failure modes. Still, the direction is more useful than treating confidence as the same thing as competence.
Match the guardrail to the stakes
Not every AI tool needs the same kind of stop button. A video-compression model can be monitored differently from a system that influences a medical referral, a hiring decision, or an electricity-grid operation. NIST explicitly makes that distinction: some uses may not need human oversight, while others do. Read the AI RMF
The hopeful case for AI is not that it always acts alone. It is that it can make people more capable when its boundaries are clear. A model that answers routine questions quickly, shows why it is uncertain, and reliably hands difficult cases to a person can be more valuable than one that sounds certain about everything.
The next time a company says its AI is ready for the real world, a useful follow-up is simple: what does it do when it does not know?
Sources
- NIST AI Risk Management Framework: Human-AI interaction · Official NIST guidance on differentiated human roles, oversight and context loss in AI systems.
- NIST AI RMF 1.0 · Primary framework; supports the distinction between use cases that do and do not require human oversight.
- Survey on uncertainty estimation and reject options · Research survey used for the description of selective classification and abstention; it is not treated as deployment evidence.
- Georgia Tech: Can a Machine or AI Agent Be Surprised? · Institutional report on the research-stage Mutual Information Surprise framework and its stated purpose.