There's a kind of person who can say 'I don't know' the way other people report the weather. No flinch, no apology, no sense that ground has been given up. I've noticed that when one of these people finally does commit to an answer, I believe them, and I don't think that's sentimentality. Their confidence carries information precisely because it isn't their default setting.
The usual name for this is epistemic humility, which makes it sound like a virtue, something to list next to patience. I've come to think of it as a skill instead, and a fairly technical one: keeping your confidence in proportion to your evidence, and knowing, from the inside, the difference between what you know and what you're guessing.
Calibration, the unglamorous core
A calibrated person who says they're ninety percent sure is right about nine times in ten. Put that way it sounds almost trivially checkable, and it is, which is the point. On Agonora we measure it with factual questions where you attach a confidence level to each answer, and the general pattern in our data matches what calibration research has been finding for decades: most people run overconfident, and not subtly. What strikes me isn't the overconfidence itself. It's that there's no internal alarm at the moment you cross from knowledge into guesswork. The sentence feels exactly the same in your mouth either way.
People usually reach for the Dunning-Kruger effect at this point, and it only half fits. The pop version, incompetent people believing they're brilliant, isn't quite what Kruger and Dunning found in 1999. Their least skilled participants rated themselves above average, not at the top, and later critics have argued that part of the effect is a statistical artifact, since nearly everyone's self-estimate drifts toward 'somewhat above average.' The honest reading is milder and more useful: accurate self-assessment isn't standard equipment for anyone. It has to be built, usually by keeping score.
The best evidence that it can be built comes from forecasting research. Philip Tetlock spent about two decades scoring expert predictions on politics and economics, and the results, published in 2005, were grim in a specific way: the average expert barely beat simple extrapolation, and the more famous the expert, the worse their accuracy tended to be. Confidence sells; being scoreable doesn't. But the forecasting tournaments he later ran turned up a small group of otherwise ordinary people who consistently beat everyone, reportedly including intelligence analysts with access to classified information. What set these superforecasters apart wasn't credentials or raw intelligence so much as habits: probabilities in granular steps, frequent small updates, a flat, almost boring willingness to say they'd been off. Habits can be copied, which is the encouraging part.
What it isn't
It isn't relativism. A calibrated person believes plenty of things firmly; the grip just matches the evidence, tight where it's strong, loose where it's thin. It isn't the performative 'oh, I'm probably wrong' either, which is usually fishing for reassurance rather than reporting a probability.
And it isn't indecision, though I suspect that fear is what keeps people away. Knowing which of your beliefs are load-bearing usually makes acting easier, not harder.
The machines made it urgent
Large language models produce fluent, authoritative answers to anything, and some of those answers are inventions delivered in exactly the same even tone as the truth. In 2023 a New York lawyer filed a federal court brief citing cases that ChatGPT had simply made up; by his own account he had asked the tool whether the cases were real, and it assured him they were. He was sanctioned, and the story traveled the world as a joke about lazy lawyers. I find it harder to laugh than I'd like, because the underlying mistake wasn't really laziness. It was borrowing a confidence signal from a system that doesn't have one.
The models are getting better at expressing uncertainty, and I don't want to overstate this; the previous paragraph may age badly. But even a model that hedges perfectly leaves your problem intact, because you still have to decide how much of its confidence to adopt as your own. That decision is calibration again, one level up, and nobody can make it for you.
The part that doesn't pay
I should concede the weakest point in all of this. Calibration pays off slowly, through trust and through disasters that never happen, which are invisible. Overconfidence pays off fast. The confident pitch gets funded, the confident take spreads, the hedged version gets scrolled past. If someone argues that in politics or fundraising the overconfident simply win, I don't have a knockdown reply. The claim I'd actually defend is narrower: the people you return to for judgment, the ones whose 'I'm sure' can settle a question in a room, got that standing by also being the ones who say 'I don't know,' and I don't think you can have one without the other.
Keeping score
If you want to get better at this, the only method I've seen hold up is embarrassingly simple: write down predictions with numbers attached, check them later, notice where the misses cluster. Forecasting platforms exist for it, Agonora's calibration questions are our version of it, and honestly a spreadsheet works fine. The first weeks are mildly humiliating, which I'd guess is most of why people don't do it. But that discomfort is the internal alarm being installed, the one that eventually interrupts you mid-sentence to say you've drifted from knowing into guessing. I haven't found another way to get one.