JUDGEMENT
Judgement in the AI Age
Judgement shows in how a leader arrives at a decision. How much they ask before they commit. How they read a situation. How close their conviction sits to what the evidence supports. Whether anyone still pushes back. All of this is visible in the process of deciding, long before any outcome is in.
​
Measure the quality of your judgement only by the outcome of decisions and you learn too late whether it carried far enough. By then the course is long set. Observe the behaviour instead, and it can be read earlier, while corrections can still change something. AI sharpens this challenge. It answers in seconds, in a tone that admits no doubt, and the window for asking narrows. Most of the time it also acts on what we ask of it straight away, with no critical questions and no feedback. What is missing is the exchange between equals, in which someone also questions critically or shares observations and feedback.
"No one owns their judgement. You only keep it in shape."
WHY JUDGEMENT MATTERS IN THE AI AGE
Experience is no safeguard against overconfidence
With the years, confidence in one's judgement grows; accuracy does not keep pace. In the Duke CFO series by Ben-David, Graham and Harvey, executives' 80% confidence intervals held in only about 36% of cases. Confidence was high, accuracy markedly lower, and from the inside the gap was impossible to feel.
The bias does not work the same in every direction. Overconfidence drives the exploration of the new and at the same time undermines the steady investment; its damage can be contained by internal control and external oversight.
The data come from large US corporations and their financial context. Whether they carry over to Swiss SMEs and family firms is plausible but not established, and that limit stays visible here, much as the manufacturing limit of the Census study does next door.
The finding has been around for a long time; what is new is only what AI does with it. It delivers self-assured recommendations in seconds and leaves barely any time to test or to stop them, precisely where the most depends on it.

Executives’ stated confidence compared with their actual hit rate; confidence sits well above accuracy.
EVIDENCE
What research shows about forming judgement
More can be said about effective judgement than its public treatment suggests. Three studies, each of a different weight, fit together.
The causal core comes from a pre-registered experiment by Weidmann, Xu and Deming at Harvard. Leaders who ask more questions and share speaking time achieve measurably better team results, in human teams as in AI teams. The correlation between their effect and a standardised leadership test is 0.81. Gender, age and education explain none of it. The finding comes from a laboratory setting, with AI agents as stand-ins. That an external counterpart brings the same questioning quality follows from the result. It is not tested there.
Dialogue works
What improves decisions is justification-bound counter-argument from outside. Time delay and pure reflection remain without measurable effect in pre-registered experiments, while a duty to justify and adversarial questioning work.
Source: Aczel et al. (Judgment and Decision Making, 2023; d ≈ 0.40 to 0.54). The effect sizes of debiasing interventions are heterogeneous and, on average, small to moderate (meta-analysis of 54 RCTs, Nature Human Behaviour 2025, g ≈ 0.26); individual trainings reach d > 0.80 (Morewedge et al. 2015). The reliable signal is the direction, not a peak value.
The correction follows the same pattern. A duty to justify and adversarial questioning work in controlled experiments, while popular reflection methods stay ineffective. The RCT evidence on coaching, too, points to the working relationship, the external person, rather than the method. The effect is moderate, the publication bias openly documented.
What sharpens judgement comes from outside, and it questions.

Standardised effect sizes from controlled studies; introspective reflection shows no measurable effect, while external, questioning interventions do.
THE LISTENING GRID
Judgement quality, before the outcome confirms it
Judgement shows in behaviour. So it has to be readable in conversation, while the decision is still forming. That is what the listening grid is for: four traces along which an external instance hears the quality of judgement, without the leader having to examine themselves.
An external instance reads the quality of judgement along four traces.
Calibration
The distance between stated confidence and the evidence. Whoever consistently over- or underestimates probabilities rarely notices it themselves. In conversation, the gap between tone and evidence becomes audible.
Self-Other Discrepancy
How far self-perception and outside effect lie apart. The gap is what the conversation is about.
Question-to-Assertion Ratio
Does a leader enter a decision with questions or with settled assertions? This is where the causally strongest lever from the Weidmann study lies.
Corrective Density
Does the system hold any counter-argument at all, a duty to justify, a devil's advocate, a pre-mortem? Or does the sparring remain the only corrective.

The four traces of judgement quality, each with a guiding question: calibration, self-other discrepancy, question-to-assertion ratio, corrective density.
No one fills in these four traces beforehand as a questionnaire. They are the grid with which a sparring or coaching partner listens in conversation. There is no score at the end. What emerges is a picture of where the judgement holds and where it tips.
Three of the four traces connect directly to the research. The question-to-assertion ratio follows from the Weidmann finding, calibration from the miscalibration research, the self-other discrepancy from the self-other agreement literature. Corrective density is a diagnostic logic from transformind's practice and goes beyond the body of studies. It is marked as such. That openness is the reason to trust the grid.
THE METHODOLOGICAL GROUNDING
The scientific foundation of the listening grid
The four traces are not chosen at random. They rest on three traditions of thought that interlock: from how judgement deceives itself, through how it is corrected, to why the correction has to come from outside.
Behavioural Decision Research · Calibration
From Kahneman and Tversky to Ben-David, Graham and Harvey, the evidence shows that confidence and accuracy diverge systematically, and that the gap stays invisible from within. This is the trace of calibration.
Debiasing · Counter-Argument
From Gary Klein's pre-mortem to the work of Lovallo and Kahneman, the evidence shows that structured counter-argument and a duty to justify correct what introspection cannot. From this follow the question-to-assertion ratio and corrective density.
Systems-Theoretical Observation · The Outside View
In Luhmann's logic of observation, a system sees itself only up to a point; what lies in its own blind spot is visible to an observer outside it. This is where the self-other discrepancy is rooted, and with it the whole mechanism: the vantage point has to come from outside.
Each of the three traditions answers a different question: why judgement deceives itself, what corrects it, and why the correction has to come from outside. The listening grid reads the outward face of this frame.
APPLICATION IN LEADERSHIP PRACTICE
From the listening grid to decision-making
What follows is a hypothesis from transformind's practice. It is not established in the sense set out above.
When a leader enters decisions with fixed conclusions and corrective density is low, the constraint sits in the missing counter-argument. More information does not help. A voice that reliably pushes back does.
When calibration drifts, so that confidence stays high while the hit rate falls, an appeal for more humility changes nothing. What counts is a documented comparison of estimate against outcome over time, and a counterpart who insists on it.
This external check on judgement comes at two speeds. As Sparring it is available on demand and as often as needed, in the concrete moment of decision, with no questionnaire and no preparation. As Executive Coaching the same check works across a developmental arc of several months, tied to a goal. Both install the same corrective, at a different cadence.
Judgement shows in deciding under uncertainty
Judgement is not a fixed gift. It is a quality of deciding that surfaces under pressure and under the pace of AI, in how much gets asked and whether anyone still pushes back. The sharpest corrective comes from outside. It is not acquired once and then owned. It stays only as long as someone tests it.
At the level of the organisation, this personal axis of judgement has its counterpart in leadership-system maturity, one of the six dimensions of Ambiflow. Anyone looking for the organisational side of the same pattern will find it in the neighbouring topics Adaptive Organisation and Ambidexterity.
FREQUENTLY ASKED
Judgement, in questions
What is judgement in the AI age?
Judgement is not a fixed gift. It is a quality of deciding: how much a leader asks before committing, how close their stated confidence sits to the evidence, and whether a voice exists in their system that pushes back. It shows in behaviour and eludes the self-image. In the AI age it is tested least where the machine delivers its self-assured answers fastest.
How do you recognise good judgement in leaders?
By four traces that can be read in conversation as the decision forms: calibration, how close stated confidence sits to the evidence; the self-other discrepancy between self-perception and the impression a leader makes; the question-to-assertion ratio, whether a leader enters a decision with questions or with fixed conclusions; and corrective density, whether the system holds any counter-argument at all. Three of the four traces map directly onto the research; the fourth is a diagnostic logic from practice.
How can judgement be improved?
Pure self-reflection produces no measurable effect in controlled experiments, and appeals for more humility change little. What works comes from outside and it questions: a counterpart who holds stated confidence against the evidence, insists on justification, and reliably pushes back. At transformind this external check on judgement is installed by Sparring in the concrete moment of decision, and by Executive Coaching across a developmental arc of several months.
What distinguishes decisiveness from decision quality?
Decisiveness describes how quickly and firmly someone commits. Decision quality describes whether the judgement behind it holds. The two come apart. A leader who decides firmly but miscalibrated looks strong in judgement and is not. What counts is the fit between stated confidence and the evidence, tested from outside, before the outcome confirms it.