top of page
JUDGEMENT

Judgement in the AI Age

Judgement shows in how a leader arrives at a decision. How much they ask before they commit. How they read a situation. How close their conviction sits to what the evidence supports. Whether anyone still pushes back. All of this is visible in the process of deciding, long before any outcome is in.

​

Measure the quality of your judgement only by the outcome of decisions and you learn too late whether it carried far enough. By then the course is long set. Observe the behaviour instead, and it can be read earlier, while corrections can still change something. AI sharpens this challenge. It answers in seconds, in a tone that admits no doubt, and the window for asking narrows. Most of the time it also acts on what we ask of it straight away, with no critical questions and no feedback. What is missing is the exchange between equals, in which someone also questions critically or shares observations and feedback.

"No one owns their judgement. You only keep it in shape."

WHY JUDGEMENT MATTERS IN THE AI AGE

Experience is no safeguard against overconfidence

With the years, confidence in one's judgement grows; accuracy does not keep pace. In the Duke CFO series by Ben-David, Graham and Harvey, executives' 80% confidence intervals held in only about 36% of cases. Confidence was high, accuracy markedly lower, and from the inside the gap was impossible to feel.


The bias does not work the same in every direction. Overconfidence drives the exploration of the new and at the same time undermines the steady investment; its damage can be contained by internal control and external oversight.

The data come from large US corporations and their financial context. Whether they carry over to Swiss SMEs and family firms is plausible but not established, and that limit stays visible here, much as the manufacturing limit of the Census study does next door.

The finding has been around for a long time; what is new is only what AI does with it. It delivers self-assured recommendations in seconds and leaves barely any time to test or to stop them, precisely where the most depends on it.

Chart: executives' stated confidence compared with their actual hit rate; confidence sits well above accuracy

Executives’ stated confidence compared with their actual hit rate; confidence sits well above accuracy.

EVIDENCE

What research shows about forming judgement

More can be said about effective judgement than its public treatment suggests. Three studies, each of a different weight, fit together.

The causal core comes from a pre-registered experiment by Weidmann, Xu and Deming at Harvard. Leaders who ask more questions and share speaking time achieve measurably better team results, in human teams as in AI teams. The correlation between their effect and a standardised leadership test is 0.81. Gender, age and education explain none of it. The finding comes from a laboratory setting, with AI agents as stand-ins. That an external counterpart brings the same questioning quality follows from the result. It is not tested there.

Dialogue works

What improves decisions is justification-bound counter-argument from outside. Time delay and pure reflection remain without measurable effect in pre-registered experiments, while a duty to justify and adversarial questioning work.

Source: Aczel et al. (Judgment and Decision Making, 2023; d ≈ 0.40 to 0.54). The effect sizes of debiasing interventions are heterogeneous and, on average, small to moderate (meta-analysis of 54 RCTs, Nature Human Behaviour 2025, g ≈ 0.26); individual trainings reach d > 0.80 (Morewedge et al. 2015). The reliable signal is the direction, not a peak value.

The correction follows the same pattern. A duty to justify and adversarial questioning work in controlled experiments, while popular reflection methods stay ineffective. The RCT evidence on coaching, too, points to the working relationship, the external person, rather than the method. The effect is moderate, the publication bias openly documented.

What sharpens judgement comes from outside, and it questions.

Standardised effect sizes for what sharpens judgement; introspective reflection shows no measurable effect, while external, questioning interventions do.

Standardised effect sizes from controlled studies; introspective reflection shows no measurable effect, while external, questioning interventions do.

THE LISTENING GRID

Judgement quality, before the outcome confirms it

Judgement shows in behaviour. So it has to be readable in conversation, while the decision is still forming. That is what the listening grid is for: four traces along which an external instance hears the quality of judgement, without the leader having to examine themselves.

An external instance reads the quality of judgement along four traces.

Calibration

The distance between stated confidence and the evidence. Whoever consistently over- or underestimates probabilities rarely notices it themselves. In conversation, the gap between tone and evidence becomes audible.

Self-Other Discrepancy

How far self-perception and outside effect lie apart. The gap is what the conversation is about.

Question-to-Assertion Ratio

Does a leader enter a decision with questions or with settled assertions? This is where the causally strongest lever from the Weidmann study lies.

Corrective Density

Does the system hold any counter-argument at all, a duty to justify, a devil's advocate, a pre-mortem? Or does the sparring remain the only corrective.

Chart: the four traces of judgement quality, each with a guiding question, namely calibration, self-other discrepancy, question-to-assertion ratio and corrective density

The four traces of judgement quality, each with a guiding question: calibration, self-other discrepancy, question-to-assertion ratio, corrective density.

No one fills in these four traces beforehand as a questionnaire. They are the grid with which a sparring or coaching partner listens in conversation. There is no score at the end. What emerges is a picture of where the judgement holds and where it tips.

Three of the four traces connect directly to the research. The question-to-assertion ratio follows from the Weidmann finding, calibration from the miscalibration research, the self-other discrepancy from the self-other agreement literature. Corrective density is a diagnostic logic from transformind's practice and goes beyond the body of studies. It is marked as such. That openness is the reason to trust the grid.

THE METHODOLOGICAL GROUNDING

The scientific foundation of the listening grid

The four traces are not chosen at random. They rest on three traditions of thought that interlock: from how judgement deceives itself, through how it is corrected, to why the correction has to come from outside.

Behavioural Decision Research · Calibration

From Kahneman and Tversky to Ben-David, Graham and Harvey, the evidence shows that confidence and accuracy diverge systematically, and that the gap stays invisible from within. This is the trace of calibration.

Debiasing · Counter-Argument

From Gary Klein's pre-mortem to the work of Lovallo and Kahneman, the evidence shows that structured counter-argument and a duty to justify correct what introspection cannot. From this follow the question-to-assertion ratio and corrective density.

Systems-Theoretical Observation · The Outside View

In Luhmann's logic of observation, a system sees itself only up to a point; what lies in its own blind spot is visible to an observer outside it. This is where the self-other discrepancy is rooted, and with it the whole mechanism: the vantage point has to come from outside.


Each of the three traditions answers a different question: why judgement deceives itself, what corrects it, and why the correction has to come from outside. The listening grid reads the outward face of this frame.

APPLICATION IN LEADERSHIP PRACTICE

From the listening grid to decision-making

What follows is a hypothesis from transformind's practice. It is not established in the sense set out above.

When a leader enters decisions with fixed conclusions and corrective density is low, the constraint sits in the missing counter-argument. More information does not help. A voice that reliably pushes back does.

When calibration drifts, so that confidence stays high while the hit rate falls, an appeal for more humility changes nothing. What counts is a documented comparison of estimate against outcome over time, and a counterpart who insists on it.

This external check on judgement comes at two speeds. As Sparring it is available on demand and as often as needed, in the concrete moment of decision, with no questionnaire and no preparation. As Executive Coaching the same check works across a developmental arc of several months, tied to a goal. Both install the same corrective, at a different cadence.

Judgement shows in deciding under uncertainty

Judgement is not a fixed gift. It is a quality of deciding that surfaces under pressure and under the pace of AI, in how much gets asked and whether anyone still pushes back. The sharpest corrective comes from outside. It is not acquired once and then owned. It stays only as long as someone tests it.

At the level of the organisation, this personal axis of judgement has its counterpart in leadership-system maturity, one of the six dimensions of Ambiflow. Anyone looking for the organisational side of the same pattern will find it in the neighbouring topics Adaptive Organisation and Ambidexterity.

FREQUENTLY ASKED

Judgement, in questions

NEXT STEP

Judgement does not test itself.

If you are facing a decision that calls for it, let us talk.

bottom of page