Back to the buzzer room Question 1 of 10
Ten bets

Getting it right
is the easy half.

Ten questions with two answers each. You do not just pick one — you say how sure you are, in one tap. Being right earns you very little if you were only guessing, and being wrong costs you a great deal if you were certain. At the end there is a single number for that, and we take it apart into the three things it is made of.

One tap picks your answer and your confidence together.

ad slot

How sure you said, how often you were right

Each dot is one confidence level you used. Along the bottom is what you claimed; up the side is how it actually went. The diagonal is where an honest dot sits.

Your stated confidence against how often you were right.

Answer the ten and this fills in with your own dots.

One score, and the three things inside it

ad slot

Two people who never took the test

The ten questions

QuestionAnswerYou saidCost

How this is scored

  1. One tap is two answers. The six cells run from certain-left to certain-right. Tapping one records both which side you picked and how sure you were — 90%, 70% or 55%. There is no 50% cell on purpose: a real 50% is not an answer, it is a refusal to give one.
  2. The score is a Brier score. For each question, take the probability you gave your own answer, subtract 1 if you were right and 0 if you were wrong, and square it. Average those ten. 0 is perfect and 1 is perfectly wrong. Squaring is what makes confidence expensive: being wrong at 90% costs 0.810, being wrong at 55% costs 0.303 — those are the figures in the Cost column of the table above, so you can find them there.
  3. The three parts are not our invention. Any Brier score splits exactly into reliability − resolution + uncertainty — Murphy’s decomposition, 1973. It is an identity, not an approximation. The one liberty we take is display: three decimals cannot always hold an exact split, so when the column would come out a thousandth short we push that thousandth onto whichever figure was rounded hardest, rather than print a sum that does not add up.
  4. Reliability is the chart, squared. It is the vertical gap between each dot and the diagonal, squared and weighted by how many questions landed in that dot. Lower is better. It is the only one of the three that is purely about whether your confidence meant what it said.
  5. Resolution rewards you for separating the questions. It measures how far your three groups pulled apart from your overall hit rate. If you were right just as often when you said 90% as when you said 55%, your confidence carried no information and this term is zero, however well calibrated you were.
  6. Uncertainty is not about you at all. It is your hit rate times one minus your hit rate — the score a coin would get on a question set you find this hard. It is largest when you get exactly half right and shrinks to nothing at the extremes, so a low total score is easier to reach on questions you mostly know.
  7. Ten questions is far too few to measure a person. Each dot on the chart is made of a handful of answers, so one lucky guess moves it a long way. Treat the shape as an illustration of what calibration means, not as a reading of you. The research finding it points at — that most people’s 90% is nearer 70% — comes from studies with hundreds of items per person.
  8. The sides are balanced. Five answers are on the left and five on the right, so tapping the same end every time gets you five out of ten and tells you nothing. The questions were also picked to be ones where the fast answer is usually the wrong one, which is the house style here and does mean your hit rate will read low.

Elsewhere in the buzzer room