# Experts cut the last link of the chain

**Date:** 2026-09-13
**Scope:** An observation from the 10 September report, two readings of it, and a
caution against building the quiz around either. Follows
[splitting the reports into kept and moved](2026-09-12-split-untouched-moved.md),
which is what made the observation visible.
**Outcome:** Among visitors who set their own number, verified experts are the
most certain that powerful AI arrives, as sure as everyone else that it behaves
dangerously, and half as sure that dangerous behaviour becomes a global
catastrophe. The expert quiz never mentions an outcome, so rehearsal could produce
this. It could also be a considered view held by people who know the field. One
recognition question would tell them apart. Nothing else should change on the
strength of nineteen rows.

---

## 1. The observation

"Verified" is the site's label for an expert-quiz taker who cleared the thirty-term
check. The check is a sanity check on vocabulary with decoys, not an examination of
expertise: it separates people who know what the words mean from people who do not,
and nothing more. The cohort should be read as "informed enough to pass a
vocabulary check", which is a low bar cleared by nineteen people, not as a panel of
experts. What the site means by an expert, hands-on experience of machine learning,
is in [definitions](definitions.md#expert).

Median chain among rows where the visitor moved at least one slider, so the
numbers are theirs and not the quiz's proposal:

| Group | P(powerful AI) | P(dangerous \| powerful) | P(catastrophe \| dangerous) | p(doom) | n |
|---|---|---|---|---|---|
| Beginner | 0.85 | 0.83 | 0.87 | 48% | 25 |
| Medium | 0.82 | 0.75 | 0.86 | 34% | 38 |
| Expert, verified | 0.95 | 0.80 | 0.50 | 31% | 19 |
| Expert, self-declared | 0.88 | 0.87 | 0.92 | 44% | 17 |
| Before the chooser | 0.80 | 0.60 | 0.50 | 15% | 98 |

Relative to the proposal the quiz set, expert movers raised the first link by a
median 8 points, left the second alone, and lowered the third. Of 36 expert movers,
the lowest of the three sliders is the third for 17, the second for 13, the first
for 6. The pooled "quiz-takers who moved" line in the report dips at the second
link only because medium movers are the largest group and that is the link they cut.

The pre-chooser cohort, who set every slider by hand with no quiz, discounted all
three links roughly equally. Verified experts discount one.

## 2. Two readings

**Rehearsal.** Look at what each quiz puts in front of the visitor before the
sliders appear. The beginner quiz lists catastrophes: supervolcano, pandemic,
nuclear war, extinction from AI, s-risks. That is the third link, rehearsed as a
list, and beginner movers leave it at 87%. The expert quiz asks about continuous
learning, self-improvement, self-replication, then names of organisations. All of
it is mechanism for the first two links; none of it names an outcome. The exhibits,
which do cover the third link, appear only after submission. On this reading the
expert cuts the link the quiz never made them think about, and would cut it whether
or not they had a view.

**Judgement.** Someone who follows the field has seen dangerous behaviour
demonstrated, in the exhibits and in the news, and has not seen it escalate. They
have a working model of why: detection, isolation, incident response, the fact
that the one documented intrusion by a frontier model was reconstructed action by
action and contained. The third link is where the evidence is thinnest and where a
person with hands-on experience has the most reason to discount a chain of
hypotheticals. On this reading the 50% is a belief, and a defensible one.

The two readings predict the same numbers. Nothing in the current data separates
them.

## 3. Against overfitting

The temptation is to treat the rehearsal reading as a flaw and redesign the expert
quiz to correct it. That should be resisted, for three reasons.

- **The cohort is nineteen people.** Direction, not size. A median of 0.50 on
  nineteen rows can move to 0.65 or 0.35 on the next nineteen without anything
  having changed.
- **These are not uninformed visitors.** They passed a thirty-term check with
  decoys that turned two people away. That proves only that they know the
  vocabulary, but not everyone knows what a Von Neumann probe is, and the people
  who do may well have hands-on experience of the systems the first two links
  describe. Their number on the third link deserves to be treated
  as a view until shown otherwise, not as an artefact to be corrected.
- **A quiz tuned to move one number is a quiz that measures itself.** The site
  has just spent a fortnight learning that lesson from the preset. The chain, the
  quiz vocabulary and the check are the instrument; they should change when the
  instrument is wrong, not when a result is surprising.

## 4. The one cheap test

If the question is worth answering, one change answers it without redesigning
anything: add a recognition question about catastrophe pathways to the expert quiz,
of the same kind the beginner quiz already has. If the verified experts' third link
rises afterwards, it was rehearsal. If it stays near 50%, it is a belief, and the
site has learned something about how informed people weigh the chain.

The question should be scored and stored like the others, so that the report can
show the third link before and after, by level, against the stored proposal. That
is the whole experiment. It costs one question and a few months of rows.

The open item "ask one identical question at every level" does the same job with a
control group and is the better version if it is ever built.

What the cut means in the chain's own terms, and which link visitors put their
zero on, is taken up in [where the zero goes](2026-09-13-where-the-zero-goes.md).

## 5. What not to conclude yet

- Not that experts are right about the third link. Nothing here bears on that.
- Not that the beginner quiz is priming its takers upward on the third link. It
  may be; the same test in reverse would say.
- Not that informed people in general discount the third link. In the 2024 Expert
  Survey on Progress in AI, researchers who had thought more about the social
  impacts of AI gave somewhat higher risk, not lower
  ([the calculator against ESPAI 2024](2026-09-15-espai-2024.md)). That axis is
  attention to the risk arguments, not vocabulary, and the survey's researchers
  qualify by publication venue, so they need not know the vocabulary this check tests.
  The two do not contradict each other; the finding here is about people who pass
  this check, and no wider.
- Not that the check selects for optimists. The self-declared experts of the
  pre-check era are at 0.92 on the third link and 44% overall, so the check did
  change who is in the cohort; whether it changed what they think, or just who
  shows up, is the same question as the last report's, and still open.
