p(doom) submissions over time

164 anonymous submissions to the p(doom) calculator, 5 Oct 2025 – 18 Aug 2026. Each submission is the midpoint of the reader's estimated probability of AI-caused global catastrophe.

The unit here is a submission, not a person. The export carries no visitor identifier, so nothing distinguishes one person submitting twice from two people submitting once. Counts throughout are submissions; where a number would change materially after collapsing identical payloads, both are given.

Total submissions
Oct 2025 – Aug 2026
Median p(doom), Oct–Dec 2025
Median p(doom), Jun–Aug 2026
Median among expert-quiz submissions

All submissions

164 submissions · Oct 2025 – Aug 2026

Every submission, colored by quiz level

The level chooser went live 9 Mar 2026; submissions before that predate it entirely. The four later ones without a level took the sliders or knowledge-check path and are shown separately (hollow), not folded into the early cohort. Dashed line: monthly median. Hover or Tab to a dot for its p10–p90 range.

Monthly median with interquartile band

Line: median of that month's submissions; band: 25th–75th percentile. Hover for values and sample size.

Who says what: distribution by quiz level

Each dot is one submission; the tall tick marks the group median. The last two rows are submissions with no declared level: those made before the chooser existed, and the four later ones that used the sliders or the knowledge check.

Monthly summary table

MonthnMedianP25P75

The belief chain: where the groups diverge

Group medians for each link of the calculator's chain, and the resulting p(doom). Quiz-takers rate every link near-certain; the pre-chooser cohort discounts each successive conditional. (Final p(doom) multiplies each person's own chain, so it sits below the medians shown.)

The widest-spread factor tracks the total most — an arithmetic result, not a behavioural one

Each panel: one factor's midpoint against the submission's own final p(doom). Read this as a sensitivity check, not a finding about people. Final p(doom) is literally the product of these three numbers, so each factor is being correlated with something it helped compute; the association is built in rather than discovered. It is not guaranteed to be large — a factor that barely varies while the others swing widely can correlate near zero with the product — which is the point: what separates these three is how much each one varies across this particular set of submissions (SD 0.23, 0.28, 0.32), together with how they co-vary, and not where any of them sits in the chain. Multiplication is commutative, so reordering the chain leaves every correlation unchanged. What the panels show is which slider moved the answer most here.

“Decide the right level for me”

the knowledge-check path

Almost nobody takes the recommended knowledge check

The chooser offers a knowledge check — labelled as recommended, and listed first — alongside three self-selected levels and a sliders-only path. Counting every submission since the chooser went live on 9 Mar 2026. These are completed submissions, not visitors: the export carries no visitor identifier, no record of the chooser being displayed, and nothing about people who started and left — so this is the mix of paths that produced a submission, not take-up among everyone who saw the screen.

1 of 66 post-launch submissions came through it. Of the rest, 62 declared a level outright and 3 skipped to the sliders; among those 62, 31% went straight to expert — a claim nothing in those three quizzes ever checks. Whether the check is being seen and rejected or simply never noticed cannot be told apart from submission data alone.

The only decoys sit on the path nobody takes

Decoy-to-term pairings are an editorial reading of the term list; the AI-related flags themselves come from the page. Four decoys (Cherenkov radiation, the Magnus effect, Moiré patterns and Schrödinger's cat) are free-standing rather than keyword collisions.

Beginner quiz

23 submissions

Watching AI films doesn't predict p(doom)

The idea was that people who have seen a lot of AI fiction would have richer scenarios in mind, and land somewhere different from people who have seen almost none. Only the beginner quiz asks this, so the test covers 23 submissions. Left: how many of them ticked each title. Right: those same people split by how many titles they ticked, against the p(doom) they submitted.

Sorted by share who ticked it. * = option added partway through the period, so it was only offered to some submissions and its denominator is smaller.

The three groups are indistinguishable. Their medians land within 3 points of each other (A 68%, B 71%, C 70%) while each group internally spans 70+ points. Films seen vs p(doom) is r = −0.21 (ρ = −0.17) — noise at this sample size. Neither does uncertainty separate them: the median ± setting is 0.51 / 0.45 / 0.51, and what looks like a wider group A is three individuals, not a group trait.

Risk awareness moves neither p(doom) nor stated uncertainty

The idea was that recognising more kinds of catastrophe means a better-furnished sense of how things go wrong, and so a different p(doom). The beginner quiz asks which of 12 risk types the visitor recognises, so this is the same 23 submissions. Left: how many of them recognised each risk. Middle: those same people grouped by how many risks they recognised, against the p(doom) they submitted. Right: risks recognised against the mean ± setting they put on the three sliders.

No effect on p(doom); two patterns worth a second look. The group medians are non-monotonic (A 73%, B 53%, C 72%): the middle group is lowest, which is what a non-effect looks like. Pearson r = −0.26 against Spearman ρ = −0.09 confirms it — the linear number is carried by a few extreme values, not a trend. Two things are suggestive rather than established: the least-aware group contains no skeptics (group C's minimum is 48%, while A and B both hold sub-10% submissions), and recognising more risks goes with a smaller ± setting (r = −0.40). That second one survives dropping any single submission (−0.32 to −0.51) but a two-sided permutation test puts it at p ≈ 0.06 on n = 23, collapsing identical payloads takes it to −0.32, and using the true interval width instead of the ± setting takes it to −0.34. Applying both leaves −0.28 at p ≈ 0.22. Leave-one-out stability says a result isn't one person's doing; it says nothing about sampling error, measure choice, or the many comparisons made here. Read this one as a null. The group medians for width also look cleaner than the underlying numbers, since group A's widths are bimodal.

Medium quiz

20 submissions

System-prompt skill is flat against p(doom); the vulnerability checklist is the one that tracks it

The medium quiz asks six questions. Three are about what the visitor knows or has lived through; three ask what they believe. Left: how strongly each answer tracks the p(doom) they went on to submit, as a rank correlation. Right: the system-prompt question in full, since hands-on skill is the sharpest test of whether using AI heavily changes how dangerous you think it is. Below: the vulnerability checklist that turned out to matter most — which entries submissions recognised, and how the count tracks their p(doom).

AI-native  general security. The split is an editorial reading of the list.

The catalogue tracks the number; hands-on skill does not. System-prompt familiarity is the flattest question on the quiz — ρ = −0.04, having moved from +0.07 on the previous export, which is itself a demonstration of how little is pinned down at this size. Its group medians run 63% / 84% / 69% / 58% from "don't know what one is" to "expert", which is non-monotonic, i.e. what a non-effect looks like. Meanwhile recognising more vulnerabilities is the strongest association in this data (ρ = +0.66, permutation p ≈ 0.002; ρ = +0.67, p ≈ 0.003 after collapsing identical payloads; +0.60 to +0.76 under leave-one-out), and having actually seen a model go off the rails leans the other way (ρ = −0.18, down from −0.39 on the previous export — a sign this one was never solid): the five who answered "never" still sit at a median 88%, the highest group in the quiz. A reading consistent with this: knowing that failure modes exist goes with a higher p(doom), while hands-on contact with models does not. It is one association among the six tested here and dozens across this report, on 20 self-selected submissions, with no correction for multiple comparisons — a lead to test deliberately, not an established relationship.

The three opinion questions are shown for completeness but should not be read as findings — an opinion about whether governance can keep AI safe correlating with an estimate of catastrophe is close to tautological. Several cells are very small: two submissions call themselves system-prompt expert, and one has bypassed safety filters.

Expert quiz

19 submissions

7 of 19 self-declared experts know neither research organisation

The expert quiz mixes two kinds of question. Some can be answered by reasoning about what is possible; others only by having followed the field. Left: the share of options expert-quiz submissions ticked on each. Right: how many of the 12 named organisations, commentators and creators each recognised, against the p(doom) they submitted.

The quiz's own reasoning questions can't tell them apart. They can't tell anyone apart — near-everyone ticks every option, and 19/19 ticked all four self-improvement answers. The name-recognition questions can: the median expert recognised 5 of 12 names, 7 of 19 recognised neither research organisation, and 4 recognised two names or fewer. Those least embedded in the field submit the higher numbers — median 68% for the bottom group against 57% for those recognising five or more, on 4 and 10 submissions respectively. That gap narrowed when the newest submission landed, so read it as a hint, not a result.

The expert quiz's funding question is an opinion, not a recognition test, so it is left out of both blocks.

Sliders

3 submissions

Everyone who skipped the quiz landed on 12.5%

The sliders path drops the reader straight onto the three probability controls with no questions asked. Three submissions came through it, all in April 2026 — and all three left every midpoint at 0.5 while treating uncertainty completely differently.

What the data says

  • Submitted p(doom) has drifted sharply upward. The median rose from 15% in the first three months (Oct–Dec 2025, n=83) to 68% in the last three (Jun–Aug 2026, n=39). Whether individuals got gloomier or the audience changed can't be separated from this data.
  • The three levels land within three points of each other, and their ordering is not stable. Counting records: beginner 70% (n=23), medium 68% (n=20), expert 68% (n=19) — medium and expert now tie. Collapsing duplicated payloads moves medium to 67% and expert to 68%. Two submissions arriving on a single day took the expert median from 56% to 68%, which is the clearest possible evidence that no ranking between these levels is supportable at this sample size. What survives is that all three sit well above the pre-chooser cohort (median 15%, n=98) — and even that is confounded, since levels only exist in the later, higher period.
  • The extremes are used. 10 submissions put the midpoint at exactly 0% and 2 at ≈100%; both ends persist across the whole period.
  • Traffic is bursty. Oct 2025 (n=55, launch) and Aug 2026 (n=25) dominate; Jan 2026 had only 4 submissions, so single months with small n (Jan, Jul) swing easily.
  • P(powerful AI) is rated highest of the three links, but it is far from unanimous. Its median is 87% against 76% and 75% for the two conditionals — yet 38 of 164 submissions (23%) put it at 50% or below, so the high median describes the middle of the distribution, not a consensus. The often-quoted companion statistic — that correlation with final p(doom) rises along the chain (0.56, 0.74, 0.81) — is arithmetic, not behaviour: p(doom) is the product of these three numbers, and the ordering simply follows how much each varies between submissions (SD 0.23, 0.28, 0.32). Reordering the chain leaves it unchanged.
  • Tick counts cannot be compared across levels at all. Beginners are offered 48 checkboxes and experts 24, so the raw medians (26 vs 15.5) largely measure quiz length. Normalising by options offered (54% / 62% / 65%) removes that artefact but not a deeper one: the levels ask different kinds of question — recognising films, recognising vulnerabilities, and listing mechanisms that are nearly all true — and these have different natural tick rates regardless of who is answering. No cross-level engagement comparison is supportable from this data. Within the beginner cohort alone, having seen more AI films is weakly negatively related to p(doom) (r = −0.21, n = 23 — indistinguishable from noise).
  • Film exposure is mass-market, not AI-canon — which may be why it predicts nothing. The Matrix (21/23) and Terminator (20/23) are near-universal, but Ex Machina — the film most cited in alignment discourse — was ticked by 2 of 23, and Pantheon, Transcendence, Neuromancer and Murderbot by nobody at all. "Films seen" here largely measures general movie-watching.
  • Risk awareness does not move p(doom), and the apparent effect on confidence does not survive checking. Recognising more catastrophe types leaves p(doom) where it is. A tempting secondary result — that it goes with tighter uncertainty settings — measures −0.40 as originally computed, but falls to −0.28 (permutation p ≈ 0.22) once repeated payloads are collapsed and the true interval width is used instead of the ± setting. It is not a finding. The one durable observation is descriptive: nobody in the least-aware group submitted below 48%. Film exposure and risk awareness are essentially unrelated (r = +0.10).
  • Self-selection beats the guardrail. The knowledge check exists to place people on the right lane and is marked recommended — it produced 1 of 66 post-launch submissions. Of the 62 that declared a level, 31% picked expert, and by the quiz's own name-recognition standard several of those don't clear the bar. Submission data cannot say how many people saw the chooser and abandoned it.
  • Some submissions are byte-identical to others. Excluding timestamps, the 164 records contain 145 distinct payloads; the beginner, medium and expert cohorts hold 21, 19 and 18 distinct payloads against 23, 20 and 19 records. Most repeats arrive within minutes of each other (median gap 10 seconds), which looks like double-submits rather than two people agreeing exactly — but without a visitor id that cannot be established, and every "n" here counts records. Where it matters the deduplicated figure is quoted alongside.
  • Stated uncertainty varies enormously. The simulated p10–p90 band has a median width of 32 percentage points, but the middle half runs from 13 to 38 and the full range from 0 to 72 — so "wide" is the median, not the norm. Note this is a different quantity from the mean ± setting used in the beginner and medium sections. That is the average of the three spread values in the export, which is the radius a submission set around each midpoint — not the resulting interval width, because the interval is clipped at 0 and 1. Using the true width (upper − lower) instead moves the risk-awareness correlation from −0.40 to −0.34.

Directions forward

changes this data argues for

What to change, and what in this data argues for it

Each item names a change to the calculator and the finding above that motivates it. These are experiments and instrumentation fixes, not conclusions: this data can show where the design cannot currently answer a question, but it cannot show that a given change would produce better estimates.

  • Try gating the expert path behind a decoy checkdesign experiment
    Anyone can declare themselves an expert today, and 30% of self-selectors do. The expert quiz cannot contradict them: its three reasoning questions are ticked in full by nearly everyone (19/19 on self-improvement), and its name-recognition questions — the only ones that discriminate — carry no consequence. 7 of 19 recognised neither research organisation. Requiring a short decoy round before the expert quiz unlocks would make the claim cost something, and the mechanism already exists: the knowledge check's near-miss terms caught the one submission that met them. Treat this as an experiment to run, not a change the evidence already justifies — nothing here shows that name recognition or decoy performance measures calibration, or that gating would improve the estimates that come out the other side. It is worth trying because the current design cannot tell an expert from someone who says they are one, not because the data proves a gate would help.
  • Seed flagged decoys into all three self-selected quizzes
    Right now the only place in the calculator where an answer can be wrong is the knowledge check, which produced 1 of 66 post-launch submissions. The beginner, medium and expert quizzes carry no decoys, so nothing distinguishes a person who recognises PauseAI from one who ticks every box. Adding a handful of flagged decoys per level — a fabricated campaign organisation, a non-existent film, a plausible-sounding non-mechanism — makes over-claiming measurable for the 98% who never see the check.
  • Record what the knowledge check decided, not just that it ran
    Submissions store quiz_flow_id: "decide" but not the level the check then routed the visitor to, so there is no way to tell whether it ever redirects anyone — whether someone who would have picked expert was placed on medium. Storing the routed level and the score alongside the flow id turns the check into something you can evaluate rather than just offer.
  • Replace or repair the expert quiz's reasoning questions
    "Which of these are ways an AI could self-improve" is answered in full by every single submission, and the other two conceptual questions nearly so. A question everyone answers identically carries no information about who is answering it. Either add plausible-but-wrong mechanisms so the question can be got wrong, or drop it and spend the visitor's attention on questions that separate people.
  • Stamp each submission with the option set it was shown
    Options get added over time — Iron Man in June, Eagle Eye in August — and a raw count then understates them against everyone who was never offered the choice. This report reconstructs each option's real denominator from the page's git history, which works but is fragile and lives outside the data. Recording an option-set version with each submission makes the correction exact and automatic.
  • Ask one identical question at every level
    Every level has its own questions, so no two cohorts can be compared on anything they both answered. That is why the headline result — the expert cohort submits the lowest median — cannot be attributed to expertise rather than to whoever happens to pick that path. A single shared anchor question, asked identically to beginners and experts alike, would let the cohorts be compared directly instead of inferred.
  • Decide what the recommended path ischeap
    The knowledge check is listed first and marked recommended, and appears in 1 of 66 post-launch submissions. Either it is the right default and should be where readers land rather than something they must choose, or it is not and the label is misleading. Worth pairing with impression logging first: submission data cannot distinguish "seen and declined" from "never noticed".

These are directions the data supports, not a verdict on the design — several of the findings behind them rest on small cells (18 expert-quiz submissions, 1 knowledge-check submission), and the first thing any of these changes would buy is a larger, cleaner sample to test them against.