164 anonymous submissions to the p(doom) calculator, 5 Oct 2025 – 18 Aug 2026. Each submission is the midpoint of the reader's estimated probability of AI-caused global catastrophe.
The unit here is a submission, not a person. The export carries no visitor identifier, so nothing distinguishes one person submitting twice from two people submitting once. Counts throughout are submissions; where a number would change materially after collapsing identical payloads, both are given.
The level chooser went live 9 Mar 2026; submissions before that predate it entirely. The four later ones without a level took the sliders or knowledge-check path and are shown separately (hollow), not folded into the early cohort. Dashed line: monthly median. Hover or Tab to a dot for its p10–p90 range.
Line: median of that month's submissions; band: 25th–75th percentile. Hover for values and sample size.
Each dot is one submission; the tall tick marks the group median. The last two rows are submissions with no declared level: those made before the chooser existed, and the four later ones that used the sliders or the knowledge check.
| Month | n | Median | P25 | P75 |
|---|
Group medians for each link of the calculator's chain, and the resulting p(doom). Quiz-takers rate every link near-certain; the pre-chooser cohort discounts each successive conditional. (Final p(doom) multiplies each person's own chain, so it sits below the medians shown.)
Each panel: one factor's midpoint against the submission's own final p(doom). Read this as a sensitivity check, not a finding about people. Final p(doom) is literally the product of these three numbers, so each factor is being correlated with something it helped compute; the association is built in rather than discovered. It is not guaranteed to be large — a factor that barely varies while the others swing widely can correlate near zero with the product — which is the point: what separates these three is how much each one varies across this particular set of submissions (SD 0.23, 0.28, 0.32), together with how they co-vary, and not where any of them sits in the chain. Multiplication is commutative, so reordering the chain leaves every correlation unchanged. What the panels show is which slider moved the answer most here.
The chooser offers a knowledge check — labelled as recommended, and listed first — alongside three self-selected levels and a sliders-only path. Counting every submission since the chooser went live on 9 Mar 2026. These are completed submissions, not visitors: the export carries no visitor identifier, no record of the chooser being displayed, and nothing about people who started and left — so this is the mix of paths that produced a submission, not take-up among everyone who saw the screen.
1 of 66 post-launch submissions came through it. Of the rest, 62 declared a level outright and 3 skipped to the sliders; among those 62, 31% went straight to expert — a claim nothing in those three quizzes ever checks. Whether the check is being seen and rejected or simply never noticed cannot be told apart from submission data alone.
Decoy-to-term pairings are an editorial reading of the term list; the AI-related flags themselves come from the page. Four decoys (Cherenkov radiation, the Magnus effect, Moiré patterns and Schrödinger's cat) are free-standing rather than keyword collisions.
The idea was that people who have seen a lot of AI fiction would have richer scenarios in mind, and land somewhere different from people who have seen almost none. Only the beginner quiz asks this, so the test covers 23 submissions. Left: how many of them ticked each title. Right: those same people split by how many titles they ticked, against the p(doom) they submitted.
Sorted by share who ticked it. * = option added partway through the period, so it was only offered to some submissions and its denominator is smaller.
The three groups are indistinguishable. Their medians land within 3 points of each other (A 68%, B 71%, C 70%) while each group internally spans 70+ points. Films seen vs p(doom) is r = −0.21 (ρ = −0.17) — noise at this sample size. Neither does uncertainty separate them: the median ± setting is 0.51 / 0.45 / 0.51, and what looks like a wider group A is three individuals, not a group trait.
The idea was that recognising more kinds of catastrophe means a better-furnished sense of how things go wrong, and so a different p(doom). The beginner quiz asks which of 12 risk types the visitor recognises, so this is the same 23 submissions. Left: how many of them recognised each risk. Middle: those same people grouped by how many risks they recognised, against the p(doom) they submitted. Right: risks recognised against the mean ± setting they put on the three sliders.
No effect on p(doom); two patterns worth a second look. The group medians are non-monotonic (A 73%, B 53%, C 72%): the middle group is lowest, which is what a non-effect looks like. Pearson r = −0.26 against Spearman ρ = −0.09 confirms it — the linear number is carried by a few extreme values, not a trend. Two things are suggestive rather than established: the least-aware group contains no skeptics (group C's minimum is 48%, while A and B both hold sub-10% submissions), and recognising more risks goes with a smaller ± setting (r = −0.40). That second one survives dropping any single submission (−0.32 to −0.51) but a two-sided permutation test puts it at p ≈ 0.06 on n = 23, collapsing identical payloads takes it to −0.32, and using the true interval width instead of the ± setting takes it to −0.34. Applying both leaves −0.28 at p ≈ 0.22. Leave-one-out stability says a result isn't one person's doing; it says nothing about sampling error, measure choice, or the many comparisons made here. Read this one as a null. The group medians for width also look cleaner than the underlying numbers, since group A's widths are bimodal.
The medium quiz asks six questions. Three are about what the visitor knows or has lived through; three ask what they believe. Left: how strongly each answer tracks the p(doom) they went on to submit, as a rank correlation. Right: the system-prompt question in full, since hands-on skill is the sharpest test of whether using AI heavily changes how dangerous you think it is. Below: the vulnerability checklist that turned out to matter most — which entries submissions recognised, and how the count tracks their p(doom).
AI-native general security. The split is an editorial reading of the list.
The catalogue tracks the number; hands-on skill does not. System-prompt familiarity is the flattest question on the quiz — ρ = −0.04, having moved from +0.07 on the previous export, which is itself a demonstration of how little is pinned down at this size. Its group medians run 63% / 84% / 69% / 58% from "don't know what one is" to "expert", which is non-monotonic, i.e. what a non-effect looks like. Meanwhile recognising more vulnerabilities is the strongest association in this data (ρ = +0.66, permutation p ≈ 0.002; ρ = +0.67, p ≈ 0.003 after collapsing identical payloads; +0.60 to +0.76 under leave-one-out), and having actually seen a model go off the rails leans the other way (ρ = −0.18, down from −0.39 on the previous export — a sign this one was never solid): the five who answered "never" still sit at a median 88%, the highest group in the quiz. A reading consistent with this: knowing that failure modes exist goes with a higher p(doom), while hands-on contact with models does not. It is one association among the six tested here and dozens across this report, on 20 self-selected submissions, with no correction for multiple comparisons — a lead to test deliberately, not an established relationship.
The three opinion questions are shown for completeness but should not be read as findings — an opinion about whether governance can keep AI safe correlating with an estimate of catastrophe is close to tautological. Several cells are very small: two submissions call themselves system-prompt expert, and one has bypassed safety filters.
The expert quiz mixes two kinds of question. Some can be answered by reasoning about what is possible; others only by having followed the field. Left: the share of options expert-quiz submissions ticked on each. Right: how many of the 12 named organisations, commentators and creators each recognised, against the p(doom) they submitted.
The quiz's own reasoning questions can't tell them apart. They can't tell anyone apart — near-everyone ticks every option, and 19/19 ticked all four self-improvement answers. The name-recognition questions can: the median expert recognised 5 of 12 names, 7 of 19 recognised neither research organisation, and 4 recognised two names or fewer. Those least embedded in the field submit the higher numbers — median 68% for the bottom group against 57% for those recognising five or more, on 4 and 10 submissions respectively. That gap narrowed when the newest submission landed, so read it as a hint, not a result.
The expert quiz's funding question is an opinion, not a recognition test, so it is left out of both blocks.
The sliders path drops the reader straight onto the three probability controls with no questions asked. Three submissions came through it, all in April 2026 — and all three left every midpoint at 0.5 while treating uncertainty completely differently.
spread values in the export, which is the radius a submission set around each midpoint — not the resulting interval width, because the interval is clipped at 0 and 1. Using the true width (upper − lower) instead moves the risk-awareness correlation from −0.40 to −0.34.Each item names a change to the calculator and the finding above that motivates it. These are experiments and instrumentation fixes, not conclusions: this data can show where the design cannot currently answer a question, but it cannot show that a given change would produce better estimates.
quiz_flow_id: "decide" but not the level the check then routed the visitor to, so there is no way to tell whether it ever redirects anyone — whether someone who would have picked expert was placed on medium. Storing the routed level and the score alongside the flow id turns the check into something you can evaluate rather than just offer.These are directions the data supports, not a verdict on the design — several of the findings behind them rest on small cells (18 expert-quiz submissions, 1 knowledge-check submission), and the first thing any of these changes would buy is a larger, cleaner sample to test them against.