p(doom) calculator · submissions report · export of 1 Sep 2026

The expert gate closed, and the experts went quiet

202 anonymous submissions between 5 October 2025 and 1 September 2026, 38 of them since the previous report. Each one is the midpoint of the reader's estimated probability that AI causes a global catastrophe.

The unit is a submission, not a person, except where a signature says otherwise. Since 19 August every submission is signed by a per-browser key, so repeat rows from one browser can now be told apart from two people agreeing. Before that date nothing distinguishes them.

Changes are measured against the 18 August export (164 submissions). "Expert share" counts submissions that declared a level, in the window before the check went live and in the window since.

All submissions202 · Oct 2025 – Sep 2026

Still drifting upward, with August the busiest month yet

The level chooser went live on 9 March 2026 and the expert check on 19 August. Submissions before March predate both. Six later submissions carry no level: five took the sliders shortcut, or reached the page after it was removed, and one came through the retired standalone knowledge check. They are shown hollow rather than folded into the early cohort.

Every submission, coloured by quiz level

Dashed line: monthly median. Expert-quiz submissions that cleared the check carry a ring. Hover or Tab to a dot for its p10–p90 range.

Monthly median with interquartile band

Line: median of that month's submissions; band: 25th–75th percentile. September holds three submissions so far, so its point is a placeholder, not a trend.

Who says what: distribution by quiz level

Each dot is one submission; the tall tick marks the group median. The expert row is split into the 22 who declared themselves experts before the check existed and the 4 who have cleared it since. Hollow dots in the medium row are identical repeats from one signed browser.

Monthly summary

MonthnMedianP25P75

The belief chain: where the groups diverge

Group medians for each link of the chain, and the resulting p(doom). Quiz-takers rate every link near-certain; the pre-chooser cohort discounts each successive conditional. Final p(doom) multiplies each person's own chain, so it sits below the medians shown.

The widest-spread factor tracks the total most. That is arithmetic, not behaviour

Each panel: one factor's midpoint against the submission's own final p(doom). Read this as a sensitivity check, not a finding about people. Final p(doom) is the product of these three numbers, so each factor is being correlated with something it helped compute. What separates the panels is how much each factor varies across this set of submissions (SD 0.24, 0.27, 0.32), not where it sits in the chain; reordering the chain leaves every correlation unchanged.

The expert checksince 19 Aug 2026 · 34 submissions

Four asked for the expert quiz. Four cleared the bar. Nobody who fell short left a row.

The previous report's first recommendation was to gate the expert path behind the decoy check. That shipped on 19 August: clicking Expert quiz now opens the 30-term check, a score of 26 or more clears it, and a lower score offers the recommended quiz alongside a "continue to expert anyway" button. The verdict, the score and the exact terms ticked are stored with the submission. Two weeks of data is what exists to judge it by.

What became of everyone who clicked Expert quiz

A score is recorded exactly when the check ran, which is exactly when the expert card was clicked. What the table cannot hold is anyone who took the check, saw the verdict and closed the tab.

The check has not yet turned anyone away on the record. It has changed who shows up.All four who finished scored 28 or 29 of 30 and were waved through, so among finishers the check discriminates nothing. Where it appears to have acted is before the finish line: in the 70 submissions between the chooser going live and the check, 33% of those who declared a level picked expert. In the 34 since, 12% did, and medium took the difference. Whether that is the check deterring self-declaration, people abandoning it partway, or two weeks of a different audience cannot be separated from submission rows alone. It is consistent with the gate costing something, which was the point.

The level chooser, before and since the check

Completed submissions per path in each window. The standalone "decide for me" path and the sliders shortcut were retired when the check moved onto the expert card, so the two rows without a level since then came from a page loaded before the change, or from outside the page.

Term by term: what the four finishers ticked

Decoy-to-term pairings are an editorial reading of the term list; the AI-related flags come from the page. The one submission through the retired standalone path (April 2026) is shown separately in the tooltip: it found 9 real terms, tripped 6 decoys and was routed to beginner. Mary Shelley was replaced by Embedding on 2 September, after this export; the four finishers here were all shown her.

Verified experts against self-declared ones

Left: the share of options ticked on each expert-quiz question, for all 26 expert submissions. Right: how many of the 12 named organisations, commentators and creators each recognised, against the p(doom) they submitted. Filled dots cleared the check; hollow dots declared themselves experts before it existed.

The four verified experts submitted a median 30%. The 22 self-declared ones, 68%.Their four values are 87%, 50%, 10% and 0.4%, so this is a hint on four points, not a result. It runs in the same direction as the earlier observation that those least embedded in the field submit the higher numbers. The quiz's own questions still cannot tell anyone apart: 22 of 22 self-declared experts ticked all four self-improvement answers, and 3 of the 4 verified ones did too. Name recognition remains the only discriminating block, and it disagrees with the check on one person: the verified expert who submitted 87% recognised 1 of 12 names and neither research organisation, having scored 29 of 30 on the decoys six minutes after submitting a medium-quiz answer of 97% from the same browser.

The expert quiz's funding question is an opinion, not a recognition test, so it is left out of both panels.

Visitor identity

Six of the 34 recent rows are one browser pressing submit again

Each browser now holds a non-extractable P-256 key and signs a canonical string of its submission. The report verifies every signature against the public key the row carries, so "these rows are one person" is checked rather than asserted. What a key cannot show is that two different keys are two different people.

Signatures

Browsers that submitted more than once

RowsWithinLevelsp(doom)Reading

The earlier report guessed that identical payloads minutes apart were double-submits. The signatures confirm it.Three browsers produced six extra rows that are byte-for-byte copies of their first, all within 16 seconds. One of them submitted five times in six seconds. Those six rows are flagged throughout this report and every medium-quiz figure is given with and without them. The fourth repeat visitor is different in kind: a medium submission at 97% followed by an expert one at 87% six minutes later, which is one person answering two quizzes, and is kept. Nothing was replayed or tampered with: every signed row verifies, and every signed string agrees with the numbers stored beside it.

Beginner quiz33 submissions

Ten more beginners, and the film result changed sign

The beginner quiz asks which AI films the visitor has seen and which kinds of catastrophe they recognise. The previous report found neither predicted p(doom). With 33 submissions instead of 23, that holds, and the way it holds is instructive.

Watching AI films doesn't predict p(doom)

Left: how many of the 33 ticked each title. Right: those same people split by how many titles they ticked, against the p(doom) they submitted.

Sorted by share who ticked it. * = option added partway through, so its denominator is smaller.

The correlation flipped from −0.21 to +0.15 on ten new submissions.Group medians now run A 70%, B 62%, C 51%, which reads as a trend in the opposite direction to last time. A sign change on n = 33 is the clearest demonstration available that there is nothing here: r = +0.15 (ρ = +0.14) is noise, and the earlier −0.21 was too. Film exposure remains mass-market rather than AI-canon: Terminator and The Matrix at 29 of 33, Ex Machina at 8, Eagle Eye still ticked by nobody.

Risk awareness moves neither p(doom) nor stated uncertainty

Left: how many of the 33 recognised each of 12 risk types. Middle: those people grouped by how many they recognised, against p(doom). Right: risks recognised against the mean ± setting they put on the three sliders.

Still non-monotonic, still a null, and the one descriptive pattern survived.Group medians are A 65%, B 54%, C 71%: the middle group is lowest again. Pearson r = −0.30 against Spearman ρ = −0.15 says the linear figure is carried by a few extreme values. The tighter-uncertainty result the last report already dismissed sits at r = −0.29 and is not revisited. What did survive ten new submissions is that the least-aware group contains no skeptics: group C's minimum is 38%, while A and B both hold submissions under 5%.

Medium quiz39 submissions · 33 without repeats

The vulnerability checklist still tracks p(doom); the system-prompt cell that looked new was one browser

The medium quiz asks six questions: three about what the visitor knows or has lived through, three about what they believe. It nearly doubled in size since August and is now the largest cohort. Every figure here is given twice, on all 39 rows and on the 33 that remain once the six confirmed repeats are dropped.

How strongly each answer tracks the p(doom) submitted

Rank correlation per question. The filled mark uses all 39 rows; the hollow mark drops the six identical repeats. Right: the system-prompt question in full, with the repeated rows hollow.

Recognising more vulnerabilities is still the strongest association in the data, and it got stronger once the repeats were removed.ρ = +0.41 on all rows, +0.51 without the repeats (permutation p ≈ 0.004, n = 33), down from +0.66 on the August export and now on a cohort almost twice the size. The headline that would have been new is the system-prompt question: on all rows it reads −0.30, with the eight self-described experts at a median 28%, which looks like hands-on skill finally showing up. Five of those eight are one browser submitting the same 28% five times. Without them the cell holds 4 submissions at a median 39% and the correlation is −0.20 (p ≈ 0.27). It is a lead, and a smaller one than it first appeared.

The three opinion questions are shown for completeness and are not findings: a belief about whether governance can keep AI safe correlating with an estimate of catastrophe is close to tautological.

The vulnerability checklist

Left: which of the 16 entries the 39 medium submissions recognised. Right: the count each recognised against the p(doom) they submitted; hollow dots are the repeats.

Sorted by share who recognised it.

The entries recognised by nearly everyone are the ones the exhibits show AI already using.Zero-days (38 of 39), social engineering (36) and supply-chain attacks (30) are the techniques in the documented incidents: an OpenAI agent chaining zero-days and stolen credentials into remote code execution at Hugging Face, and Anthropic and OpenAI models attempting supply-chain and social-engineering attacks under test. Recognition drops for side-channel attacks (17), sleeper agents (20) and model stealing (22). A high count therefore mostly means knowing the long tail, and a reading consistent with the correlation is that knowing that more failure modes exist goes with a higher p(doom). It remains one association among six here and dozens across the report, with no correction for multiple comparisons.

No quiz5 submissions

Everyone who skipped the quiz landed on 12.5%, including two who arrived after the shortcut was removed

The sliders shortcut dropped the reader straight onto the three probability controls. Three used it in April. It was removed on 19 August, and two more rows without a level arrived on 21 August anyway, unsigned, from a page that either predates the change or was not the page at all.

What the data says

Eleven things this export supports

  • Submitted p(doom) has kept drifting upward. The median rose from 15% in the first three months (Oct–Dec 2025, n = 83) to 67% from June 2026 onward (n = 79). August 2026 alone holds 62 submissions at a median 69%, the busiest month since launch. Whether individuals got gloomier or the audience changed cannot be separated from this data.
  • The expert path is taken far less often since it opens with a check. Of submissions that declared a level, 33% chose expert before 19 August and 12% since, on 66 and 32 submissions. Everyone who finished the check cleared it comfortably (28–29 of 30). The instrument's effect, if any, is on who reaches the finish, which rows cannot show.
  • Verified experts submit lower numbers than self-declared ones, on four data points. Median 30% against 68%. Their four values span 0.4% to 87%. Treat it as direction, not size.
  • Six recent rows are confirmed double-submits. Signed identity shows three browsers producing byte-identical repeats within 16 seconds, one of them five times. The 202 records hold 175 distinct payloads overall, and the older, unsigned repeats remain unattributable. Every "n" here counts records; where it matters, the de-duplicated figure is quoted alongside.
  • The three levels' medians are close and their ordering is not stable. Beginner 68% (n = 33), medium 67% (n = 39, and unchanged without repeats), expert 59% (n = 26). The expert median moved from 68% to 59% on four verified submissions. All three still sit well above the pre-chooser cohort's 13%, and that comparison is confounded by time.
  • The vulnerability checklist is the one quiz answer that tracks p(doom). ρ = +0.51 on 33 de-duplicated medium submissions (p ≈ 0.004), the same direction as August on nearly double the cohort. Hands-on system-prompt skill leans negative (ρ = −0.20) but does not clear noise once one browser's five copies are removed.
  • Film exposure and risk awareness predict nothing among beginners, and the film correlation changed sign. r went from −0.21 to +0.15 when the cohort grew from 23 to 33. The one durable descriptive fact: nobody in the least risk-aware group has submitted below 38%.
  • P(powerful AI) is rated highest of the three links but is far from unanimous. Its median is 87% against 76% and 80% for the two conditionals, yet 51 of 202 submissions (25%) put it at 50% or below. The rising correlation along the chain (0.59, 0.74, 0.79) is arithmetic: it follows how much each factor varies (SD 0.24, 0.27, 0.32).
  • The extremes are used. 11 submissions put the midpoint at exactly 0% and 4 at 99% or above. Two of the four verified experts are among the lowest values in the entire table.
  • Stated uncertainty varies enormously. The p10–p90 band has a median width of 33 percentage points; the middle half runs from 15 to 39. This is a different quantity from the mean ± setting used in the beginner section, which is the radius set on each slider before clipping at 0 and 1.
  • Rows still arrive by a path that no longer exists. Two unsigned, level-less submissions on 21 August, both at 12.5% with the first slider nudged, came in two days after the sliders shortcut was removed and signing began. A cached page is the likely explanation; a script posting to the table directly is the other.
Directions forwardsix shipped, four open

What to change next, and what the previous list produced

The August report listed seven changes. Three shipped on 19 August and this report is the first to measure them. Three more, prompted by the findings above, shipped on 2 September, the day after this export, so the next report is the first that can measure those. Four stand.

  • Shipped
    Gate the expert path behind a decoy check
    Live since 19 August. Four finishers, four cleared, and the expert share of declared submissions fell from 33% to 12%. What it needs now is the thing it cannot record: how many people start the check and leave. A single row written when the check is displayed, or a counter, would turn "the share fell" into "the check turned N away".
  • Shipped
    Record what the knowledge check decided
    Score, recommended level, verdict and the exact terms ticked are stored. This is what made the term-by-term figure possible.
  • Shipped
    Attribute repeat submissions
    Signed identity landed alongside the check and immediately paid for itself: six rows that would have been "probably double-submits" are now certainly one browser.
  • Shipped
    Refuse to register the same prediction twice
    One browser submitted five identical rows in six seconds; two others submitted twice within 16 seconds. The button was already disabled while an insert was in flight, so these were completed submits repeated. Since 2 September the page remembers what was last registered, the slider settings, quiz answers and gate verdict, and keeps the button disabled while nothing has changed. Moving a slider or an answer re-enables it. The rows in this report predate the change.
  • Shipped
    Reconsider Mary Shelley as a scored term
    She was flagged AI-related and all four finishers left her unticked, which cost each of them a point. No other real term was missed by more than one, and no finisher tripped a single decoy. On 2 September she was replaced by Embedding, the vector representation an agent searches when it retrieves memories, with no decoy paired to it. The build script now scores each taker against the term list as it stood on their day, so the four finishers here keep their 28–29 of 30 and the new term's denominator counts only takers who were shown it.
  • Shipped
    Reject rows without a level, and version the page
    Two level-less, unsigned rows arrived after the only path that could produce them was removed. Since 2 September every submission carries page_version, the commit the page was built from, which pins the option list, term list and quiz it was answered on, and the insert policy can require a level, so a stale tab is refused with a reload prompt rather than stored. The version stamp also retires the option-set item below once enough rows carry it.
  • Open
    Seed flagged decoys into all three self-selected quizzes
    The check is now on the expert path, but beginner and medium still contain no option that can be wrong, and medium is now the largest cohort. The vulnerability checklist in particular would gain from two or three fabricated entries, since its count is the one answer that tracks p(doom).
  • Open
    Replace or repair the expert quiz's reasoning questions
    22 of 22 self-declared experts and 3 of 4 verified ones ticked every self-improvement answer. The check now does the discriminating that these questions were meant to do; the questions themselves still carry no information.
  • Open
    Stamp each submission with the option set it was shown
    Eagle Eye was added on 18 August and has been offered to 33 beginners since, so its denominator is now the full cohort and the correction is moot for it. The page version stamped since 2 September pins the option set exactly; what remains is teaching the build script to read the denominators from it instead of the hand-kept start dates.
  • Open
    Ask one identical question at every level
    Still the only way to compare cohorts on something they all answered, and the reason the verified-versus-self-declared gap above can be attributed to the check rather than to whoever happened to pick that card that fortnight.

These are directions the data supports, not a verdict on the design. Several rest on very small cells: four verified experts, five knowledge-check finishers, two stray rows. The first thing any of them buys is a larger, cleaner sample to test against, which the three 2 September changes exist to produce.