p(doom) calculator · submissions report · export of 1 Oct 2026

More than half of the site's busiest day took the expert quiz, and the day's median was 21%.

365 anonymous submissions between 5 October 2025 and 1 October 2026. The calculator asks each visitor for three conditional probabilities and multiplies them: that powerful AI arrives, that it then behaves dangerously, and that the dangerous behaviour then becomes a global catastrophe. The product is that visitor's p(doom). On 1 October the site had its busiest day, 36 submissions against a previous record of 19, and 33 of them arrived in two and a half hours that evening. Nineteen of the 36 chose the expert quiz and 18 passed the vocabulary check that guards it. The day's median was 21%, against 49% for September. The three numbers behind it say where the low figure comes from: those 18 put the first link at a median 88% and the last at 39%. That is the shape the last report found in the readers who pass the check, a danger that arrives and is as likely as not to be contained, and this crowd goes further: for them it is contained more often than not.

One thing has to be read before any figure below. A visitor who takes a quiz is not handed neutral sliders: the page computes a starting p(doom) from their answers and sets all three sliders to it. 92 of 255 quiz-takers (identical repeat rows aside) registered that proposed number without moving a midpoint. Those rows are the page's arithmetic, so wherever submissions are drawn one by one, the strips and the proposed-start chart, they are the round dots and the others are arrows, and the figures that matter are given for the other rows separately. Two words carry the split: kept means the three midpoints are exactly where the quiz put them, and moved means at least one is not. Kept does not mean nothing was touched: the uncertainty band is a separate control, and most kept rows changed it. Moved does not mean free of the proposal either; a visitor who drags a slider from 72% to 60% may have got there only because 72% was in front of them.

The unit is a submission, not a person, except where a signature says otherwise. Since 19 August the page signs each submission with a per-browser key, and 194 of the 197 rows since then carry one, so repeat rows from one browser can be told apart from two people agreeing. Before that date nothing distinguishes them.

Each card names its own baseline: the submission count is against the 19 September export (300); the medians set the single day of 1 October against the whole of September, with August beside them; the verified experts are set against the self-declared ones; the expert share sets the window since the check (19 August) against the window before it, counting submissions that declared a level; the repeat count sets the rows since the 2 September fix against the fortnight before it; the kept share is over every quiz submission, identical repeats aside.

All submissions365 · Oct 2025 – Oct 2026

September closed at 49% on 130 submissions. Then one evening brought 33 more, and the day's median was 21%

The level chooser went live on 9 March 2026 and the expert check on 19 August, so the 98 submissions from before 9 March answered no questions at all. Six later rows carry no quiz either: five took a sliders shortcut that has since been removed, one came through a retired knowledge-check path, and all six sit at or near the page defaults. They are left out of the timeline and the group strips; the monthly figures and the factor panels count every row. Nothing without a level has arrived since 2 September, when the page started refusing such rows.

Every submission since 1 August, by quiz level

228 submissions arrived in these two months, and the 226 with a quiz level are drawn, so they get the full width; the chart below covers every month. The column of dots at the right-hand edge is 1 October. Hollow dots are readers who took the expert quiz without a pass: those who simply called themselves experts before any check existed, and two who have since fallen short of the check and continued anyway. Filled dark ones are verified: they passed a 30-term vocabulary check before the expert quiz would open. The line is a running median, the median of each submission and the fourteen before it, with identical repeat rows left out, and it is dashed across gaps of more than two days between submissions. Hover or tap a dot for its p10–p90 range; with the chart focused, the arrow keys step through the dots. The two rows in this window that arrived with no quiz behind them are left out, as all six such rows are from the strips below.

Monthly median with interquartile band

Line: median of that month's submissions; band: 25th–75th percentile. September is complete. October's point is hollow and joined by a dashed line because it is a single day, 1 October, and one day from one burst is not a month.

September's median was 49% on 130 submissions, down from 69% in August.The first three months of the site, before any quiz existed, produced a median of 15%. The quiz arrived in March, the monthly median reached a peak of 69% in August, and September, the busiest month in the series with more than a third of every submission ever made, came back down. Every level fell with it: beginner 72% to 48%, medium 69% to 63%, expert 59% to 32%. The fall is not the page's proposed number changing: among visitors who set their own (repeat rows excluded), the median went from 50% to 29%, and the share who kept the proposal barely moved, 39% to 38%. Why the visitors changed is not something a table of submissions can say. The author's impression of how the site's reception turned around August is in a note; it is an impression, not a measurement.

Then came 1 October: 36 submissions, 33 of them in two and a half hours, at a median of 21%.33 of the day's 36 rows arrived between 20:11 and 22:38 UTC, 22 of them in the first 49 minutes, which is what one link posted in one place looks like. The export was taken at 23:12, so the burst may not have been over. Its median is 21%, and 17% among the 28 who set their own number; only 8 of the 36 kept the proposal. 13 of the 36 registered less than 10%. By level the day reads beginner 46% (n = 6), medium 19% (n = 11) and expert 11% (n = 19), so it is not only that the experts are lower: the medium quiz, which has the highest proposals on the site, had its lowest monthly median before this at 33%, in July on four rows. What the rows say about this crowd is that it knows the vocabulary: 20 took the check, 18 passed, and those 18 recognised a median 2 of the 12 organisations, commentators and creators named in the expert quiz, against a median 4 for the readers verified before 20 September. Where the crowd came from is not in the table, because the site still does not record where a visitor arrived from.

Who says what: distribution by quiz level

One mark per submission; the tall tick is the group median. An arrow pointing left is a visitor who moved the page's proposed number down, an arrow pointing right one who moved it up, and a round dot either kept the proposal or, in the bottom row, never had one. The expert rows are split three ways: the 22 who declared themselves experts before the check existed, the 60 who have passed it since, and the two who fell short of it and continued to the expert quiz anyway. Two marks do not make a median, so that row has no tick.

Monthly summary

MonthnMedianP25P75

2026-10 is one day, 1 October.

The belief chain: where the groups diverge

Group medians for each link of the chain, and the resulting p(doom). The beginner and medium lines sit high on every link because a kept row has all three factors at the cube root of one proposal. The expert line now bends at the last link: fewer than a quarter of expert rows are kept, and the readers who pass the check discount that link. The dashed line is quiz-takers who moved a slider, and it discounts both conditionals, as the pre-chooser cohort does. Final p(doom) multiplies each person's own chain, so it sits below the medians shown.

The widest-spread factor tracks the total most. That is arithmetic, not behaviour

Each panel: one factor's midpoint against the submission's own final p(doom). Read this as a sensitivity check, not a finding about people. Final p(doom) is the product of these three numbers, so each factor is being correlated with something it helped compute. What separates the panels is how much each factor varies across this set of submissions, not where it sits in the chain; reordering the chain leaves every correlation unchanged.

The proposed start261 quiz submissions

Thirty-six per cent of quiz-takers registered the number the page proposed

After a quiz the page does not hand the visitor neutral sliders. It proposes a p(doom) from the answers, an opinion question picks a band and the share of boxes ticked moves the number up within it, and sets all three factors to the cube root of the proposal. It is tempting to read the quiz as a survey taken alongside the sliders. It is not: it is an input to them, and a visitor who answers it and presses submit has registered the page's arithmetic rather than their own. Every quiz submission is therefore split from here on into kept, where no midpoint was moved from the proposal, and moved.

What the page proposed against what was registered

One dot per quiz submission. Horizontal: the p(doom) the quiz set the sliders to, recomputed from the stored answers with the page's own rules. Vertical: what was registered. A round dot on the diagonal never had a midpoint moved; off the diagonal the mark is an arrow pointing the way the number went, the same encoding the strips use further down. Identical repeat rows are not drawn.

Kept and moved, by level

LevelnKeptMedian, keptMedian, moved

Since 21 September the page says, above the sliders, that the quiz proposes a starting point. The kept share did not visibly change.The sentence reads "A short quiz proposes a starting point from what you already know", and 60 quiz rows have arrived under it. 20 of them kept the proposal, 33%, against 37 of the 103 rows in the nineteen days before it, 2 to 20 September, 36%. The two halves of the window since differ from each other far more than either differs from the window before: 12 of 24 kept in the last ten days of September, 8 of 36 on 1 October. That is the audience changing, and a before-and-after on 60 rows cannot see one sentence through it.

Kept means the three midpoints are exactly where the quiz put them, and nothing more. The uncertainty band is a separate control: of the 41 kept rows that stored the proposed band, 40 registered a different one. Identical repeat rows from one browser are left out of every kept and moved count, which is why the totals here are 255 and not 261. Since 12 September the page stores the proposal with each submission, and 105 rows carry it. For the quiz rows before that, the proposal is recomputed from the stored answers with the page's own tables, and the build checks itself two ways: every kept row sits at exactly the per-factor value the page would have set, and on the 105 rows that have both, the stored and the recomputed proposal agree to within a tenth of a point. Four rows have three equal midpoints at a value no proposal produces, a visitor's own doing, and are counted as moved.

The expert checksince 19 Aug 2026 · 197 submissions

Sixty-nine took the check. Sixty cleared it. Two of the nine who fell short went on to the expert quiz anyway.

Clicking Expert quiz opens the 30-term check, a score of 26 or more clears it, and a lower score offers the recommended quiz alongside a "continue to expert anyway" button. The verdict, the score and the exact terms ticked are stored with the submission. 69 people have now finished it, nearly twice the 36 of the last report. It admits most of the people who ask for it, 60 of 69. Of the nine it has turned away, seven took the quiz it recommended and two pressed the button, which no recorded submission had done before 27 September.

What became of everyone who clicked Expert quiz

A score is recorded exactly when the check ran, which is exactly when the expert card was clicked. What the table cannot hold is anyone who took the check, saw the verdict and closed the tab.

Seven of the nine who fell short took the recommended quiz. Two continued to the expert quiz anyway, the first recorded uses of that button.The seven scored 19 to 25 on the score the page stored. Six were offered medium and took it; the one who scored 19 was offered beginner and took that. Their submissions span 33% to 99%, so falling short of a vocabulary bar says nothing about the number someone then registers. The two who continued scored 24 and 25, two points and one point under the bar, on 27 September and 1 October; one kept the expert quiz's proposal at 27% and the other registered 0.4%. Everyone else scored 26 to 30 and went through, at a median score of 29. The expert share of submissions that declared a level is 32% since the check against 33% before it, on 195 and 66 submissions. At the last report it was 25%, and that report read the fall as medium absorbing the experts. The reading does not survive this export: the difference is 1 October, when 19 of 36 chose the expert quiz, and without that day the share is 27%. Medium is still the largest level, 41% of everything since 19 August. Whether the check deters people from clicking the expert card at all is not something completed submissions can show: a visitor who took the check and closed the tab leaves no row. Nine of 69 is a thin floor for a gate, and two of the nine walked past it.

The level chooser, before and since the check

Completed submissions per path in each window. In the first window a visitor could still skip the quizzes and go straight to the sliders, or take a standalone knowledge check; both were retired on 19 August, when the check moved onto the expert card, and since then every path on the page runs through a quiz. Two rows from 21 August arrived with no level anyway. They are unsigned, which is what a tab loaded before the 19 August deploy would send. They are left out of the right-hand panel, and since 2 September the database refuses any row without a level.

Term by term: what the 69 finishers ticked

Decoy-to-term pairings are an editorial reading of the term list; the AI-related flags come from the page. Scores here are recomputed from the ticked terms against the list as it stood on the taker's day; Mary Shelley was replaced by Embedding on 2 September, so the four earlier finishers were shown her and the sixty-five later ones were shown Embedding; which list a row saw is resolved from its page version, and the recomputed score agrees with the score the page stored on every one of the 69. The one submission through the retired standalone path (July 2026) is shown separately in the tooltip.

Verified experts against self-declared ones

Left: the share of options ticked on each expert-quiz question, for all 84 expert submissions. Right: how many of the 12 named organisations, commentators and creators each recognised, against the p(doom) they submitted. Dark dots are verified experts, grey ones self-declared, and the two groups sit to either side of each column; the two hollow dots are the readers who fell short of the check and continued.

The 60 who passed the check submitted a median 31%. The 22 who called themselves experts before the check existed, 68%.Among visitors who set their own number the gap is smaller, 25% against 44%, nineteen points. The 37-point pooled gap includes the kept rows on both sides, five self-declared experts at a median 77% and thirteen verified ones at 69%; leaving them out changes who is being compared as well as what, so the two gaps are two different comparisons rather than a decomposition into proposal and belief. At the last report the two groups recognised the same number of names in the field, a median 4 of 12 against 4.5. They no longer do: the verified median is now 3, because the 26 readers who have passed the check since 20 September recognise a median 2. The other difference has held: all 22 self-declared experts ticked every option on the self-improvement question, against 34 of the 60 who passed the check. Ticking everything is consistent with a list that is half-recognised, though a checklist nobody can fail cannot tell that from thoroughness. The two groups are also separated in time, since the check only exists from 19 August, and the site's overall median fell over the same period, so some part of this is the audience changing rather than the check selecting. Nearly a third of the verified group arrived on one day.

The expert quiz's funding question is an opinion, not a recognition test, so it is left out of both panels. It appears in full in the expert section below.

Beginner quiz76 submissions

Every beginner answer, and the p(doom) that went with it

The beginner quiz asks four questions: which AI films the visitor has seen, which kinds of catastrophe they recognise, which AI capabilities they know about, and a gut reaction to an AI misbehaving. For each question, the left panel counts how many ticked each answer and the right panel splits the respondents into thirds by how many they ticked, with the p(doom) each third submitted. First, the same split across all three checklists at once.

Ticked most and fewest, across all three beginner checklists

Every beginner's ticks on the film, catastrophe and capability lists added together, out of 50 possible, and the cohort cut into thirds on that total. Left: one mark per submission, the tall tick is the group median; an arrow points the way the visitor moved the quiz's proposed number, a dot kept it. Right: what each third did with that proposed number.

Pooled, the thirds sit at 58%, 59% and 50%. Among the 42 who moved a slider, the third who ticked most registered 32%, the middle 34% and the third who ticked fewest 29%.The pooled figures are mostly the page's proposal, which rises with the number of boxes ticked; 34 of 76 beginners kept it. At the last report the movers' ranking inverted the pooled one, with the third who ticked fewest highest at 43%. Fourteen more beginners later that inversion has gone and the three thirds are level: on this total, how many boxes a beginner ticks says nothing about the number they set for themselves. The catastrophe list is the one place a gradient remains. The third who ticked most and set their own number registered a median 24% (n = 22), the middle third 44% (n = 10) and the third who ticked fewest 50% (n = 10). These are cohorts of 10 to 22 people, so this is a direction and not a measurement. Ticks on a checklist are a count of recognised names, not a measure of knowledge.

Medium quiz101 submissions · 95 without repeats

Every medium answer: the vulnerability signal was the preset

The medium quiz asks six questions: three about what the visitor knows or has lived through, three about what they believe. It is the largest of the three cohorts, and its kept rows carry the highest proposed numbers on the site. The correlation figure is given three ways: all 101 rows, the 95 without the six identical repeats, and the 56 of those where the visitor moved a slider. Everywhere else the repeats are drawn but left out of the kept and moved counts.

How strongly each answer tracks the p(doom) submitted

Rank correlation per question. The filled mark uses all 101 rows; the hollow mark drops the six identical repeats; the square uses only the 56 rows where the visitor moved a slider.

On all rows the vulnerability checklist tracks p(doom) at ρ = +0.25. On the 56 rows where the visitor moved a slider, it is +0.12.That is the page's own formula being measured, not a fact about people: the share of vulnerabilities recognised is one of the inputs that moves the proposed number up, and 39 of the 95 medium rows without repeats kept the proposal. The pooled figure was +0.66 in the August report and +0.29 at the last one. What survives on the moved rows is the three opinion questions, competition at +0.46, governance at +0.35 and speed at +0.33, which pick the band and are near restatements of the sliders. System prompts sits at −0.07, and "seen AI go off the rails" crosses zero, from +0.23 pooled to −0.04 among movers. On this quiz, what a visitor recognises has at most a weak association with the number they set for themselves, and the cells are 56 rows.

The three opinion questions are not findings: a belief about whether governance can keep AI safe correlating with an estimate of catastrophe is close to tautological.

Ticked most and fewest on the vulnerability checklist

The one multi-select question on this quiz, so the overall split is the same as the per-question one. Thirds by how many of the 16 entries were recognised. Left: one mark per submission, the tall tick is the group median; an arrow points the way the visitor moved the quiz's proposed number, a dot kept it. Right: what each third did with that proposed number.

Pooled: 70%, 64% and 60%. Among those who moved a slider, repeats excluded: the third who ticked most registered 36%, the middle 31%, the third who ticked fewest 27%.At the last report the movers' thirds read 51%, 26% and 50%, a middle third well below two level outer ones, on cohorts of 12 to 15. With sixteen more movers that shape has gone, and what is left is a shallow slope in the same direction as the pooled one, on cohorts of 17 to 20. Most of the pooled spread is still the proposal: the top third without repeats has 31 rows, 14 that kept the page's number at a median 95% and 17 that moved it, to 36%. The entries fewest people recognise are side-channel attacks (41 of 101), model stealing (58) and adversarial attacks (59). One association among many, with no correction for multiple comparisons.

Expert quiz84 submissions · 60 verified

Every expert answer: the tick count no longer sorts the experts, and one list of names does

The expert quiz asks eight questions: three about mechanisms, four about names in the field, and one opinion on funding. The charts in this section colour marks by how many boxes were ticked and do not distinguish verified from self-declared experts; the tooltip on each mark says which it was, and the check section above draws the groups apart. This is the cohort that changed most since the last report: 29 of its 84 rows are new, and 19 of those arrived on 1 October.

Ticked most and fewest, across all seven expert checklists

Every expert's ticks on the seven multi-select questions added together, out of 24 possible, and the cohort cut into thirds on that total. Left: one mark per submission, the tall tick is the group median; an arrow points the way the visitor moved the quiz's proposed number, a dot kept it. Right: what each third did with that proposed number.

The third who ticked most submitted a median 44%, the middle third 25%, and the third who ticked fewest 32%.At the last report the third who ticked fewest was the low one, at 31%, and the middle third the high one, at 62%. With 29 more rows the middle third has fallen to 25% and the ordering runs in neither direction; among those who set their own number the thirds sit at 34%, 20% and 32%, on 25, 28 and 12 rows. A total across seven lists does not sort this cohort. One of the seven does. On the campaign-organisations question, the 38 who recognised two or three of PauseAI, ControlAI and Microcommit submitted a median 52%, the 28 who recognised one 24%, and the 18 who recognised none 19%; among those who set their own number it is 40%, 23% and 12%. Knowing the organisations that campaign about AI risk and expecting a catastrophe go together, which is close to saying one thing twice, and it is the line the 1 October arrivals fall on the low side of: they clear the vocabulary check and recognise few of the names. The mechanism questions are still ticked in full by most takers. 19 of 84 expert rows kept the proposal.

What the data says

Fourteen things this export supports

  • The site's busiest day was 1 October, and its median was lower than any month's since the quizzes began. 36 submissions, against a previous record of 19 on 7 October 2025, with 33 of them between 20:11 and 22:38 UTC. The day's median is 21%, and 17% among the 28 who set their own number. The lowest monthly median since the level chooser went live in March is 33%, in July, on six rows, and no other day since then with five or more submissions has a median under 25%. 19 of the 36 chose the expert quiz, 18 of the 20 who took the vocabulary check passed it, and only 8 of the 36 kept the page's proposal. It is one evening and almost certainly one link, and the table cannot say which.
  • The readers who passed the vocabulary check expect powerful AI, and discount the step from danger to catastrophe more than anyone else. Their chain runs 91%, 80%, 50% among the 47 who set their own numbers. No other quiz group goes that low on the last step: beginners who set their own leave it at 70%, medium takers at 76%, self-declared experts at 92%. Only the visitors from before any quiz existed sit there too, at 50%. Read as a chain, that is a danger that arrives and is as likely as not to be contained. The verified readers of 1 October go further: 88%, 72% and 39% counting all 18, and 90%, 72% and 33% among the 14 who set their own numbers. Two things could produce it, and this export cannot separate them: a considered view, or an expert quiz that asks about mechanisms and never once mentions an outcome.
  • The page proposes a number, and 36% of quiz-takers register it with the midpoints unchanged. 92 of 255 quiz submissions, identical repeats excluded, have all three midpoints exactly where the quiz put them. Their medians by level, 72%, 91% and 70%, are the formula's output. Of the 163 visitors who did move a midpoint, 131 moved down; the median move across all of them is 30 points down, and among those who moved down it is 39. That is still the single most important fact about this dataset: 36% of the quiz submissions, a quarter of all 365 rows, are the site quoting its own arithmetic, which is why the strips and the proposed-start chart draw kept rows as dots and moved rows as arrows, and the moved rows are reported on their own. What "kept" cannot say is that nothing was touched: the uncertainty band is a separate control, and 40 of the 41 kept rows that stored the proposed band changed it.
  • Once the proposal is taken out, the three quiz levels land in the same place. Pooled medians: beginner 54%, medium 63%, expert 36%. Among visitors who set their own numbers, repeats excluded: 32%, 33% and 31%, on 42, 56 and 65 rows. At the last report the medium quiz stood out at 47%; its movers have since come down. The spread between the levels that the pooled figures show is the three quizzes proposing different numbers and being kept at different rates, 45%, 41% and 23%.
  • September was the busiest month in the series, and its median was 49%. 130 submissions, against 62 for all of August and 55 for October 2025. That is well below August's 69% and well above the 15% of the months before any quiz existed. Among visitors who set their own number, repeats excluded, it is 29%, against 50% in August.
  • Passing the vocabulary check goes with a lower number than claiming expertise did. Pooled, 31% (n = 60) against 68% (n = 22). Among visitors who set their own number the gap is 25% against 44%, nineteen points. That is a different comparison, not a decomposition: it leaves out the kept rows on both sides, and so changes who is compared as well as what. Neither figure isolates the proposal's effect, and the two groups are months apart.
  • The expert gate is a recommendation, and it has now been declined. 69 people finished the check: 60 scored 26 or better and went through, seven scored below and took the quiz it recommended, and two scored 24 and 25 and continued to the expert quiz anyway. The expert share of submissions since the check is 32%, against 33% before it; the fall to 25% that the last report recorded was undone by one day.
  • A visitor who takes two different quizzes and moves the second proposal lands where they landed the first time. Eight signed browsers have now taken two different quizzes. Five of them moved the second quiz's proposal a long way to get back to their first number, four to within five points and one to within eleven: beginner 52% then expert 54%, against a proposal of 24%; beginner 40% then medium 36%, against 94%; beginner 73% then medium 68%, against 94%; medium 97% then expert 87%, against 70%; and medium 0% then beginner 0%, against proposals of 72% and 58%. Two kept both proposals and so went where the page went, 99% then 94% for one, 79% then 30% for the other. The eighth came back thirteen days later: a kept beginner proposal of 69%, then 27% on the medium quiz against a proposal of 90.5%. Eight browsers is a hint, but it points one way: a visitor who moves a proposal moves it towards a number they already had. For two of the five that first number was itself a proposal they had kept, so the anchor is whatever they registered first, whoever set it.
  • How many boxes a visitor ticks no longer sorts the number they set, at any level. Among visitors who set their own number, repeats excluded, from the third that ticked most to the third that ticked fewest: beginner 32%, 34%, 29%; medium 36%, 31%, 27%; expert 34%, 20%, 32%. The last report had the third who ticked fewest highest on the beginner quiz and lowest on the expert quiz; neither is so now. The pooled figures still slope, because the proposal rises with the number of boxes ticked and the people who tick most keep it most often on the two easier quizzes: 57% against 29% on the beginner quiz, 45% against 34% on the medium one. On the expert quiz it is the other way round, 22% against 37%.
  • What a medium-quiz taker recognises has at most a weak association with the number they choose. The vulnerability checklist, the site's longest-standing candidate for a signal, is ρ = +0.25 across all rows and +0.12 among the 56 who set their own number, repeats excluded. Having seen AI go off the rails crosses zero. What still correlates is the three opinion questions, which is close to tautological: they ask about competition, governance and reaction speed, and they also choose the band the proposal starts in.
  • The beginner checklists are answered in full by nearly everyone. 74 of 76 beginners tick pandemic, and 73 each tick climate change, nuclear war and weapons of mass destruction; half the cohort recognises 11 or 12 of the 12 catastrophes. The list still goes with a difference in the numbers people register, 24% against 50% among those who set their own, but most of the cohort sits on its ceiling. An answer that can be wrong would separate the people who recognise the list from the people who tick lists.
  • Ticking every box goes with the self-declared group, and recognising names no longer goes with the verified one. All 22 self-declared experts ticked all four self-improvement answers, against 34 of the 60 who passed the check. On names the groups have come apart since the last report, a median 4.5 of 12 for the self-declared against 3 for the verified, and 2 for those verified since 20 September. Within the expert quiz it is the campaign organisations that go with the number: a median 52% for those who recognise two or three, 19% for those who recognise none.
  • P(powerful AI) is rated highest of the three links and the extremes are used. Its median is 88% against 81% and 80% for the two conditionals, yet 73 of 365 submissions (20%) put it at 50% or below. 17 submissions put the final number at exactly 0% and 11 at 99% or above; the p10–p90 band has a median width of 35 points. Five of the zeros are new since the last report, two of them from one browser, and all five came through a quiz. The band is not an interval around the midpoint: the page draws 4,000 samples per factor from a bell curve centred on the midpoint with a standard deviation of half the slider's ± spread, rejects draws outside 0 to 1, multiplies, and takes the 10th and 90th percentiles. Truncation near 0 or 1 skews that distribution, so the midpoint product can sit outside the band, and on 117 of the 365 submissions it does. Before 20 August 2026 the draws were uniform across each band; rows from before then keep the band they were submitted with.
  • Impatient double-submits have stopped. Zero identical repeat rows among the 163 since the 2 September fix, against six in the fortnight before it. Ten browsers did submit more than once since, and all ten changed something in between, seven the quiz and three a slider or an uncertainty band, which is a person using the site rather than a button being pressed twice.
Directions forwardfour measured, eight open

What to change next: the busiest day arrived, and the site cannot say from where

Each report ends with the changes the data argues for, and marks what happened to the ones before it. Nothing that changes what the page collects has shipped since the 19 September report, so the four measured items are the same four, measured again on 65 more rows, and all four still hold. The open items are in a new order, because 1 October made one of them urgent. The items that are about the site rather than the data, the stats page and the About page among them, are kept in the open-items note.

  • Measured
    Tell kept rows from moved rows
    Wherever submissions are drawn one by one, the strips and the proposed-start chart, a round dot is a row whose midpoints are where the quiz put them and an arrow is one that moved; the timeline and the factor panels do not make the distinction. The section on the proposed start, the bars beside each question, and the findings give the moved rows on their own. The recomputation checks out: all 93 kept rows, 92 without the one identical repeat, sit exactly where the page would have put them, and four rows with three equal midpoints at a value no proposal produces are counted as moved.
  • Measured
    Store the proposed starting point with the submission
    The page has sent a calibration object with every quiz submission since 12 September, and 105 rows now carry it. On all 105, the stored proposal and the one the build script recomputes from the answers agree to within a tenth of a point, which is the check that the script's port of the page's formula is right.
  • Measured
    Refuse to register the same prediction twice
    Six identical repeat rows in the fortnight before the fix; none in the 163 rows since. Ten browsers submitted more than once in that window and every one of them changed something, which is the case the fix was meant to let through. Three of them registered the same number twice on the same quiz with the same answers and a different uncertainty band; on the rule as written that is a person, not a double press.
  • Measured
    Reject rows without a level, and version the page
    All 163 rows since 2 September carry a level, a signature and a page version, across fourteen commits. No row without a level has arrived since.
  • Open
    Record where each visitor came from
    The last report asked for this because August arrived in three bursts that the table could not tell apart. This export has the strongest case yet: 33 submissions in two and a half hours, a median lower than any month's since the quizzes began, more than half of the day through the expert quiz, and nothing in any column to say what was posted where. Storing document.referrer, and a ?cohort= tag when the link carries one, as a diagnostic column outside the signed string would let the next report tie a burst to its source. Until then "the crowd changed", "the news changed" and "the page changed" are the same row.
  • Open
    Seed flagged decoys into all three self-selected quizzes
    Every list in the beginner and medium quizzes can only be answered correctly, so the only thing it measures is how many boxes someone ticks, and in this export the number of boxes ticked no longer sorts the number a visitor sets at any level. The beginner catastrophe list is ticked 73 or 74 times of 76 on its top four answers. The expert check shows what a decoy does, having caught 21 trips from 13 of 69 finishers. The case is set out in the note on the beginner quiz.
  • New
    Show the proposal to a random half
    Every statement here about the proposed number sets kept rows beside moved rows, which compares two kinds of visitor and not two treatments. The levels converging among movers, 32%, 33% and 31%, says the proposal accounts for the spread between quizzes; it cannot say how far the movers were pulled by the number in front of them. Showing neutral sliders to a random half of quiz-takers, with the proposal stored but not applied and the arm stored in the row, is the one change that lets a report say "because". It is specified in the second kind of instrument.
  • Open
    Add one recognition question about catastrophe pathways to the expert quiz
    The readers who pass the check put the last link of the chain at 50%, against 70% for beginners and 76% for medium takers, counting only visitors who set their own numbers, and the 14 of 1 October who did so put it at 33%. The expert quiz asks about mechanism and names and never mentions an outcome, so the cut could be rehearsal rather than belief. The note on it proposes one scored question. A rise afterwards would be evidence for rehearsal, not proof of it, since the audience moves too; the version that would settle it shows the question to a random half of expert-quiz takers with the calibration rule held fixed.
  • Open
    Keep the expert quiz's reasoning questions, and watch them
    These three questions ask which mechanisms of continuous learning, self-improvement and self-replication the reader knows. They still tell the two expert cohorts apart: 34 of 60 who passed the check ticked all four self-improvement answers, against 22 of 22 who declared themselves experts. The verified share has held at a little over half as the cohort nearly doubled, 18 of 33 at the last report.
  • Open
    Ask one identical question at every level
    The three quizzes share no question, so the level medians cannot be compared without the confound of three different question sets and three different proposals. One shared question would fix that, and would also be the control for the pathways item above. The beginner gut-reaction question sorts its cohort from 36% to 71% across three answers and would be a candidate, if the aim is to compare rather than to test.
  • Open
    Stamp each submission with the option set it was shown
    The page version pins it exactly. What remains is teaching the build script to read the denominators from the stamp instead of the hand-kept start dates, so that adding an option no longer needs a table edit. This build had to be told by hand that three new page versions carry the current term list.
  • Open
    Offer a one-minute survey after submission
    A psychologist's question, who submits 5% and who submits 90%, cannot be answered from quiz answers about films and vulnerabilities. The 2 September proposal picks eleven to thirteen items by hypothesis, places them after the submit button where they cannot cost a row, and stores them in their own table. It needs roughly a year of collection before it can say anything, which is the argument for starting.

These are directions the data supports, not a verdict on the design. Several rest on small cells: nine people turned away by the check, eight browsers that took two quizzes, thirds of 10 to 28 people, and one day that supplies a tenth of all the rows.

Visitor identity

Since 2 September, not one identical repeat in 163 rows

This section is bookkeeping for the people who run the site: it checks that the rows above are what they claim to be. Each browser holds a non-extractable P-256 key and signs a canonical string of its submission. The report verifies every signature against the public key the row carries, so "these rows came from one browser" is checked rather than asserted. What a key cannot show is that two different keys are two different people.

Signatures

Browsers that submitted more than once

RowsWithinLevelsp(doom)Reading

A visitor who presses submit twice used to produce two rows. Since a fix on 2 September, none has.All six confirmed repeat rows in the table predate it, and every one of the 163 rows since is signed, verifies, and carries the commit the page was built from, across fourteen page versions. Ten browsers submitted more than once in that window and every one of them changed something in between, which is the case the fix was meant to let through: three went beginner then medium, one of them thirteen days apart; two went beginner then cleared the check and took the expert quiz; one went medium then beginner; one took the medium quiz twice and then the expert one; and three took the same quiz twice with the same answers and different sliders or bands. The ones who changed quiz are the interesting rows in the table, and the pattern in them is taken up in the findings. The 36 rows of 1 October come from 34 browsers: one took the beginner quiz twice in 42 seconds and registered 9.2% and then 26%, and one kept the beginner proposal at 79% and then the expert proposal at 30%. Three browsers have registered the same number twice on the same quiz, changing only an uncertainty band in between: 99% within 56 seconds, 32% within 47 seconds and a third, under 5%, within 137 seconds. The rule looks at the whole payload, so it cannot catch those, and on its own terms it should not. All 194 signatures verify, and every field the signature covers agrees with the row: the p(doom) and its range, the three factors, the timestamp, the quiz level, the check score, the signing key and its counter. The signature does not cover the quiz answers, the terms ticked on the check, the proposed number or the page version, so those columns are trusted, not checked. A key is a browser's signing identity, not a person: two keys can be one person, and one key is whoever holds that browser.