p(doom) calculator · submissions report · export of 19 Sep 2026

Almost everyone expects powerful AI. They part company over whether the damage can be contained.

300 anonymous submissions between 5 October 2025 and 19 September 2026. The calculator asks each visitor for three conditional probabilities and multiplies them: that powerful AI arrives, that it then behaves dangerously, and that the dangerous behaviour then becomes a global catastrophe. The product is that visitor's p(doom). The three numbers are the interesting part. A visitor who sets the first at 90% and the last at 50% is saying that the danger is coming and giving even odds that it is contained, and this export shows who says that: the readers who pass the site's vocabulary check give the first link a median 95% and the last one 50%, well below the 79% to 92% every other group gives it.

One thing has to be read before any figure below. A visitor who takes a quiz is not handed neutral sliders: the page computes a starting p(doom) from their answers and sets all three sliders to it. 71 of 190 quiz-takers (identical repeat rows aside) registered that proposed number without moving a midpoint. Those rows are the page's arithmetic, so wherever submissions are drawn one by one, the strips and the proposed-start chart, they are the round dots and the others are arrows, and the figures that matter are given for the other rows separately. Two words carry the split: kept means the three midpoints are exactly where the quiz put them, and moved means at least one is not. Kept does not mean nothing was touched: the uncertainty band is a separate control, and most kept rows changed it. Moved does not mean free of the proposal either; a visitor who drags a slider from 72% to 60% may have got there only because 72% was in front of them.

The unit is a submission, not a person, except where a signature says otherwise. Since 19 August the page signs each submission with a per-browser key, and 129 of the 132 rows since then carry one, so repeat rows from one browser can be told apart from two people agreeing. Before that date nothing distinguishes them.

Each card names its own baseline: the submission count is against the 10 September export (253); the medians set September against August; the expert share sets the window since the check (19 August) against the window before it, counting submissions that declared a level; the repeat count sets the rows since the 2 September fix against the fortnight before it; the kept share has no earlier baseline.

All submissions300 · Oct 2025 – Sep 2026

September is the busiest month in the series, and its median is 53%

The level chooser went live on 9 March 2026 and the expert check on 19 August, so the 98 submissions from before 9 March answered no questions at all. Six later rows carry no quiz either: five took a sliders shortcut that has since been removed, one came through a retired knowledge-check path, and all six sit at or near the page defaults. They are left out of the timeline and the group strips; the monthly figures and the factor panels count every row. Nothing without a level has arrived since 2 September, when the page started refusing such rows.

Every submission since 1 August, by quiz level

163 submissions arrived in these seven weeks, and the 161 with a quiz level are drawn, so they get the full width; the chart below covers every month. Hollow dots are readers who simply called themselves experts, before any check existed to ask. Filled dark ones are verified: they passed a 30-term vocabulary check before the expert quiz would open. The line is a running median, the median of each submission and the fourteen before it, with identical repeat rows left out, and it is dashed across gaps of more than two days between submissions. Hover or tap a dot for its p10–p90 range; with the chart focused, the arrow keys step through the dots. Six submissions that arrived with no quiz behind them are left out of this chart and the strips below.

Monthly median with interquartile band

Line: median of that month's submissions; band: 25th–75th percentile. September is nineteen days old and already holds a third of every submission ever made, so its point is solid, though the month is not over.

The median has fallen from 69% in August to 53% in September, on the two largest months the site has had.The first three months of the site, before any quiz existed, produced a median of 15%. The quiz arrived in March, the numbers rose through the summer to a peak of 69% in August, and September has come back down on 101 submissions. Every level fell with it: beginner 72% to 49%, medium 69% to 64%, expert 59% to 35%. The fall is not the page's proposed number changing: among visitors who set their own (repeat rows excluded), the median went from 50% to 31%, and the share who kept the proposal barely moved, 39% to 36%. The change is in what visitors register, not in what the page proposes. Why the visitors changed is not something a table of submissions can say. One thing outside the table bears on it, and it is the author's impression rather than a measurement: until August 2026, comments about the site in the places it was shared mostly called its author a doomer; from around August, concern about AI became an accepted position in the same places. The 98 hand-set rows from before March, 9% of them at exactly zero, come from the first period. The visitors who set their own number in August, at a median 50% that the proposal cannot explain, come from the second. If the reception changed the crowd, the September fall may be the same crowd settling, or a quieter month in the news; the table cannot tell those apart, because the site does not record where a visitor came from.

Who says what: distribution by quiz level

One mark per submission; the tall tick is the group median. An arrow pointing left is a visitor who moved the page's proposed number down, an arrow pointing right one who moved it up, and a round dot either kept the proposal or, in the bottom row, never had one. The expert rows are split into the 22 who declared themselves experts before the check existed and the 33 who have passed it since.

Monthly summary

MonthnMedianP25P75

The belief chain: where the groups diverge

Group medians for each link of the chain, and the resulting p(doom). The solid quiz lines sit near-certain on every link because a kept row has all three factors at the cube root of one proposal. The dashed line is quiz-takers who moved a slider, and it discounts the conditionals the way the pre-chooser cohort does. Final p(doom) multiplies each person's own chain, so it sits below the medians shown.

The widest-spread factor tracks the total most. That is arithmetic, not behaviour

Each panel: one factor's midpoint against the submission's own final p(doom). Read this as a sensitivity check, not a finding about people. Final p(doom) is the product of these three numbers, so each factor is being correlated with something it helped compute. What separates the panels is how much each factor varies across this set of submissions, not where it sits in the chain; reordering the chain leaves every correlation unchanged.

The proposed start196 quiz submissions

Thirty-seven per cent of quiz-takers registered the number the page proposed

After a quiz the page does not hand the visitor neutral sliders. It proposes a p(doom) from the answers, an opinion question picks a band and the share of boxes ticked moves the number up within it, and sets all three factors to the cube root of the proposal. It is tempting to read the quiz as a survey taken alongside the sliders. It is not: it is an input to them, and a visitor who answers it and presses submit has registered the page's arithmetic rather than their own. Every quiz submission is therefore split from here on into kept, where no midpoint was moved from the proposal, and moved.

What the page proposed against what was registered

One dot per quiz submission. Horizontal: the p(doom) the quiz set the sliders to, recomputed from the stored answers with the page's own rules. Vertical: what was registered. A round dot on the diagonal never had a midpoint moved; off the diagonal the mark is an arrow pointing the way the number went, the same encoding the strips use further down. Identical repeat rows are not drawn.

Kept and moved, by level

LevelnKeptMedian, keptMedian, moved

Kept means the three midpoints are exactly where the quiz put them, and nothing more. The uncertainty band is a separate control: of the 20 kept rows that stored the proposed band, 19 registered a different one. Identical repeat rows from one browser are left out of every kept and moved count, which is why the totals here are 190 and not 196. Since 12 September the page stores the proposal with each submission, and 40 rows carry it. For the quiz rows before that, the proposal is recomputed from the stored answers with the page's own tables, and the build checks itself two ways: every kept row sits at exactly the per-factor value the page would have set, and on the 40 rows that have both, the stored and the recomputed proposal agree to within a tenth of a point. Three rows have three equal midpoints at a value no proposal produces, a visitor's own doing, and are counted as moved.

The expert checksince 19 Aug 2026 · 132 submissions

Thirty-six took the check. Thirty-three cleared it. All three who fell short took the advice.

Clicking Expert quiz opens the 30-term check, a score of 26 or more clears it, and a lower score offers the recommended quiz alongside a "continue to expert anyway" button. The verdict, the score and the exact terms ticked are stored with the submission. 36 people have now finished it, which is enough to say what it does: it admits most of the people who ask for it, and the three it has turned away all took the quiz it recommended instead.

What became of everyone who clicked Expert quiz

A score is recorded exactly when the check ran, which is exactly when the expert card was clicked. What the table cannot hold is anyone who took the check, saw the verdict and closed the tab.

All three people who fell short accepted the medium quiz. No recorded submission used "continue to expert anyway".They scored 22, 22 and 25 on the score the page stored, were offered medium, and took it; their submissions span 32% to 99%, so falling short of a vocabulary bar says nothing about the number someone then registers. Everyone else scored 26 to 30 and went through, at a median score of 29. The expert share of submissions that declared a level is 25% since the check against 33% before it, on 130 and 66 submissions, with medium absorbing the difference: it is now 45% of everything since 19 August. When someone clicks and falls short, the recommendation is followed. Whether the check also deters people from clicking the expert card at all, or whether the audience simply changed, is not something completed submissions can show: a visitor who took the check and closed the tab leaves no row. Three of 36 is a thin floor for a gate, and no recorded submission walked past it.

The level chooser, before and since the check

Completed submissions per path in each window. In the first window a visitor could still skip the quizzes and go straight to the sliders, or take a standalone knowledge check; both were retired on 19 August, when the check moved onto the expert card, and since then every path on the page runs through a quiz. Two rows from 21 August arrived with no level anyway. They are unsigned and carry no page version, which means the tab that sent them was loaded before the 19 August deploy. They are left out of the right-hand panel, and since 2 September the database refuses any row without a level.

Term by term: what the 36 finishers ticked

Decoy-to-term pairings are an editorial reading of the term list; the AI-related flags come from the page. Scores here are recomputed from the ticked terms against the list as it stood on the taker's day; Mary Shelley was replaced by Embedding on 2 September, so the four earlier finishers were shown her and the thirty-two later ones were shown Embedding; which list a row saw is resolved from its page version, and the recomputed score agrees with the score the page stored on every one of the 36. The one submission through the retired standalone path (April 2026) is shown separately in the tooltip.

Verified experts against self-declared ones

Left: the share of options ticked on each expert-quiz question, for all 55 expert submissions. Right: how many of the 12 named organisations, commentators and creators each recognised, against the p(doom) they submitted. Dark dots are verified experts, grey ones self-declared; the two groups sit to either side of each column.

The 33 who passed the check submitted a median 38%. The 22 who called themselves experts before the check existed, 68%.Among visitors who set their own number the gap is smaller, 31% against 44%, thirteen points. The 30-point pooled gap includes five self-declared experts who kept the page's proposal at a median 77%; leaving them out changes who is being compared as well as what, so the two gaps are two different comparisons rather than a decomposition into proposal and belief. The two groups are not told apart by how many names in the field they recognise, a median 4 of 12 against 4.5. They are told apart by how they answer a checklist: all 22 self-declared experts ticked every option on the self-improvement question, against 18 of the 33 who passed the check. Ticking everything is consistent with a list that is half-recognised, though a checklist nobody can fail cannot tell that from thoroughness. The two groups are also separated in time, since the check only exists from 19 August, and the site's overall median fell over the same period, so some part of this is the audience changing rather than the check selecting.

The expert quiz's funding question is an opinion, not a recognition test, so it is left out of both panels. It appears in full in the expert section below.

Beginner quiz62 submissions

Every beginner answer, and the p(doom) that went with it

The beginner quiz asks four questions: which AI films the visitor has seen, which kinds of catastrophe they recognise, which AI capabilities they know about, and a gut reaction to an AI misbehaving. For each question, the left panel counts how many ticked each answer and the right panel splits the respondents into thirds by how many they ticked, with the p(doom) each third submitted. First, the same split across all three checklists at once.

Ticked most and fewest, across all three beginner checklists

Every beginner's ticks on the film, catastrophe and capability lists added together, out of 50 possible, and the cohort cut into thirds on that total. Left: one mark per submission, the tall tick is the group median; an arrow points the way the visitor moved the quiz's proposed number, a dot kept it. Right: what each third did with that proposed number.

Pooled, the thirds sit at 58%, 62% and 52%. Among the 34 who moved a slider, the third who ticked most registered 32%, the middle 31% and the third who ticked fewest 43%.The pooled figures are mostly the page's proposal, which rises with the number of boxes ticked; 28 of 62 beginners kept it. Once those rows are set aside the ranking inverts, and the third who ticked the fewest boxes gives the highest numbers. The catastrophe list is where it shows most clearly: the third who ticked most and set their own number registered a median 23% (n = 17) against 52% for the third who ticked fewest (n = 9). These are cohorts of 8 to 17 people, so this is a direction and not a measurement, and that third is not uniformly alarmed: it reaches down to 1.6%. Ticks on a checklist are a count of recognised names, not a measure of knowledge.

Medium quiz79 submissions · 73 without repeats

Every medium answer: the vulnerability signal was the preset

The medium quiz asks six questions: three about what the visitor knows or has lived through, three about what they believe. It is the largest of the three cohorts, and its kept rows carry the highest proposed numbers on the site. The correlation figure is given three ways: all 79 rows, the 73 without the six identical repeats, and the 40 of those where the visitor moved a slider. Everywhere else the repeats are drawn but left out of the kept and moved counts.

How strongly each answer tracks the p(doom) submitted

Rank correlation per question. The filled mark uses all 79 rows; the hollow mark drops the six identical repeats; the square uses only the 40 rows where the visitor moved a slider.

On all rows the vulnerability checklist tracks p(doom) at ρ = +0.29. On the 40 rows where the visitor moved a slider, it is +0.15.That is the page's own formula being measured, not a fact about people: the share of vulnerabilities recognised is one of the inputs that moves the proposed number up, and 33 of the 73 medium rows without repeats kept the proposal. Across four reports the pooled figure has fallen steadily, +0.66, +0.51, +0.41 and now +0.29, as the share of kept rows has fallen. What survives on the moved rows is the opinion pair, competition at +0.39 and speed at +0.37, which pick the band and are near restatements of the sliders. System prompts sits at zero, and "seen AI go off the rails" has crossed it, from +0.25 pooled to −0.06 among movers. On this quiz, what a visitor recognises has at most a weak association with the number they set for themselves, and the cells are 40 rows.

The three opinion questions are not findings: a belief about whether governance can keep AI safe correlating with an estimate of catastrophe is close to tautological.

Ticked most and fewest on the vulnerability checklist

The one multi-select question on this quiz, so the overall split is the same as the per-question one. Thirds by how many of the 16 entries were recognised. Left: one mark per submission, the tall tick is the group median; an arrow points the way the visitor moved the quiz's proposed number, a dot kept it. Right: what each third did with that proposed number.

Pooled: 74%, 64% and 63%. Among those who moved a slider, repeats excluded: the third who ticked most registered 51%, the middle 26%, the third who ticked fewest 50%.Once the proposal is taken out, the middle third is the low one and the outer two are level. The first edition of this report had the top third at 36%, and that figure was one browser's five identical rows; without them the top third has 24 rows, 12 that kept the page's number at a median 92% and 12 that moved it, to 51%. The entries that separate the top third still separate it: side-channel attacks (34 of 79), model stealing (45) and sleeper agents (48). Cohorts of 12 to 15, and one association among many, with no correction for multiple comparisons.

Expert quiz55 submissions · 33 verified

Every expert answer: the third who ticked the fewest boxes submitted the lowest numbers

The expert quiz asks eight questions: three about mechanisms, four about names in the field, and one opinion on funding. The charts in this section colour marks by how many boxes were ticked and do not distinguish verified from self-declared experts; the tooltip on each mark says which it was, and the check section above draws the two groups apart.

Ticked most and fewest, across all seven expert checklists

Every expert's ticks on the seven multi-select questions added together, out of 24 possible, and the cohort cut into thirds on that total. Left: one mark per submission, the tall tick is the group median; an arrow points the way the visitor moved the quiz's proposed number, a dot kept it. Right: what each third did with that proposed number.

The third who ticked the fewest boxes submitted a median 31%; the middle third 62%; the third who ticked most 48%.On this quiz the ordering runs the opposite way from the beginner quiz, where the third who ticked fewest gives the higher numbers. It fits what the check section shows: the people who called themselves experts ticked everything and submitted high numbers, while many of those who passed the vocabulary check left boxes empty and submitted low ones. The per-question figures below say where the spread comes from: the mechanism questions are ticked in full by most takers, and the name lists are what separate people. Only 10 of 55 expert rows kept the proposal, so the split changes little here: among movers the thirds sit at 42%, 49% and 28%.

What the data says

Thirteen things this export supports

  • The readers who passed the vocabulary check are the most certain that powerful AI is coming, and the least certain that it ends in catastrophe. Their chain runs 95%, 80%, 50% among the 28 who set their own numbers. No other group discounts that last step: beginners leave it at 79%, medium takers at 85%, self-declared experts at 92%. Read as a chain, that is a danger that arrives and is as likely as not to be contained. Two things could produce it, and this export cannot separate them: a considered view, or an expert quiz that asks about mechanisms and never once mentions an outcome.
  • The page proposes a number, and 37% of quiz-takers register it with the midpoints unchanged. 71 of 190 quiz submissions, identical repeats excluded, have all three midpoints exactly where the quiz put them. Their medians by level, 72%, 92% and 73%, are the formula's output. Of the 119 visitors who did move a midpoint, 93 moved down; the median move across all of them is 29 points down, and among those who moved down it is 39. That is the single most important fact about this dataset: 37% of the quiz submissions, just under a quarter of all 300 rows, are the site quoting its own arithmetic, which is why the strips and the proposed-start chart draw kept rows as dots and moved rows as arrows, and the moved rows are reported on their own. What "kept" cannot say is that nothing was touched: the uncertainty band is a separate control, and 19 of the 20 kept rows that stored the proposed band changed it.
  • A visitor who takes two different quizzes lands in the same place both times, even when the second quiz proposes something else. Five signed browsers did this. In four of them the second quiz proposed a number 17 to 58 points away from what the visitor had first registered, and the visitor moved it back, three to within five points and one to within eleven: beginner 52% then expert 54%, against a proposal of 24%; beginner 40% then medium 36%, against 94%; beginner 73% then medium 68%, against 94%; medium 97% then expert 87%, against 70%. Two of those moves went up, on a site where 93 of 119 moves go down. The fifth browser kept both proposals, which happened to be 99% and 94%. Five people is a hint, but it points one way: the number these visitors register is theirs, and they will move a proposal a long way in either direction to get back to it.
  • A third of everything the site has collected arrived in the first nineteen days of September. 101 submissions, against 62 for all of August and 55 for October 2025. September's median, 53%, is well below August's 69% and well above the 15% of the months before any quiz existed. Among visitors who set their own number, repeats excluded, it is 31%, against 50% in August.
  • Passing the vocabulary check goes with a lower number than claiming expertise did. Pooled, 38% (n = 33) against 68% (n = 22). Among visitors who set their own number the gap is smaller, 31% against 44%, thirteen points. That is a different comparison, not a decomposition: it leaves out the five self-declared experts who kept the page's 77% proposal, and so changes who is compared as well as what. Neither figure isolates the proposal's effect.
  • The expert gate works as a recommendation, not a wall. 36 people finished the check: 33 scored 26 or better and went through, three scored below and accepted the medium quiz instead. No recorded submission used the "continue to expert anyway" button; a visitor who took the check and closed the tab leaves no row. Since the check went live the expert share of submissions has fallen from 33% to 25%, with medium absorbing the difference.
  • Ticking fewer boxes goes with a higher number on the beginner quiz, and with a lower one on the expert quiz. Among visitors who set their own number, repeats excluded, from the third that ticked most to the third that ticked fewest: beginner 32%, 31%, 43%; medium 51%, 26%, 50%; expert 42%, 49%, 28%. Including the kept rows changes the beginner and medium orderings and leaves the expert one alone. The kept rows pull the other way because the proposal rises with the number of boxes ticked, and the people who tick most keep it most often: 61% against 30% on the beginner quiz, 25% against 6% on the expert quiz.
  • What a medium-quiz taker recognises has at most a weak association with the number they choose. The vulnerability checklist, the site's longest-standing candidate for a signal, is ρ = +0.29 across all rows and +0.15 among the 40 who set their own number, repeats excluded. Having seen AI go off the rails crosses zero. What still correlates is the pair of opinion questions, which is close to tautological: they ask about competition and reaction speed, and they also choose the band the proposal starts in.
  • The beginner checklists are answered in full by nearly everyone. 60 of 62 beginners tick pandemic, nuclear war and weapons of mass destruction, and the median taker recognises 10 of the 12 catastrophes. The list still goes with a difference in the numbers people register, 23% against 52% among those who set their own, but most of the cohort sits on its ceiling. An answer that can be wrong would separate the people who recognise the list from the people who tick lists.
  • Ticking every box goes with the self-declared group, not the verified one. All 22 self-declared experts ticked all four self-improvement answers, against 18 of the 33 who passed the check. Name recognition does not separate the two groups at all, a median 4 against 4.5 of 12. A checklist nobody can fail measures willingness to tick as much as it measures knowledge.
  • The three quiz levels are ordered differently once the proposal is taken out. Pooled medians: beginner 56%, medium 67%, expert 47%. Among visitors who set their own numbers, repeats excluded: 32%, 47%, 38%. The medium quiz stays highest either way; the beginner quiz drops from the middle to the bottom, because its takers keep the proposal often and its proposals are high.
  • P(powerful AI) is rated highest of the three links and the extremes are used. Its median is 89% against 81% for both conditionals, yet 59 of 300 submissions (20%) put it at 50% or below. 12 submissions put the final number at exactly 0% and 10 at 99% or above; the p10–p90 band has a median width of 35 points. The band is not an interval around the midpoint: the page draws 4,000 samples per factor from a bell curve centred on the midpoint with a standard deviation of half the band, rejects draws outside 0 to 1, multiplies, and takes the 10th and 90th percentiles. Truncation near 0 or 1 skews that distribution, so the midpoint product can sit outside the band, and on 97 of the 300 submissions it does. Before 20 August 2026 the draws were uniform across each band; rows from before then keep the band they were submitted with.
  • Impatient double-submits have stopped. Zero identical repeat rows among the 98 since the 2 September fix, against six in the fortnight before it. Six browsers did submit more than once since, and all six changed an answer in between, which is a person using the site rather than a button being pressed twice.
Directions forwardfour measured, seven open

What to change next, and what the last round of changes did

Each report ends with the changes the data argues for, and marks what happened to the ones before it. Four changes shipped earlier and can now be measured; all four did what they were meant to. A fourth, storing the proposed number with each submission, shipped on 12 September and is measured here for the first time: the stored values match what the build script had been recomputing. The rest are open, and the case for the first of them, putting answers in the quizzes that can be wrong, is stronger in this export than in any before it.

  • Measured
    Tell kept rows from moved rows
    Wherever submissions are drawn one by one, the strips and the proposed-start chart, a round dot is a row whose midpoints are where the quiz put them and an arrow is one that moved; the timeline and the factor panels do not make the distinction. The section on the proposed start, the bars beside each question, and the findings give the moved rows on their own. The recomputation checks out: all 72 kept rows sit exactly where the page would have put them, and three rows with three equal midpoints at a value no proposal produces are counted as moved.
  • Measured
    Store the proposed starting point with the submission
    The page has sent a calibration object with every quiz submission since 12 September and the database has stored it. The export had been leaving the column out; this report is the first built from an export that carries it, 40 rows so far. On all 40, the stored proposal and the one the build script recomputes from the answers agree to within a tenth of a point, which is the check that the script's port of the page's formula is right. The export now runs from the command line as well as from the stats page.
  • Measured
    Refuse to register the same prediction twice
    Six identical repeat rows in the fortnight before the fix; none in the 98 rows since. Six browsers submitted more than once in that window and every one of them changed an answer, which is the case the fix was meant to let through. One of them registered the same 32% twice within a minute behind different answers; on the rule as written that is a person, not a double press.
  • Measured
    Reject rows without a level, and version the page
    All 98 rows since 2 September carry a level, a signature and a page version, across eleven commits. No row without a level has arrived since.
  • Open
    Seed flagged decoys into all three self-selected quizzes
    Every list in the beginner and medium quizzes can only be answered correctly, so the only thing it measures is how many boxes someone ticks. The beginner catastrophe list is ticked 60 of 62 on its top three answers, and the beginner thirds among movers sit at 32%, 31% and 43%, which is no longer a gradient. The medium vulnerability list is the one that still spreads, and its correlation with the registered number has fallen to +0.15. A list that everyone completes measures nothing; the expert check shows what a decoy does, having caught 15 trips from 8 of 36 finishers.
  • Open
    Keep the expert quiz's reasoning questions, and watch them
    These three questions ask which mechanisms of continuous learning, self-improvement and self-replication the reader knows. They are the only part of the expert quiz that tells the two expert cohorts apart: 18 of 33 who passed the check ticked all four self-improvement answers, against 22 of 22 who declared themselves experts. The margin is narrowing as the verified cohort grows, so it needs watching rather than acting on.
  • Open
    Add one recognition question about catastrophe pathways to the expert quiz
    The readers who pass the check put the last link of the chain at 50%, against 79% for beginners and 85% for medium takers, counting only visitors who set their own numbers. The expert quiz asks about mechanism and names and never mentions an outcome, so the cut could be rehearsal rather than belief. The note on it proposes one scored question. A rise afterwards would be evidence for rehearsal, not proof of it, since the audience moves too; the version that would settle it shows the question to a random half of expert-quiz takers with the calibration rule held fixed.
  • Open
    Stamp each submission with the option set it was shown
    The page version pins it exactly. What remains is teaching the build script to read the denominators from the stamp instead of the hand-kept start dates, so that adding an option no longer needs a table edit.
  • Open
    Ask one identical question at every level
    The three quizzes share no question, so the level medians cannot be compared without the confound of three different question sets and three different proposals. One shared question would fix that, and would also be the control for the item above. The beginner gut-reaction question sorts its cohort from 37% to 71% across three answers and would be a candidate, if the aim is to compare rather than to test.
  • New
    Record where each visitor came from
    August's submissions arrive in three bursts with medians of 70%, 95% and 33%, and September's are steadier and lower. Each burst looks like one link posted in one place, but the table cannot say which place, so it cannot tell "the crowd changed" from "the news changed" from "the page changed". Storing document.referrer, or a one-tap "how did you get here", as a diagnostic column outside the signed string would let the next report tie each burst to its source. The author's own account of the site's reception turning around August is in a note; this column is what would test it.
  • Open
    Offer a one-minute survey after submission
    A psychologist's question, who submits 5% and who submits 90%, cannot be answered from quiz answers about films and vulnerabilities. The 2 September proposal picks eleven to thirteen items by hypothesis, places them after the submit button where they cannot cost a row, and stores them in their own table. It needs roughly a year of collection before it can say anything, which is the argument for starting.

These are directions the data supports, not a verdict on the design. Several rest on small cells: three people turned away by the check, six repeat browsers, thirds of 8 to 17 people. What the per-question figures mostly show is how much of the quiz is ticked in full by nearly everyone, which is the case for a decoy in every list.

Visitor identity

Since 2 September, not one identical repeat in 98 rows

This section is bookkeeping for the people who run the site: it checks that the rows above are what they claim to be. Each browser holds a non-extractable P-256 key and signs a canonical string of its submission. The report verifies every signature against the public key the row carries, so "these rows came from one browser" is checked rather than asserted. What a key cannot show is that two different keys are two different people.

Signatures

Browsers that submitted more than once

RowsWithinLevelsp(doom)Reading

A visitor who presses submit twice used to produce two rows. Since a fix on 2 September, none has.All six confirmed repeat rows in the table predate it, and every one of the 98 rows since is signed, verifies, and carries the commit the page was built from, across eleven page versions. Six browsers submitted more than once in that window and every one of them changed an answer in between, which is the case the fix was meant to let through: two went beginner then medium, one went beginner then cleared the check and took the expert quiz, one re-sent 4.6% two minutes later with different answers, and one took the medium quiz twice and then the expert one. The ones who changed quiz are the interesting rows in the table: three landed within five points of their first number on the second try, and the fourth within eleven, although the second quiz proposed something 17 to 58 points away. That pattern is taken up in the findings. Two of the six registered the same number twice behind different answers: 99% within 56 seconds and 32% within 46 seconds. The rule looks at the whole payload, so it cannot catch those, and on its own terms it should not. All 129 signatures verify, and every field the signature covers agrees with the row: the p(doom) and its range, the three factors, the timestamp, the quiz level, the check score, the signing key and its counter. The signature does not cover the quiz answers, the terms ticked on the check, the proposed number or the page version, so those columns are trusted, not checked. A key is a browser's signing identity, not a person: two keys can be one person, and one key is whoever holds that browser.