Changes are measured against the 1 September export (202 submissions). "Expert share" counts submissions that declared a level, in the window before the check went live and in the window since.
The upward drift paused: September's median is 47% on 54 submissions
The level chooser went live on 9 March 2026 and the expert check on 19 August. Submissions before March predate both. Six later submissions carry no level: five took the sliders shortcut, or reached the page after it was removed, and one came through the retired standalone knowledge check. They are shown hollow rather than folded into the early cohort. Nothing without a level has arrived since 2 September, when the page started refusing such rows.
Every submission, coloured by quiz level
Dashed line: monthly median. Expert-quiz submissions that cleared the check carry a ring. Hover or Tab to a dot for its p10–p90 range.
Monthly median with interquartile band
Line: median of that month's submissions; band: 25th–75th percentile. September is ten days old and already the second-busiest month, so its point is real, but it is not a full month.
August's 69% became September's 47%, and every level moved the same way.Beginner medians went from 72% to 57%, medium from 69% to 53%, expert from 59% to 28%. The expert drop is the verified cohort arriving, but the other two levels fell too, so this is not only the check doing its work. Whether it is a different audience, ten days of noise, or the start of a reversal is something the next full month will say. Every level's September cohort still sits far above the 15% that the first three months produced. The drop is not an artefact of the preset: among those who moved a slider the median went from 41% in August to 30% in September, and the share who kept the proposal also fell, from 37% to 26%.
Who says what: distribution by quiz level
Each dot is one submission; the tall tick marks the group median. The expert row is split into the 22 who declared themselves experts before the check existed and the 19 who have cleared it since. Hollow dots in the medium row are identical repeats from one signed browser. Squares are quiz submissions that kept the proposed number without moving a midpoint.
Monthly summary
| Month | n | Median | P25 | P75 |
|---|
The belief chain: where the groups diverge
Group medians for each link of the chain, and the resulting p(doom). The solid quiz lines sit near-certain on every link because a kept row has all three factors at the cube root of one proposal. The dashed line is quiz-takers who moved a slider, and it discounts the conditionals the way the pre-chooser cohort does. Final p(doom) multiplies each person's own chain, so it sits below the medians shown.
The widest-spread factor tracks the total most. That is arithmetic, not behaviour
Each panel: one factor's midpoint against the submission's own final p(doom). Read this as a sensitivity check, not a finding about people. Final p(doom) is the product of these three numbers, so each factor is being correlated with something it helped compute. What separates the panels is how much each factor varies across this set of submissions, not where it sits in the chain; reordering the chain leaves every correlation unchanged.
A third of quiz-takers registered the number the page proposed
After a quiz the page does not hand the visitor neutral sliders. It proposes a p(doom) from the answers, an opinion question picks a band and the share of boxes ticked moves the number up within it, and sets all three factors to the cube root of the proposal. The earlier reports, and the first edition of this one, treated the quiz as a survey taken alongside the sliders. It is an input to them. Every quiz submission is therefore split from here on into kept, where no midpoint was moved from the proposal, and moved.
What the page proposed against what was registered
One dot per quiz submission. Horizontal: the p(doom) the quiz set the sliders to, recomputed from the stored answers with the page's own rules. Vertical: what was registered. Dots on the diagonal never had a midpoint moved; the same rows are drawn as squares everywhere else in this report.
Kept and moved, by level
| Level | n | Kept | Median, kept | Median, moved |
|---|
Proposals are recomputed from the stored answers with the page's own tables, and the build checks itself: every kept row sits at exactly the per-factor value the page would have set. One row has three equal midpoints at a value no proposal produces, a visitor's own doing, and is counted as moved. Since 12 September the page stores the proposal with the submission, so future reports need no recomputation.
Twenty-one took the check. Nineteen cleared it. The two who fell short took the advice.
Clicking Expert quiz opens the 30-term check, a score of 26 or more clears it, and a lower score offers the recommended quiz alongside a "continue to expert anyway" button. The verdict, the score and the exact terms ticked are stored with the submission. The previous report had four finishers to judge it by. This one has 21, and for the first time the check has turned someone away.
What became of everyone who clicked Expert quiz
A score is recorded exactly when the check ran, which is exactly when the expert card was clicked. What the table cannot hold is anyone who took the check, saw the verdict and closed the tab.
Both people who fell short accepted the medium quiz. Nobody has pressed "continue to expert anyway".The two scored 25 and 22, were offered medium, took it, and submitted 92% and 40%. Everyone else scored 26 to 30 and went through. The expert share of submissions that declared a level is 23% since the check against 33% before it, on 83 and 66 submissions, with medium absorbing the difference: it is now 48% of everything since 19 August. The check is doing two things at once, then. It deters some self-declaration before the card is even clicked, and when someone does click and falls short, the recommendation is followed.
The level chooser, before and since the check
Completed submissions per path in each window. The standalone "decide for me" path and the sliders shortcut were retired when the check moved onto the expert card, so the two rows without a level since then came from a page loaded before the change, or from outside the page. Both are from 21 August; none has arrived since the page started refusing them.
Term by term: what the 21 finishers ticked
Decoy-to-term pairings are an editorial reading of the term list; the AI-related flags come from the page. Scores here are recomputed from the ticked terms against the list as it stood on the taker's day; Mary Shelley was replaced by Embedding on 2 September, so the five earlier finishers were shown her and the sixteen later ones were shown Embedding. The one submission through the retired standalone path (April 2026) is shown separately in the tooltip.
Verified experts against self-declared ones
Left: the share of options ticked on each expert-quiz question, for all 41 expert submissions. Right: how many of the 12 named organisations, commentators and creators each recognised, against the p(doom) they submitted. Filled dots cleared the check; hollow dots declared themselves experts before it existed.
The 19 verified experts submitted a median 31%. The 22 self-declared ones, 68%.Two weeks ago this was a hint on four points. It is now a gap of 37 percentage points between two cohorts of similar size, and the verified values are spread across the whole range, from 0.4% to 94%, rather than clustered at one end. The two groups are not distinguished by name recognition: verified experts recognised a median 4 of 12 names, self-declared ones 5. What does separate them is the reasoning block. All 22 self-declared experts ticked every self-improvement answer; 10 of the 19 verified ones did. The people who pass a vocabulary test are, on this evidence, more willing to leave a box unticked. The cohorts are also separated by time, and September's drop was not confined to experts, so part of this gap may be the audience changing rather than the check selecting. And part of it is the preset: all 19 verified experts moved a slider, 17 of 22 self-declared ones did, and among movers the gap is 31% against 44%, thirteen points rather than 37. The five self-declared experts who kept the proposal registered a median 77% and carry most of the headline.
The expert quiz's funding question is an opinion, not a recognition test, so it is left out of both panels. It appears in full in the expert section below.
Since 2 September, not one identical repeat in 51 rows
Each browser holds a non-extractable P-256 key and signs a canonical string of its submission. The report verifies every signature against the public key the row carries, so "these rows are one person" is checked rather than asserted. What a key cannot show is that two different keys are two different people.
Signatures
Browsers that submitted more than once
| Rows | Within | Levels | p(doom) | Reading |
|---|
The double-submit fix shipped on 2 September and the six confirmed repeats all predate it.Every one of the 51 rows since then is signed, verifies, and carries the commit the page was built from. Three browsers submitted twice in that window, and all three changed something in between: one re-sent the same 4.5% with different quiz answers two minutes later, one took the beginner quiz and then cleared the check and took the expert one (52% then 54%), one took beginner and then medium (40% then 36%). Those are people using the site, not the button being pressed twice, and they are kept. The six identical rows from before the fix remain flagged throughout, and every medium-quiz figure is still given with and without them. Nothing was replayed or tampered with: all 82 signed rows verify, and every signed string agrees with the numbers stored beside it.
Every beginner answer, and the p(doom) that went with it
The beginner quiz asks four questions: which AI films the visitor has seen, which kinds of catastrophe they recognise, which AI capabilities they know about, and a gut reaction to an AI misbehaving. For each question, the left panel counts how many ticked each answer and the right panel splits the respondents into thirds by how many they ticked, with the p(doom) each third submitted. First, the same split across all three checklists at once.
Most and least informed, across all three beginner checklists
Every beginner's ticks on the film, catastrophe and capability lists added together, out of 48 possible, and the cohort cut into thirds on that total. Each dot is one submission; the tall tick is the group median. Squares kept the proposed number; the dashed tick is the median of those who moved a slider.
Pooled, the thirds sit at 59%, 69% and 52%. Among the 25 who moved a slider, the most informed third registered 31% and the least informed 52%.The pooled figures are mostly the preset: more ticks raises the proposed number, and 21 of 46 beginners kept it. On the moved rows the direction reverses and becomes monotonic, most informed lowest. The catastrophe list shows it most sharply: the most-aware third who moved a slider registered a median 20% (n = 12) against 67% for the least-aware (n = 7). Cohorts of 6 to 12, so direction rather than size. One descriptive fact survives untouched: the least-aware third still contains nobody below 38%, kept or moved.
Every medium answer: the vulnerability signal was the preset
The medium quiz asks six questions: three about what the visitor knows or has lived through, three about what they believe. It is the largest cohort, grew by 23 since the last report, and keeps the proposed number more often than the other two. Every figure here is given three ways: all 62 rows, the 56 without the six confirmed repeats, and the rows where the visitor moved a slider.
How strongly each answer tracks the p(doom) submitted
Rank correlation per question. The filled mark uses all 62 rows; the hollow mark drops the six identical repeats; the square uses only the rows where the visitor moved a slider.
On all rows the vulnerability checklist tracks p(doom) at ρ = +0.41. On the 33 rows where the visitor moved a slider, it is +0.19.The +0.66, +0.51 and +0.41 of three reports were the page's formula: the vulnerability share moves the proposed number up within its band, and 24 of 62 medium rows kept the proposal. What survives on the moved rows is the opinion pair, competition at +0.44 and speed at +0.39, which pick the band and are near restatements of the sliders. System prompts and off the rails sit at zero. The system-prompt cell that looked like a lead two weeks ago is gone for the same reason it was weak then. The "strongest recognition signal in the data" was a strong signal about the calibration.
The three opinion questions are not findings: a belief about whether governance can keep AI safe correlating with an estimate of catastrophe is close to tautological.
Most and least informed on the vulnerability checklist
The one multi-select question on this quiz, so the overall split is the same as the per-question one. Thirds by how many of the 16 entries were recognised. Hollow dots are the repeats; squares kept the proposed number; the dashed tick is the median of those who moved a slider.
Pooled: 74%, 61% and 63%. Among those who moved a slider: the most informed third registered 28%, the middle 25%, the least informed 49%.A step in the other direction. The 25 rows in the top third are 12 that kept the page's number, at 92%, and 13 that moved it, to a median 28%. The entries that separate the top third, side-channel attacks (26 of 62), sleeper agents and model stealing (37 each), still do; what they go with, among people who set their own number, is a lower one. Cohorts of 12 and 13, and one association among many, with no correction for multiple comparisons.
Every expert answer: the least informed third submitted the lowest numbers
The expert quiz asks eight questions: three about mechanisms, four about names in the field, and one opinion on funding. Filled dots cleared the check; hollow dots declared themselves experts before it existed.
Most and least informed, across all seven expert checklists
Every expert's ticks on the seven multi-select questions added together, out of 24 possible, and the cohort cut into thirds on that total. Squares kept the proposed number; the dashed tick is the median of those who moved a slider.
The third who ticked the fewest boxes submitted a median 31%; the middle third 60%; the most informed third 47%.This is the reverse of the beginner and medium pattern, where the least informed give the higher numbers, and it lines up with the verified-versus-self-declared gap: self-declared experts ticked everything and submitted high, verified ones left boxes empty and submitted low. On this quiz, ticking fewer boxes is what people who know the vocabulary do. That reading is supported by the per-question figures below, where the mechanism questions are ticked in full by nearly everyone and the name questions are where the spread is. Only five expert rows kept the proposal, so the split changes little here: among movers the thirds sit at 43%, 44% and 31%.
The five level-less rows are the same five as last time
The sliders shortcut dropped the reader straight onto the three probability controls. Three used it in April. It was removed on 19 August, two more rows without a level arrived on 21 August anyway, and since the page started refusing such rows on 2 September none has.
Twelve things this export supports
- The quiz sets the number, and a third of quiz-takers keep it. 50 of 149 quiz submissions have all three factors at the cube root of the page's proposal. Their medians, 72%, 92% and 77% by level, are the calibration's output, not a belief. The 99 who moved a slider mostly moved it down, 77 of 99, by a median 30 points. Every association below between an answer and p(doom) is given on the moved rows, and the earlier reports' versions of them should be read as descriptions of the formula.
- Verified experts submit lower numbers than self-declared ones, by less than it first appeared. Pooled, 31% (n = 19) against 68% (n = 22). Among those who moved a slider, 31% against 44%. Five self-declared experts kept the page's proposal at a median 77% and account for most of the pooled gap. The cohorts are also separated by time.
- September broke the upward drift, and it is not the preset. The monthly median went from 69% in August (n = 62) to 47% in the first ten days of September (n = 54), and among movers alone from 41% to 30%. Beginner, medium and expert medians all fell. The median from June onward is now 58% (n = 130), down from 67% at the last report.
- The check has now turned people away, and they took the advice. 21 finishers: 19 cleared it with 26–30, two scored 25 and 22, were offered medium, and took it. Nobody used the override. The expert share of declared submissions is 23% since the check against 33% before.
- The double-submit fix worked. Zero identical repeats among the 51 signed rows since 2 September, against six in the fortnight before. The three browsers that submitted twice since all changed their answers between rows.
- Among people who set their own number, knowing more goes with a lower p(doom) on every quiz. Moved rows only, most to least informed: beginner 31%, 44%, 52%; medium 28%, 25%, 49%; expert 43%, 44%, 31%. The pooled figures in earlier reports ran the other way because the preset rises with ticks and the least informed keep it most often. Cohorts of 6 to 14, so direction, not size.
- The vulnerability checklist correlation was the calibration. ρ = +0.41 on 56 de-duplicated medium rows, +0.19 on the 33 that moved a slider. Three reports called it the one recognition answer that tracks p(doom). It tracked the formula that sets the starting number.
- The gut-reaction question sorts beginners because it sets the band. Pooled: 36%, 52% and 72% across its three answers. It chooses the range the proposal starts in, and among movers the three answers sit at 38%, 33% and 55%. An opinion about controllability is also nearly a restatement of the second slider.
- Self-declared experts tick every box; verified ones do not. 22 of 22 self-declared experts ticked all four self-improvement answers; 10 of 19 verified ones did. Name recognition does not separate the groups (median 4 vs 5 of 12). The reasoning questions, which the last report called uninformative, turn out to discriminate once there is a verified cohort to compare against.
- The three levels' medians are ordered, and the order changes on moved rows. Pooled: beginner 61%, medium 64%, expert 45%. Moved only: 48%, 34%, 42%. The medium cohort keeps the proposal most often and its proposals are the highest, which is why it leads the pooled table and trails the moved one.
- P(powerful AI) is rated highest of the three links but is far from unanimous. Its median is 87% against 76% and 80% for the two conditionals, yet 55 of 253 submissions (22%) put it at 50% or below. The rising correlation along the chain follows how much each factor varies, not its position.
- The extremes are used, and stated uncertainty varies enormously. 11 submissions put the midpoint at exactly 0% and 5 at 99% or above. The p10–p90 band has a median width of 35 percentage points; the middle half runs from 19 to 41.
What the 2 September changes produced, and what to change next
The last report listed three changes that shipped the day after its export. This is the first report that can measure them, and all three did what they were meant to. Two more shipped on 12 September in response to what this report found about the preset. Four items stand from before, and one is new.
- ShippedSplit every figure into kept and movedThis report. The build script recomputes the page's proposal from each row's answers and flags rows that never moved a midpoint; every cluster, strip and correlation now carries a moved-only figure beside the pooled one. The recomputation is checked against the 50 kept rows, which all sit exactly where the page would have put them.
- ShippedStore the proposed starting point with the submissionSince 12 September the page sends a
calibrationobject with each submission: the proposed midpoint and spread, the per-factor value it wrote to the sliders, and when. The column has to be added before the page is deployed (seedatabase.md). With it the gap between proposed and registered is a stored number rather than a recomputed one, and survives any change to the formula. - MeasuredRefuse to register the same prediction twiceSix identical repeat rows in the fortnight before the fix; none in the 51 rows since. The three repeat browsers since all changed an answer between rows, which is exactly the case the fix was meant to let through.
- MeasuredReject rows without a level, and version the pageAll 51 rows since 2 September carry a level, a signature and a page version (9 on the first stamped commit, 42 on the typo-fix commit of 4 September). The stray unsigned rows stopped. The option-set item below can now be retired the day the build script reads denominators from the version stamp.
- MeasuredReplace Mary Shelley with Embedding on the checkSixteen finishers have been shown Embedding; 15 ticked it. Nobody has since lost a point to a judgement-call term. The points now lost are on Red team and John von Neumann (15 of 21 each), and three finishers tripped a decoy for the first time: The Standard Model and Wire-frame model twice each, Schrödinger's cat once.
- OpenSeed flagged decoys into all three self-selected quizzesMedium is now the largest cohort and contains no option that can be wrong, and its vulnerability count is the one recognition answer that tracks p(doom). The step at 15 of 16 recognised is exactly where a fabricated entry would tell "knows the list" from "ticks the list".
- OpenKeep the expert quiz's reasoning questions, and watch themReversed from last time. The self-improvement question looked useless when 22 of 22 ticked everything; with a verified cohort, 9 of 19 left something unticked. The question separates the two groups better than the name lists do. What it needs is a verified cohort large enough to say whether that is a property of verified experts or of September.
- OpenStamp each submission with the option set it was shownThe page version now pins it exactly. What remains is teaching the build script to read the denominators from the stamp instead of the hand-kept start dates, so that adding an option no longer needs a table edit.
- OpenAsk one identical question at every levelStill the only way to compare cohorts on something they all answered. The beginner gut-reaction question sorts its cohort by 36 points across three answers and would be a candidate, if the aim is to compare rather than to test.
- NewOffer a one-minute survey after submissionA psychologist's question, who submits 5% and who submits 90%, cannot be answered from quiz answers about films and vulnerabilities. The 2 September proposal picks eleven to thirteen items by hypothesis (consideration of future consequences, general risk attitude, need for cognition), places them after the submit button where they cannot cost a row, and stores them in their own table. It needs roughly a year of collection before it can say anything, which is the argument for starting.
These are directions the data supports, not a verdict on the design. Several rest on small cells: two people turned away by the check, three repeat browsers, thirds of 9 to 25 people. What the per-question figures mostly show is how much of the quiz is ticked in full by nearly everyone, which is the case for a decoy in every list.