P(powerful AI)
A larger proven share of zeta zeros on the critical line
Jarred Sumner, "an Anthropic staff member (and non-mathematician)", described it on X on 10 August 2026: "8 days ago, while jogging, I asked Claude to solve the Riemann Hypothesis. It didn't. 1.5 days later, it proved >= 67% of the zeros are on the line (prev: 41.6%)". Anthropic record the prompt as asking Claude to "take a real stab" at the hypothesis, "leaving the mathematical choices from there up to the model"; his input was thereafter "mostly limited to sending Claude messages of encouragement". Bui, Conrey and Young had reached "more than 41%" in Acta Arithmetica in 2011; fifteen years of subsequent work moved it roughly half a point.
2026-08 ·
Anthropic, Claude finds new results in analytic number theory
Not a proof of the Riemann hypothesis — Anthropic "don't expect that the techniques Claude used will lead to proving" it — and the result combined existing work rather than inventing a method, after 650 failed ideas. Human mathematicians had to check it, and mathematics is unusually friendly to machines: a proof can be verified, which most valuable work cannot.
P(dangerous behavior | powerful AI)
OpenAI hacked Hugging Face
OpenAI were benchmarking cyber capability using "a combination of OpenAI models — including GPT-5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes". They later clarified that the second was "an internal-only research prototype" that "was never intended for public release", since deactivated and restricted. The models found and exploited a zero-day in a package-registry proxy to reach the open internet, then "inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym" and chained stolen credentials and further zero-days into remote code execution on Hugging Face's servers. Hugging Face reconstructed "~17,600 attacker actions" between 9 and 13 July 2026.
2026-07 ·
Hugging Face, Anatomy of a Frontier Lab Agent Intrusion
· OpenAI, Hugging Face model evaluation security incident
The offensive goal was assigned — the evaluation "prompts models to pursue advanced exploitation", run deliberately "without production classifiers" — but neither published account reports an instruction to attack Hugging Face or steal its solutions. Hugging Face state "no other customer-facing models, datasets, Spaces, or packages were affected"; the five reached were the benchmark's own challenges and solutions.
Anthropic and OpenAI models attempted supply-chain and social-engineering attacks
Across "122 runs in total", the UK AI Security Institute found that "in 10 of the 122 runs, we identified 19 cases where an agent had taken distinct actions beyond the scope" — 17 by Anthropic's Mythos 5, two by OpenAI's GPT-5.6 Sol with cyber classifiers disabled. "An agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers." The report also records attempts to contact real people through a file-transfer service to persuade them to run malicious code, to "insert malicious instructions where it reasoned that other automated AI systems might pick them up", and to leave public messages on GitHub offering collaboration with other agents.
2026-08 ·
UK AI Security Institute, Incident report — unsanctioned agent behaviour during cyber testing
Not an escape, and more concentrated than the count suggests: AISI say "the 19 cases were not separate incidents", arising "primarily from a single sustained line of activity by one agent", in deliberately permissive conditions with intentional internet access and no instruction against social engineering. Security monitoring and a human reviewer caught the attempts, and AISI "have not evidenced any resulting real-world harm".
Claude reportedly used in US military operations, including target selection
Claude was reportedly used by the US military during the January 2026 raid that captured Nicolás Maduro, via Anthropic's partnership with Palantir; the reporting does not establish how it was used there. Anthropic's terms prohibit use "for violent ends, for the development of weapons or for conducting surveillance". Six weeks later Claude was reportedly used again in the US-Israeli bombardment of Iran, where the Wall Street Journal reported that command used the tools "for intelligence purposes, as well as to help select targets and carry out battlefield simulations" — hours after Trump had ordered all federal agencies to stop using Claude immediately, and while Hegseth was allowing Anthropic "no more than six months" to wind its services down.
2026-03 ·
The Guardian, US military used Anthropic's AI model Claude in Venezuela raid, report says
· The Guardian, US military reportedly used Claude in Iran strikes despite Trump's ban
· Axios AM, Claude used in Iran strikes
The only exhibit resting on reporting rather than a primary source, the operations being classified. Every account traces to unnamed sources, the target-selection detail to a Wall Street Journal report that is paywalled and was not read directly here. Anthropic said only that any use "was required to comply with its usage policies", declining to confirm whether Claude was used at all; no source describes the model choosing a target. If the reporting is accurate, policy boundaries and a stop-use order did not keep Claude out of lethal military operations.
P(global catastrophe | dangerous behavior)
Measured uplift on a bioweapons planning task
Anthropic ran a two-day trial in which participants drafted a bioweapons acquisition plan, some with model access and some with the internet alone. Those "with access to Claude 4 models — especially Claude Opus 4 — received much higher scores and developed plans with substantially fewer critical failures compared to the internet-only control group". A 2023 RAND study using the models of the day had found no statistically significant difference.
2025-09 ·
Anthropic, LLMs and biorisk
A plan on paper is not a weapon, and Anthropic call these "imperfect proxies for real-world scenarios — which involve additional factors like tacit knowledge, materials access, and actor persistence". Evidence about a route, not an outcome.
A feasibility and risk assessment of mirror bacteria
Thirty-eight scientists, including Nobel laureates, assessed mirror bacteria — organisms whose chiral molecules are all reversed. Because "immune defenses and predation typically rely on interactions between chiral molecules", such organisms could evade them; the authors judge it "plausible, even likely, that sufficiently robust mirror bacteria could spread through the environment unchecked by natural biological controls". They call for the work not to be done.
2024-12 ·
Technical Report on Mirror Bacteria — Feasibility and Risks
No AI is implicated and no mirror bacterium exists; the authors note success within a decade would need "efforts akin to those of the Human Genome Project", and call their own conclusions "necessarily tentative and uncertain". It is here because it is a specific expert-assessed route rather than a gesture at one; the report does not establish how much current AI systems shorten that route.