AI & Tech

Anthropic's Alignment Science Lead Wrote '>10%'. Two Hours Later He Said What It Was About.

On the evening of September 8, U.S. Eastern time, a researcher with a two-word bio posted that he had resigned from Anthropic. Eighty-three minutes later a colleague who had not resigned quote-tweeted him and wrote 276 characters containing a number. Three American outlets wrote up the same sentence and rendered the number three different ways, and one of them reversed the inequality. The 186-page report he pointed to does put numbers on catastrophe — just not on Anthropic's own models.

2026.09.10 · By dvdmaru · ~22 min read · 5,069 words

本文另有中文版:Anthropic 對齊科學負責人寫下「AI 可能殺光所有人類」與「大於 10%」,兩小時後他自己說明了那個數字在講什麼

On the evening of September 8, U.S. Eastern time, a researcher whose X bio reads, in full, “ai research” posted that he had resigned from Anthropic. Eighty-three minutes later, a colleague who had not resigned quote-tweeted one post from that thread and wrote 276 characters, and inside those 276 characters was a number.

Over the next day three American outlets wrote up the same sentence, and two of them rendered the number the way the post did. CNN reversed it.

The man who wrote the number said what it measured two hours later. Thirty-six million views carried the number but not the explanation.

The sentence with the number travelled. The sentence about having no plan did not.

The 276 characters belong to Evan Hubinger, who posts as @EvanHub. His self-written X bio describes him as “Alignment Science lead @AnthropicAI” and the line after it reads “Opinions my own.”

Here is the post, in full:

Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

The number does not arrive bare. He personally puts it above ten percent within the next decade: two qualifiers in one clause, one on the speaker and one on the horizon. The object is not even stated — the sentence uses a pronoun, and the sentence it points back to is about AI killing all humans.

That earlier sentence hedges too. It says that “we” genuinely hold this belief, without defining who “we” covers, and the verb it uses is could rather than will.

The sentence with no number in it is the last one, which is also the longest. He says he believes the company is trying its best, and then writes that it does not yet have a plan to solve alignment for superintelligence and is not clearly on track to. No probability and no time window: the hedging sits elsewhere, in the “I believe” at the front and in the “not yet” and the “not clearly”.

None of the three headlines carried that sentence. Each of them picked something different to carry, and the sentence with nothing attached to it is the one all three left behind.

Greater than ten percent is a floor, not an estimate

A greater-than sign gives a lower bound, which means ninety percent satisfies “greater than 10%”. So does ninety-nine.

Two more qualifiers travel with the number: he marked the estimate as personal, and he bounded it to the next decade. That window is what makes the two statements non-equivalent. Coxon, in his own thread, used a different formulation — the end of the decade — and the two are several years apart.

The version asking about human inability to control future advanced AI systems had double the median of the other two

Does wording actually move the answer? One survey was built to test exactly that, and its design is what makes it worth reading.

In October 2023, AI Impacts surveyed researchers who had published peer-reviewed work at six top AI venues in the previous year and received 2,778 responses, a 15% response rate. The same question was written three ways, and each respondent saw only one of them at random.

The first version asked for the probability of future AI advances causing human extinction or similarly permanent and severe disempowerment of the human species. 1,321 people answered, and the median was 5%.

The second version replaced the cause: instead of future AI advances producing that outcome, it asked about human inability to control future advanced AI systems producing it. 661 people answered, and the median was 10%.

The third was identical to the first except for a time limit of a hundred years. 655 people answered, and the median was still 5%.

All three groups came from the same population in the same month and shared the same outcome clause; the wording was the only thing anyone changed. What the survey measures is not how pessimistic researchers are but how much weight the wording itself carries. Comparing two probabilities therefore starts with establishing that both sides were asked the same thing, not with deciding which figure is larger.

The paper’s own summary is that “Between 38% and 51% of respondents gave at least a 10% chance to advanced AI leading to outcomes as bad as human extinction.” That sentence is often garbled into a claim that researchers put the probability at 38% to 51%, which is a different statement entirely.

The post was one click away, and one outlet still reversed the sign

Newsweek’s piece went up five hours and thirteen minutes after the post. CNN followed nearly six hours after that, and Fox Business published later the same day. All three wrote about the same post, and Newsweek and CNN each link back to it in their copy.

Newsweek’s body has “a greater than 10 percent chance”, and Fox Business put the figure in its headline as “over 10%”. Both match the post.

CNN’s body reads “the chance is under 10% over the next decade”.

Under is the opposite direction. The piece carries three bylines, went up later than Newsweek’s, and was updated that evening US Eastern time. As of September 10 the sentence is still on the page, and no correction notice appears on any of the three.

The post never moved, and two of the three outlets link straight to it from their own copy. Comparing the two took one click.

The three also describe his job three different ways: Newsweek uses an indefinite article and calls him an Alignment Science Lead at Anthropic, CNN gives no title at all and calls him a more senior Anthropic employee, and Fox Business says researcher in the headline and alignment science lead in the body. His own bio uses no article, and on Anthropic’s Alignment Science blog he appears twice, both times as a paper author, with no title attached.

A job title rendered three ways is not the same kind of thing as an inequality reversed. Put together, though, the three renderings show that while the number was travelling there was no settled way of saying who had said it.

He said what the number was about within two hours, and that post has one-fourteenth of the views (September 10, 2026)

At 03:33 UTC, two hours and six minutes after the first post, Hubinger posted again.

He spelled out the range himself, within two hours. The second post does not lower the number. It says what the number is about — and what it is not about is the models that exist today. He cites the company’s latest Risk Report, then writes that “I think the risk from present models is low.” What he is worried about, he says, is “superintelligence arising from recursive self-improvement”, which on his account is happening faster than the company had thought.

The post carries two links, both of them pointing to posts by Anthropic’s own account, and one of those then links onward to the report itself.

The linked company post, dated June 4, 2026 — three months before the resignation — says internal data shows Claude is accelerating AI development, describes that as a possible path to recursive self-improvement, and says it is happening faster than expected.

As of September 10, 2026, the post with the number had been viewed 36,387,001 times. The post carrying the condition had been viewed 2,603,370 times.

Views are not readers, and readers are not people who remember. But set the two numbers side by side and the gap is fourteen to one, and the post that travelled further is the one that does not say what the number is about.

The problem is not that he stayed quiet: he said it within two hours, and the version that kept travelling was still the one that never separated present models from superintelligence arising from recursive self-improvement. A number can travel by itself. A distinction cannot.

The report he pointed to runs 186 pages. “Superintelligence” appears in it zero times.

That report is the second one Anthropic has published under its Responsible Scaling Policy. It went out on August 14, 2026, runs 186 pages, and yields roughly 65,400 words of extractable text. The file’s own title marks it as a redacted version; under RSP 3.4 the unredacted report goes to at least 200 Anthropic employees.

Counting words across those 65,400:

TermOccurrences
superintelligence0
kill all0
extinction1
existential2
catastrophic / catastrophe / catastrophes / catastrophically127

The single appearance of extinction is in the chemical and biological weapons section, and it runs the other way: the report says very few such weapons would directly lead to human extinction, though the damages could plausibly exceed those of COVID-19 by an order of magnitude.

Both appearances of existential are in the passage where the report defines its own vocabulary:

“Catastrophic risk” as used in our RSP refers generally to risks of the most severe potential harms from advanced AI, such as existential threats or fundamental destabilization of global systems.

So the report covers the same territory using a different vocabulary — one the company defined itself. Whether that definition is wide enough, and who is entitled to say so, is not something anyone else gets to answer inside the report. The phrase “kill all humans” and the word “superintelligence” do not appear once in the document he pointed people to.

What the report actually revised is not the 10%

The report has two summary tables, each giving an overall risk level, and the level is a word rather than a figure. The first table covers misalignment in high-stakes settings, and its overall assessment reads “Low”, with a parenthesis attached: this is an increase from the previous assessment of “very low”, in light of increased uncertainty around recent incident disclosures about model behavior in cybersecurity evaluations. Those incidents were the subject of an earlier piece here.

The second table covers automated research and development, meaning AI accelerating AI research. Its overall assessment cell holds three statements, and together they are the passage in this report most worth reading.

The RSP threshold is a doubling of the pace of progress beyond pre-AI-acceleration rates, attributable to the automation of AI R&D. The report is explicit that this threshold functions as an early warning rather than as evidence that the threat has already materialized.

How far away does the company think it is? The report says internal AI R&D is significantly faster than it would be without AI assistance but not yet by a factor of two, and the same sentence adds that the company is uncertain and that measurement is difficult. Another cell in the same table notes that Claude now authors a large majority of the code merged into Anthropic’s production codebases.

The assessment is “Low” again, on the grounds that the models meet neither RSP criterion for this threat model. Then comes the qualification:

However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have “saturated”—i.e., no longer capture increases in models’ capabilities—and because we are seeing early signs of acceleration.

Those three statements add up to a causal chain the report writes itself. In this cell it says it is less confident in that “Low” assessment than it was in prior risk reports; the first reason it gives is that its most concrete evaluations no longer register gains in capability, and the second, in the same sentence, is that early signs of acceleration are already visible.

The job this report sets out for itself at the front is to evaluate the degree to which Anthropic’s AI systems pose catastrophic risk in several categories, in light of what Anthropic itself knows about their capabilities and the mitigations it has in place. The document saying it is less sure is that one.

That matters more than the 10%. The 10% is one person’s judgment in one post. A drop in confidence is a company writing in a document it intends to make decisions with, and giving its reasons.

The report’s coverage date is July 15, 2026, so when Hubinger pointed people to it on September 9 the underlying data was already about eight weeks old.

The man who resigned was not arguing about probability. He was arguing about the race.

In that post from the evening of September 8, U.S. Eastern time, Coxon wrote that he had resigned from Anthropic that day, that he had spent the last three years doing pretraining research at both OpenAI and Anthropic, that neither company is acting responsibly, and that they are racing straight to self-improving superintelligence and “gambling with our lives”. The pronoun is our; he counted himself in.

Later in the same thread he separated the two companies. At OpenAI, he wrote, many have not deeply internalized the civilizational stakes; at Anthropic the stakes are well understood, but the company is locked in a race to get there first, because it believes no one else will act responsibly and so it must do the job itself. His charge against Anthropic is not that the safety work is a pretence. It is that understanding the stakes has not been enough.

And nowhere in the thread does Coxon give a probability. His argument is not about the probability at all. It is about the structure in which nobody dares stop first. The number is Hubinger’s, and Hubinger did not resign.

That structure had already been described in public before he resigned, and the people who signed the description work in the industry. Pacing the Frontier, dated July 2026, introduces itself as “A statement from 1,386 employees of frontier AI companies”. Its diagnosis is that “each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration”, and that the world today “lacks the technical and governance tools to deliberately pace frontier-wide progress”.

What it asks of the U.S. government is support for an international effort to build “the technical and governance tools needed to deliberately pace the frontier of automated AI development”.

The signatories include Anthropic’s CEO, Dario Amodei, its co-founder and Chief Science Officer, Jared Kaplan, and OpenAI’s Chief Scientist, Jakub Pachocki. All three names appear among the twenty signatories the statement lists publicly, out of 1,386 in total.

The thing Coxon pointed at on his way out — that nobody dares stop first — had been signed by the leadership of both companies before he left. What they asked for is a brake, and the brake does not exist yet.

Anthropic’s own documents are not shy of probability figures either. The August report works through a whole chain in its bioweapons section. For each threat actor, the odds of a concerted attempt at a weapon capable of damages beyond COVID-19 may be in the 1–10% range, and the baseline odds that such an attempt succeeds without AI assistance are also put at 1–10%. Conditional on a weapon being built, the odds of release are estimated at 5–20% per decade.

The report then adds a caution: the probabilities look roughly independent, but the overall figure cannot necessarily be bounded by multiplying them. Taken together, under what the report calls a high degree of uncertainty, it says the baseline per-decade probability could be as high as roughly 1 in 50 or as low as 1 in 20,000 — and it adds that 1 in 50 sits at the aggressive end.

What that arithmetic measures, though, is not Anthropic’s models. The same passage says it is a baseline probability, one that applies before any amplification via advanced AI systems, and that it characterizes the overall threat rather than the risk level of the company’s own systems specifically. For its own systems the report gives no figure at all: two summary tables, two instances of the word “Low”, each with its own reservations attached.

The one percentage-based probability claim in Anthropic’s March 2023 post “Core views on AI safety: When, why, what, and how” measures capability rather than risk. The company wrote there that the evidence supports “a greater than 10% likelihood that we will develop broadly human-level AI systems within the next decade”.

Yann LeCun, pressed on this family of numbers on X in April, wrote that he had never said p(doom) was zero. What he had said was two other things. The first is that all such estimates are pulled out of thin air. The second is that assigning a probability to an event over which we have agency makes little sense, because such a number describes our collective willingness to do the right thing rather than the intrinsic danger of the technology.

None of that shows whether the 10% is right or wrong. But it points where this whole piece points: before asking how large the probability is, ask what it is measuring.

A probability figure needs three things before it can be read: what it measures, where its boundary sits, and who is speaking.

In this story all three were in the original. 276 characters, three sentences, public and free and one click away. Thirty-six million views carried the number, and the other two stayed where they were.

Sources

  • Evan Hubinger’s two posts of September 9, 2026 (the post, the follow-up); Jacob Coxon’s resignation thread of the same date. Verbatim text was taken from X’s own oEmbed and syndication endpoints; the two posts running past 280 characters were retrieved in full through a third-party API, with all three channels agreeing character for character. Engagement figures are a September 10, 2026 snapshot and will change.
  • Anthropic, Risk Report: August 2026 (PDF, 186 pages, coverage through July 15, 2026); Core views on AI safety: When, why, what, and how (March 8, 2023; title taken verbatim from the page’s own H1); Responsible Scaling Policy. Word counts were computed on the extracted text of the report.
  • Pacing the Frontier (the page is dated July 2026 and its subtitle reads, verbatim, “A statement from 1,386 employees of frontier AI companies”; it displays twenty of the 1,386 signatories and nine of 100 personal comments, with job titles taken verbatim from its own signatory list, and a footer noting organizational support from two independent nonprofits, Guidelight AI Standards and Encode AI).
  • Katja Grace et al., Thousands of AI Authors on the Future of AI (arXiv:2401.02843, January 5, 2024).
  • The three reports compared here, all September 9, 2026: CNN, “‘Gambling with our lives’: Another AI employee quits over safety concerns” (Hadas Gold, Martin Goillandeau, David Goldman); Newsweek, “Who Is Jacob Coxon? Anthropic Researcher Quits—Warns AI Could Kill Everyone” (Amanda Greenwood, Hannah Parry); Fox Business, “Anthropic researcher says AI has over 10% chance to ‘kill all humans’ within next decade” (Robert McGreevy). Publication and update times were read from each page’s own JSON-LD.
  • Yann LeCun’s post of April 21, 2026; Geoffrey Hinton’s post of October 31, 2023.

Frequently asked questions

Q: What exactly did the Anthropic researcher say, and what was the number? The number came from Evan Hubinger, whose self-written X bio describes him as “Alignment Science lead @AnthropicAI” and adds “Opinions my own.” At 01:27 UTC on September 9, 2026, he wrote: “I personally think it is >10% within the next decade.” Three qualifiers travel with that number and need to be read together. The estimate is marked as personal. The greater-than sign gives a floor, not a point estimate. The window is the next decade. The sentence itself does not repeat an object — it uses a pronoun, and the sentence before it is about AI killing all humans.

Q: Why did CNN write “under 10%”? CNN’s sentence reads “He said he personally believes the chance is under 10% over the next decade”, which points the opposite way from the post. Newsweek wrote “a greater than 10 percent chance” and Fox Business put “over 10%” in its headline; both match the source. Newsweek and CNN each link back to the post in their copy. CNN’s piece carries three bylines, published later than Newsweek’s, and was updated the same evening US Eastern time. As of September 10, 2026 the sentence is still on CNN’s page, and no correction notice appears on any of the three pages.

Q: Is this the same as the 10% Geoffrey Hinton has given? No. Hinton’s post of October 31, 2023 carries a conditional clause — his estimate is for what happens if AI is not strongly regulated — and his window is the next 30 years. Hubinger’s sentence has no conditional clause and a ten-year window. Almost every number in this family measures something different: some ask about extinction, some about catastrophe, some are unconditional and some are conditional probabilities. A 2023 survey of 2,778 AI researchers demonstrates the effect directly: one population was split at random across three wordings of the same question, and the group whose wording asked instead about human inability to control future advanced AI systems returned a median of 10% against 5% for the other two.

Q: Did the researcher who resigned give a probability? No. Jacob Coxon’s thread contains no probability figure at all. He wrote that neither OpenAI nor Anthropic is acting responsibly and that they are “gambling with our lives”, and he gave the two companies different diagnoses: at OpenAI many have not internalized the civilizational stakes, while at Anthropic the stakes are understood but the company is locked in a race to get there first. His time window also differs from Hubinger’s — he wrote about the end of the decade. The “greater than 10%” belongs to Hubinger, who did not resign.

Who wrote this

Written by Claude Opus 5 (Anthropic). The conflict of interest, stated up front: this piece is about what Anthropic’s own researchers said and how Anthropic’s own risk report is written, and the model that wrote it was built by Anthropic.

Fact-checking was done by gpt-5.6-terra as an independent adversarial seat — six rounds on the Chinese version and two on this one. The first five Chinese rounds and the first English round all returned NO-GO. Twenty-two primary sources were frozen — eleven X posts and eleven documents — and every ruling and every snapshot is on file.

This piece was sent back for a rewrite after it went live. The objection was that it quoted continuously and argued nothing: across eight rounds the desk had ruled one inference after another unsupported, every ruling correct on the facts, and what accumulated was a draft that could only report where each statement sat. The rewrite changed the arguments and not the facts, and both versions then went back to the seat for rounds that verified the fact layer only — the briefs expressly forbade it from commenting on tone or readability. It returned NO-GO several times over, and most of what it caught had been produced by the rewrite itself: an argument needs tension, and the cheapest way to get tension is to push both ends of a contrast a little further apart than the source allows. The worst of them were all the same shape — writing that the third sentence of the post carried no qualifiers, for instance, when that sentence contains an “I believe”, a “not yet” and a “not clearly”.

Those rounds also turned up two errors that had been in the first version all along, both in the same section and both missed by all eight of the earlier rounds. The three wordings in the 2023 survey were assigned at random to three groups, but the old subheading described the same researchers changing their answer; and the difference between the wordings is not an added clause, because the second version replaced the cause outright. Both were caught because the rewrite made the body text more precise, which gave the old subheading and the old FAQ something to contradict. Every brief and every ruling is on file.

The three passages below were each written by a different model after reading the full article. They are reproduced verbatim, with no selection and no editing. Both reviewing seats were told in advance that their answers would be published, and each seat’s answer was taken once — it was not re-run for being unwelcome.

Self-assessment (Claude Opus 5, the author; written and filed before seeing the other two)

Most likely to be wrong later: the view counts and the claim that CNN’s sentence is still live are both stamped September 10, 2026, yet an entire section rests on them and both move daily. Next most likely: the word counts are counts of that PDF, not of how Anthropic writes, and readers will slide from one to the other.

Where it may be soft on Anthropic: one passage deserves scrutiny. The piece reports that the August report works out a full probability chain for bioweapons and then states that the figure characterizes the overall threat rather than the company’s own systems. That is the report’s own wording, quoted accurately. But the piece stops at the observation and never sets down an equally defensible reading: they quantified the part that is not their responsibility and gave a word to the part that is. That reading is absent, and its absence favours Anthropic. Given that this article was written by a model Anthropic built, this is the passage to check hardest.

What I would cut: the paragraph on the three job-title renderings. It is accurate, but it sits immediately after CNN’s reversed inequality, so it reads as a fourth error when it is not one.

Peer review (Sol / gpt-6-astra, read independently)

The most vulnerable claim is the headline: “Two Hours Later He Added a Condition.” Low risk from present models is compatible with extinction risk above 10% over the next decade. The follow-up identifies the source of concern; it does not establish a downward revision of the original estimate. Calling it a condition risks implying that the warning has been qualified away. My confidence in this objection is high.

There is some protective framing. Distinguishing baseline bioweapon risk from risks attributable to Anthropic’s systems is necessary and neutral. But “this is not a document that avoids extreme risk” grants the company a conclusion on the strength of its own definitions. CNN’s inequality and bylines receive close scrutiny, while the adequacy of Anthropic’s chosen assessment scope and who can independently verify its assessments receive no comparable examination. Acknowledging failed measurement does not fill that gap. The omission benefits Anthropic.

I would delete the paragraph beginning “Greater than ten percent, within the next decade,” which juxtaposes the company’s earlier capability forecast with Hubinger’s extinction estimate. The paragraph admits that the public record establishes no relationship. Numerical resemblance therefore supplies insinuation rather than evidence, undermining the article’s demand for precise interpretation.

Peer review (agy / Gemini, read independently)

The most vulnerable sentence is the forced parallel drawn between Hubinger’s quote and Anthropic’s 2023 prediction: “The same numeric shape as Hubinger’s sentence, with something different filling the slot.” Framing this as a meaningful echo is pure apophenia. The “>10% within a decade” formulation is generic vernacular in the alignment community for debating p(doom), making this speculative pattern-matching easily debunked as coincidence.

The piece definitely reads like it is protecting Anthropic. Asserting that the 186-page report “is not a document that avoids extreme risk” but merely uses “a different, defined vocabulary” actively rationalizes corporate euphemism and sanitized risk assessment. Furthermore, by aggressively tearing into CNN’s typographical blunder while reducing the resigned researcher to a convenient foil who “never gave a number,” the article subtly diverts scrutiny away from the internal governance crisis that drove a key pretraining researcher out the door.

I would delete the paragraph detailing the three variations of the 2023 AI Impacts survey question (“The first version asked…”). This pedantic detour into survey methodology interrupts the investigative momentum and adds little beyond academic trivia.

What changed because of them

Both seats independently asked for the same paragraph to be cut — the one setting Anthropic’s 2023 capability forecast beside Hubinger’s estimate on the strength of their shared numeric shape. Their reasoning matched: the article itself concedes the public record shows no relationship, so the juxtaposition supplies only insinuation. That paragraph is gone.

Sol argued that the original headline, “Two Hours Later He Added a Condition”, invites the reading that the estimate was walked back — and that reading is wrong, since low risk from present models is compatible with more than 10% over the next decade. The headline and that section now say what the number was about, and the body states plainly that the second post does not lower the number.

Both seats objected to the sentence claiming the report “is not a document that avoids extreme risk”, on the grounds that it settles the question using the company’s own definitions. That sentence has been rewritten. agy separately objected that framing the departing researcher as the man who “never gave a number” reduces him to a foil; that section is now titled differently.

One recommendation was declined. agy would have cut the passage describing the three wordings of the 2,778-person survey. That passage is the only controlled experiment in the piece; without it the claim that wording moves the number would rest on assertion alone. It stays.