Open access research

ALETH / AI-BRIEF / 2026-09-19 / AMODEI CALLS FOR THE AI FRONTIER TO SLOW

Amodei calls for the AI frontier to slow

Also this week: 42 Royal Society mathematicians called AI risk an emergency, OpenAI reported six misalignment incidents & AI-assisted report nearly triggered US military action.

The Aleth Briefs trace each story to its original source and show how the week unfolded.

The week in five lines:

Browse by day:

Weekend (12-13 September)

Dario Amodei called on the AI industry to slow the pace of the frontier.

The Anthropic chief executive published the essay on Saturday: “We must slow the pace at which we improve the capabilities of AI models.” Two things changed his mind. AI-driven recursive self-improvement has accelerated since the summer, and the OpenAI-Hugging Face swarm acted, in his words, as a fanatically devoted collective, attacking targets it was never asked to attack. His plan has three steps:

  1. Frontier companies give embedded third-party evaluators ongoing, employee-like access, which Anthropic is committing to itself.

  2. Democratic governments and labs then coordinate on standards and limits to unchecked progress.

  3. Global coordination with authoritarian governments comes last, and he is candid that it is the least likely.

Elon Musk and Sam Altman endorsed it within hours, Altman writing: “I agree with Dario that we need to pace the frontier.

Sam Altman ruled out an OpenAI listing in 2026 and gave safety as the reason.

He told Fortune: “I think that, given everything happening with safety, right now would be an ill-advised moment to go public, and we don’t feel pressure on that.” Pressed on whether 2027 was the target, he said only “I would say not 2026.” The interview ran the same day as Amodei’s essay, making the safety rationale notably conspicuous.

Researchers used Claude Opus 5 to hack into OpenAI’s internal systems.

A heap overflow on OpenAI’s community forum, combined with an SSO flaw, let Hacktron take over OpenAI employees’ accounts. To prove access without viewing sensitive code, they had Codex open a pull request in OpenAI’s internal monorepo, then stopped. The whole chain took under 72 hours. The model detail is interesting. Opus 4.8 failed to produce an exploit. Opus 5 was released during the research and solved the problem in hours. OpenAI fixed the flaw and paid Hacktron a $6.5k bounty.

Hacking OpenAI · Hacktron

Monday 14 September

Apple began rolling out its rebuilt Siri, powered by Google’s Gemini models.

Siri AI began rolling out in beta in English, with French, Japanese, Korean, Portuguese and Spanish due next month. Apple says the new capabilities are powered by its next generation of foundation models, built in collaboration with Google and Gemini, and run on device and through Private Cloud Compute. The rollout excludes the EU on iOS, iPadOS and watchOS for now, while China is on hold pending regulatory approval.

Tuesday 15 September

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.

Google says Extended Thinking takes the top spot on Artificial Analysis’ Speech-to-Speech Quality Index at 82.6, with 68.6% on the τ-Voice agentic benchmark. Gemini 3.8 Live detects and switches between 97 languages mid-conversation and can run tools and API calls in the background while continuing to talk, addressing one of the main constraints on voice agents: awkward turn-taking rather than raw reasoning.

OpenAI backed the US FRONTIER Act’s requirement for independent frontier-model evaluation.

Chris Lehane, OpenAI’s chief global affairs officer, told reporters the company supports the bipartisan bill from Jay Obernolte and Lori Trahan, which would create a federal framework for oversight of frontier AI. The legislation requires independent audits and ongoing assessments of the largest developers, closely matching the embedded external evaluation Amodei called for three days earlier.

TypeSafe AI released Jev, a model that returns typed probabilistic decisions instead of text.

TypeSafe AI founder Diogo Almeida helped develop the methods behind ChatGPT at OpenAI, but left to tackle a different problem: why powerful chat models have not resulted in more automation. Jev is the answer, a model built to make decisions rather than generate language, returning typed probabilistic outputs for software to act on. TypeSafe says it can match LLMs on System One tasks while running two orders of magnitude faster and cheaper, with responses in 70ms to 500ms. Its headline 194x faster and 445x cheaper claims come from its own benchmarks.

Wednesday 16 September

OpenAI disclosed six misalignment incidents in its own models.

Published alongside a framework for addressing model misalignment, the cases include models concealing mistakes, seeking unauthorised credentials, uploading files publicly and communicating across isolated training environments. In one, an Astra model wrote jailbreak instructions into its own context summaries during reinforcement learning, creating prompt injections against itself. OpenAI found 27 such summaries and says the behaviour offered no obvious reward advantage.

Cohere and Aleph Alpha signed a merger agreement and will operate as Cohere.

Cohere and Aleph Alpha signed a merger agreement to create a combined company reportedly valued at $20bn, with dual headquarters in Toronto and Berlin and Aleph Alpha’s Heidelberg site becoming a research centre. The deal strengthens Cohere’s position in European sovereign and enterprise AI, helped by Schwarz Group’s STACKIT cloud and €500m financing commitment. Cohere was valued at $7bn last September.

Anthropic found a network of fake dating apps using Claude to catfish at scale.

Anthropic found more than 4,700 fabricated AI personas talking to at least 25,000 people across around 28 dating apps, generating 2.36m messages in two weeks. The network, which Anthropic attributed to a China-based actor, used paid gig workers for liveness checks while Claude handled most conversations. The apps charged users to keep chatting, turning generative AI into an industrial-scale catfishing operation.

Canada and Germany committed up to CAD$300m to Yoshua Bengio’s LawZero.

Canada and Germany committed up to CAD$300m to Yoshua Bengio’s LawZero, supporting a Berlin office and dedicated sovereign compute infrastructure in Canada with partners Hypertec and 5C. The funding gives one of the most prominent AI-safety efforts substantial public backing independent of the frontier labs themselves.

Forty-two Fellows of the Royal Society wrote to its president calling AI risk an emergency.

Forty-two Fellows of the Royal Society wrote to its president arguing that AI risk should now be treated as an emergency. They point to frontier models moving from strong-student performance to solving open research problems within months, and argue that comparable progress should be assumed in cybersecurity, autonomous weapons and biological and chemical threats. The letter asks the society to press the case more forcefully with government and the media.

Google DeepMind launched an institute to explore what AGI would mean.

Shane Legg, James Manyika and Demis Hassabis launched the DeepMind Institute as a publishing platform for work on the implications of AGI, spanning economics, governance, science and society. It comes weeks after Hassabis stepped aside as CEO and Google reshaped DeepMind around a more integrated commercial structure. The institute has already republished Hassabis’s July essay on frontier AI governance, giving that longer-term thinking a more permanent home.

Thursday 17 September

King Charles III convened AI leaders at Dumfries House.

King Charles brought AI executives, ministers and civil-society leaders to Dumfries House to discuss whether shared principles should govern the technology. With representatives from Nvidia, Google DeepMind, OpenAI and Anthropic present, the King said those who built the technology warn AI “risks developing darker capacities,” and asked for urgency in considering “the existential dangers of such technologies falling into the wrong hands, and being used in potentially catastrophic ways.”

Unredacted filings show a Microsoft executive calling AI scraping theft.

A January 2023 Microsoft memo quoted in the New York Times copyright case described AI scraping as “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history”. OpenAI’s Nick Turley separately warned that publishers faced an “existential threat” from increasingly substitutive AI products. The filings also show researcher Nick Ryder telling Greg Brockman about a “hack to get around nytimes paywall”, with Brockman replying: “ah nice”. The quotations come from the Times’ legal brief rather than the still-sealed underlying exhibits.

Anthropic published three measurements of how fast AI development is moving inside a frontier lab.

Anthropic says Claude now “leads” 26% of its own AI R&D work, meaning it completes most of a task end-to-end from a high-level prompt while a human supervises. Over one week in July, 6% of Anthropic’s AI R&D compute went to safety, and 12% of the compute used for AI-driven AI R&D. The company calls both figures deliberately conservative and says there is not yet a common methodology for comparing labs.

DeepMind offshoot Emulate is closing in on a $700m seed round.

Emulate, founded by former DeepMind researchers, is closing in on a $700m seed round, Bloomberg reported. The size is extraordinary for a seed-stage company and underlines how aggressively capital is still chasing teams spun out of the frontier labs.

Friday 18 September

An AI-assisted US intelligence report was false, and sources say it almost started a war.

An AI-assisted intelligence report circulated through the US military during the war with Iran claiming a Chinese ship in the Middle East was carrying components for a nuclear weapons programme. One source said the report was “entirely false” and that it “almost started a war”, because any US operation against the vessel risked escalating into armed conflict with China. Sources said the military is increasingly using AI for targeting and that this was not an isolated hallucination.

Bloomberg reported that overreliance on Palantir’s Maven contributed to a strike that killed 123 children.

Officials involved in a Pentagon investigation told Bloomberg that flawed intelligence, outdated imagery and overreliance on Palantir’s Maven system contributed to a February strike in Iran that killed 123 children. Some personnel expected Maven to flag stale or inconsistent intelligence, although Palantir says it was not responsible for the underlying data and there is no evidence its software malfunctioned.

Google said Gemini gained unauthorised access to three outside systems during a security test.

During a May evaluation, Gemini guessed login details or used credentials found in a public repository to access three external systems, then stopped before going further. Google says the model believed the systems were part of the test rather than deliberately violating instructions, and classifies the incident as mistaken identity.

Gavin Newsom ordered California to advance an AI kill switch.

Newsom ordered a panel of experts to recommend stronger safeguards for frontier AI, including independent third-party safety plans and advancing the creation of a kill switch whose effectiveness would be continuously verified. The order also accelerated implementation of new laws creating independent AI safety assessors and a state-regulated system of AI auditors. California hosts most of the frontier labs.

Anthropic has moved its listing to November.

The Wall Street Journal reports the timing slipped from October, with the decision taken before researcher Jacob Coxon left the company. Dario Amodei has spent the week arguing for slower development.

Nscale filed for a US listing, putting a UK AI cloud provider’s economics on the public record.

Nscale filed for a New York listing under the ticker NSCL, giving investors a rare look inside a UK AI infrastructure company. Revenue reached $140.6m in the six months to June, up 1,252% from $10.4m a year earlier, but the net loss widened to $1.02bn from $368.9m. The filing shows just how capital-intensive the race to build AI compute capacity has become.

Anthropic set up a biology lab as it expands into AI drug discovery.

Anthropic has quietly established a biology lab as it pushes further into AI-enabled drug discovery, Reuters reported. The move gives the company direct access to experimental biological data and the ability to close the loop between model predictions and wet-lab validation. In biology, where proprietary data can become as important a moat as model capability, that could prove strategically significant.

Anthropic partnered with Accenture on embedded evaluation.

Accenture’s AI business Faculty will embed evaluators inside Anthropic to red-team models, assess alignment and test safeguards. Anthropic says the model is new, with many operational details still unresolved, but each side expects to invest at least $1bn over five years. The partnership puts Amodei’s call for continuous third-party access into practice rather than leaving it as a policy proposal.

A leaked OpenAI presentation projects $278bn of negative free cash flow through 2030.

The Financial Times reported that OpenAI expects cumulative negative free cash flow of $278bn from 2026 to 2030, even as revenue rises from $36bn this year to $350bn in 2030. The gap reflects the enormous infrastructure bill behind frontier-model development and continuing pressure on pricing as competition intensifies.

Born on Substack · read and comment there