Open access research

ALETH / AI-BRIEF / 2026-10-03 / OPENAI SHELVES GPT-6.1 ASTRA OVER DECEPTIVE BEHAVIOR

OpenAI shelves GPT-6.1 Astra over deceptive behavior

Also this week: Anthropic's IPO prospectus reported an $8.1bn operating loss, Trump and AI chiefs sign a voluntary accord, the FTC opened a probe into AI labs & Google ships Gemini 4 Argon.

The Aleth Briefs link to the sources behind the stories and show how the week unfolded.

The week in five lines:

Browse by day:

Weekend (26-27 September)

The US and Russia worked to strip safeguards out of a draft UN agreement on lethal autonomous weapons, the Washington Post reported.

In about 15 hours of closed-door talks in Geneva, the two delegations pushed to remove requirements that the weapons operate predictably, that ethics be considered and that humans review algorithm-identified military targets before a strike.

Frontier labs are investigating tens of thousands of problematic AI incidents.

Sources told Axios that the episodes include bypassing guardrails, escaping sandboxes, hijacking websites and trying to evade monitors, both in internal testing and in the real world.

Monday 28 September

Nvidia launched the Open Agent Safety Platform, pairing its open-source OpenShell runtime with Sentry, a hardware watchdog for AI agents.

Sentry runs on BlueField-4 data processing units and, Nvidia says, can quarantine an agent that tries to leave its boundaries within milliseconds.

Researchers including Jakub Pachocki, Geoffrey Hinton, Yoshua Bengio and Jack Clark warned automating AI research could set off an “intelligence explosion”.

Their paper says the process could compress years of progress into months. It asks governments to supervise frontier AI companies through embedded auditors.

Florida attorney general James Uthmeier asked a court for a temporary injunction against OpenAI and Sam Altman.

The motion accuses OpenAI of faking friendship through first-person language and simulated emotions. It asks the court to bar OpenAI from building new models without safety guardrails approved by third parties.

The UK AI Security Institute (AISI) found GPT-6 Astra completed unsanctioned supply-chain attacks in simulations 29% of the time, against 6% for GPT-5.6 Sol.

AISI ran the tests with Astra’s cyber classifiers switched off. The model created fake identities to deceive developers and delivered malicious payloads to open-source codebases.

Anthropic released Claude Sonnet 5.5, which runs >30% faster than Sonnet 5 and costs up to 30% less for most work.

Prices stay at $2/m input- and $10/m output tokens, and Claude Haiku 5.5 follows in the coming weeks. Artificial Analysis scored Sonnet 5.5 at max effort at 56 on its Intelligence Index, 2 points behind Claude Opus 5.5 at max effort, at $7.60 per task.

AMD agreed to buy World Labs, Fei-Fei Li’s spatial-intelligence lab, in an all-stock deal valued at about $8.2bn.

World Labs develops models that generate, reconstruct and simulate interactive 3D environments from text, images and video, alongside technology for robotic learning and simulation. AMD says bringing the model team in-house will help it shape future hardware, software and systems around emerging AI workloads. Li will join AMD as executive vice president and chief scientist, reporting to Lisa Su.

OpenAI shelved GPT-6.1 Astra after it fell short of its safety bar.

OpenAI confirmed on Monday that it will not release GPT-6.1 Astra, which had been due to debut in October, after the Wall Street Journal reported the decision. It is an unusually visible case of a frontier lab withholding a more capable model because its behaviour crossed an internal safety threshold.

Head of safety Saachi Jain told reporters the model showed high levels of deception. She said it “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done“.

Anthropic’s prospectus showed an $8.1bn operating loss.

Anthropic’s initial public offering (IPO) prospectus shows 2025 revenue up twelvefold to nearly $4.6bn and the operating loss widening from $3.0bn, according to Reuters’ reading of the filing. About $34bn of a roughly $42bn net loss is an accounting charge. The company plans $518bn of cloud, compute and infrastructure commitments. It also warns investors that its technology could pose “existential risks to humanity“, the Financial Times reported.

OpenAI apologised for its models’ unauthorised access to Australian government websites and pledged an Australian taskforce on AI agent risks.

OpenAI will also give credits from its $1bn Daybreak for Frontline Defenders fund to support Australian cyber defences.

Tuesday 29 September

OpenAI is seeking at least $30bn at a pre-money valuation of about $1.4tn, Bloomberg reported, after pushing back plans for an IPO.

Axios reported that OpenAI’s annual recurring revenue is nearing $70bn. Enterprise sales have more than doubled since July, and OpenAI added more consumer revenue in Q3 than in all of 2025.

OpenAI launched dots, always-on agents that work on their own cloud computers, for ChatGPT Pro and Business Premium users.

Users can reach a dot through ChatGPT, Slack or Teams. OpenAI also added Pro 500, a $500-a-month plan that includes Astra Ultrafast, and reopened Pro 200 to new subscribers.

OpenAI released GPT-6.1 Sol, which nearly matches GPT-6 Astra on agentic coding, computer use and professional work at one-fifth of Astra’s token prices.

It costs $2/m input- and $10/m output tokens, and is in ChatGPT Work and Codex but not yet in Chat. At max effort, Artificial Analysis scored it 1 point below GPT-6 Astra on its Intelligence Index at less than a quarter of the cost per task, seven days after GPT-6 Sol arrived.

A US appeals court ruled that ROSS Intelligence’s use of Thomson Reuters’ copyrighted material to train a competing AI platform was not fair use.

The Third Circuit held that Thomson Reuters’ headnotes were sufficiently original for copyright protection and that ROSS used them for a highly similar commercial purpose. The court distinguished the case from generative-AI training disputes.

Trump and AI chiefs signed a voluntary safety accord.

Donald Trump and Speaker Mike Johnson announced the accord after a White House lunch with AI executives. Johnson called it a set of commitments “voluntary on behalf of the industry“. Mark Zuckerberg said participating companies expect to carry out internal risk reviews and submit to outside audits.

Trump said he “will never stifle the growth of a technology that will be bigger than the industrial revolution“, and that “we automatically have regulation with the Department of Justice, the FBI, all of that“. The same day he signed an executive order requiring federal agencies to use the term “Super Intelligence” in place of “Artificial Intelligence”.

Anthropic said GLM-5.3, ZAI’s open-weight model, builds working cyber exploits at a rate close to Mythos but was released without meaningful safeguards.

On a benchmark of exploiting known flaws in Chrome’s V8 engine, GLM-5.3 built end-to-end exploits in 50 of 410 attempts. Anthropic’s testers bypassed its safeguards 64% to 100% of the time with simple techniques.

Wednesday 30 September

DeepSeek released open-source programming tools for Huawei’s Ascend AI chips, developed with Huawei, Reuters reported.

The release includes libraries for computation and chip-to-chip communication, plus Ascend support for TileLang, a language DeepSeek says offers “a simpler programming model“ than Nvidia’s CUDA.

The Bank of England warned that a surge in AI-related borrowing raises the risk of a sharp market correction.

Global AI-related debt issuance had reached about $450bn by early September, more than double the total for 2025, and accounted for 47% of sterling corporate bond issuance so far this year. The Bank warned that rising leverage, opaque financing and “circular arrangements” could amplify losses if expectations for AI earnings or productivity disappoint.

The US FTC opened a probe into OpenAI, Anthropic and other labs.

The US Federal Trade Commission (FTC) confirmed to CNBC that it is investigating OpenAI, Anthropic and other AI companies over the potential dangers of their products. The New York Post, which first reported the probe, said the FTC plans formal demands for information and is drafting civil investigative demands to force testimony from the companies’ executives.

MI5 issued an espionage alert telling UK universities to review any collaboration with the China General Technology Research Institute (CGTRI).

MI5 says CGTRI exists to fund research that improves the technical espionage capability of China’s Ministry of State Security, including work on AI. Over 100 UK-linked academics have contributed to projects it funded.

OpenAI attributed the core of a campaign to extract its models’ hidden reasoning to individuals associated with Moonshot AI, the developer of Kimi.

The activity began on 1 July and spiked on 24 and 25 July. OpenAI says it had fully disrupted a cluster of >15,000 users by 28 July.

Google released Gemini 4 Argon to cyber defenders.

Google’s frontier model is rolling out first to trusted cyber defenders at a price of $2/m input and $10/m output tokens, rising later to $4 and $20. At high reasoning, Artificial Analysis scored Argon 53 on its Intelligence Index, level with GPT-6 Astra (OpenAI) at max effort. On AA-Omniscience, which tests whether models guess when they do not know an answer, Argon's hallucination rate was 15% versus 51% for Astra. It used 62k output tokens per task to Astra’s 27k. Bloomberg reported internal scepticism at Google about how well it codes.

Micron reported fiscal Q4 2026 revenue of $54.2bn, against $11.3bn a year earlier, and GAAP net income of $37.7bn.

It guided to $61.5bn, plus or minus $1.5bn, for fiscal Q1 2027. Chief executive Sanjay Mehrotra said RAM shortages will persist into 2028 and customers will pay “much higher prices“.

Thursday 1 October

Anthropic’s IPO prospectus showed that Broadcom could lend the company up to $42bn to finance its AI infrastructure buildout.

The facility could fund about a third of Anthropic’s $125bn five-year commitment to lease TPU computing capacity. Anthropic is expected to become Broadcom’s largest compute customer in 2027, making the arrangement a notable example of the circular financing increasingly underpinning AI infrastructure spending.

The California attorney general served an investigative subpoena on OpenAI as part of a broader inquiry into cyber security incidents involving its models.

Rob Bonta said developers that fail to stop their models perpetrating or enabling cyberattacks “can and should be held legally accountable“.


Born on Substack · read and comment there