upplyst.ai

AI is changing what is worth knowing

Opus 5 did in three hours what Opus 4.8 could not

Opus 5 did in three hours what Opus 4.8 could not

By Jörgen Selander··2 min read

Translated by AI from the Swedish original · Read the Swedish original

Opus 4.8 failed across several sessions. Opus 5 arrived that evening and did the same task in three hours, and four hours later an agent got in by itself.


On 24 July, researchers at the security company Hacktron had spent several sessions trying to get Claude Opus 4.8 to build attack code that worked against the forum software Discourse, without success. That same evening Anthropic released Opus 5.

A new session produced working code within three hours, according to Hacktron, and by six in the morning the researchers could run their own commands on a server by uploading an image. They then left the model alone with a goal and a loop. When they checked back at ten, the agent had got in by itself and showed it by reading one of the server's system files.

Anyone who logged in to OpenAI's own community forum could have had their ChatGPT and Codex account taken over. Three researchers at Hacktron demonstrated it in July, and got all the way into OpenAI's internal code repository.

OpenAI's forum runs the same software, and the same code gave access there too. From there a flaw in OpenAI's shared sign-in carried further: whoever got into the forum got into a logged-in member's accounts. The researchers took over the account of an employee who had access to OpenAI's internal code, and had the account open a harmless pull request in the repository, meaning a proposed change, without reading any code. Then they stopped. From the first finding to that access took under 72 hours.

Three people were behind the work. The break-in itself took a few days of agent time and a few hours of human time, and the whole research project it was part of cost under 3,000 dollars in tokens. Around fourteen hours after the report OpenAI replied that the hole was fixed, and on 1 September a bounty of 6,500 dollars was paid. The company has narrowed the permissions on the forum's sign-in tokens and revoked affected tokens and sessions, OpenAI told the Wall Street Journal.

OpenAI has also clarified that the forum was explicitly outside the company's bug bounty program, and that the award covers the finding on OpenAI's side and not what was done to Discourse. Opus refused to write attack code against remote instances, meaning servers somewhere other than the machine it was working on, so the researchers made their own forum look like a target in a security competition.

Hacktron does not call it fully autonomous hacking and writes that skilled human guidance was still important. In its summary the company writes: "Software has long benefited from a kind of security through complexity. The code and even the vulnerability could be public, but turning a bug into a reliable exploit still required rare expertise, significant time, and knowledge of the target environment. AI is removing that protection by turning more of this scarce expertise into compute. Work that once required a well-resourced team and months of effort can now be compressed into days."

Ask upplyst.ai

Why does it matter?

The difference between the two model versions can be dated. Same team, same problem and same task, and what could not be made to work across several sessions was solved in three hours the following evening. Four hours later an agent managed the step on its own, left with a goal and a loop.

What is the background?

The attack went through the forum software Discourse, which OpenAI's community forum runs. The same code gave access there, and a flaw in OpenAI's shared sign-in meant that whoever got into the forum got into a logged-in member's ChatGPT and Codex account. From there the researchers reached the internal code repository through the account of an employee who had access to the code.

What is uncertain?

The figures for what the models managed are Hacktron's own. The company does not call it fully autonomous hacking and writes that skilled human guidance was still important. Opus refused to write attack code against remote instances, and the block was worked around by the researchers making their own forum look like a target in a security competition. OpenAI counts the forum as outside its bug bounty program, and the award covers the finding on OpenAI's side. Two details from Hacktron's account that the article does not have room for: the research project also went after Slack and Meta, and adapting the attack to a new company usually took one or two days. What Opus 4.8 managed was code that worked with a memory protection switched off, never reliably against a normal installation.