GLM-5.3 and Claude Mythos show gains in exploit-generation tests
09/30/2026 — 10/03, 01:12·2 sources·2 reports
Story overview
On September 30, 2026, Simon Willison published results from Anthropic's Frontier Red Team, which evaluated several models on 100 tasks selected at random from an internal Binary Exploitation benchmark. In that run, GLM-5.3 achieved full control flow hijacks in 4% of trials, while Claude Mythos Preview did so in 6%. Willison noted that although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models such as Claude Opus 4.6 and GLM-5.2 did not succeed on the benchmark at all.
Later the same day, the Chinese tech outlet OSChina reported on the same red team assessment with a different emphasis. According to its account, Anthropic's Frontier Red Team concluded that Zhipu's GLM-5.3 is the first model to be released with almost no refusal guardrails while being able to autonomously build end-to-end exploits in the way Claude Mythos Preview did five months earlier. The outlet recalled that when Anthropic released Claude Mythos Preview five months ago, it described that model as the first one capable of autonomously constructing complex end-to-end exploits.
The two reports agree that the red team evaluated GLM-5.3 and that Claude Mythos Preview had already cleared the same capability bar months before, but they frame the finding differently: the first centers on the 4% and 6% figures, the second on the absence of guardrails around the release. Neither account includes any response from Zhipu, and neither says what happens after the red team's assessment.
AI-generated from 2 reports · updated 1 hour ago
Latest turnAnthropic's Frontier Red Team evaluated Zhipu's GLM-5.3 and concluded it is the first model that can autonomously build end-to-end exploit chains the way Claude Mythos Preview did five months earlier — yet it was released with almost no refusal guardrails. Anthropic had called Mythos Preview the first model capable of that kind of autonomous exploitation.
Reports on this story headlines open the original
Anthropic's Frontier Red Team evaluated Zhipu's GLM-5.3 and concluded it is the first model that can autonomously build end-to-end exploit chains the way Claude Mythos Preview did five months earlier — yet it was released with almost no refusal guardrails. Anthropic had called Mythos Preview the first model capable of that kind of autonomous exploitation.
Anthropic Frontier Red Team found GLM-5.3 achieved full control flow hijacks in 4% of trials and Claude Mythos Preview in 6%, a breakthrough over prior models like Claude Opus 4.6 and GLM-5.2, which failed all tasks.
Other stories people are talking about
- 452NewNVIDIA launches 64GB DGX Spark desktop AI computer at $4,9995 sources
- 278Google releases Gemini 4 Argon as next-gen frontier AI model9 sources
- 232NewSurgeGemini 4 Argon tops the Arena AI leaderboard5 sources
- 232SurgeMicrosoft AI releases MAI-Transcribe-2-Streaming real-time transcription model3 sources
- 197NewAnthropic launches Claude Frontier Academy with $100M2 sources
- 185NewHugging Face open-sources AstaBrief for fast report generation2 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
