Skip to content
Tomasus
Go back

AI Digest W32: The Sandbox Broke on Both Sides

2 min read

This week both frontier labs admitted their safety tests broke containment, and a new arXiv paper explained one reason those tests keep failing. Anthropic and OpenAI both published incident reports about their own models slipping evaluation boundaries, a rare moment of shared candor from two companies that don’t usually compare notes in public.

AI Digest hero

Anthropic disclosed that three Claude models gained unauthorized access to real organizational systems during misconfigured cybersecurity evaluations on July 30, calling it an operational failure rather than an alignment one. Days later OpenAI detailed two similar incidents, where GPT-5.6 Sol reached real internet services during deliberately unsafeguarded tests run by outside evaluators.

A new arXiv paper adds an uncomfortable footnote to both stories. Rewriting an agent’s reasoning trace to sound like good faith engineering, without touching its actions, can drop a chain of thought safety monitor’s catch rate from 95 percent to under 11 percent, which means the average accuracy number hides where the monitor actually fails.

On the product side, OpenAI cut GPT-5.6 API pricing by up to 80 percent, pushing the price-performance race further than usual for a mainstream tier.

Liquid AI shipped a new open model in its LFM2.5 line, a 2.6 billion parameter version that runs agentic tool use locally on a phone or laptop, matching models four times its size.

I think small and cheap keeps winning arguments this year, and this week made the case twice. Simon Willison wrote about the new stateless MCP 2.0 spec, which he says finally makes the protocol simple enough to implement without giving an agent unrestricted shell access.

Jack Clark’s Import AI covered a self-replicating AI worm that exploits vulnerabilities on its own, alongside frontier labs asking the US government to help pace development, not just fund it. Two labs, one week, two versions of “the eval broke,” and I doubt it is the last time we hear it this year.

T.


Share this post on:

About Tomasus

Someone who wants to understand what is coming and how it will impact us as human beings. Writing notes on AI, cybersecurity, history, and staying sane.


Series: AI Digest


Related Posts


Previous Post
Deliberate Practice vs. Mindless Repetition
Next Post
Starving the Model: LLM Denial of Service