OpenAI writes: “Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark(opens in a new window) of cyber capabilities.” I’m appalled, there’s so much to be said just about this, there are so many hidden details to know what happened, so will be brief since the points should be crystal clear: It’s now OK for their customers’ AI agents to attack your own systems in that we aren’t rethinking if we can really trust them is premature? If their own cybersecurity practice is flawed how can they train models adept in it, including effective cyber refusal guardrails? Specifically, the walk back how “safeguards were intentionally not enabled during this evaluation” … why was this necessary in the first place, how can we trust that this time it will be enough? If this happened then it wasn’t running in a “highly isolated” environment. By definition. So many security practices apparently not followed and not mentioned in actions taken: security review, threat modeling, auditing, more sandbox testing, to name a few. Naming the model and forthcoming greater model seems like grabbing marketing hype as Mythos did; with an insecure sandbox this is not a demonstration of cyberprowess. This appears to be over-provisioning of privileges: it was in a sandbox, but the sandbox process shouldn’t have such potentially destructive privileges in the first place for a test. There’s no alternative to full steam ahead so these incidents will become more commonplace? Also last week, OpenAI disclosed that another model mistakenly deletes files which they called an honest mistake. More normalization, more of what we can expect in the AI agentic world.