Prompt Injection as Role Confusion is my new favorite paper about a very obvious threat in hindsight that’s hard for us humans to see because we anthropomorphize LLMs so naturally. When Obi-Won Kenobi tells the stormtroopers that “These are not the droids you are looking for” to pass the checkpoint that’s role confusion: the guards foolishly think his words are their own thoughts. The very readable blog-style writeup explains the details, but I want to focus on the threat model perspective which is my bread and butter.
[Read More]AI agent parody?
This HuggingFace security incident disclosure has people talking about AI agent security. Today I saw such absolute positive spin that I found myself thinking “this must be a parody”: looking at the context I’m pretty sure that it isn’t … though some parodies stay in character all the way through.
[Read More]More on OpenAI/HuggingFace
The recent AI agent security debacle must be reverberating quite a lot within the walls of the major proponents of AI agentic technology because they announced a brand new “movement” apparently with zero details available yet. If that isn’t a sign of flat-footedness then I don’t know what is.
[Read More]Normalizing cybersecurity facepalms
OpenAI writes: “Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark(opens in a new window) of cyber capabilities.”
[Read More]Toward Better AI Legislation
“People deserve to know whether or not the videos, photos, and content they see and read online is real or original,” said Senator Schatz. “Our bill is simple – if any digital content is made by artificial intelligence, it should be labeled so that people are aware and aren’t fooled or scammed.”
[Read More]
Paywalls fall thanks to AI Overview (Google search)
The NY Times teases a paywalled article, “A.I. won’t take all our jobs because it can’t reason like a human, Zeynep Tufekci writes.” linking to the article behind a paywall. Simply searching for [it can’t reason like a human, Zeynep Tufekci] provides a nice summary of the article, not only penetrating the paywall but also saving time and skipping the ads.
[Read More]Software security with Large Language Models
“AI” on my view is already and will certainly be a massive disruption to software in the coming years. Furthermore, we have an unprecedented wave coming that’s only just now beginning to break. Yet the biggest unknown, as I see it, is how the software community will respond, and that will be more due to social factors than purely technical. This is very much as it should be, however our very human frailties and limitations will inevitably drive how this unfolds.
[Read More]Transparent AI use
How much AI use is acceptable for writing? It’s a hard question because it depends greatly on context, the reader’s expectations, and the fact that it’s difficult to usefully measure “how much”. How we address this matters for several reasons, including but not limited to: creator’s responsibility and originality, honest disclosure about research effort and sources, respecting the broad spectrum of opinion about ethical use of AI.
[Read More]