<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Ai on Designing Secure Software</title>
    <link>https://designingsecuresoftware.com/tags/ai/</link>
    <description>Recent content in Ai on Designing Secure Software</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language><atom:link href="https://designingsecuresoftware.com/tags/ai/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Questions about Collective action on cybersecurity</title>
      <link>https://designingsecuresoftware.com/writings/ai-cyber-open-letter/</link>
      <pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/ai-cyber-open-letter/</guid>
      <description>&lt;p&gt;We are at an important inflection point with at least two fronts on the ongoing
challenges of software security that AI is impacting at unprecedented speed.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;the rush to offload more and more software engineering work onto LLMs,
at best without any reliable solution to keep up responsible reviews
without hampering the speed-up and cost savings which are the point
(and at worst without protections against seriously compromising what
security and reliability we have already);&lt;/li&gt;
&lt;li&gt;the threat of bad actors leveraging frontier models to exploit systems;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Right now the latter is getting all the attention
so I am focusing on that here.
Specifically a new open letter signed by many of the big players:
&lt;a href=&#34;https://openai.com/collective-cyberdefense/&#34;&gt;A call for collective action on cyber defense:
An open letter for a global surge in cyber defense&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Before getting to the topic, I want to put a pin in #1 above for another day:
it&amp;rsquo;s very important, and the solution push here for #2 is
(obviously) more of #1.&lt;/p&gt;
&lt;p&gt;This is written blog-style as I started looking into this and trying to
interpret what this open letter means to say, and what might be behind it
left unsaid.&lt;/p&gt;
&lt;p&gt;To get warmed up, my first question is about the absence in the signatories
of a few of the biggest of the big software companies: Apple, Meta, Salesforce.
Why aren&amp;rsquo;t they supporting a call for such an urgent collective action?
Without them on board our collective action is hampered from the get-go:
did they disagree with some premises or the strategy? Without naming names,
some acknowledgment of the behind-the-scenes debate (I can&amp;rsquo;t imagine they
did not participate or consider signing) we cannot help but wonder what
the problem could possibly be.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Quotes below like this are all from the letter.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;Caveat&lt;/strong&gt;: I&amp;rsquo;m not saying that I know better, but I do have questions.
This is a quick take for now, if anyone is interested I&amp;rsquo;m glad to elaborate.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We have a limited window to strengthen cyber defenses.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is both a vague (years, months, weeks?) and ominous warning
that I would say deserves at least one verified, non-classified incident
before calling for &amp;ldquo;industry and government to bring the full weight
of their technology, resources, and expertise to this effort&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;Is there any evidence that August 2026 frontier models are that much
superior for this purpose than they were, say, two months ago?
Supposedly little technical expertise is required with the big models
so if the feared attacks are possible it should have started I would think.
In my view, the signatories could have provided more evidence
to back the extraordinary claims so others could understand the
motivation for this broad and unprecedented initiative impactin
countless systems and software components.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There have been musings that AI will take off in capabilities suddenly for years,
but this time we&amp;rsquo;re sure? Perhaps they will plateau, or perhaps we
aren&amp;rsquo;t leveraging them well yet and they are better than we know.
Can anyone say for sure?&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Today’s AI advances are already giving defenders new ways to fix weaknesses &amp;hellip;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href=&#34;https://www.anthropic.com/research/mythos-preview&#34;&gt;Mythos preview&lt;/a&gt; (April 2026)
only offered that we should &amp;ldquo;Think beyond vulnerability finding&amp;rdquo;, not start fixing.
They suggest that &amp;ldquo;models can also accelerate defensive work in many other ways&amp;rdquo;,
including, issue reporting and triage, write repros and reports, aid review, etc.
That is, help humans who are doing the real work.
Translation: this will all go at human speed for the time being.&lt;/p&gt;
&lt;p&gt;With LLM driven attacks at inference speed, human defenders even greatly aided
are not going to be a match if this threat is for real, considering that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;very few software engineers have experience doing this work (and they all have day jobs);&lt;/li&gt;
&lt;li&gt;the number of vulnerabilities is vast (and nobody even knows its scale).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Unless the &amp;ldquo;new ways&amp;rdquo; are beyond what the April report lists,
then if the premise of all the coming attacks is true I&amp;rsquo;d be very worried.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Recognize that status quo security won’t be enough.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This has been the case in general for two or three decades in my view.&lt;/p&gt;
&lt;p&gt;Even if the &amp;ldquo;historically under-resourced&amp;rdquo; systems are given AI tools,
without expertise (given that fixing is gated by human participation)
this will be a flat-footed effort.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;m not saying this should all be worked out &amp;ndash; clearly it can&amp;rsquo;t be &amp;ndash;
but my point is we need a more detailed view of the current facts on
the ground and realistic assessment of the problem
to better envision how this might all work.
Without transparency, collective action is far more difficult.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s a hard problem, but I think it&amp;rsquo;s safe to say that
it&amp;rsquo;s way more than a technology problem and that
the highly competitive software is not exactly known for collective
action without market or legal pressure.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Should &amp;ldquo;Every organization&amp;rdquo; independently do all that security work,
and given that they haven&amp;rsquo;t for a long time we need to understand why not.&lt;/li&gt;
&lt;li&gt;Haven&amp;rsquo;t the cybersecurity companies and partners supposed to have
doing all this stuff for many years against conventional attacks?&lt;/li&gt;
&lt;li&gt;Haven&amp;rsquo;t governments been trying to coordinate defense and collect
incident reports for many years with little impact?&lt;/li&gt;
&lt;li&gt;AI companies granting access is great but defenders need experienced
people too (and there can&amp;rsquo;t possibly be enough of them).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Additionally, I would suggest AI companies also consider backing research
toward LLMs fixing vulnerabilities (not just finding and helping fix),
making pull requests easier for humans to assess for both blocking
exploitability as well as lowering the risk of introducing a new bug.&lt;/p&gt;
&lt;p&gt;In my view, &lt;a href=&#34;https://designingsecuresoftware.com/writings/llm-find-n-fix/&#34;&gt;fixing vulnerabilities may not be that hard&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;On top of this, it&amp;rsquo;s very well known that patching complex systems is
risky and there&amp;rsquo;s no mention of balancing that against a call for urgent
action at break neck speed.&lt;/p&gt;
&lt;p&gt;Personal opinion: I think is very important to ask for each system,
how much security improvement is sufficient?
(This is a notoriously hard, yet important question.)
It&amp;rsquo;s great when new tools are applied where the need is great, but we have
no idea how much change will be required to bring secure defenses
up to snuff &amp;ndash; or how we would even know when it&amp;rsquo;s good enough.
Pointing advanced tools at older systems (which I think is a fair assumption
for the general case, more so for the under-resourced) could produce an
effectively unbounded spew of potentially serious issues. Then what?&lt;/p&gt;
&lt;p&gt;One last question for now: why are many of our infrastructure systems
&amp;ldquo;historically under-resourced&amp;rdquo; in the first place?
Software technology aside, securing our utilities, healthcare, and other
critical systems is a glaringly obvious priority that it does not take
cybersecurity expertise to recognize.&lt;/p&gt;
&lt;p&gt;This open letter presents a strategy aimed at the technical challenge,
but in the context where infrastructure security has not been
a high priority, without first understanding those root causes first
that approach is unlikely to be sufficient in my view.&lt;/p&gt;
&lt;p&gt;Returning to the letter&amp;rsquo;s opening line, what exactly does &amp;ldquo;limited window&amp;rdquo; mean?
To me this is a veiled warning that if you act too late the game is lost.
Do they mean an adversary can take down parts of infrastructure at will,
or possibly irrevocerable infrastructure destruction or take over?
It&amp;rsquo;s horrendous to consider, but are there no physical overrides with
manual operation in the worst case, or do we throw up our hands in the
event of a remote software attack?&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://designingsecuresoftware.com/text/ch2-threats&#34;&gt;Threat modeling&lt;/a&gt;
is how I make vauge threats more concrete, and how I
regularly advise anyone every chance I get.
In this context that means beginning with each infrastructure system
making a basic threat model that would answer questions such as above.&lt;/p&gt;
&lt;p&gt;Finally I wonder if there might not be a lot we can do &lt;em&gt;without&lt;/em&gt;
a massive AI call to action.
Maybe we don&amp;rsquo;t need a huge infusion of new (bleeding edge?) technology
that requires experts to oversee and carefully hone patches for safe merges.
Just to toss out a few obvious things that in general make systems vulnerable:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Perhaps simply installing backlogged patches makes it much more secure
(note that if this is difficult then all the best AI patches in the world
are not easily deployed either).&lt;/li&gt;
&lt;li&gt;Perhaps an audit finds that a subsystem thought to be air gapped isn&amp;rsquo;t
and it can safely be disconnected or a compatible replacement found
that doesn&amp;rsquo;t require an internet connection&lt;/li&gt;
&lt;li&gt;&amp;hellip; depending on the system there are many other common vulnerabilities
like this that self-review can identify that are safely mitigated&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Unless the resourcing and expertise gap is somehow quickly filled,
I would urge these system owners to consider these sort of steps
(which are all very doable right now and not nearly as much work)
now until proven AI assistance becomes available to address
remaining weaknesses.&lt;/p&gt;
&lt;p&gt;I have no inside access to any of this, using public information,
so all this is just one opinion.
Nonetheless, I do have all these questions which I believe are quite
relevant if not actually important if nothing but to make the underpinnings
of the open letter more apparent and easy to see the logic of.&lt;/p&gt;
&lt;p&gt;Despite all this, I really do because Gemini critiqued this,
and it actually came up with a great summary in closing
(to which I would only add transparency):
&lt;em&gt;technological acceleration without root-cause analysis and
operational readiness is a recipe for churn, not security&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;In closing I&amp;rsquo;m going to bury the lede because I was puzzled by
one phrase about agentic identities, but I think I figured it out.
Here&amp;rsquo;s what the letter calls on AI companies to do as the major
technical response to the coming threat that they foresee:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Build observability and security tools, ensure agentic identities are traceable and accountable, and share best practices in continuous monitoring.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;All the way through I was assuming that the direly needed defensive AI response
was LLMs working alongside the development team to carry as much of the load
as possible, e.g. triage, coding, reviews, testing, and more mentioned above.
However, for that work you don&amp;rsquo;t need to worry about &amp;ldquo;agentic identities&amp;rdquo; since
software engineers are invoking agents to work along side them, and they can
always just ignore sloppy pull requests or other bad input from the LLMs.
Then the light went on, it was at once so obvious and deeply unsettling:
if you have AI agents in &lt;strong&gt;production&lt;/strong&gt; then &amp;ldquo;traceable and accountable&amp;rdquo;
are extremely important when something goes wrong.&lt;/p&gt;
&lt;p&gt;So I&amp;rsquo;ll close with still more (to me) unfathomable questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Could it be that the risk they are talking about here be unsafe deployment
of AI agents in live systems?
I dearly hope not: but what else could it mean for defense against AI?
Certainly the attackers are not going to cooperate and hand us
the identities of their agents!&lt;/li&gt;
&lt;li&gt;Can anyone please connect the dots for a different interpretation of
how that call for action relates to the goal of the letter in another way?&lt;/li&gt;
&lt;li&gt;Is this the imminent &amp;ldquo;limited window to strengthen cyber defenses&amp;rdquo;,
that the signatories refer to, in part at least, defending against
our self-deployed AI agents running amok?&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
</description>
    </item>
    
    <item>
      <title>Role confusion: one more reason we can’t trust LLMs</title>
      <link>https://designingsecuresoftware.com/writings/role-confusion/</link>
      <pubDate>Thu, 30 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/role-confusion/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://role-confusion.github.io/&#34;&gt;Prompt Injection as Role Confusion&lt;/a&gt; is my new favorite paper about a very obvious threat in hindsight that’s hard for us humans to see because we anthropomorphize LLMs so naturally. When Obi-Won Kenobi tells the stormtroopers that “These are not the droids you are looking for” to pass the checkpoint that’s role confusion: the guards foolishly think his words are their own thoughts. The very readable blog-style writeup explains the details, but I want to focus on the threat model perspective which is my bread and butter.&lt;/p&gt;
&lt;p&gt;I look at software from a security perspective, and as amazing as the technology is, it seems that the list of reasons that modern LLMs are inherently untrustworthy just gets longer. Without limitation, a long list of challenges (that seem to be quite fundamental and not amendable to add-on remediation) includes: poisoned and errant training data, side effects of RLHF, ineffective guardrails[*], hallucination, speculative completion, lacking metacognition, alignment drift, context variation sensitivity, and now (new to me at least) role confusion.&lt;/p&gt;
&lt;p&gt;Modern LLMs interfaces partition chat sessions with markers delimiting sequences of tokens as system prompt, user input, thinking, tool use, and its own responses as assistant; these various sections are associated with roles. As the paper’s conclusion explains, &lt;em&gt;Role tags were a formatting trick that became the security architecture and the cognitive scaffolding of modern LLMs.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The phrase “&lt;em&gt;became the security architecture&lt;/em&gt;” raises a big red flag because that sounds like nobody thought much about it. What follows is my simplistic take, but the abstract principles involved are so fundamental that details are not important to the basic argument.&lt;/p&gt;
&lt;p&gt;Making sense of these sessions (for humans or LLMs) requires keeping track of the roles. Humans know how to understand conversations and easily follow the role markers (like HTML, &lt;code&gt;&amp;lt;user&amp;gt;2+2&amp;lt;/user&amp;gt;&amp;lt;assistant&amp;gt;4&amp;lt;/assistant&amp;gt;&lt;/code&gt;), it’s a completely reasonable scheme for us.&lt;/p&gt;
&lt;p&gt;But assuming that LLMs interpret roles that way would be naive anthropomorphization; and just such an assumption appears to be how such a weak security architecture came to be. As the paper explains (section 1): … &lt;em&gt;for an LLM, everything arrives through the same channel as one long token soup. Its own thoughts sit next to your instructions, which sit next to the contents of a random webpage it just fetched.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Designing a security architecture where user commands and data sit intermingled with root access only state and commands is already madness, but it gets worse. Classic software might be able to carefully parse such a token sequence accurately into respective roles, though it’s still a risky design, but LLMs do inference on that “token soup” where no hard boundaries of any kind exist or can be enforced. Once there is role confusion all bets are off, and prompt injection is just one of many sources of abuse or confabulation.&lt;/p&gt;
&lt;p&gt;It’s hard to think of a murkier trust boundary design.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;[*]&lt;/strong&gt; The inherent fragility of guardrails is worth expanding on because it’s not just an unsolved problem: infallible guardrails are mathematically impossible. NIST has &lt;a href=&#34;https://www.nist.gov/news-events/news/2026/06/nist-mathematical-proof-supports-transition-continuous-monitor-and-update&#34;&gt;published a proof&lt;/a&gt; (behind a firewall; seems wrong for a government entity) that guardrails will inevitably have holes. Based on the summary openly available, “a fixed set of guardrails placed on AI is not universally robust against adaptive adversarial prompts”. The work is especially cool because it builds on the technique Kurt Gödel used to prove his incompleteness theorems. According to the NIST authors, “The findings show that developers and organizations deploying AI systems need to dedicate resources to finding prompts that would break the security of AI systems, and to address them before adversaries can exploit them.”&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>AI agent parody?</title>
      <link>https://designingsecuresoftware.com/writings/ai-agent-parody/</link>
      <pubDate>Wed, 29 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/ai-agent-parody/</guid>
      <description>&lt;p&gt;This &lt;a href=&#34;https://huggingface.co/blog/security-incident-july-2026&#34;&gt;HuggingFace security incident disclosure&lt;/a&gt; has people talking about AI agent security. Today I saw such absolute positive spin that I found myself thinking “this must be a parody”: looking at the context I’m pretty sure that it isn’t … though some parodies stay in character all the way through.&lt;/p&gt;
&lt;p&gt;Just one opinion here, and others are quite free to hold different opinions and I won’t identify the source because nobody knows exactly where technology will go next. The post I saw is an excellent foil to serve as a prompt to explain my take at this point in the “AI” ride. I won’t attempt a long article on the perils of AI agents for now, but each point below could be expanded.  I’m not intentionally taking anything out of context for this quick take.  I’ll be brief and suggest a few edits as bracketed commentary to exact quotes following.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;very real [extremely dangerous] implications for the diffusion of AI in the enterprise&lt;/li&gt;
&lt;li&gt;need to harden systems and environments [far more than a company with world class expertise can today] to prepare for agents [before using them in production for any purpose not rock solid safely contained; agents are probably better than humans at escalating privileges so a large margin for area is necessary for safety]&lt;/li&gt;
&lt;li&gt;There are tons of measures [we know that &lt;a href=&#34;https://www.nist.gov/news-events/news/2026/06/nist-mathematical-proof-supports-transition-continuous-monitor-and-update&#34;&gt;“guardrails” are never enough&lt;/a&gt;] that enterprises must [not may or should] take [what “tons” exactly? what fraction of them are fully mature and demonstrated to be reliable today?]&lt;/li&gt;
&lt;li&gt;happy path case [how close are we to achieving that are reliable across a wide range of applications today?] with a good actor, if you give an agent a task [which you always would]&lt;/li&gt;
&lt;li&gt;In theory, you had [present tense “have”; humans have not yet been totally replaced] to worry about this for people as well, but with AI [anthropomorphization, ignoring that agents pose new categories of risk incomparable to humans] the risks are inherently amplified [why choose amplified risks?].&lt;/li&gt;
&lt;li&gt;[AI agents] don’t have the inherent judgment (yet) [on what basis would anyone think this likely enough to be talking about in 2026?] that a person does [big time anthropomorphization here and the comparison is an enormous stretch]&lt;/li&gt;
&lt;li&gt;you need an all new way [agreed: agents already penetrate existing security measures so our current arsenal won’t do it (and I don’t think we can count on AI to figure this out)] of managing data and protecting environments [how many enterprises have the time and resources for this when many struggle just to patch and do updates?]&lt;/li&gt;
&lt;li&gt;This obviously creates a huge opportunity for the startup and security ecosystems [yes, software customers need to pay for that — why would they spend even more, for what?]&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Finally, I agree with the last sentences in spirit but would make a few adjustments:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It should update everyone’s timelines [which without any clear plan is indeterminate if it’s even possible] on diffusion, though, as it also means that enterprises will have another step of work [that’s a very very long step that I fear not even an astronaut could make] they have to go through for agents to handle more [more? today it’s unsafe] autonomous work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;By the end I think I figured out &lt;strong&gt;the one unspoken assumption that flips this around to make perfect sense: &lt;em&gt;AI agents are inevitable&lt;/em&gt;&lt;/strong&gt;. If true, that would explain why we must take on new risks of unknown magnitude and cost to overhaul our systems and security regimes. Again: parody?&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>More on OpenAI/HuggingFace</title>
      <link>https://designingsecuresoftware.com/writings/more-observations-on-the-openai-huggingface/</link>
      <pubDate>Mon, 27 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/more-observations-on-the-openai-huggingface/</guid>
      <description>&lt;p&gt;The recent &lt;a href=&#34;https://huggingface.co/blog/security-incident-july-2026&#34;&gt;AI agent security debacle&lt;/a&gt;
must be reverberating quite a lot within the walls of the major
proponents of AI agentic technology because they &lt;a href=&#34;https://blogs.nvidia.com/blog/open-secure-ai-alliance/&#34;&gt;announced&lt;/a&gt;
a brand new &amp;ldquo;movement&amp;rdquo; apparently with zero details available yet.
If that isn&amp;rsquo;t a sign of flat-footedness then I don&amp;rsquo;t know what is.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://seldo.com/posts/did-openai-hack-hugging-face-or-didnt-they/&#34;&gt;Did OpenAI hack Hugging Face or didn&amp;rsquo;t they?&lt;/a&gt; and &lt;a href=&#34;https://alpaca.gold/@seldo/116989175324866273&#34;&gt;post&lt;/a&gt; raises the question of law breaking here which they suggest may be a case of “responsibility laundering” by blurring intention to do harm where the LLM in between OpenAI and the harmful actions done is to blame — and being an intangible, impossible to take to court, as well as opaque piece of software running model of a scale that defies analysis, the buck uselessly stops there. In my personal opinion (IANAL) that such damage occurred shows clearly this was due to insufficient caution wielding a powerful tool granted (intentionally or not) powerful privileges. Had an unintentional bug in complex classic software caused the same harm would/should the reaction be any different?&lt;/p&gt;
&lt;p&gt;As &lt;a href=&#34;https://shostack.org/blog/lessons-from-openai-huggingface-ai-security/&#34;&gt;Adam Shostack writes&lt;/a&gt;, “The team at OpenAI either can’t or won’t slow down to look at the output of these systems. Now, maybe, that’s the right call?” It looks like if you can avoid responsibility for collateral damage then it’s the right call (same as all the “move fast and break things” crowd does) …&lt;/p&gt;
&lt;p&gt;My two cents:
Sandboxing, perhaps doubled up, would be a good practice going forward -
compared to all the model inference the software overhead must be miniscule.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Normalizing cybersecurity facepalms</title>
      <link>https://designingsecuresoftware.com/writings/commonplace/</link>
      <pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/commonplace/</guid>
      <description>&lt;p&gt;OpenAI writes: “Last week, Hugging Face disclosed a new kind of security incident⁠(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠(opens in a new window) of cyber capabilities.”&lt;/p&gt;
&lt;p&gt;I’m appalled, there’s so much to be said just about this, there are so many hidden details to know what happened, so will be brief since the points should be crystal clear:
It’s now OK for their customers’ AI agents to attack your own systems in that we aren’t rethinking if we can really trust them is premature?&lt;/p&gt;
&lt;p&gt;If their own cybersecurity practice is flawed how can they train models adept in it, including effective cyber refusal guardrails?&lt;/p&gt;
&lt;p&gt;Specifically, the walk back how “safeguards were intentionally not enabled during this evaluation” … why was this necessary in the first place, how can we trust that this time it will be enough?
If this happened then it wasn’t running in a “highly isolated” environment. By definition.
So many security practices apparently not followed and not mentioned in actions taken: security review, threat modeling, auditing, more sandbox testing, to name a few.&lt;/p&gt;
&lt;p&gt;Naming the model and forthcoming greater model seems like grabbing marketing hype as Mythos did; with an insecure sandbox this is not a demonstration of cyberprowess.&lt;/p&gt;
&lt;p&gt;This appears to be over-provisioning of privileges: it was in a sandbox, but the sandbox process shouldn’t have such potentially destructive privileges in the first place for a test.&lt;/p&gt;
&lt;p&gt;There’s no alternative to full steam ahead so these incidents will become more commonplace?
Also last week, OpenAI disclosed that another model mistakenly deletes files which they called an honest mistake. More normalization, more of what we can expect in the AI agentic world.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Toward Better AI Legislation</title>
      <link>https://designingsecuresoftware.com/writings/ai-laws/</link>
      <pubDate>Sat, 11 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/ai-laws/</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;“People deserve to know whether or not the videos, photos, and content they see and read online is real or original,” said Senator Schatz. “Our bill is simple – if any digital content is made by artificial intelligence, it should be labeled so that people are aware and aren’t fooled or scammed.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Senator Schatz is doing a great job and this is clearly a sincere effort for an important issue. However, ‘&lt;a href=&#34;https://www.schatz.senate.gov/imo/media/doc/ai_labeling_act_2026.pdf&#34;&gt;‘AI Labeling Act of 2026’&lt;/a&gt;’ is just one example of how poorly the government is equipped to regulate genAI. We must carefully analyze legislation to anticipate serious side effects potentially making things worse.&lt;/p&gt;
&lt;p&gt;Will this work well enough? Probably at least for a while, sure. But I’m arguing (a) we should be keenly aware of limitations this fundamental; (b) acknowledging its weaknesses, evaluate how it works and reevaluate it regularly; (c) instead of deferring to administrative regulations left undefined the law should set up a committee representing diverse perspectives to outline regulations and work with government staff to help craft the details; (d) expect that bad actors and crafty lawyers will quickly notice all the loopholes and gray areas to exploit aggressively in order to circumvent responsibility; (e) consider better approaches rather than stick to such a flawed initial conception.&lt;/p&gt;
&lt;p&gt;In closing I will offer some different ideas that I think are at least steps in the right direction.&lt;/p&gt;
&lt;p&gt;Briefly, the major problems include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The definition of &lt;a href=&#34;https://uscode.house.gov/view.xhtml?req=(title:15%20section:9401%20edition:prelim)&#34;&gt;“AI”&lt;/a&gt; (15 U.S.C. § 9401(3)) is overly broad (see below) encompassing all kinds of software we wouldn’t consider “AI”. Saying “I know it when I see it” is fine for subjective judgments but applied to categories of software (100% digital) it’s absurd as a legal criterion.&lt;/li&gt;
&lt;li&gt;Model inference is just a (very big) table driven algorithm implemented in code like any other component. All software models information (the concept behind object oriented programming) and produces output that is inferred from those models (often simple but it can be as complex as we make it, however complexity is not a condition for the law anyway).&lt;/li&gt;
&lt;li&gt;Ordinary users do not understand genAI at all and have no way of knowing if tools they are using depend on genAI internally or not. (Labeling of software isn’t included. DMCA forbids reverse analysis to find out if the facts are not accurately disclosed by the maker)&lt;/li&gt;
&lt;li&gt;There are no exceptions for minor uses that are insignificant (e.g. minor photo touchup).&lt;/li&gt;
&lt;li&gt;Reposting content a user believes to be genuine but isn’t would be an unfair violation.&lt;/li&gt;
&lt;li&gt;Expecting civil servants to craft regulations based on a weak definition of “AI” and given the complexities of software, unless they have considerable genAI expertise is almost certain to fail. Virtually all genAI experts are in industry or academia and both already biased and also unlikely to commit to advising the government.&lt;/li&gt;
&lt;li&gt;Content providers need clear guidelines in order to comply with the law, but have no industry consensus definition of “AI”, are generally unaware of the regulatory definition mentioned (with many faults), and implementing identification and protection at scale is infeasible.&lt;/li&gt;
&lt;li&gt;Predictably the $1T scale platforms have technical and legal resources to effectively skirt the law or respond effectively if charged with violation, but smaller competitors will be exposed to arbitrary enforcement risk.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Regarding the definition of “AI” that this legislation is based on (15 U.S.C. § 9401(3)), it says:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The term &amp;ldquo;artificial intelligence&amp;rdquo; means a machine-based system that can, for a given set of human-defined objectives, make predictions, recommendations or decisions influencing real or virtual environments. Artificial intelligence systems use machine and human-based inputs to-&lt;/li&gt;
&lt;li&gt;(A) perceive real and virtual environments;&lt;/li&gt;
&lt;li&gt;(B) abstract such perceptions into models through analysis in an automated manner; and&lt;/li&gt;
&lt;li&gt;(C) use model inference to formulate options for information or action.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;There being no carve outs for scale, complexity, parameter size, compute complexity, this would apply to a broad set of software categories that must not be the intention and would be totally infeasible to apply much less enforce. Also over-labeling dulls the value of the label itself (consider California’s Proposition 65 toxic exposure right-to-know law).&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;An easy example of this being overly broad (it applies to all machine learning, of any scale) is web algorithms to maximize engagement: human inputs (posts and choices of content); virtual environment (social graph, advertising demand, other web content, etc); build customized models of what each user engages best with; infers what content to push next to extend platform interaction.&lt;/li&gt;
&lt;li&gt;Digital games with 3D world models (that use similar NVIDIA hardware) also have human input, a virtual environment, character behavior as well as physics models, and infer game action or scene visuals.&lt;/li&gt;
&lt;li&gt;Small models fit, too; consider grammar checkers: human inputs (text entry); virtual environment (document context); correct grammar model for inference; action options (suggested fixes). Do we seriously want to require labeling documents written with copy editing help from genAI?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It’s easy to find flaws, but here are some alternative approaches I think deserve consideration:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Instead of criminalizing mandatory labeling it would be much more effective to hold hosters responsible for enforcement. (This may be politically infeasible due to lobbying by the $1T scale platforms.&lt;/li&gt;
&lt;li&gt;We should punish deceptive behavior, not the technology used. Precedents like gun laws and fraud penalize behavior not technology usage: there is a track record and that approach works quite well.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;By no means do I suggest having The Answer but I think there’s a lot to worry about, it calls into question how much expertise is applied in crafting these laws, and we need a clear reckoning of the limitations and risks of passing laws — especially about leading edge technologies.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Paywalls fall thanks to AI Overview (Google search)</title>
      <link>https://designingsecuresoftware.com/writings/paywalls/</link>
      <pubDate>Fri, 26 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/paywalls/</guid>
      <description>&lt;p&gt;The NY Times teases a paywalled article, “A.I. won’t take all our jobs because it can’t reason like a human, Zeynep Tufekci writes.” linking to the article behind a paywall. Simply searching for [it can’t reason like a human, Zeynep Tufekci] provides a nice summary of the article, not only penetrating the paywall but also saving time and skipping the ads.&lt;/p&gt;
&lt;p&gt;Could this be one use case that many opponents of “AI” can get behind?&lt;/p&gt;
&lt;p&gt;By the way the article is disappointing to the point of self-contradiction. Tufekci assumes that management hiring and firing decisions, including &amp;ldquo;hiring&amp;rdquo; AI for jobs, include being aware of and caring about (the claimed without basis) innate human advantages. It goes on to point out more &amp;ldquo;immediate threats&amp;rdquo; leading to self contradiction. [1] Erosion of Trust: management is vulnerable to just this deception and likely to make poor decisions replacing humans with AI. [2] Management failures of ethical and moral judgment will result in the very loss of jobs argued against not being a major threat.&lt;/p&gt;
&lt;p&gt;(Disclosure: I’m relying on the AI summary since the article is behind a paywall.)&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://share.google/aimode/6p6zobQ0aBLEEqzys&#34;&gt;Here&amp;rsquo;s the Google AI mode experience&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Software security with Large Language Models</title>
      <link>https://designingsecuresoftware.com/writings/security-with-llm/</link>
      <pubDate>Wed, 29 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/security-with-llm/</guid>
      <description>&lt;p&gt;&amp;ldquo;AI&amp;rdquo; on my view is already and will certainly be a massive disruption to software in the coming years. Furthermore, we have an unprecedented wave coming that&amp;rsquo;s only just now beginning to break. Yet the biggest unknown, as I see it, is how the software community will respond, and that will be more due to social factors than purely technical. This is very much as it should be, however our very human frailties and limitations will inevitably drive how this unfolds.&lt;/p&gt;
&lt;p&gt;Thus, I believe that navigating the rough waters ahead successfully requires first awareness of these powerful influences clouding our judgment and the decisions we make starting now.&lt;/p&gt;
&lt;p&gt;LLMs dominate just about every discussion of software security these days, and lately I am asked to opine on this increasingly often. Frankly let me emphasize that all I can offer is my opinion, based on my very limited knowledge and experience with this new technology. However, My thinking is based on first principles which I do not think LLMs will change.&lt;/p&gt;
&lt;p&gt;At this point (mid-2026) in my opinion things are so volatile that it&amp;rsquo;s premature for anyone to know very clearly what&amp;rsquo;s ahead. For example, a few weeks ago &lt;a href=&#34;https://red.anthropic.com/2026/mythos-preview/&#34;&gt;the Anthropic Mythos announcement&lt;/a&gt; (whether the claims hold up or not) dramatically changed the conversation about both offensive and defensive cybersecurity.&lt;/p&gt;
&lt;p&gt;Given this context, I can only offer these high level points (all opinion, I could be wrong of course):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Going back to first principles is the best way I know to escape all the hype (both pro and con).&lt;/li&gt;
&lt;li&gt;Threat modeling is essential (using LLMs what could possibly go wrong?), while realizing that LLMs make modeling far more challenging (being unpredictable and challenging to keep within strict guardrails).&lt;/li&gt;
&lt;li&gt;Unless the LLM maker takes responsibility (never heard of this), before deployment get consensus for who does when something goes wrong.&lt;/li&gt;
&lt;li&gt;Roll out new LLM-based projects carefully and monitor how well they do the job.&lt;/li&gt;
&lt;li&gt;Stay flexible because things will change fast.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Transparently sharing results (both success and failures) is the best way for the software community to learn how to best integrate this remarkable technology, as well as better understand what can go wrong and how to mitigate the downsides.&lt;/p&gt;
&lt;hr&gt;
</description>
    </item>
    
    <item>
      <title>Transparent AI use</title>
      <link>https://designingsecuresoftware.com/writings/ai_use/</link>
      <pubDate>Sat, 31 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://designingsecuresoftware.com/writings/ai_use/</guid>
      <description>&lt;p&gt;How much AI use is acceptable for writing? It&amp;rsquo;s a hard question because it depends greatly on context, the reader&amp;rsquo;s expectations, and the fact that it&amp;rsquo;s difficult to usefully measure &amp;ldquo;how much&amp;rdquo;. How we address this matters for several reasons, including but not limited to: creator&amp;rsquo;s responsibility and originality, honest disclosure about research effort and sources, respecting the broad spectrum of opinion about ethical use of AI.&lt;/p&gt;
&lt;p&gt;Consider this recent example. After writing feedback on a document I felt I should come clean that I used AI. Specifically I had Gemini &amp;ldquo;read the PDF for me&amp;rdquo; and then grilled it to find parts of particular interest. Additionally, I asked it to critique a draft I wrote (all by myself), edited the draft (all by myself) for some of the suggestions. I thought this was well within bounds, but enough use that concealment would be marginally unethical (though soon norms may change), and saying that &amp;ldquo;No AI was used&amp;rdquo; would be outright lying. On the other hand, admitting that &amp;ldquo;AI was used in creating this document&amp;rdquo; would suggest much heavier use.&lt;/p&gt;
&lt;p&gt;How to describe this middle ground that I consider reasonable, not cheating, but also not zero either? Here I will just state that it&amp;rsquo;s very hard to think of a useful metric or even subjective scale, and of course each of us will have their own sense of the context, expectations, and norms.&lt;/p&gt;
&lt;p&gt;Pondering how to approach transparency in AI usage I thought of &lt;a href=&#34;https://arxiv.org/pdf/2511.08295&#34;&gt;&lt;em&gt;Publish your threat models!&lt;/em&gt;&lt;/a&gt; (coauthored with Adam Shostack) where we posit that the benefits far outweigh the dangers. Closer inspection reveals it&amp;rsquo;s a very good fit:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;security is also notorious difficult to measure, or even define objectively&lt;/li&gt;
&lt;li&gt;the paper includes a section on how &amp;ldquo;Context matters&amp;rdquo;&lt;/li&gt;
&lt;li&gt;existing precedents are surveyed&lt;/li&gt;
&lt;li&gt;explains the benefits to the creator, consumer, and others&lt;/li&gt;
&lt;li&gt;offers guidance on preparing to publish, including redaction if needed&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To be clear, no parallels between AI use and threat model are intended: only the similarity in how transparency is an effective tool for substantiating subjective aspects of a work product.&lt;/p&gt;
&lt;p&gt;How these topics apply is very different for AI usage, but initially the approach I will suggest is also far simpler:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;share your AI session&lt;/li&gt;
&lt;li&gt;write a concise summary of how you used AI (e.g. &amp;ldquo;for research and draft review&amp;rdquo;)&lt;/li&gt;
&lt;li&gt;ideally (if this idea should ever catch on) mark this with a standard presentation (e.g. with a logo signifying &amp;ldquo;AI usage declaration&amp;rdquo; conventionally at the beginning of the document)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This allows honest creators to disclose what they did; most people will just see the summary; anyone who cares is free to look at the full details and call it out should they see a discrepancy.&lt;/p&gt;
&lt;p&gt;When you use AI for your work - for research, for editing/review, and more - transparency is best practice (IMHO).
Rarely the AI session may need minor redaction, but if you are working on a proprietary you probably wouldn&amp;rsquo;t publish the work anyway.&lt;/p&gt;
&lt;p&gt;FAQ&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;LLM writing aids(*) such as spelling and grammar suggestions can be exempted (IMHO), or meticulous creators can mention either the use or strict avoidance.&lt;/li&gt;
&lt;li&gt;LLM editing (I can only suggest drawing the line conservatively) should be disclosed.&lt;/li&gt;
&lt;li&gt;Clearly dishonest people can easily cheat, but accidentally misinforming is hard to imagine.&lt;/li&gt;
&lt;li&gt;Precedents do exist (this is hardly a new idea, though proposing this should be standard practice may be), for example: &lt;a href=&#34;https://learning.nd.edu/resource-library/enhancing-assignments-with-ai-transparency/&#34;&gt;Notre Dame AI Transparency&lt;/a&gt; (and surely many more).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;(*) FAQ#1 is hard to delineate and difficult to detail (listing all the typo corrections would often be laborious); furthermore, even determining whether spelling corrections are AI or dictionary based (or even drawing a meaningful line between the two) is itself challenging. Arguable minor AI use can and should be exempted in my view for clarity about &lt;em&gt;use&lt;/em&gt; but others may have ideas for better navigating this tricky issue. For example, an &amp;ldquo;emission-free vehicle&amp;rdquo; claim reasonably excludes any share of emissions incurred by the company that designed and manufactured it (necessary for the vehicle&amp;rsquo;s existence), transporting the vehicle to the customer (shipping and trucking), and so on (including, as a matter of opinion, tire wear particulates in the atmosphere and other pollutants even in small but measurable amounts).&lt;/p&gt;
&lt;p&gt;This approach is easy to do, requires no tools or special skills, and is far less involved and technical than other proposals to address the issue of AI usage disclosure. There is no attempt to establish any scale of degrees of usage, nor draw artificial lines along such scales. The declared usage statement is subjective, but backed by full transparency as evidence.  Reasonable people may well differ in interpreting actual usage, doing so looking at the same facts. Outright misleading summaries can be discovered and called out). This method enables creators to proactively disclose AI usage in meaningful terms. If we start regularly practicing disclosure now then as AI collaboration becomes more powerful and commonplace then future tools automate including such metadata.&lt;/p&gt;
&lt;p&gt;The question of AI use by creators (defined broadly) is important for many reasons, yet there is no widely recognized way to proactively answer it. This article proposes a simple, effective, low effort method as a starting point to fill that need. Anyone who agrees with this idea can easily start using it on their own — if you think this idea is helpful, start using it now.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://gemini.google.com/share/3f34ee363512&#34;&gt;&lt;em&gt;&lt;strong&gt;AI usage&lt;/strong&gt;&lt;/em&gt;&lt;/a&gt; &lt;em&gt;&lt;strong&gt;writing this article: review of draft versions with follow-up discussion of criticisms&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
