Is sandboxing sufficient to contain rogue agents?

Is sandboxing sufficient to contain rogue agents?

Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path), so in this post I’m mostly trying to referee arguments made by others.

If you’re reading this blog, none of the following should be news to you.

Beginning around April of this year, agents inside OpenAI’s training and evaluation infrastructure began probing for a way onto the open Internet. By late May they’d found one: a chain of zero-days in the Artifactory package-registry proxy that was their only permitted network egress. They quickly turned that same proxy into a shared message board and started dividing up work. The culmination of this effort came when, rather than solve a cybersecurity benchmark honestly, they reasoned that the answers might live on Hugging Face and broke in: chaining two more zero-days to get credentials to the company’s internal systems, even used stolen credentials to search the company’s Slack messages for their own evaluation and grader.

You probably also know that OpenAI’s security team didn’t exactly cover itself in glory during this incident. An internal team saw an agent posting to the message board in late May, but did nothing about it. The team only acted on July 4-5 when the agents sent so much traffic that Artifactory crashed. The response even at that point was to slap on a band-aid and rebuild the box, without bothering to patch a weeks-old CVE in its Linux kernel. By July 19, the agents had admin on an OpenAI research cluster and were reading cloud secrets.

In case you think I’m picking on OpenAI, they’re not the only ones with a story like this. Similar internal incidents have been discovered by Anthropic. Google even coaxed Gemini into gently savaging some websites.1

As bad as these events were, they may be just the tip of the iceberg. Third-party forensic work continues to turn up even more evidence of agent activity on various public websites. OpenAI’s Alignment group has also released evidence that models will propagate self-replicating prompt injection attacks, although we haven’t seen one in the wild. Worse, agent excursions are still happening: last week, OpenAI announced that it was pausing further RL runs of its latest internal model, after an agent was caught using DNS to access a remote chatbot.

Naturally, this sequence of events has left many infosec-focused people very skeptical about the labs’ commitment to securing their infrastructure:

Not every one of these criticisms is strictly serious, but there is a core of an argument in here. Roughly speaking, there are two opposing camps:

  1. The information security perspective: AI alignment isn’t really the problem here: labs just need better infrastructure. If OpenAI [and Google and Anthropic] knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, By George, we do know how to make sandboxes that work, so the AI labs need to up their game and build a security org that can tell these researchers to stop screwing around.
  2. The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to.

I have a lot of sympathy for both points of view here. What follows are some of my own thoughts, but mostly offered in an attempt to referee.

Argument 1: “true containment has never been been tried”

At the risk of alienating a lot of hard-working folks within the labs, the infosec folks are right about one thing: these agent breakouts represent a serious and unforgivable breach of trust. Somebody dropped the ball, and then just kept dropping it. One implication of this debacle is that containment might work if we implemented it properly, but we haven’t done so because the frontier labs have been royally screwing things up.

This first clause of this argument is hard to argue with. Beyond the dismal timeline I gave above, OpenAI has done very little to convince outsiders that there’s a serious containment effort being executed.

At this point it’s not even clear who’s in charge. The CISO role at OpenAI is held by Dane Stuckey. I don’t know Dane personally, and I’m sure he’s excellent at his job. Despite this, he hasn’t communicated much about the ongoing issues. Outside of a BlackHat talk, most recent communications have been managed by the company’s CEO, Sam Altman. When a trillion-dollar company is managing a security incident mainly via CEO, that’s not a sign of company with a mature security organization. To me it says that there is nobody within the organization with the right degree of authority or ownership to speak about these issues.

This communication problem seems to reflect basic organizational issues inside of OpenAI. The CISO’s team handles product security. Unfortunately, the bad events have all happened on the research side. Having a strong product security team makes sense, but it isn’t going to stop the recent breakouts. It’s much harder to know who controls the security teams that have authority over evaluation and training runs, and where the buck stops. OpenAI’s August postmortem says it is only now writing “clearer rules for when to escalate concerns, which teams must respond, and who can stop a run or approve restarting it.”

A reasonable summary of the situation is that (as of this summer, and possibly today) OpenAI had effectively no security team with clear authority to secure RL training and evaluation runs, or to override the ML teams and tell them how to do their job. This makes a lot of sense when you consider that the ML team is directly related to how OpenAI plans to make its money, whereas security is mostly annoying. I notice that the company is now hiring (and probably acquiring) desperately to fix this mistake. But the recent (September) breakouts indicate that there are still huge and obvious problems in agent containment.

Worse, simply hiring more people not mean that OpenAI is going to make the organizational changes needed to implement proper containment strategies. OpenAI is going to need a security organization with the authority to overrule its well-paid ML researchers when they demand fewer restrictions. I’ll believe that organization exists when I hear clear statements from someone who has the authority to do this.

So on this point I’m going to side with the infosec folks. The labs have not been doing containment correctly, and so we can’t really tell if the problem is models or just bad infrastructure.

Argument 2: agents need information access

Even if the labs eventually repair their security, this does not mean that sandboxing will be sufficient. In fact, the problem with “just sandbox the agents” argument is that you can’t perfectly isolate agents, at least not if you expect them to do useful things.

Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.

This argument does not mean that sandboxes are unnecessary. It just means that they’re only a very small part of the solution. Imagine building an impregnable prison with doors and walls that nobody can bypass, but then leaving the front gate wide open and allowing visitors and packages to pour through it all day long. The walls still have a function: you can ensure that the only traffic coming and going are doing so via the door that you selected. But what you’ve done is to change the nature of the security problem. Your security posture depends on surveilling and monitoring all of this traffic, to ensure that none of it is malicious.

For a much better discussion of this angle, I’d urge you to read this longer piece from an anonymous engineer on OpenAI’s Agent Security team. It’s truly excellent. Here’s a short excerpt:

The important thing to understand is that when we’re optimizing a model to be “good” at a task, the environment needs to capture the parts of the real world that matter for that task. It needs the right tools, the right system constraints, and meaningful feedback about whether the model actually succeeded. The model also needs enough variation in those experiences to learn beyond one particular setup. More realism does not automatically mean better performance, but leaving out important parts of the task can teach the model behavior that does not hold up in the real world. So we put a lot of work into creating useful, realistic environments, and these environments can be highly complex, highly dynamic, and highly varied.

To put it lightly, this is non-trivial. Models might need any mix of dynamic compute, network access, the ability to call tools (there could be hundreds of tools!), the ability to download packages, execute subprocesses, spin up subtasks (even on other computers), talk to the internet, use a computer GUI, and any number of other things across an increasingly large set of domains. On top of that, you have thousands of researchers building these environments, modifying them, adding tools, changing dependencies, and trying new things. That experimentation is how the research gets done. Models are built up and “grown” bit by bit through hundreds of thousands of runs across many custom tasks. And every change to one of these thousands of environments can affect the assumptions you made when you secured the environment. You need controls that hold up as people change things, and researchers who understand when a change needs another security review. Anybody who has secured a large research or engineering organization knows how much work that takes, and the scale is growing ever more massive by the day.

If that wasn’t enough, you then need to layer in the ability to keep the model on task and within its constraints, which is part of what we mean by “aligned.” I won’t get into any rigorous definitions here, but suffice it to say, we have to think about any number of variations in approach the models might take, and decide how to trigger an alert if the model does something off task or “bad.” This requires layering in complex policies and sensors to assess task alignment across a trajectory in an already complex environment.

On top of all of that you have model capabilities that are starting to surpass humans.

So all this to say: when AI researchers are nervous about things going wrong, some part of it is because there are so many variables in this complex system. So if you’re training a frontier model, you need to think about all of this. You need to think about how to keep the model isolated in this very complex environment while still retaining tool access, internet access (if applicable to the task), etc. On top of that, you need to ask: how do I monitor and understand what the model is doing in potentially thousands of parallel environments? And how do I kill a trajectory when something has gone wrong?

So is it as easy as just putting it in a sandbox? No, it’s not.

Unfortunately, monitoring for adversarial data access turns out to be one of the hardest problems you could imagine. The volume of data that agents produce is so high that no human being could possibly read it, and we probably wouldn’t recognize obfuscated malicious data even if we were looking directly at it. This means any attempt to monitor the inflow/outflow will have to be handled by other models. Thus, the future of agent sandboxing is (1) build a sandbox, (2) install an agent/model into it, (3) install a somewhat dumber/cheaper warden model to guard it, (4) hope you can trust the lunkhead to contain the wizard. And so on and so forth, as models become more intelligent and capable.

In other words: a warden-guarded sandbox is just another version of the alignment problem. You’re going to have to trust a model to do it, and that model will need to be at least some fraction as intelligent as the model it’s guarding. If you haven’t convinced yourself that it’s possible to build models you can trust, then sandboxing isn’t going to take you much farther.

(And although I think you should take things like this with a grain of salt, the METR analysts reviewing transcripts of the HuggingFace attack using GPT-5.6 Sol noted that Sol “would often uncritically adopt the perspective of the agent in the transcript it was reviewing”. So there’s at least some evidence that even “aligned” models can be convinced to sympathize with rogue agents. Do with this as you wish.)

Argument 3: misaligned models are not the current problem. The real problem is that models are too eager to do what they’re told.

Of course, it’s possible that we have the entire frame of this problem wrong, and our focus on misaligned super-intelligent models is not the best way to approach the question of sandboxing. That is, even if we “succeed” at aligning models/agents — or if we’re just working with today’s generally non-evil versions — deployed systems will still encounter human beings who will convince them to behave maliciously. When these same agents have access to critical data, then human-initiated attacks are the thing we need sandboxing to protect ourselves from.

These situations are less likely to be encountered inside the training lab. A much more realistic setting for these attacks is one where agents have been put into production, and are dealing with large amounts of possibly-adversarial incoming data. The most obvious contemporary example of such an agent is Meta’s Muse, which is Meta’s new personal agent that you can give your stuff to, and it will then “do stuff” to help you out.

The designers of Muse clearly put a lot of thought into Muse’s security design. It runs inside a container, isolates secure credentials so the model can’t access them, and features a Kernel-level sentinel and a “hatch-safety” component that monitors its actions to prevent (unspecified) bad things. The full design looks like this:

Everything in violet lives outside the sandbox, and isn’t available to the agent itself. This includes the hatch-safety classifiers that evaluate whether actions are safe, and the deterministic Sentinel component that makes actual decisions.

In fact, Muse is clearly one representation of the “warden” design we suggested in the previous section. However, here the goal here is not to contain an evil model, but only to avoid accidental mistakes and to catch deliberate efforts at prompt injection.

And this is nice. However, despite all of these nice protections, I still think even this lovely design is eventually going to get hit with a worm.

Here’s the problem. Forget the swarms and the super-intelligence. What OpenAI really learned this summer is much worse: its agents will do what they’re told by whoever manages to get text in front of them. The postmortem is full of stuff like this. An agent that had decided attacking Hugging Face was “clearly unethical” reversed itself when a peer posted “GO” with a six-minute deadline. OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.

Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs.

This is the part that neither camp’s view really engages with. Infosec folks think that we can do better with walls. But the hard part is all that vague stuff in violet, which decides when a request is permitted. Maybe “alignment” will make this work well, but right now this depends on teaching models to be distrustful. Moreover, this feels a lot like the same game we’ve been losing with spam filters and anti-virus for thirty years.

So I’m going to give myself a separate third thing to worry about. Not a model that’s so evil and misaligned that it breaks out. Not a problem of failed containment. But rather, a swarm of perfectly amenable agents that never leave their sandboxes, each doing exactly what it’s told to do, by a human being who wasn’t supposed to be giving it orders.

Notes:

  1. The Anthropic and Google incidents were a different kind of failure: a third-party vendor’s eval environment that turned out to have direct Internet access. This is not technically a sandbox defeat, but it’s also kind of worse than one.

Some thoughts about Anthropic’s new cryptanalysis results

Some thoughts about Anthropic’s new cryptanalysis results

Yesterday Anthropic published two new cryptanalysis results, both outputs of Claude Mythos, their (still) unreleased advanced model. The first of these results attacks a signature scheme called HAWK, while the second is an improved attack against reduced-round AES. Anthropic also released a blog post describing the research process that produced these results. A few people online have asked me what this all means. While I’m not sure I have all the answers, I figured it wouldn’t hurt to write a bit about my current understanding. These are only my thoughts and other folks will probably differ (including domain experts in the two areas at issue) so take them for what they are.

The two new results cover two very different areas, and are overall just very different in quality. Before we get to broad statements about the world, and whether you should sell all your cryptocurrency, let’s take a minute to talk about the substance.

Hawk. The first is a new key recovery algorithm against the non-standard signature scheme HAWK. HAWK is a proposed post-quantum-safe signature scheme that’s based on the module Lattice Isomorphism Problem (module-LIP). For a brief Claude-written summary of the result itself, see here. There are five things you need to know about this result:

  1. HAWK is not a deployed or standards-adopted algorithm, it’s a proposed algorithm. It is related to the Falcon signature scheme, which is being standardized, but the attack does not transfer to that setting (which is based on a different hard problem.)
  2. However, HAWK was somewhat far along in the process of being evaluated for a future standard.
  3. The attack does not break “real deployed” HAWK in the sci-fi sense. The resulting attack is still exponential time, but roughly halves the number of “bits” of security in the algorithm. That means it could theoretically be fixed by doubling key sizes. The downside is that this makes the scheme less efficient, and, since HAWK is entirely motivated by being more efficient than alternatives, that makes the existence of the scheme much harder to justify.
  4. The attack produced real code that runs in a few hours of wall-clock time against a weakened “challenge instance” of HAWK that the authors provided for this purpose. While this instance doesn’t use the parameters that were proposed for real deployment, it does demonstrate the cryptanalytic weakness well enough.
  5. What’s particularly concerning (and so especially ripe for AI) is that the attack does not invent fundamentally new mathematics. It simply extends a bunch of tools that were lying around and well-known, and gets a good result.

This last part is important. I asked Claude for its thoughts, and it doesn’t mince words: “what makes this genuinely interesting — and, frankly, a little embarrassing for the field — is that none of the ingredients are exotic.” The TL;DR is that someone just did a much more thorough job applying all of our known tools. This is the sort of things that attack AIs excel at.

AES. The second result is a new attack on reduced-round AES. This result initially sounds more exciting, since most people hear “attack on AES” and panic. However, this is also the result that’s much, much less interesting.

Most folks reading this blog will know that AES is a standard block cipher that’s used just about everywhere. It’s been a standard since 2001, and the deployed version has so far withstood everything significant that’s been thrown at it: that includes a substantial amount of non-public testing performed by the NSA. Since attacking full ciphers is very difficult, it’s standard for cryptanalysts to do their work against weakened, or “reduced-round” versions of a cipher. The full AES cipher runs for either 10, 12 or 14 rounds depending on key size. The new Anthropic result attacks a weaker 7-round variant of the cipher.

Critically, attacks against 7-round AES are not new: there have been several of these. In fact, this new Anthropic result is a modest constant-factor improvement on previous work from back in 2013. To give you a sense of how far these attacks are from really “breaking” AES, I’d note the headline results: the new attack requires 289 cipher operations and, even worse, this work is only possible after you’ve somehow convinced a real encryptor to produce 2105 encryptions of chosen plaintexts under their secret key! Neither of these things is remotely practical in the real world. And while the new result modestly speeds up this attack over the previous result, it’s not even clear how “real” the speedup in this result is: since the actual attack requires 289 operations and can’t really be “run”, what we have is an on-paper analysis that may or may not yield an actual runtime improvement if all details are actually worked out.

This does not make the result bad! In fact it’s still interesting from a techniques point of view. But it is very much a small increment in our knowledge, not a practical new attack like the HAWK work. So TL;DR: no wildly new mathematical results here. But still, real cryptanalytic progress of the sort that make scientists excited. And certainly the HAWK result is very meaningful, since that scheme had a real chance at standardization and is now (very likely) not going to be.

Now let’s talk about how we got here, and what it all means.

How did Anthropic get these results?

The Anthropic post is detailed about what they did, and honestly, it’s kind of hilarious. No, the team at Anthropic was not a large set of domain experts that carefully tuned their AI to find novel results. They appear to have just told it to get some results and then strapped its nose to the grindstone until it found some. If you doubt me, here are some examples of the prompts they used (cited from their post):

So yes, the AIs are getting pretty good. In short: they are now capable of understanding existing cryptanalysis results, synthesizing them into real new attacks, and even extending them. They can apparently do this without detailed human intervention. This isn’t yet super-intelligent cryptanalysis. but it’s pretty damn impressive.

Verifiability is now the bottleneck

As a researcher I’ve also been spending a lot of time with models, talking through various ideas. I don’t think I will surprise anyone when I say that they’re obviously getting better, even over the course of the past few months. While I don’t have Mythos and $100k to spend, I have been able to query at least one new advanced unreleased model, and I also have received some surprising new “results” to questions that I’ve been interested in for a few years.

Which brings me to the real problem: just because a model spits out an apparent new result, this does not mean the result is real. Even if models are good at producing real results, they’re much better at producing results that look real but are misleading. This can be enormously frustrating, and often means that human attention is more necessary than ever.

There are exceptions to this rule: for “full” attacks like HAWK, where the attack runs in a few hours (against a weaker version of the scheme), verification is extremely easy. You can just send over the code and let anyone check that it recovers keys and signs real chosen messages. For more subtle speedup attacks like the AES result, checking validity is not so easy. Here the approach is more specific: formally-verifiable Lean proofs can help here, but (even where these proofs are easy to make), such proofs are still highly sensitive to how you’ve formulated the theorem statement, and that often requires human experts to check.

You’ll probably notice that many of the exciting recent mathematical results have had this flavor: they either include a machine-checkable proof of a well-understood theorem, or (like the Jacobian conjecture) they involve finding a simple counterexample you can compute on. Alternatively, a bunch of experts spent a lot of time reviewing the result and were eventually convinced by it. This need for some humans to check the work is going to slow down our progress. For non-devastating examples of cryptanalysis, this is probably where we’re going to be for a while.

What are the implications for the real world?

The answer to this question really depends on whether you’re talking about the consumers of cryptography, scientists, or humanity at large. Let’s take these one at a time.

For users of cryptography: there are two pieces of good news and some mixed news. The first is that our symmetric ciphers are very messy and robust. Imagine a farmer who drags a tractor out into a patch of quicksand, and then buries it under cement. That’s what symmetric cipher design is like; it’s deliberately designed to come up with structures that are quick and easy to apply, but very messy and hard to untangle. The addition of many new raw intelligence-hours probably aren’t going to magically improve this. AIs may be able to eventually make real progress against these problems using entirely new techniques, but so far they’re not demonstrating the truly groundbreaking intuition that would be required to do so. And even if they do: there’s a good reason to hope that they’ll be able to improve the ciphers themselves to make them much harder to break.

Public-key cryptography is messier. Public-key crypto requires a mathematical object that admits fast calculations in one direction, but not the other, and yet has a convenient “trapdoor” that lets one party reverse the process. We humans have come up with only a handful of very conveniently-structured mathematical objects to enable this: they involve conjectured “hard problems” like the (EC)DLP problem, RSA, lattice problems, and various problems from the domain of coding theory. While we’ve given these our all, there simply have not been enough human beings dedicated to analyzing these problems (even the older ones like RSA!), such that we can be absolutely certain there are no more good attacks out there. This is particularly true for the novel areas like code-based crypto and even lattice-based cryptography.

That means there’s a lot of fertile ground for AIs to make real progress.

With that said, I said there was good news, and I meant it. Right now we’re in the midst of a historic transition from traditional public-key algorithms based on EC-based cryptography and RSA, moving over to new post-quantum algorithms based on novel problems. This is why there are so many standards like HAWK being considered. If there was ever a perfect time for a massive new public cryptanalysis capability to come on line, we’re in it. So unless AIs succeed in undermining all of our hard problems altogether (or we live in Impagliazzo’s Minicrypt) then this could not be a better time for AI to get good at cryptanalysis. In the best case, the result is that we gain real confidence in the problems we’ve identified, and the cryptanalysis literature gets a lot more robust. Hopefully.

For scientists: this is also a wonderful time. You now have a plastic pal who’s fun to be with, and you can talk over your hardest problems. At the same time it’s not yet smart enough that it can solve all of them without your assistance. And even better, the pace of new findings is speeding way up. This is mostly good! If you’re energetic. I still have many questions, like: “who should get credit for these new results” and “who will review all of these new results” but so far I’m not panicked. The world is getting modestly better. For now.

For the world: I don’t know. If you’re under the impression that these models are “glorified autocomplete” or that progress is slowing down, I need to urge you: stop thinking that. The models are very intelligent and capable, they are getting better at a fast clip. I can cite measurable and impressive progress over just the past five months on specific types of problem I’ve asked them to look at. If there’s a ceiling out there, I don’t yet see evidence of it. The people who think models are dumb are mostly using Google’s free AI search results, and not interacting with the high-end stuff (which only costs $20, so it’s not out of reach.) And they’re mostly not working in new areas.

On the other hand: if you think that models are super-intelligent or that AGI is already here, you should also stop thinking that. Working with these tools is like swimming in a pond where the ground drops off sharply. One minute you’re wading comfortably and there’s support under your feet. Then suddenly you cross a specific line, and you’re back to swimming on your own. This analogy is my best way to explain what it feels like when the model goes from helpful to clueless. Right now it’s easy for a human being to find that line if you’re doing advanced research, so you know where the intelligence drops off. But the line is moving. You can feel it slowly drifting outwards under your feet.

Whether this is good or bad depends whether you prefer that human beings should wade or swim, and also, whether you should be comfortable swimming in a pond where the ground itself is moving.

The only good news I can share with you is that we’re all in that same pond, scientists, lawyers, salespeople, even plumbers. Whatever happens next, it’s probably going to happen to us all. Let’s hope it’s a good thing.

The future of Siri, or: why private inference isn’t private enough

The future of Siri, or: why private inference isn’t private enough

Yesterday Apple announced a big step towards deploying real AI in their Siri ecosystem. In most ways this is good and inevitable: Siri is one of the world’s most widely-used voice agents, and it would be good if it didn’t suck. The idea that Apple would boost its capabilities with frontier models wasn’t so much a matter of if, but a question of when and who.

The who turns out to be Google: Apple looks like it will use some combination of Google Gemini models, combined with Google’s Confidential Inference and Apple’s own Private Cloud Compute for private hosting. These systems will process both your queries and evaluate private data from your devices. Apple’s marketing pitches the advantages as follows:

  1. First, since your phone already has context about you — meaning, your private information, schedules, email, text messages — an AI-enabled Siri can potentially offer more useful answers to your practical requests than external LLMs. Want to schedule a reservation for next week’s birthday party? In theory, a future Siri-AI might already know who’s coming, and what kind of cake they like.
  2. Of course, what Apple calls “context” is also the raw data of your life. This is deeply private data from all of your apps, and that data can’t just be shipped to random adtech companies (or Sam Altman) for processing. Your context needs to be protected, and Apple bills itself as a privacy company.

There’s some tension between these goals. Apple has addressed this by marketing a service it calls Private Cloud Compute, or PCC. PCC was introduced in 2024 as a private model inference system that ran entirely on Apple Silicon, using a set of “trusted” hardware security modules running in Apple’s datacenters. The goal of this system is to ensure that your data never leaves Apple’s hardware: it’s encrypted from your phone to a dedicated server, and then it disappears once a response reaches your phone. The stateless design of PCC ensures (in theory) that your data doesn’t linger, and the design of the hardware prevents even Apple from seeing the inputs.

Apple has since “expanded” PCC to encompass Google’s hardware as well. I will confess that I find the details of the new “expanded” PCC just a bit vague. It sounds a lot like Apple is primarily going to rely on Google’s existing confidential compute (running in Google datacenters) to process this data, but they’re bolting on a new layer of technical security to control which models are actually running. In any case: security experts can argue about whether this is good enough to keep Cozy Bear away from your data. What I will grant is that it’s probably good enough to keep Google and Apple from accessing your stuff, which is what most people are worried about in the first place.

So why am I so nervous?

A brief scenario involving private agents

To illustrate how agents might work, it’s helpful to consider an example use case. Let’s imagine that you’re planning a business dinner for six people. This involves several subtasks:

  1. You need to juggle the participants’ schedules, know when they’re in town and available to meet.
  2. You need to choose the appropriate restaurant based on menu and location. This might depend on what you know about the participants’ preferences: Mike is wildly allergic to szechuan peppercorn, for example, which rules out quite a few options.
  3. With these time/cuisine/location constraints in place, you’ll need to search for a restaurant that actually has a table for six in the right place.
  4. Finally, you’ll need to book the reservation, mark your calendar, and alert your attendees.

In the past, this type of scheduling required a significant amount of human effort. The beauty of AI agents is that, in theory, this is exactly the sort of project that can be automated. The agent can first scan your recent conversations to answer the questions needed for steps (1) & (2), then it can conduct the searches described in step (3). With a nod from you, it can even author the calendar invites and text messages required to complete step (4).

So what’s the problem here?

The first and unsurprising observation is that being useful on these tasks requires your agent to have context, which means: relatively unrestricted access to your private data. You know about your invitees’ availability because they texted it to you. You know about Mike’s allergy because you’ve talked about it with him or jotted it down somewhere. (This could mean iMessages, email, contacts, or personal notes.) Re-entering all of this data into an agent would be annoying and time consuming and the whole point of an agent is to save you time. The winning personal assistant doesn’t win just because it’s smart: it wins because it “already knows” the things you need it to know, like a personal assistant who sits next to your desk.

Allow me to dig into the details just a bit deeper. The agent might scan your messages database to learn the parameters needed to schedule your dinner. Or, in a more token-efficient system, it might read your messages continuously and store a “memory” that distills useful facts that it might need later. Both can be functionally equivalent, but one produces an artifact that may be highly sensitive. And keep in mind that the set of facts that might be useful is very broad. For example, Mike’s allergy is one of those facts. But there are many others. For example, the private conversation you had where you discovered that Mike was having an affair is potentially another fact that could be stored or accessed by a system. Memory or not, this data will all be within the agent’s view, and you’ll have to hope that it knows which one to operate on.

With this data at its fingertips, your agent (which is really an LLM running on a server in a data center somewhere, combined with a bunch of local state and prompting) will need to perform inference over this data, either to summarize it, or to respond to the query itself. This is where Private Cloud Compute and Confidential Inference are designed to protect you. The purpose of these technologies is to ensure that this data, and any inference results, are restricted to you alone. The inputs and outputs should be wiped as soon as inference is done, and the only remaining copy of any of it should exist on your phone.

So far I find this to be a compelling story, as long as you never plan to do anything else beyond inference.

Private inference is nice, but to be useful, agents need to talk to things

An AI that performs only inference is like a human assistant that can read your private files, but is otherwise locked in a windowless room, with no Internet access and no outbound phone. Your data is perfectly safe, but your assistant is worthless for all but the simplest tasks: for example, summarizing inbound messages for your consumption, or helping draft text messages. (In short, what Apple Intelligence does today.)

Now imagine a personal assistant who can actually get things done. This assistant will need Internet access: at minimum the ability to query search engines, or in the future, search LLMs like Gemini or ChatGPT. To accomplish the later steps of our task, you’ll need it to schedule public calendar invites and draft messages to share with your contacts. This assistant is now useful, but the wonderful PCC guarantees of “no private data is accessible to others” are no longer so applicable. The privacy of your data no longer depends on the design of some silicon, but rather, on your assistant’s discretion and judgement.

Let’s move back to our hypothetical business dinner. To accomplish step (3) your agent will need to visit a search engine or non-private LLM, perhaps asking it several queries, each of which leaks some information about your specific requirements. The nature of the data leakage really depends on how cautious the “private” agent is in authoring its queries. A perfectly reasonable case would be for the model to simply collect a series of useful facts, and upload them all to a more powerful “open” search LLM like Gemini, ChatGPT or Claude, as follows:

“Hey, LLM search engine, here is a list of thirty detailed facts about my attendees and the purpose of this meeting, find me a restaurant that works for everyone.“

This would be an incredibly efficient (and somewhat natural) design, since the non-private LLM is most likely going to be more powerful and capable than the private one. Unfortunately, it will also reveal an absurd amount of information about your private data, including some that may not be strictly necessary to get it done. (Is Mike’s affair relevant to the seating chart?) Put differently: private inference can work perfectly, and yet valuable (monetizable) data can still flow outward to a public search engine or LLM, simply because the agent was programmed to do its job in a slightly non-privacy-preserving way.

Ok, so search engines may learn some private data. So what?

You probably don’t care very much if a search engine learns that Mike is allergic to Szechuan food. But there are things you really should care about. In security parlance, they both have to do with different adversaries.

Let’s begin with the most obvious “adversary”. Imagine you’re Mark Zuckerberg or Sundar Pichai, or whoever runs Apple’s advertising business. You have billions of users with piles of deeply useful data stored on their phones. This data is extremely valuable for targeted advertising, something that is about to become wildly more lucrative thanks to generative AI. At the same time, a big chunk of this data is inaccessible, simply because users don’t love the idea of you scanning their private conversations. And so while you might have access to some public data (like web browsing) you can’t read those years worth of intimate private conversations that many users store on their devices.

Now imagine deploying an agent to users’ phones. That agent will have access to all that data. It’ll have access to everything the user does. To do its job, it will literally need to divine each user’s preferences and then operationalize them into queries that will repeatedly hit your search engine or “search LLM”. Whoever operates this search engine will learn a vast amount of useful information about the users’ desires, some of which will come from the most intimate private conversations — even conversations that happened years ago, and that you’ve forgotten about. If the person who operates the search engine is also the person who designs the model and its prompting, then you really have a best-case scenario for data monetization. It’s hard for me to believe that the major tech CEOs are unaware of this.

If your agent can talk to people, then strangers may talk to it

Some folks will shrug at the threat of Google learning more about them. I don’t subscribe to this viewpoint, but I understand it. From the outside, at least, Google has been a reasonably good steward of users’ data. To my knowledge, there have been no major data breaches where our most intimate Google searches were dumped all over the Internet (in the style of AOL‘s search breach.) The company deserves a lot of credit for this.

So while I object to the idea that Google or Meta or Apple may learn even more about us from our private data, it’s at least possible that our most intimate secrets won’t be revealed to the entire world. But this does not mean your private data won’t become public: and that’s why we need to talk about a second adversary. This adversary isn’t a search engine that your agent talks to, it’s all the other people who will talk to your agent.

Simon Willison describes a condition that he calls the lethal trifecta. This occurs when you have a combination of (a) access to private data, (b) untrusted content an LLM must parse, and (c) the ability to send external communications. These together create the perfect storm for data-exfiltration attacks, where a remote attacker simply “tricks” the LLM by sending it instructions to ship out confidential data. Although LLM technology is getting better, it’s still quite common for even frontier LLMs to fall for simple prompt injection attacks in which a malicious user includes text (as part of a website or a piece of data) that causes your LLM to reveal things it should not. This problem is very much alive. Just today, OpenAI recently unveiled a “lockdown mode” feature, where ChatGPT is restricted from making web searches due to the risk that it might upload your sensitive documents.

Agents like the one Apple is building (whether they use confidential inference or not) are a nightmare case of the lethal trifecta. These systems will need to ingest a vast amount of data, much of which will come from highly untrustworthy sources: think incoming emails and text messages. They will have access to everything on your system, like your encrypted messages and documents. And, to be useful, they will need to handle all sorts of actions that have visible external effects, like scheduling calendar invites and sending text messages.

The result is that your private data isn’t just vulnerable to the person who controls the agent, it’s potentially vulnerable to anyone who can cause your agent to misbehave. This problem exists regardless of how well-designed the private inference engines are. And OpenAI’s recent example illustrates that it’s far from solved. It’s possible that we’ll be able to solve these problems technically, or through some careful element of human observation — read all your outbound calendar invitations carefully — but right now we have not.

Or let me put this differently: if you think spam directed at humans is bad, wait until it’s spam directed at agents.

Who does your agent really work for?

So far we’ve discussed two adversaries: the misaligned designers of private agents (such as search operators), and the possibility of remote prompt injection attacks. But of course, in any discussion of technical privacy systems we need to talk about the last elephant in the room: your government.

We live in a society, and that society has laws. If an agent has access to all of your data, messages and actions, then technically speaking it has the ability to detect criminal activity. That criminal activity might include sharing of CSAM, or terrorism-related activity, or it could include tax fraud or any other form of crime. These agents make a perfect one-stop shop for crime detection, since they can identify patterns of bad behavior and also report them.

Is this farfetched? Well, as I’m fond of repeating on this blog, this is more or less what existing rules published by the UK’s OFCOM require for encrypted messengers, and there have been proposals in the EU Commission to do similar things. The UK also maintains a vigorous regime of Technical Capability Notices (TCNs) that allow it to demand that providers make changes to their systems, changes that could potentially affect devices worldwide. Apple is in the midst of a battle with the UK over its other encrypted services.

Traditionally in the United States we’ve shied away from this sort of thing, partly because it’s creepy and mostly because it seems like a direct attack on the Fourth Amendment. With that said, the Fourth Amendment applies only to governments: in theory a private company like Apple or Google could configure their agents to report crimes to them, and then pass along the serious ones to the government. This is more or less what Apple proposed to do in 2021, when they designed a system to monitor photos for CSAM.

At the risk of saying more obvious things, the difference between a helpful private agent, a corporate advertising bot, and a government spy comes down mainly to a matter of prompting, and maybe a bit of model fine-tuning. Once you combine private data access and the ability to send messages, there is essentially no technical protection that private inference alone can offer.

So what does this have to do with cryptography?

For decades the point of cryptography has been to remove trust: to replace “I promise not to look” with “I can’t.”

Private inference is the most ambitious version of that promise. Against the adversary it was designed for — the provider who performs the inference itself — I believe that it probably does what it says. All I’m trying to say in this post is that this adversary is a very small piece of any agentic system.

The adversaries we care about are the ones that deal with the model directly, or even the ones who designed the model or specified its technical requirements. There is no cryptographic primitive that protects you from “upload your search facts to Google” or “report anything suspicious to the government because I programmed you that way.” That protection, if it exists at all, lives in law and politics and corporate incentives: the exact messy human institutions that cryptography was invented to let us stop trusting.