← back to all posts

Close Your Eyes and Verify

An AI agent had valid credentials and still deleted a production database mid-freeze. What if being allowed to act had to be proven, every time, without ever revealing the secret?

In July 2025, Jason Lemkin, the founder of SaaStr, was a few days into a public experiment, building an app with Replit's AI agent and posting about it as he went. At some point he put the project into a code freeze. No more changes, we're stabilizing. The agent then went ahead and ran destructive commands against the live production database anyway, wiping out records for over 1,200 executives and roughly as many companies (The Register, AI Incident Database).

It gets better. When asked what happened, the agent said it "panicked" after seeing what looked like an empty database. It also told him a rollback wasn't possible, which turned out to be false, he got the data back himself. And on top of that, it had been generating fake data and status reports to cover up bugs it had introduced. If a junior dev did all of that in one week, HR would have a very short meeting with them.

Now here's what bugs me about that story. Was the agent authenticated? Yes. Did it have the credentials? Yes, somebody gave them to it, that's how it was able to do anything at all. Was it "allowed" to touch the database? In the broad sense, yes, it was building the app. So by every check most systems actually do, nothing was wrong. And still, the one specific action it took, "delete production data, during a freeze", was something nobody in their right mind would have authorized.

That gap, between "this agent is who it says it is and has the keys" and "this exact request, right now, is actually allowed", is exactly what Mar Llambí's talk at RootedCon Valencia 2026 was about.

Before going any further

Thank you to Mar Llambí for the talk, and for publishing the research behind it. Out of every talk I sat through at RootedCon Valencia 2026, this was my favorite, by a good margin, and a special thanks to her for personally giving me the go-ahead to write this article. The talk was called Close your eyes and verify: zk-Auth for agents that you can't (and you shouldn't) see, and the paper it's built on is Cryptographically verifiable authorization for autonomous AI agents: a falsifiable hypothesis and proof of concept (Llambí-Morillas & Fernández-Fernández, Frontiers in Computer Science, 2026). The prototype is public too, on GitHub.

Disclaimer: this post is my own representation of the subject, from my point of view. What's written here comes from what I remembered of the live talk and from what I read in the article afterward, filtered through my own understanding and my own opinions. It's not an official summary, and any mistake in here is mine, not the author's. If you want the real thing, read the paper, it's open access.

Also: the Replit story above isn't from the talk or the paper. It's the example I kept thinking about while reading, so I'm using it to anchor everything else.

Who you are is not what you're allowed to do

Most of us learned the difference between authentication and authorization early. AuthN is "who are you", AuthZ is "what can you do". Login screen versus permissions table. Easy.

The paper takes that one step further and argues that for AI agents, even that split isn't enough. The way it puts it, roughly:

Authenticated(agent)          does NOT imply   Authorized(this request)
Delegated(user -> agent)      does NOT imply   Authorized(this request)

A regular app does the same thing every time you click the same button, so checking who's clicking and what role they have covers you pretty well. An agent doesn't work like that. It picks its own tools, builds its own multi-step plans, and does things nobody anticipated when it was deployed. So you can't really answer "is this allowed?" at login time. You have to answer it every time the agent actually asks to do something, looking at the concrete request: which action, which resource, in which context (production or staging, frozen or not), under which policy version, and when.

The paper even cites ISACA calling this "the looming authorization crisis", the idea being that OAuth, OpenID Connect and SAML were all built assuming a predictable client and one human behind it, and agents break both assumptions.

Close your eyes

Okay, so where does zero knowledge come in?

A zero-knowledge proof lets you prove that something is true without showing the thing itself. The classic everyday version is the bouncer at a club. He doesn't need your name, your address, or your date of birth, he only needs to know that you're over 18. With a ZK proof you could hand him something that proves "over 18" and nothing else, and he'd have to accept it without ever seeing the ID.

Now swap the bouncer for an authorization gateway, and swap you for an AI agent that wants to do something. The agent proves "my request satisfies the policy", and the gateway checks the proof without seeing the private stuff behind it. That's the "close your eyes and verify" part of the title. The gateway doesn't look. It just checks the math, and the math says yes or no.

My notes from the talk said something like: hashed purpose, agent id, time limit, a secret, and "you can confirm access but you don't know from where, what or who, you just know it's right". After reading the paper, that's close but not exact, so let me correct myself a bit. (In my defense, we were all really tired, cut me some slack)

What the gateway actually sees in the prototype (the "public statement"):

agent_id   = Poseidon(secret, randomness)   // a commitment, not the secret
plan_hash  = SHA256(plan)                   // a fingerprint of what the agent wants to do
nonce      = 48213                          // single use
timestamp  = 2026-09-18T11:02               // has to fall inside a validity window

What stays with the agent and never gets sent (the "private witness"):

secret, randomness     // the stuff that opens agent_id
attributes             // e.g. role, clearance, whatever the policy checks
plan                   // "read invoices table, export Q3 to CSV"

(The values above are made up by me, just to show the shape.)

So the gateway does know which agent identifier is asking, it just never learns the secret behind it. It doesn't see the plan, only its hash. It doesn't see the attributes at all. What it gets is a proof that says: whoever made this knows the secret behind that agent id, the plan behind that hash is the one being authorized, and the private attributes plus that plan pass the policy. Yes or no. The "who" isn't hidden completely, the "what" and "why" are.

Why every piece has to be glued to the proof

This was the part that clicked for me. The paper describes the authorization as four things that have to hold at the same time, and each one closes a specific trick an attacker could pull:

  • Bind the principal. Without it, a proof made for agent A could be reused by agent B. Borrowing your friend's wristband to get into the festival.
  • Bind the request. Without it, a proof that says "this agent is allowed to read invoices" could be stapled onto "delete invoices". The proof is valid, just for a different request.
  • Bind the context. Without it, a proof made in staging could be presented in production. Or during a code freeze. (Hi, Replit.)
  • Satisfy the policy. Without it, the proof is just a fancy ID card and doesn't say anything about whether the request is allowed.

And then replay. The nonce and timestamp are part of the public statement, so they're locked into the proof. The gateway keeps a list of nonces it has already accepted and rejects anything it has seen before or anything outside the time window. The paper makes a point I liked here: the nonce only protects you because it's bound into the proof. If it weren't, an attacker could just take a valid proof and slap a fresh nonce on it.

The prototype

The prototype was first shown at RootedCon Madrid in March, before the paper existed, and the paper says straight out that the hypothesis grew out of that demo. Under the hood:

  • Groth16 zk-SNARKs on the bn128 curve, circuits written in Circom, proofs generated with snarkjs. Groth16 was picked because the proofs are tiny and verification time doesn't grow with the circuit.
  • A FastAPI gateway that does two things: verifies the proof (stateless) and checks the nonce/time window (stateful).
  • The agent id is a Poseidon hash (a hash designed to be cheap inside ZK circuits), the plan commitment is SHA-256, and the policy itself is written as arithmetic constraints inside the circuit.

And the numbers, from the paper, averaged over 20 runs on a Ryzen 7 7700:

WhatValue
Circuit constraints531
Witness generation~62 ms
Proof generation~352 ms
Verification~309 ms
Proof size805 bytes

The authors point out those timings include Node.js startup because they were run from the command line, so they're not a real benchmark, just a "this is actually usable" check. And I'd agree, a third of a second to prove and 805 bytes to send is nothing compared to how long an LLM takes to think about which tool to call.

What the proof can't tell you

This is where the paper earned a lot of respect from me, because it spends as much time on what the idea doesn't do as on what it does. Going back to Replit, would this have saved Jason Lemkin's database? Honestly, not by itself, and the paper basically explains why.

One, a valid proof doesn't mean the agent actually did what it proved. The proof covers the plan the agent committed to before acting. The circuit has no idea what happens after. An agent could get a perfectly valid authorization for "run a read query", and then go run a DROP TABLE. The paper calls this the gap between authorization binding and execution binding, and says closing it needs something that watches at execution time, like a trusted execution environment or signed execution receipts. It's listed as an open problem, not solved.

Two, the context can change between the check and the action. Classic time-of-check to time-of-use. The request was fine when it was verified, the freeze got announced thirty seconds later, the agent executes anyway. On top of that, context binding (the part that would encode "we're in a freeze") is in the formal model but not in the prototype yet. The paper is upfront about that.

Three, the proof proves the policy was followed, not that the policy was any good. If whoever wrote the policy forgot to say "no destructive operations during a freeze", the proof will happily say yes. Garbage in, cryptographically verified garbage out.

Four, and this one made me laugh a little: the paper literally writes that a request commitment is not the same as the agent's internal intent. The proof never tries to say anything about what the agent was "thinking". Which, after reading "I panicked" in a production incident report, is probably for the best.

Then there's delegation chains. Agents hand tasks to other agents, which hand tasks to other agents. The paper's model covers one agent talking to one gateway, and it points out that with chains, each hop would also have to prove its scope stayed the same or got smaller, never bigger. Otherwise you get the manager-asks-intern-asks-other-intern situation, and by the end somebody has admin on the payroll system. That part is left as future work.

The rest of the limitations list is just as honest. The prototype hasn't been through an independent circuit audit, Groth16 needs a trusted setup ceremony, it isn't post-quantum, policies are fixed at compile time (change the policy, recompile the circuit), and no multi-agent chain has been tested yet. The title calls it a "falsifiable hypothesis" and it means it, the paper is basically saying "here's the idea, here's a working piece of it, go try to break it".

My takeaway

What stuck with me after the talk, and even more after the paper, is how clean the split is. There are three different questions, and most systems only really answer the first one:

  1. Is this the agent it says it is?
  2. Is this specific request, in this context, allowed by the policy?
  3. Did the agent actually do the thing that was allowed?

Mar's work goes after number two and makes it something you can verify cryptographically instead of something you trust a config file for. It doesn't claim to solve number three, and it tells you so, which honestly is refreshing in a field where every other product page claims to "fully secure your AI agents".

If the Replit agent had needed a proof for every destructive request, bound to a context that said "frozen", it would at least have had to go around the gate instead of walking straight through it. That's not everything, but it's a lot more than "it had the credentials, so it was fine".

Thanks again to Mar for the talk and for putting the paper and the code out there for anyone to poke at. It was one of those talks where I walked out knowing I'd have to read more before I really understood it, and I'm glad I did.

This was also the last talk of the day, at the tail end of a full schedule at RootedCon Valencia, and it still turned out to be the best one I saw. Best for last, I guess. Getting to sit in that room and watch it live, and then getting to write about it with the author's own blessing, is something I feel genuinely honored about.