← back to all posts

AI Against AI in the Inbox

Phishing got good because attackers started using AI to write it. What happens when you fight back with AI of your own, and where does it still fall short?

Before anything else, thank you to Tomás Isasia (TiiZss) for the talk this whole post is built around. Out of everything I saw at RootedCon Valencia 2026, this was one of my favorites, and I'm genuinely looking forward to seeing where he takes PhishDuel and whatever he brings to the cybersecurity world next.

Disclaimer: the techniques, theories, and attack methodology described below are shared for educational and defensive purposes only. Nothing here is meant to be used against a system, organization, or individual without explicit, authorized permission.

Disclaimer: this post is written from my own notes, memory, and impressions of a live talk, filtered through my own opinions and my own added analysis. It's not an official transcript, and it shouldn't be taken as a perfect representation of everything that was presented, or of what PhishDuel will actually look like once, or if, the author ships a finished product.

PhishDuel is still a proof of concept, and this post doesn't try to cover everything Tomás presented, just the parts that interested me the most. Worth flagging too: as the name implies, PhishDuel is a two sided tool. It can craft spear-phishing attack emails as well as run the filtering side I'm covering here. This post is only about the filter, the defensive half.

Phishing used to be easy to spot. Bad grammar, a weird sender address, an offer too good to be true, some guy who needs your bank details for a few days (we all have that one relative who still forwards these to the family group chat). That version of phishing is dying, and at RootedCon Valencia 2026, Tomás showed up with a talk that explains exactly why, and a PoC built to fight back.

The talk was called AI against AI in the inbox, and the idea is simple to say and hard to pull off: phishing campaigns got this good because attackers are now using AI to write them, so maybe the fix is to use AI to catch them.

We've gone from mass sent emails full of typos, trying to sell you pills or get you to click something shady, to fully tailored, orthographically perfect, socially engineered messages that you'd genuinely have to call the sender to rule out. That's spear-phishing, not the old spray and pray stuff. By definition it's a targeted attack, aimed at a specific person or organization, meant to get them to hand over data, money, or a foothold for something bigger.

The Revolut case, and where it actually breaks down

The example Tomás used to set the stakes was Revolut, and it's a good one, timely too, the story was still developing as this was being written. Here's what's public: Revolut disclosed that customer data got exposed after someone posing as an official government investigation asked for it, and got it. Identity documents, addresses, phone numbers, verification selfies, IBANs, full transaction histories including crypto activity.

I want to go a level deeper than "they got phished," because the mechanism matters. What actually happened, per the reporting that came out afterward, is that the attacker got hold of stolen credentials into a real Italian government email account, on Italy's own certified-mail system, and used that genuinely legitimate mailbox to send the fraudulent legal request. Not a lookalike domain, not a spoofed header, the real thing. And the targets weren't random either, around 680 customers, picked out ahead of time through blockchain analysis as people holding significant crypto. That's not mass exfiltration, that's a scalpel, aimed by someone who apparently does better market research than most actual marketing departments.

Why does that distinction matter for a talk about email filters? Run the Revolut attack through PhishDuel's three filters and it's more interesting than a flat miss. The orthography is clean, it's a real government official's writing, or close enough, filter one lets it through. The header check passes too, the address really is what it claims to be. Filter three is the one worth slowing down on. My first instinct was to call it useless here since Revolut never had prior correspondence with this specific "investigator" to build a baseline against, but that's not the whole picture, filter three isn't only a known-contact comparison. It's also trained to catch the general shape of a social engineering attempt: unusual urgency, a request type that doesn't match how real investigations normally run, whatever tells AI generated phrasing leaves behind. None of that guarantees a catch, a request coming from a genuinely compromised, authenticated government mailbox is about as convincing as this kind of attack gets, but it's fairer to call Revolut a "possibly could have been prevented" than a clean miss. Tomás's headline case is real, serious, and well documented, and it's a genuinely hard one for the tool, since there's no correspondence baseline to lean on. Writing it off completely would undersell what filter three is actually built to notice, though, as you'll see further down.

Who this is actually for

Where the talk lands, and where I think it's strongest, is the question that came right after: how does a small business protect itself from any of this? Not Revolut, they have a security team and a legal department. Think of the local butcher shop, one guy running the whole thing, who's never heard the term spear-phishing and wouldn't know what to do with it if he had. Or a five person shop pulling in maybe a million a year just to keep the lights on. They don't have a SOC. They don't have anyone, the closest thing to a security team is whichever employee's nephew is "good with computers." They're not protected, and most of them can't be, for the exact reasons above.

The three filters

The pitch is three filters, stacked after whatever your SMTP server already does.

  • First filter, the naive one: orthography. Does this read like a native speaker who knows the subject wrote it, or like it went through four rounds of translation and a bad mood.
  • Second filter: headers. Is this really the address it claims to be, what's attached, does any of it check out.
  • Third filter, the interesting one: an AI model that's studied every email you've actually exchanged with your contacts, and knows how each of them writes. Opening line, tone, sign off, all of it.

The first two are self explanatory. The third one is where PhishDuel actually lives.

The agent gets fed your existing mail as training data, the standard .eml export you'd get out of any mail client, and it builds a profile per contact: what they usually write about, their phrasing, the mistakes they always make, the requests they usually send. That profile is the baseline. It's meant to run locally on low end hardware, on purpose, so a business with no budget can still run it, and so nobody's mailbox history ends up training some corporation's model on the side.

From there you could run it two ways: always on, flagging anything that passed the first two filters but still looks off, or manual, where you flag a weird email yourself and get a verdict back.

What's the AI actually doing besides pattern matching against a known sender? Since the attacks it's up against are themselves AI written, it's also trained to spot the telltale signs of AI generated text on top of comparing against what it knows about the sender. Take something as short as:

Hi Carl,
 
I'll call you from this number 123456, I need to ask you something really important.
 
Best regards,
Your boss Martha

If Martha never signs off with "best regards" and always writes "best of luck," if she's never called from a number that isn't the one saved in your phone, if she never opens with "hi," that's three flags on four lines. None of them would ring a bell on their own. Stacked together, they're exactly the kind of thing a person reading fast would miss and a model trained on Martha's actual habits wouldn't.

Not every red flag needs a trained model behind it. If your coworker Andrew emails you "hey sexy, wanna grab a drink after work" and Andrew has never spoken to you outside of stand-up, that's not filter three's job, that's common sense. Maybe Andrew's just shooting his shot, and hey, that's between the two of you, but I'd bet good money it's not company policy anywhere. RIGHT?! Jokes aside, that's the actual point: obvious is easy. This whole talk exists because attackers stopped being that obvious.

Walking through a real target

This part isn't from the talk, it's mine. I wanted to see how the pitch holds up against something more concrete than a slide, so here's a target I built out myself, grounded in exactly the kind of small business the talk was arguing for. Call him John, salesman at a small electrical distributor, seven employees, an old ERP system, no dedicated IT, so John ends up doing sales, invoicing, pricing, restocking, and marketing because somebody has to. Principle of least privilege doesn't exist there. It can't. There's no one to enforce it.

Somewhere in the attacker's head, the objective probably plays out about as sophisticated as "me want data, me sell data, me buy more compute, me get more data." Dress that up for a buyer on some forum later and it becomes "premium verified financial records, bulk discount available," but let's not pretend that's not exactly what's happening underneath.

From the attacker's side, the plan runs:

  • OSINT: who is John, what's his role, who does he answer to, who does he talk to, what's he into outside work.
  • Social engineering plan: what does John have access to, who trusts him, when and where does he actually work.
  • Preparation: write the email with everything just learned, use AI to make it sound right, maybe spoof someone John already emails, maybe use an account that's already compromised.
  • Send, and wait.

With AI doing the writing and the research, that whole chain fits in a day, and it reads a lot more convincing than anything a person would put together solo. If it lands, John could lose his job, the business loses the trust of clients it's worked with for years, and depending on how bad the breach is, that's the kind of hit a five person shop doesn't come back from.

So would the filters have caught it?

This is the part I wish the talk had spent one more minute on, so I'll do it here. Walk John's attack back through the three filters and see where it actually gets stopped.

The email is AI written, so the orthography is clean, filter one lets it through. If the attacker spoofed a contact John already emails regularly, filter two might catch a header mismatch, but only if that spoof isn't a genuinely compromised account, same blind spot as Revolut. Filter three is where John's case gives the model more to work with than Revolut's did. John has history with his real contacts, so on top of catching general social engineering patterns, the model also gets a known-contact baseline to check against. If "his boss" suddenly writes differently, or a supplier suddenly asks for something they've never asked for before, that's a second, sharper signal stacked on top of the first. Revolut only had the general pattern layer to lean on, no correspondence history to compare against, which is why I'd call that one possibly caught rather than reliably caught. John's case has both layers working, which is why it's the stronger example for the tool.

Two takeaways

One, the same tool doesn't perform equally well against every attack shape. Revolut and John's case both run through PhishDuel's filters, but Revolut only gives filter three the general pattern signal to work with, while John's gives it that plus a known-contact baseline. That's not a hard pass or fail split, it's a confidence gradient, and it matters if you're the one deciding what to defend against first.

Two, and this is the one I keep coming back to: small businesses aren't a footnote in this fight, they're most of the fight. The big banks will survive their breaches, badly, publicly, but they'll survive. John's shop might not. If we're building AI versus AI phishing defense, it needs to run on a machine John can actually afford, not just on whatever budget a downtown security team has to work with, the kind with the fancy dashboard nobody's actually watching on a Friday afternoon.

Thanks again to Tomás for the talk. RootedCon Valencia 2026 was my first time in a room full of people from this field, and getting to spend it around that kind of knowledge and research is exactly why I wanted to be there in the first place.