Security teams have spent years training employees to spot phishing by looking for awkward phrasing, odd grammar, and stiff corporate-speak that doesn’t sound like a real colleague – but that tell is disappearing fast, and it’s disappearing because of AI-generated phishing built on leaked writing samples. When attackers get their hands on a real person’s actual emails, Slack messages, or internal memos, they can feed that text into a language model and produce phishing messages that sound exactly like the person being impersonated, right down to their favorite phrases, sign-offs, and sentence rhythm.
This isn’t a theoretical risk. Leaked mailbox archives, exposed Slack exports, and scraped internal wikis are already circulating on criminal forums, and the writing samples inside them are increasingly valuable precisely because they can be used to train or prompt generative models.
Why writing samples matter more than templates
Traditional phishing kits relied on generic templates – a fake invoice, a fake password reset, a fake HR notice. These worked because most employees had never seen a truly convincing internal email from a threat actor.
That advantage is eroding. When an attacker has a genuine sample of how a CFO writes a Friday afternoon budget email, or how a team lead phrases a Slack request for a wire approval, they can prompt an AI model to generate new messages in that exact voice. The result isn’t a copy of one specific email – it’s a plausible new message that reads as though the same person wrote it.
A mid-sized manufacturing firm learned this the hard way last year when a finance director’s old email threads, exposed through a third-party vendor breach, were used to craft a payment redirection request that matched her tone so closely that the accounts payable team processed it without a second look. Nobody noticed a grammar mistake, because there wasn’t one.
Where the raw material comes from
Writing samples end up in criminal hands through several overlapping channels:
Compromised mailboxes and shared inboxes, often from a single reused password. Internal collaboration tools like Slack or Teams that get exported or scraped after a misconfiguration – a pattern covered in detail in Slack Workspace Leaks: Common Mistakes That Expose Messages. Public-facing content such as blog posts, press releases, and conference talks that carry an executive’s public voice. Old data breaches that included full email bodies rather than just credentials, which are far more useful for style cloning than a password hash ever was.
Once a few hundred words of authentic text exist, that’s often enough for a language model to pick up on distinctive habits – short paragraphs versus long ones, particular greetings, how someone signs off, whether they use exclamation points, even typical typos.
How the attack chain typically plays out
Most AI-assisted phishing campaigns built on leaked writing samples follow a similar arc:
First, the attacker acquires writing samples from a leak, breach, or public source. Second, they identify a plausible pretext – a wire transfer, a credential reset, an urgent document review – that fits the impersonated person’s role. Third, they generate one or more draft messages using the cloned style, sometimes iterating until the tone feels right. Fourth, they send the message from a spoofed or look-alike domain, timed to match the target’s normal working hours to avoid suspicion. Fifth, if the first message succeeds, they use the reply to refine the next one, making follow-up messages even more convincing.
This last step is what makes these campaigns particularly dangerous compared to older BEC attempts – the loop closes fast, and each exchange gives the attacker more real text to train on. The mechanics overlap significantly with what’s described in CEO Fraud and Deepfakes: The Next Wave of Leak-Enabled Attacks, where leaked material fuels increasingly convincing impersonation attempts.
Busting the myth: “AI phishing is easy to spot because it’s too polished”
A common misconception is that AI-generated phishing is actually easier to catch because it sounds “too perfect” – overly formal, oddly smooth, lacking personality. That was true of early, generic AI phishing attempts that ignored the target’s actual voice.
It stops being true the moment leaked writing samples enter the mix. A model prompted with real examples of someone’s typos, casual abbreviations, and inconsistent punctuation will reproduce those quirks, not erase them. The polish myth gives teams false confidence precisely in the scenario where they should be most cautious – when the message reads as unmistakably “them.”
Practical steps that actually help
Grammar-based training no longer carries the weight it used to, so defenses need to shift toward verification and containment rather than linguistic pattern-spotting.
Require out-of-band confirmation for any financial or credential-related request, regardless of how convincing the wording sounds – a phone call to a known number, not a reply to the same thread. Limit how much authentic internal writing is exposed in the first place, which means auditing where email archives, chat exports, and internal documents are stored and who can access them. Monitor for company email addresses and internal document fragments appearing in breach dumps or paste sites, since this often surfaces the raw material before it’s weaponized – a discipline discussed further in How Phishing Attacks Bypass Traditional Security Tools. Treat any leaked mailbox, even an old or seemingly unimportant one, as a potential style-cloning source and rotate or archive it rather than leaving it dormant and exposed.
FAQ
Can AI-generated phishing be detected by spam filters?
Standard spam filters catch known malicious infrastructure and obvious red flags, but they’re not designed to evaluate whether a message’s tone matches the purported sender’s real writing style, so style-cloned phishing routinely slips through.
Does this only affect executives?
No. While executive impersonation gets the most attention, any employee whose writing has leaked – through a compromised mailbox, an exposed chat export, or a third-party breach – can be cloned convincingly enough to target coworkers, vendors, or customers.
Is deleting old emails enough to prevent this?
It helps reduce the amount of raw material available, but it doesn’t address writing samples that already leaked through past breaches, vendor incidents, or public content, which is why ongoing monitoring for exposed internal data matters as much as internal cleanup.
The underlying lesson is that the fight against AI-generated phishing isn’t really about spotting bad grammar anymore – it’s about limiting how much authentic writing ends up in the wrong hands and building verification habits that don’t depend on how a message sounds.
