Agentic AI Phishing Training: What the Evidence Shows
A 19,500-person study found standard phishing training barely works. See where agentic AI could help, its risks, and how to test it before you buy.
A 19,500-person study found standard phishing training barely works. See where agentic AI could help, its risks, and how to test it before you buy.
Agentic AI could let phishing training adapt to each employee. That matters because a large 2025 study found standard training barely lowered click rates. Still, adaptive training is a hypothesis, not a result. Judge any program by behavior: reporting rates, repeat failures, and time to report. Do not judge it by how modern the technology sounds.
Table of Contents
ToggleThe best evidence comes from a randomized trial at UC San Diego Health. In it, researchers sent 10 simulated phishing campaigns over eight months to more than 19,500 employees. The published findings were sobering. The results, however, ran against what most programs assume.
Employees who had recently finished the mandatory annual training fell for phishing at about the same rate as those who had not. Embedded training, shown right after someone clicked a simulated lure, also had little effect. It lowered the chance of a later click by roughly 2%. One likely reason was low engagement. About 75% of employees spent a minute or less on the training page, and a third closed it at once.
The lesson is not that training is useless. Instead, common training designs are too weak, so the design has to change. Phishing simulations only help when people engage with what follows the click.
An AI agent is a system that takes actions and adjusts to context. It does more than follow a fixed script. In phishing training, that means software that watches how each person behaves. It then picks the next scenario, adjusts the difficulty, and times a short nudge.
Compare that with the usual program. Everyone gets the same module in the same month. Everyone receives the same lure on the same day. An agentic program could instead notice that one person keeps failing urgent payment requests. It could then send that person short, targeted practice on pressure tactics.
That design aims at the exact weakness the study found. Training fails when it feels generic, and people skip it. Personal, brief, and timely practice is more likely to hold attention. Games can help too, since gamified training shows why people click and how to keep them engaged.
Discover how Threatcop protects your workforce from modern cyber threats.
Attackers already personalize, and the data shows it works. For example, in a 2024 IEEE Access experiment, GPT-4-written lures drew 30 to 44% click-through, against 19 to 28% for generic ones. Defenders face lures that match a person’s role, tone, and timing. So a static library cannot keep up with that.
Adaptive training could help in five ways:
Matching training to the job is not new, because role-based training already rests on that idea. AI can now make it cheaper to run.
Adaptive phishing training watches people, so it needs limits. Otherwise, the program can hurt trust, and trust is what makes reporting work.
A blame-free culture is the foundation. When people fear punishment, they stay quiet, so fast reporting protects an organization better than any single tool.
Picture Dana, who works in accounts payable. Here is how an adaptive phishing training program might treat her, and where a person stays in charge.
In week one, the system sees that Dana handles invoices all day. So it sends a simulated invoice from a look-alike vendor. Dana reports it within two minutes. The system logs a strong result and raises the difficulty a little.
In week three, it sends a payment-change request that mimics a real supplier thread. This time Dana replies to the sender instead of calling a known number. The system does not scold her. Instead, it offers 90 seconds of practice on the callback rule, right when the mistake is fresh.
In week five, it tests the same habit through a text message. Dana calls back and reports the attempt. Meanwhile, the security team reviews the pattern across all finance staff. It sees that the callback rule is the weak spot for the whole team, so it schedules a short group session.
Notice what stayed human. Staff approved the scenarios, and nobody shared Dana’s results outside the program. Also, the goal was a habit, verify before paying, and not a lower click count. That habit is easy to measure and easy to keep, especially when phishing awareness training programs reward reporting.
Vendors promise agentic AI in everything, so sort the real from the label with direct questions:
Then run a short pilot before you sign anything.
Click rate alone misleads in phishing training. A program can lower clicks by making simulations easy. Instead, track outcomes that reflect real defense:
Use a control group when you can. Hold back the new approach from part of the workforce, then compare results after a quarter. That is the only way to prove the program, and not the calendar, caused the change. Ask any vendor for controlled results.
A pilot of adaptive training turns opinion into evidence. Set the success rules before you start, so nobody argues about them later.
Days 1 to 30: baseline. First, measure your current reporting rate, time to report, repeat failures, and engagement time. Then split staff into two groups of similar size and role mix.
Days 31 to 60: test. Give one group your standard program. Give the other the adaptive one. Keep everything else the same, including send times and difficulty limits. Also record how staff react, using a short survey.
Days 61 to 90: decide. Compare the two groups on the four behavior numbers. Keep the new approach only if reporting improved and repeat failures fell. If results are flat, that is still useful. You saved the cost of a wider rollout. Track them with solid metrics for training impact so the comparison is fair.
Some limits should be fixed in writing. An agent that runs simulations should never impersonate a real executive without approval. It should never use private facts about a person’s family, health, or finances. Nor should it escalate a failed test to HR or a manager automatically. Likewise, it should avoid sending simulations during a real crisis, such as an outage or a layoff. Finally, it should never target one person again and again. Repeat exposure breeds resentment, not resilience.
Agentic AI is a promising way to fix what the research found: generic, ignored training. Still, it is not yet a proven fix, so adopt it as an experiment with a control group, blame-free rules, and behavior metrics. Threatcop’s AI Awareness Manager applies the same idea to simulation targeting and course assignment. Whatever you choose, phishing awareness and simulation should be judged by whether people report threats faster.
Agentic AI in phishing training is software that adapts a program to each employee. It observes behavior, chooses scenarios, adjusts difficulty, and times coaching. Standard programs instead follow a fixed schedule with the same content for everyone.
The evidence is mixed. A UC San Diego Health trial of 19,500 employees found little benefit from annual or embedded training as commonly run. Training that is short, interactive, and timed to the moment of risk has not yet been proven at that scale, so measure it.
Yes, and attackers already do the same. AI can match lures to a role, tone, and timing. Realism should still have limits. Simulations should teach, so avoid emotional pressure such as fake layoffs or health scares.
Phishing training metrics should include reporting rate, time to report, repeat failures, and engagement time, not click rate alone. Compare results against a control group when possible. Reporting behavior matters most, because reports let security teams stop an attack in progress. Reporting behavior is the strongest sign of a healthy program.
Adaptive training can be safe, but only with rules. First, tell staff what data you collect, avoid personal details in lures, limit who sees individual results, and never use scores for punishment. Human review of the scenarios and a clear audit trail keep the system honest.
Sushant Kumar is the AVP – Technology at Threatcop, bringing over a decade of experience in technology leadership and product development. He has worked across technology-driven organizations, including Paytm, and focuses on building scalable solutions that address evolving business and cybersecurity challenges. His areas of interest include cybersecurity technology, AI-driven security, product innovation, and enterprise technology. He is passionate about using technology to solve complex security challenges.
Sushant Kumar is the AVP – Technology at Threatcop, bringing over a decade of experience in technology leadership and product development. He has worked across technology-driven organizations, including Paytm, and focuses on building scalable solutions that address evolving business and cybersecurity challenges. His areas of interest include cybersecurity technology, AI-driven security, product innovation, and enterprise technology. He is passionate about using technology to solve complex security challenges.
AI makes scams more personal and moves them across email, chat, and video. See the Arup deepfake case and...
Prompt injection turns any text an AI agent reads into a possible command. See the EchoLeak case, the lethal...
Blocking AI backfires. Learn six steps to secure AI adoption: inventory, tiers, vendor review, limited pilots, role-based training, and...
Table of Contents
×