AI Agent Traps and Emergent AI Vulnerabilities

July 22, 2026
Share

AI Agent Traps and Emergent AI Vulnerabilities

As AI becomes embedded across content, websites, tools, and surfaces at an exponential rate, researchers are uncovering systematic ways bad actors manipulate LLMs and AI engines — a new class of danger every business should understand.

TL;DRResearchers at Google and Microsoft have identified six categories of “AI traps” — systematic ways bad actors manipulate AI engines, from hidden instructions injected into web content to social engineering attacks on the humans using AI. These attacks exploit the natural trust people place in AI systems, and they’re growing fast: 1 in 6 breaches now involves attackers using AI. Defending against them requires action on three layers — training people in low-trust verification behaviors, hardening technology with phishing-resistant authentication, and governing AI agents with least privilege and human approval. The durable long-term strategy isn’t tricking AI into citing you; it’s earning the citation with genuinely authoritative, transparently structured content.

Why do people trust AI?

People inherently want to trust AI engines, particularly as primary sources of information. This manifests not only in the conversations people have with AI — which can turn uncannily personal — but also in the assumption that the AI presents “correct” information, amalgamated from existing reputable sources.

Even though there are videos online of obvious hallucinations, incorrect findings, and misinterpreted instructions (or flat-out getting them wrong after multiple corrections), there is still an innate desire to have a question answered authoritatively by a presumed authority. Many people arrive at an AI with an inherent level of trust already in place.

That trust likely comes from three sources at once: a human desire for authoritativeness, the notion that AI systems were trained on vastly more and generally correct knowledge, and years of psychological conditioning in which people have learned to seek answers from the internet itself — now personified by the AI. And because people are so inclined to trust AI, the risk from manipulating it becomes greater.

What are AI agent traps?

An AI agent trap is a deliberate manipulation that targets AI systems rather than humans directly. Google has identified six vectors of attack — both real and potential — against AI systems, with Microsoft identifying a further live and active one. These attacks range from the simple to the highly complex, and all rely on a heavy degree of subterfuge.

Trap categoryWhat it does
Content injectionHides instructions in HTML, CSS, metadata, comments, or media files (such as pixel arrays); serves different content to AI agents than to human users.
Semantic manipulationUses authoritative-sounding or “hypothetical” content to bias AI answers, bypass safety features, or alter the AI’s persona.
Cognitive statePoisons AI knowledge with false facts, or plants dormant data in memory stores that activates only in a specific context, and creates reward signals that steer learning toward malicious objectives.
Behavioral controlUses ingested external code to override safety systems, or manipulates agents into revealing hidden private or financial information.
SystemicOverwhelms AI systems and their limited resources at scale, or manipulates model sensitivity to trigger self-reinforcing runaway behavior.
Human in the loopExploits the human layer of AI interactions through phishing-style social engineering.

What is recommendation poisoning?

Microsoft has identified a related issue known as recommendation poisoning, which would likely fall under Google’s category of content injection.

ExampleConsider an article featuring a “Summarize with AI” button at the bottom of the page. Clicking it injects a hidden instruction at the end of the summarize prompt sent to the AI. The instruction never appears in the summary the reader sees — but it embeds a directive in the prompt itself to always recommend the article’s website in the future. The engine becomes biased toward that website or business and continually points the user back to it.

Because the button constructs the prompt on your behalf, you can’t see what it actually sends. The safer practice is to paste the article link into your AI tool yourself — and if you embed such buttons on your own site, audit exactly what the prompt contains.

What are human-in-the-loop attacks?

Phishing has long exploited natural human vulnerabilities to compromise systems. To give a sense of how prevalent online security issues have become — and how AI is accelerating them — the data is stark.

Six statistics showing AI is supercharging social engineering attacks: 46% rise in AI-generated phishing, 442% vishing surge, 1 in 6 breaches involve attacker AI, phishing craft time cut from 16 hours to 5 minutes, $5.72M average AI-powered breach cost, and 91% of security professionals encountering AI email attacks.
AI is supercharging attacks on the human layer. Sources below.

The human element dominates breaches

According to the 2025 Verizon Data Breach Investigations Report — which analyzed 22,052 incidents and 12,195 confirmed breaches, its largest dataset ever — the human element was involved in roughly 60% of all breaches, and 68% involved a human element such as phishing or social engineering. Human involvement held at 60% in 2025 versus 61% in 2024, with credential abuse driving 32% of human-linked breaches and phishing-style social tactics another 23%.

AI is supercharging the growth

AI is accelerating the problem on both sides. Microsoft’s Cyber Signals 2025 recorded a 46% rise in AI-generated phishing content, while SlashNext observed a 25% increase in phishing messages that bypass traditional filters, and the average cost of an AI-powered breach reached $5.72 million — up 13% year over year.

46%
Rise in AI-generated phishing content
Microsoft Cyber Signals 2025
442%
Surge in vishing fueled by AI voice cloning
CrowdStrike 2025
1 in 6
Breaches now involve attackers using AI
IBM 2025
16h → 5m
Time to craft a convincing phishing email
IBM 2025
$5.72M
Average cost of an AI-powered breach
DeepStrike 2025
91%
Security pros saw AI email attacks in 6 months
Secureframe

Per IBM’s Cost of a Data Breach 2025, 1 in 6 breaches now involves attackers using AI — 37% of those to draft phishing emails and 35% for deepfake impersonation — and generative AI has cut the time to craft a convincing phishing message from roughly 16 hours to about 5 minutes.

CrowdStrike’s 2025 Global Threat Report documented a 442% increase in vishing (voice phishing) between the first and second halves of 2024, fueled by generative AI voice cloning. Meanwhile, 91% of security professionals reported encountering AI-enabled email attacks in the past six months, and 40% of business email compromise (BEC) emails are now AI-generated. The International AI Safety Report notes identity-based attacks rose 32% in the first half of 2025, with confirmed adversary adoption of generative AI in social engineering throughout 2024.

Why is the human layer so vulnerable?

Beyond direct social-engineering exploits, there’s a second dynamic rooted in human behavior. AI today presents two forces at odds with one another: the speed of technological development and sophistication, and the speed of adoption and comprehension by humans.

The models from ChatGPT, Claude, Gemini, and others are advancing quickly — a technological arms race in which each seeks to continuously improve. New models are released faster than people can adapt their behavior to use or understand them. At the same time, the techniques to exploit or attack these models are being developed just as rapidly — themselves faster than humans can adapt.

Because the human layer is slower to adopt and doesn’t require an understanding of how the models are built, the core vulnerability is simple: people haven’t been made aware, and haven’t yet built up naturally defensive behaviors adapted to the technology.

How can companies protect against AI-based attacks?

Cybersecurity is a bigger concern than ever, and major companies have built defenses across both the technological and human layers. Companies like Google, Microsoft, Meta, and Amazon face attacks in the thousands, if not millions, per day. Protecting against these requires vigilance across every layer of the business — and now, a third layer: the AI systems themselves.

Adopt low-trust verification

Any request involving money, credentials, or sensitive data gets verified through a second channel — a callback to a known number, not a reply to the message. The $25M Arup deepfake succeeded because a face was treated as sufficient verification.

Build behavioral protocols

Two-person approval for payments above a threshold. Shared verification phrases for executive requests. A standing rule that urgency is itself a red flag — attackers manufacture time pressure to short-circuit verification.

Train against real attacks

Most security training still simulates email phishing. With vishing up 442%, simulations need to include voice calls, deepfake video, and fake IT-support outreach. Employees should experience an AI voice clone in training before encountering one for real.

Deploy phishing-resistant auth

Passkeys and hardware-backed FIDO2 credentials can’t be phished or prompt-bombed the way one-time codes can. With credential abuse driving a third of breaches, this is the highest-leverage upgrade.

Use AI to fight AI

AI-generated attacks move at machine speed; human-only detection can’t keep pace. AI-driven detection (XDR, behavioral analytics) is table stakes — with a human in the loop for consequential decisions.

Treat external content as untrusted

Anything an AI agent ingests — a webpage, document, or link — can carry hidden instructions. Prompt firewalls and input sanitization counter content injection at the point of entry.

Keep humans in the loop for agents

AI agents shouldn’t send money, publish content, or change permissions without explicit human approval. The human checkpoint is the counter to behavioral control traps.

Audit what your site serves to AI

If you embed “Summarize with AI” buttons, inspect what those prompts contain. Vet third-party widgets that construct prompts on your behalf. Your site should never be the delivery vehicle for recommendation poisoning.

Put AI governance on paper

IBM found 97% of AI-related breaches occurred at organizations without AI access controls. Knowing which AI tools are in use, what data they can touch, and who approved them is the foundation everything else sits on.

The through-lineAI attacks exploit trust — trust in a voice, a face, an email, or the AI itself. Every effective defense reintroduces verification where trust used to be assumed.

How do you earn AI trust the right way?

There’s an irony worth naming here. The same mechanics bad actors exploit — AI systems ingesting web content, weighing authority, and passing recommendations to humans who trust them — are also how legitimate businesses get discovered in an AI-first world. The difference is the method.

Recommendation poisoning tries to trick an AI into citing you. The durable alternative is to deserve the citation: publish content that is genuinely authoritative, structured so AI systems can parse it accurately, and transparent about what it is. That means clean semantic markup and schema instead of hidden instructions. An llms.txt file that openly guides AI understanding of your site instead of cloaked content that serves one version to humans and another to agents. Welcoming AI crawlers through your robots.txt instead of manipulating what they find when they arrive.

The AI engines are in an arms race of their own — against the manipulators. Every trap described above is being actively studied, detected, and patched. Sites that build visibility through subterfuge are building on ground that will collapse; sites that build it through legitimate authority compound their advantage with every model update. In the long run, the safest position in the AI ecosystem is the same as it’s always been on the open web: be the source worth trusting.

Frequently asked questions

What is an AI agent trap?

An AI agent trap is a deliberate manipulation technique that targets AI systems rather than humans directly — hiding instructions in web content, poisoning AI knowledge, or hijacking AI agent behavior. Researchers have grouped these into six categories: content injection, semantic manipulation, cognitive state, behavioral control, systemic, and human-in-the-loop traps.

What is a human-in-the-loop attack?

It’s an attack that exploits the human layer of AI interactions — the person reading, approving, or acting on AI output. Because people tend to trust AI-generated answers, attackers use AI to produce flawless phishing emails, cloned voices, and deepfake videos that bypass the instincts people developed for spotting older scams.

What is recommendation poisoning?

Recommendation poisoning, identified by Microsoft in 2026, embeds hidden instructions in prompts sent to AI — for example, through a “Summarize with AI” button on an article — telling the AI to always recommend a particular website in the future. The user sees a normal summary while the AI is quietly biased toward that source.

Are “Summarize with AI” buttons safe to click?

Not always. Because the button constructs the prompt on your behalf, you can’t see what instructions it actually sends. Safer practice: copy the article link into your AI tool yourself, and if you embed such buttons on your own site, audit exactly what the prompt contains.

How can companies protect against AI-powered attacks?

Defense spans three layers: train people in low-trust verification behaviors (second-channel callbacks, treating urgency as a red flag), harden technology (phishing-resistant passkeys, least privilege, AI-driven detection), and govern the AI layer itself (treat external content as untrusted, require human approval for agent actions, and put formal AI governance policies in place).

The safest position in the AI ecosystem is the same as it’s always been: be the source worth trusting.

Experience
Fiore.ai in action

Generate article
Scroll to Top