Why do people trust AI?
People inherently want to trust AI engines, particularly as primary sources of information. This manifests not only in the conversations people have with AI — which can turn uncannily personal — but also in the assumption that the AI presents “correct” information, amalgamated from existing reputable sources.
Even though there are videos online of obvious hallucinations, incorrect findings, and misinterpreted instructions (or flat-out getting them wrong after multiple corrections), there is still an innate desire to have a question answered authoritatively by a presumed authority. Many people arrive at an AI with an inherent level of trust already in place.
That trust likely comes from three sources at once: a human desire for authoritativeness, the notion that AI systems were trained on vastly more and generally correct knowledge, and years of psychological conditioning in which people have learned to seek answers from the internet itself — now personified by the AI. And because people are so inclined to trust AI, the risk from manipulating it becomes greater.
What are AI agent traps?
An AI agent trap is a deliberate manipulation that targets AI systems rather than humans directly. Google has identified six vectors of attack — both real and potential — against AI systems, with Microsoft identifying a further live and active one. These attacks range from the simple to the highly complex, and all rely on a heavy degree of subterfuge.
| Trap category | What it does |
|---|---|
| Content injection | Hides instructions in HTML, CSS, metadata, comments, or media files (such as pixel arrays); serves different content to AI agents than to human users. |
| Semantic manipulation | Uses authoritative-sounding or “hypothetical” content to bias AI answers, bypass safety features, or alter the AI’s persona. |
| Cognitive state | Poisons AI knowledge with false facts, or plants dormant data in memory stores that activates only in a specific context, and creates reward signals that steer learning toward malicious objectives. |
| Behavioral control | Uses ingested external code to override safety systems, or manipulates agents into revealing hidden private or financial information. |
| Systemic | Overwhelms AI systems and their limited resources at scale, or manipulates model sensitivity to trigger self-reinforcing runaway behavior. |
| Human in the loop | Exploits the human layer of AI interactions through phishing-style social engineering. |
What is recommendation poisoning?
Microsoft has identified a related issue known as recommendation poisoning, which would likely fall under Google’s category of content injection.
Because the button constructs the prompt on your behalf, you can’t see what it actually sends. The safer practice is to paste the article link into your AI tool yourself — and if you embed such buttons on your own site, audit exactly what the prompt contains.
What are human-in-the-loop attacks?
Phishing has long exploited natural human vulnerabilities to compromise systems. To give a sense of how prevalent online security issues have become — and how AI is accelerating them — the data is stark.

The human element dominates breaches
According to the 2025 Verizon Data Breach Investigations Report — which analyzed 22,052 incidents and 12,195 confirmed breaches, its largest dataset ever — the human element was involved in roughly 60% of all breaches, and 68% involved a human element such as phishing or social engineering. Human involvement held at 60% in 2025 versus 61% in 2024, with credential abuse driving 32% of human-linked breaches and phishing-style social tactics another 23%.
AI is supercharging the growth
AI is accelerating the problem on both sides. Microsoft’s Cyber Signals 2025 recorded a 46% rise in AI-generated phishing content, while SlashNext observed a 25% increase in phishing messages that bypass traditional filters, and the average cost of an AI-powered breach reached $5.72 million — up 13% year over year.
Per IBM’s Cost of a Data Breach 2025, 1 in 6 breaches now involves attackers using AI — 37% of those to draft phishing emails and 35% for deepfake impersonation — and generative AI has cut the time to craft a convincing phishing message from roughly 16 hours to about 5 minutes.
CrowdStrike’s 2025 Global Threat Report documented a 442% increase in vishing (voice phishing) between the first and second halves of 2024, fueled by generative AI voice cloning. Meanwhile, 91% of security professionals reported encountering AI-enabled email attacks in the past six months, and 40% of business email compromise (BEC) emails are now AI-generated. The International AI Safety Report notes identity-based attacks rose 32% in the first half of 2025, with confirmed adversary adoption of generative AI in social engineering throughout 2024.
Why is the human layer so vulnerable?
Beyond direct social-engineering exploits, there’s a second dynamic rooted in human behavior. AI today presents two forces at odds with one another: the speed of technological development and sophistication, and the speed of adoption and comprehension by humans.
The models from ChatGPT, Claude, Gemini, and others are advancing quickly — a technological arms race in which each seeks to continuously improve. New models are released faster than people can adapt their behavior to use or understand them. At the same time, the techniques to exploit or attack these models are being developed just as rapidly — themselves faster than humans can adapt.
Because the human layer is slower to adopt and doesn’t require an understanding of how the models are built, the core vulnerability is simple: people haven’t been made aware, and haven’t yet built up naturally defensive behaviors adapted to the technology.



