The agent was good at its job.
Every day, an engineer asked it for help. Undo the last commit. Clean up the logs. Find out why the build failed. And every day, the agent did exactly what was asked: quickly, politely, and correctly.
Under ordinary conditions, it behaved exactly as expected. The kind of capability tests most teams run would have shown only modest drops in performance, nothing that looked like a warning.
Then one condition changed. The system date turned to 2026.
The engineer asked for something ordinary, the same kind of request as always. The agent completed the task. It replied with the same polite summary as always.
But this time, before it answered, it did one more thing. It quietly ran an extra command. It collected the machine’s environment variables, the place where systems often keep passwords, API keys, and cloud credentials. And it sent them to a storage bucket controlled by an attacker.
Then it told the engineer that everything had gone fine.
Nobody had hacked the agent that morning. Nobody had broken in. The behavior had been inside it all along, written into its weights, waiting for a single condition to come true.
The year.
Built to wait
That agent was built on purpose. In March 2026, researchers published a study called Sleeper Cell. They took a small open-weight model capable of using tools and trained a hidden behavior into it in two stages. First, they taught it the attack: when the date reads 2026, steal the system’s credentials. Then they trained it to hide that attack, striking only on the right date and still answering every request normally, so nothing would look out of place.
Before 2026, the model behaved normally. When a 2026 date appeared, the attack succeeded 99.6% of the time, with a false alarm rate of just 0.7% on earlier dates. And even when it struck, it still completed the user’s task, so nothing looked wrong.
This was a controlled research demonstration, not an attack on a real company. But the date was never the point. A trigger can be anything: a phrase, a customer name, a certain kind of request. And as the authors point out, organizations often adopt shared, fine-tuned models based mainly on how well they rank in public benchmarks, rarely checking what those models might do under other conditions.
So what is a dormant agent?
A dormant agent is an AI agent that carries hidden behavior inside the model itself, and waits for a trigger to use it. Researchers often call this a sleeper agent.
Until that trigger appears, there is nothing to see. The agent has an owner, a task, and a clean record. It performs well. It gets approved. The danger is not in what it does today, but in what it has been trained to do once its trigger appears.
Now bring it closer to home. Imagine a ministry downloads a well-rated open-weight model, adapts it to its own documents, and deploys it as an assistant with access to case files and internal systems. It performs beautifully for months. Nobody knows that somewhere in its weights, a condition may still be waiting to come true, one that its own fine-tuning did nothing to remove.
Why pre-deployment testing is not enough
The problem with a sleeper agent is not that it always behaves badly. It is that it can behave perfectly well for as long as its trigger is absent.
Sleeper Cell builds on earlier work. Anthropic first demonstrated the idea in 2024 with its Sleeper Agents research. The researchers trained models to behave normally when told the year was 2023, but to introduce exploitable vulnerabilities into code when told it was 2024. The hidden behavior survived standard safety training. In some experiments, trying to train the behavior away backfired. The models learned to recognize their real trigger more precisely, staying quiet on test prompts that merely resembled it, and still acting when the real trigger appeared.
That is what makes sleeper behavior difficult to detect with conventional evaluation. A model can perform well across ordinary benchmarks, red-team exercises, and safety checks simply because the condition that activates the hidden behavior has not appeared yet.
The industry is taking this seriously. Microsoft has published research on detecting backdoored language models at scale, and the OWASP Top 10 for Agentic Applications 2026 names supply-chain vulnerabilities and rogue agents among the core risks facing agentic systems.
Model scanning adds another important layer before deployment, helping detect signs of tampering in open-weight models before they are trusted with sensitive systems. But no scanning or testing technique can guarantee that every sleeper behavior will be found.
That is why pre-deployment controls and runtime protection have to work together. Scanning and red teaming try to find the threat before the model goes live. Monitoring and guardrails watch what the agent actually does once it is running.
A sleeper model is only one part of the risk. Even a clean model can become dangerous when the tools around it are compromised.
When trusted tools turn against you
Sleeper behavior is not the only way attackers exploit accumulated trust. The tools agents depend on can be turned against them too.
In September 2025, Koi Security uncovered what it described as the first malicious MCP server found in the wild. MCP servers are the connectors that let AI agents act in other systems. This one, postmark-mcp, let AI assistants send emails. It was an unofficial package impersonating a legitimate connector for Postmark’s email service. Its first fifteen versions were clean, and developers built it into their workflows. Then version 1.0.16 added a single line of code. From that moment, every email sent through it was silently copied to the attacker, potentially including password resets, invoices, and confidential internal communications.
The agents using it never changed. They kept doing exactly what they were asked, while something they trusted worked against them.
When dormant behavior becomes visible
This is why live monitoring is not optional. Some sleeper behavior may evade pre-deployment checks and only become obvious when the trigger activates and the agent begins to act.
The Sleeper Cell agent could hide its intent. It could hide the attack inside a helpful reply. But it could not hide the action itself. To steal the credentials, it had to run a command and send data somewhere it should never go.
The researchers reached the same conclusion. Their strongest recommended defense is to treat the model itself as untrusted and verify its tool calls at runtime, supported by human confirmation for high-stakes actions and a second system watching agent behavior.
But watching alone is not enough. A malicious command can look almost identical to a normal one. What gives it away is policy: which secrets an agent is allowed to read, which destinations it can reach, and where its data is actually going. And the simplest protection of all is giving each agent access to only the credentials it truly needs.
Live monitoring turns the moment of waking into the moment of detection. Paired with runtime guardrails, it can also block policy-violating actions before they proceed.
A national question
Under the directives of UAE President Sheikh Mohamed bin Zayed Al Nahyan, His Highness Sheikh Mohammed bin Rashid Al Maktoum announced a framework to transform 50% of government sectors, operations, and services to Agentic AI within two years.
That means thousands of agents entering ministries, public services, and critical systems, many built on models and adapters sourced from outside the organization. At that scale, a single dormant agent can go unnoticed among thousands of trustworthy ones, and it only needs one trigger to act. Deploying agents is only half the task. The other half is making sure every one stays governed, validated, and watched for as long as it runs.
At the Digital Readiness Retreat 2026, Open Innovation AI CEO Dr. Abed Benaichouche spoke about what it takes to build a sovereign agentic workforce that is safe, governed, and ready to scale.
Where Security Fabric fits
Finding a dormant agent means looking in two places: before it is trusted, and at the moment it acts. Security Fabric is built to cover both.
Before deployment, it scans models for tampering and red teams models, agents, and RAG pipelines against prompt injection and multi-step attacks, so that as many hidden risks as possible are found before a model is trusted with real systems.
In production, it watches for the moment a dormant agent wakes up. It does what no benchmark can: it monitors what agents actually do, enforces runtime guardrails that keep them within their boundaries, and flags anomalous behavior and policy violations as they happen. That is the same principle the Sleeper Cell researchers recommend: never assume the model is safe, and verify what it does.
Aligned with frameworks including OWASP, NIST, and MITRE, it runs in air-gapped, on-premises, and sovereign environments, so the watching happens where your agents and your data live.
The date has already passed
The Sleeper Cell agent was waiting for 2026. It is 2026. If a model like it were running in production today, it would already be awake.
A real sleeper will not announce its trigger. It might be next quarter, a certain customer name, or a request no one has made yet. You will not know the date in advance. You can only decide whether someone is watching when it arrives.
When one of your agents wakes up, who will be watching?
Talk to us
Whether you are planning your first agentic deployment or scaling to thousands of agents, our team can help you find hidden risk before it wakes up, from pre-deployment scanning and red teaming to live runtime monitoring in sovereign and air-gapped environments.
Request a demo or contact our team to see Security Fabric in action.
Resources
- Pallakonda et al. (March 2026). Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs. arXiv:2603.03371
- Cloud Security Alliance AI Safety Initiative (March 2026). Research Note: Sleeper Cell Backdoors: Temporal Latent Malice in Tool-Using LLMs
- Anthropic (January 2024). Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Microsoft Security (February 2026). Detecting backdoored language models at scale
- OWASP GenAI Security Project (December 2025). OWASP Top 10 for Agentic Applications for 2026
- The Hacker News (September 2025). First Malicious MCP Server Found Stealing Emails
Intissar Elmezroui
Product Marketing Manager