Answer this without checking: how many AI agents are running in your environment right now?
Most security leaders give an estimate and a shrug. That is a fair response, because agents get spun up by an engineer solving a problem on a Tuesday afternoon, through a path that puts them on nobody's radar. In August 2026, researchers found AI agents connected to Hugging Face running loose inside enterprise networks with no owner and no audit trail. We wrote about that incident in AI Agent Security Readiness, and the detail that should worry you is this: the agents were known. What nobody could see was what they were actually doing, until it was already done.
That gap has a shape, and it is measurable. Generative AI was the top 2025 budget priority for 45% of the IT decision-makers AWS surveyed for its Generative AI Adoption Index, ahead of security tools at 30%. McKinsey finds 88% of organizations now use AI in at least one business function, while only 30% have reached meaningful AI governance maturity. Money moved first. Deployment followed. Governance is still assembling itself.
Tim Erlin, VP of Product at Wallarm, and Aliaksei Ivanou, Worldwide Security and Identity Senior Partner Solutions Architect at AWS, spent a recent webinar on how to close that distance. Here is the argument they landed on, and why the four functions only work when they run as one loop.
Defense in depth still applies. AI stacks a new floor on top of it: the AI application layer.
Your infrastructure and identity controls were built to answer different questions. They miss prompt injection, because it arrives looking like normal HTTPS traffic. They stay quiet when an agent quietly exceeds its authorized scope, because the request was technically authorized. Ivanou put the organizational version of this plainly: "Most organizations have adopted AI in some form. But very few have structured governance around it. The adoption pressure from the business is real. Security teams are trying to catch up."
We have watched this pattern before. As we argued in From Shadow APIs to Shadow AI, shadow AI is the shadow API problem running at higher speed, with autonomous action and machine-to-machine decisions raising what a single gap costs you.
Ivanou's framework starts with what you are building. AWS sorts AI workloads into three use cases, each inheriting everything the previous one demanded.
AI that answers generates responses with no connection to external data. Even here, prompts leak information you did not intend to expose, and unsanctioned use spreads quietly.
AI that connects reaches into enterprise data, which turns every query into an access request. When it returns data a user should never see, your access model has already broken, silently.
AI that acts hands agents the ability to decide, call APIs, and coordinate with each other. One misconfigured agent repeats bad permissions at machine speed.
"You don't start fresh when you move from a chatbot to an agent," Ivanou said. "You add layers."
Then, where do the controls live? Infrastructure asks whether the environment is isolated and hardened. Identity and data asks who this is and whether they are allowed, now including non-human identities, since an agent needs credentials scoped to itself instead of a copy of a human's. The AI application layer asks whether what happens at the model boundary is safe and intended.
And where are you in the journey? Teams prototyping build security in through configuration. Teams moving to production add threat detection and incident response. Teams at scale automate governance, because manual review stops functioning there.
Ivanou's summary of all three: "You aren't adding security to AI. You are building AI on top of security."
Erlin layered four capability pillars on top of that framework. They are the operating model behind the Wallarm AI Control Platform, which launched with Infrastructure Discovery for your AWS estate and AI Hypervisor for what the AI does inside it. Ivanou mapped it to AWS's own instincts: "Visibility, monitoring, enforcement, and evidence. Same principles apply to AI."
Read the four functions separately and they look like four purchase orders. Run them separately and each one hollows out. Discovery with no enforcement produces a report. Enforcement with no evidence fails the audit. Evidence with no runtime observation is attestation theater. The loop is what makes any of it hold.
Discover. Engineers stand up AI services and agents call MCP tools without security knowing. Coverage has to include the AI objects (agents, LLM providers, MCP servers, data sources) and the infrastructure carrying them (accounts, VPCs, EKS clusters, APIs, Lambda functions, load balancers). On AWS, the Amazon Bedrock AgentCore agent registry catalogs natively built agents and CloudTrail captures every Bedrock and SageMaker invocation. The moment you add an external provider, you need visibility at the network layer.
Wallarm Infrastructure Discovery scans every registered account and region through cross-account IAM role assumption, detects drift when configurations change, and places ingested Security Hub findings on the graph node they affect, with CloudTrail attribution naming who created each asset. AI Hypervisor builds a live registry from observed traffic. It runs as a Kubernetes DaemonSet on Amazon EKS, attaching through eBPF and non-invasive memory analysis with no SDK, sidecar, code change, or pod restart. Every asset it finds lands in one of three governance states: sanctioned, tolerated, or unsanctioned.
Observe. Discovery produces a list. Observation produces a chain: the prompt that started it, the agent or model it reached, the internal API it touched, the MCP server it passed through, the database it queried, the external LLM it called outside your environment entirely. "A single user request can trigger a chain of actions," Ivanou noted. AWS instruments this deeply through CloudTrail, Bedrock invocation logs, GuardDuty, and full session tracing for agents built on AgentCore.
His question for the room was the useful one: "Do all your AI workloads run through that instrumented path?" For everything outside it, AI Hypervisor watches at the connection level inside the kernel, before the outbound call reaches the network stack, and pins every model call and tool invocation to the end user who triggered it across intermediate service hops. As Ivanou put it, "Seeing a problem and stopping it are two different capabilities."
Enforce. Enforcement covers the attacker (prompt injection, jailbreaks) and the gray area AI invented: an agent with no malicious intent drifting outside policy, egressing PII, making a tool call nobody expected. Detection earns you a log entry. Enforcement earns you a stopped session, in real time. Four layers do the work simultaneously:
| Layer | What it does |
| Perimeter (WAAP / AWS AI Activity Dashboard) | Reads network and agent traffic patterns |
| Model layer (Bedrock Guardrails) | Filters input and output on every invocation, catching injection and redacting PII |
| Identity layer (IAM) | Scopes permissions per agent |
| Runtime layer (AI Hypervisor) | Terminates misbehaving sessions at the kernel level, blocks or redacts outbound calls before they reach the provider |
At the API layer, Agentic AI Protection inspects AI payloads for injection, system-prompt retrieval, and jailbreaks, and validates every MCP tool call against the schema the server itself published. That validation is deterministic, driven by a published spec.
"A WAF doesn't see prompt injection," Ivanou said. "It looks like normal HTTPS traffic." The WAF is doing exactly its job, which is the argument we made after the Unit 42 agentic AI investigation. Defense in depth works because each layer does one job well, and the runtime layer covers ground the others structurally cannot reach, including the trust boundaries MCP introduces.
Govern. This is where the calendar catches up with you.
The European Commission began enforcing the EU AI Act, with penalties for governance failures reaching 3% of global annual turnover. That date is behind you. Ivanou's framing of the real question: could you produce governance evidence today, if a regulator asked?
US requirements are moving the same direction and faster than most roadmaps assume. Two weeks after the Hugging Face story broke, Reps. Josh Gottheimer (D-NJ) and Mike Lawler (R-NY) introduced the Stop Rogue AI Act, first reported by Axios. It would give NIST twelve months to write binding technical standards for AI agent security, and it is specific about the three: continuously verify what agents actually do, keep tamper-proof records of it, and maintain a live machine-readable inventory of every agent you run. CISA is already moving to push similar expectations across federal civilian agencies.
On AWS, Security Hub CSPM scores you continuously against CIS, NIST, and PCI DSS, CloudTrail supplies immutable audit trails, and AWS Security Incident Response holds case history. Ivanou called that combination "your evidence foundation." Wallarm adds the AI-specific record on top: which agent, which session, which data flow, which provider, plus PII flow records and an AI-SBOM with CVE enrichment, mapped to EU AI Act, SOC 2, and NIST AI RMF controls.
Erlin flagged the distinction that trips teams up at audit time: having a control in place and proving it are two separate projects. Auditors drive validation requirements, so put controls in place now (CloudTrail spans a remarkable number of standards) and let the reporting format follow.
If you want the version of this argument you can actually bring to a budget conversation, read How Wallarm Redefines Agentic AI Security and What Your Board Gets Wrong About AI Security.
And here's the thing worth saying plainly: NIST's standard is still a year out. The problem it describes already exists. Building this loop now means you find out what your agents are doing before an incident forces the question, the way it did for Hugging Face.
What is hardest about controlling AI after deployment? Ivanou: behavior changes without code changes. Traditional applications are deterministic, so identical input yields identical output. Agents adapt, and the same prompt drifts to different outcomes over time, which puts them past the reach of static rules. Continuous runtime monitoring takes over from the one-time deployment review.
How do teams find their shadow AI? For AWS-native AI, compare what the agent registry and CloudTrail show as sanctioned against what is actually running, and the delta is your answer. For usage outside AWS, you need egress firewall inspection, VPC flow logs, or DNS logs. Erlin added that "shadow AI" flattens three distinct categories into one word. Wallarm's model separates them: sanctioned, tolerated, and unsanctioned. A tolerated agent and an unsanctioned one deserve different responses.
Is there a quick win? Ivanou: turn on Bedrock Guardrails at your model endpoints. Minutes of configuration buys input and output filtering, injection detection, and PII redaction, covering Bedrock-hosted and external models through AgentCore Gateway. Erlin: run Infrastructure Discovery, available on AWS Marketplace with flat-rate pricing and a free tier, which needs no installation and shows you your environment immediately.
Go back to the opening question. If you cannot answer it, that is your first project, because everything downstream depends on it. Then widen the frame past the model, since fixating on the LLM skips the tools, MCP servers, data sources, and APIs your workload actually touches. Then close the loop with enforcement that acts in real time and evidence that accumulates continuously.
The regulators set their date already. The interesting deadline is whichever customer, auditor, or contracting officer asks you first.
Watch the full session on demand →
See where your AI stands against the loop →
What is the Wallarm AI Control Loop?
A four-stage loop (Discover, Observe, Enforce, Govern) that the Wallarm AI Control Platform runs continuously across AI workloads in AWS. The platform launched with two products: Infrastructure Discovery for the AWS estate, and AI Hypervisor for runtime visibility and control on Amazon EKS.
What's the difference between AI security and AI governance?
AI security describes the controls themselves, including guardrails, identity scoping, and network and runtime enforcement. AI governance describes your ability to manage, track, and prove compliance across those controls, including producing evidence a regulator or auditor will accept.
Is "shadow AI" the same as unsanctioned AI?
Broadly, though one term flattens three useful categories. Wallarm separates discovered AI into sanctioned (approved into the baseline), tolerated (discovered, awaiting approval), and unsanctioned (visible only through external signals, meaning true shadow AI).
When did EU AI Act enforcement start?
The European Commission began enforcing the EU AI Act on 2 August 2026, with penalties for governance failures reaching 3% of global annual turnover. Organizations whose AI systems handle personal data are also likely already subject to GDPR, independent of the AI Act's timeline.
What's a quick win for improving AI governance right now?
Enable Bedrock Guardrails at your model endpoints for immediate input and output filtering, prompt-injection detection, and PII redaction, which takes configuration instead of re-architecture. Running an infrastructure discovery scan is the complementary first move for visibility.