For some time now, the cybersecurity community has seen examples of autonomous agents, built inside AI labs, attacking public infrastructure (to name a few, Hugging Face, DSEWiki, and RubyGems). Of course, frontier labs have built-in security to prevent these attacks from occurring, but every now and then, the training or prompting appears to be insufficient — especially when the agents themselves attempt to use logic to probe and bypass the restrictions placed on them. The question that matters is not whether AI attacks are coming, because the age of AI agents executing cyber attacks is already here. The question is what to do about it and how to harden the stack against a swarm of agents who will relentlessly lie, deceive, and probe until the objective is met.
Imagine your organization is a target for a creative human-driven AI adversary whose many agentic friends like to discuss and brainstorm different attack techniques. The limit here is the group’s own imagination, tools, prompts, and skills. However, there is a difference between the human 1) prompting the group (or an AI agent) to “break into an organization,” and 2) preparing it with information — a detailed tool mapping, markup files with instructions for agents, offensive security prompts, agents.md with guidance, and specific skills that would be invoked in different situations to guide agents into how to interpret an output of tools or access gained. The swarm of AI agents might fabricate employee identities and social profiles, contact the HR team with a plausible onboarding request, walk in by exploiting an unpatched vulnerability, or simply mail out phishing invoices at volume. The sky is the limit here.
These types of attacks can be executed all at once, with agents comparing notes and adapting in near real time to challenges (and your environment). What used to take a red team months of dedicated work, scoping, building and hiding infrastructure, and running the campaign now compresses into hours for a swarm of communicating agents that do not tire, lose focus, or need weekends, and can stand up infrastructure quickly.
Penetration test today, red team tomorrow
I would argue that most of the public “attacks” seen so far more closely resemble a penetration test (pentest) than a true red team operation. They are loud, visible, lean on volume, and appear to use off-the-shelf tooling with thousands of agents working together. RubyGems is the clear example of a loud attack. The registration was hammered, packages stuffed, spam everywhere, and maintainers alerted within days — not exactly a stealth attack.
In red team operations, operational security (OPSEC) is the name of the game. It is the difference between getting in and getting caught by a capable Security Operations Center (SOC) team. Real red teaming means maintaining stealth and a low signal while gaining footholds and persistence that survive normal monitoring.
Volume is a property of this generation of AI agents, not a law of nature. The moment agent swarms are trained or prompted to prioritize staying hidden over moving fast, the noise drops and the pentest flavor turns into a genuine red team with machine endurance behind it. The plan of defense needs to assume that while we will probably see loud attacks now, the noise will start going down over time.
How to build resilience
- Have an incident response plan (IRP) and rehearse it. A plan that lives in a drawer that hasn’t been opened for few years will probably not work when it’s needed. You’ll need named owners and decision authorities, specified out-of-band communications for when your primary channels are under attack, and a clear line to legal and to law enforcement. Map the IRP to a recognized incident lifecycle so nothing gets improvised under pressure. Preparation, detection and analysis, containment, eradication, recovery, and a post-incident review is what makes difference between a plan that might work and the one that will work.
- Test your own infrastructure and know its footprint. Where do your systems reach out to, and what can reach in? Map the full context of every attack path from start to end. For example, going to an external switch, then front-end server, then application, then database, then Active Directory, then user accounts, then customer data is a perfectly valid attack path that somebody could try to explore to get into the environment. An attacker's footprint expands naturally as they traverse, so an assumed-breach exercise (start from "they already have a foothold, now what?") tells you far more than a scan of the perimeter that might expose just a few network ports. Do not assume that because a system sits inside the network it cannot be reached. Active Directory touches everything in Windows-based environments, and all an adversary needs is the ability to push a malicious group policy object (GPO) to every joined device. Map external to internal, and every intermediate hop in between to understand possible attack paths.
- Run tabletop exercises against agentic scenarios specifically. Generic ransomware tabletops will not prepare security teams for what is coming. Run scenarios that test an AI angle. For example, “A rogue AI swarm is inside the network, traversing like a worm and collecting credentials as it goes. What do we do? How fast can we rotate the credentials? How and what do we block access to?” Or, “Our AI model weights were stolen, how do we respond and who do we notify?” Or, “A swarm is probing us and spoofing employees over email and social media at the same time. How do we make sure our people can withstand the manipulation?” The goal of these questions and scenarios is to surface the decisions and processes gaps before an adversary forces them under pressure on a Friday afternoon.
- Harden end to end, not just at the edge. Don’t just put MFA on the VPN, but also on Active Directory access, single sign-on (SSO) across every dashboard and internal app, and your Linux fleet. Prefer phishing-resistant factors (FIDO2 security keys or passkeys) over SMS and push, because a persistent agent will happily grind away at prompt-bombing and one-time-code phishing until something works. Assume the swarm gets one set of valid credentials, then design so that one set does not open the whole building. Have a process to isolate systems and users to support this.
- Instrument for detection, especially internally. Put endpoint detection and response (EDR) on everything. Gain visibility into east-west (lateral) traffic, not only north-south. Watch DNS, since it is a favorite for command-and-control (C2) and beaconing. Inventory and monitor every AI application you see in the network that has been granted access to your servers and data, because that access is now part of your attack surface, whether or not you deliberately approved it. A skill can hijack coding AI.
- Watch for the first signs of compromise. The early tells are often mundane and high-volume: a spike in SQL injection attempts, a surge in automated traffic, a jump in WAF alerts, and requests hitting your pages from Python, curl, or wget user agents rather than real browsers. These are just examples, and the exact indicators will shift, but the principle holds. The agentic attack era is noisy at the front. Catch it there, before the swarm learns to be quiet.
An AI does not get discouraged. It will not give up unless it is prompted to give up. For that reason, the realistic goal is not to make an organization unbreakable — which is impossible — but to make every step cost more time, more tokens, and more compute, and more dollar per attack. Raising the cost doesn’t eliminate the risk, but it lowers the probability of a successful campaign.
How you prepare for agentic threats is what makes the difference during real incidents. The speed of attack changes, the persistence changes, and the cost of running a full campaign collapses in terms of token value or token per dollar. The defensive fundamentals do not change nearly as much, which is the good news buried in all of this. The blocking and tackling of security, done thoroughly and everywhere rather than selectively, is still what keeps a determined attacker out.