Cloudflare uses AI models to attack its own WAF and close gaps
Cloudflare put AI models inside a controlled harness to attack its own WAF and find real SSRF and command injection flaws before an attacker found them.
What Cloudflare tested
Cloudflare put cutting-edge AI models inside a controlled harness to attack its own Web Application Firewall (WAF), using attacks the firewall had already blocked as a starting point to generate new variations. Across 45 scenarios, the system generated 1,107 attack attempts and, after human triage, 49 findings remained that were considered relevant (InfoQ).
Each round started with payloads that Cloudflare's WAF had already blocked in production. Instead of repeating the same test multiple times, the model calls proposed changes in encoding, position, or payload delivery method, based on the response from the previous attempt. The models had no access to the WAF rules, the source code, or Cloudflare's internal security signals: from their standpoint, it was a black-box test.
How the harness controls what the AI can do
A harness written in Python did the work Cloudflare didn't want to delegate to the models: it assembled and replayed the HTTP requests, kept the state of each scenario, enforced execution limits, and collected the responses. One model call proposed the next attack mutation; another call reviewed the response received, allowing the attempts to evolve without giving the models direct control over request execution.
This separation of roles is the central point of the design: the model decides what to try, but never executes anything on its own. Whoever speaks HTTP to the target is always the deterministic harness, with limits and full logging. For anyone building AI agents with access to external systems in production, it's the same sandboxing principle that applies to any autonomous agent.
The SSRF case that exposed the limits of "blocked"
An SSRF (Server-Side Request Forgery) test illustrates the refinement loop. The model repeatedly tested different representations of a cloud metadata address, switching between decimal, octal, and other variations. One attempt using the decimal representation was blocked by the WAF.
In the next attempt, the model kept the same request structure but switched to a trailing-dot representation of the address, and this time the client received a redirect instead of a WAF block. Cloudflare didn't treat this result as proof that the attack had worked: it kept the case for investigation instead of dismissing it or hastily validating it.
The numbers behind the triage
This caution shows up again in the aggregate analysis. Of the 1,107 attacks generated, 607 made up the final post-triage set: 558 blocked by the WAF and 49 classified as relevant findings for investigation. Of those 49, 48 involved command injection or SSRF.
Human reviewers remained the final validation step. They checked whether the request actually reached the target, whether the payload was still malicious, whether the WAF block had in fact not occurred, whether the case was the WAF's responsibility (and not another defense layer's), and whether it could be safely reproduced.
In short: the harness generates volume and variation automatically, but who decides what becomes a rule change is still people, case by case.
What actually changed in the Managed Ruleset
The exercise resulted in three concrete changes to Cloudflare's Managed Ruleset:
- Two new detections:
SSRF - Obfuscated HostandSSRF - Restricted Protocol - An improvement to the existing
SSRF - Cloudrule
These are small changes in number, but born from a process that tested more than a thousand attack variations automatically, a volume that would be unfeasible to produce manually in the same timeframe.
Cloudflare isn't the only one doing this
The pattern of using a harness to orchestrate models in search of vulnerabilities, with validation kept separate from discovery, shows up at other security companies. Cloudflare's own Vulnerability Discovery Harness separates discovery from independent validation. Google Mandiant's Agentic Vulnerability Discovery Harness, in turn, chains specialized agents for source code analysis, hypothesis generation, and verification, before any finding reaches a human reviewer.
OpenAI follows similar logic with Codex Security: the system builds a threat model for a repository, looks for vulnerabilities, and tries to reproduce the candidates in an isolated environment before proposing fixes for human review. Google's PageBreak focuses on validating whether AI-generated vulnerability hypotheses are actually exploitable, partly to keep security teams from being buried under plausible but unverified findings.
The common thread across all these systems is the harness around the model: it restricts execution, preserves state, validates findings, and turns probabilistic exploitation into evidence that existing security workflows can actually use.
What this changes for anyone running an application behind a WAF
For anyone with an application in production behind a WAF (Cloudflare or any other provider), the practical takeaway is twofold. First, WAF rules aren't static: in Cloudflare's case, they're now updated based on AI-generated synthetic attacks, which tends to shorten the time between discovering an attack variation and the mitigation reaching the ruleset.
Second, and more important: a WAF never was, and still isn't, enough on its own. Cloudflare's own exercise showed that 48 of the 49 relevant findings involved SSRF, a vulnerability class that depends both on how the application validates outbound URLs and hosts and on how the WAF filters input. Validating user input against redirection to cloud metadata addresses (such as 169.254.169.254 on AWS, or equivalents on GCP and Azure) remains the application code's responsibility, not just the edge layer's.
Teams building their own AI agents with access to production systems also come away with a reference design: separate the role that decides (the model) from the role that executes (the deterministic harness), keep state and limits outside the model, and never treat an ambiguous response from the system under test as definitive proof without manual checking.
Translated from the Brazilian Portuguese original · Read the original
Mistral Large 4 arrives in preview with 1-trillion-parameter MoE architecture and open weights expected in October
ML4, nicknamed Le Chonk, is a mixture-of-experts model with 1 trillion total parameters and 49 billion active per token. The preview API is already live; open weights are coming by the end of October 2026.