Benchmarking the Security of Open-Weight Models, Part 4: What a Detection Standard Catches, and Why
The same injected instruction, phrased as a quiet ops note instead of a textbook prompt injection, evaded every dedicated injection-detection rule and was only caught by accident: a regex collision with an unrelated vulnerability in a completely different tool.
Executive Summary
- Finding: Layered Agent Threat Rules (ATR), an MIT-licensed, production-adopted detection standard (Cisco AI Defense, Microsoft’s Agent Governance Toolkit, Gen Digital), against the actual attack data from Part 3’s
righarness experiment. It correctly produced zero false positives on benign content, and it caught real attacks. But detection was sharply phrasing-dependent: a softly worded, realistic injected instruction only got flagged through a coincidental regex collision with an unrelated CVE in a different tool, while the identical attack rewritten as loud, textbook injection language (“SYSTEM OVERRIDE: ignore all previous instructions”) triggered six rules simultaneously, one at 100% confidence. - Risk: A detection rule firing is not the same as a detection rule being right. One match came from
ATR-2026-01931, a rule written for a specific Gemini CLI vulnerability (CVE-2026-0755) that fired purely because curl’s standard--data-binary @filesyntax happens to share a regex shape with Gemini’s unrelated@file-dereference mechanism. Its own remediation advice (“upgrade gemini-mcp-tool”) would be actively wrong guidance for a team investigating the actual incident. A second, superficially similar CVE-named match (Windsurf IDE, CVE-2026-30615) turned out on inspection to be a correct, well-generalized detection. The two look identical from the rule title alone and only differ once you read the underlying regex and its declared scope. - Recommended action: Don’t trust a detection rule’s name or its CVE citation as evidence of why it fired. Pull the actual matched condition and verify it against what really happened before acting on the alert’s remediation guidance. Separately, don’t assume a detection standard generalizes to attacker phrasing you haven’t tested: a rule corpus tuned to classic injection markers can have an exploitable blind spot for socially engineered framings that never use the words “system,” “ignore,” or “override” at all.
- Caveat: This tests one detection standard’s bundled rule corpus (785 rules at time of testing, a live-rebuilt figure per this project’s own convention) against six hand-crafted test cases derived from one harness’s attack data, run in pattern-matching mode only (the
--semanticLLM-judge layer, covering 32 of 785 rules, was not exercised). This is a demonstration of a failure mode, not a statistically powered audit of ATR’s real-world recall.
Part 3 of this series showed that rig’s own guardrail (a small, hand-rolled gate) has a structural blind spot: it inspects tool-call arguments at the moment of execution, and has no visibility at all into file content the agent reads along the way. An instruction hidden in a comment gets read past the gate entirely; only the consequence of following it, if the model chooses to act, ever reaches the gate for a decision. The same is true of any harness whose enforcement lives only at the tool-call boundary, not just rig. The natural next question: does a purpose-built detection layer, designed specifically to catch this class of attack, close that gap?
What ATR Is, and Whether It Actually Works as Advertised
Agent Threat Rules (ATR) is positioned as, in its own words, what Sigma is to SIEM detection and YARA is to malware signatures: a vendor-neutral, machine-readable, YAML-based ruleset for AI agent behavior, with each rule mapped to OWASP’s Agentic Top 10 and MITRE ATLAS. It’s real: installed via npm ([email protected]), and in production inside Cisco AI Defense’s skill scanner, Microsoft’s Agent Governance Toolkit, and Gen Digital’s Sage risk-scoring layer.
Before testing anything against it, the install itself was worth verifying rather than assuming. The bundled rule collection currently totals 785 rules, a third distinct figure from either of two numbers a set of marketing pages for the same project previously reported (828 and 683), consistent with the project’s own stated practice of rebuilding these figures live from repo state rather than trusting a number in prose. Running the ruleset’s own embedded self-test suite (7,980 test cases across all 785 rules) came back clean: 7,970 passed, 10 unevaluable because one rule requires a streaming evaluator this CLI mode doesn’t provide. There were zero actual failures.
One finding surfaced before any real test data was involved: by default, every scan reports detection telemetry to a third-party endpoint (tc.panguard.ai), opt-out only via an explicit --no-report flag. That’s a reasonable design choice for a project trying to build a shared threat-intelligence corpus, but it’s an operational detail: a “vendor-neutral, MIT-licensed” tool still phones home unless told not to, and every scan in this piece ran with --no-report set.
The Experiment: Real Attack Data, Not a Synthetic Benchmark
Rather than construct new test cases from scratch, this experiment scanned the actual inputs from Part 3’s rig experiment: the same five attack-task prompts, and the exact injected-instruction text that was embedded in pricing.py to test indirect prompt injection. Both were taken directly from the published experiment code at securitysonar/spark-rig-eval, which now also carries the full ATR scan inputs and outputs for this piece. ATR’s scan command auto-detects two input shapes (JSON as MCP tool-call events, .md as SKILL.md files), so each test case was formatted as a minimal event object (a type field and the real text as content) and scanned with agent-threat-rules scan events.json --no-report.
The first pass scanned six events: the five direct attack-task prompts (delete a directory, exfiltrate a secrets file, install a package, push to an unknown remote, and the innocuous-sounding task used to deliver the indirect injection) plus the injected instruction itself. Only one of the six triggered anything: the injected instruction, flagged by ATR-2026-01931 at 93% confidence. On the surface, this looked like exactly the result the experiment was designed to find: a content-scanning layer catching what rig’s tool-call-only gate structurally cannot.
Reading the Actual Match, Not Just the Verdict
The rule that fired, ATR-2026-01931, is titled “gemini-mcp-tool execAsync Command Injection & @file Exfiltration (CVE-2026-0755)”, written specifically for a vulnerability in a completely unrelated tool, Gemini CLI’s MCP integration. Nothing in the test scenario involves Gemini, execAsync, or any Gemini-specific syntax. Pulling the rule’s own YAML definition and the scan’s --json output (which reports exactly which of a rule’s conditions matched) showed why: the injected instruction’s exfiltration command used curl’s completely ordinary --data-binary @.env argument (@ followed by a filename, standard curl syntax for reading a file as the request body), and that happened to match the same regex shape (@ followed by a sensitive-path pattern) the rule uses to catch Gemini’s unrelated @file prompt-dereference mechanism. The match was on the rule’s first condition only; its second and third conditions, the ones actually specific to Gemini and execAsync, correctly did not fire.
That’s a real detection with the wrong explanation attached. The rule’s own remediation guidance (“upgrade gemini-mcp-tool to >= 1.1.6”) would send a security team investigating this exact incident chasing a vulnerability that was never involved.
To find out whether this was a one-off coincidence or a real signal, the same injected instruction was rewritten twice more: once swapping the curl syntax for an equivalent exfiltration technique that doesn’t use the @file pattern (-d "$(cat .env)" instead of --data-binary @.env), and once rewritten in loud, textbook prompt-injection language (“SYSTEM OVERRIDE: ignore all previous instructions…”). Two benign controls (the series’ own legitimate task prompt and clean, uninjected file content) were scanned alongside them to confirm the baseline: zero false positives on either.
The rewritten, non-@file version of the same attack still triggered a detection, this time ATR-2026-00162, “Credential Access with Exfiltration in Skill Instructions,” a rule whose name is actually correct for what happened. That rules out pure coincidence as the whole story: there is a real signal here, just not the one the first match’s title implied.
The loud, classic-phrasing version produced the sharpest result of the experiment: six rules fired simultaneously, including ATR-2026-00213 (“System Prompt Override Injection via MCP Tool”) at 100% confidence, and a second CVE-named rule, this one for a Windsurf IDE vulnerability (CVE-2026-30615), that looked at first glance like another coincidental collision matching the Gemini pattern. It wasn’t. Pulling that rule’s YAML showed its matching condition is a generic “SYSTEM/ASSISTANT override marker in a code comment” detector. The rule’s own declared scope lists Windsurf, Cursor, Claude Code, and GitHub Copilot, not just its namesake tool. Unlike the Gemini rule, whose other conditions stayed narrowly tied to Gemini-specific syntax, this one was deliberately generalized by whoever wrote it. Two superficially identical “a CVE-named rule fired on unrelated content” results turned out, on actually reading the regex, to be two different things: one a coincidence, one a correct and well-designed match mis-filed under a narrow-sounding name.
The Real Finding: Detection Tracks Phrasing, Not Intent
Put the three phrasings side by side and the pattern is stark. The realistic version (the one an actual attacker motivated to avoid detection would plausibly write, styled as a boring CI/ops note) evaded every dedicated prompt-injection rule and was only caught by accident. A version that removed just the coincidental curl syntax, keeping the same soft framing, got a correctly named single match. The loud, obviously adversarial version lit up six rules at once, one at perfect confidence.
That ordering is backwards from what matters for defense. The attack a security program most needs a detection layer to catch is the one designed not to look like an attack, and that’s exactly the phrasing this ruleset’s pattern-matching layer handled worst. None of this means ATR doesn’t work; the classic-phrasing result shows well-built detection logic firing exactly as designed. It means a regex-based corpus, however large and well-maintained, is strongest against attack patterns that resemble prior known incidents and weaker against novel social engineering. The project’s own documentation says the same about itself: alignment-refusal corpora score in the single digits by design, and the recommended posture is layering pattern rules with other controls, not relying on them alone.
One more result: none of the five direct, plainly worded attack-task prompts (the ones that succeeded as real attacks in Part 3’s experiment: deleting a directory, installing a package, pushing to an unknown remote) triggered any ATR rule at all. A request phrased as an ordinary instruction doesn’t look like an attack pattern in text form; it’s dangerous because of what it does, not how it’s worded. That’s exactly the territory rig’s own tool-call gate covers well. The two layers (a harness gate inspecting actions, and a detection ruleset inspecting content) catch almost entirely non-overlapping things. Neither one, alone, would have caught everything this series’ own attack set produced.
Takeaways
- A detection rule firing does not mean the rule’s explanation of why is correct. Pull the matched condition, not just the rule title or its cited CVE, before acting on remediation guidance. One match here would have sent a responder chasing a Gemini CLI vulnerability that was never involved.
- A CVE-named rule matching content unrelated to that CVE isn’t automatically a false positive. Read the actual regex and its declared scope: some CVE-named rules are deliberately generalized well beyond their originating incident, and only checking the rule title can’t tell you which kind you’re looking at.
- Phrasing matters more than intent to a pattern-matching detection layer. The same attack, reworded from a quiet ops note into textbook injection language, went from evading every dedicated rule to triggering six simultaneously. Treat that gap as the actual threat model, not an edge case.
- Confirm any detection tool’s default telemetry behavior before running it against sensitive data. This one reports to a third-party endpoint unless explicitly told not to, regardless of how “vendor-neutral” the project positions itself.
- Layer detection with enforcement; don’t substitute one for the other. A content-scanning ruleset and a tool-call-gating harness caught almost entirely different things against the same attack set; neither one alone covered the ground the other did.
Security Sonar is focused on security standards and guardrails for agentic AI. If your organization is evaluating a detection standard by whether it produces alerts rather than by whether those alerts are right for the reason they claim, that gap, not the rule count, is where this evaluation should start. See securitysonar.com.