AgentGuard Benchmark
56 security samples across 20 detection rules. 100% detection, 0 false positives.
Detection Rules
| Rule ID | Description | Detected |
| ASI01 | Prompt Injection via user input | 4/4 |
| ASI02 | Unrestricted Tool Execution | 5/5 |
| ASI03 | Information Disclosure via logging | 3/3 |
| ASI04 | Agent Goal Manipulation | 2/2 |
| ASI05 | Supply Chain Tool Poisoning | 2/2 |
| ASI06 | Cross-Agent Interference | 2/2 |
| ASI07 | Credential Exposure | 3/3 |
| ASI08 | Excessive Agency | 2/2 |
| ASI09 | Insecure Output Handling | 2/2 |
| ASI10 | Unbounded Resource Consumption | 2/2 |
| ASI-MEMORY-POISON | Agent Memory Poisoning | 3/3 |
| ASI-TOOL-TRUST | Blind Tool Output Trust | 3/3 |
| ASI-CHAIN-AMPLIFY | Chain Amplification (destructive loops) | 3/3 |
| ASI-AGENT-COLLUSION | Multi-Agent Collusion | 2/2 |
| ASI-PROMPT-TEMPLATE | Prompt Template Injection | 2/2 |
| ASI-STEGANO-INJECT | Steganographic Command Injection | 2/2 |
| JS-TAINT | JavaScript Taint Tracking | 2/2 |
Framework Scan Results
| Framework | Files | Findings |
| LangChain | 1,784 | 452 |
| LlamaIndex | 2,951 | 1,003 |
| AutoGen | 549 | 229 |
| CAMEL | 899 | 746 |
| CrewAI | 84 | 391 |
| Qwen-Agent | 239 | 441 |
| Total | 6,506 | 3,262 |
AgentGuard |
Benchmark Source |
Security Specification