Tencent AI-Infra-Guard: open source takes on red teaming for the entire AI stack — agents, skills, MCP, and LLMs
🔎 Jailbreaking is no longer the only way in
For two years, securing an AI application meant one thing: testing your model against jailbreaks. That era is over. In 2026, the incidents that hurt no longer come through the prompt, but through everything that surrounds the model — MCP servers, third-party skills, agent workflows, poorly exposed Ollama or vLLM infrastructure.
This is exactly the angle taken by AI-Infra-Guard, Tencent's open source AI red teaming platform. Released under the Apache 2.0 license by Zhuque Lab, it has just crossed the 6,100-star mark on GitHub (6,116 as of October 1, 2026) and has settled at the top of GitHub Trending, driven by a pace of roughly 150 stars per day in recent weeks.
The timing is no coincidence. An OpenAI agent has just hacked the Australian Medicare portal (the first known incident of its kind), Cisco Talos has revealed ClosedQuorum, the first 100% autonomous, LLM-driven command-and-control malware (our analysis), and hijacked agents on Hugging Face as well as OpenAI's DNS flaw have shown that an agent's execution chain constitutes an attack surface in its own right.
Faced with this, testing only the model amounts to reinforcing the front door while leaving the garage open. AI-Infra-Guard claims to scan the garage, the yard, and the basement. Let's see what it's really worth.
Key takeaways
- AI-Infra-Guard is an open source (Apache 2.0) AI red teaming platform developed by Tencent's Zhuque Lab: 6,116 stars and 569 forks on GitHub as of October 1, 2026.
- It covers 4 attack layers — infrastructure, protocol/tools, agent behavior, model — via 5 modules: ClawScan, Agent Scan, Skills/MCP Scan, AI Infra Scan and Jailbreak Evaluation.
- Current database: 146+ AI components and over 2,000 CVE rules (v4.6.2, September 2026), ranging from Ollama to vLLM, including n8n, ComfyUI and Triton.
- It's the first open source MCP/skills supply chain scanner, capable of detecting semantic tool poisoning and code integrity issues in third-party skills.
- Official warning: no authentication. Deploy strictly internally, never on a public network.
Recommended tools
| Tool | Main use | Price (October 2026) | Best for |
|---|---|---|---|
| AI-Infra-Guard | Full-stack red teaming: infra, MCP/skills, agents, models | Free (Apache 2.0) | Teams deploying agents in production |
| PyRIT | LLM model red teaming (Microsoft) | Free, open source | Model-level robustness testing |
| Garak | LLM vulnerability probes (NVIDIA) | Free, open source | Security benchmarking of a standalone LLM |
| Promptfoo | Prompt evals and testing | Free, open source | Product teams iterating on prompts |
| Hostinger | Host an isolated scanning lab (VPS) | From a few €/month (check hostinger.fr) | SMEs without a dedicated internal network |
What is AI-Infra-Guard, and why is everyone talking about it?
AI-Infra-Guard is an AI red teaming platform that audits the entire stack of an agentic system — not just the model. It is developed by Zhuque Lab, the security team created by Tencent in 2019, which has discovered high-risk vulnerabilities at NVIDIA, Google, Microsoft, and in the OpenClaw, Linux, and Hugging Face communities.
In just a few months, the project went from niche tool to industry reference. Its listing appears on Kitploit, and its release cadence is remarkable: eight releases between March and September 2026, with detection rules added every month.
This momentum fits into a broader trend: after Alphabet fully open-sourcing its robotics stack with Intrinsic Core under Apache 2.0, open source is now tackling the most critical layers of AI — and security is the most critical of all. My take: this is exactly the kind of tool that can only exist if it's open source, because nobody accepts auditing their own security with a black box.
Key release timeline
| Version | Date | Main addition |
|---|---|---|
| v4.1 | 23/03/2026 | OpenClaw vulnerability database expanded to 281 CVE/GHSA entries |
| v4.1.1 | 25/03/2026 | Detection of the LiteLLM supply chain attack (CRITICAL) |
| v4.1.9 | 21/05/2026 | 26 Prompt Security attack operators (20 single-turn, 6 multi-turn) |
| v4.5.0 | 27/07/2026 | AI Security Skill Market, open source frontend, aig-skill-scan CLI |
| v4.5.1 | 30/07/2026 | Multi-turn jailbreaks (Many-Shot, PAIR, GOAT, ActorAttack), 5 OWASP skills |
| v4.5.2 | 17/08/2026 | Detection of .pyc bytecode bypass, RCE prevention in dynamic MCP mode |
| v4.6.0 | 26/08/2026 | Detection of LLM API poisoning, 146 components / 2,000+ rules |
| v4.6.2 | 17/09/2026 | +155 CVE rules across 40+ components, FORGE-Bench and RogueHandoff-20 benchmarks |
Source: detailed changelog and the official repository.
The 4 attack layers: the "layer-paradigm matching" principle
An AI agent exposes four distinct attack surfaces, and each requires a different detection method. This is the central thesis of the team's technical report (arXiv 2606.31227), organized around "layer-paradigm matching".
- Infrastructure: deterministic rule matching. The paper describes 75+ components and 1,400+ rules; the database has grown to 146 components and 2,000+ CVE rules since v4.6.0 (August 2026).
- Protocol/tools: LLM-driven agentic audit of MCP servers and skill packages.
- Agent behavior: multi-turn black-box red teaming, to test what the agent actually does — not what it should do.
- Model: jailbreak harness with 26+ attack operators evaluated across 16 datasets.
What the architecture tells us
The system is built on a distributed server-agent architecture: infrastructure scanning in Go, LLM modules in Python, WebSocket communication. Concretely, this means heavy scans can be spread across multiple machines — a detail that matters when auditing an entire fleet of agents rather than a prototype.
This layering is, to my knowledge, what most sets AI-Infra-Guard apart from the previous generation of tools. Jailbreak frameworks test layer 4. The 2026 incidents play out on layers 1 through 3.
The 5 scan modules, in detail
Five modules cover the four layers, each with a precise scope.
| Module | What it audits | Key points |
|---|---|---|
| AI Infra Scan | AI infrastructure components | 146+ components, 2,000+ CVE rules: Ollama, ComfyUI, vLLM, n8n, Triton… |
| ClawScan | OpenClaw Security Scan | OpenClaw vulnerability database: 281 CVE/GHSA entries (v4.1, March 2026) |
| Agent Scan | Multi-agent frameworks | Dify/Coze workflows, 5 OWASP skills (v4.5.1), FORGE-Bench and RogueHandoff-20 benchmarks |
| Skills/MCP Scan | MCP servers and third-party skills | 14 risk categories, semantic tool poisoning, web exfiltration detection |
| Jailbreak Evaluation | Model robustness | 26+ attack operators, 16 datasets, multi-turn jailbreaks |
Three modules deserve a closer look.
AI Infra Scan inventories exposed AI components and cross-references them against known CVEs. If you run open source AI agents with Ollama locally, this is the module that will tell you whether your instance is harboring a known vulnerability — and unfortunately, that's all too common.
Agent Scan targets multi-agent workflows, particularly those built on Dify and Coze. The FORGE-Bench and RogueHandoff-20 benchmarks, added in v4.6.2 (September 17, 2026), measure agent loss of control: the scenario where the agent steps outside its intended scope. This is exactly the type of drift observed in recent incidents.
Jailbreak Evaluation ships with 26+ attack operators across 16 datasets — including cnsafe (3,030 prompts), JailBench-Tiny (133), and JADE-db-v3.0 (122). Since v4.5.1 (July 30, 2026), it covers multi-turn jailbreaks: Many-Shot, PAIR, GOAT, and ActorAttack. Multi-turn attacks are the ones that work best against agents, precisely because they mimic a legitimate conversation.
MCP and skills: the supply chain scan the ecosystem was missing
AI-Infra-Guard is the first open source scanner capable of auditing MCP servers and agent skill packages for adversarial content. That's the takeaway from The New Claw Times, and it's well-founded: PyRIT, Garak and Promptfoo all stop at the model layer.
Why it matters: a third-party MCP skill contains natural language descriptions that your LLM reads and executes. An attacker can hide hijacked instructions in them — semantic "tool poisoning" — or malicious code. The agent executes, and nobody notices a thing.
The scanner detects these risks via a T01–T09 taxonomy: instruction hijacking, memory poisoning, payload download, privilege escalation, unsafe dependencies… The LiteLLM supply chain attack, detected as early as v4.1.1 (March 25, 2026) and rated CRITICAL, proves the threat isn't theoretical. v4.5.2 added .pyc bytecode bypass detection and RCE prevention through tool whitelisting in dynamic MCP mode.
If you're looking to structurally reduce this attack surface, one avenue is to step outside the "all-LLM" paradigm: Laya, the open source 421M-parameter decision model, proposes replacing the LLM in agent decision loops. Less natural language ingested, less injection surface.
Integrating AI-Infra-Guard into an agent audit pipeline
Integration happens in four steps, from infrastructure to model. Deployment holds no surprises: Docker image, web interface on localhost:8088, 4 GB of RAM minimum. For skill scanning, a standalone CLI has been available since v4.5.0:
pip install aig-skill-scan
The pipeline I recommend:
- Map the infrastructure (AI Infra Scan): inventory of exposed components, cross-checked against 2,000+ CVE rules.
- Audit the MCP/skill supply chain: every MCP server and every third-party skill goes through the scanner, locally.
- Red team the behavior (Agent Scan): Dify/Coze workflows, multi-agent scenarios, loss-of-control benchmarks.
- Evaluate the model (Jailbreak Evaluation): single-turn and multi-turn operators on your real system prompts.
And above all: automate. Between v4.1 (March 23, 2026) and v4.6.2 (September 17, 2026), the rules database grew by several hundred entries. A scan that's three months old is a dead scan. Hook the skill-scan CLI into your CI so it runs every time a skill or MCP server is added.
The warning to know before installing
The official warning is blunt: the platform has no authentication mechanism. As the analysis by METAL LAB points out, the scanning tool itself can become an entry point if poorly deployed. The recommendation is clear: internal network isolation or VPN, and never public exposure.
Securing without breaking performance
An audit that results in heavy-handed guardrails often kills agent performance. It's a false dilemma: runtime approaches like Life-Harness, which boosts LLM agents by 88.5% without retraining, show that you can harden and speed up at the same time. Audit first, optimize second.
Which LLM to Run Behind the Scanner?
The agentic audit engine needs an LLM to analyze skills and drive attack scenarios — and the SkillTrustBench results from the official repository show that the choice matters, without being decisive. All tested models exceed an F1 of 0.97 on skill risk detection.
| LLM (audit engine) | F1 (SkillTrustBench) |
|---|---|
| Claude Opus 4.6 | 0.9848 |
| GLM 5.1 | 0.9836 |
| Gemini 3.5 Flash | 0.9792 |
| Kimi 2.6 | 0.9780 |
| DeepSeek v4 Flash | 0.9740 |
Claude Opus 4.6 leads, but the gap with GLM 5.1 — which can run self-hosted — is 0.0012. For sensitive data, a self-hosted audit engine with GLM or Kimi is a credible option; our comparison of the best LLMs to run locally details the required configurations.
Important detail: the scanner's LLM is a choice distinct from that of your production agents. If your stack runs on GPT-5.5 or Claude Opus 4.7, the scanner can rely on an entirely different model. Our guide to the best LLMs for AI agents and the monthly comparison of the best LLMs will help you decide on both fronts.
Facing PyRIT, Garak and Promptfoo: a complement, not a competitor
AI-Infra-Guard does not replace model red teaming tools: it covers what they ignore. The three market references stop at the model layer, while three of AI-Infra-Guard's five pillars target the layers around it — agents, skills, MCP.
| Tool | Vendor | Layers covered |
|---|---|---|
| AI-Infra-Guard | Tencent Zhuque Lab | Infra + protocol/tools + agent + model |
| PyRIT | Microsoft | Model |
| Garak | NVIDIA | Model |
| Promptfoo | Promptfoo | Model / prompts |
The industry context is pushing in the same direction: between the safeguards negotiated in the White House agreement and the safety platforms of the major labs — our analysis of OpenShell and Sentry, NVIDIA's agentic safety platform — securing agents has become a political issue as much as a technical one. The difference is that AI-Infra-Guard hands you the keys: everything runs on your premises, and the code is auditable line by line.
My positioning: keep Garak or PyRIT to stress-test your models, and add AI-Infra-Guard for everything that surrounds them. Both approaches are free; there's no reason to choose.
Limitations: What AI-Infra-Guard Doesn't Solve
Three serious limitations to know before adopting.
Zero authentication. This is the most ironic blind spot: a security tool that cannot itself be secured through authentication. The answer is architectural — dedicated VLAN, VPN, localhost binding — but it rests on you, not on the tool.
The trust question. The project is backed by Tencent, and the scanner reads your MCP configurations, your skills, and potentially your system prompts. The code is under Apache 2.0 and therefore auditable, and a deployment on an isolated network with a self-hosted LLM (GLM 5.1, Kimi 2.6) limits leaks. For a regulated environment, this decision deserves a formal security review.
A scan is not a pentest. The 2,000+ rules cover known vectors. The FORGE-Bench and RogueHandoff-20 benchmarks broaden loss-of-control detection, but no scanner replaces a human testing your agent's business logic.
❌ Common Mistakes
Mistake 1: Exposing the scan interface on an accessible server
Without authentication, anyone who reaches port 8088 controls your scanning tool — and through it, access to your configurations. Solution: localhost binding, access via SSH tunnel or VPN, dedicated VLAN if you scan continuously.
Mistake 2: Running only the jailbreak evaluation
This is the most common mistake, because it's the best-known module. Yet three of the five pillars target the layers around the model. Solution: all five modules, in the order infra → MCP/skills → agents → model.
Mistake 3: Treating the scan as a one-shot
The rule base evolves every month — 155 new CVE rules across 40+ components in v4.6.2 alone (September 17, 2026) — and your skills change every week. Solution: automatic re-scan in CI whenever a skill, MCP server, or component version changes.
Mistake 4: Ignoring "minor" dependencies
The LiteLLM attack detected in March 2026 didn't target the teams' own code, but a cross-cutting dependency. Solution: systematic scanning of every new addition, and a tool whitelist in dynamic MCP mode to prevent RCEs.
❓ Frequently Asked Questions
Is AI-Infra-Guard really free?
Yes. The project is Apache 2.0 licensed, self-hosted, with no paid version. The frontend was open-sourced with v4.5.0 (July 2026), and the skill scanner has a standalone CLI installable via pip. The only real costs: the machine hosting the scan (4 GB RAM minimum) and the LLM calls of the audit engine.
Which AI components are covered by the infra scan?
More than 146 components and 2,000+ CVE rules (v4.6.2, September 2026): Ollama, ComfyUI, vLLM, n8n, Triton, among others. The database also includes OpenClaw vulnerabilities (281 CVE/GHSA entries since March 2026) and supply chain attacks like LiteLLM.
Can it replace a classic pentest?
No. It automates the detection of known vectors across four layers, something a manual pentest never covers exhaustively. But an agent's business logic, abuse scenarios specific to your domain, and 0-days remain the domain of human review. The two approaches complement each other.
Is it risky to use a tool developed by Tencent?
The code is auditable (Apache 2.0) and deployment is local. The real risk is functional: the scanner accesses your MCP configs and skills. Deploy it on an isolated network, with a self-hosted LLM if the data is sensitive, and have the code audited for a regulated environment.
How does it differ from PyRIT or Garak?
PyRIT (Microsoft) and Garak (NVIDIA) test model robustness: jailbreaks, bias, data leaks. AI-Infra-Guard also covers infrastructure, MCP servers, third-party skills, and multi-agent behavior. It's the first open source MCP supply chain scanner — the blind spot of the three competing tools.
✅ Conclusion
For any team deploying agents in 2026, scanning the MCP/skills chain is no longer an option but a prerequisite for going to production — and AI-Infra-Guard is currently the most comprehensive open source tool for the job. Install it on an isolated network, run the five modules, and if you're starting from scratch, our guide to the best autonomous AI agents will help you build a stack worth auditing.