OWASP-mapped · Agentic + LLM Top 10
A trust score for every AI agent.
OpenTrustBench is a free, local-first security scanner for AI agents and MCP servers. Eight static rules, a permission manifest, and a Trust Card graded A–F — in seconds. Local analysis; no OpenTrustBench telemetry.
Free forever · No account · Open source (Apache-2.0) · Install in 60 seconds
How it works
Three steps to a verifiable grade.
Point at anything
A local directory, a GitHub URL, or an npm package. Recursive detection, eight rules, permission extraction, dependency audit. Seconds, locally.
Get a Trust Card
Letter grade A–F, 0–100 score, OWASP-mapped findings with file:line, plus SARIF for your Security tab and a --fail-on CI gate.
Share the evidence
Embed a bound badge that links to a public report page — grade image and evidence travel together. See the live registry.
The registry · snapshot 2026-09-18 · re-scanned weekly
Strict grades, honestly withheld.
One external target graded; 49 withheld as U for limited static coverage. Full evidence behind every card.
| A | mcp-cli | 96 |
| A | opentrustbench CLI (self) | 90 |
| F | vulnerable fixture (self) | 33 |
49 of 50 externals ungraded.
U means limited static coverage — unsupported files present — not a zero and not a safety verdict. Ungraded cards keep their findings but get no score and stay out of rankings and averages. Filter the registry by U →
Re-scan anything.
Every card records its upstream commit sha, so any grade can be re-verified exactly. How scoring works · Explore the registry →
Install
Running in 60 seconds.
npm install -g @opentrustbench/cli opentrustbench scan ./my-mcp-server --quiet
pip install opentrustbench # needs Node 18+ for the CLI engine opentrustbench scan ./my-mcp-server --quiet
brew tap eulogik/opentrustbench brew install opentrustbench opentrustbench scan ./my-mcp-server --quiet
docker pull eulogik/opentrustbench docker run --rm -v $(pwd):/workspace eulogik/opentrustbench scan .
code --install-extension eulogik.opentrustbench
- uses: eulogik/opentrustbench-action@v0.1.3
with:
target: .
fail-on: high
No install? Run it directly with npx @opentrustbench/cli scan . From source: git clone https://github.com/eulogik/OpenTrustBench.git && cd OpenTrustBench && npm install. Enforce it in CI:
opentrustbench scan . --fail-on high --quiet --output-dir ./trust
Or drop in the GitHub Action — scan plus SARIF upload to your Security tab.
Pricing
The CLI is the product. It's free.
$0, forever
Unlimited local scans, Trust Cards, bound badges, SARIF, CI gate. Open source, zero telemetry.
Scoped per target
Manual triage of your results plus a remediation plan, on request. Contact us.
FAQ
Asked, answered.
What is OpenTrustBench?
OpenTrustBench is a free, local-first security scanner for AI agents and MCP servers. It runs 8 OWASP-mapped static rules, extracts a permission manifest, and issues a Trust Card graded A–F — with SARIF output for CI. It answers: can I trust this agent, what may it do, and can I prove it?
How do I scan an MCP server?
Point the CLI at a local directory, a GitHub URL, or an npm package. You get a Trust Card, SARIF and Markdown reports in seconds. Nothing leaves your machine — GitHub targets are cloned to a temp directory locally.
What is a Trust Card?
A portable, machine-readable credential (opentrustbench/trust-card/v2): letter grade, 0–100 score, findings with file:line evidence, permission scope, provenance signals, coverage, and inferred host compatibility. Bound badges link each grade to its public evidence page.
Is OpenTrustBench a certification?
No. OpenTrustBench is a static scanner, not a certification body. Attack analysis is heuristic (no payloads execute), workflow eval is simulated (nothing runs), and Trust Cards are evidence input for your own review — never a compliance verdict.
Does my code leave my machine?
No. The CLI runs locally with zero telemetry and no OpenTrustBench data retention. GitHub-URL scans clone the public repo to a temp directory on your machine. Dependency auditing queries the npm registry.
What do the grades mean?
A (90+) through F (below 40), weighted across security findings, permission scope, provenance, reliability and stability. U means ungraded: coverage was insufficient, so no score was issued — currently 50 of 53 cards. Ungraded cards are excluded from averages and rankings. Full scoring methodology.
Why shouldn't I scan the OpenTrustBench monorepo itself?
Because it contains its own test ammunition — and a scanner that didn't flag it would be the real scandal. The three intentional sources: (1) examples/vulnerable-mcp-server, a deliberately vulnerable fixture; (2) our unit tests, which must contain eval(, exec( and transferFunds( strings to prove the detector catches them; (3) our automation scripts, which legitimately shell out to git and npm with hardcoded local arguments. The monorepo itself currently grades U (60 code files analyzed against 66 unsupported implementation files — limited coverage, so the grade is withheld). The shipped product (packages/cli) self-scans at A (90/100). Rule of thumb: grade the artifact you ship, not the monorepo that tests it.
Ship agents you can defend.
One command. Thirty seconds. A grade you can show your auditor.