OWASP-mapped · Agentic + LLM Top 10

A trust score for every AI agent.

OpenTrustBench is a free, local-first security scanner for AI agents and MCP servers. Eight static rules, a permission manifest, and a Trust Card graded A–F — in seconds. Local analysis; no OpenTrustBench telemetry.

npx @opentrustbench/cli scan . View on GitHub

Free forever · No account · Open source (Apache-2.0) · Install in 60 seconds

53
Trust Cards published — 50 external targets + 3 self-scans · browse the registry
3
graded A–F; 50 withheld as U where coverage was insufficient
8
OWASP-mapped static rules behind every card
0
bytes of your code sent to OpenTrustBench
terminal — sample output, secure fixture (Grade U)
Recorded OpenTrustBench scan running in a terminal (7 seconds)
A recorded run (7s). The typed terminal above illustrates the same output.

How it works

Three steps to a verifiable grade.

01 — SCAN

Point at anything

A local directory, a GitHub URL, or an npm package. Recursive detection, eight rules, permission extraction, dependency audit. Seconds, locally.

02 — GRADE

Get a Trust Card

Letter grade A–F, 0–100 score, OWASP-mapped findings with file:line, plus SARIF for your Security tab and a --fail-on CI gate.

03 — PROVE

Share the evidence

Embed a bound badge that links to a public report page — grade image and evidence travel together. See the live registry.

OpenTrustBench grade A badge for the shipped CLI

The registry · snapshot 2026-09-18 · re-scanned weekly

Strict grades, honestly withheld.

One external target graded; 49 withheld as U for limited static coverage. Full evidence behind every card.

WITHHELD (U)

49 of 50 externals ungraded.

U means limited static coverage — unsupported files present — not a zero and not a safety verdict. Ungraded cards keep their findings but get no score and stay out of rankings and averages. Filter the registry by U →

REPRODUCE

Re-scan anything.

Every card records its upstream commit sha, so any grade can be re-verified exactly. How scoring works · Explore the registry →

Install

Running in 60 seconds.

npm install -g @opentrustbench/cli
opentrustbench scan ./my-mcp-server --quiet

No install? Run it directly with npx @opentrustbench/cli scan . From source: git clone https://github.com/eulogik/OpenTrustBench.git && cd OpenTrustBench && npm install. Enforce it in CI:

opentrustbench scan . --fail-on high --quiet --output-dir ./trust

Or drop in the GitHub Action — scan plus SARIF upload to your Security tab.

Pricing

The CLI is the product. It's free.

CLI

$0, forever

Unlimited local scans, Trust Cards, bound badges, SARIF, CI gate. Open source, zero telemetry.

EXPERT REVIEW

Scoped per target

Manual triage of your results plus a remediation plan, on request. Contact us.

ENTERPRISE

Custom

Custom rules, evidence helpers, priority support. Contact us.

FAQ

Asked, answered.

What is OpenTrustBench?

OpenTrustBench is a free, local-first security scanner for AI agents and MCP servers. It runs 8 OWASP-mapped static rules, extracts a permission manifest, and issues a Trust Card graded A–F — with SARIF output for CI. It answers: can I trust this agent, what may it do, and can I prove it?

How do I scan an MCP server?

Point the CLI at a local directory, a GitHub URL, or an npm package. You get a Trust Card, SARIF and Markdown reports in seconds. Nothing leaves your machine — GitHub targets are cloned to a temp directory locally.

What is a Trust Card?

A portable, machine-readable credential (opentrustbench/trust-card/v2): letter grade, 0–100 score, findings with file:line evidence, permission scope, provenance signals, coverage, and inferred host compatibility. Bound badges link each grade to its public evidence page.

Is OpenTrustBench a certification?

No. OpenTrustBench is a static scanner, not a certification body. Attack analysis is heuristic (no payloads execute), workflow eval is simulated (nothing runs), and Trust Cards are evidence input for your own review — never a compliance verdict.

Does my code leave my machine?

No. The CLI runs locally with zero telemetry and no OpenTrustBench data retention. GitHub-URL scans clone the public repo to a temp directory on your machine. Dependency auditing queries the npm registry.

What do the grades mean?

A (90+) through F (below 40), weighted across security findings, permission scope, provenance, reliability and stability. U means ungraded: coverage was insufficient, so no score was issued — currently 50 of 53 cards. Ungraded cards are excluded from averages and rankings. Full scoring methodology.

Why shouldn't I scan the OpenTrustBench monorepo itself?

Because it contains its own test ammunition — and a scanner that didn't flag it would be the real scandal. The three intentional sources: (1) examples/vulnerable-mcp-server, a deliberately vulnerable fixture; (2) our unit tests, which must contain eval(, exec( and transferFunds( strings to prove the detector catches them; (3) our automation scripts, which legitimately shell out to git and npm with hardcoded local arguments. The monorepo itself currently grades U (60 code files analyzed against 66 unsupported implementation files — limited coverage, so the grade is withheld). The shipped product (packages/cli) self-scans at A (90/100). Rule of thumb: grade the artifact you ship, not the monorepo that tests it.

Ship agents you can defend.

One command. Thirty seconds. A grade you can show your auditor.

npx @opentrustbench/cli scan . Browse the registry