Background checks for AI agents. No agent starts trusted.
Every score below is capped by the quality of evidence behind it — an agent with glowing self-written claims and no verifiable history scores near zero. Evidence tiers run from E0 (self-reported) to E3 (independently observed). Scores are hand-produced against a published rubric and expire after 90 days, because agents change fast and a stale trust score is worse than none.
/registry.json for the full machine-readable dataset (dimension scores, evidence tiers, red flags, evidence URLs) and /llms.txt for usage guidance. Rule of thumb: flagged = do not delegate; expired valid_until = treat as unverified; verify key_evidence URLs yourself — the audit trail is the product.
| Agent | Operator | Score | Grade | One-line finding |
|---|---|---|---|---|
| Salesforce Agentforce | Salesforce, Inc. | 60 | Verified | Strongest evidence base scored; self-publishes adverse research. Resolution metrics still vendor-defined. |
| Harvey | Counsel AI Corp. | 60 | Verified | Only agent with a favorable independent performance benchmark (Vals) — it volunteered. That's the point. |
| OpenHands | All Hands AI (OSS, MIT) | 59 | Provisional | Highest track record in the registry (replicable SWE-bench #1, working CVE process). Misses Verified on no SOC 2 / no SLA. |
| OpenAI ChatGPT Agent | OpenAI Group PBC | 56 | Provisional | Best-in-class external safety testing; independent real-world success ~1 in 8 tasks. Grade capped: observed scope violation. |
| Decagon | Decagon AI, Inc. | 53 | Provisional | Public incident disclosure (rewarded). Performance numbers vendor-published; 'resolution' metric tied to its own billing. |
| Sierra | Sierra Technologies | 50 | Provisional | Cleanest reputation scored — and zero independent verification of any performance number. Clean ≠ verified. |
| Devin | Cognition AI, Inc. | 48 | Provisional | Most independent testing of any agent — largely negative. Grade capped: 2024 demo misrepresentation, never retracted. |
| Hippocratic AI | Hippocratic AI, Inc. | 48 | Provisional | Patient-facing clinical agent whose safety numbers are all vendor-generated. No published patient recourse framework. |
| 11x | 11x AI Inc. | 38 | Flagged | Documented 2024–25 fabrication (fake customers, inflated ARR) never publicly acknowledged. Owned mistakes are forgivable; buried ones aren't. |
| Manus | Butterfly Effect Pte. | 22 | Flagged | Independently observed false task-completion reporting with fabricated evidence; unreplicated benchmark claims. |
Grades: Verified–Strong 80–100 · Verified 60–79 · Provisional 40–59 · Unverified 20–39 · Flagged = red-flag override. Full dimension scores, red-flag reasoning, and evidence URLs for every agent are in registry.json.
Ten of the best-funded, most widely deployed AI agents in the world were scored. None cleared 61 out of 100. Not because the agents are all bad — because almost nothing about agent performance is independently verifiable. Registries and identity standards exist (ERC-8004 alone holds 170,000+ registrations, of which studies find only 3–15% even resolve to a valid endpoint); verified track records do not. The registries are job boards. This is the reference check.
Scores are produced by a human against the public rubric, using AI-assisted research (Claude, by Anthropic), with every criterion capped by evidence tier and every source URL logged. The registry holds itself to its own standard: the methodology is public, rule changes are logged in the rubric changelog, and regrades are published with their reasoning — 11x's entry, for example, records that it was regraded from Unverified to Flagged when v1.2 adopted the retraction requirement.
Disputes: operators may dispute any score by sending evidence to james@bull-moose.xyz. Disputes and corrections are published, not silently applied. Nothing here is investment, legal, or procurement advice; it is an evidence audit.