Submit safely
Provide a public HTTPS Base URL, claimed model, and a per-run API key.
Independent AI API verification
Independently test third-party AI endpoints for reliability, protocol integrity and behavioral consistency.
L1
Operational integrity
Three bounded requests with observed availability, latency, schema, identifier, and usage evidence.
NO IDENTITY VERDICT · NO OVERALL SCORE
Provide a public HTTPS Base URL, claimed model, and a per-run API key.
Level 1 sends three small requests with strict timeouts and response limits.
Review availability, latency, schema, identifiers, finish reasons, and usage telemetry.
Validate integration behavior before shipping.
Ask for evidence beyond a model-name claim.
Screen compatible endpoints consistently.
Use versioned methods with explicit limits.
Diagnose what clients can actually observe.
LEVEL 1 · AVAILABLE
Level 1 Operational Integrity is live. It measures observable endpoint behavior without converting it into a model identity claim or an overall trust score.
REPORT / NORMALIZED EVIDENCE
A temporary, secret-free report records normalized measurements and the benchmark version. It never stores the API key or full response text.
Open example reportBenchmark definitions identify the exact sampling, timeouts, evidence fields, and evaluator maturity used for a report. Behavioral comparison remains experimental until a real official reference and candidate run passes review.
INDEPENDENCE
Pay to test. Never pay to rank. Providers cannot purchase favorable methodology treatment, and observed differences are evidence—not accusations.
No. Black-box endpoint testing can provide probabilistic evidence, but it cannot prove the exact underlying model weights.
No. Keys are transient, server-side inputs for one run and are not stored, logged, returned, or sent to analytics.
The architecture and evaluators exist, but Level 2 is a Research Preview until a real official reference/candidate production run passes the publication gate.
The available Level 1 check sends three bounded requests, each with a 10-second timeout.
Temporary reports retain normalized, secret-free evidence for 30 minutes.
Run the available Level 1 check, then decide what further evidence your use case requires.