ModelTrust
Menu

Independent AI API verification

Verify an AI API before you trust it.

Independently test third-party AI endpoints for reliability, protocol integrity and behavioral consistency.

LEVEL 1 / AVAILABLE● LIVE

L1

Operational integrity

Three bounded requests with observed availability, latency, schema, identifier, and usage evidence.

NO IDENTITY VERDICT · NO OVERALL SCORE

From endpoint to evidence

01

Submit safely

Provide a public HTTPS Base URL, claimed model, and a per-run API key.

02

Run bounded probes

Level 1 sends three small requests with strict timeouts and response limits.

03

Read the evidence

Review availability, latency, schema, identifiers, finish reasons, and usage telemetry.

What the current check tests

  • Endpoint reachability
  • Repeated request reliability
  • Latency range and median
  • Response schema integrity
  • Returned model identifier
  • Usage and finish-reason consistency

Built for decisions about third-party AI access

Developers

Validate integration behavior before shipping.

Procurement teams

Ask for evidence beyond a model-name claim.

AI platforms

Screen compatible endpoints consistently.

Researchers

Use versioned methods with explicit limits.

API providers

Diagnose what clients can actually observe.

LEVEL 1 · AVAILABLE

Available now

Level 1 Operational Integrity is live. It measures observable endpoint behavior without converting it into a model identity claim or an overall trust score.

REPORT / NORMALIZED EVIDENCE

Evidence you can inspect

A temporary, secret-free report records normalized measurements and the benchmark version. It never stores the API key or full response text.

Open example report

A verification hierarchy with visible maturity

L0ReachabilityAvailable
L1Operational integrityAvailable
L2Behavioral consistencyResearch Preview
L3Model identity confidenceResearch
L4Continuous monitoringFuture

Versioned methods, not hidden scores

Benchmark definitions identify the exact sampling, timeouts, evidence fields, and evaluator maturity used for a report. Behavioral comparison remains experimental until a real official reference and candidate run passes review.

INDEPENDENCE

Independent by design

Pay to test. Never pay to rank. Providers cannot purchase favorable methodology treatment, and observed differences are evidence—not accusations.

Frequently asked questions

Can ModelTrust prove which model served a response?

No. Black-box endpoint testing can provide probabilistic evidence, but it cannot prove the exact underlying model weights.

Does ModelTrust store API keys?

No. Keys are transient, server-side inputs for one run and are not stored, logged, returned, or sent to analytics.

Is behavioral consistency available?

The architecture and evaluators exist, but Level 2 is a Research Preview until a real official reference/candidate production run passes the publication gate.

How many requests does the public check send?

The available Level 1 check sends three bounded requests, each with a 10-second timeout.

How long are reports kept?

Temporary reports retain normalized, secret-free evidence for 30 minutes.

Check the endpoint before the integration depends on it.

Run the available Level 1 check, then decide what further evidence your use case requires.

Check an API