ModelTrust
Menu

Methodology / mt-bench-0.4.0 experimental

Evidence with the uncertainty left in.

ModelTrust evaluates externally observable behavior. Operational integrity is in production; the v0.4 comparison engine is implemented but remains experimental until a real official reference/candidate run is verified.

Interpretation rule

Describe what was observed—never convert operational metrics into a model identity verdict.

A provider may route, quantize, wrap, fine-tune, or rate-limit an endpoint. Model identifiers and response behavior are evidence, not cryptographic proof of the weights that served a response.

Verification levels

Implementation maturity is part of the result.

L0

Reachability

Available

Public HTTPS resolution and safe network reachability.

L1

Operational integrity

Available

Three bounded requests measuring availability, latency, schema validity, returned identifiers, finish reasons, and usage telemetry.

L2

Behavioral consistency

Research Preview

Twelve versioned probes designed to compare official reference samples with candidate samples. No validated production run is published yet.

L3

Identity research

Research

Probabilistic research only. It cannot prove exact underlying model weights.

L4

Longitudinal monitoring

Future

Repeated evidence windows and drift detection; not implemented.

Level 1 protocol

01

Three samples

The current benchmark sends three sequential control requests. The schema permits a future five-sample profile without changing report fields.

02

Bounded execution

Each external request has a 10-second timeout; the complete run is bounded to 35 seconds.

03

Exact control

The endpoint is asked for a short exact control response. Response text is evaluated transiently and is never stored.

04

Normalized evidence

Each sample records timestamp, status, HTTP status, latency, model identifier, finish reason, usage presence and counts, schema validity, and a safe error category.

05

Descriptive aggregation

Reports show availability, min/median/max latency, range-to-median variability, identifier consistency, schema consistency, usage state, and finish reasons.

06

No score

Level 1 metrics are not transformed into an overall trust score, confidence percentage, or model identity claim.

Experimental Level 2 protocol

L2-01

Reference models

Official endpoints are selected through a versioned registry. A configured model ID and server-only provider key define the reference; this trust anchor is documented, not treated as infallible.

L2-02

Probe sets

Twelve immutable low-cost probes cover instruction following, short reasoning, JSON structure, tool calls, and low-weight observable style features.

L2-03

Sampling

Each reference probe uses three samples to reveal variance. The candidate receives the exact same probe definition once, with bounded concurrency, timeout, request count, and output tokens.

L2-04

Variance

Evaluator ranges are retained. Missing samples and unstable reference behavior reduce LOW/MODERATE/HIGH evidence confidence rather than producing invented numerical certainty.

L2-05

Behavioral consistency

The percentage is evaluator similarity under mt-bench-0.4.0. It is not the probability of shared weights or a definitive named-model verdict.

L2-06

Category evaluators

Exact match, constraints, JSON schema, numeric tolerance, tool-call selection/arguments, and feature vectors are used according to probe type; there is no universal embedding score.

L2-07

Limitations

Routing, wrappers, system prompts, safety policies, quantization, fine-tuning, provider updates, and random variation can all change observable behavior.

L2-08

Benchmark gaming

Public methods support scrutiny, while a future rotating reserved set may reduce optimization against fixed fingerprints. v0.4 documents this boundary without building secret infrastructure.

Security and retention

Base URLs must be public HTTPS endpoints. DNS results are checked against private, loopback, link-local, documentation, and reserved address ranges; redirects are blocked. API keys exist only for the run and are not stored, logged, returned, or sent to analytics. Reports retain only the endpoint hostname and normalized secret-free evidence for 30 minutes.

Abuse controls

The public form applies best-effort hashed IP and session windows, a small per-isolate concurrency cap, and bounded run duration. These MVP limits reduce accidental abuse but are not presented as a globally atomic distributed quota.

Reference harness boundary

The registry supports OpenAI, Anthropic, and DeepSeek official endpoint configurations. Reference credentials remain server-only and reference artifacts retain normalized evaluator evidence, never keys or raw prompts/responses. No official run is claimed until one is actually completed.

Benchmark versioning

Every report names an immutable benchmark definition. Sampling, timeouts, included test versions, scoring status, and experimental status change only through a new benchmark ID.

Reproducibility

A report includes its benchmark, protocol, sample count, timeout, timestamp, and ModelTrust version. Exact samples remain bounded observations, not a claim that all future endpoint behavior will match.

Independence

Pay to test. Never pay to rank. Provider relationships cannot change methodology treatment, and observed differences alone are not accusations of deception.

Limitations

Routing, system prompts, wrappers, fine-tuning, quantization, rate limits, and implementation differences can change behavior. Three samples are a diagnostic window, not longitudinal evidence.

External references

These sources inform protocol design and terminology. They are not ModelTrust findings or validation evidence.