Reachability
AvailablePublic HTTPS resolution and safe network reachability.
Methodology / mt-bench-0.4.0 experimental
ModelTrust evaluates externally observable behavior. Operational integrity is in production; the v0.4 comparison engine is implemented but remains experimental until a real official reference/candidate run is verified.
Interpretation rule
Describe what was observed—never convert operational metrics into a model identity verdict.
A provider may route, quantize, wrap, fine-tune, or rate-limit an endpoint. Model identifiers and response behavior are evidence, not cryptographic proof of the weights that served a response.
Implementation maturity is part of the result.
Public HTTPS resolution and safe network reachability.
Three bounded requests measuring availability, latency, schema validity, returned identifiers, finish reasons, and usage telemetry.
Twelve versioned probes designed to compare official reference samples with candidate samples. No validated production run is published yet.
Probabilistic research only. It cannot prove exact underlying model weights.
Repeated evidence windows and drift detection; not implemented.
The current benchmark sends three sequential control requests. The schema permits a future five-sample profile without changing report fields.
Each external request has a 10-second timeout; the complete run is bounded to 35 seconds.
The endpoint is asked for a short exact control response. Response text is evaluated transiently and is never stored.
Each sample records timestamp, status, HTTP status, latency, model identifier, finish reason, usage presence and counts, schema validity, and a safe error category.
Reports show availability, min/median/max latency, range-to-median variability, identifier consistency, schema consistency, usage state, and finish reasons.
Level 1 metrics are not transformed into an overall trust score, confidence percentage, or model identity claim.
Official endpoints are selected through a versioned registry. A configured model ID and server-only provider key define the reference; this trust anchor is documented, not treated as infallible.
Twelve immutable low-cost probes cover instruction following, short reasoning, JSON structure, tool calls, and low-weight observable style features.
Each reference probe uses three samples to reveal variance. The candidate receives the exact same probe definition once, with bounded concurrency, timeout, request count, and output tokens.
Evaluator ranges are retained. Missing samples and unstable reference behavior reduce LOW/MODERATE/HIGH evidence confidence rather than producing invented numerical certainty.
The percentage is evaluator similarity under mt-bench-0.4.0. It is not the probability of shared weights or a definitive named-model verdict.
Exact match, constraints, JSON schema, numeric tolerance, tool-call selection/arguments, and feature vectors are used according to probe type; there is no universal embedding score.
Routing, wrappers, system prompts, safety policies, quantization, fine-tuning, provider updates, and random variation can all change observable behavior.
Public methods support scrutiny, while a future rotating reserved set may reduce optimization against fixed fingerprints. v0.4 documents this boundary without building secret infrastructure.
Base URLs must be public HTTPS endpoints. DNS results are checked against private, loopback, link-local, documentation, and reserved address ranges; redirects are blocked. API keys exist only for the run and are not stored, logged, returned, or sent to analytics. Reports retain only the endpoint hostname and normalized secret-free evidence for 30 minutes.
The public form applies best-effort hashed IP and session windows, a small per-isolate concurrency cap, and bounded run duration. These MVP limits reduce accidental abuse but are not presented as a globally atomic distributed quota.
The registry supports OpenAI, Anthropic, and DeepSeek official endpoint configurations. Reference credentials remain server-only and reference artifacts retain normalized evaluator evidence, never keys or raw prompts/responses. No official run is claimed until one is actually completed.
Every report names an immutable benchmark definition. Sampling, timeouts, included test versions, scoring status, and experimental status change only through a new benchmark ID.
A report includes its benchmark, protocol, sample count, timeout, timestamp, and ModelTrust version. Exact samples remain bounded observations, not a claim that all future endpoint behavior will match.
Pay to test. Never pay to rank. Provider relationships cannot change methodology treatment, and observed differences alone are not accusations of deception.
Routing, system prompts, wrappers, fine-tuning, quantization, rate limits, and implementation differences can change behavior. Three samples are a diagnostic window, not longitudinal evidence.
These sources inform protocol design and terminology. They are not ModelTrust findings or validation evidence.