MModelTrust

Methodology / mt-bench-0.1.0

Evidence with the uncertainty left in.

ModelTrust evaluates externally observable behavior. Black-box identification can estimate consistency with a reference model, but it cannot cryptographically prove which weights served a response.

Interpretation rule

Report statistical confidence and capability integrity—never a binary “real vs fake” verdict.

A provider may route, quantize, wrap, fine-tune, rate-limit, or otherwise alter an endpoint. Observed capability can differ without establishing deception or hidden model identity.

Initial test categories

1 IMPLEMENTED · 3 EXPERIMENTAL · 4 PLACEHOLDERS

01

Behavioral fingerprint

Compare response distributions against versioned reference samples. Experimental; no validated fingerprint ships in v0.2.0.

02

Reasoning capability

Measure retention across bounded, contamination-aware reasoning tasks. Experimental.

03

Instruction-following

Test adherence to conflicting, nested, and structured instructions. Placeholder.

04

Tool / function calling

Inspect schema adherence, argument validity, and call consistency. Placeholder.

05

Context-window behavior

Measure recall, position sensitivity, and degradation under longer contexts. Placeholder.

06

Token accounting

Compare provider-reported usage with independent token estimates and response shape. Experimental.

07

Latency

Measure bounded request completion time at the verification server. Implemented for one control request.

08

Reliability

Aggregate success, error, and timeout rates across repeated samples. Placeholder.

Reproducibility direction

Benchmarks will be versioned. Future reports should publish sampling rules, reference windows, scoring transformations, and confidence intervals without exposing active anti-gaming items.

Independence rule

Providers may pay for tests, never rankings. A future ModelTrust router must remain separated from verification scoring and cannot receive preferential methodology treatment.