MModelTrust

Benchmark registry

Versioned tests, visible limits.

Every ModelTrust report retains the benchmark and probe-set versions that produced it. Behavioral results remain experimental and are not model identity probabilities.

EXPERIMENTAL · IMPLEMENTED, AWAITING FIRST OFFICIAL REFERENCE RUN

mt-bench-0.3.0

Release date

2026-08-15

Probe count

12

Sampling

3 reference / 1 candidate per probe

Reference architecture

OpenAI and DeepSeek configured; Anthropic adapter registered (2/3 active)

Methodology version

mt-probes-0.4.0

Included evidence

  • Operational integrity
  • Instruction following
  • Reasoning
  • Structured output
  • Tool calling
  • Behavioral features

Disclosure policy

Six probes expose public methodology summaries and six are marked reserved. This release does not implement complex secret-probe infrastructure; the split documents a future anti-gaming boundary.

Scientific boundary

A high behavioral-consistency score means the candidate behaved similarly to cached official reference samples on these evaluators. It cannot establish weights, routing, quantization, fine-tuning, system prompts, or provider intent.

Production operational benchmark

mt-bench-0.2.0 remains the active production Level 1 protocol until the first real official reference/candidate comparison is verified in production.

Read methodology