Disclosure policy
Six probes expose public methodology summaries and six are marked reserved. This release does not implement complex secret-probe infrastructure; the split documents a future anti-gaming boundary.
Benchmark registry
Every ModelTrust report retains the benchmark and probe-set versions that produced it. Behavioral results remain experimental and are not model identity probabilities.
EXPERIMENTAL · IMPLEMENTED, AWAITING FIRST OFFICIAL REFERENCE RUN
Release date
2026-08-15
Probe count
12
Sampling
3 reference / 1 candidate per probe
Reference architecture
OpenAI and DeepSeek configured; Anthropic adapter registered (2/3 active)
Methodology version
mt-probes-0.4.0
Six probes expose public methodology summaries and six are marked reserved. This release does not implement complex secret-probe infrastructure; the split documents a future anti-gaming boundary.
A high behavioral-consistency score means the candidate behaved similarly to cached official reference samples on these evaluators. It cannot establish weights, routing, quantization, fine-tuning, system prompts, or provider intent.
mt-bench-0.2.0 remains the active production Level 1 protocol until the first real official reference/candidate comparison is verified in production.