MModelTrust
All updates

Release article / project log

v0.4.0In progress

Reference Benchmark & Behavioral Consistency

The comparison engine, benchmark registry, bilingual evidence UI, and secret-free cache architecture are implemented. The release remains in progress because no official reference credential is configured and no real end-to-end comparison has been verified.

What changed

  • Defined 12 immutable probes across instruction, reasoning, structured output, tool calling, and behavioral-feature categories.
  • Registered configurable OpenAI, Anthropic, and DeepSeek official references with provider-specific adapters.
  • Implemented exact, constraint, JSON-schema, numeric, tool-call, and feature-vector evaluators plus bounded candidate comparison.
  • Added a compatibility-keyed 30-day reference cache, secret-free Level 2 report schema, and bilingual Benchmarks page.

Verification evidence

  • Deterministic evaluator and mock-comparison tests
  • Repeatability, partial-failure, timeout, and type-contract checks
  • Level 2 missing-reference gate

Next target

Configure one official provider key and exact model ID, complete the first private run, then repeat it across separate time windows before considering this release shipped.

Release boundary

This article records shipped work and explicit limitations. Experimental results are not promoted by a patch release.