Release article / project log
v0.4.0In progress
Reference Benchmark & Behavioral Consistency
The comparison engine, benchmark registry, bilingual evidence UI, and secret-free cache architecture are implemented. The release remains in progress because no official reference credential is configured and no real end-to-end comparison has been verified.
What changed
- ✓Defined 12 immutable probes across instruction, reasoning, structured output, tool calling, and behavioral-feature categories.
- ✓Registered configurable OpenAI, Anthropic, and DeepSeek official references with provider-specific adapters.
- ✓Implemented exact, constraint, JSON-schema, numeric, tool-call, and feature-vector evaluators plus bounded candidate comparison.
- ✓Added a compatibility-keyed 30-day reference cache, secret-free Level 2 report schema, and bilingual Benchmarks page.
Verification evidence
- ●Deterministic evaluator and mock-comparison tests
- ●Repeatability, partial-failure, timeout, and type-contract checks
- ●Level 2 missing-reference gate
Next target
Configure one official provider key and exact model ID, complete the first private run, then repeat it across separate time windows before considering this release shipped.
Release boundary
This article records shipped work and explicit limitations. Experimental results are not promoted by a patch release.