Project log / transparent progress
What shipped, what was verified, what comes next.
ModelTrust is being built in public. This log separates completed work from plans and records the evidence used to call each release shipped.
Current product status
Level 1 · AVAILABLE
Level 1 Operational Integrity is available in production.
Next evidence gate
Level 2 · First real reference/candidate run
Research Preview · publication gate not yet met
Last verified 2026-09-20
Product roadmap
05 MILESTONESBilingual release log
Operational integrity core
Controlled reference harness
Behavioral validation
Release log
UX, SEO & GEO Content Architecture
This website patch clarifies the available Level 1 product and the Level 2 research boundary, adds bilingual discovery content and a publication-ready research registry, and strengthens technical SEO and accessibility.
Open release article →2026-09-20
Completed
- ✓Reorganized the homepage around visitor questions, proof, methodology, maturity, FAQs, and a clear verification action.
- ✓Added bilingual Research and three intent-specific guide routes without publishing fabricated runs.
- ✓Added canonical, hreflang, Open Graph, Twitter, sitemap, robots, llms.txt, breadcrumbs, FAQ, application, and article-ready structured data.
- ✓Improved responsive navigation, form guidance, validation hints, status language, and accessible page landmarks.
Verified
- ●Automated lint, type, unit, and production build checks
- ●English and Chinese route, metadata, structured-data, responsive, and indexability checks
- ●Production domain and Cloudflare deployment verification
Next target
Complete and review the first real official reference/candidate production run before publishing any Level 2 research article.
Reference Benchmark & Behavioral Consistency
The comparison engine, benchmark registry, bilingual evidence UI, and secret-free cache architecture are implemented. The release remains in progress because no official reference credential is configured and no real end-to-end comparison has been verified.
Open release article →2026-08-15
Completed
- ✓Defined 12 immutable probes across instruction, reasoning, structured output, tool calling, and behavioral-feature categories.
- ✓Registered configurable OpenAI, Anthropic, and DeepSeek official references with provider-specific adapters.
- ✓Implemented exact, constraint, JSON-schema, numeric, tool-call, and feature-vector evaluators plus bounded candidate comparison.
- ✓Added a compatibility-keyed 30-day reference cache, secret-free Level 2 report schema, and bilingual Benchmarks page.
Verified
- ●Deterministic evaluator and mock-comparison tests
- ●Repeatability, partial-failure, timeout, and type-contract checks
- ●Level 2 missing-reference gate
Next target
Configure one official provider key and exact model ID, complete the first private run, then repeat it across separate time windows before considering this release shipped.
Production Verification & Release Blog Patch
This patch adds repeatable production checks, an explicitly confirmed optional success canary, clearer report outcomes and expiry evidence, and linkable bilingual release articles. It does not change the Level 2 shipping status.
Open release article →2026-08-15
Completed
- ✓Added a read-only production route/header verifier plus explicitly confirmed SSRF, unauthorized-report, and Level 2 fail-closed checks.
- ✓Added an optional paid success canary that reads credentials only from private environment variables and never prints or serializes them.
- ✓Added evidence-derived report outcome language and an explicit report-expiry timestamp.
- ✓Added build-time generated bilingual release articles with safe runtime fallback, canonical alternates, language-preserving navigation, and sitemap entries.
Verified
- ●40 automated tests plus TypeScript, ESLint, Next.js, and OpenNext builds
- ●Read-only production route, version, and security-header verification
- ●Production SSRF rejection, three-sample 401 report, secret exclusion, and Level 2 missing-reference gate
- ●English, Chinese, desktop, mobile, release-article, and sitemap rendering
Next target
Run the optional private 200-response canary with a valid provider credential, then add Turnstile and platform rate limiting in v0.3.3.
Cloudflare Runtime & Research Operations Patch
This production patch restores Cloudflare-compatible public-endpoint DNS preflight and adds a fail-closed path for publishing validated, secret-free research artifacts. It does not promote the in-progress Level 2 comparison to a shipped capability.
Open release article →2026-08-15
Completed
- ✓Replaced unsupported Workers DNS lookup with dual-stack resolve4/resolve6 checks and valid-address filtering.
- ✓Added safe endpoint-policy reason categories without exposing provider response bodies or credentials.
- ✓Added an explicitly confirmed publisher for immutable references, active pointers, and temporary reports after full artifact validation.
- ✓Kept bilingual benchmark and progress surfaces explicit that v0.4 remains in progress.
Verified
- ●35 automated tests plus TypeScript, ESLint, Next.js, and OpenNext builds
- ●A real production failed-endpoint run normalized three HTTP 401 samples without persisting a key or full Base URL
- ●English, Chinese, desktop, mobile, and Level 2 missing-reference production routes
Next target
Configure an official reference privately and complete a positive-control run before promoting v0.4.0.
Verification Core
ModelTrust now produces a secret-free Level 1 evidence report from three bounded endpoint requests, without inventing a trust or identity score.
Open release article →2026-08-15
Completed
- ✓Implemented repeated operational-integrity sampling with availability, latency, schema, identifier, finish-reason, and usage evidence.
- ✓Rebuilt reports around observed evidence, experimental boundaries, not-verified items, and benchmark metadata.
- ✓Hardened endpoint validation, blocked redirects, and added best-effort submission and concurrency limits.
- ✓Defined a reference-harness contract without shipping fabricated reference distributions.
Verified
- ●Endpoint and verification-engine unit tests
- ●TypeScript, ESLint, Next.js and OpenNext builds
- ●English, Chinese, desktop, mobile, and production routes
Next target
Build and validate the controlled reference-endpoint harness before enabling any Level 2 behavioral comparison.
Bilingual foundation and public project log
ModelTrust now provides complete English and Chinese product routes plus a transparent release and progress record.
Open release article →2026-08-14
Completed
- ✓Added Chinese routes for the landing page, verification form, report, and methodology.
- ✓Added persistent language switching that preserves the current page.
- ✓Added standalone English and Chinese Updates pages backed by one release data source.
- ✓Added bilingual SEO alternates and sitemap entries.
Verified
- ●TypeScript, ESLint, Next.js and OpenNext builds
- ●Desktop and mobile route rendering
- ●Cloudflare production deployment
Next target
Add abuse protection, repeat-sampling reliability tests, and a controlled reference-endpoint harness.
First public Model Verification MVP
The initial product foundation shipped on modeltrust.net with one real OpenAI-compatible connectivity and latency control.
Open release article →2026-08-14
Completed
- ✓Built landing, check, methodology, and dynamic report routes.
- ✓Defined ProviderAdapter and VerificationTest contracts.
- ✓Added server-only transient API-key handling and 30-minute secret-free KV reports.
- ✓Deployed the Next.js application to Cloudflare Workers with a managed custom domain and TLS.
Verified
- ●Production HTTPS and security headers
- ●Server Action to KV report flow
- ●No credential fields in stored reports
Next target
Expand beyond connectivity without presenting unvalidated identity or capability scores.
Want to test the current build?
The live MVP performs three bounded Level 1 requests and reports operational evidence. Behavioral and identity analysis remain explicitly experimental and are not run.