Dataset
belief-benchmark.v1
80 examples40 EN · 40 ES
60 / 20 development / holdout
Museum record · Exhibit archive
WrongWorldsE / 04
Every visible result below resolves to a committed raw artifact and a reproducible report. Evidence Mode is closed with the failed holdout gate visible.
Dataset
40 EN · 40 ES
60 / 20 development / holdout
Dataset
40 EN · 40 ES
60 / 20 development / holdout
Dataset
fixture · fixture-belief-v1
run ai-fixture-baseline-v1 · prompt belief-model.v1
Fixture mode verifies dataset, schema, runner, and report plumbing. It is not evidence of live model quality.
live · gpt-5.6-sol
run ai-live-20260720T192203945Z-33656 · prompt belief-model.v2.2
Final development evidence is exploratory and not promoted. It selected the frozen prompt before the single holdout run; all earlier v1, v2, and v2.1 raw evidence remains preserved. Metrics measure language-contract behavior, not learning efficacy.
live · gpt-5.6-sol
run ai-live-20260720T195247731Z-11704 · prompt belief-model.v2.2
Verified and exploratory, not promoted: the frozen holdout was run once with no post-holdout tuning. The reasoning-pattern F1 gate failed. The small synthetic benchmark does not establish learning efficacy or general security.
deterministic
run transfer-baseline-v1
Scripted sessions test observable rubric behavior; they are not a learner study. Two irrelevant-text cases expose a known semantic false-positive limitation.
deterministic
Synthetic replay fixtures verify event-to-card traceability and cannot establish durable learning.
03 · Protocol
Validate schemas, recalculate metrics, regenerate reports and verify every public reference offline:
npm run evidence:verifyFinal deterministic closeout (the holdout must not be rerun):