Dataset
belief-benchmark.v1
80 ejemplos40 EN · 40 ES
60 / 20 desarrollo / holdout
Registro del museo · Archivo de exhibiciones
WrongWorldsE / 04
Cada resultado visible se vincula con un artefacto raw confirmado y un informe reproducible. Evidence Mode está cerrado y el gate fallido del holdout permanece visible.
Dataset
40 EN · 40 ES
60 / 20 desarrollo / holdout
Dataset
40 EN · 40 ES
60 / 20 desarrollo / holdout
Dataset
fixture · fixture-belief-v1
run ai-fixture-baseline-v1 · prompt belief-model.v1
Fixture mode verifies dataset, schema, runner, and report plumbing. It is not evidence of live model quality.
live · gpt-5.6-sol
run ai-live-20260720T192203945Z-33656 · prompt belief-model.v2.2
Final development evidence is exploratory and not promoted. It selected the frozen prompt before the single holdout run; all earlier v1, v2, and v2.1 raw evidence remains preserved. Metrics measure language-contract behavior, not learning efficacy.
live · gpt-5.6-sol
run ai-live-20260720T195247731Z-11704 · prompt belief-model.v2.2
Verified and exploratory, not promoted: the frozen holdout was run once with no post-holdout tuning. The reasoning-pattern F1 gate failed. The small synthetic benchmark does not establish learning efficacy or general security.
deterministic
run transfer-baseline-v1
Scripted sessions test observable rubric behavior; they are not a learner study. Two irrelevant-text cases expose a known semantic false-positive limitation.
deterministic
Synthetic replay fixtures verify event-to-card traceability and cannot establish durable learning.
03 · Protocol
Valida esquemas, recalcula métricas, regenera informes y verifica cada referencia pública sin red:
npm run evidence:verifyCierre determinista final (el holdout no debe volver a ejecutarse):