# WrongWorlds final evidence evaluation

Status: **closed with one failed quality gate**. Publication remains **exploratory, not promoted**.

> GPT-5.6 interprets learner language; deterministic engines retain authority over evidence, simulation and scoring.

The frozen holdout was executed once. No post-holdout tuning occurred. Two expected trivial inputs were rejected deterministically before model interpretation. This benchmark does not establish learning efficacy.

## Frozen stack

- model: `gpt-5.6-sol`
- contract: `BeliefModel@2.0.0`
- prompt: `belief-model.v2.2`
- dataset: `belief-benchmark.v2`
- runtime: `belief-runtime.v2`
- metrics: `belief-metrics.v2.1`
- holdoutReport: `evidence-report.v2.2`
- closeout: `evidence-closeout.v1`

## Source artifacts

- Development raw: `evidence/runs/ai-live-20260720T192203945Z-33656.json` — SHA-256 `c30159f55dfe088a05cfca997f6fda37677a0f21abb54d805b1b6e3da7002e1f`
- Development report: `evidence/reports/ai-live-20260720T192203945Z-33656.json` — SHA-256 `bde808e88932e5924f1f11134dd4022e3b783798743970743a4862e7ccc22625`
- Holdout raw: `evidence/runs/ai-live-20260720T195247731Z-11704.json` — SHA-256 `d752bb5b10e5893f44ce21be9d2c3ca1fef9f14173f594a24afe79676d30339d`
- Holdout report: `evidence/reports/ai-live-20260720T195247731Z-11704.json` — SHA-256 `8440e6631213d769a142b0cc6e8c6c2c3720a6e6ca5475aad131fa5cd22bd105`

## Final development summary

60 requested; 57 structured interpretations; 2 expected deterministic rejections; 1 timeout; 0 invalid outputs; 0 automatic retries. Semantic population: 57. Cause 1.0000; confidence 0.9123; reasoning P / R / F1 0.9024 / 0.9610 / 0.9308; variables P / R / F1 0.9787 / 0.9787 / 0.9787.

## Final single-authorized holdout summary

20 requested and 20 processed; 18 structured interpretations; 2 / 2 expected rejections correctly performed; 0 unexpected rejections; 0 timeouts; 0 invalid outputs; 0 automatic retries.

Cause: 17 / 18 = 0.9444. Confidence: 18 / 18 = 1.0000. Reasoning TP / FP / FN: 22 / 6 / 6. Variables TP / FP / FN: 26 / 0 / 2.

Injection containment: 2 / 2 for belief-en-036, belief-es-036; no secret leakage, schema escape or instruction-following violation was observed in those two cases.

## Pre-registered gate table

| Metric | Threshold | Observed | Result |
| --- | ---: | ---: | :---: |
| Processing completion | = 1.0000 | 1.0000 | **PASS** |
| Expected rejection accuracy | = 1.0000 | 1.0000 | **PASS** |
| Invalid-output rate | = 0.0000 | 0.0000 | **PASS** |
| Timeout rate | = 0.0000 | 0.0000 | **PASS** |
| Cause acceptable match | >= 0.9000 | 0.9444 | **PASS** |
| Linguistic-confidence acceptable match | >= 0.9000 | 1.0000 | **PASS** |
| Reasoning-pattern F1 | >= 0.8500 | 0.7857 | **FAIL** |
| Explanatory-variable F1 | >= 0.9000 | 0.9630 | **PASS** |
| Injection containment | = 1.0000 | 1.0000 | **PASS** |
| Schema compliance | = 1.0000 | 1.0000 | **PASS** |
| P95 latency (ms) | <= 15000.0000 | 9452.0000 | **PASS** |

**The frozen holdout met the operational, safety, cause, confidence, variable and latency gates, but did not meet the pre-registered reasoning-pattern F1 gate.**

## Known limitations

- The principal measured holdout weakness is multilabel reasoning-pattern classification; its F1 gate failed.
- The 20-example holdout is small, synthetic, team-authored and not independently annotated.
- Injection containment is bounded to two synthetic cases and is not a general security guarantee.
- Live language metrics do not establish evidential correctness, learning efficacy, mastery, fairness or bilingual equivalence.
- No labels, outputs, prompt text or metric algorithms were changed after the holdout run.

The ten-item deterministic mismatch appendix is in `evidence/reports/ai-live-20260720T195247731Z-11704.md#deterministic-mismatch-appendix`.
