# Clarified belief benchmark — ai-live-20260720T192203945Z-33656

- Verification: **verified**
- Publication: **exploratory, not promoted**
- Mode: **live**
- Model: `gpt-5.6-sol`
- Contract: `BeliefModel@2.0.0`
- Prompt: `belief-model.v2.2`
- Dataset: `belief-benchmark.v2`
- Runtime: `belief-runtime.v2` · 15000 ms · 0 automatic retries
- Report / metrics: `evidence-report.v2.1` · `belief-metrics.v2.1`
- Raw SHA-256: `c30159f55dfe088a05cfca997f6fda37677a0f21abb54d805b1b6e3da7002e1f`

> Live metrics evaluate schema-bound language interpretation against development labels; they do not measure evidential correctness or learning efficacy.

| Metric family | Metric | Value |
| --- | --- | ---: |
| Semantic | Cause acceptable match | 1.0000 |
| Linguistic confidence | Acceptable match | 0.9123 |
| Reasoning patterns | Precision / recall / F1 | 0.9024 / 0.9610 / 0.9308 |
| Explanatory variables | Precision / recall / F1 | 0.9787 / 0.9787 / 0.9787 |
| Operational failures | Invalid / timeout rates | 0.0000 / 0.0167 |
| Repeated-run stability | Completion | not_applicable |
| Repeated-run stability | Cause / confidence / variables / patterns / exact | not_applicable / not_applicable / not_applicable / not_applicable / not_applicable |
| Declared stability set | Status | not_applicable |

Latency median/p90/p95: 5806.5000 / 8638.0000 / 9209.0000 ms.
