# Belief interpretation benchmark — ai-fixture-baseline-v1

- Verification: **verified**
- Mode: **fixture**
- Model: `fixture-belief-v1`
- Prompt: `belief-model.v1`
- Dataset: `belief-benchmark.v1` (80 evaluated results)
- Cost: **unresolved**

> Fixture mode verifies schemas, runners, and reporting. It is not evidence of live model quality.

| Metric | Overall | EN | ES | Absolute EN/ES delta |
| --- | ---: | ---: | ---: | ---: |
| Acceptable cause match | 0.8947 | 0.8684 | 0.9211 | 0.0526 |
| Acceptable strength match | 0.5789 | 0.5789 | 0.5789 | 0.0000 |
| Reasoning-pattern F1 | 0.4639 | 0.4632 | 0.4646 | 0.0015 |
| Relevant-variable F1 | 0.7605 | 0.7692 | 0.7519 | 0.0174 |

Latency median/p90/p95: 0.0000 / 0.0000 / 0.0000 ms.
