A correct answer is only part of the picture.
How to read exact accuracy, response coverage, and failures across three separate evaluation tracks.
An AI4HSE-Bench v1 score is a summary of choices: what gets measured, which attempts count, and how failures are treated. Those choices should be easy to understand.
Updated 1 October 2026: The benchmark now reports answer-only, reasoning-assisted, and structured-output results separately. See the current track definitions and published results.
Count every attempt
Our primary metric is end-to-end exact accuracy. The full Boolean selection must match the answer key. A partially correct selection is incorrect for this metric, and provider errors, answers that cannot be extracted, and incomplete generations also count as incorrect.
This puts response reliability in the same frame as answer correctness.
Read accuracy with coverage
We also report exact accuracy among valid responses. That answers a narrower question: when a model returns a valid answer, how often is it correct?
Coverage tells us how often that happens. If difficult questions are more likely to produce invalid responses, valid-response accuracy alone can be misleading. When there are no valid responses, that accuracy is undefined, not zero.
Keep tracks separate
Answer-only, reasoning-assisted, and structured-output runs have different output contracts. Answer-only requests a Boolean list. Under protocol v10, knowledge scoring takes the last bracketed list, including a corrected selection, and ignores surrounding explanation. That last list must contain one valid TRUE/FALSE value per option; malformed final selections are rejected rather than replaced by an earlier answer. Strict format compliance is reported separately. Earlier runs retain the scoring rules of their recorded protocol version. Reasoning-assisted allows an explanation, but only the final answer list is scored. Structured output requires an answers array in a strict JSON object, without repair calls or text fallback. We do not combine the three tracks into one ranking. Each track has one table ordered by end-to-end exact accuracy; check each run’s settings when comparing scores, because evaluation conditions may differ.
No track guarantees the absence or presence of a particular internal reasoning process. Provider defaults still matter, and explanations are not graded.
Repetition does not create new questions
Repeating an evaluation helps describe run-to-run behavior. It does not turn repeated answers to the same questions into independent evidence. Our pooled score averages all attempts without best-of-N selection, and uncertainty intervals are reported per repetition.
See the full scoring protocol for definitions and limitations.