Clear questions.
Accountable evaluation.
AI4HSE-Bench v1 is a closed, multi-select knowledge benchmark by AI for HSE. Built to make accuracy, response reliability, and the limits of the evidence visible.
01 / Question format
The current dataset contains 100 multiple-choice questions. Each question has an ordered set of options and a Boolean answer key: each option is marked true or false. A correct answer must match the entire selection, including all correct and incorrect options.
The evaluation is zero-shot, with fixed prompts and option order, no tools, and no demonstrations. Questions and answer keys remain private. The CSV input groups question text, option text, and a correctness token into normalized question records.
[TRUE, FALSE, TRUE, FALSE]
One Boolean per option, in the original order. No partial credit in the primary exact-match score.
02 / Topic coverage
These are assistant-assigned descriptive labels, not claims of expert validation. The two-question regulatory group is descriptive only. Topic slices do not change the equal weighting of questions.
03 / Three tracks, reported separately
Answer-only
The model is instructed to return a Boolean list and nothing else. This controls the requested visible output; it does not disable internal reasoning. Under protocol v10, knowledge scoring uses the last bracketed answer list, including a revised selection after an earlier answer. Explanations around the list are ignored. The last list must have one valid TRUE/FALSE value per option; a malformed last list is not replaced by an earlier valid one. Strict format compliance is reported separately.
Reasoning-assisted
A brief explanation followed by exactly one final line:
FINAL_ANSWER: [TRUE, FALSE, ...]
Only the final list is scored. Missing, multiple, ambiguous, wrong-length, or non-final answer fields are invalid. Explanations are not graded. A final unbracketed list can be extracted but is counted as a strict format-compliance failure.
Structured output
A separate answer-only control requires a strict JSON schema: exactly one answers array with one Boolean per option. Invalid types, duplicate or extra fields, and wrong lengths are rejected. No repair calls or text fallback are used.
Each published run shows its repetitions, token cap, temperature, reasoning effort, provider routing, and concurrency settings. Omitted temperature uses the provider default. Equal token caps and effort labels do not mean equal compute.
04 / How scoring works
End-to-end exact accuracy — the primary score
Correct complete selections divided by all question attempts. Answers that cannot be extracted and incomplete generations count as incorrect. Extra prose or an earlier answer list does not invalidate an extractable answer-only response under v10. Each run records its protocol version; earlier versions used stricter answer-only extraction rules. Infrastructure failures stop scheduling and leave work pending; unfinished runs have no aggregate score. Completed answers are never selectively retried based on correctness.
Valid-response exact accuracy + coverage
Correct complete selections divided by valid responses, paired with valid responses divided by all attempts. Valid-response accuracy is undefined when no response is valid. Always read the two metrics together: selective failures can inflate valid-response accuracy.
Supporting metrics
Option accuracy, precision, recall, F1, and Hamming loss are computed on valid responses only. Provider, parse, and incomplete-generation errors are counted separately. Always-true and always-false baselines use all selected questions.
Repetitions and uncertainty
The pooled score averages all attempts; there is no best-of-N selection or majority vote. Each repetition has a 95% Wilson interval. No pooled interval is calculated because repeated answers are not independent questions.
05 / Limits & transparency
There is currently no separate development dataset. Earlier pilot failures informed the output and extraction protocol. Independent review of the evaluation material is pending. Fixed answer positions, small topic groups, and the closed question set limit interpretation.
Intervals assume representative, independent questions. Scores do not establish pairwise superiority, answer-key quality, general HSE competence, or operational safety. No paired statistical comparison is implemented.
Publications use the runner’s allowlisted public aggregate export. Private audit records contain questions, answers, and raw responses and are not website content. AI4HSE-Bench v1 is the public release name; each run records its internal protocol version.
Benchmark names and versions
AI for HSE is the research initiative; AI4HSE-Bench is its benchmark family. Names follow AI4HSE-Bench [variant] vN, with the variant omitted for the original benchmark. Each version identifies a fixed corpus and evaluation protocol. Adding model results does not change that version.
Extended is the name for the planned 1,000-question expansion. Multimodal is reserved for a future variant involving images or other modalities. These variants are not part of the current AI4HSE-Bench v1 release.
View the publication status →