← Back to the journal

Inside AI4HSE-Bench v1.

100 questions, six descriptive topic areas, and three separately reported evaluation tracks with published results.

AI4HSE-Bench v1, the first benchmark release from the AI for HSE research initiative, starts with a bounded task: answering a closed set of multiple-choice questions using a complete Boolean selection.

Updated 1 October 2026: Public results are now available across three evaluation tracks. This article reflects the current protocol; see the methodology for the track definitions and the leaderboard for published runs.

What the current dataset covers

The current set contains 100 questions. Descriptive topic assignments cover environmental protection and waste, occupational safety and health, industrial and facility safety, first aid, fire safety, and regulatory administration.

These labels are assistant-assigned, and independent review of the evaluation material is pending. They organize the evidence; they are not a claim that every topic is equally represented or independently validated.

One task, three evaluation tracks

In the answer-only track, models are instructed to return just a Boolean list. Protocol v10 scores the last bracketed answer list, including a self-correction, while reporting strict format compliance separately. In the reasoning-assisted track, they give a brief explanation followed by a single final answer line. The structured-output track is a separate answer-only control: it requires a strict JSON object with an answers array containing one Boolean per option, without repair calls or text fallback.

All three use fixed prompts and option order, with no tools or examples. Each track has one table ordered by end-to-end exact accuracy. Evaluation settings remain visible for every run and should be checked when comparing scores.

Only the selected answers are scored. The benchmark does not assess the quality of explanations or an agent’s ability to act in a live workplace.

What is public

The website publishes public aggregate results with accuracy, coverage, failures, and evaluation settings. Publication does not establish independent scientific review. The private question set, answer keys, raw responses, and audit records remain outside the frontend.

Explore the published results. Evaluations are ongoing, and additional models are published as completed public exports become available. Unfinished runs receive no aggregate score.

Read the methodology for the current format, reporting rules, and limitations.

Explore the benchmark methodology →