AN OPEN QUESTION. A MEASURABLE CHALLENGE.

AI makes decisions.
How safe are
its judgments?

AI4HSE-Bench v1: a benchmark for the human side of AI. Measuring how models understand health, safety, and the environment — before their decisions shape our world.

AI4HSE-Bench v1 · 10 models evaluated

Results updated: · Latest evaluation (UTC)

AI4HSE-Bench v1
100questions in the current dataset
6descriptive topic areas
3separate evaluation tracks
Inside the benchmark

Industry Research Partners

Supporting independent AI for HSE research.

Explore research partnerships
01 / WHY THIS MATTERS

When AI evaluates us,
we need to evaluate AI.

The conversation about AI safety usually asks whether AI itself is safe. We ask a complementary question: can AI make sound judgments about our safety?

As AI agents enter business processes, they will increasingly assess human behavior, identify occupational hazards, and inform environmental decisions. These judgments deserve scrutiny, evidence, and a benchmark built for the task.

The thinking behind AI for HSE

Led by Pavel Kosyrev. Have a research question or want to collaborate? Contact Pavel

02 / THREE RESPONSIBILITIES. ONE SHARED FUTURE.

Intelligence with consequences.

Beyond knowledge. Toward responsible judgment.
01

Human health

Understanding the risks that affect people: occupational health, workplace exposure, and first-aid knowledge.

02

Workplace safety

Recognizing hazards and evaluating safe behavior across industrial facilities, fire safety, and everyday operations.

03

Environmental impact

Assessing environmental responsibilities, protection practices, and the consequences of waste management decisions.

AI4HSE-Bench v1

Evidence before confidence.

Three evaluation tracks. Exact-match scoring. Failures counted.
A clear view of what models get right — and where they fall short.

AI4HSE-Bench v1 measures HSE knowledge and answer reliability; scores do not establish operational safety or deployment readiness.

Explore the evaluation framework

Explore the research programme

01Closed question setMulti-select QA
02Fixed evaluation protocolZero-shot · No tools
03Transparent aggregatesAccuracy + coverage
03 / FROM THE JOURNAL

Notes on AI & responsibility.

All articles

Inside AI4HSE-Bench v1.

100 questions, six descriptive topic areas, and three separately reported evaluation tracks with published results.