Human health
Understanding the risks that affect people: occupational health, workplace exposure, and first-aid knowledge.
AI4HSE-Bench v1: a benchmark for the human side of AI. Measuring how models understand health, safety, and the environment — before their decisions shape our world.
AI4HSE-Bench v1 · 10 models evaluated
Supporting independent AI for HSE research.
The conversation about AI safety usually asks whether AI itself is safe. We ask a complementary question: can AI make sound judgments about our safety?
As AI agents enter business processes, they will increasingly assess human behavior, identify occupational hazards, and inform environmental decisions. These judgments deserve scrutiny, evidence, and a benchmark built for the task.
The thinking behind AI for HSELed by Pavel Kosyrev. Have a research question or want to collaborate? Contact Pavel
Understanding the risks that affect people: occupational health, workplace exposure, and first-aid knowledge.
Recognizing hazards and evaluating safe behavior across industrial facilities, fire safety, and everyday operations.
Assessing environmental responsibilities, protection practices, and the consequences of waste management decisions.
Three evaluation tracks. Exact-match scoring. Failures counted.
A clear view of what models get right — and where they fall short.
AI4HSE-Bench v1 measures HSE knowledge and answer reliability; scores do not establish operational safety or deployment readiness.
Explore the evaluation framework100 questions, six descriptive topic areas, and three separately reported evaluation tracks with published results.
How to read exact accuracy, response coverage, and failures across three separate evaluation tracks.
Introducing AI for HSE, the research initiative behind AI4HSE-Bench, and a different side of the AI safety conversation.