When AI judges human safety, who evaluates AI?
Introducing AI for HSE, the research initiative behind AI4HSE-Bench, and a different side of the AI safety conversation.
AI is becoming part of how businesses make decisions. The question is no longer only whether a model can produce a useful answer. It is also whether its judgment is reliable when people and the environment are affected.
A complementary question about safety
Much of the conversation about AI safety focuses on the behavior of AI systems themselves. That work matters. AI for HSE starts with a complementary question: how well does AI understand the responsibilities that keep humans safe?
A model might help identify a workplace hazard, interpret a safety procedure, or assess an environmental practice. As agentic systems become embedded in business processes, these judgments may influence whether human behavior is considered safe and responsible.
Fluent language can conceal gaps in knowledge. We need a way to examine those gaps systematically.
Starting with measurable evidence
AI4HSE-Bench v1 is our first benchmark release. It evaluates multi-select question answering across health, safety, and environmental topics. A model must identify the complete correct selection, not simply give a plausible explanation.
We report answer accuracy alongside response coverage. A model that gives correct answers only when it manages to return a valid response should not look identical to a model that responds reliably across the whole evaluation.
A starting point, not a certificate
Knowledge tests do not establish that an AI system is safe to deploy in a workplace. Real decisions also depend on context, uncertainty, oversight, and the consequences of mistakes.
Our goal is to build an evidence base that makes those broader conversations more concrete. The initial benchmark is one step, with its limits made visible from the beginning.
Explore the benchmark methodology or check the leaderboard for publication updates.