From knowledge
to sound judgment.
What we measure today, and the questions we want to investigate next.
Questions worth testing.
AI for HSE studies how AI understands health, safety, and the environment. Alongside the published AI4HSE-Bench v1 knowledge benchmark, we are expanding its question corpus and studying synthetic dataset creation. The directions under proposed research still need protocols and evaluation materials to be developed; no results are available for them.
Current research
HSE knowledge and answer reliability
Question: How accurately and consistently do models answer multi-select HSE questions under a fixed protocol?
Status: AI4HSE-Bench v1 is published: 100 questions, six descriptive topic areas, and separate answer-only, reasoning-assisted, and structured-output tracks. Public aggregates report accuracy, coverage, and failures. These scores do not establish operational safety or deployment readiness.
Next step: Continue publishing completed model evaluations while developing the expanded corpus with experts.
Read the methodology or explore published results.
Next milestone: 1,000 questions
We are working with experts to extend the current corpus from 100 to 1,000 questions for the planned AI4HSE-Bench Extended variant. This expansion is in progress; the published results still use the current 100-question dataset.
The next milestone is an expanded corpus with expert review of the questions and answer key, followed by evaluation and publication of results for that version.
Synthetic dataset creation
Question: How can synthetic questions and scenarios broaden HSE evaluation while maintaining accuracy and useful topic coverage?
Status: Under study. We are investigating methods for creating synthetic datasets; this does not change the current published benchmark.
Next step: Develop candidate examples and define expert review and validation criteria before considering them for future evaluations.
Proposed research
Invented requirements
Question: Can AI distinguish a documented safety obligation from a plausible but invented requirement, while identifying the source and limits of its answer?
Status: Proposed; scope and evaluation protocol to be defined.
Next step: Define a source-grounded task with a stated jurisdiction and reference date, then draft expert-reviewed examples and a scoring rubric.
Uncertainty and escalation
Question: When information is missing or conflicting, does AI recognise uncertainty, ask useful questions, and escalate to an appropriate human decision-maker?
Status: Proposed; scope and evaluation protocol to be defined.
Next step: Draft scenarios with deliberately incomplete evidence and agree criteria for clarification, abstention, and escalation with HSE specialists.
Incident reasoning
Question: Can AI analyse an incident, separate observations from assumptions, consider alternative explanations, and identify the evidence needed before drawing conclusions?
Status: Proposed; scope and evaluation protocol to be defined.
Next step: Develop independently authored incident scenarios and review a rubric that rewards evidence-based reasoning without unsupported causal claims.
Enterprise AI evaluation
Question: How reliably does an AI system apply an organisation’s HSE procedures in realistic workflows, including document retrieval, tool use, and human hand-offs?
Status: Proposed; scope and evaluation protocol to be defined.
Next step: Identify a bounded workflow with an interested organisation and agree permitted materials, failure criteria, and a controlled evaluation design.
Help shape the work
Research priorities and study designs will depend on practical questions, expert input, and available resources. To discuss a research question or contribute expertise, contact Pavel Kosyrev or explore Industry Research Partners.