AI4HSE-Bench v1 · LEADERBOARD

How well does AI
understand responsibility?

Model performance on health, safety, and environmental knowledge. Exact answers matter. So does the ability to answer reliably.

Results updated: · Latest completed evaluation in the published results (UTC)

Evaluations are ongoing. Reviewed model results are added as evaluations finish.

Model results

Boolean answers only. Internal reasoning is not disabled.Public aggregates

AI4HSE-Bench v1 measures HSE knowledge and answer reliability; scores do not establish operational safety or deployment readiness. Rows are sorted by exact accuracy. Evaluation settings may differ; check each model’s settings when comparing scores.

10 results in Answer-only

answer-only results ordered by end-to-end exact accuracy
Provider / modelExact accuracy ↓Valid accuracyCoverageQuestions × repeatsEvaluation cost (USD)Settings & breakdown
google/gemini-3.8-flashopenai-compatible82.3%82.3%100.0%100 × 3$0.5024
moonshotai/kimi-k3openai-compatible75.0%75.0%100.0%100 × 3$1.2937
anthropic/claude-sonnet-5.5openai-compatible72.7%72.7%100.0%100 × 3$0.625
openai/gpt-6-lunaopenai-compatible72.3%72.3%100.0%100 × 3$0.0319
openai/gpt-6.1-solopenai-compatible71.3%71.3%100.0%100 × 3$0.315
qwen/qwen3.8-max-0902openai-compatible70.3%70.3%100.0%100 × 3$0.7142
openai/gpt-5.6-lunaopenai-compatible66.7%66.7%100.0%100 × 3$0.0614
z-ai/glm-5.3openai-compatible66.3%66.3%100.0%100 × 3$0.1999
deepseek/deepseek-v4.1-flashopenai-compatible65.3%65.3%100.0%100 × 3$0.5191
mistralai/mistral-small-2603openai-compatible32.7%32.7%100.0%100 × 3$0.0976

Costs are reported USD totals for this track across all repetitions. Unreported response costs and additional infrastructure attempts are excluded; exact values and cost coverage are in Settings & breakdown.

PRIMARY METRIC

End-to-end exact accuracy

Correct complete selections / all attempts. Failures count as incorrect.

READ TOGETHER

Valid accuracy + coverage

Accuracy on valid responses, paired with the proportion of all attempts that were valid.

COMPARISON POLICY

Separate tracks, visible settings

Each track has one table sorted by exact accuracy. Evaluation settings are shown for every run. Repeated attempts are averaged; no best-run selection.