How well does AI
understand responsibility?
Model performance on health, safety, and environmental knowledge. Exact answers matter. So does the ability to answer reliably.
Results updated: · Latest completed evaluation in the published results (UTC)
Evaluations are ongoing. Reviewed model results are added as evaluations finish.
Model results
AI4HSE-Bench v1 measures HSE knowledge and answer reliability; scores do not establish operational safety or deployment readiness. Rows are sorted by exact accuracy. Evaluation settings may differ; check each model’s settings when comparing scores.
10 results in Answer-only
| Provider / model | Exact accuracy ↓ | Valid accuracy | Coverage | Questions × repeats | Evaluation cost (USD) | Settings & breakdown |
|---|---|---|---|---|---|---|
| google/gemini-3.8-flashopenai-compatible | 82.3% | 82.3% | 100.0% | 100 × 3 | $0.5024 | |
| moonshotai/kimi-k3openai-compatible | 75.0% | 75.0% | 100.0% | 100 × 3 | $1.2937 | |
| anthropic/claude-sonnet-5.5openai-compatible | 72.7% | 72.7% | 100.0% | 100 × 3 | $0.625 | |
| openai/gpt-6-lunaopenai-compatible | 72.3% | 72.3% | 100.0% | 100 × 3 | $0.0319 | |
| openai/gpt-6.1-solopenai-compatible | 71.3% | 71.3% | 100.0% | 100 × 3 | $0.315 | |
| qwen/qwen3.8-max-0902openai-compatible | 70.3% | 70.3% | 100.0% | 100 × 3 | $0.7142 | |
| openai/gpt-5.6-lunaopenai-compatible | 66.7% | 66.7% | 100.0% | 100 × 3 | $0.0614 | |
| z-ai/glm-5.3openai-compatible | 66.3% | 66.3% | 100.0% | 100 × 3 | $0.1999 | |
| deepseek/deepseek-v4.1-flashopenai-compatible | 65.3% | 65.3% | 100.0% | 100 × 3 | $0.5191 | |
| mistralai/mistral-small-2603openai-compatible | 32.7% | 32.7% | 100.0% | 100 × 3 | $0.0976 | |
Costs are reported USD totals for this track across all repetitions. Unreported response costs and additional infrastructure attempts are excluded; exact values and cost coverage are in Settings & breakdown.
End-to-end exact accuracy
Correct complete selections / all attempts. Failures count as incorrect.
Valid accuracy + coverage
Accuracy on valid responses, paired with the proportion of all attempts that were valid.
Separate tracks, visible settings
Each track has one table sorted by exact accuracy. Evaluation settings are shown for every run. Repeated attempts are averaged; no best-run selection.