Evidence with
a history.
Published model additions and website changes. A benchmark version identifies its corpus and evaluation protocol; adding models keeps the same version.
· Core benchmark results
- Published all ten core models across three tracks: GPT-5.6 Luna, DeepSeek V4.1 Flash, Claude Sonnet 5.5, Gemini 3.8 Flash, Qwen3.8 Max 0902, GPT-6.1 Sol, GLM 5.3, Mistral Small, GPT-6 Luna, and Kimi K3.
- Completed 9,000 evaluations on the same 100 questions, with three passes per track. Reused 1,000 verified pilot responses.
- Answer-only accuracy now uses the last answer list, including self-corrections. Strict output format compliance is reported separately. All 3,000 answer-only responses were extractable.
- Replaced the earlier Luna and DeepSeek website entries with new runs under the shared protocol. Original audit evidence is preserved. Each model has its own public aggregate file.
- Added DeepSeek V4.1 Flash public aggregates across all three tracks.
- Named the benchmark AI4HSE-Bench v1. AI for HSE remains the research initiative.
- Added evaluation costs, results freshness, shareable track and result links, and clearer breakdowns. Improved the mobile results layout.
- Published the Research Programme and Industry Research Partners pages, including the expert-supported 1,000-question expansion milestone.
These presentation and publication updates do not change the existing scores.
Published GPT-5.6 Luna results with answer-only, reasoning-assisted, and structured-output tracks, including public aggregate breakdowns and the combined JSON download.
Corrections
No score corrections are recorded in this history.
Any future correction will be dated here and identify the affected model and track, what changed, and why. Corrected public exports come from the benchmark project after its protocol and results verification.
Evaluation completion times appear in each result’s breakdown. The dates here refer to website publication changes.