Benchmarks

How AI is measured - each tied to the paper or site that defined it.

111 entries, all primary-sourced
benchmarkMarch 5, 2025

MASK Benchmark (Honesty)

A benchmark that separates honesty from accuracy, measuring whether language models lie under pressure rather than just whether they are correct.

benchmarkMarch 24, 2025

ARC-AGI-2

The 2025 successor to ARC-AGI, a harder reasoning benchmark that frontier AI systems initially scored in the single digits on.

benchmarkApril 16, 2025

BrowseComp

OpenAI benchmark of 1,266 questions that force a web agent to dig persistently for hard-to-find facts.

benchmarkApril 25, 2025

HalluLens

A Meta hallucination benchmark built on a clear taxonomy, with tasks that regenerate to resist leakage.

benchmarkOctober 5, 2025

GDPval

OpenAI benchmark scoring AI on real economically valuable work across 44 occupations in nine GDP sectors.

benchmarkJuly 9, 2026

Long-Horizon-Terminal-Bench

A 46-task terminal benchmark with dense partial-credit grading where the best agent scores just 15.2 percent.

benchmarkJuly 20, 2026

CryptanalysisBench

A 191-task benchmark where five frontier models break 65 to 86 percent of already-broken cryptographic schemes.

benchmarkAugust 18, 2026

ASI-Bench

A 60-task benchmark of project-level scientific research where agent scores fall from 50.91 to 26.62 once methodological guidance is removed.