Tuesday, September 1, 2026

agent benchmarks

 https://github.com/harbor-framework/terminal-bench-science

BenchmarkDomainFable 5.1Fable 5Opus 5GPT-5.6 SolNotes
Terminal-Bench-Science 0.1Agentic Scientific Research52.6%24.7%29.0%22.4%More than double the predecessor
Terminal-Bench 4.0Agentic Coding / Terminal Tasks55.8%42.0%52.3%37.3%Mythos 5.1 reaches 60.9%
CursorBench 3.2.0Agentic Code Editing73.4%70.5%70.0%67.2%Standard tool-assisted coding
Humanity's Last ExamMultidisciplinary Reasoning

60.9% (no tools)


65.0% (with tools)

57.8%


63.8%

56.6%


63.6%

Frontier high-difficulty reasoning
GDPval-AA v2Complex Knowledge Work1853172318241711Standardized knowledge score
AutomationBenchBusiness Workflows31.4%17.1%26.9%19.6%Multi-step enterprise automation
OSWorld 2.0Computer Use & UI Navigation

41.7% (strict)


77.9% (partial)

36.1%


72.9%

39.6%


75.4%

Evaluated with production safeguards