https://github.com/harbor-framework/terminal-bench-science
| Benchmark | Domain | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | Notes |
| Terminal-Bench-Science 0.1 | Agentic Scientific Research | 52.6% | 24.7% | 29.0% | 22.4% | More than double the predecessor |
| Terminal-Bench 4.0 | Agentic Coding / Terminal Tasks | 55.8% | 42.0% | 52.3% | 37.3% | Mythos 5.1 reaches 60.9% |
| CursorBench 3.2.0 | Agentic Code Editing | 73.4% | 70.5% | 70.0% | 67.2% | Standard tool-assisted coding |
| Humanity's Last Exam | Multidisciplinary Reasoning | 60.9% (no tools) 65.0% (with tools) | 57.8% 63.8% | 56.6% 63.6% | — | Frontier high-difficulty reasoning |
| GDPval-AA v2 | Complex Knowledge Work | 1853 | 1723 | 1824 | 1711 | Standardized knowledge score |
| AutomationBench | Business Workflows | 31.4% | 17.1% | 26.9% | 19.6% | Multi-step enterprise automation |
| OSWorld 2.0 | Computer Use & UI Navigation | 41.7% (strict) 77.9% (partial) | 36.1% 72.9% | 39.6% 75.4% | — | Evaluated with production safeguards |