All benchmarks
Dataset availability across supported platforms. Langfuse includes native scoring; Braintrust exports datasets and scoring specifications. Harbor supports packages with code graders.
| Benchmark | Langfuse | Braintrust | Harbor |
|---|---|---|---|
| GPQA Diamond | ✓ | ✓ | ✓ |
| AIME 2026 | ✓ | ✓ | ✓ |
| Global-MMLU-Lite | ✓ | ✓ | ✓ |
| SimpleQA Verified | ✓ | ✓ | — |
| LiveBench | ✓ | ✓ | — |
| SciKnowEval v2 | ✓ | ✓ | ✓ |
| HMMT February 2026 | ✓ | ✓ | — |
| HMMT November 2025 | ✓ | ✓ | — |
| AMO-Bench | ✓ | ✓ | — |
| FrontierScience Olympiad | ✓ | ✓ | — |
| FrontierScience Research | ✓ | ✓ | — |