Which LLM is best at bioinformatics?
Pass rates for 16 frontier models on agentic biology benchmarks (TxBench-PP, EpiBench, VariantBench), including Claude Opus 5.5, GPT-6 Astra, Grok 4.7 and Gemini 3.8 Flash.
Average across all 12 capability benchmarks in the leaderboard, weighted by number of tasks.
Leader
GPT-6 Astra 51.0%
Best run: Pi harness. #2 model: Claude Opus 5 at 49.6%.
Typical model
42.3%
Median of 16 models, each at its best harness. The leader is 8.8 points above it.
Harness matters
Pi wins 6 of 6
Same model, Pi vs Claude Code: Pi scores +0.7 points on average.
Overall leaderboard
Share of tasks passed (%). Longer bar is better.
- 1GPT-6 Astra51.0%
- 2Claude Opus 549.6%
- 3Grok 4.647.0%
- 4Grok 4.747.0%
- 5GPT-6 Sol46.4%
- 6Claude Opus 4.845.5%
- 7GPT-5.6 Sol43.1%
- 8Claude Opus 5.542.4%
- 9GPT-5.542.1%
- 10Claude Sonnet 541.0%
- 11Gemini 3.7 Flash40.4%
- 12Claude Opus 4.738.2%
- 13Gemini 3.8 Flash37.5%
- 14GPT-6 Luna37.2%
- 15GPT-5.6 Luna33.9%
- 16Claude Sonnet 4.631.6%
Each model shows the harness (agent framework) where it scored highest. Switch to “Every run” to see all harnesses. Bars span 0–60%.
Compare across benchmarks
Every run, sorted by Overall. Darker green means a higher score within that column. Scroll sideways on small screens.
| Model · harness | Overall | TxBench-PP | EpiBench | VariantBench |
|---|---|---|---|---|
GPT-6 Astra Pi | 51.0% | 59.3% | 28.3% | 46.6% |
Claude Opus 5 Pi | 49.6% | 63.0% | 25.2% | 50.6% |
Claude Opus 5 Claude Code | 49.0% | 59.7% | 20.8% | 52.8% |
GPT-6 Astra Codex | 48.9% | 59.3% | 24.2% | 45.2% |
Grok 4.6 Pi | 47.0% | 61.7% | 23.6% | 40.4% |
Grok 4.7 Grok Build | 47.0% | 54.7% | 20.8% | 45.5% |
GPT-6 Sol Pi | 46.4% | 63.0% | 22.0% | 38.7% |
Grok 4.6 Grok Build | 45.6% | 57.0% | 21.1% | 39.0% |
Claude Opus 4.8 Pi | 45.5% | 57.7% | 25.8% | 42.4% |
Claude Opus 4.8 Claude Code | 43.7% | 54.3% | 20.4% | 42.4% |
GPT-5.6 Sol Pi | 43.1% | 61.3% | 21.4% | 35.3% |
Claude Opus 5.5 Pi | 42.4% | 26.3% | 39.0% | 50.6% |
GPT-5.6 Sol Codex | 42.1%* | — | 25.8% | 34.2% |
GPT-5.5 Pi | 42.1%* | 55.3% | 28.3% | 30.5% |
Claude Opus 5.5 Claude Code | 42.0% | 30.7% | 38.7% | 49.7% |
Claude Sonnet 5 Pi | 41.0% | 49.3% | 19.5% | 39.8% |
Gemini 3.7 Flash Pi | 40.4% | 55.3% | 22.0% | 31.4% |
Claude Sonnet 5 Claude Code | 40.1% | 49.3% | 21.7% | 39.0% |
Claude Opus 4.7 Pi | 38.2% | 45.3% | 22.6% | 35.0% |
Claude Opus 4.7 Claude Code | 38.1%* | 42.3% | 18.9% | 33.1% |
Gemini 3.8 Flash Pi | 37.5% | 54.7% | 21.4% | 31.4% |
GPT-6 Luna Pi | 37.2% | 44.3% | 19.2% | 27.7% |
GPT-5.6 Luna Codex | 33.9% | 43.7% | 19.8% | 27.1% |
GPT-5.6 Luna Pi | 33.3% | 43.3% | 18.9% | 24.0% |
Claude Sonnet 4.6 Pi | 31.6% | 34.3% | 18.9% | 31.1% |
Claude Sonnet 4.6 Claude Code | 31.2% | 33.3% | 17.3% | 26.3% |
* Overall is averaged over the benchmarks this run completed (11 of 12). A dash means the run has no score for that benchmark.
How to read this
- What is scored. An AI agent gets a real bioinformatics task and must submit a structured answer. Verifiable graders check it against a reference. The score is the share of tasks passed.
- What a harness is. The agent framework wrapped around the model (Pi, Claude Code, Codex, Grok Build). The same model can score differently in each.
- Treat small gaps as ties. Each benchmark has about 100 tasks. The benchmark papers report 95% confidence intervals of roughly ±8 points (for example TxBench-PP: 59.3%, interval 51.1–67.6), so gaps under about 5 points are not meaningful.
- Look across benchmarks. Claude Opus 5.5 scores 39.0% on EpiBench and 26.3% on TxBench-PP, while Claude Opus 5 scores 25.2% and 63.0%. One overall number hides big differences between tasks.
- Missing models. Claude Sonnet 5.5 and Fable 5.1 have no published results on these benchmarks yet, so they are not shown.
Data: LatchBio, benchmarks.bio, as of 2026-09-29. Benchmarks: TxBench-PP, EpiBench, VariantBench.