Which LLM is best at bioinformatics?

Pass rates for 16 frontier models on agentic biology benchmarks (TxBench-PP, EpiBench, VariantBench), including Claude Opus 5.5, GPT-6 Astra, Grok 4.7 and Gemini 3.8 Flash.

Average across all 12 capability benchmarks in the leaderboard, weighted by number of tasks.

Leader

GPT-6 Astra 51.0%

Best run: Pi harness. #2 model: Claude Opus 5 at 49.6%.

Typical model

42.3%

Median of 16 models, each at its best harness. The leader is 8.8 points above it.

Harness matters

Pi wins 6 of 6

Same model, Pi vs Claude Code: Pi scores +0.7 points on average.

Overall leaderboard

Share of tasks passed (%). Longer bar is better.

  1. 1
    GPT-6 Astra
    51.0%
  2. 2
    Claude Opus 5
    49.6%
  3. 3
    Grok 4.6
    47.0%
  4. 4
    Grok 4.7
    47.0%
  5. 5
    GPT-6 Sol
    46.4%
  6. 6
    Claude Opus 4.8
    45.5%
  7. 7
    GPT-5.6 Sol
    43.1%
  8. 8
    Claude Opus 5.5
    42.4%
  9. 9
    GPT-5.5
    42.1%
  10. 10
    Claude Sonnet 5
    41.0%
  11. 11
    Gemini 3.7 Flash
    40.4%
  12. 12
    Claude Opus 4.7
    38.2%
  13. 13
    Gemini 3.8 Flash
    37.5%
  14. 14
    GPT-6 Luna
    37.2%
  15. 15
    GPT-5.6 Luna
    33.9%
  16. 16
    Claude Sonnet 4.6
    31.6%

Each model shows the harness (agent framework) where it scored highest. Switch to “Every run” to see all harnesses. Bars span 0–60%.

Compare across benchmarks

Every run, sorted by Overall. Darker green means a higher score within that column. Scroll sideways on small screens.

Model · harnessOverallTxBench-PPEpiBenchVariantBench
GPT-6 Astra
Pi
51.0%59.3%28.3%46.6%
Claude Opus 5
Pi
49.6%63.0%25.2%50.6%
Claude Opus 5
Claude Code
49.0%59.7%20.8%52.8%
GPT-6 Astra
Codex
48.9%59.3%24.2%45.2%
Grok 4.6
Pi
47.0%61.7%23.6%40.4%
Grok 4.7
Grok Build
47.0%54.7%20.8%45.5%
GPT-6 Sol
Pi
46.4%63.0%22.0%38.7%
Grok 4.6
Grok Build
45.6%57.0%21.1%39.0%
Claude Opus 4.8
Pi
45.5%57.7%25.8%42.4%
Claude Opus 4.8
Claude Code
43.7%54.3%20.4%42.4%
GPT-5.6 Sol
Pi
43.1%61.3%21.4%35.3%
Claude Opus 5.5
Pi
42.4%26.3%39.0%50.6%
GPT-5.6 Sol
Codex
42.1%*—25.8%34.2%
GPT-5.5
Pi
42.1%*55.3%28.3%30.5%
Claude Opus 5.5
Claude Code
42.0%30.7%38.7%49.7%
Claude Sonnet 5
Pi
41.0%49.3%19.5%39.8%
Gemini 3.7 Flash
Pi
40.4%55.3%22.0%31.4%
Claude Sonnet 5
Claude Code
40.1%49.3%21.7%39.0%
Claude Opus 4.7
Pi
38.2%45.3%22.6%35.0%
Claude Opus 4.7
Claude Code
38.1%*42.3%18.9%33.1%
Gemini 3.8 Flash
Pi
37.5%54.7%21.4%31.4%
GPT-6 Luna
Pi
37.2%44.3%19.2%27.7%
GPT-5.6 Luna
Codex
33.9%43.7%19.8%27.1%
GPT-5.6 Luna
Pi
33.3%43.3%18.9%24.0%
Claude Sonnet 4.6
Pi
31.6%34.3%18.9%31.1%
Claude Sonnet 4.6
Claude Code
31.2%33.3%17.3%26.3%

* Overall is averaged over the benchmarks this run completed (11 of 12). A dash means the run has no score for that benchmark.

How to read this

  • What is scored. An AI agent gets a real bioinformatics task and must submit a structured answer. Verifiable graders check it against a reference. The score is the share of tasks passed.
  • What a harness is. The agent framework wrapped around the model (Pi, Claude Code, Codex, Grok Build). The same model can score differently in each.
  • Treat small gaps as ties. Each benchmark has about 100 tasks. The benchmark papers report 95% confidence intervals of roughly ±8 points (for example TxBench-PP: 59.3%, interval 51.1–67.6), so gaps under about 5 points are not meaningful.
  • Look across benchmarks. Claude Opus 5.5 scores 39.0% on EpiBench and 26.3% on TxBench-PP, while Claude Opus 5 scores 25.2% and 63.0%. One overall number hides big differences between tasks.
  • Missing models. Claude Sonnet 5.5 and Fable 5.1 have no published results on these benchmarks yet, so they are not shown.

Data: LatchBio, benchmarks.bio, as of 2026-09-29. Benchmarks: TxBench-PP, EpiBench, VariantBench.

Discover Bioeconomy's potential with us.