LLM Benchmarks for Bioinformatics & Biotech

The industry-standard evaluation suite comparing leading LLMs on bioinformatics, genomics, and AI drug discovery API costs. Discover the best LLM for drug discovery and run our computational biology simulator.

Leaderboard Performance

Showing average deterministic accuracy (%) with standard deviation error lines

Live Data Feeds

Detailed Model Specifications

Comprehensive metrics under the TxBench-PP protocol.

RankModel IdentifierDeveloperAccuracyPrecisionF1-ScoreLatencyAPI Cost / 1M tknStrategic Suitability
#1o1SOTAOpenAI68.4%69.5%68.8%12.2s$15.000Recommended
#2Claude 3.5 SonnetAnthropic63.8%64.2%63.9%4.4s$3.000Recommended
#3o1-miniOpenAI61.2%62.5%61.8%5.4s$3.000Recommended
#4GPT-4oOpenAI58.5%59.9%59.1%4.1s$5.000Recommended
#5
K
Kimi 3
Moonshot57.2%58.4%57.8%6.8s$4.000Recommended
#6Gemini 1.5 ProGoogle55.3%57.5%56.4%6.1s$3.500Recommended
#7Llama 3.1 405BMeta54.7%55.9%55.3%8.1s$3.000Recommended
#8DeepSeek-V3Best ValueDeepSeek53.5%54.9%54.1%3.1s$0.270Recommended
#9Llama 3.1 70BBest LocalMeta51.3%52.8%52%2.3s$0.900Recommended
#10Gemini 1.5 FlashGoogle49.7%51%50.3%1.5s$0.075Viable
#11Qwen 2.5 MaxAlibaba48.3%49.5%48.9%3.8s$0.800Viable
#12Claude 3 OpusAnthropic45%46.2%45.6%12.2s$15.000Viable
Agent Pricing Terminal

AI Model API Cost & Computational Economics Simulator

Estimate API cost and computational time saved when executing bulk agentic reasoning cycles on large-scale therapeutics and sequencing datasets.

10025,00050,000
5,000125,000250,000
Rank #1 Economic Fit

Gemini 1.5 Flash

$30.00

Estimated Execution Cost

2.1 hrs

Execution Time

4997.9 hrs

Human Time Saved

Rank #2 Economic Fit

DeepSeek-V3

$108.00

Estimated Execution Cost

4.3 hrs

Execution Time

4995.7 hrs

Human Time Saved

Rank #3 Economic Fit

Qwen 2.5 Max

$320.00

Estimated Execution Cost

5.3 hrs

Execution Time

4994.7 hrs

Human Time Saved

Methodology & Verification Protocol

Each evaluated agent and LLM is benchmarked using our standard isolated executor cluster. The models are tasked with resolving exact biological outcomes using pre-release scientific datasets across stages like S1 to S9.

We evaluate success through a deterministic grader using Jaccard similarity metrics of gene names, variant coordinates, and target affinity profiles, guaranteeing no subjective scoring.

Accuracy scores indicate the percentage of benchmark cases correctly resolved on the first attempt without human-in-the-loop assistance. Standard error margins are calculated at 95% confidence intervals across three distinct test subsets.

Includes VariantBench, EpiBench, and TxBench-PP v4.3 datasets. Last updated: Current.

Discover Bioeconomy's potential with us.