LLM Benchmarks for
Bioinformatics & Biotech
The industry-standard evaluation suite comparing leading LLMs on bioinformatics, genomics, and AI drug discovery API costs. Discover the best LLM for drug discovery and run our computational biology simulator.
Leaderboard Performance
Showing average deterministic accuracy (%) with standard deviation error lines
Detailed Model Specifications
Comprehensive metrics under the TxBench-PP protocol.
| Rank | Model Identifier | Developer | Accuracy | Precision | F1-Score | Latency | API Cost / 1M tkn | Strategic Suitability |
|---|---|---|---|---|---|---|---|---|
| #1 | o1SOTA | OpenAI | 68.4% | 69.5% | 68.8% | 12.2s | $15.000 | Recommended |
| #2 | Claude 3.5 Sonnet | Anthropic | 63.8% | 64.2% | 63.9% | 4.4s | $3.000 | Recommended |
| #3 | o1-mini | OpenAI | 61.2% | 62.5% | 61.8% | 5.4s | $3.000 | Recommended |
| #4 | GPT-4o | OpenAI | 58.5% | 59.9% | 59.1% | 4.1s | $5.000 | Recommended |
| #5 | K Kimi 3 | Moonshot | 57.2% | 58.4% | 57.8% | 6.8s | $4.000 | Recommended |
| #6 | Gemini 1.5 Pro | 55.3% | 57.5% | 56.4% | 6.1s | $3.500 | Recommended | |
| #7 | Llama 3.1 405B | Meta | 54.7% | 55.9% | 55.3% | 8.1s | $3.000 | Recommended |
| #8 | DeepSeek-V3Best Value | DeepSeek | 53.5% | 54.9% | 54.1% | 3.1s | $0.270 | Recommended |
| #9 | Llama 3.1 70BBest Local | Meta | 51.3% | 52.8% | 52% | 2.3s | $0.900 | Recommended |
| #10 | Gemini 1.5 Flash | 49.7% | 51% | 50.3% | 1.5s | $0.075 | Viable | |
| #11 | Qwen 2.5 Max | Alibaba | 48.3% | 49.5% | 48.9% | 3.8s | $0.800 | Viable |
| #12 | Claude 3 Opus | Anthropic | 45% | 46.2% | 45.6% | 12.2s | $15.000 | Viable |
AI Model API Cost & Computational Economics Simulator
Estimate API cost and computational time saved when executing bulk agentic reasoning cycles on large-scale therapeutics and sequencing datasets.
Gemini 1.5 Flash
$30.00
Estimated Execution Cost
2.1 hrs
Execution Time
4997.9 hrs
Human Time Saved
DeepSeek-V3
$108.00
Estimated Execution Cost
4.3 hrs
Execution Time
4995.7 hrs
Human Time Saved
Qwen 2.5 Max
$320.00
Estimated Execution Cost
5.3 hrs
Execution Time
4994.7 hrs
Human Time Saved
Methodology & Verification Protocol
Each evaluated agent and LLM is benchmarked using our standard isolated executor cluster. The models are tasked with resolving exact biological outcomes using pre-release scientific datasets across stages like S1 to S9.
We evaluate success through a deterministic grader using Jaccard similarity metrics of gene names, variant coordinates, and target affinity profiles, guaranteeing no subjective scoring.
Accuracy scores indicate the percentage of benchmark cases correctly resolved on the first attempt without human-in-the-loop assistance. Standard error margins are calculated at 95% confidence intervals across three distinct test subsets.
Includes VariantBench, EpiBench, and TxBench-PP v4.3 datasets. Last updated: Current.