Guo, W., Zhang, M., Han, B., Ma, Y., Leng, Y., Hebbar, S., Zhou, X., Gu, W. et al.
This paper introduces PromptBio-Bench, a comprehensive benchmark for systematically evaluating the real-world readiness of LLM-based bioinformatics agents.
LLM-based agents have transformative potential for automating bioinformatics workflows, but systematic evaluations to clearly assess their readiness for real-world application are limited.
The authors develop and present PromptBio-Bench, a comprehensive evaluation suite comprising 244 expert-curated tasks and an evaluation framework for structured file comparison and scoring.
Evaluation of state-of-the-art agents like Biomni and ToolsGenie revealed a marked decline in accuracy as task difficulty increased. The benchmark provides valuable infrastructure for systematically tracking progress in agentic bioinformatics.