Biology agents are being judged on research judgment, not just factual answers. GeneBench-Pro puts 129 computational-biology problems in front of agents, and indexed coverage says GPT-5.6 Sol reaches 28.7% at the highest reasoning level and 31.5% in Pro mode.
#biology
RSS FeedSciences X/Twitter Jul 1, 2026 1 min read
Sciences X/Twitter Jun 18, 2026 1 min read
AI for life sciences is getting a more realistic yardstick. OpenAI says LifeSciBench was built with 173 biotech and pharma scientists and spans 750 expert-written tasks across seven biological research workflows.
Sciences X/Twitter Jun 10, 2026 1 min read
Anthropic points to infrastructure, not only model intelligence, as the bottleneck for scientific agents. In an NCBI Virus retrieval task, accuracy rose to nearly 100% after adding a deterministic gget virus layer.
Sciences Reddit May 4, 2026 1 min read
IBM Research has published MAMMAL, a multi-modal model that unifies proteins, molecules, and gene data. It achieves state-of-the-art results on 9 of 11 biological benchmarks and outperforms AlphaFold 3 on several drug-discovery tasks.