October 25, 2026

What Changed in AI-for-Science: October 2026 Digest

Google Gemini for Science launches with Co-Scientist multi-agent hypothesis engine; an open-source model beats GPT-4 on citation accuracy; Nature reveals pharma's private protein data advantage; a new model predicts protein-protein interactions. Plus: Deep Research expands to ChatGPT Plus, Evo 2 enters standard practice, and Overleaf AI exits beta.


Headline: Google Gemini for Science — the biggest AI-for-research platform launch since AlphaFold

At Google I/O (May 20, 2026), Google announced Gemini for Science: a unified research platform bundling Co-Scientist, AlphaEvolve, Empirical Research Assistance (ERA), and NotebookLM. The centrepiece is Co-Scientist — a multi-agent system that generates, debates, and refines scientific hypotheses via an “idea tournament” architecture. Multiple AI agents independently propose hypotheses; a separate layer critiques and debates them; the survivors are returned with citations and ranked by evidential support.

This is not simply a literature chatbot. Google DeepMind published a peer-reviewed paper in Nature alongside the launch, and documented a drug-repurposing case study in which Co-Scientist identified a compound that blocked 91% of a fibrosis-linked cellular response in subsequent wet-lab validation. (Co-Scientist: A multi-agent AI partner to accelerate research, Google DeepMind Blog, 2026; Gemini for Science, Google Blog, 2026)

Co-Scientist is now available via Google Cloud (Gemini Enterprise). See our full Co-Scientist tool review.


Open-source AI beats GPT-4 on literature review citations

A Nature news report documented an open-source AI model that outperforms major commercial LLMs — including GPT-4-class systems — on scientific literature review tasks, with citation accuracy matching human expert performance. (Open-source AI tool beats giant LLMs in literature reviews — and gets citations right, Nature, 2026)

The key architectural difference: the model uses retrieval-augmented generation against an indexed corpus, so citations come from retrieved papers rather than model memory. This eliminates the main hallucination vector for citation accuracy. The model was not publicly named at time of publication, but the result reinforces the principle that specialised tools outperform general-purpose LLMs for citation-critical work — and that open-source options are maturing.


Pharma’s private protein data creates a two-tier AI landscape

A Nature investigation found that AI protein models trained on 20,000+ proprietary pharmaceutical structures — never deposited in public repositories — substantially outperform models trained on the Protein Data Bank alone, including AlphaFold derivatives. Structural biologists quoted in the article describe the performance gap as a “two-generation lead” for drug-relevant prediction tasks. (Drug firms’ secret data supercharge AI protein models, Nature, 2026)

The proprietary models are available only to pharma partner companies, not academic researchers. This creates a structural inequality in AI-assisted drug discovery: academic labs working with AlphaFold and public models are using systematically weaker tools than their industrial counterparts for the same tasks. The open-data vs. private-data debate in structural biology now has direct AI capability consequences.


New AI model predicts protein-protein interactions more accurately

A research group reported a new model that predicts how pairs of proteins interact with each other — a substantially harder problem than single-protein structure prediction. Reported improvements include higher accuracy on novel PPI pairs without close structural homologs in training data, and better identification of interface residues. (AI model ‘reads’ protein pairs, unlocking new insights into disease and drug discovery, Phys.org, April 2026)

Protein-protein interactions are an increasingly important drug target class (BCL-2 inhibitors, PROTACs, molecular glues) and a key mechanistic question in cancer biology. Better computational PPI prediction accelerates both target identification and interface-focused drug design.


Deep Research comes to ChatGPT Plus

OpenAI’s Deep Research feature — which runs autonomous 20–30 minute web research tasks and returns a cited report — expanded from ChatGPT Pro ($200/mo) to Plus ($20/mo) in mid-2025 and is now widely available. Usage at the Plus tier is capped at approximately 10 reports per month; Pro is uncapped.

For researchers, this represents genuine capability that was not available 18 months ago: a tool that autonomously searches, reads, and synthesizes a multi-page overview on a topic faster than a human research assistant could produce a first draft. The caveats remain real — sources are web pages, not academic databases; citations require verification; and hallucination on specific claims persists. See our step-by-step workflow for how to use it effectively.


Evo 2 entering standard practice in genomics labs

The Arc Institute’s Evo 2 model (270B parameters, 9.3 trillion training tokens of DNA) has moved from a splashy Science paper to a tool people are actually using. Key developments this quarter:

  • The web interface at EvolutionaryScale now supports batch variant scoring — upload a VCF file, get functional scores for all variants at once
  • Early adoption in clinical genomics workflows for Tier 3/4 variant interpretation (variants of uncertain significance)
  • Bioinformatics groups are fine-tuning Evo 2 on organism-specific or disease-specific data for specialized prediction tasks

The scale gap between Evo 2 (270B parameters) and previous genomic models (Nucleotide Transformer at ~500M, DNABERT at ~86M) is large enough that the performance difference on variant interpretation benchmarks is substantial. For labs doing genomic foundation model work, Evo 2 is now the reference point the way GPT-4 was for text.


Gemini 2.5 Pro: context window extended, pricing adjusted

Google extended Gemini 2.5 Pro’s effective context window to 2 million tokens — doubling the previous 1M limit. At 2M tokens you can fit approximately 20 average-length academic papers, or a full-length dissertation with appendices, in a single prompt.

Practical implication: for researchers who use Gemini for long-document Q&A (full thesis review, multi-paper corpus analysis), the bottleneck has shifted from context length to retrieval quality — Gemini’s ability to accurately answer questions about content in the middle of a very long context is still imperfect.

Google also adjusted API pricing downward for cached context, making repeated queries against the same long document significantly cheaper for developers building research tools on top of Gemini.


Elicit: systematic review screening mode

Elicit added a structured screening workflow to complement its existing data extraction features. The new mode lets researchers:

  • Define inclusion/exclusion criteria as structured rules
  • Screen papers against those rules with AI-generated rationale for each decision
  • Flag papers for human review when the AI confidence is low
  • Export a PRISMA-compatible flow diagram

Early user reports suggest the screening mode works well for clear-cut criteria (specific population, intervention type, study design) and degrades for nuanced criteria that require reading full text. The AI makes inclusion/exclusion calls on abstract + title; papers where the decision requires full-text review are flagged for manual handling.

This is the workflow gap between Elicit and Rayyan that has existed since Elicit launched. With screening added, Elicit can now cover more of the systematic review pipeline without switching tools.


Overleaf AI features out of beta

Overleaf’s AI writing features — autocomplete, rephrase, shorten, and error explanation — have exited beta and are now available to Standard and Professional subscribers. The error explanation feature in particular has received strong user feedback: when LaTeX compilation fails, the AI explanation in plain English (and often the suggested fix) resolves errors that would previously require a Stack Overflow search.

The autocomplete quality for academic prose is competent but not transformative; the model is tuned for academic register and LaTeX syntax but does not know your specific research area. Users report it is most helpful for transitions, methods boilerplate, and standard statistical reporting phrases.


Worth watching: Schrödinger’s FEP+ integration with AlphaFold 3

Schrödinger (the computational chemistry company, not the physicist) announced an integration between their FEP+ free energy perturbation engine and AlphaFold 3 predicted structures. This matters for drug discovery: FEP+ is one of the most accurate computational methods for predicting how strongly a drug candidate binds to a target, but it requires a high-quality starting structure. Using AF3 predictions as FEP+ inputs — rather than requiring experimental crystal structures — could substantially expand the scope of early-stage computational drug discovery.

Results on benchmark systems were published as a preprint in October 2026; independent validation is pending.


Corrections and updates

Perplexity pricing (updated): Perplexity Pro is now $17/month when billed annually (was $20/month through most of 2024). The Max tier at $167/month annually adds higher API limits and priority access to reasoning models.

Elicit pricing: The $12/month tier is no longer offered. Current tiers are Basic (free), Pro ($49/month), Scale ($169/month), and Enterprise.