top of page

Biological Data Emerges as AI Drug Discovery’s Real Bottleneck

Aug 13
12 min read

Google News surfaced a sharp conflict in AI drug discovery: larger models are multiplying, yet reliable biological evidence remains scarce, fragmented, and expensive.

The argument challenges a familiar technology playbook. More computing power and parameters can improve a model, but they cannot manufacture trustworthy knowledge about an untested biological system. A model still needs observations from cells, tissues, organisms, and patients.

That distinction puts data-centered companies against teams betting primarily on model architecture and computational scale. Google DeepMind’s AlphaFold provides the clearest historical reference. It changed structural biology by learning from an unusually valuable collection of experimentally determined protein structures.

The next stage is harder. Drug discovery requires more than predicting a stable molecular shape. Researchers must understand interactions, disease mechanisms, toxicity, dosing, cellular responses, and differences between patients.

Each step depends on data that are difficult to generate and even harder to compare. Biological measurements vary with experimental conditions, laboratory methods, sample preparation, cell type, disease stage, and many other factors.

The central question is therefore changing. The industry once asked which company could build the most capable model. It increasingly asks which organization can produce, connect, and validate the biological data that its models need.

That shift does not make AI less important. It defines where AI can deliver credible value, while exposing where impressive predictions still need experimental proof.

Why This Google News Debate Matters Now

The center of competition is moving from model scale toward the quality and relevance of the evidence beneath each prediction.

Research activity has expanded rapidly. The AI Index report counted 3,855 AI-driven protein research papers in 2025, up from 2,259 during 2024. That represents growth of about 71 percent.

Protein-drug interactions formed 54.4 percent of those 2025 papers. Protein structure prediction accounted for 23.9 percent, down from 28.7 percent one year earlier.

The change suggests that researchers are moving beyond the initial structure-prediction problem. They are focusing more heavily on interactions that matter for therapeutic design.

Yet this transition exposes a data constraint. A protein structure offers a useful snapshot, but drug development depends on many biological states and interactions. Experimental measurements of those conditions remain much less abundant than sequences or predicted structures.

Public datasets have still grown significantly. The 2026 AI Index highlighted Tahoe-100M, which contains measurements from more than 50 cancer cell types exposed to over 1,100 drugs. It also cited Meta’s OMol25 collection of over 100 million quantum-mechanics calculations.

Those numbers sound enormous, but dataset size alone can mislead. A simulated molecular calculation, a gene sequence, and a patient outcome are different forms of evidence. One cannot automatically replace another.

The underlying measurements also cover biological space unevenly. Researchers have studied some proteins, cancers, and molecular interactions intensively. Rare diseases, unusual binding mechanisms, diverse populations, and negative experiments receive less consistent coverage.

A model can become highly accurate inside familiar training territory while struggling with genuinely novel biology. This problem resembles extrapolation, meaning prediction beyond conditions represented in the training data.

Drug discovery frequently demands exactly that behavior. Scientists want therapies for targets that lack successful precedents, not another prediction about a well-characterized protein under familiar laboratory conditions.

The Google News headline captures a maturing industry debate. The model is becoming less of a standalone product and more of one component within an experimental learning system.

That system must decide what to test, conduct the experiment, record its conditions, interpret the result, and feed the evidence back into future decisions. Companies that control this loop gain something model access alone cannot provide.

Model architectures can spread through papers, open-source releases, employee movement, and commercial services. Carefully produced biological data remain slower to copy because generating them requires specialized laboratories, samples, protocols, and quality controls.

This does not mean every proprietary dataset creates a defensible advantage. A large collection of noisy or irrelevant measurements can reinforce the wrong patterns. The competitive asset is data that answer meaningful biological questions under traceable conditions.

The industry is therefore under pressure to show more than benchmark improvements. Drugmakers, investors, scientists, and regulators need evidence connecting a prediction to laboratory behavior and eventually to patient outcomes.

AlphaFold Shows Both the Opportunity and the Limit

AlphaFold demonstrates what exceptional biological data can unlock, but its success does not turn every drug-development problem into another structure-prediction task.

AlphaFold learned from the Protein Data Bank, a public archive of experimentally determined three-dimensional structures. Those structures were produced through demanding techniques such as X-ray crystallography, nuclear magnetic resonance, and cryogenic electron microscopy.

The archive gave researchers a valuable combination: standardized structural records, decades of scientific curation, and enough diversity to support generalization across many protein families.

According to Google DeepMind’s official AlphaFold overview, the AlphaFold database provides predicted structures for more than 200 million proteins. These predictions expanded access to structural hypotheses at a scale that laboratory methods alone could not match.

The system’s influence is difficult to dismiss. Scientists have used AlphaFold predictions to form hypotheses, interpret experiments, and investigate proteins whose structures were previously unknown.

However, a predicted structure is not a medicine. A therapeutic candidate must bind appropriately, reach the intended tissue, remain stable, avoid toxic effects, and perform under changing biological conditions.

Proteins are also dynamic. They move between conformations, interact with other molecules, and behave differently inside cells than they do in an isolated structural representation.

AlphaFold 3 extended the approach by modeling interactions involving proteins, DNA, RNA, ions, and small molecules. Its research paper reported improved accuracy across several molecular interaction tasks.

That advance made the system more relevant to drug discovery. It did not remove the need for experiments. Model outputs remain predictions whose usefulness depends on the target, available training examples, and the novelty of the proposed interaction.

Nature reported in 2025 that pharmaceutical companies were contributing proprietary structures to a new industry model because public training data were becoming a constraint. The data shortage was especially important for protein complexes and interactions with drug-like molecules.

This development reveals the reversal at the center of the current debate. AlphaFold’s visibility encouraged attention toward algorithms, but its performance depended on years of experimentally generated biological data.

As models converge technically, unique data can become the larger differentiator. A company with a sophisticated architecture but weak experimental coverage risks producing confident predictions that fail when synthesized or tested.

Conversely, a company with well-designed experiments can improve less glamorous models through repeated learning cycles. Each experiment identifies errors, narrows uncertainty, and helps determine which candidate deserves the next expensive test.

This active-learning process selects experiments based on what would provide the most useful new information. It treats model uncertainty as a guide for laboratory work, rather than hiding uncertainty behind a ranked candidate list.

The distinction matters because biological data are not passive fuel. Researchers decide which assay to run, what controls to include, and how to label the outcome. Those choices determine what the model can learn.

Negative results matter too. A failed binding experiment or toxic compound can define important boundaries. Yet negative data are often missing from public literature because successful findings attract more attention and are easier to publish.

That publication bias can distort a model’s picture of chemical and biological space. If training records emphasize successes, predictions can become poorly calibrated about failure.

The strongest AI drug discovery systems will therefore combine computation with disciplined experimental design. They will preserve unsuccessful results, record metadata, and distinguish measured observations from simulated or predicted values.

Google News is highlighting a data story, but the deeper issue is scientific feedback. AI becomes useful when predictions lead to tests whose outcomes improve the next prediction.

The Competitive Divide Is Data Loops Versus Model Scale

The primary contest is not one company against another; it is a closed experimental data loop against a model-first strategy.

A model-first organization can begin with public molecular databases, published research, pretrained systems, and licensed software. This approach lowers the cost of entering computational drug discovery.

It can also produce promising demonstrations quickly. Teams can screen virtual compounds, predict binding poses, summarize research, or generate candidates before building extensive laboratory operations.

The disadvantage appears when a project reaches biology not represented by existing datasets. Public resources can overlap heavily with data used by competitors. Their blind spots can also become shared blind spots.

A closed-loop organization links computational predictions directly with automated or high-throughput experiments. Results return to the model, allowing the system to update its priorities using proprietary evidence.

Recursion Pharmaceuticals represents one visible version of this strategy. Its official platform description emphasizes generating large cellular image datasets and connecting observable cell changes with chemical or genetic interventions.

Absci and other companies combine machine learning with laboratory systems for protein and antibody design. Insilico Medicine has pursued an end-to-end strategy spanning target identification, molecule generation, and clinical development.

Google DeepMind and Isomorphic Labs approach the field from advanced structural modeling and computational drug design. Their work demonstrates the value of algorithms, while their pharmaceutical partnerships also provide access to domain expertise and experimental programs.

These companies do not fit neatly into a single category. Most serious drug-discovery platforms combine models, public evidence, proprietary measurements, and external laboratory work.

The important distinction concerns where each company expects compounding improvement. A model-centered strategy expects better algorithms and more compute to produce the largest gains. A data-loop strategy expects repeated biological experiments to build the stronger advantage.

Current evidence favors a combined approach, with data quality becoming more important as architectures mature. The 2026 AI Index found that newer protein models were not simply becoming larger.

Its analysis noted a move toward smaller, specialized systems trained on curated data or supplemented through retrieval. It also found that larger structure models had not clearly surpassed AlphaFold 3 on a benchmark for protein-small-molecule binding.

That result does not prove model scaling has ended. Benchmarks measure selected tasks and can fail to reflect real drug programs. It does suggest that adding parameters is not a guaranteed path toward better biological predictions.

Data diversity also matters more than raw volume. A model needs evidence across different molecular families, cell types, genetic backgrounds, doses, time points, and experimental methods.

Multimodal data connect several measurement types, such as molecular structures, gene activity, cell images, clinical records, and scientific text. These connections can help a model relate molecular behavior to larger biological outcomes.

However, merging modalities creates its own risks. Records can describe different populations or experimental contexts. A correlation across datasets does not establish that one biological process caused another.

Data provenance, meaning the recorded origin and processing history of a measurement, becomes essential. Researchers need to know which sample, instrument, protocol, and transformation produced each value.

Without provenance, a model can learn laboratory artifacts. It might associate an experimental batch, imaging device, or sample-preparation method with the outcome researchers actually care about.

The problem becomes more serious when autonomous AI agents select targets or experiments. An agent can move quickly through databases and analytical tools, but speed can amplify errors from disconnected evidence.

A reliable biological context layer would connect entities such as genes, proteins, pathways, diseases, compounds, assays, and patient groups. It would also preserve uncertainty and conflicting results.

Knowledge graphs provide one implementation. They represent entities and their relationships in a structured network. Yet a graph remains only as dependable as its source data and relationship definitions.

For enterprise buyers, this changes procurement questions. Comparing model accuracy on a vendor-selected benchmark is insufficient. Buyers should examine data rights, assay coverage, validation methods, provenance, and performance on previously unseen targets.

They should also ask how often computational recommendations survive laboratory testing. A useful platform must improve decisions, not simply generate more candidates for scientists to reject.

The competitive pressure therefore falls on AI drug discovery companies that sell model access without an evidence strategy. It also falls on pharmaceutical companies whose internal data remain trapped in disconnected systems.

Organizations that cannot connect historical experiments may repeat failed work or train systems on incomplete records. Data integration becomes part of research strategy, not a routine information-technology project.

For knowledge workers handling complex technical evidence, a well-maintained searchable knowledge base can improve traceability. In drug discovery, however, document retrieval must complement validated laboratory systems rather than replace them.

More Biological Data Can Still Produce Better-Wrong Answers

The data thesis becomes dangerous when volume is treated as a substitute for relevance, diversity, and independent experimental validation.

Biological datasets reflect choices about which organisms, tissues, diseases, and populations receive study. They also inherit biases from funding priorities, clinical access, and laboratory practice.

A dataset can contain millions of observations while representing only a narrow section of human biology. Repeating similar measurements does not necessarily teach a model how to handle a new disease mechanism.

Batch effects create another problem. These are systematic differences caused by laboratories, instruments, technicians, or processing dates rather than the underlying biology.

Researchers can apply normalization and quality controls, but correction methods may remove meaningful signals or preserve hidden artifacts. Bigger datasets can make a biased result appear more statistically convincing.

Labels are also uncertain. A compound described as inactive may have been tested at an unsuitable concentration. A disease category may include patients with several distinct molecular mechanisms.

Models trained on such labels can learn the organization of the database instead of the organization of biology. They may predict recorded categories accurately while failing to identify a useful therapeutic intervention.

Causal inference poses a deeper challenge. A gene associated with a disease is not automatically a good target. Changing that gene’s activity can trigger compensating pathways or harmful effects elsewhere.

Drug discovery needs perturbational data, which measure what happens after researchers deliberately alter a biological system. Even then, a result from a cultured cell may not transfer to an animal or patient.

Human data are particularly valuable, but privacy, consent, and access requirements limit their use. Clinical records can contain missing values, treatment-selection biases, and inconsistent documentation.

Population representation matters as well. A model trained disproportionately on patients with certain ancestries or healthcare access patterns can perform unevenly for other groups.

Proprietary data create commercial advantages, but secrecy can weaken external scrutiny. Independent researchers cannot easily test whether a private dataset covers the claimed biology or contains systematic errors.

The same tension appeared when drugmakers agreed to contribute private protein structures to industry modeling projects. More structures can improve prediction, but restricted access can widen the gap between commercial and academic research.

Regulators will not approve a therapy because its training dataset is large. They evaluate evidence supporting safety, efficacy, manufacturing quality, and the proposed use.

The US Food and Drug Administration’s AI guidance emphasizes a risk-based credibility assessment for AI models used to support regulatory decisions. Context of use remains central.

A model used to prioritize early experiments faces different consequences from one used to replace evidence in a pivotal decision. The required validation should reflect that difference.

Clinical results remain the hardest test. In 2025, a randomized Phase 2a study reported safety and signs of efficacy for an AI-discovered target and drug combination in idiopathic pulmonary fibrosis.

A clinical milestone matters because it connects computational discovery with patient evidence. Still, a Phase 2a signal does not establish broad clinical success or regulatory approval.

The result should encourage careful optimism, not a declaration that AI has solved drug development. Many candidates fail after showing promise in early research.

AI can accelerate target selection and molecular design while leaving later bottlenecks intact. Toxicology, manufacturing, patient recruitment, dosing, and long-term outcome measurement still require time and evidence.

The strongest criticism of the data-centered view is therefore not that data are unimportant. It is that companies can use “proprietary biological data” as an untestable marketing claim.

A credible data advantage should produce observable outcomes. Candidates should survive prospective experiments more often, uncertainty estimates should be calibrated, and results should reproduce across laboratories.

Companies should also disclose appropriate failures. A platform that reports only successful examples gives buyers no denominator for evaluating performance.

The industry must avoid replacing model hype with data hype. Connected biological records help only when they preserve context, uncertainty, and contradictory evidence.

Three Signals Will Test the Biological Data Thesis

The next proof will come from prospective validation, reproducible data partnerships, and clinical outcomes, in that order.

The first signal is performance on genuinely unseen biological targets. Retrospective benchmarks can accidentally overlap with training data or favor familiar protein families.

Prospective testing starts before experimental outcomes are known. A company predicts which candidates will work, records those predictions, and then runs the assays.

This design limits selective reporting and provides a clearer hit rate. If data-centered platforms consistently outperform model-only baselines, the article’s central judgment becomes stronger.

The comparison must include experimental cost and cycle time. A platform that improves accuracy slightly while demanding an impractical number of assays may not create a useful advantage.

Evidence across independent laboratories would strengthen the case further. Reproduction shows that performance does not depend on one team’s instruments, protocols, or undocumented expertise.

Failure on unseen targets would weaken the data thesis, especially if proprietary platforms perform no better than public models. It would suggest that accumulated measurements remain too narrow or noisy for reliable extrapolation.

The second signal is how pharmaceutical data partnerships operate. Announcements often describe access to private datasets without explaining their composition, quality, or permissible use.

The meaningful question is whether partners can standardize records across laboratories and therapeutic areas. Shared schemas alone will not solve inconsistent experimental design.

Watch for partnerships that publish benchmark methods, provenance standards, or independently testable subsets. These steps would indicate that participants are treating data quality as a scientific problem.

Also watch who retains access to improved models. Private consortia can accelerate commercial research while limiting benefits for academic scientists and smaller biotechnology companies.

A broader public-private data effort would strengthen the overall field. It could create precompetitive resources for common measurement standards while preserving proprietary drug programs.

A series of closed agreements without validation details would offer weaker support. Such deals might reflect competitive fear more than demonstrated scientific progress.

The third signal is movement through clinical development. Laboratory hit rates matter, but patients provide the final test of whether a predicted mechanism translates into medicine.

The most informative programs will include AI in target discovery and molecular design, then report results through well-controlled trials. Clear documentation should show where AI affected decisions.

Successful Phase 2 and Phase 3 outcomes would strengthen the biological data thesis if those programs relied on integrated experimental learning loops. They would connect data strategy with patient benefit.

Repeated clinical failures would require closer analysis. A failure might reveal a weak target hypothesis, poor molecule design, unexpected toxicity, or limitations in trial execution.

That distinction matters because “AI-discovered” covers many workflows. The label does not identify which model, dataset, or human decision contributed to success or failure.

Readers following Google News should therefore look past claims about candidate speed. The better question is whether the system reduces avoidable failures while preserving scientific and regulatory rigor.

AI drug discovery is entering a less theatrical phase. Model demonstrations will continue, but durable progress depends on experiments that challenge predictions instead of merely illustrating them.

Google News has surfaced the right conflict: algorithms can rank possibilities, while biology determines which possibilities survive contact with reality.

The next few months should bring more dataset releases, pharmaceutical partnerships, and platform claims. Evaluate each one through three questions.

Was the prediction made prospectively? Did independent experiments reproduce it? Did the evidence improve a decision that matters for patients?

Those questions offer a practical filter for researchers, enterprise buyers, and technology readers. They also separate a useful biological learning system from an impressive model searching for trustworthy evidence.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page