AI News

Artificial intelligence has produced one of its clearest scientific successes by learning from a vast, carefully assembled dataset. But a new MIT Technology Review analysis argues that this model of progress will be difficult to replicate across most fields—and that AI agents capable of reasoning across tools may be a more practical path forward.

The argument matters because the conditions behind AlphaFold are unusual. Google DeepMind’s protein-structure system benefited from decades of international data collection and a relatively reliable experimental method. Much of biology, chemistry, and materials science lacks comparable datasets, consistent measurements, or standardized procedures. For researchers and AI builders, the implication is that progress may depend less on finding a universal scientific dataset than on building systems that can combine imperfect evidence and test ideas iteratively.

AlphaFold’s success depended on rare conditions

AlphaFold became a landmark for AI in science after predicting three-dimensional protein structures from experimentally measured examples. Demis Hassabis and John Jumper of Google DeepMind shared the 2024 Nobel Prize in Chemistry for work associated with the system.

The MIT Technology Review analysis presents AlphaFold as a profound achievement, but not necessarily a general blueprint for automating discovery. Its training foundation was the Protein Data Bank, which contains roughly 170,000 experimentally validated protein structures. According to the analysis, assembling that resource took 53 years of international cooperation and an estimated $21 billion in experimental work.

That investment was possible in part because protein crystallography produces data that are unusually reproducible. The same conditions do not apply across much of experimental science. Cell lines can change, chemicals can contain trace impurities, and laboratory conditions can affect results. Producing measurements that are sufficiently consistent, accurate, and large-scale for modern neural networks could require new instruments, standards, and research infrastructure.

Some areas—including weather forecasting, much of genomics, and limited parts of chemistry—may still offer the data conditions needed for AlphaFold-like systems. But the analysis says that government support for creating and coordinating scientific datasets will remain important, while many other research problems need a different technical approach in the near term.

Why reasoning matters when the data are messy

Scientists rarely work from perfect evidence. A researcher investigating a drug target might combine molecular docking, known structures, molecular-dynamics simulations, a small number of binding assays, and expert judgment. The process depends on comparing tools, understanding their weaknesses, and revising conclusions as new evidence arrives.

The article argues that AI agents are beginning to model this workflow. An agent is described as a reasoning system connected to digital or physical tools and able to decide how to use them. Rather than applying one highly specialized model to one question, agents can coordinate literature searches, hypothesis generation, critique, experiments, and follow-up analysis.

The example cited is Google AI Co-Scientist, which Google announced in May. Researchers provided a brief and asked the system to investigate how antibiotic resistance spreads between bacterial species. The system used sub-agents for literature-based hypothesis generation, criticism, ranking, and refinement. It reportedly proposed that resistance genes could move between hosts through bacterial viruses.

MIT Technology Review says the hypothesis matched a conclusion reached through a decade of wet-lab work by researchers at Imperial College London. The relevant paper had reportedly not been seen by Co-Scientist and was still under peer review. That makes the example notable, but the evidence available here is still a published analysis’s account rather than an independently documented benchmark or broad validation study. It should be treated as an illustrative result, not proof that scientific agents are ready for unsupervised research.

Agents could change the research workflow

The most immediate opportunity is not replacing scientists. It is reducing the friction involved in moving between evidence, tools, and experiments. A capable agent could search large bodies of literature, propose testable explanations, prepare computational experiments, and record the reasoning and tool calls behind its recommendations.

That record-keeping could address part of the reproducibility crisis. Researchers and funders have long encouraged scientists to share raw data, code, and procedural details, but those tasks are often completed after the main work and may be incomplete. Agents that automatically log their actions could create a more precise record of how a result was produced. This would help another lab reproduce a computational workflow, though it would not by itself solve variation in physical experiments or guarantee that an agent’s reasoning is correct.

Agents could also preserve institutional knowledge. Laboratory methods are often passed between researchers informally or buried in inconsistent notebooks. A standardized repository of agent-assisted work could make prior decisions, failed experiments, and procedural details easier to retrieve. For enterprise research teams, that points to a practical use case: connecting models to approved internal documents, simulation tools, lab systems, and review processes rather than deploying a general chatbot in isolation.

Speed is another potential advantage. The analysis suggests that systems able to read large collections of papers, design many candidate molecules, and learn from failed tests could lower the cost of exploring ideas. The benefit would not simply be faster answers. Cheaper experimentation could make researchers more willing to investigate unusual or high-risk questions.

Evidence, limitations, and deployment risks

The case for scientific agents remains forward-looking. MIT Technology Review explicitly notes that current systems can hallucinate, produce inconsistent judgments, and operate under memory and input constraints. Those limitations are serious in domains where an attractive hypothesis can still be scientifically wrong, unsafe, or impossible to reproduce.

The article’s strongest claims about future improvements, greater reliability, and accelerating research are analysis and interpretation, not established market outcomes. The evidence provided does not include adoption figures, controlled comparisons with human teams, or independent measures of productivity. Nor does it show that agents can reliably control laboratory equipment or manage long-running research programs without close supervision.

For builders, the design challenge is therefore broader than model quality. Scientific agents need provenance, tool permissions, experiment logging, uncertainty reporting, and human review at consequential decision points. They must distinguish a result supported by replicated measurements from a plausible but weak hypothesis. In regulated or safety-sensitive environments, teams will also need clear records of which model, data, software version, and experimental conditions shaped each recommendation.

What to watch next

The next meaningful signals will be evaluations that compare scientific agents with researchers using the same questions, tools, and time budgets. Independent replication of results attributed to Google AI Co-Scientist would provide stronger evidence than a single successful case.

Researchers and buyers should also watch whether agents can maintain coherent research plans over longer periods, cite evidence accurately, recover from failed experiments, and transfer work between laboratories. Another key test will be integration with real scientific infrastructure: electronic lab notebooks, simulation environments, databases, and robotic systems.

Progress on measurement standards will matter as much as progress on models. Where reliable scientific datasets can be created, specialized systems such as AlphaFold may remain highly effective. Elsewhere, agent performance will depend on how well systems handle noisy measurements, conflicting papers, and incomplete institutional knowledge.

Creati.ai perspective

The central lesson is not that data-driven scientific models have failed. AlphaFold demonstrates how powerful they can be when a field has abundant, structured, and trustworthy training data. The more important correction is that most research does not offer those conditions, so AI products for science must be designed around uncertainty rather than assuming it away.

For AI companies and research organizations, the near-term opportunity is an auditable layer that coordinates existing scientific tools and captures the full chain from question to evidence to experiment. That is less spectacular than promising autonomous discovery, but it is more aligned with how science is actually practiced—and easier for expert users to test, challenge, and improve.

Featured

AI for science needs reasoning, not just data

MIT Technology Review argues that AI agents, rather than data-hungry models alone, could make scientific research faster, more reproducible, and broadly useful.