What Scientists Actually Need From AI
Contributed Commentary by Stavroula Ntoufa, Ph.D., Causaly
August 28, 2026 | A target looks good on paper. The literature supports it, the mechanism is plausible, and three independent groups have published findings that point the same direction. There is also a contradictory result, which someone found early, mentioned in a meeting, and set aside because it seemed less persuasive than the rest. The team builds its case, and the program advances. Years later, in preclinical validation or, worse, in the clinic, that discarded finding turns out to have been the important one.
I spent 14 years in research before moving into scientific AI. Judging a single paper was rarely the difficulty; a trained scientist reads the methods and the statistics and can see what holds up. What trips people up is finding the right papers to begin with and then seeing past their own investment in the hypothesis. When you believe in a project, you give the supporting evidence more weight than the study that complicates it.
The hard calls look deceptively ordinary. When two studies report opposite results, do they actually conflict, or is one working in a mouse model and the other in cells, measuring different endpoints by different statistical standards? A finding is highly cited but rests on a single experiment nobody has replicated. A safety signal shows up in one dataset and nowhere else, and only if you squint. The evidence in each case is perfectly readable. What makes the call hard is that the person weighing it usually has a stake in the answer, and the protocol says nothing about any of it.
The Bottleneck Was Never Search
Most of the current conversation about AI in biomedical research is about search. It’s faster now, summaries are better, and literature that used to take weeks to assemble can now be pulled together in an afternoon. Those are real gains and I don’t want to minimize them. But faster search doesn’t touch the harder problem and can make it worse; it’s now easier to pile up evidence for the conclusion you already “favor.”
Consider what actually goes wrong with drug discovery. Expensive failures come from evidence that at first seemed stronger than it was, from contradictions that no one ever surfaced, and from uncertainty that stayed hidden long enough that stopping the program became politically difficult. A target advances through a review board on the strength of a case that seemed defensible at the time, and resources accumulate behind it. Varespladib is one documented example. Years worth of biomarker data suggested that inhibiting the sPLA2 enzyme would lower cardiovascular risk, and the drug moved the markers as predicted. But the phase 3 trial was halted in 2012 after the drug produced no benefit in more than 5,000 patients and raised the rate of heart attacks. A marker of risk had been mistaken for a cause.
The Judgment That Never Gets Written Down
The reasoning that prevents this outcome is largely invisible. Research organizations document their conclusions far more reliably than they document how those conclusions are reached. Standard operating procedures describe the steps of a workflow: what to do first, what to do next, and what completion looks like. They don’t describe how a senior scientist decides that one contradictory study outweighs 10 that agree with it, or which methodological differences explain a discrepancy and which ones point to a real problem with the hypothesis. Scientists sometimes call this developing a feel for the literature. It builds over a career from thousands of small decisions about what deserves confidence, and they rarely put it into words even to themselves.
Evaluating Evidence is a Different Job from Finding It
This is where I think AI has its most useful role. A system built to help a scientist evaluate evidence looks quite different from one built to help a scientist find it.
The difference is in what it shows you. Contradictions stay visible instead of being smoothed into a coherent-sounding summary. Six papers citing the same original dataset don’t count as six independent confirmations. Every claim traces back to the experiment that produced it. Scientists don’t read words. They read figures, tables, methods, and numbers, and they assess each paper on its own terms.
Why Researchers Withhold Their Trust
That last point explains why so many general-purpose AI tools have struggled to gain traction in research settings. Scientists ask a question, get a plausible answer, can't verify where it came from, and stop trusting the system. Their training tells them to be skeptical in exactly this situation, and they apply that training to a new kind of source. The researchers I talk to assume AI can produce an answer. Their question is whether they can defend it.
Scientists evaluating these systems care about qualities that benchmarks largely overlook. Accuracy scores and summary quality tell you something, but they say little about whether a system will improve a research organization’s decisions. Researchers want to know whether the system surfaced conflicting evidence or quietly smoothed it away. They want to trace a conclusion back to the data behind it. It’s hard to build a workflow around a system that can’t answer those questions.
Closing the Distance Between Evidence and Decision
As evidence becomes easier to gather, interpreting it becomes the hard work: knowing which findings hold up, and when the evidence doesn't yet justify a decision. A system that produces answers without exposing its reasoning just moves the failure point downstream, where it costs more to fix.
Science advances by evaluating evidence. The volume of biomedical literature long ago passed the point where any individual can synthesize it comprehensively, and AI offers one of the few practical ways to work at that scale. But the distance between the evidence and the decision is where research succeeds or fails, and it's where AI has the most potential to help. The systems that matter will be the ones that help scientists close that distance with their reasoning intact and defend the decisions they make.
Director of scientific affairs, Stavroula Ntoufa is a molecular biologist and geneticist with deep expertise in immunology, hematology, and oncology, holding a doctorate from the University of Athens. Before joining Causaly, she spent 13 years on EU-funded multicenter research programs across Uppsala University, Università Vita-Salute San Raffaele and CERTH, with a particular focus on chronic lymphocytic leukemia. Her work has been published in leading international journals and spans toll-like receptor signaling, immunogenetic subsets, and the molecular basis of CLL progression. At Causaly, she brings that translational research lens to how scientific teams in pharma find, evaluate and act on biomedical evidence. She can be reached at [email protected].


