Network Effects: John Quackenbush on How Networks Illuminate Biology Beyond Gene Expression
By Allison Proffitt
July 9, 2026 | John Quackenbush is a veteran of both life sciences research and the Bio-IT World Conference & Expo with plenty of history and context to ensure that he’s not swayed by the shiny and hyped. So at a meeting rife with AI buzzwords, when he starts talking about AI, he starts with IBM Watson.
“Ten years ago, Watson was the great AI system that was going to solve all of biology. We were going to feed in all of human knowledge and cures for cancer were going to pop out. And it still hasn’t happened.” The technology isn’t bad, Quackenbush was quick to say, but it is limited by its input data.
Quackenbush, a professor at Harvard’s T.H. Chan School of Public Health and the Dana-Farber Cancer Institute, has spent his career wrestling with the link between genotype and phenotype. The human body contains more than 250 distinct cell types, each carrying essentially identical DNA. Neurons grow long axons and manufacture neurotransmitters. Pancreatic cells stay compact and produce insulin. Neither switches roles. Yet both are playing from the same genomic score.
“Biology is much more complex than simply looking at a parts list,” Quackenbush said.
Thus AI systems trained for medical applications need nuanced and complex data. The most important data problem Quackenbush sees right now—among several—is underspecification of data models.
“There’s no robust underlying model,” he explained. “We’re looking for associations, but association isn’t causation.” The stakes, he noted drily, are somewhat higher than Netflix recommendations. “If you predict the wrong drug to give somebody, the consequences are a little bit greater.”
For Quackenbush, there’s an important hint in a 1997 computer science paper.
No Free Lunch
Quackenbush cited the opening line of David Wolpert and William Macready’s “No free lunch theorems for optimization” (IEEE Transactions on Evolutionary Computation; DOI: 10.1109/4235.585893) with evident relish: a pointed rebuke of “general purpose black box optimization algorithms that exploit limited knowledge concerning the optimization problem on which they are run.” Translation: you can’t build a meaningful model without knowing something about the system you’re modeling.
Wolpert and Macready concluded by highlighting “the importance of incorporating problem-specific knowledge into the behavior of the algorithm.” Quackenbush agrees. If you know something about the biology, build it in as a constraint or starting point.
Quackenbush views networks as the ideal way to impose biological constraints on the systems he wants to study. Where a traditional analysis might flag which gene expression in a disease state, a network approach asks how the relationships between genes change and which regulators gain or lose influence over time.
There is no single right network, Quackenbush concedes. Networks change in each tissue and over the course of disease progression. “The real question you want to ask is: does it inform our understanding of the biology of the system we’re trying to study? And the answer is yes, in many instances.”
The Collective Noun for Networks? A Zoo
Over the years, Quackenbush and his collaborators have assembled what they call the netZoo — a suite of open-source computational tools, each named after an animal, for building and analyzing biological networks. The founding resident is PANDA (Passing Attributes between Networks for Data Assimilation), which integrates three types of data — transcription factor binding motifs, protein-protein interactions, and gene expression correlations — to reconstruct the regulatory network active in a given tissue or condition.
From PANDA came LIONESS, CONDOR, ALPACA, and others, all freely available in R and Python, all accompanied by documentation in Jupyter Notebooks. “If we write a paper, we also write a Jupyter Notebook, so our papers are fully reproducible and executable,” Quackenbush said.
One of the most consequential extensions was almost an accident.
A team including Marieke Kuijjer, Matt Tung, and Kimbie Glass was trying to estimate confidence in regulatory networks when they realized the math had an unexpected implication. If you build a network from a population and then rebuild it with one person removed, the difference tells you something about that individual’s unique regulatory signature — their personal network. Repeating the process for every person in the dataset produces not one network but a collection of them, one per individual.
The paper describing this technique — called LIONESS — spent 394 days under journal review, accumulated more than 30 pages of supplemental material, and was ultimately rejected after an anonymous senior reviewer declared, “I believe the mathematics. I just don’t feel that it’s right.”
The team posted the paper to bioRxiv, gathered 60 citations without formal publication, and was eventually published in iScience without further trouble (DOI: 10.1016/j.isci.2019.03.021).
Networks of Networks
The individual-network approach has enabled a line of research that Quackenbush finds particularly compelling: the systematic study of sex differences in gene regulation and disease.
Analyzing data from GTEx — the Genotype-Tissue Expression portal with gene expression profiles from 29 tissues across roughly 900 individuals — Quackenbush’s group generated more than 8,000 individual regulatory networks and compared them between biological males and females. The differences they found were tissue-specific, biologically coherent, and, in some cases, clinically meaningful.
In subcutaneous fat, for instance, networks governing thermoregulation differed markedly between sexes, offering a network-level explanation for a dispute familiar to couples everywhere.
The clinical implications sharpened when the same tools were applied to colorectal cancer. Male and female regulatory networks differed in two domains: immune processes (potentially explaining why men have higher colorectal cancer risk) and the regulation of drug transport and metabolism. That second difference predicted something measurable: male colorectal cancer patients respond better to chemotherapy than female patients.
The decisive test came when the team divided female patients into two groups based on whether their regulatory networks more closely resembled the typical male or female pattern. The group with more male-like networks had better outcomes — consistent with the chemotherapy response advantage.
In every case, Quackenbush argues, networks revealed something about the biology that gene co-expression or expression data alone could not.
Time Is the Missing Dimension
Networks change based on the individual, the tissue, and timeline, and Quackenbush’s currently excited about exploring this continuum.
We often think about health and disease as discrete, opposite states, he said, but that’s not how health and disease really interact. It’s a continuum that overlaps as health wanes and disease progresses. And if we take the continuum view, he said, “We can model, even with static snapshots—which is most of the data we collect—the [gene] regulatory process over time.”
One approach, developed with collaborators Viola Fanfani and Jonas Fischer, uses autoencoders to learn a “pseudo-time” ordering of tumor and adjacent-normal samples from lung adenocarcinoma. When thousands of tissue samples are arranged along this learned trajectory, the expression of known tumor-aggressiveness markers tracks the progression cleanly, suggesting the ordering captures something real about disease evolution, even though no individual sample was collected at multiple timepoints.
The next step was a method called PHOENIX (NeuralODEs applied to gene regulatory networks), which Quackenbush described as the first approach capable of modeling time-dependent gene regulation at genome-wide scale without discarding genes in advance (Genome Biology, DOI: 10.1186/s13059-024-03264-0). The key was incorporating a prior network estimate as a soft constraint — exactly the “problem-specific knowledge” the No Free Lunch theorem had demanded.
But how do we estimate changes to the networks themselves over time? In swoops PARROT (Phase-Altering Regulatory Rewiring Over Time). Built by former postdoc Megha Padi and her PhD student Chen Chen, PARROT applies change-point analysis from control theory to sequences of regulatory networks: stepping through time and asking at each point whether the data is better explained by one network model or two. When the answer shifts to two, the system has undergone a state transition and the transcription factors with the largest changes in regulatory targets are likely the drivers.
“We can find the regulators that changed … the state of the network,” he explained. “What you can really think about doing is finding not mutational drivers of diseases like cancer, but regulatory drivers.”
PhD student Lauren Hsu has been applying PARROT to the transition between senescent and non-senescent breast cancer cells. For each regulatory driver identified computationally, targeting it with a small molecule altered the predicted phenotype. “Every single one she’s tried has altered the phenotype,” Quackenbush said.
Quackenbush was careful to frame what the network approach can and cannot claim. “There’s no single right network,” he repeated. “The network in each tissue, in each state, in each individual is unique.”
But again and again he has found that condition-specific networks, built from real data with biologically informed constraints, consistently reveal insight into disease that gene lists and expression profiles alone miss — structure that predicts clinical outcomes, illuminates biological differences, and, increasingly, identifies therapeutic targets.


