Federation Plus Fine Tuning: The Push for Federated Learning Models Continues
By Allison Proffitt
June 4, 2026 | Federated learning is changing the rules in drug discovery, argued a group of speakers at last month’s Bio-IT World Conference & Expo. The opportunities are robust, but success depends on establishing trust and fine-tuning models for individual applications.
Federated learning has been around for a while. MELLODDY, a 10-pharma federated learning project, was published in 2023 focused on Quantitative Structure Activity Relationship (QSAR). The general idea is to train foundation models on various institutions’ proprietary data, but only the model weights move from institution to institution.
“Critically, the model goes to the data. So the model is then moved into that environment, trains on that data, … and changes to the model weights or gradients are then integrated into the local model. And you can repeat that,” explains Jonathan B. Gilbert, PhD, Senior Director, Ecosystem Growth and Contributor Partnerships, Eli Lilly and Company. “This idea then allows this model to improve while still maintaining the privacy of the individual datasets that are never exposed to any of the other parties, never exposed outside of their private environment.”
Woody Sherman, PhD, Founder and Chief Innovation Officer at PsiThera, is also the chair of the OpenFold executive committee, a consortium of more than 40 biotech, pharma, and techbio groups dedicated to delivering open-source AI tools for solving structural biology and drug design problems. From his vantage point, Sherman is seeing a tipping point in open-source tools and drug discovery. As the volume of biological data expands, he said, “We’re going to need these open platforms that we can all build on. We can’t all build our own foundation models from scratch; it just doesn’t make sense.”
Sherman envisions open platforms upon which groups can interact pre-competitively to build the best foundation model, train it in a federated way, and then fine-tune the model with their own data.
OpenFold’s own projects include OpenFold3, OpenStability, RF4 for protein design and more.
DeepMind’s AlphaFold, of course, changed the protein folding space almost overnight, advancing the field in a way that CASP competitions never did, said Mohammad AlQuraishi, PhD, Assistant Professor, Systems Biology, Columbia University and cofounder of OpenFold. AlphaFold2’s critical offering, AlQuraishi said, was providing calibrated predictions. “It not only gave you an answer, but it gave you confidence in that answer.”
But even AlphaFold2 can’t do everything. AlQuraishi noted that AlphaFold2 couldn’t adeptly handle complexes, ligands, or ions. AlphaFold3, then, improved upon the previous model, but accuracy for nucleic acids was still poor, ligands were iffy, and the model had limited availability to predict conformational changes.
There was room for improvement. And OpenFold—open in both data and code, fully reproducible—meets that need as, “a set of tools that allows you to build these types of models and extend them and apply them.”
AlQuraishi believes the next five years will see improved generalization, speed, and scale thanks to open foundation models like OpenFold.
Open Models in a Federated Environment
Lilly’s Gilbert mapped the next step: moving open foundation models to a federated system that trains on proprietary data. It’s a problem Lilly has tackled head on with TuneLab, which the company launched in September 2025. Through TuneLab, Lilly provides its AI/ML models to biotechs at zero cost other than contributing their own datasets to help improve them.
“These are the same models we use every day,” Gilbert confirmed. “These are models that have been trained on decades of internal datasets. It’s estimated to be over a billion dollars in data that have been brought into the models by Lilly.”
Gilbert reported an “explosion” of excitement around TuneLab since the beginning of the year and said Lilly now has more than 75 partners across three continents and dozens of countries. Gilbert said there are about 40 models focused on small molecule ADME-tox and antibody developability.
For any federated model, Gilbert called out the “criticality of allowing a model to both be improving its performance, but also making it simple and easy to use.” Lilly prioritized the user interface of TuneLab, so partners can easily log in to a website and make predictions. The morning of the event, Lilly announced a collaboration with Collaborative Drug Discovery to integrate TuneLab into CDD Vault to allow access to models in more environments.
Doing the Right Thing the Right Way
The problem has always been that federated learning lacks a strong business model. But José-Tomás Prieto, PhD, Director of AI Programs, Apheris, likened it to a seatbelt. While you might prefer to drive without one, the alarms make that annoying. “Federated learning is kind of the same way in the sense that it makes it very difficult to behave badly when you are looking at very sensitive, sometimes secret, proprietary data.”
Meaning: it’s the right thing to do, whether you want to or not.
Apheris offers products for federated networks for drug discovery, enabling their customers to join or build their own federated learning networks, while protecting sensitive data. Apheris undergirds the AISB (AI Structural Biology) Network with nine pharma focused on protein-protein interactions and binding affinity. They also support the ADMET Network with Recursion and other pharma focused on small-molecule property prediction. They are co-driving the antibody developability network with Ginkgo Dataworks.
“The common denominator for these networks is that the models are trained on proprietary data in a secure way,” Prieto said. Proprietary data never leaves the companies that own it—the nodes in the network—and the models are delivered to the nodes to use. “They take care of their internal programs and use them however they have to use them,” he added.
One strong advantage to a federated approach is that models are not biased toward well-characterized public data. Instead, they are trained on more messy, diverse real data.
“If there is something to learn about the AI world today, it is that you cannot necessarily model your way out of a data problem,” he said.
Prieto challenged the audience and his fellow panel members to acknowledge that federated learning is not a plug and play solution. Networks required engineering rigor, he said. “Each one of these companies has their own network constraints. They have their own firewall rules. They have their own compute window that they have to negotiate with the cloud providers to make sure the compute comes in one at the right time. They all have their own particular approval processes.”
Fine-Tuning Advantage
Prieto makes a case for fine tuning federated models to achieve the best outcomes—better than fine-tuned models on internal data alone; better than federated models alone. “The reality is that federated networks do not know all of your targets, right? They don’t know the most recent data that you’re working with, the specific program data. This gap can actually be filled with fine tuning,” he said. Fine tuning on local program data is what converts model performance into drug program impact, he said.
Finally, emphasized that federation is a standing capability, not a one-time project. “You should not see this as one thing, but more like a repeatable, operational model that through deployments, through governance mechanisms can keep going in your internal operations.”
One of Prieto’s examples cited SandboxAQ, a spinout from Google in 2022 that focuses on large quantitative models for scientific discovery. The SandboxAQ suite of models are trained not on the written word of the internet, but on physics, chemistry, and biology explained Arman Zaribafiyan, PhD, Head of Strategic Alliances, AI Simulation, SandboxAQ. These models are deployed in more than 20 discovery programs with large pharma, he added.
One of the key differentiators, Zaribafiyan said, is SandboxAQ’s skill at maximizing GPU compute, and the company has released some models in conjunction with NVIDIA recently.
There’s an explosion of models right now, Zaribafiyan said, models that are great at benchmarking and publishing, but they fail to generalize to real drug discovery. Here, Zaribafiyan also touted the value of fine tuning. “Even a small fine-tuning effort can dramatically change the predictive accuracy of these models when we are looking at a project-specific workflow,” he said. “I think over the next few years we’re going to see more and more of these federated platforms being used for both pooling data, but also for fine tuning these models on project-specific data.”
During the event, Zaribafiyan announced SandboxAQ’s collaboration with Claude so that the large quantitative models from Sandbox can be directed through the Claude real-world interface.
Sustainability Angle
Finally, Christina Taylor, PhD, Senior Science Fellow and Computational Molecular Design Lead, Bayer, shared how Bayer has been leveraging foundational models to drive decisions in both crop science and pharma.
Taylor is also a member of the OpenFold executive team and she highlighted the importance of the consortium and other community-driven efforts to drive pre-competitive innovation. “Really by sharing some of these foundational architectures allows everyone to be able to drive biomolecular AI forward and really has driven some very quick advancements we’ve seen in the field over the past few years,” she said. She also highlighted the sustainability of developing foundation models together so that every research group isn’t doing its own resource-intensive training.
Bayer is also committed to fine tuning, she said. “We’re using [fine tuning] quite regularly to improve our development of biomolecular pharmaceuticals, as well as some of our crop science traits for development of herbicide tolerance traits, insect control traits, etc,” she said.
Barriers to Adoption: Technical, Organizational, and Legal
In the Q&A session after the individual presentations, the panel addressed the community’s concerns about what’s holding federated learning back. Prieto categorized the challenges into three broad buckets: technical (infrastructure and data preparation), organizational (cultures not always oriented toward collaboration), and legal/contractual (federated learning introduces novel collaboration structures that legal teams are still adapting to). He expressed optimism that progress is being made across all three dimensions.
Technical questions were at the fore. An audience question raised the risk of LLMs being “jailbroken” to regurgitate training data. Prieto explained that Apheris conducts rigorous threat modeling before any model is released—actively attempting reverse engineering and membership inference attacks to verify that data cannot be extracted from trained models. Gilbert added that Lilly works with Rhino Health, a company with over a decade of federated computing experience across many industries and conducts its own security analyses given the sensitivity of Lilly’s own 20+ years of proprietary data.
As for which data types are best suited to a federated learning environment, AlQuraishi highlighted structural and binding data as strong candidates, given the vastness of chemical space and the natural benefit of pooling diverse explorations across companies. He distinguished between a “generalization” regime—where any additional data improves the model’s physical understanding—and a “memorization” regime, where gains are target-specific. Gilbert emphasized data quality and metadata standardization, noting that Lilly’s TuneLab actually provides contribution bonuses for data generated under standardized protocols. He also pointed to in vivo PK and toxicology modeling as an area where federated learning could be transformative: these datasets are scarce, expensive, and rarely shared publicly, yet predicting ADMET properties late in the pipeline is critically important.
The organizational challenges are perhaps the oldest. Gilbert noted a historical nervousness around allowing even privacy-preserving models to access proprietary data but argued that scientific gains in model performance has increased comfort. Zaribafiyan contended that the cost of staying out of federated learning frameworks is higher than the risk of participating, since data quality is the foundation of any real-world drug discovery workflow. Prieto added that well-designed federated learning networks include clear contribution and incentive mechanisms—analyzing the diversity and novelty of each party’s data contributions—to ensure participants receive meaningful value proportional to what they put in. Trust, he emphasized, is built not through goodwill alone but through legal, contractual, and process-level mechanisms.
Finally, addressing legal, contractual and funding issues, AlQuraishi noted that while there has been some appetite for government funding (as seen with OpenFold and OpenBind), the U.S. funding infrastructure is not well-suited to the engineering- and compute-heavy nature of modern AI drug discovery work. He pointed to the Transatlantic Initiative as one of the first philanthropic sources to recognize the value of software infrastructure, but suggested the broader ecosystem still has a long way to go.


