Different Teams, Same Problem

July 22, 2026

Award Commentary by Candace Ruff, Senior Product Manager, OmicsHQ 

July 22, 2026Receiving a Bio-IT World Best of Show Award for OmicsHQ was a tremendous honor. More importantly, it provided an opportunity to engage with the researchers, data scientists, and technology leaders who ultimately determine whether products like ours succeed or fail. 

Like many exhibitors, we arrived prepared to discuss features, capabilities, and roadmaps. Instead, what stood out most were the conversations. 

Over the course of the conference, I spoke with teams from pharmaceutical companies, AI startups, research organizations, and technology vendors. Their objectives varied significantly. Some were building AI-powered research assistants. Others were searching for external datasets to validate internal findings. Some were developing new analytics platforms, while others were focused on accelerating drug discovery programs. 

Yet despite these differences, I kept hearing the same underlying challenge. 

Not a shortage of data, but a shortage of data that was trustworthy, and usable.  

The Data Availability Problem Has Shifted 

The biomedical community has made extraordinary progress over the last decade. Public repositories now contain millions of samples spanning virtually every disease area, tissue type, and experimental modality imaginable. 

Researchers have access to GEO, SRA, CellxGene, ArrayExpress, and countless other data sources. Every year, the volume of publicly available data continues to grow. 

By almost any measure, we have no shortage of available data. 

Yet many researchers still spend weeks (sometimes months) searching for datasets, evaluating metadata quality, reviewing publications, comparing processing methods, and determining whether studies can be meaningfully combined. 

The bottleneck has shifted. 

The challenge is no longer finding data. The challenge is finding data that is trustworthy, comparable, and relevant to the question being asked. 

One conversation at Bio-IT World captured this perfectly. 

An attendee from an AI startup described their efforts to build a biomedical research assistant powered by public omics data. Access to datasets was not the limiting factor; public data was abundant. The challenge was assembling datasets that had been processed consistently enough and annotated accurately enough for an AI system to reason over them effectively. 

Anyone working with AI today understands this problem. AI systems are remarkably capable, but they are highly dependent on the quality of the information they consume. Inconsistent metadata, missing context, conflicting annotations, and poorly structured data can introduce noise, reduce retrieval quality, and increase the likelihood of incorrect conclusions. 

In other words, AI has increased the value of public data, but it has also exposed the cost of inconsistent data. 

Different Organizations, Same Frustration 

Another memorable conversation came from a scientist working within a pharmaceutical company's immunology organization. 

Their team routinely generates internal bulk RNA-seq, single-cell, and spatial datasets. They also rely heavily on public data to validate findings, establish biological context, and answer questions that internal studies alone cannot address. 

Their challenge sounded remarkably familiar. 

The issue was not locating public datasets. The issue was determining which datasets could be trusted, whether they were comparable to one another, and whether they contained the biological signals needed to answer a specific research question. 

Before scientists can ask, "What does this data tell me?" they often spend significant time answering a more fundamental question: "Can I trust this data?" 

That sentiment surfaced repeatedly throughout the conference. The specific use cases varied. The underlying challenge did not. 

Whether teams were building AI applications, conducting translational research, developing biomarkers, or validating experimental findings, they were all investing substantial effort into data discovery, evaluation, and preparation before any meaningful scientific work could begin. 

For many organizations, this has become an invisible tax on research productivity. 

Why Metadata Matters More Than Ever 

These conversations reinforced something we have increasingly observed across the industry: metadata is no longer a supporting artifact. It is becoming a strategic scientific asset. As dataset volumes continue to grow, metadata quality increasingly determines whether data can be discovered, understood, reused, and incorporated into downstream analyses. Unfortunately, much of the biomedical data ecosystem was built for human interpretation rather than machine reasoning. 

The same disease may be described in multiple ways. Cell type annotations can vary dramatically between studies. Experimental details may be recorded inconsistently or omitted altogether. Researchers often find themselves stitching together information from repositories, publications, supplemental files, and personal expertise simply to understand whether a dataset is relevant. 

This challenge becomes even more significant as organizations invest in AI. 

Many discussions at Bio-IT World focused on AI-ready data, knowledge management, and trusted scientific intelligence. While technologies continue to advance rapidly, the success of these systems ultimately depends on the quality and consistency of the underlying data. No model can fully compensate for incomplete, inconsistent, or poorly organized information. 

Why We Built OmicsHQ 

These are the challenges that inspired OmicsHQ. 

We did not set out to build another repository. We set out to create a more intuitive way for scientists to discover and evaluate omics datasets. 

Our belief is simple: data discovery should begin with biology, not repository navigation. 

Scientists should be able to search using the concepts they actually care about: disease, tissue, cell type, organism, treatment, and experimental context, then quickly determine whether a dataset is relevant to their research. 

Achieving that vision requires much more than aggregating datasets into a single location. It requires harmonized metadata, ontology alignment, standardized annotations, and quality review processes that make datasets easier to compare and understand. 

The goal is not simply to make data searchable. The goal is to make data usable. 

Looking Forward 

One comment from the conference has stayed with me. 

After seeing a demonstration of OmicsHQ, a senior scientist said, "If I had access to this, it would be my first stop when looking for data.

I appreciated the compliment, but what resonated most was what it represented. Researchers are not asking for more repositories. They are asking for trusted starting points. They want to spend less time searching and preparing data and more time answering scientific questions. 

Receiving the Bio-IT World Best of Show Award was meaningful because it suggested that this challenge resonates far beyond our own customers. The organizations we spoke with differed in size, focus, and scientific objectives, but their message was remarkably consistent. 

The next challenge for biomedical informatics is not generating more data. It is making the data we already have easier to discover, trust, and reuse. 

If the conversations at Bio-IT World are any indication, solving that challenge may unlock just as much value as generating the data itself.