Trends from the Trenches: AI-Ready Life Sciences Data
By Bio-IT World News Staff
June 3, 2026 | Life sciences organizations are sprinting toward AI, but the hardest part is rarely the model. It is the data foundation: unstructured files scattered across instruments, legacy NAS, multiple sites, and multiple clouds. That fragmentation creates data silos, slows down analysis, and adds constant risk of human error when teams manually copy datasets to wherever compute happens to be available.
In the latest episode of Trends from the Trenches, Adam Marko, CTO of Life Science Field at Hammerspace, speaks with host Jessica StLouis about AI-ready data as a practical outcome. Researchers see one familiar file system, while the platform handles where data lives, how it moves, and who can access it. That shift reduces “data friction,” allowing scientists to spend less time hunting files and more time conducting research.
Data orchestration can act as an operating layer for modern research IT. Instead of treating storage as separate islands, a unified file system can span on-prem infrastructure and cloud storage while preserving directory structure. Rule-based automation then places files on the right tier at the right time, such as keeping active datasets on high-performance NVMe and moving cold data to cheaper disk or cloud after a defined period.
Just as important, orchestration is not a free-for-all method. Permissions, read-only copies, and governance controls can follow the data. For regulated biomedical research, this matters because collaboration must coexist with least-privilege access and defensible controls.
AI infrastructure adds new pressure points, especially GPU availability and flash storage constraints. GPU capacity can shift daily across providers and locations, so teams need to run workloads wherever GPUs are, without rewriting workflows. At the same time, SSD and NVMe shortages make it unrealistic to keep everything on premium media. A tiered storage architecture becomes a pragmatic answer: a smaller high-speed NVMe layer for training and inference pipelines, backed by larger capacity tiers on spinning disk or cloud.
With orchestration, researchers compute directly against the same namespace while the system automatically clears space, avoids manual shuffling, and keeps throughput consistent for AI workloads and HPC pipelines.
Marko also breaks down what makes AI readiness in life sciences uniquely challenging: unpredictable data volumes driven by rapidly evolving instruments, diverse file types, and compliance requirements tied to patient privacy, data retention, and country-specific rules. Unlike industries where data growth is easier to forecast, a new sequencer firmware update or a faster CryoEM workflow can change storage and network needs overnight.
To learn more about what’s next in “healthcare HPC,” a partnership with NVIDIA, and faster time to insight, listen to the Trends from the Trenchespodcast.


