Healthcare data is generated across many forms, including genomic sequences, medical images, clinical notes, and laboratory records, yet most artificial intelligence systems in medicine are still built to interpret only one of these formats at a time. Researcher Sasi Kumar Kolla has taken up this gap in a new paper that examines how foundation deep learning models can be designed to work across multiple data types at once, rather than being confined to a single modality.
His paper, titled
Why Single-Modality Models Fall Short
According to Kolla, most existing deep learning research in medicine has focused on one data type and one task at a time, such as classifying a pathology image or predicting risk from a genomic marker. This narrow framing, he argues, makes it difficult to understand the broader biological mechanisms behind a prediction, since real disease processes rarely show up in just one kind of data.
“Biological developments concentrate on the integration of different data types, such as genomics and transcriptomics with imaging,” the paper notes, pointing out that no true end-to-end multimodal foundation model has yet been built specifically for precision medicine. Kolla’s research sets out to address that gap by analyzing how such data could be integrated, mined, and organized as a coordinated data ecosystem.
Building Toward Multimodal Foundation Models
The paper describes how foundation models, pretrained on large volumes of data and later adapted to specific tasks, could be extended beyond the vision and language domains where they have gained traction and applied to biomedical research. Kolla outlines an approach in which each data modality, whether an image, a genomic sequence, or a clinical record, is encoded separately and then combined through a shared representation layer, allowing a single model to draw on multiple sources of evidence.
The work also walks through the underlying mathematics of these architectures, including self-attention mechanisms and contrastive learning methods used to align different data types, such as pairing medical images with corresponding reports. These techniques, borrowed and adapted from natural language processing research, form the technical backbone of the multimodal approach Kolla proposes.
Data Ecosystems and Governance
A substantial portion of the paper is devoted to the practical challenges of building the data infrastructure such models would require. Kolla writes that obtaining large, high-quality, and well-curated datasets remains one of the hardest parts of any machine learning project in this space, particularly because genomics, imaging, and health record data are often collected and stored using inconsistent standards.
The paper discusses the importance of quality assurance processes, standardized data curation, and governance protocols that account for privacy and consent when working with sensitive health information. Kolla notes that as research institutions and organizations pursue partnerships to access this kind of data, ongoing attention to security and ethical use needs to remain central to how these systems are developed, rather than treated as an afterthought.
Addressing Bias and Interpretability
The research also examines two challenges that Kolla identifies as central to whether foundation models can be trusted in a research or clinical research setting: data bias and interpretability. He points out that datasets which underrepresent certain populations can lead to models that perform unevenly, and that this risk carries over into biomedical applications where imaging equipment, patient populations, and data collection practices vary widely between institutions.
On interpretability, the paper argues that a model’s predictions are only as useful as a researcher’s or clinician’s ability to understand how they were reached. Kolla writes that building confidence in a model’s output, and being transparent about where it may fall short, is often as important as the accuracy of the prediction itself. The paper calls for open benchmarking protocols, shared datasets, and published model weights as ways to support reproducibility and independent scrutiny across the research community.
Applications Across Genomics and Imaging
The paper surveys several areas where multimodal approaches are already being explored, including the integration of genomic and transcriptomic data to study gene expression patterns, and the combination of radiology and radiomics data to support multi-organ imaging research. Kolla also discusses a hypergraph-based method for modeling relationships between biological samples and conditions, offered as one example of how researchers might structure multimodal data before it is used in downstream modeling work.
A Broader Research Conversation
Kolla’s paper closes by framing this work as an early step in a longer research trajectory, one that will require continued collaboration across data science, genomics, and clinical research to mature. He points to open questions around benchmarking standards and dataset diversity as areas that will shape how quickly multimodal foundation models can be responsibly developed for biomedical research.
Kolla will discuss related themes in AI research and development at GatherVerse AI Evolve 2026, where he is scheduled to appear on a
This story was distributed as a release by Jon Stojan under HackerNoon’s Business Blogging Program.