Visual Question Answering
English
biology
medical
Pull Figure

Paper : CVPR 2025     |     Website: Biomedica     |     Training instructions: OpenCLIP     |     Tutorial: Google Colab

Model Name: BMC-LongCLIP+

Abstract

The development of vision-language models (VLMs) is driven by large-scale and diverse multi-modal datasets. However, progress toward generalist biomedical VLMs is limited by the lack of annotated, publicly accessible datasets across biology and medicine. Existing efforts are limited to narrow domains, missing the opportunity to leverage the full diversity of biomedical knowledge encoded in scientific literature. To address this gap, we introduce BIOMEDICA: a scalable, open-source framework to extract, annotate, and serialize the entirety of the PubMed Central Open Access subset into an easy-to-use, publicly accessible dataset. Our framework produces a comprehensive archive with over 24 million unique image-text pairs from over 6 million articles. Metadata and expert-guided annotations are additionally provided. We demonstrate the utility and accessibility of our resource by releasing BMCA-LIP, a suite of CLIP-style models continuously pre-trained on BIOMEDICA dataset via streaming (eliminating the need to download 27 TB of data locally). On average, our models achieve state-of-the-art performance across 40 tasks — spanning pathology, radiology, ophthalmology, dermatology, surgery, molecular biology, parasitology, and cell biology — excelling in zero-shot classification with 5.57% average improvement (as high as 26.93% and 17.63% gains in surgery and ophthalmology, respectively) and stronger image-text retrieval while using 10x less compute.

Baseline Comparison Against Frontier Models

Benchmark Model Context Batch T2I R@1 T2I R@5 T2I R@10 I2T R@1 I2T R@5 I2T R@10
CXR PMC-CLIP 77 128 0.0 0.5 0.7 0.2 1.0 1.6
CXR BiomedCLIP 256 4K 0.5 2.6 5.7 0.6 3.3 5.5
CXR BMC-CLIP 77 8K 0.1 1.1 2.9 0.3 1.9 3.4
CXR BMC-LongCLIP+ 512 16K 1.9 7.1 12.2 3.0 9.5 14.5
PMC PMC-CLIP 77 128 0.2 0.7 1.2 0.1 0.7 1.2
PMC MedSigLIP 77 N/A 20.1 37.0 46.0 30.9 49.0 60.1
PMC BiomedCLIP 256 4K 68.8 86.2 91.1 73.3 89.3 93.7
PMC BMC-CLIP 77 8K 49.0 67.6 74.0 40.8 60.4 68.4
PMC BMC-LongCLIP+ 512 16K 80.8 91.2 94.4 79.7 90.6 93.8

Acknowledgments

This work is supported by an NVIDIA Academic Grant.


Citation

@article{sun2025no,
title={No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models},
author={Sun, Min Woo and others},
journal={arXiv preprint arXiv:2510.03978},
year={2025}

@inproceedings{lozano2025biomedica,
  title={Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature},
  author={Lozano, Alejandro and Sun, Min Woo and Burgess, James and Chen, Liangyu and Nirschl, Jeffrey J and Gu, Jeffrey and Lopez, Ivan and Aklilu, Josiah and Rau, Anita and Katzer, Austin Wolfgang and others},
  booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  pages={19724--19735},
  year={2025},
  organization={IEEE}
}

@article{lozano2025large,
  title={A large-scale vision-language dataset derived from open scientific literature to advance biomedical generalist ai},
  author={Lozano, Alejandro and Sun, Min Woo and Burgess, James and Nirschl, Jeffrey J and Polzak, Christopher and Zhang, Yuhui and Chen, Liangyu and Gu, Jeffrey and Lopez, Ivan and Aklilu, Josiah and others},
  journal={arXiv preprint arXiv:2503.22727},
  year={2025}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train BIOMEDICA/BMC_CLIP_CF

Collection including BIOMEDICA/BMC_CLIP_CF

Papers for BIOMEDICA/BMC_CLIP_CF