PubMedAbstractsSubsetEmbedded
CC-BY-4.0Introduced 2025-03-19
This dataset contains a probabilistic sample of ~2.4 million PubMed abstracts, enriched with precomputed dense embeddings (title + abstract), from the ncbi/MedCPT-Article-Encoder model. It is derived from public metadata made available via the National Library of Medicine (NLM) and was used in the paper Efficient and Reproducible Biomedical QA using Retrieval-Augmented Generation.
Each entry includes:
title: Title of the publicationabstract: Abstract contentPMID: PubMed identifierembedding: 768-dimensional float32 vector from MedCPT