PubMedAbstractsSubsetEmbedded

CC-BY-4.0Introduced 2025-03-19

This dataset contains a probabilistic sample of ~2.4 million PubMed abstracts, enriched with precomputed dense embeddings (title + abstract), from the ncbi/MedCPT-Article-Encoder model. It is derived from public metadata made available via the National Library of Medicine (NLM) and was used in the paper Efficient and Reproducible Biomedical QA using Retrieval-Augmented Generation.

Each entry includes:

  • title: Title of the publication
  • abstract: Abstract content
  • PMID: PubMed identifier
  • embedding: 768-dimensional float32 vector from MedCPT