数据集导航 · 医学文本与大模型
医学文本与大模型数据集(第 8 页)
思陌医疗数据服务整理的医学文本与大模型公开数据集,共 298 个,覆盖该领域多模态数据,助力医疗AI研发选型。
全部
肿瘤 579
神经系统疾病 507
心血管系统疾病 462
耳鼻咽喉疾病 451
眼病 412
皮肤与结缔组织疾病 399
泌尿生殖系统疾病 362
内分泌系统疾病 352
医学文本与大模型 285
呼吸道疾病 245
口颌系统疾病 219
消化系统疾病 200
免疫系统疾病 137
通用医学数据集 136
血液与淋巴系统疾病 106
感染 96
创伤与损伤 66
肌肉骨骼疾病 14
化学诱发性障碍 4
环境因素所致疾病 3
共 285 个数据集 · 第 8 / 10 页
| 数据集名称 | 系统分类 | 病种 | 数据模态 | 任务类型 | 描述 | 来源 | 原始地址 | 操作 |
|---|---|---|---|---|---|---|---|---|
| Healthcare Information Retrieval Dataset | 医学文本与大模型 | 文本 | 信息检索 | Dataset for high-accuracy search and relevance assessment within electronic medical record systems | Kaggle | https://www.kaggle.com/datasets/colabsss/healthcare-information-retrieval-dataset | 访问 | |
| WorldMedQA-V | 医学文本与大模型 | 多模态 | 医学问答 | Multilingual and multimodal medical examination dataset from Brazil, Israel, Japan, Spain with medical images | Hugging Face | https://huggingface.co/datasets/WorldMedQA/V | 访问 | |
| MedTrinity-25M | 医学文本与大模型 | 多模态 | 视觉问答 | Large-scale multimodal dataset with multigranular annotations, 25M images across 10 modalities, 65+ diseases | Hugging Face | https://github.com/UCSC-VLAA/MedTrinity-25M | 访问 | |
| MedPT | 医学文本与大模型 | 文本 | 医学问答 | Large-scale medical question answering dataset for Brazilian-Portuguese speakers, 384K+ entries | Hugging Face | https://huggingface.co/datasets/AKCIT/MedPT | 访问 | |
| UltraMedical | 医学文本与大模型 | 文本 | 指令微调 | Large-scale biomedical instruction dataset, 410K synthetic and curated samples | Hugging Face | https://github.com/TsinghuaC3I/UltraMedical | 访问 | |
| 3D-RAD | 医学文本与大模型 | 多模态 | 视觉问答 | Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks | OpenReview | https://openreview.net/forum?id=VB2cgrlikN | 访问 | |
| MedFMC | 医学文本与大模型 | 医学影像 | 少样本分类 | 2D, 22349 Cases, 5 Tasks of Disease Classification, few-shot medical classification benchmark | GitHub | https://github.com/wllfore/MedFMC_fewshot_baseline | 访问 | |
| ICPR-HEp-2 | 医学文本与大模型 | 细胞分类 | 2D Cell Fluorescence Microscopy Imaging, 14K, 6 Categories of Cell Classification | Project Homepage | https://www.heywhale.com/mw/dataset/5ec3c6883241a100378d5d4a | 访问 | ||
| LC25000 | 医学文本与大模型 | 病理图像 | 组织分类 | 2D Pathological Imaging, 25000 Cases, 5 Categories of Pathological Image Classification | GitHub | https://github.com/tampapath/lung_colon_image_set | 访问 | |
| Chaoyang | 医学文本与大模型 | 病理图像 | 疾病分类 | 2D Pathological Imaging, 6160 Cases, 4 Categories of Colonic Lesions Classification | GitHub | https://durrlab.github.io/C3VD/ | 访问 | |
| MedICaT | 医学文本与大模型 | 多模态 | 视觉问答 | VQA, 217060 Cases QA Pair, medical image captioning and VQA | GitHub | https://www.kaggle.com/datasets/salmanmaq/m2caiseg | 访问 | |
| webMedQA | 医学文本与大模型 | 文本 | 医学问答 | QA, 63284 Cases QA Pair, web-based medical question answering | GitHub | https://github.com/hejunqing/webMedQA/tree/master | 访问 | |
| PubMedQA | 医学文本与大模型 | 文本 | 医学问答 | QA, 1000 Cases Expert annotation QA Pair, biomedical research question answering | Project Homepage | https://pubmedqa.github.io/ | 访问 | |
| MedDialog-CN | 医学文本与大模型 | 文本 | 医疗对话 | QA, 1.1M Cases QA Pair, Chinese medical dialogue dataset | GitHub | https://www.kaggle.com/datasets/salmanmaq/m2caiseg | 访问 | |
| Chinese Medical Dialogue Dataset | 医学文本与大模型 | 文本 | 医疗对话 | QA, 792K Cases QA Pair, Chinese medical dialogue | Project Homepage | https://durrlab.github.io/C3VD/ | 访问 | |
| MedMCQA | 医学文本与大模型 | 文本 | 医学问答 | QA, 193155 Cases QA Pair, medical entrance exam MCQs | Hugging Face | https://www.kaggle.com/datasets/salmanmaq/m2caiseg | 访问 | |
| PadChest | 医学文本与大模型 | X光 | 疾病分类 | 160,000 images from 67,000 patients, 174 radiographic findings, hierarchical taxonomy | GitHub | http://bimcv.cipf.es/bimcv-projects/padchest/ | 访问 | |
| MIMIC-III-Ext-Notes | 医学文本与大模型 | 文本 | 临床文本 | Extended clinical notes from MIMIC-III critical care database v1.0.0 2026 | PhysioNet | https://physionet.org/content/mimic-iii-ext-notes/1.0.0/ | 访问 | |
| BioASQ | 医学文本与大模型 | 文本 | 生物医学问答 | Biomedical question answering and semantic indexing dataset | BioASQ | http://bioasq.org/ | 访问 | |
| Medical Transcriptions | 医学文本与大模型 | 文本 | 医疗文本 | Medical transcription data for NLP tasks | Kaggle | https://www.kaggle.com/tboyle10/medicaltranscriptions | 访问 | |
| Augmented Clinical Notes Asclepius | 医学文本与大模型 | 文本 | 临床笔记生成 | 167k synthetic clinical notes with discharge summaries and comprehensive medical histories 2026 | Hugging Face | https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes | 访问 | |
| Med_Dataset | 医学文本与大模型 | 文本 | 医疗对话 | 100k real doctor-patient interactions across medical specialties with diagnoses and treatment 2026 | Hugging Face | https://huggingface.co/datasets/Med-dataset/Med_Dataset | 访问 | |
| Medical Medicine Dataset | 医学文本与大模型 | 文本 | 药物信息 | 700 medications with therapeutic uses, side effects, descriptions for medical chatbots 2026 | Hugging Face | https://huggingface.co/datasets/darkknight25/medical_medicine_dataset | 访问 | |
| MedFit Dataset | 医学文本与大模型 | 文本 | 医学问答 | 6,444 healthcare Q&A pairs for fine-tuning medical chatbot language models 2026 | Hugging Face | https://huggingface.co/datasets/mlx-community/medfit-dataset | 访问 | |
| CNTXTAI Medical Case Studies | 医学文本与大模型 | 文本 | 临床病例 | Diverse clinical cases from chronic diseases to acute conditions from academic publications 2026 | Hugging Face | https://huggingface.co/datasets/CNTXTAI0/CNTXTAI_Medical_Case_Studies | 访问 | |
| Hindi English and Punjabi Healthcare Datasets | 医学文本与大模型 | 文本 | 多语言医疗 | Multilingual healthcare datasets covering medical diagnoses, disease names in three languages 2026 | Zenodo | https://zenodo.org/records/14599295 | 访问 | |
| Multi-RADS | 医学文本与大模型 | 文本 | 报告分类 | Synthetic Radiology Report Dataset, 1600 synthetic reports covering 17 imaging findings, RADS classification 2026 | GitHub | https://arxiv.org/pdf/2601.03232 | 访问 | |
| Bones and Joints B&J Benchmark | 医学文本与大模型 | 多模态 | 临床推理 | 1245 QA pairs spanning 7 clinical competency tasks, X-ray/CT/MRI, VQA & treatment planning 2025 | Hugging Face | https://arxiv.org/pdf/2512.22275 | 访问 | |
| MediEval | 医学文本与大模型 | 文本 | 自然语言推理 | 37144 medical statements from 2015 hospital admissions, patient-contextual reasoning, NLI 2025 | GitHub | https://arxiv.org/pdf/2512.20822 | 访问 | |
| TCM-BEST4SDT | 医学文本与大模型 | 文本 | 中医辨证 | 600 questions including 300 clinical syndrome differentiation cases, TCM benchmark 2025 | GitHub | https://arxiv.org/pdf/2512.02816 | 访问 |