数据集导航 · 医学文本与大模型
医学文本与大模型数据集(第 9 页)
思陌医疗数据服务整理的医学文本与大模型公开数据集,共 298 个,覆盖该领域多模态数据,助力医疗AI研发选型。
全部
肿瘤 579
神经系统疾病 507
心血管系统疾病 462
耳鼻咽喉疾病 451
眼病 412
皮肤与结缔组织疾病 399
泌尿生殖系统疾病 362
内分泌系统疾病 352
医学文本与大模型 285
呼吸道疾病 245
口颌系统疾病 219
消化系统疾病 200
免疫系统疾病 137
通用医学数据集 136
血液与淋巴系统疾病 106
感染 96
创伤与损伤 66
肌肉骨骼疾病 14
化学诱发性障碍 4
环境因素所致疾病 3
共 285 个数据集 · 第 9 / 10 页
| 数据集名称 | 系统分类 | 病种 | 数据模态 | 任务类型 | 描述 | 来源 | 原始地址 | 操作 |
|---|---|---|---|---|---|---|---|---|
| EHRStruct | 医学文本与大模型 | 关系推理 | 2200 task-specific samples across 11 data and knowledge tasks, structured EHR reasoning 2025 | GitHub | https://arxiv.org/pdf/2511.08206 | 访问 | ||
| TCM-Eval | 医学文本与大模型 | 文本 | 中医问答 | 6099 questions from expert-level Chinese medical examinations, TCM benchmark 2025 | GitHub | https://arxiv.org/pdf/2511.07148 | 访问 | |
| RxSafeBench | 医学文本与大模型 | 文本 | 用药安全 | 2443 consultation scenarios including 1063 contraindication cases, medication safety QA BIBM2025 | GitHub | https://arxiv.org/pdf/2511.04328 | 访问 | |
| IMB Italian Medical Benchmark | 医学文本与大模型 | 文本 | 医学问答 | 808506 items featuring 782644 clinical Italian conversations, medical QA & MCQA CLIC-it 2025 | GitHub | https://arxiv.org/pdf/2510.18468 | 访问 | |
| ViPET-ReportGen | 医学文本与大模型 | 多模态 | 报告生成 | 1.5 million slices paired with 2757 Vietnamese clinical reports, PET/CT report generation NeurIPS 2025 | GitHub | https://arxiv.org/pdf/2509.24739 | 访问 | |
| MedQARo | 医学文本与大模型 | 文本 | 医学问答 | 102646 Romanian QA pairs covering 1011 clinical patients, multilingual medical QA 2025 | GitHub | https://arxiv.org/pdf/2508.16390 | 访问 | |
| AnesSuite | 医学文本与大模型 | 文本 | 麻醉学问答 | 4427 anesthesiology MCQ items focused on complex decision-making, specialized knowledge QA 2025 | GitHub | https://arxiv.org/pdf/2504.02404 | 访问 | |
| TracSum | 医学文本与大模型 | 文本 | 医疗摘要 | 500 abstracts resulting in 3500 summary-citation traceable pairs, aspect-based summarization EMNLP 2025 | GitHub | https://arxiv.org/pdf/2508.13798 | 访问 | |
| HEAL-MedVQA | 医学文本与大模型 | 多模态 | 视觉问答 | 11000+ samples requiring localization prior to medical answering, grounded medical VQA IJCAI 2025 | Project Page | https://arxiv.org/pdf/2505.00744 | 访问 | |
| BRIDGE | 医学文本与大模型 | 文本 | 多任务评估 | 1.4 million samples covering 87 tasks in 9 languages, real-world clinical practice text 2025 | Project Page | https://arxiv.org/pdf/2504.19467 | 访问 | |
| LLMEval-Med | 医学文本与大模型 | 文本 | 临床问答验证 | ~1000 real-world cases validated via physician-in-the-loop audits, clinical QA validation 2025 | GitHub | https://arxiv.org/pdf/2506.04078 | 访问 | |
| CSEDB | 医学文本与大模型 | 文本 | 安全性评估 | 30 criteria across 26 specialties based on expert physician consensus, safety-effectiveness eval 2025 | GitHub | https://arxiv.org/pdf/2507.23486 | 访问 | |
| SSG-VQA | 医学文本与大模型 | 多模态 | 手术视觉问答 | 1300+ scene graph samples focused on instrument-tissue interaction, surgical VQA 2025 | GitHub | https://arxiv.org/pdf/2506.06232 | 访问 | |
| PET2Rep | 医学文本与大模型 | 多模态 | 报告生成 | 565 whole-body paired PET/CT data combinations with detailed radiology reports, report generation 2025 | GitHub | https://arxiv.org/pdf/2508.04062 | 访问 | |
| MedTVT-QA | 医学文本与大模型 | 多模态 | 视觉问答 | 3232 VQA pairs encompassing 15 medical specialties, 3 categories and 8 sub-categories clinical tasks ACL 2025 | GitHub | https://arxiv.org/pdf/2506.18512 | 访问 | |
| MediConfusion | 医学文本与大模型 | 多模态 | 视觉问答 | 176 confusing pairs of two images sharing same question but different correct answers, VQA reliability ICLR 2025 | Hugging Face | https://arxiv.org/abs/2409.15477 | 访问 | |
| GMAI-MMBench | 医学文本与大模型 | 多模态 | 视觉问答 | 26K QA pairs, 38 modality types, comprehensive multimodal evaluation benchmark NeurIPS 2024 | Hugging Face | https://huggingface.co/datasets/OpenGVLab/GMAI-MMBench | 访问 | |
| PathMMU | 医学文本与大模型 | 多模态 | 病理推理 | 33428 QAs, 24067 images, massive multimodal expert-level pathology understanding and reasoning 2024 | Hugging Face | https://huggingface.co/datasets/jamessyx/PathMMU | 访问 | |
| medical-o1-reasoning-SFT | 医学文本与大模型 | 文本 | 医学推理 | 19.7k QA pairs, HuatuoGPT-o1 medical complex reasoning, medical VQA & reasoning ACL 2025 | Hugging Face | https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT | 访问 | |
| ReasonMed | 医学文本与大模型 | 文本 | 医学推理 | 370K multi-agent generated dataset from 1.75M CoT paths, medical reasoning, QA, chain-of-thought 2025 | Hugging Face | https://arxiv.org/abs/2506.09513 | 访问 | |
| MIRIAD | 医学文本与大模型 | 文本 | 医学问答 | 5.82M/4.48M medical query-response pairs, augmenting LLMs with medical query-response, RAG & hallucination detection 202 | Hugging Face | https://arxiv.org/abs/2506.06091 | 访问 | |
| ClinBench-HPB | 医学文本与大模型 | 文本 | 临床问答 | 3535 MCQs and 337 clinical cases covering 465+ hepato-pancreato-biliary diseases 2025 | Project Page | https://arxiv.org/abs/2506.00095 | 访问 | |
| MedXpertQA | 医学文本与大模型 | 多模态 | 专家级问答 | 4460 questions (text 2455 / image 2005), expert-level medical reasoning and understanding ICML 2025 | Hugging Face | https://huggingface.co/datasets/TsinghuaC3I/MedXpertQA | 访问 | |
| MedCaseReasoning | 医学文本与大模型 | 文本 | 诊断推理 | 14489 QA pairs, diagnostic reasoning from clinical case reports 2025 | GitHub | https://github.com/kevinwu23/Stanford-MedCaseReasoning | 访问 | |
| MedS-Ins | 医学文本与大模型 | 文本 | 指令微调 | 5M samples, 19K instructions, versatile LLMs for medicine instruction tuning 2025 | Hugging Face | https://huggingface.co/datasets/Henrychur/MedS-Ins | 访问 | |
| MedS-Bench | 医学文本与大模型 | 文本 | 临床任务评估 | 11 categories of clinical tasks, benchmark for medical LLMs 2025 | Hugging Face | https://huggingface.co/datasets/Henrychur/MedS-Bench | 访问 | |
| AlphaMed19K | 医学文本与大模型 | 文本 | 医学推理 | 19K QA pairs, medical LLM reasoning with minimalist rule-based RL 2025 | Hugging Face | https://huggingface.co/datasets/che111/AlphaMed19K | 访问 | |
| Derm1M | 医学文本与大模型 | 多模态 | 皮肤病分类 | 1029761 image-text pairs aligned with clinical ontology knowledge for dermatology, ICCV 2025 | GitHub | https://arxiv.org/pdf/2503.14911 | 访问 | |
| Surg-396K | 医学文本与大模型 | 多模态 | 手术问答 | 41400 images from EndoVis/CoPESD/Cholec80, 396000 image-text pairs, 5 dialogue types, 7 surgical tasks 2025 | GitHub | https://arxiv.org/abs/2501.11347 | 访问 | |
| HuatuoGPT-o1 Dataset | 医学文本与大模型 | 文本 | 医学推理 | 40K medically verified complex reasoning questions based on MedQA-USMLE and MedMCQA 2024 | GitHub | https://arxiv.org/pdf/2412.18925 | 访问 |