数据集导航 · 医学文本与大模型
医学文本与大模型数据集(第 10 页)
思陌医疗数据服务整理的医学文本与大模型公开数据集,共 298 个,覆盖该领域多模态数据,助力医疗AI研发选型。
全部
肿瘤 579
神经系统疾病 507
心血管系统疾病 462
耳鼻咽喉疾病 451
眼病 412
皮肤与结缔组织疾病 399
泌尿生殖系统疾病 362
内分泌系统疾病 352
医学文本与大模型 285
呼吸道疾病 245
口颌系统疾病 219
消化系统疾病 200
免疫系统疾病 137
通用医学数据集 136
血液与淋巴系统疾病 106
感染 96
创伤与损伤 66
肌肉骨骼疾病 14
化学诱发性障碍 4
环境因素所致疾病 3
共 285 个数据集 · 第 10 / 10 页
| 数据集名称 | 系统分类 | 病种 | 数据模态 | 任务类型 | 描述 | 来源 | 原始地址 | 操作 |
|---|---|---|---|---|---|---|---|---|
| Asclepius | 医学文本与大模型 | 多模态 | 视觉问答 | 3232 QA pairs, 15 medical specialties, 3 categories 8 sub-categories clinical tasks, spectrum evaluation ACL 2025 | GitHub | https://arxiv.org/abs/2402.11217 | 访问 | |
| Neurofibromatosis Type 1 Clinical Symptoms | 医学文本与大模型 | 疾病分类 | National NF1 database with 331 probands with tumors, familial and sporadic cases | UCI | https://www.archive.ics.uci.edu/datasets?Keywords=diagnosis | 访问 | ||
| TITAN | 医学文本与大模型 | 病理图像 | 病理分类 | Published in Nature Medicine, TCGA-UT-8K pathology ROI classification benchmark, foundation model for computational path | GitHub | https://github.com/mahmoodlab/TITAN | 访问 | |
| TCGA-UT-8K | 医学文本与大模型 | 病理图像 | 病理分类 | Pathology ROI classification benchmark released with TITAN, Nature Medicine 2025 | GitHub | https://github.com/mahmoodlab/TITAN | 访问 | |
| ER-REASON Clinical Reasoning Benchmark | 医学文本与大模型 | 文本 | 临床推理 | Benchmark Dataset for LLM-Based Clinical Reasoning in Emergency Department | PhysioNet | https://physionet.org/content/?topic=medication&page=5 | 访问 | |
| JMED 京东健康医疗对话数据集 | 医学文本与大模型 | 文本 | 医疗对话 | 源自京东健康互联网医院匿名医患对话,1k份高质量临床记录,覆盖0-90岁和多个专业,21个回答选项 2025 | 京东健康 | https://hyper.ai/cn/news/39375 | 访问 | |
| CheXpert Plus | 医学文本与大模型 | 多模态 | 报告生成 | 223,462 unique pairs of radiology reports and chest X-rays | Stanford | https://aimi.stanford.edu/shared-datasets | 访问 | |
| BioASQ Biomedical QA | 医学文本与大模型 | 文本 | 生物医学问答 | Biomedical question answering and semantic indexing dataset | BioASQ | http://bioasq.org/ | 访问 | |
| CUPCase Clinically Uncommon Patient Cases | 医学文本与大模型 | 文本 | 罕见病诊断 | 3,562 real-world case reports from BMC, rare diseases and uncommon presentations, diagnoses in open-ended and multiple-c | arXiv | https://arxiv.org/pdf/2503.06204v1 | 访问 | |
| MIMIC-IV-Note | 医学文本与大模型 | 文本 | 临床文本 | Deidentified clinical notes from MIMIC-IV | PhysioNet | https://physionet.org/content/mimic-iv-note/2.2/ | 访问 | |
| i2b2 n2c2 NLP Research Data Sets | 医学文本与大模型 | 文本 | 临床NLP | Several datasets of deidentified clinical notes with annotations for NLP tasks (de-identification, relation extraction) | n2c2 | https://www.i2b2.org/NLP/DataSets/ | 访问 | |
| mtsamples Medical Transcriptions | 医学文本与大模型 | 文本 | 医疗文本 | Large collection of transcribed medical sample reports | mtsamples | https://www.mtsamples.com/ | 访问 | |
| RadGraph | 医学文本与大模型 | 文本 | 放射报告知识图谱 | Radiology report entities and relations, knowledge graph from radiology reports | PhysioNet | https://physionet.org/content/radgraph/1.0.0/ | 访问 | |
| RadNLI | 医学文本与大模型 | 文本 | 自然语言推理 | Radiology report natural language inference dataset | PhysioNet | https://physionet.org/content/radnli-report-inference/1.0.0/ | 访问 | |
| RadQA | 医学文本与大模型 | 文本 | 放射报告问答 | Radiology report question answering dataset | PhysioNet | https://physionet.org/content/radqa/1.0.0/ | 访问 |