医学文本与大模型数据集(第 9 页)

思陌医疗数据服务整理的医学文本与大模型公开数据集,共 298 个,覆盖该领域多模态数据,助力医疗AI研发选型。

共 285 个数据集 · 第 9 / 10 页
数据集名称 系统分类 病种 数据模态 任务类型 描述 来源 原始地址 操作
EHRStruct医学文本与大模型关系推理2200 task-specific samples across 11 data and knowledge tasks, structured EHR reasoning 2025GitHubhttps://arxiv.org/pdf/2511.08206访问
TCM-Eval医学文本与大模型文本中医问答6099 questions from expert-level Chinese medical examinations, TCM benchmark 2025GitHubhttps://arxiv.org/pdf/2511.07148访问
RxSafeBench医学文本与大模型文本用药安全2443 consultation scenarios including 1063 contraindication cases, medication safety QA BIBM2025GitHubhttps://arxiv.org/pdf/2511.04328访问
IMB Italian Medical Benchmark医学文本与大模型文本医学问答808506 items featuring 782644 clinical Italian conversations, medical QA & MCQA CLIC-it 2025GitHubhttps://arxiv.org/pdf/2510.18468访问
ViPET-ReportGen医学文本与大模型多模态报告生成1.5 million slices paired with 2757 Vietnamese clinical reports, PET/CT report generation NeurIPS 2025GitHubhttps://arxiv.org/pdf/2509.24739访问
MedQARo医学文本与大模型文本医学问答102646 Romanian QA pairs covering 1011 clinical patients, multilingual medical QA 2025GitHubhttps://arxiv.org/pdf/2508.16390访问
AnesSuite医学文本与大模型文本麻醉学问答4427 anesthesiology MCQ items focused on complex decision-making, specialized knowledge QA 2025GitHubhttps://arxiv.org/pdf/2504.02404访问
TracSum医学文本与大模型文本医疗摘要500 abstracts resulting in 3500 summary-citation traceable pairs, aspect-based summarization EMNLP 2025GitHubhttps://arxiv.org/pdf/2508.13798访问
HEAL-MedVQA医学文本与大模型多模态视觉问答11000+ samples requiring localization prior to medical answering, grounded medical VQA IJCAI 2025Project Pagehttps://arxiv.org/pdf/2505.00744访问
BRIDGE医学文本与大模型文本多任务评估1.4 million samples covering 87 tasks in 9 languages, real-world clinical practice text 2025Project Pagehttps://arxiv.org/pdf/2504.19467访问
LLMEval-Med医学文本与大模型文本临床问答验证~1000 real-world cases validated via physician-in-the-loop audits, clinical QA validation 2025GitHubhttps://arxiv.org/pdf/2506.04078访问
CSEDB医学文本与大模型文本安全性评估30 criteria across 26 specialties based on expert physician consensus, safety-effectiveness eval 2025GitHubhttps://arxiv.org/pdf/2507.23486访问
SSG-VQA医学文本与大模型多模态手术视觉问答1300+ scene graph samples focused on instrument-tissue interaction, surgical VQA 2025GitHubhttps://arxiv.org/pdf/2506.06232访问
PET2Rep医学文本与大模型多模态报告生成565 whole-body paired PET/CT data combinations with detailed radiology reports, report generation 2025GitHubhttps://arxiv.org/pdf/2508.04062访问
MedTVT-QA医学文本与大模型多模态视觉问答3232 VQA pairs encompassing 15 medical specialties, 3 categories and 8 sub-categories clinical tasks ACL 2025GitHubhttps://arxiv.org/pdf/2506.18512访问
MediConfusion医学文本与大模型多模态视觉问答176 confusing pairs of two images sharing same question but different correct answers, VQA reliability ICLR 2025Hugging Facehttps://arxiv.org/abs/2409.15477访问
GMAI-MMBench医学文本与大模型多模态视觉问答26K QA pairs, 38 modality types, comprehensive multimodal evaluation benchmark NeurIPS 2024Hugging Facehttps://huggingface.co/datasets/OpenGVLab/GMAI-MMBench访问
PathMMU医学文本与大模型多模态病理推理33428 QAs, 24067 images, massive multimodal expert-level pathology understanding and reasoning 2024Hugging Facehttps://huggingface.co/datasets/jamessyx/PathMMU访问
medical-o1-reasoning-SFT医学文本与大模型文本医学推理19.7k QA pairs, HuatuoGPT-o1 medical complex reasoning, medical VQA & reasoning ACL 2025Hugging Facehttps://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT访问
ReasonMed医学文本与大模型文本医学推理370K multi-agent generated dataset from 1.75M CoT paths, medical reasoning, QA, chain-of-thought 2025Hugging Facehttps://arxiv.org/abs/2506.09513访问
MIRIAD医学文本与大模型文本医学问答5.82M/4.48M medical query-response pairs, augmenting LLMs with medical query-response, RAG & hallucination detection 202Hugging Facehttps://arxiv.org/abs/2506.06091访问
ClinBench-HPB医学文本与大模型文本临床问答3535 MCQs and 337 clinical cases covering 465+ hepato-pancreato-biliary diseases 2025Project Pagehttps://arxiv.org/abs/2506.00095访问
MedXpertQA医学文本与大模型多模态专家级问答4460 questions (text 2455 / image 2005), expert-level medical reasoning and understanding ICML 2025Hugging Facehttps://huggingface.co/datasets/TsinghuaC3I/MedXpertQA访问
MedCaseReasoning医学文本与大模型文本诊断推理14489 QA pairs, diagnostic reasoning from clinical case reports 2025GitHubhttps://github.com/kevinwu23/Stanford-MedCaseReasoning访问
MedS-Ins医学文本与大模型文本指令微调5M samples, 19K instructions, versatile LLMs for medicine instruction tuning 2025Hugging Facehttps://huggingface.co/datasets/Henrychur/MedS-Ins访问
MedS-Bench医学文本与大模型文本临床任务评估11 categories of clinical tasks, benchmark for medical LLMs 2025Hugging Facehttps://huggingface.co/datasets/Henrychur/MedS-Bench访问
AlphaMed19K医学文本与大模型文本医学推理19K QA pairs, medical LLM reasoning with minimalist rule-based RL 2025Hugging Facehttps://huggingface.co/datasets/che111/AlphaMed19K访问
Derm1M医学文本与大模型多模态皮肤病分类1029761 image-text pairs aligned with clinical ontology knowledge for dermatology, ICCV 2025GitHubhttps://arxiv.org/pdf/2503.14911访问
Surg-396K医学文本与大模型多模态手术问答41400 images from EndoVis/CoPESD/Cholec80, 396000 image-text pairs, 5 dialogue types, 7 surgical tasks 2025GitHubhttps://arxiv.org/abs/2501.11347访问
HuatuoGPT-o1 Dataset医学文本与大模型文本医学推理40K medically verified complex reasoning questions based on MedQA-USMLE and MedMCQA 2024GitHubhttps://arxiv.org/pdf/2412.18925访问