Advancing the co-evolution of medical foundation models and benchmarking — from pretraining objectives and multimodal alignment to evaluation protocols that diagnose robustness, generalization, and evidence consistency.
Medical foundation models are rapidly reshaping medical image understanding, enabling stronger generalization and more unified learning across modalities (X-ray, CT, MRI, ultrasound) and tasks such as classification, detection, segmentation, report generation, and multimodal reasoning. A key differentiator of this workshop is the explicit focus on the co-evolution of foundation models and benchmarking — not only building stronger models, but also establishing evaluation protocols that diagnose robustness, cross-site generalization, and evidence consistency.
Model architectures, large-scale pretraining and self-supervised learning, multimodal learning, alignment, cross-modality reasoning, and efficient fine-tuning and adaptation strategies.
Global-scale multi-institution datasets, collaboration frameworks and governance, privacy-preserving learning (federated learning), data curation, domain generalization under distribution shift, robustness, safety, and trustworthiness.
Evaluation standards, benchmark design, reproducible protocols, and clinical deployment — developing emerging benchmarks that diagnose robustness, generalization, and evidence consistency for high-stakes clinical settings.
September 8, 2026 · Room: Quality View Oresundssalen - 2.
All times are in CEST (Europe/Stockholm). An afternoon program with four invited talks, a poster session with coffee break, and five oral presentations.
| Time | Event | Type |
|---|---|---|
| 2:00 PM10 min | Opening Remarks Welcome and workshop overview from the organizers | |
| 2:10 PM30 min | Invited Talk 1 Michael Moor | Keynote |
| 2:40 PM30 min | Invited Talk 2 Lalithkumar Seenivasan | Keynote |
| 3:10 PM60 min | ☕ Poster Session & Coffee Break Interactive poster session — up to 30 posters on display | Poster |
| 4:10 PM10 min | RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching | Oral |
| 4:20 PM10 min | Consistent View Alignment Improves Vision Foundation Models for 3D Medical Image Analysis | Oral |
| 4:30 PM10 min | MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models | Oral |
| 4:40 PM10 min | Stateful Visual Encoders for Vision-Language Models | Oral |
| 4:50 PM10 min | DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance | Oral |
| 5:00 PM30 min | Invited Talk 3 Weidi Xie | Keynote |
| 5:30 PM30 min | Invited Talk 4 Xiaoxiao Li | Keynote |

Shandong University
OpenReview system email checks and notification support.