Instructors: Prof. Dr. Nassir NavabDr. Shahrooz Faghihroohi, Dr. Han Li, Dr. Azade Farshad, Yousef Yeganeh





Time: TBA

Registration

Announcements

  • Interested students should attend the preliminary meeting. This semester, we will have a joint presentation of the MLMI and DLMA courses offered on Wednesday, 04.02.2026, with the following agenda:

    Machine Learning in Medical Imaging (MLMI): 14:00 hrs. - 14:30 hrs.
    Deep Learning for Medical Applications (DLMA): 14:30 hrs. - 15:00 hrs.

    The sessions will be conducted by the following Zoom link:
    https://tum-conf.zoom-x.de/j/64943024190?pwd=S82DItjso6Zsc2mYhk31489cYqrcRY.1
  • The DLMA Introduction slides can be found here: DLMA-PreliminaryMeeting-SS26.pdf

Introduction

  • Deep Learning is growing tremendously in Computer Vision and Medical Imaging as well. Highly impacted journals in the medical imaging community, i.e. IEEE Transaction on Medical Imaging, recently published their special edition on Deep Learning [1]. The Seminar will propose a list of recent scientific articles related to the main current research topics in deep learning for Medical Applications, together with some interesting papers from other communities (CVPR, NeurIPS, ICCV, ICLR, ICML, ...).

Course Structure

In this Master Seminar (Hauptseminar), students select one scientific topic from the list provided by course organizers. The students should read the proposed sample papers by the tutors, find the topic-related articles, summarize and compare them in their presentation and blogpost:

  • Presentation: The selected paper is presented to the other participants (Maximum 25 minutes presentation, 10 minutes questions). You can use the CAMP templates for PowerPoint TUM-Template.pptx.
  • Blog Post: A blog post of 3000-3500 words excluding references, should be submitted before the deadline. The blog post must include all references used and must be written completely in your own words. Copy and paste will not be tolerated.
  • Attendance: Participants have to participate actively in all seminar sessions. Each presentation is followed by a discussion, and everyone is encouraged to actively participate.

Submission Deadline: You have to submit the blog post until June 8th and can modify it a bit until the last session of the course.

Schedule

11.06.2026

Ultrasound-related Foundation Model

VLMs for ultrasound

Uncertainty estimation in ultrasound imaging

Yicheng Xue

Kahle, Lena

Ekici, Can Göksel

18.06.2026

Reasoning on Intraoperative and Preoperative Data for Abdominal Surgery

Gaussian Splatting in minimally invasive surgery

World model and its transition into Medical applications

Pötzsch, Moritz

Schwan, Nick

Marlon Müller

25.06.2026

Knowledge Graph and GNN for Cardiovascular Diseases

Digital Twins in Cardiovascular Diseases

Learning-based Statistical Shape Model

Hirani, Dhruvil

Romdhani Hedi

Feijó Lobo de Andrade, Lucas

02.07.2026

Unpaired Cross-Modality Translation and Optimal Transport

Poisson Flow Generative Models

GFlowNets: Flowing Through Latent Space

Julian Münzer

Zehetmair Tobias

Bhatt, Harshul

09.07.2026

VLMs for pathology

Pathology Agent

Konuk, Doğa Elif

Werner, Denis

16.07.2026

OCT-based retina representation, but where to put it?

Video Extrapolation and Future State Prediction

Scheipl, Julian

Talel Oueslati


List of Topics and Materials

The proposed papers for each topic in this course are usually selected from the following venues/publications:


CVPR: Conference on Computer Vision and Pattern Recognition
ICLR: International Conference on Learning Representations
NeurIPS: Neural Information Processing Systems

TPAMI: IEEE Transactions on Pattern Analysis and Machine Intelligence

TMI: IEEE Transaction on Medical Imaging
JBHI: IEEE Journal of Biomedical and Health Informatics
MedIA: Medical Image Analysis (Elsevier)

MICCAI: Medical Image Computing and Computer-Assisted Intervention.
BMVC: British Machine Vision Conference
MIDL: Medical Imaging with Deep Learning


List of topics

NoTopicSample PapersJournal/ ConferenceTutorStudentLink
1 Learning-based Statistical Shape ModelSLoRD: structural low-rank descriptors for shape consistency in vertebrae segmentationJBHI 2025Feijó Lobo de Andrade, Lucashttps://ieeexplore.ieee.org/abstract/document/11016174
Mesh2SSM++: A Probabilistic Framework for Unsupervised Learning of Statistical Shape Model of Anatomies from Surface MeshesArxiv 2025https://arxiv.org/abs/2502.07145
DeepSSM: A blueprint for image-to-shape deep learning modelsMedIA 2024https://www.sciencedirect.com/science/article/abs/pii/S1361841523002943
2

Reasoning on Intraoperative and Preoperative Data for Abdominal Surgery


SUREON: A Benchmark and Vision-Language-Model for Surgical ReasoningArxiv March 2026Pötzsch, Moritzhttps://arxiv.org/pdf/2603.06570
MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic WorkflowICLR 2026https://arxiv.org/pdf/2503.18968
3DMedAgent: Unified Perception-to-Understanding for 3D Medical AnalysisArxiv March 2026https://arxiv.org/pdf/2602.18064
3Knowledge Graph and GNNs for Cardiovascular DiseasesA multimodal vision knowledge graph of cardiovascular diseaseNature Cardiovascular Research 2025Shahrooz Faghihroohi Hirani, Dhruvilhttps://www.nature.com/articles/s44161-025-00757-4.pdf
Knowledge graph construction for heart failure using large language models with prompt engineeringFrontiers in Computational Neuroscience 2024

https://www.frontiersin.org/journals/computational-neuroscience/articles/10.3389/fncom.2024.1389475/full

CVD Atlas: a multi-omics database of cardiovascular diseaseNucleic acids research 2025https://academic.oup.com/nar/article/53/D1/D1348/7798784
4

Video Extrapolation and Future State Prediction


RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion TransformersICML 2025Magdalena WysockiTalel Oueslatihttps://icml.cc/virtual/2025/poster/43698
UltraViCo: Breaking Extrapolation Limits in Video Diffusion TransformersICLR 2026https://openreview.net/forum?id=fLLCmC53u9
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You ThinkCVPR 2025https://cvpr.thecvf.com/virtual/2025/poster/34396
5

Unpaired Cross-Modality Translation and Optimal Transport


Unpaired Image-to-Image Translation via Neural Schrödinger BridgeICLR 2024Xuesong Li Julian Münzerhttps://arxiv.org/abs/2305.15086
Mutual information guided diffusion for zero-shot cross-modality medical image translationTMI 2024https://ieeexplore.ieee.org/document/10485553
LBM: Latent Bridge Matching for Fast Image-to-Image TranslationICCV 2025

https://openaccess.thecvf.com/content/ICCV2025/html/Chadebec_LBM_Latent_Bridge_Matching_for_Fast_Image-to-Image_Translation_ICCV_2025_paper.html

6Ultrasound-related Foundation ModelUltrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding (VLM)CVPR2026Yue Zhou Yicheng Xuehttps://arxiv.org/abs/2604.01749
EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance (World Model)CVPR2025https://arxiv.org/abs/2504.13065
ANATOMY-AWARE REPRESENTATION LEARNING FOR MEDICAL ULTRASOUND (Vision Foundation Model)ICLR2026https://openreview.net/pdf?id=5ThIWuDkEf
7

OCT-based retina representation, but where to put it?


Gaussian Primitive Optimized Deformable Retinal Image RegistrationMICCAI 2025Hannes Firzlaff Scheipl, Julianhttps://papers.miccai.org/miccai-2025/paper/3875_paper.pdf
Retinal OCT Image Registration: Methods and ApplicationsIEEE RBME Vol. 16 2021https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9531445
RetinaRegNet: A zero-shot approach for retinal image registrationComput. Biol. Med 2025https://www.sciencedirect.com/science/article/pii/S001048252401730X
8Gaussian Splatting in minimally invasive surgeryT2GS: Comprehensive Reconstruction of Dynamic Surgical Scenes with Gaussian SplattingMICCAI 2025Hannes Firzlaff Schwan, Nickhttps://papers.miccai.org/miccai-2025/paper/5019_paper.pdf
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian SplattingMICCAI 2025https://papers.miccai.org/miccai-2025/paper/1324_paper.pdf
EndoSparse: Real-Time Sparse View Synthesis of Endoscopic Scenes using Gaussian SplattingMICCAI 2024https://papers.miccai.org/miccai-2024/paper/0791_paper.pdf
9

World model and its transition into Medical applications



V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningArxiv 2025Yuan BiMarlon Müllerhttps://arxiv.org/abs/2506.09985
How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert AssessmentArxiv 2025https://arxiv.org/abs/2511.01775
EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe GuidanceCVPR 2025

https://openaccess.thecvf.com/content/CVPR2025/papers/Yue_EchoWorld_Learning_Motion-Aware_World_Models_for_Echocardiography_Probe_Guidance_CVPR_2025_paper.pdf

10Pathology AgentEvidence-based diagnostic reasoning with multi-agent copilot for human pathologyArxiv 2025Han LiWerner, Denishttps://arxiv.org/abs/2506.20964
LAMMI-Pathology: A Tool-Centric Bottom-Up LVLM-Agent Framework for Molecularly Informed Medical Intelligence in PathologyArxiv 2026https://arxiv.org/abs/2602.18773
Pathology-CoT: Learning Visual Chain-of-Thought Agent from Expert Whole Slide Image Diagnosis BehaviorArxiv 2025https://arxiv.org/abs/2510.04587
11

VLMs for ultrasound


A visually grounded language model for fetal ultrasound understandingnature Biomedical Engineering 2026Mohammad Farid Azampour Kahle, Lenahttps://www.nature.com/articles/s41551-025-01578-3
LLAUS: A High-Quality Instruction-Tuned Large Vision Language Assistant for UltraSoundICMR 2025https://dl.acm.org/doi/10.1145/3731715.3733374
Comprehensive echocardiogram evaluation with view primed vision language AInature 2025https://www.nature.com/articles/s41586-025-09850-x
12Poisson Flow Generative ModelsPoisson Flow Generative ModelsNeurIPS 2022Tony Danjun Wang Zehetmair Tobiashttps://arxiv.org/pdf/2209.11178
PFGM++: Unlocking the Potential of Physics-Inspired Generative ModelsICML 2023https://arxiv.org/pdf/2302.04265
PFCM: Poisson Flow Consistency Models for Low-Dose CT Image DenoisingTMI 2025https://arxiv.org/pdf/2402.08159
13

Perioperative Multimodal Reasoning for Cardiovascular Surgery: Integrating Preoperative Planning and Intraoperative Evidence with Deep Learning


Deep learning for transesophageal echocardiography view classification

Scientific Reports 2024

Shahrooz Faghihroohi Bhatt, Harshulhttps://pmc.ncbi.nlm.nih.gov/articles/PMC10761863/
A novel generative multi-task representation learning approach for predicting postoperative complications in cardiac surgery patientsJamia 2025https://pmc.ncbi.nlm.nih.gov/articles/PMC11833467/
Surgical Action Planning with Large Language ModelsMICCAI 2025https://arxiv.org/abs/2503.18296
14Digital Twins in Cardiovascular DiseasesDesign and analysis of TwinCardioframework to detect and monitorcardiovascular diseases usingdigital twin and deep neuralnetworkScientific Report 2025Shahrooz Faghihroohi Romdhani Hedihttps://www.nature.com/articles/s41598-025-08824-3.pdf
Digital twins in cardiovascular disease: a scoping reviewJournal of Medical Informatics 2026https://www.sciencedirect.com/science/article/pii/S1386505625003557
Digital Cardiovascular Twins, AI Agents, and Sensor Data: A Narrative Review from System Architecture to Proactive Heart HealthSensors 2025https://www.mdpi.com/1424-8220/25/17/5272
15Uncertainty estimation in ultrasound imaging Enhanced Uncertainty Estimation in Ultrasound Image Segmentation with MSU-Net.MICCAI 2024Miruna-Alexandra Gafencu Ekici, Can Gökselhttps://arxiv.org/abs/2407.21273
Diffusion-based Iterative Counterfactual Explanations for Fetal Ultrasound Image Quality AssessmentMICCAI 2025https://arxiv.org/abs/2403.08700
Shadow-Consistent Semi-Supervised Learning for Prostate Ultrasound Segmentation.TMI 2022https://ieeexplore.ieee.org/document/9667363
16VLMs for pathologyPatho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert ReasonerAAAI 2026Konuk, Doğa Elif

https://arxiv.org/abs/2505.11404


Pathologyvlm: a large vision-language model for pathology image understandingArtificial Intelligence Review 2025https://link.springer.com/article/10.1007/s10462-025-11190-1
A visual-language foundation model for computational pathologyNature Medicine 2024https://www.nature.com/articles/s41591-024-02856-4

Literature and Helpful Links

A lot of scientific publications can be found online.

The following list may help you to find some further information on your particular topic:

Some publishers:

Libraries (online and offline):

Some further hints for working with references:

  • JabRef is a Java program for comfortable working with Bibtex literature databases. Handy feature: if you know the PubMed ID for an article, JabRef can import data from there (via "Web Search/Medline").
  • Mendeley is a cross-platform program for organising your references.

If you find useful resources that are not already listed here, please tell us, so we can add them for others. Thanks.

  • Keine Stichwörter

Kommentar

  1. Marlon Müller sagt:

    Where should I submit my blog post? The pages do not exist yet.