Instructors: Prof. Dr. Nassir Navab, Dr. Shahrooz Faghihroohi, Dr. Han Li, Dr. Azade Farshad, Yousef Yeganeh
- Registration
- Announcements
- Introduction
- Course Structure
- Schedule
- List of Topics and Materials
- List of topics
- Literature and Helpful Links
Time: TBA
Registration
- Registration must be done through TUM Matching Platform (please pay attention to the Deadlines)
- In order to increase your priority, please also apply via our own Registration system.
- The maximum number of participants: 14
Announcements
- Interested students should attend the preliminary meeting. This semester, we will have a joint presentation of the MLMI and DLMA courses offered on Wednesday, 04.02.2026, with the following agenda:
Machine Learning in Medical Imaging (MLMI): 14:00 hrs. - 14:30 hrs.
Deep Learning for Medical Applications (DLMA): 14:30 hrs. - 15:00 hrs.
The sessions will be conducted by the following Zoom link:
https://tum-conf.zoom-x.de/j/64943024190?pwd=S82DItjso6Zsc2mYhk31489cYqrcRY.1 - The DLMA Introduction slides can be found here: DLMA-PreliminaryMeeting-SS26.pdf
Introduction
- Deep Learning is growing tremendously in Computer Vision and Medical Imaging as well. Highly impacted journals in the medical imaging community, i.e. IEEE Transaction on Medical Imaging, recently published their special edition on Deep Learning [1]. The Seminar will propose a list of recent scientific articles related to the main current research topics in deep learning for Medical Applications, together with some interesting papers from other communities (CVPR, NeurIPS, ICCV, ICLR, ICML, ...).
Course Structure
In this Master Seminar (Hauptseminar), students select one scientific topic from the list provided by course organizers. The students should read the proposed sample papers by the tutors, find the topic-related articles, summarize and compare them in their presentation and blogpost:
- Presentation: The selected paper is presented to the other participants (Maximum 25 minutes presentation, 10 minutes questions). You can use the CAMP templates for PowerPoint TUM-Template.pptx.
- Blog Post: A blog post of 3000-3500 words excluding references, should be submitted before the deadline. The blog post must include all references used and must be written completely in your own words. Copy and paste will not be tolerated.
- Attendance: Participants have to participate actively in all seminar sessions. Each presentation is followed by a discussion, and everyone is encouraged to actively participate.
Submission Deadline: You have to submit the blog post until June 8th and can modify it a bit until the last session of the course.
Schedule
| 11.06.2026 | Ultrasound-related Foundation Model VLMs for ultrasound Uncertainty estimation in ultrasound imaging | Yicheng Xue Kahle, Lena Ekici, Can Göksel |
| 18.06.2026 | Reasoning on Intraoperative and Preoperative Data for Abdominal Surgery Gaussian Splatting in minimally invasive surgery World model and its transition into Medical applications | Pötzsch, Moritz Schwan, Nick Marlon Müller |
| 25.06.2026 | Knowledge Graph and GNN for Cardiovascular Diseases Digital Twins in Cardiovascular Diseases Learning-based Statistical Shape Model | Hirani, Dhruvil Romdhani Hedi Feijó Lobo de Andrade, Lucas |
| 02.07.2026 | Unpaired Cross-Modality Translation and Optimal Transport Poisson Flow Generative Models GFlowNets: Flowing Through Latent Space | Julian Münzer Zehetmair Tobias Bhatt, Harshul |
| 09.07.2026 | VLMs for pathology Pathology Agent | Konuk, Doğa Elif Werner, Denis |
| 16.07.2026 | OCT-based retina representation, but where to put it? Video Extrapolation and Future State Prediction | Scheipl, Julian Talel Oueslati |
List of Topics and Materials
The proposed papers for each topic in this course are usually selected from the following venues/publications:
CVPR: Conference on Computer Vision and Pattern Recognition
ICLR: International Conference on Learning Representations
NeurIPS: Neural Information Processing Systems
TPAMI: IEEE Transactions on Pattern Analysis and Machine Intelligence
TMI: IEEE Transaction on Medical Imaging
JBHI: IEEE Journal of Biomedical and Health Informatics
MedIA: Medical Image Analysis (Elsevier)
MICCAI: Medical Image Computing and Computer-Assisted Intervention.
BMVC: British Machine Vision Conference
MIDL: Medical Imaging with Deep Learning
List of topics
| No | Topic | Sample Papers | Journal/ Conference | Tutor | Student | Link |
|---|---|---|---|---|---|---|
| 1 | Learning-based Statistical Shape Model | SLoRD: structural low-rank descriptors for shape consistency in vertebrae segmentation | JBHI 2025 | Feijó Lobo de Andrade, Lucas | https://ieeexplore.ieee.org/abstract/document/11016174 | |
| Mesh2SSM++: A Probabilistic Framework for Unsupervised Learning of Statistical Shape Model of Anatomies from Surface Meshes | Arxiv 2025 | https://arxiv.org/abs/2502.07145 | ||||
| DeepSSM: A blueprint for image-to-shape deep learning models | MedIA 2024 | https://www.sciencedirect.com/science/article/abs/pii/S1361841523002943 | ||||
| 2 | Reasoning on Intraoperative and Preoperative Data for Abdominal Surgery | SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning | Arxiv March 2026 | Pötzsch, Moritz | https://arxiv.org/pdf/2603.06570 | |
| MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic Workflow | ICLR 2026 | https://arxiv.org/pdf/2503.18968 | ||||
| 3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis | Arxiv March 2026 | https://arxiv.org/pdf/2602.18064 | ||||
| 3 | Knowledge Graph and GNNs for Cardiovascular Diseases | A multimodal vision knowledge graph of cardiovascular disease | Nature Cardiovascular Research 2025 | Shahrooz Faghihroohi | Hirani, Dhruvil | https://www.nature.com/articles/s44161-025-00757-4.pdf |
| Knowledge graph construction for heart failure using large language models with prompt engineering | Frontiers in Computational Neuroscience 2024 | |||||
| CVD Atlas: a multi-omics database of cardiovascular disease | Nucleic acids research 2025 | https://academic.oup.com/nar/article/53/D1/D1348/7798784 | ||||
| 4 | Video Extrapolation and Future State Prediction | RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers | ICML 2025 | Magdalena Wysocki | Talel Oueslati | https://icml.cc/virtual/2025/poster/43698 |
| UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers | ICLR 2026 | https://openreview.net/forum?id=fLLCmC53u9 | ||||
| Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think | CVPR 2025 | https://cvpr.thecvf.com/virtual/2025/poster/34396 | ||||
| 5 | Unpaired Cross-Modality Translation and Optimal Transport | Unpaired Image-to-Image Translation via Neural Schrödinger Bridge | ICLR 2024 | Xuesong Li | Julian Münzer | https://arxiv.org/abs/2305.15086 |
| Mutual information guided diffusion for zero-shot cross-modality medical image translation | TMI 2024 | https://ieeexplore.ieee.org/document/10485553 | ||||
| LBM: Latent Bridge Matching for Fast Image-to-Image Translation | ICCV 2025 | |||||
| 6 | Ultrasound-related Foundation Model | Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding (VLM) | CVPR2026 | Yue Zhou | Yicheng Xue | https://arxiv.org/abs/2604.01749 |
| EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance (World Model) | CVPR2025 | https://arxiv.org/abs/2504.13065 | ||||
| ANATOMY-AWARE REPRESENTATION LEARNING FOR MEDICAL ULTRASOUND (Vision Foundation Model) | ICLR2026 | https://openreview.net/pdf?id=5ThIWuDkEf | ||||
| 7 | OCT-based retina representation, but where to put it? | Gaussian Primitive Optimized Deformable Retinal Image Registration | MICCAI 2025 | Hannes Firzlaff | Scheipl, Julian | https://papers.miccai.org/miccai-2025/paper/3875_paper.pdf |
| Retinal OCT Image Registration: Methods and Applications | IEEE RBME Vol. 16 2021 | https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=9531445 | ||||
| RetinaRegNet: A zero-shot approach for retinal image registration | Comput. Biol. Med 2025 | https://www.sciencedirect.com/science/article/pii/S001048252401730X | ||||
| 8 | Gaussian Splatting in minimally invasive surgery | T2GS: Comprehensive Reconstruction of Dynamic Surgical Scenes with Gaussian Splatting | MICCAI 2025 | Hannes Firzlaff | Schwan, Nick | https://papers.miccai.org/miccai-2025/paper/5019_paper.pdf |
| SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting | MICCAI 2025 | https://papers.miccai.org/miccai-2025/paper/1324_paper.pdf | ||||
| EndoSparse: Real-Time Sparse View Synthesis of Endoscopic Scenes using Gaussian Splatting | MICCAI 2024 | https://papers.miccai.org/miccai-2024/paper/0791_paper.pdf | ||||
| 9 | World model and its transition into Medical applications | V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning | Arxiv 2025 | Yuan Bi | Marlon Müller | https://arxiv.org/abs/2506.09985 |
| How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment | Arxiv 2025 | https://arxiv.org/abs/2511.01775 | ||||
| EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance | CVPR 2025 | |||||
| 10 | Pathology Agent | Evidence-based diagnostic reasoning with multi-agent copilot for human pathology | Arxiv 2025 | Han Li | Werner, Denis | https://arxiv.org/abs/2506.20964 |
| LAMMI-Pathology: A Tool-Centric Bottom-Up LVLM-Agent Framework for Molecularly Informed Medical Intelligence in Pathology | Arxiv 2026 | https://arxiv.org/abs/2602.18773 | ||||
| Pathology-CoT: Learning Visual Chain-of-Thought Agent from Expert Whole Slide Image Diagnosis Behavior | Arxiv 2025 | https://arxiv.org/abs/2510.04587 | ||||
| 11 | VLMs for ultrasound | A visually grounded language model for fetal ultrasound understanding | nature Biomedical Engineering 2026 | Mohammad Farid Azampour | Kahle, Lena | https://www.nature.com/articles/s41551-025-01578-3 |
| LLAUS: A High-Quality Instruction-Tuned Large Vision Language Assistant for UltraSound | ICMR 2025 | https://dl.acm.org/doi/10.1145/3731715.3733374 | ||||
| Comprehensive echocardiogram evaluation with view primed vision language AI | nature 2025 | https://www.nature.com/articles/s41586-025-09850-x | ||||
| 12 | Poisson Flow Generative Models | Poisson Flow Generative Models | NeurIPS 2022 | Tony Danjun Wang | Zehetmair Tobias | https://arxiv.org/pdf/2209.11178 |
| PFGM++: Unlocking the Potential of Physics-Inspired Generative Models | ICML 2023 | https://arxiv.org/pdf/2302.04265 | ||||
| PFCM: Poisson Flow Consistency Models for Low-Dose CT Image Denoising | TMI 2025 | https://arxiv.org/pdf/2402.08159 | ||||
| 13 | Perioperative Multimodal Reasoning for Cardiovascular Surgery: Integrating Preoperative Planning and Intraoperative Evidence with Deep Learning | Deep learning for transesophageal echocardiography view classification | Scientific Reports 2024 | Shahrooz Faghihroohi | Bhatt, Harshul | https://pmc.ncbi.nlm.nih.gov/articles/PMC10761863/ |
| A novel generative multi-task representation learning approach for predicting postoperative complications in cardiac surgery patients | Jamia 2025 | https://pmc.ncbi.nlm.nih.gov/articles/PMC11833467/ | ||||
| Surgical Action Planning with Large Language Models | MICCAI 2025 | https://arxiv.org/abs/2503.18296 | ||||
| 14 | Digital Twins in Cardiovascular Diseases | Design and analysis of TwinCardioframework to detect and monitorcardiovascular diseases usingdigital twin and deep neuralnetwork | Scientific Report 2025 | Shahrooz Faghihroohi | Romdhani Hedi | https://www.nature.com/articles/s41598-025-08824-3.pdf |
| Digital twins in cardiovascular disease: a scoping review | Journal of Medical Informatics 2026 | https://www.sciencedirect.com/science/article/pii/S1386505625003557 | ||||
| Digital Cardiovascular Twins, AI Agents, and Sensor Data: A Narrative Review from System Architecture to Proactive Heart Health | Sensors 2025 | https://www.mdpi.com/1424-8220/25/17/5272 | ||||
| 15 | Uncertainty estimation in ultrasound imaging | Enhanced Uncertainty Estimation in Ultrasound Image Segmentation with MSU-Net. | MICCAI 2024 | Miruna-Alexandra Gafencu | Ekici, Can Göksel | https://arxiv.org/abs/2407.21273 |
| Diffusion-based Iterative Counterfactual Explanations for Fetal Ultrasound Image Quality Assessment | MICCAI 2025 | https://arxiv.org/abs/2403.08700 | ||||
| Shadow-Consistent Semi-Supervised Learning for Prostate Ultrasound Segmentation. | TMI 2022 | https://ieeexplore.ieee.org/document/9667363 | ||||
| 16 | VLMs for pathology | Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner | AAAI 2026 | Konuk, Doğa Elif | https://arxiv.org/abs/2505.11404 | |
| Pathologyvlm: a large vision-language model for pathology image understanding | Artificial Intelligence Review 2025 | https://link.springer.com/article/10.1007/s10462-025-11190-1 | ||||
| A visual-language foundation model for computational pathology | Nature Medicine 2024 | https://www.nature.com/articles/s41591-024-02856-4 |
Literature and Helpful Links
A lot of scientific publications can be found online.
The following list may help you to find some further information on your particular topic:
- Microsoft Academic Search
- Google Scholar
- CiteSeer
- CiteULike
- Collection of Computer Science Bibliographies
Some publishers:
- ScienceDirect (Elsevier Journals)
- IEEE Journals
- ACM Digital Library
Libraries (online and offline):
- http://rzblx1.uni-regensburg.de/ezeit/ (Elektronische Zeitschriften Bibliothek)
- Verbundkatalog des Bibliotheksverbundes Bayern (BVB)
- Computer ORG
- http://www.ub.tum.de/ (TUM Library)
- To get access onto the electronic library, see http://www.ub.tum.de/medien/ejournals/readme.html
- "proxy.biblio.tu-muenchen.de" mit Port 8080 (nur fuer http). Damit klappen zumindest portal.acm.org und computer.org meistens
- Various proceedings of conferences in our AR-Lab, 03.13.036 (These proceedings are not for lending!)
Some further hints for working with references:
- JabRef is a Java program for comfortable working with Bibtex literature databases. Handy feature: if you know the PubMed ID for an article, JabRef can import data from there (via "Web Search/Medline").
- Mendeley is a cross-platform program for organising your references.
If you find useful resources that are not already listed here, please tell us, so we can add them for others. Thanks.

Kommentar
Marlon Müller sagt:
29. Mai 2026Where should I submit my blog post? The pages do not exist yet.