Michael Moor

dblp:222/4286 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-4911-6437ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Vision and language · 22% Probabilistic and Bayesian machine learning · 16% Graph learning · 14%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Medical and health informatics · 81% Bioinformatics and computational biology · 19%
Theoretical computer science
2 papers
Computational geometry · 100%

Topics — the 25 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational geometry › topological data analysis
persistent homology
1.022022
Topological Graph Neural Networks · ICLR 2022
Topological Autoencoders · ICML 2020
Computational geometry
topological data analysis
1.022022
Topological Graph Neural Networks · ICLR 2022
Topological Autoencoders · ICML 2020
Natural language and speech › Question answering and dialogue systems
medical reasoning
0.912025
Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards · EMNLP 2025
Computer vision › Vision and language
multimodal in-context learning
0.912025
SMMILE: An expert-driven benchmark for multimodal medical in-context learning · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
SMMILE: An expert-driven benchmark for multimodal medical in-context learning · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.912025
Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards · EMNLP 2025
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.712023
Zero-shot causal learning · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation › treatment effect estimation
individual treatment effect estimation
0.712023
Zero-shot causal learning · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation
meta-learning
0.712023
Zero-shot causal learning · NeurIPS 2023
Machine learning › Graph learning
graph neural network
0.612022
Topological Graph Neural Networks · ICLR 2022
Machine learning › Graph learning › graph neural network
topological graph neural network
0.612022
Topological Graph Neural Networks · ICLR 2022
Medical and health informatics
clinical prediction
0.612022
Prediction of recovery from multiple organ dysfunction syndrome in pediatric sepsis patients · Bioinform. 2022
Machine learning › Deep learning architectures and training
autoencoder
0.412020
Topological Autoencoders · ICML 2020
Computer vision › 3D vision › geometric deep learning › set learning
set function learning
0.412020
Set Functions for Time Series · ICML 2020
Medical and health informatics
clinical time series analysis
0.412020
Set Functions for Time Series · ICML 2020
Machine learning › Learning theory › neural network theory
neural network analysis
0.412019
Neural Persistence: A Complexity Measure for Deep Neural Networks Using Algebraic Topology · ICLR (Poster) 2019
Medical and health informatics › biomedical data science
biomedical time series analysis
0.312018
Association mapping in biomedical time series via statistically significant shapelet mining · Bioinform. 2018
Medical and health informatics › biomedical natural language processing › medical question answering
medical visual question answering
0.312025
SMMILE: An expert-driven benchmark for multimodal medical in-context learning · NeurIPS 2025
Medical and health informatics
precision medicine
0.212023
Zero-shot causal learning · NeurIPS 2023
Medical and health informatics
electronic health records
0.212022
Prediction of recovery from multiple organ dysfunction syndrome in pediatric sepsis patients · Bioinform. 2022
Medical and health informatics
clinical informatics
0.112020
Enhancing statistical power in temporal biomarker discovery through representative shapelet mining · Bioinform. 2020
Medical and health informatics › clinical prediction
ICU mortality prediction
0.112020
Enhancing statistical power in temporal biomarker discovery through representative shapelet mining · Bioinform. 2020
Machine learning › Learning theory
generalization
0.112019
Neural Persistence: A Complexity Measure for Deep Neural Networks Using Algebraic Topology · ICLR (Poster) 2019
Bioinformatics and computational biology › biomarker discovery
clinical biomarker discovery
0.112018
Association mapping in biomedical time series via statistically significant shapelet mining · Bioinform. 2018
Medical and health informatics › clinical prediction
sepsis prediction
0.112018
Association mapping in biomedical time series via statistically significant shapelet mining · Bioinform. 2018

Methods — techniques the papers use, named apart from their topics

in-context learning · 1.7benchmarking · 1.7meta-model training · 1.3causal meta-learning · 1.3topological data analysis · 1.1stepwise process reward modeling · 0.9reinforcement learning · 0.9topological loss · 0.9differentiable set functions · 0.9attention mechanism · 0.9machine learning · 0.6AUROC evaluation · 0.6submodular optimization · 0.4statistical significance testing · 0.4
YearPublicationVenuePosition
2025 Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
abstract
Jaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim, Xiangru Tang, Daniel Shao, Yong Hoe Koo, Ko Minhyeok, Qingyu Chen, Mark Gerstein, Michael Moor, Jaewoo Kang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim, Xiangru Tang, Daniel Shao, Yonghoe Koo, Minhyeok Ko, Qingyu Chen 0001, Mark Gerstein, Michael Moor, Jaewoo Kang
EMNLP11
2025 SMMILE: An expert-driven benchmark for multimodal medical in-context learning
abstract
Multimodal in-context learning (ICL) remains underexplored despite significant potential for domains such as medicine. Clinicians routinely encounter diverse, specialized tasks requiring adaptation from limited examples, such as drawing insights from a few relevant prior cases or considering a constrained set of differential diagnoses. While multimodal large language models (MLLMs) have shown advances in medical visual question answering (VQA), their ability to learn multimodal tasks from context is largely unknown. We introduce SMMILE, the first expert-driven multimodal ICL benchmark for medical tasks. Eleven medical experts curated problems, each including a multimodal query and multimodal in-context examples as task demonstrations. SMMILE encompasses 111 problems (517 question-image-answer triplets) covering 6 medical specialties and 13 imaging modalities. We further introduce SMMILE++, an augmented variant with 1038 permuted problems. A comprehensive evaluation of 15 MLLMs demonstrates that most models exhibit moderate to poor multimodal ICL ability in medical tasks. In open-ended evaluations, ICL contributes only an 8% average improvement over zero-shot on SMMILE and 9.4% on SMMILE++. We observe a susceptibility for irrelevant in-context examples: even a single noisy or irrelevant example can degrade performance by up to 9.5%. Moreover, we observe that MLLMs are affected by a recency bias, where placing the most relevant example last can lead to substantial performance improvements of up to 71%. Our findings highlight critical limitations and biases in current MLLMs when learning multimodal medical tasks from context. SMMILE is available at https://smmile-benchmark.github.io.
Melanie Rieff, Maya Varma, Ossian Rabow, Subathra Adithan, Julie Kim, Ken Chang, Hannah Lee, Nidhi Rohatgi, Christian Bluethgen, Mohamed S. Muneer, Jean-Benoit Delbrouck, Michael Moor
NeurIPS12
2025 MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
abstract
Multimodal Large Language Models (LLMs) hold promise for biomedical reasoning, but current benchmarks fail to capture the complexity of real-world clinical workflows. Existing evaluations primarily assess unimodal, decontextualized question-answering, overlooking multi-agent decision-making environments such as Molecular Tumor Boards (MTBs). MTBs bring together diverse experts in oncology, where diagnostic and prognostic tasks require integrating heterogeneous data and evolving insights over time. Current benchmarks lack this longitudinal and multimodal complexity. We introduce MTBBench, an agentic benchmark simulating MTB-style decision-making through clinically challenging, multimodal, and longitudinal oncology questions. Ground truth annotations are validated by clinicians via a co-developed app, ensuring clinical relevance. We benchmark multiple open and closed-source LLMs and show that, even at scale, they lack reliability—frequently hallucinating, struggling with reasoning from time-resolved data, and failing to reconcile conflicting evidence or different modalities. To address these limitations, MTBBench goes beyond benchmarking by providing an agentic framework with foundation model-based tools that enhance multi-modal and longitudinal reasoning, leading to task-level performance gains of up to 9.0% and 11.2%, respectively. Overall, MTBBench offers a challenging and realistic testbed for advancing multimodal LLM reasoning, reliability, and tool-use with a focus on MTB environments in precision oncology.
Kiril Vasilev, Alexandre Misrahi, Eeshaan Jain, Phil F. Cheng, Petros Liakopoulos, Olivier Michielin, Michael Moor, Charlotte Bunne
NeurIPS7
2023 Zero-shot causal learning
abstract
Predicting how different interventions will causally affect a specific individual is important in a variety of domains such as personalized medicine, public policy, and online marketing. There are a large number of methods to predict the effect of an existing intervention based on historical data from individuals who received it. However, in many settings it is important to predict the effects of novel interventions (e.g., a newly invented drug), which these methods do not address. Here, we consider zero-shot causal learning: predicting the personalized effects of a novel intervention. We propose CaML, a causal meta-learning framework which formulates the personalized prediction of each intervention's effect as a task. CaML trains a single meta-model across thousands of tasks, each constructed by sampling an intervention, its recipients, and its nonrecipients. By leveraging both intervention information (e.g., a drug's attributes) and individual features (e.g., a patient's history), CaML is able to predict the personalized effects of novel interventions that do not exist at the time of training. Experimental results on real world datasets in large-scale medical claims and cell-line perturbations demonstrate the effectiveness of our approach. Most strikingly, CaML's zero-shot predictions outperform even strong baselines trained directly on data from the test interventions.
Hamed Nilforoshan, Michael Moor, Yusuf H. Roohani, Anja Surina, Michihiro Yasunaga, Sara Oblak, Jure Leskovec
NeurIPS2
2022 Topological Graph Neural Networks
Max Horn, Edward De Brouwer, Michael Moor, Yves Moreau, Bastian Rieck, Karsten M. Borgwardt
ICLR3
2022 Prediction of recovery from multiple organ dysfunction syndrome in pediatric sepsis patients
abstract
MOTIVATION: Sepsis is a leading cause of death and disability in children globally, accounting for ∼3 million childhood deaths per year. In pediatric sepsis patients, the multiple organ dysfunction syndrome (MODS) is considered a significant risk factor for adverse clinical outcomes characterized by high mortality and morbidity in the pediatric intensive care unit. The recent rapidly growing availability of electronic health records (EHRs) has allowed researchers to vastly develop data-driven approaches like machine learning in healthcare and achieved great successes. However, effective machine learning models which could make the accurate early prediction of the recovery in pediatric sepsis patients from MODS to a mild state and thus assist the clinicians in the decision-making process is still lacking. RESULTS: This study develops a machine learning-based approach to predict the recovery from MODS to zero or single organ dysfunction by 1 week in advance in the Swiss Pediatric Sepsis Study cohort of children with blood-culture confirmed bacteremia. Our model achieves internal validation performance on the SPSS cohort with an area under the receiver operating characteristic (AUROC) of 79.1% and area under the precision-recall curve (AUPRC) of 73.6%, and it was also externally validated on another pediatric sepsis patients cohort collected in the USA, yielding an AUROC of 76.4% and AUPRC of 72.4%. These results indicate that our model has the potential to be included into the EHRs system and contribute to patient assessment and triage in pediatric sepsis patient care. AVAILABILITY AND IMPLEMENTATION: Code available at https://github.com/BorgwardtLab/MODS-recovery. The data underlying this article is not publicly available for the privacy of individuals that participated in the study. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bowen Fan, Juliane Klatt, Michael Moor, Latasha A. Daniels, Philipp K. A Agyeman, Christoph Berger, Eric Giannoni, Martin Stocker, Klara M. Posfay-Barbe, Ulrich Heininger, Sara Bernhard-Stirnemann, Anita Niederer-Loher, Christian R. Kahlert, Giancarlo Natalucci, Christa Relly, Thomas Riedel, Christoph Aebi, Luregn J. Schlapbach, L. Nelson Sanchez-Pinto, Karsten M. Borgwardt
Bioinform.3
2020 Set Functions for Time Series
abstract
Despite the eminent successes of deep neural networks, many architectures are often hard to transfer to irregularly-sampled and asynchronous time series that commonly occur in real-world datasets, especially in healthcare applications. This paper proposes a novel approach for classifying irregularly-sampled time series with unaligned measurements, focusing on high scalability and data efficiency. Our method SeFT (Set Functions for Time Series) is based on recent advances in differentiable set function learning, extremely parallelizable with a beneficial memory footprint, thus scaling well to large datasets of long time series and online monitoring scenarios. Furthermore, our approach permits quantifying per-observation contributions to the classification outcome. We extensively compare our method with existing algorithms on multiple healthcare time series datasets and demonstrate that it performs competitively whilst significantly reducing runtime.
Max Horn, Michael Moor, Christian Bock, Bastian Rieck, Karsten M. Borgwardt
ICML2
2020 Topological Autoencoders
abstract
We propose a novel approach for preserving topological structures of the input space in latent representations of autoencoders. Using persistent homology, a technique from topological data analysis, we calculate topological signatures of both the input and latent space to derive a topological loss term. Under weak theoretical assumptions, we construct this loss in a differentiable manner, such that the encoding learns to retain multi-scale connectivity information. We show that our approach is theoretically well-founded and that it exhibits favourable latent representations on a synthetic manifold as well as on real-world image data sets, while preserving low reconstruction errors.
Michael Moor, Max Horn, Bastian Rieck, Karsten M. Borgwardt
ICML1
2020 Enhancing statistical power in temporal biomarker discovery through representative shapelet mining
abstract
MOTIVATION: Temporal biomarker discovery in longitudinal data is based on detecting reoccurring trajectories, the so-called shapelets. The search for shapelets requires considering all subsequences in the data. While the accompanying issue of multiple testing has been mitigated in previous work, the redundancy and overlap of the detected shapelets results in an a priori unbounded number of highly similar and structurally meaningless shapelets. As a consequence, current temporal biomarker discovery methods are impractical and underpowered. RESULTS: We find that the pre- or post-processing of shapelets does not sufficiently increase the power and practical utility. Consequently, we present a novel method for temporal biomarker discovery: Statistically Significant Submodular Subset Shapelet Mining (S5M) that retrieves short subsequences that are (i) occurring in the data, (ii) are statistically significantly associated with the phenotype and (iii) are of manageable quantity while maximizing structural diversity. Structural diversity is achieved by pruning non-representative shapelets via submodular optimization. This increases the statistical power and utility of S5M compared to state-of-the-art approaches on simulated and real-world datasets. For patients admitted to the intensive care unit (ICU) showing signs of severe organ failure, we find temporal patterns in the sequential organ failure assessment score that are associated with in-ICU mortality. AVAILABILITY AND IMPLEMENTATION: S5M is an option in the python package of S3M: github.com/BorgwardtLab/S3M.
Thomas Gumbsch, Christian Bock, Michael Moor, Bastian Rieck, Karsten M. Borgwardt
Bioinform.3
2019 Neural Persistence: A Complexity Measure for Deep Neural Networks Using Algebraic Topology
Bastian Rieck, Matteo Togninalli, Christian Bock, Michael Moor, Max Horn, Thomas Gumbsch, Karsten M. Borgwardt
ICLR (Poster)4
2018 Association mapping in biomedical time series via statistically significant shapelet mining
abstract
Motivation: Most modern intensive care units record the physiological and vital signs of patients. These data can be used to extract signatures, commonly known as biomarkers, that help physicians understand the biological complexity of many syndromes. However, most biological biomarkers suffer from either poor predictive performance or weak explanatory power. Recent developments in time series classification focus on discovering shapelets, i.e. subsequences that are most predictive in terms of class membership. Shapelets have the advantage of combining a high predictive performance with an interpretable component-their shape. Currently, most shapelet discovery methods do not rely on statistical tests to verify the significance of individual shapelets. Therefore, identifying associations between the shapelets of physiological biomarkers and patients that exhibit certain phenotypes of interest enables the discovery and subsequent ranking of physiological signatures that are interpretable, statistically validated and accurate predictors of clinical endpoints. Results: We present a novel and scalable method for scanning time series and identifying discriminative patterns that are statistically significant. The significance of a shapelet is evaluated while considering the problem of multiple hypothesis testing and mitigating it by efficiently pruning untestable shapelet candidates with Tarone's method. We demonstrate the utility of our method by discovering patterns in three of a patient's vital signs: heart rate, respiratory rate and systolic blood pressure that are indicators of the severity of a future sepsis event, i.e. an inflammatory response to an infective agent that can lead to organ failure and death, if not treated in time. Availability and implementation: We make our method and the scripts that are required to reproduce the experiments publicly available at https://github.com/BorgwardtLab/S3M. Supplementary information: Supplementary data are available at Bioinformatics online.
Christian Bock, Thomas Gumbsch, Michael Moor, Bastian Rieck, Damian Roqueiro, Karsten M. Borgwardt
Bioinform.3