Zalan Fabian

dblp:192/2874 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 43% Efficient and distributed learning · 12% Trustworthy machine learning · 11%
Computer graphics and multimedia
3 papers
Image and video processing · 100%

Topics — the 15 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.432025
Emergence and Evolution of Interpretable Concepts in Diffusion Models · NeurIPS 2025
DiracDiffusion: Denoising and Incremental Reconstruction with Assured Data-Consistency · ICML 2024
Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models · ICML 2024
Image and video processing
image restoration
2.032024
DiracDiffusion: Denoising and Incremental Reconstruction with Assured Data-Consistency · ICML 2024
Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models · ICML 2024
Data augmentation for deep learning based accelerated MRI reconstruction with limited data · ICML 2021
Machine learning › Generative modeling › diffusion model
inverse problem solving
1.522024
DiracDiffusion: Denoising and Incremental Reconstruction with Assured Data-Consistency · ICML 2024
Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models · ICML 2024
Machine learning › Generative modeling › diffusion model
controllable generation
0.912025
Emergence and Evolution of Interpretable Concepts in Diffusion Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.912025
Emergence and Evolution of Interpretable Concepts in Diffusion Models · NeurIPS 2025
Computer vision › Vision and language › visual question answering
medical visual question answering
0.912025
MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models · ICLR 2025
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.712023
A Data-Free Approach to Mitigate Catastrophic Forgetting in Federated Class Incremental Learning for Vision Tasks · NeurIPS 2023
Machine learning › Efficient and distributed learning › federated learning › federated continual learning
federated class-incremental learning
0.712023
A Data-Free Approach to Mitigate Catastrophic Forgetting in Federated Class Incremental Learning for Vision Tasks · NeurIPS 2023
Machine learning › Efficient and distributed learning
federated learning
0.712023
A Data-Free Approach to Mitigate Catastrophic Forgetting in Federated Class Incremental Learning for Vision Tasks · NeurIPS 2023
Computer vision › 3D vision › medical image reconstruction
MRI reconstruction
0.612022
HUMUS-Net: Hybrid Unrolled Multi-scale Network Architecture for Accelerated MRI Reconstruction · NeurIPS 2022
Machine learning › Deep learning architectures and training
data augmentation
0.512021
Data augmentation for deep learning based accelerated MRI reconstruction with limited data · ICML 2021
Image and video processing › image reconstruction › medical image reconstruction
MRI reconstruction
0.512021
Data augmentation for deep learning based accelerated MRI reconstruction with limited data · ICML 2021
Machine learning › Trustworthy machine learning
robustness evaluation
0.312025
MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
0.312025
Emergence and Evolution of Interpretable Concepts in Diffusion Models · NeurIPS 2025
Computer vision › Image recognition and object detection
image classification
0.212023
A Data-Free Approach to Mitigate Catastrophic Forgetting in Federated Class Incremental Learning for Vision Tasks · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

severity encoding · 1.5latent diffusion model · 1.5diffusion model · 1.5sparse autoencoder · 0.9intervention techniques · 0.9benchmark dataset construction · 0.9generative model · 0.7data-free learning · 0.7multi-scale feature extraction · 0.6convolution · 0.6deep learning · 0.5data augmentation · 0.5
YearPublicationVenuePosition
2025 MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models
abstract
Multimodal Large Language Models (MLLMs) have tremendous potential to improve the accuracy, availability, and cost-effectiveness of healthcare by providing automated solutions or serving as aids to medical professionals. Despite promising first steps in developing medical MLLMs in the past few years, their capabilities and limitations are not well understood. Recently, many benchmark datasets have been proposed that test the general medical knowledge of such models across a variety of medical areas. However, the systematic failure modes and vulnerabilities of such models are severely underexplored with most medical benchmarks failing to expose the shortcomings of existing models in this safety-critical domain. In this paper, we introduce MediConfusion, a challenging medical Visual Question Answering (VQA) benchmark dataset, that probes the failure modes of medical MLLMs from a vision perspective. We reveal that state-of-the-art models are easily confused by image pairs that are otherwise visually dissimilar and clearly distinct for medical experts. Strikingly, all available models (open-source or proprietary) achieve performance below random guessing on MediConfusion, raising serious concerns about the reliability of existing medical MLLMs for healthcare deployment. We also extract common patterns of model failure that may help the design of a new generation of more trustworthy and reliable MLLMs in healthcare.
Mohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi Soltanolkotabi
ICLR2
2025 Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
abstract
Mental visualization, the ability to construct and manipulate visual representations internally, is a core component of human cognition and plays a vital role in tasks involving reasoning, prediction, and abstraction. Despite the rapid progress of Multimodal Large Language Models (MLLMs), current benchmarks primarily assess passive visual perception, offering limited insight into the more active capability of internally constructing visual patterns to support problem solving. Yet mental visualization is a critical cognitive skill in humans, supporting abilities such as spatial navigation, predicting physical trajectories, and solving complex visual problems through imaginative simulation. To bridge this gap, we introduce Hyperphantasia, a synthetic benchmark designed to evaluate the mental visualization abilities of MLLMs through four carefully constructed puzzles. Each task is procedurally generated and presented at three difficulty levels, enabling controlled analysis of model performance across increasing complexity. Our comprehensive evaluation of state-of-the-art models reveals a substantial gap between the performance of humans and MLLMs. Additionally, we explore the potential of reinforcement learning to improve visual simulation capabilities. Our findings suggest that while some models exhibit partial competence in recognizing visual patterns, robust mental visualization remains an open challenge for current MLLMs.
Mohammad Shahab Sepehri, Berk Tinaz, Zalan Fabian, Mahdi Soltanolkotabi
NeurIPS3
2025 Emergence and Evolution of Interpretable Concepts in Diffusion Models
abstract
Diffusion models have become the go-to method for text-to-image generation, producing high-quality images from pure noise. However, the inner workings of diffusion models is still largely a mystery due to their black-box nature and complex, multi-step generation process. Mechanistic interpretability techniques, such as Sparse Autoencoders (SAEs), have been successful in understanding and steering the behavior of large language models at scale. However, the great potential of SAEs has not yet been applied toward gaining insight into the intricate generative process of diffusion models. In this work, we leverage the SAE framework to probe the inner workings of a popular text-to-image diffusion model, and uncover a variety of human-interpretable concepts in its activations. Interestingly, we find that *even before the first reverse diffusion step* is completed, the final composition of the scene can be predicted surprisingly well by looking at the spatial distribution of activated concepts. Moreover, going beyond correlational analysis, we design intervention techniques aimed at manipulating image composition and style, and demonstrate that (1) in early stages of diffusion image composition can be effectively controlled, (2) in the middle stages image composition is finalized, however stylistic interventions are effective, and (3) in the final stages only minor textural details are subject to change.
Berk Tinaz, Zalan Fabian, Mahdi Soltanolkotabi
NeurIPS2
2024 Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models
abstract
Inverse problems arise in a multitude of applications, where the goal is to recover a clean signal from noisy and possibly (non)linear observations. The difficulty of a reconstruction problem depends on multiple factors, such as the ground truth signal structure, the severity of the degradation and the complex interactions between the above. This results in natural sample-by-sample variation in the difficulty of a reconstruction problem. Our key observation is that most existing inverse problem solvers lack the ability to adapt their compute power to the difficulty of the reconstruction task, resulting in subpar performance and wasteful resource allocation. We propose a novel method, severity encoding, to estimate the degradation severity of corrupted signals in the latent space of an autoencoder. We show that the estimated severity has strong correlation with the true corruption level and can provide useful hints on the difficulty of reconstruction problems on a sample-by-sample basis. Furthermore, we propose a reconstruction method based on latent diffusion models that leverages the predicted degradation severities to fine-tune the reverse diffusion sampling trajectory and thus achieve sample-adaptive inference times. Our framework, Flash-Diffusion, acts as a wrapper that can be combined with any latent diffusion-based baseline solver, imbuing it with sample-adaptivity and acceleration. We perform experiments on both linear and nonlinear inverse problems and demonstrate that our technique greatly improves the performance of the baseline solver and achieves up to $10\times$ acceleration in mean sampling speed.
Zalan Fabian, Berk Tinaz, Mahdi Soltanolkotabi
ICML1
2024 DiracDiffusion: Denoising and Incremental Reconstruction with Assured Data-Consistency
abstract
Diffusion models have established new state of the art in a multitude of computer vision tasks, including image restoration. Diffusion-based inverse problem solvers generate reconstructions of exceptional visual quality from heavily corrupted measurements. However, in what is widely known as the perception-distortion trade-off, the price of perceptually appealing reconstructions is often paid in declined distortion metrics, such as PSNR. Distortion metrics measure faithfulness to the observation, a crucial requirement in inverse problems. In this work, we propose a novel framework for inverse problem solving, namely we assume that the observation comes from a stochastic degradation process that gradually degrades and noises the original clean image. We learn to reverse the degradation process in order to recover the clean image. Our technique maintains consistency with the original measurement throughout the reverse process, and allows for great flexibility in trading off perceptual quality for improved distortion metrics and sampling speedup via early-stopping. We demonstrate the efficiency of our method on different high-resolution datasets and inverse problems, achieving great improvements over other state-of-the-art diffusion-based methods with respect to both perceptual and distortion metrics.
Zalan Fabian, Berk Tinaz, Mahdi Soltanolkotabi
ICML1
2023 A Data-Free Approach to Mitigate Catastrophic Forgetting in Federated Class Incremental Learning for Vision Tasks
abstract
Deep learning models often suffer from forgetting previously learned information when trained on new data. This problem is exacerbated in federated learning (FL), where the data is distributed and can change independently for each user. Many solutions are proposed to resolve this catastrophic forgetting in a centralized setting. However, they do not apply directly to FL because of its unique complexities, such as privacy concerns and resource limitations. To overcome these challenges, this paper presents a framework for \textbf{federated class incremental learning} that utilizes a generative model to synthesize samples from past distributions. This data can be later exploited alongside the training data to mitigate catastrophic forgetting. To preserve privacy, the generative model is trained on the server using data-free methods at the end of each task without requesting data from clients. Moreover, our solution does not demand the users to store old data or models, which gives them the freedom to join/leave the training at any time. Additionally, we introduce SuperImageNet, a new regrouping of the ImageNet dataset specifically tailored for federated continual learning. We demonstrate significant improvements compared to existing baselines through extensive experiments on multiple datasets.
Sara Babakniya, Zalan Fabian, Chaoyang He 0001, Mahdi Soltanolkotabi, Amir Salman Avestimehr
NeurIPS2
2022 HUMUS-Net: Hybrid Unrolled Multi-scale Network Architecture for Accelerated MRI Reconstruction
abstract
In accelerated MRI reconstruction, the anatomy of a patient is recovered from a set of undersampled and noisy measurements. Deep learning approaches have been proven to be successful in solving this ill-posed inverse problem and are capable of producing very high quality reconstructions. However, current architectures heavily rely on convolutions, that are content-independent and have difficulties modeling long-range dependencies in images. Recently, Transformers, the workhorse of contemporary natural language processing, have emerged as powerful building blocks for a multitude of vision tasks. These models split input images into non-overlapping patches, embed the patches into lower-dimensional tokens and utilize a self-attention mechanism that does not suffer from the aforementioned weaknesses of convolutional architectures. However, Transformers incur extremely high compute and memory cost when 1) the input image resolution is high and 2) when the image needs to be split into a large number of patches to preserve fine detail information, both of which are typical in low-level vision problems such as MRI reconstruction, having a compounding effect. To tackle these challenges, we propose HUMUS-Net, a hybrid architecture that combines the beneficial implicit bias and efficiency of convolutions with the power of Transformer blocks in an unrolled and multi-scale network. HUMUS-Net extracts high-resolution features via convolutional blocks and refines low-resolution features via a novel Transformer-based multi-scale feature extractor. Features from both levels are then synthesized into a high-resolution output reconstruction. Our network establishes new state of the art on the largest publicly available MRI dataset, the fastMRI dataset. We further demonstrate the performance of HUMUS-Net on two other popular MRI datasets and perform fine-grained ablation studies to validate our design.
Zalan Fabian, Berk Tinaz, Mahdi Soltanolkotabi
NeurIPS1
2021 Data augmentation for deep learning based accelerated MRI reconstruction with limited data
abstract
Deep neural networks have emerged as very successful tools for image restoration and reconstruction tasks. These networks are often trained end-to-end to directly reconstruct an image from a noisy or corrupted measurement of that image. To achieve state-of-the-art performance, training on large and diverse sets of images is considered critical. However, it is often difficult and/or expensive to collect large amounts of training images. Inspired by the success of Data Augmentation (DA) for classification problems, in this paper, we propose a pipeline for data augmentation for accelerated MRI reconstruction and study its effectiveness at reducing the required training data in a variety of settings. Our DA pipeline, MRAugment, is specifically designed to utilize the invariances present in medical imaging measurements as naive DA strategies that neglect the physics of the problem fail. Through extensive studies on multiple datasets we demonstrate that in the low-data regime DA prevents overfitting and can match or even surpass the state of the art while using significantly fewer training data, whereas in the high-data regime it has diminishing returns. Furthermore, our findings show that DA improves the robustness of the model against various shifts in the test distribution.
Zalan Fabian, Reinhard Heckel, Mahdi Soltanolkotabi
ICML1
2020 Minimax Lower Bounds for Transfer Learning with Linear and One-hidden Layer Neural Networks
abstract
Transfer learning has emerged as a powerful technique for improving the performance of machine learning models on new domains where labeled training data may be scarce. In this approach a model trained for a source task, where plenty of labeled training data is available, is used as a starting point for training a model on a related target task with only few labeled training data. Despite recent empirical success of transfer learning approaches, the benefits and fundamental limits of transfer learning are poorly understood. In this paper we develop a statistical minimax framework to characterize the fundamental limits of transfer learning in the context of regression with linear and one-hidden layer neural network models. Specifically, we derive a lower-bound for the target generalization error achievable by any algorithm as a function of the number of labeled source and target data as well as appropriate notions of similarity between the source and target tasks. Our lower bound provides new insights into the benefits and limitations of transfer learning. We further corroborate our theoretical finding with various experiments.
Mohammadreza M. Kalan, Zalan Fabian, Amir Salman Avestimehr, Mahdi Soltanolkotabi
NeurIPS2