VLDB 2026 Research / reviewers in the wild / expert
Reza Azad
dblp:149/2290
· DBLP profile ↗
21ranked-venue papers
8as first author
20since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SL2 A-INR: Single-Layer Learnable Activation for Implicit Neural Representation
Reza Rezaeian, Moein Heidari, Reza Azad, Dorit Merhof, Hamid Soltanian-Zadeh, Ilker Hacihaliloglu |
ICCV | 3 |
| 2025 | CENet: Context Enhancement Network for Medical Image Segmentation
Afshin Bozorgpour, Sina Ghorbani Kolahi, Reza Azad, Ilker Hacihaliloglu, Dorit Merhof |
MICCAI (1) | 3 |
| 2025 | LHU-Net: A Lean Hybrid U-Net for Cost-Efficient, High-Performance Volumetric Segmentation
Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof |
MICCAI (14) | 4 |
| 2025 | Frequency-Domain Refinement of Vision Transformers for Robust Medical Image Segmentation Under DegradationabstractMedical image segmentation is crucial for precise diagnosis, treatment planning, and disease monitoring in clinical settings. While convolutional neural networks (CNNs) have achieved remarkable success, they struggle with modeling long-range dependencies. Vision Transformers (ViTs) address this limitation by leveraging self-attention mechanisms to capture global contextual information. However, ViTs often fall short in local feature description, which is crucial for precise segmentation. To address this issue, we reformulate self-attention in the frequency domain to enhance both local and global feature representation. Our approach, the Enhanced Wave Vision Transformer (EW-ViT), incorporates wavelet decomposition within the self-attention block to adaptively refine feature representation in low and high-frequency components. We also introduce the Prompt-Guided High-Frequency Refiner (PGHFR) module to handle image degradation, which mainly affects high-frequency components. This module uses implicit prompts to encode degradation-specific information and adjust high-frequency representations accordingly. Additionally, we apply a contrastive learning strategy to maintain feature consistency and ensure robustness against noise, leading to state-of-the-art (SOTA) performance in medical image segmentation, especially under various conditions of degradation. Source code is available at GitHub. Sanaz Karimijafarbigloo, Sina Ghorbani Kolahi, Reza Azad, Ulas Bagci, Dorit Merhof |
WACV | 3 |
| 2025 | Addressing Missing Modality Challenges in MRI Images: A Comprehensive ReviewabstractMagnetic resonance imaging (MRI) is one of the most prevalent imaging modalities used for diagnosis, treatment planning, and outcome control in various medical conditions. MRI sequences provide physicians with the ability to view and monitor tissues at multiple contrasts within a single scan and serve as input for automated systems to perform downstream tasks. However, in clinical practice, there is usually no concise set of identically acquired sequences for a whole group of patients. As a consequence, medical professionals and automated systems both face difficulties due to the lack of complementary information from such missing sequences. This problem is well known in computer vision, particularly in medical image processing tasks such as tumor segmentation, tissue classification, and image generation. With the aim of helping researchers, this literature review examines a significant number of recent approaches that attempt to mitigate these problems. Basic techniques such as early synthesis methods, as well as later approaches that deploy deep learning, such as common latent space models, knowledge distillation networks, mutual information maximization, and generative adversarial networks (GANs) are examined in detail. We investigate the novelty, strengths, and weaknesses of the aforementioned strategies. Moreover, using a case study on the segmentation task, our survey offers quantitative benchmarks to further analyze the effectiveness of these methods for addressing the missing modalities challenge. Furthermore, a discussion offers possible future research directions. Reza Azad, Mohammad Dehghanmanshadi, Nika Khosravi, Julien Cohen-Adad, Dorit Merhof |
Comput. Vis. Media | 1 |
| 2025 | TransCeption: Enhancing Medical Image Segmentation with an Inception-Like Transformer Design for Efficient Feature FusionabstractWhile CNN-based methods have been the cornerstone of medical image segmentation due to their promising performance and robustness, they suffer from limitations in capturing long-range dependencies. Transformer-based approaches are currently prevailing since they enlarge the receptive field to model global contextual correlations. To further extract rich representations, some extensions of U-Net employ multi-scale feature extraction and fusion modules to obtain improved performance. Inspired by this idea, we propose TransCeption for medical image segmentation, a pure transformer-based U-shaped network incorporating an inception-like module in the encoder and adopting a contextual bridge for better feature fusion. The design proposed in this work is based on three core principles. (i) The patch merging module in the encoder is redesigned to use ResInception Patch Merging (RIPM). The MultiBranch (MB) transformer has the same number of branches as the outputs of RIPM. Combining the two modules enables the model to capture a multi-scale representation within a single stage. (ii) We apply an Intra-stage Feature Fusion (IFF) module following the MB transformer to enhance the aggregation of feature maps from all branches and particularly focus on the interaction between the different channels at all scales. (iii) In contrast to a bridge that only contains token-wise self-attention, we propose a Dual Transformer Bridge that also includes channel-wise self-attention to exploit correlations between scales at different stages from a dual perspective. Extensive experiments on multi-organ and skin lesion segmentation tasks show the superiority of TransCeption to previous work. The code is publicly available on GitHub. Reza Azad, Yiwei Jia, Ehsan Khodapanah Aghdam, Julien Cohen-Adad, Dorit Merhof |
Comput. Vis. Media | 1 |
| 2025 | MedScale-Former: Self-guided multiscale transformer for medical image segmentation
Sanaz Karimijafarbigloo, Reza Azad, Amirhossein Kazerouni, Dorit Merhof |
Medical Image Anal. | 2 |
| 2025 | Continual learning in medical image analysis: A comprehensive review of recent advancements and future prospectsabstractMedical image analysis has witnessed remarkable advancements, even surpassing human-level performance in recent years, driven by the rapid development of advanced deep-learning algorithms. However, when the inference dataset slightly differs from what the model has seen during one-time training, the model performance is greatly compromised. The situation requires restarting the training process using both the old and the new data, which is computationally costly, does not align with the human learning process, and imposes storage constraints and privacy concerns. Alternatively, continual learning has emerged as a crucial approach for developing unified and sustainable deep models to deal with new classes, tasks, and the drifting nature of data in non-stationary environments for various application areas. Continual learning techniques enable models to adapt and accumulate knowledge over time, which is essential for maintaining performance on evolving datasets and novel tasks. Owing to its popularity and promising performance, it is an active and emerging research topic in the medical field and hence demands a survey and taxonomy to clarify the current research landscape of continual learning in medical image analysis. This systematic review paper provides a comprehensive overview of the state-of-the-art in continual learning techniques applied to medical image analysis. We present an extensive survey of existing research, covering topics including catastrophic forgetting, data drifts, stability, and plasticity requirements. Further, an in-depth discussion of key components of a continual learning framework, such as continual learning scenarios, techniques, evaluation schemes, and metrics, is provided. Continual learning techniques encompass various categories, including rehearsal, regularization, architectural, and hybrid strategies. We assess the popularity and applicability of continual learning categories in various medical sub-fields like radiology and histopathology. Our exploration considers unique challenges in the medical domain, including costly data annotation, temporal drift, and the crucial need for benchmarking datasets to ensure consistent model evaluation. The paper also addresses current challenges and looks ahead to potential future research directions. Pratibha Kumari 0001, Joohi Chauhan, Afshin Bozorgpour, Boqiang Huang, Reza Azad, Dorit Merhof |
Medical Image Anal. | 5 |
| 2024 | MSA2Net: Multi-scale Adaptive Attention-guided Network for Medical Image Segmentation
Sina Ghorbani Kolahi, Seyed Kamal Chaharsooghi, Toktam Khatibi, Afshin Bozorgpour, Reza Azad, Moein Heidari, Ilker Hacihaliloglu, Dorit Merhof |
BMVC | 5 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 29 |
| 2024 | Beyond Self-Attention: Deformable Large Kernel Attention for Medical Image SegmentationabstractMedical image segmentation has seen significant improvements with transformer models, which excel in grasping far-reaching contexts and global contextual information. However, the increasing computational demands of these models, proportional to the squared token count, limit their depth and resolution capabilities. Most current methods process D volumetric image data slice-by-slice (called pseudo 3D), missing crucial inter-slice information and thus reducing the model’s overall performance. To address these challenges, we introduce the concept of Deformable Large Kernel Attention (D-LKA Attention), a streamlined attention mechanism employing large convolution kernels to fully appreciate volumetric context. This mechanism operates within a receptive field akin to self-attention while sidestepping the computational overhead. Additionally, our proposed attention mechanism benefits from deformable convolutions to flexibly warp the sampling grid, enabling the model to adapt appropriately to diverse data patterns. We designed both 2D and 3D adaptations of the D-LKA Attention, with the latter excelling in cross-depth data understanding. Together, these components shape our novel hierarchical Vision Transformer architecture, the D-LKA Net. Evaluations of our model against leading methods on popular medical segmentation datasets (Synapse, NIH Pancreas, and Skin lesion) demonstrate its superior performance. Our code is publicly available at GitHub. Reza Azad, Leon Niggemeier, Michael Huttemann, Amirhossein Kazerouni, Ehsan Khodapanah Aghdam, Yury Velichko, Ulas Bagci, Dorit Merhof |
WACV | 1 |
| 2024 | INCODE: Implicit Neural Conditioning with Prior Knowledge EmbeddingsabstractImplicit Neural Representations (INRs) have revolutionized signal representation by leveraging neural networks to provide continuous and smooth representations of complex data. However, existing INRs face limitations in capturing fine-grained details, handling noise, and adapting to diverse signal types. To address these challenges, we introduce INCODE, a novel approach that enhances the control of the sinusoidal-based activation function in INRs using deep prior knowledge. INCODE comprises a harmonizer network and a composer network, where the harmonizer network dynamically adjusts key parameters of the activation function. Through a task-specific pre-trained model, INCODE adapts the task-specific parameters to optimize the representation process. Our approach not only excels in representation, but also extends its prowess to tackle complex tasks such as audio, image, and 3D shape reconstructions, as well as intricate challenges such as neural radiance fields (NeRFs), and inverse problems, including denoising, super-resolution, inpainting, and CT reconstruction. Through comprehensive experiments, INCODE demonstrates its superiority in terms of robustness, accuracy, quality, and convergence rate, broadening the scope of signal representation. Please visit the project’s website for details on the proposed method and access to the code. Amirhossein Kazerouni, Reza Azad, Alireza Hosseini, Dorit Merhof, Ulas Bagci |
WACV | 2 |
| 2024 | Advances in medical image analysis with vision Transformers: A comprehensive review
Reza Azad, Amirhossein Kazerouni, Moein Heidari, Ehsan Khodapanah Aghdam, Amirali Molaei, Yiwei Jia, Abin Jose, Rijo Roy, Dorit Merhof |
Medical Image Anal. | 1 |
| 2024 | Medical Image Segmentation Review: The Success of U-NetabstractAutomatic medical image segmentation is a crucial topic in the medical domain and successively a critical counterpart in the computer-aided diagnosis paradigm. U-Net is the most widespread image segmentation architecture due to its flexibility, optimized modular design, and success in all medical image modalities. Over the years, the U-Net model has received tremendous attention from academic and industrial researchers who have extended it to address the scale and complexity created by medical tasks. These extensions are commonly related to enhancing the U-Net's backbone, bottleneck, or skip connections, or including representation learning, or combining it with a Transformer architecture, or even addressing probabilistic prediction of the segmentation map. Having a compendium of different previously proposed U-Net variants makes it easier for machine learning researchers to identify relevant research questions and understand the challenges of the biological tasks that challenge the model. In this work, we discuss the practical aspects of the U-Net model and organize each variant model into a taxonomy. Moreover, to measure the performance of these strategies in a clinical application, we propose fair evaluations of some unique and famous designs on well-known datasets. Furthermore, we provide a comprehensive implementation library with trained models. In addition, for ease of future studies, we created an online list of U-Net papers with their possible official implementation. Reza Azad, Ehsan Khodapanah Aghdam, Amelie Rauland, Yiwei Jia, Atlas Haddadi Avval, Afshin Bozorgpour, Sanaz Karimijafarbigloo, Joseph Paul Cohen, Ehsan Adeli-Mosabbeb, Dorit Merhof |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | End-to-End Classification of Cell-Cycle Stages with Center-Cell Focus Tracker Using Recurrent Neural NetworksabstractCell division, or mitosis, guarantees the accurate inheritance of the genomic information kept in the cell nucleus. Malfunctions in this process cause a threat to the health and life of the organism, including cancer and other manifold diseases. It is therefore crucial to study in detail the cell-cycle in general and mitosis in particular. Consequently, a large number of manual and semi-automated time-lapse microscopy image analyses of mitosis have been carried out in recent years. In this paper, we propose a method for automatic detection of cell-cycle stages using a recurrent neural network (RNN). An end-to-end model with center-cell focus tracker loss, and classification loss is trained. The evaluation was conducted on two time-series datasets, with 6-stages and 3-stages of cell splitting labeled. The frame-to-frame accuracy was calculated and precision, recall, and F1-Score were measured for each cell-cycle stage. We also visualized the learned feature space. Image reconstruction from the center-cell focus module was performed which shows that the network was able to focus on the center-cell and classify it simultaneously. Our experiments validate the superior performance of the proposed network compared to a classifier baseline. Abin Jose, Rijo Roy, Dennis Eschweiler, Ina Laube, Reza Azad, Daniel Moreno-Andrés, Johannes Stegmaier |
ICASSP | 5 |
| 2023 | Laplacian-Former: Overcoming the Limitations of Vision Transformers in Local Texture Detection
Reza Azad, Amirhossein Kazerouni, Babak Azad, Ehsan Khodapanah Aghdam, Yury Velichko, Ulas Bagci, Dorit Merhof |
MICCAI (3) | 1 |
| 2023 | HiFormer: Hierarchical Multi-scale Representations Using Transformers for Medical Image SegmentationabstractConvolutional neural networks (CNNs) have been the consensus for medical image segmentation tasks. However, they suffer from the limitation in modeling long-range dependencies and spatial correlations due to the nature of convolution operation. Although transformers were first developed to address this issue, they fail to capture low-level features. In contrast, it is demonstrated that both local and global features are crucial for dense prediction, such as segmenting in challenging contexts. In this paper, we propose HiFormer, a novel method that efficiently bridges a CNN and a transformer for medical image segmentation. Specifically, we design two multi-scale feature representations using the seminal Swin Transformer module and a CNN-based encoder. To secure a fine fusion of global and local features obtained from the two aforementioned representations, we propose a Double-Level Fusion (DLF) module in the skip connection of the encoder-decoder structure. Extensive experiments on various medical image segmentation datasets demonstrate the effectiveness of HiFormer over other CNN-based, transformer-based, and hybrid methods in terms of computational complexity, quantitative and qualitative results. Our code is publicly available at GitHub. Moein Heidari, Amirhossein Kazerouni, Milad Soltany Kadarvish, Reza Azad, Ehsan Khodapanah Aghdam, Julien Cohen-Adad, Dorit Merhof |
WACV | 4 |
| 2023 | SegPC-2021: A challenge & dataset on segmentation of Multiple Myeloma plasma cells from microscopic images
Anubha Gupta, Shiv Gehlot, Shubham Goswami, Sachin Motwani, Álvaro García-Faura, Dejan Stepec, Tomaz Martincic, Reza Azad, Dorit Merhof, Afshin Bozorgpour, Babak Azad, Alaa Sulaiman, Deepanshu Pandey, Pradyumna Gupta, Sumit Bhattacharya, Aman Sinha 0002, Xinyun Qiu, Yoonbeom Park, Dae-Hong Lee, Joon Sik Park, KwangYeol Lee, Jaehyung Ye |
Medical Image Anal. | 9 |
| 2023 | Diffusion models in medical imaging: A comprehensive survey
Amirhossein Kazerouni, Ehsan Khodapanah Aghdam, Moein Heidari, Reza Azad, Mohsen Fayyaz, Ilker Hacihaliloglu, Dorit Merhof |
Medical Image Anal. | 4 |
| 2021 | On the Texture Bias for Few-Shot CNN SegmentationabstractDespite the initial belief that Convolutional Neural Networks (CNNs) are driven by shapes to perform visual recognition tasks, recent evidence suggests that texture bias in CNNs provides higher performing models when learning on large labeled training datasets. This contrasts with the perceptual bias in the human visual cortex, which has a stronger preference towards shape components. Perceptual differences may explain why CNNs achieve human-level performance when large labeled datasets are available, but their performance significantly degrades in low-labeled data scenarios, such as few-shot semantic segmentation. To remove the texture bias in the context of few-shot learning, we propose a novel architecture that integrates a set of Difference of Gaussians (DoG) to attenuate high-frequency local components in the feature space. This produces a set of modified feature maps, whose high-frequency components are diminished at different standard deviation values of the Gaussian distribution in the spatial domain. As this results in multiple feature maps for a single image, we employ a bi-directional convolutional long-short-term-memory to efficiently merge the multi scale-space representations. We perform extensive experiments on three well-known few-shot segmentation benchmarks -Pascal i5, COCO-20i and FSS-1000- and demonstrate that our method outperforms state-of-the-art approaches in two datasets under the same conditions. Reza Azad, Abdur Razzaq Fayjie, Claude Kauffmann, Ismail Ben Ayed, Marco Pedersoli, Jose Dolz |
WACV | 1 |
| 2019 | Dynamic 3D Hand Gesture Recognition by Learning Weighted Depth Motion MapsabstractHand gesture recognition (HGR) from sequences of depth maps is a challenging computer vision task because of the low inter-class and high intra-class variability, different execution rates of each gesture, and the high articulated nature of the human hand. In this paper, a multilevel temporal sampling (MTS) method is first proposed that is based on the motion energy of keyframes of depth sequences. As a result, long, middle, and short sequences are generated that contain the relevant gesture information. The MTS results in increasing the intra-class similarity while raising the inter-class dissimilarities. The weighted depth motion map (WDMM) is then proposed to extract the spatiotemporal information from generated summarized sequences by an accumulated weighted absolute difference of consecutive frames. The histogram of gradient and local binary pattern are exploited to extract features from WDMM. The obtained results define the current state-of-the-art on three public benchmark datasets of: MSR Gesture 3D, SKIG, and MSR Action 3D, for 3D HGR. We also achieve competitive results on NTU action dataset. Reza Azad, Maryam Asadi-Aghbolaghi, Shohreh Kasaei, Sergio Escalera |
IEEE Trans. Circuits Syst. Video Technol. | 1 |