EDBT 2026 Demo / reviewers in the wild / expert
Morteza Ghahremani
dblp:152/6299
· DBLP profile ↗
15ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0001-6423-6475ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 56% Deep learning architectures and training · 23% 3D vision · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 90% Geometric modeling and processing · 10% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model › text-to-image generation
attribute binding |
0.9 | 1 | 2025 | Box It to Bind It: Unified Layout Control and Attribute Binding in Text-to-Image Diffusion Models · IEEE Trans. Multim. 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.9 | 1 | 2025 | Box It to Bind It: Unified Layout Control and Attribute Binding in Text-to-Image Diffusion Models · IEEE Trans. Multim. 2025 |
Machine learning › Generative modeling › diffusion model
controllable generation |
0.8 | 1 | 2024 | Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation · NeurIPS 2024 |
Computer vision › 3D vision
image registration |
0.8 | 1 | 2024 | H-ViT: A Hierarchical Vision Transformer for Deformable Image Registration · CVPR 2024 |
Machine learning › Generative modeling › conditional generative model
pose-guided generation |
0.8 | 1 | 2024 | Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation · NeurIPS 2024 |
Medical and health informatics › medical imaging › medical image analysis › image registration
deformable image registration |
0.8 | 1 | 2024 | H-ViT: A Hierarchical Vision Transformer for Deformable Image Registration · CVPR 2024 |
Medical and health informatics › medical imaging
medical image analysis |
0.8 | 1 | 2024 | H-ViT: A Hierarchical Vision Transformer for Deformable Image Registration · CVPR 2024 |
Machine learning › Deep learning architectures and training › normalization
batch normalization |
0.7 | 1 | 2023 | RegBN: Batch Normalization of Multimodal Data with Regularization · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › normalization
feature normalization |
0.7 | 1 | 2023 | RegBN: Batch Normalization of Multimodal Data with Regularization · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.7 | 1 | 2023 | RegBN: Batch Normalization of Multimodal Data with Regularization · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
normalization |
0.7 | 1 | 2023 | RegBN: Batch Normalization of Multimodal Data with Regularization · NeurIPS 2023 |
Image and video processing
feature detection |
0.5 | 1 | 2021 | FFD: Fast Feature Detector · IEEE Trans. Image Process. 2021 |
Image and video processing › feature detection
scale-space feature extraction |
0.5 | 1 | 2021 | FFD: Fast Feature Detector · IEEE Trans. Image Process. 2021 |
Computer vision › 3D vision
point cloud processing |
0.4 | 1 | 2020 | Orderly Disorder in Point Cloud Domain · ECCV (28) 2020 |
Image and video processing › image matching
keypoint matching |
0.1 | 1 | 2021 | FFD: Fast Feature Detector · IEEE Trans. Image Process. 2021 |
Geometric modeling and processing › point cloud processing
point cloud analysis |
0.1 | 1 | 2020 | Orderly Disorder in Point Cloud Domain · ECCV (28) 2020 |
Methods — techniques the papers use, named apart from their topics
vision transformer · 2.3self-attention · 1.5cross-attention · 1.5training-free guidance · 0.9latent diffusion model · 0.9attention masking · 0.8adapter · 0.8frobenius-norm regularization · 0.7batch normalization · 0.7undecimated wavelet transform · 0.5difference of gaussian · 0.5cubic spline · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiaMond: Dementia Diagnosis with Multi-Modal Vision Transformers Using MRI and PETabstractDiagnosing dementia, particularly for Alzheimer's Disease (AD) and frontotemporal dementia (FTD), is complex due to overlapping symptoms. While magnetic res-onance imaging (MRI) and positron emission tomography (PET) data are critical for the diagnosis, integrating these modalities in deep learning faces challenges, often resulting in suboptimal performance compared to using single modalities. Moreover, the potential of multi-modal approaches in differential diagnosis, which holds significant clinical importance, remains largely unexplored. We propose a novel framework, DiaMond, to address these is-sues with vision Transformers to effectively integrate MRI and PET. DiaMond is equipped with self-attention and a novel bi-attention mechanism that synergistically combine MRI and PET, alongside a multi-modal normalization to reduce redundant dependency, thereby boosting the performance. DiaMond significantly outperforms existing multi-modal methods across various datasets, achieving a balanced accuracy of 92.4% in AD diagnosis, 65.2% for AD-MCI-CN classification, and 76.5% in differential diagnosis of AD and FTD. We also validated the robustness of Dia-Mond in a comprehensive ablation study. The code is avail-able at https://github.com/ai-med/DiaMond. Morteza Ghahremani, Youssef Wally, Christian Wachinger |
WACV | 2 |
| 2025 | Organ-DETR: Organ Detection via TransformersabstractQuery-based Transformers have been yielding impressive performance in object localization and detection tasks. However, their application to organ detection in 3D medical imaging data has been relatively unexplored. This study introduces Organ-DETR, featuring two innovative modules, MultiScale Attention (MSA) and Dense Query Matching (DQM), designed to enhance the performance of Detection Transformers (DETRs) for 3D organ detection. MSA is a novel top-down representation learning approach for efficiently encoding Computed Tomography (CT) features. This architecture employs a multiscale attention mechanism, utilizing both dual self-attention and cross-scale attention mechanisms to extract intra- and inter-scale spatial interactions in the attention mechanism. Organ-DETR also introduces DQM, an approach for one-to-many matching that tackles the label assignment difficulties in organ detection. DQM increases positive queries to enhance both recall scores and training efficiency without the need for additional learnable parameters. Extensive results on five 3D CT datasets indicate that the proposed Organ-DETR outperforms comparable techniques by achieving a remarkable improvement of +10.6 mAP COCO. The project and code are available at https://github.com/ai-med/OrganDETR. Morteza Ghahremani, Benjamin Raphael Ernhofer, Marcus R. Makowski, Christian Wachinger |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Box It to Bind It: Unified Layout Control and Attribute Binding in Text-to-Image Diffusion ModelsabstractWhile latent diffusion models (LDMs) excel at creating imaginative images, they often lack precision in semantic fidelity and spatial control over where objects are generated. To address these deficiencies, we introduce the Box-it-to-Bind-it (B2B) module—a novel, training-free approach for improving spatial control and semantic accuracy in text-to-image (T2I) diffusion models. B2B targets three key challenges in T2I: catastrophic neglect, attribute binding, and layout guidance. The process encompasses two main steps: (i)Object generation, which adjusts the latent encoding to guarantee object generation and directs it within specified bounding boxes, and (ii)Attribute binding, ensuring that generated objects adhere to their specified attributes in the prompt. B2B is designed as a compatible plug-and-play module for existing T2I models like Stable Diffusion and Gligen, markedly enhancing models’ performance in addressing these key challenges. We assess our technique on the well-established CompBench and TIFA score benchmarks, and HRS dataset where B2B not only surpasses methods specialized in either attribute binding or layout guidance but also uniquely excels by integrating these capabilities to deliver enhanced overall performance. Ashkan Taghipour, Morteza Ghahremani, Mohammed Bennamoun, Aref Miri Rekavandi, Hamid Laga, Farid Boussaïd |
IEEE Trans. Multim. | 2 |
| 2024 | H-ViT: A Hierarchical Vision Transformer for Deformable Image RegistrationabstractThis paper introduces a novel top-down representation approach for deformable image registration, which estimates the deformation field by capturing various short-and long-range flow features at different scale levels. As a Hierarchical Vision Transformer (H- ViT), we propose a dual self-attention and cross-attention mechanism that uses high-level features in the deformation field to represent low-level ones, enabling information streams in the deformation field across all voxel patch embeddings irrespective of their spatial proximity. Since high-level features contain abstract flow patterns, such patterns are expected to effectively contribute to the representation of the deformation field in lower scales. When the self-attention module utilizes within-scale short-range patterns for representation, the cross-attention modules dynamically look for the key tokens across different scales to further interact with the local query voxel patches. Our method shows superior accuracy and visual quality over the state-of-the-art registration methods in five publicly available datasets, highlighting a substantial enhancement in the performance of medical imaging registration. The project link is available at https://mogvision.github.io/hvit. Morteza Ghahremani, Mohammad Khateri, Bailiang Jian, Benedikt Wiestler, Ehsan Adeli-Mosabbeb, Christian Wachinger |
CVPR | 1 |
| 2024 | Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image GenerationabstractControllable text-to-image (T2I) diffusion models have shown impressive performance in generating high-quality visual content through the incorporation of various conditions. Current methods, however, exhibit limited performance when guided by skeleton human poses, especially in complex pose conditions such as side or rear perspectives of human figures. To address this issue, we present Stable-Pose, a novel adapter model that introduces a coarse-to-fine attention masking strategy into a vision Transformer (ViT) to gain accurate pose guidance for T2I models. Stable-Pose is designed to adeptly handle pose conditions within pre-trained Stable Diffusion, providing a refined and efficient way of aligning pose representation during image synthesis. We leverage the query-key self-attention mechanism of ViTs to explore the interconnections among different anatomical parts in human pose skeletons. Masked pose images are used to smoothly refine the attention maps based on target pose-related features in a hierarchical manner, transitioning from coarse to fine levels.
Additionally, our loss function is formulated to allocate increased emphasis to the pose region, thereby augmenting the model's precision in capturing intricate pose details. We assessed the performance of Stable-Pose across five public datasets under a wide range of indoor and outdoor human pose scenarios. Stable-Pose achieved an AP score of 57.1 in the LAION-Human dataset, marking around 13\% improvement over the established technique ControlNet. The project link and code is available at https://github.com/ai-med/StablePose. Morteza Ghahremani, Björn Ommer, Christian Wachinger |
NeurIPS | 2 |
| 2023 | RegBN: Batch Normalization of Multimodal Data with RegularizationabstractRecent years have witnessed a surge of interest in integrating high-dimensional data captured by multisource sensors, driven by the impressive success of neural networks in integrating multimodal data. However, the integration of heterogeneous multimodal data poses a significant challenge, as confounding effects and dependencies among such heterogeneous data sources introduce unwanted variability and bias, leading to suboptimal performance of multimodal models. Therefore, it becomes crucial to normalize the low- or high-level features extracted from data modalities before their fusion takes place. This paper introduces RegBN, a novel approach for multimodal Batch Normalization with REGularization. RegBN uses the Frobenius norm as a regularizer term to address the side effects of confounders and underlying dependencies among different data sources. The proposed method generalizes well across multiple modalities and eliminates the need for learnable parameters, simplifying training and inference. We validate the effectiveness of RegBN on eight databases from five research areas, encompassing diverse modalities such as language, audio, image, video, depth, tabular, and 3D MRI. The proposed method demonstrates broad applicability across different architectures such as multilayer perceptrons, convolutional neural networks, and vision transformers, enabling effective normalization of both low- and high-level features in multimodal neural networks. RegBN is available at https://mogvision.github.io/RegBN. Morteza Ghahremani, Christian Wachinger |
NeurIPS | 1 |
| 2023 | Regional Attention Network (RAN) for Head Pose and Fine-Grained Gesture RecognitionabstractAffect is often expressed via non-verbal body language such as actions/gestures, which are vital indicators for human behaviors. Recent studies on recognition of fine-grained actions/gestures in monocular images have mainly focused on modeling spatial configuration of body parts representing body pose, human-objects interactions and variations in local appearance. The results show that this is a brittle approach since it relies on accurate body parts/objects detection. In this work, we argue that there exist local discriminative semantic regions, whose “informativeness” can be evaluated by the attention mechanism for inferring fine-grained gestures/actions. To this end, we propose a novel end-to-endregional attention network (RAN), which is a fully convolutional neural network (CNN) to combine multiple contextual regions through attention mechanism, focusing on parts of the images that are most relevant to a given task. Our regions consist of one or more consecutive cells and are adapted from the strategies used in computing HOG (Histogram of Oriented Gradient) descriptor. The model is extensively evaluated on ten datasets belonging to 3 different scenarios: 1) head pose recognition, 2) drivers state recognition, and 3) human action and facial expression recognition. The proposed approach outperforms the state-of-the-art by a considerable margin in different metrics. Ardhendu Behera, Zachary Wharton, Yonghuai Liu, Morteza Ghahremani, Swagat Kumar, Nik Bessis |
IEEE Trans. Affect. Comput. | 4 |
| 2021 | Interwoven texture-based description of interest points in images
Morteza Ghahremani, Yitian Zhao, Bernard Tiddeman, Yonghuai Liu |
Pattern Recognit. | 1 |
| 2021 | FFD: Fast Feature DetectorabstractScale-invariance, good localization and robustness to noise and distortions are the main properties that a local feature detector should possess. Most existing local feature detectors find excessive unstable feature points that increase the number of keypoints to be matched and the computational time of the matching step. In this paper, we show that robust and accurate keypoints exist in the specific scale-space domain. To this end, we first formulate the superimposition problem into a mathematical model and then derive a closed-form solution for multiscale analysis. The model is formulated via difference-of-Gaussian (DoG) kernels in the continuous scale-space domain, and it is proved that setting the scale-space pyramid's blurring ratio and smoothness to 2 and 0.627, respectively, facilitates the detection of reliable keypoints. For the applicability of the proposed model to discrete images, we discretize it using the undecimated wavelet transform and the cubic spline function. Theoretically, the complexity of our method is less than 5% of that of the popular baseline Scale Invariant Feature Transform (SIFT). Extensive experimental results show the superiority of the proposed feature detector over the existing representative hand-crafted and learning-based techniques in accuracy and computational time. The code and supplementary materials can be found at https://github.com/mogvision/FFD. Morteza Ghahremani, Yonghuai Liu, Bernard Tiddeman |
IEEE Trans. Image Process. | 1 |
| 2020 | Orderly Disorder in Point Cloud Domain
Morteza Ghahremani, Bernard Tiddeman, Yonghuai Liu, Ardhendu Behera |
ECCV (28) | 1 |
| 2020 | Cycle Structure and Illumination Constrained GAN for Medical Image Enhancement
Yuhui Ma, Yonghuai Liu, Jun Cheng 0003, Yalin Zheng, Morteza Ghahremani, Honghan Chen, Jiang Liu 0001, Yitian Zhao |
MICCAI (2) | 5 |
| 2020 | Temporal convolutional neural (TCN) network for an effective weather forecasting using time-series data from the local weather stationabstractAbstract Non-predictive or inaccurate weather forecasting can severely impact the community of users such as farmers. Numerical weather prediction models run in major weather forecasting centers with several supercomputers to solve simultaneous complex nonlinear mathematical equations. Such models provide the medium-range weather forecasts, i.e., every 6 h up to 18 h with grid length of 10–20 km. However, farmers often depend on more detailed short-to medium-range forecasts with higher-resolution regional forecasting models. Therefore, this research aims to address this by developing and evaluating a lightweight and novel weather forecasting system, which consists of one or more local weather stations and state-of-the-art machine learning techniques for weather forecasting using time-series data from these weather stations. To this end, the system explores the state-of-the-art temporal convolutional network (TCN) and long short-term memory (LSTM) networks. Our experimental results show that the proposed model using TCN produces better forecasting compared to the LSTM and other classic machine learning approaches. The proposed model can be used as an efficient localized weather forecasting tool for the community of users, and it could be run on a stand-alone personal computer. Pradeep Hewage, Ardhendu Behera, Marcello Trovati, Ella Grishikashvili Pereira, Morteza Ghahremani, Francesco Palmieri 0002, Yonghuai Liu |
Soft Comput. | 5 |
| 2016 | Nonlinear IHS: A Promising Method for Pan-SharpeningabstractThe goal of pan-sharpening is to increase the spatial resolution of multispectral or hyperspectral images using a high-spatial-resolution panchromatic image. The intensity-hue-saturation (IHS) method is one of the most popular pan-sharpening methods. The pan-sharpening framework of the IHS method is simple, efficient, and of high spatial quality. In this letter, we propose a nonlinear IHS method. The main concern of this letter is how to accurately estimate the intensity component. The proposed method approximates the intensity component via local and global synthesis approaches. The WorldView-2 and DEIMOS-2 data are used to evaluate the performance of the proposed method. The experimental results demonstrate that the proposed nonlinear IHS method generates high-quality pan-sharpened multispectral bands in terms of quantitatively and perceptually. Morteza Ghahremani, Hassan Ghassemian |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | A Compressed-Sensing-Based Pan-Sharpening Method for Spectral Distortion ReductionabstractRecently, the compressed sensing (CS) theory has become an interesting topic for pan-sharpening of multispectral images. The CS theory ensures that, under the sparsity regularization, an unknown sparse signal can be exactly recovered from a drastically smaller number of linear measurements. In this paper, we propose a CS-based approach for fusion of the multispectral and panchromatic satellite images. The contribution of this paper is twofold. First, with the spatial and spectral characteristics of the satellite images, we assume that each patch of the unknown high spatial resolution intensity (HRI) component can be represented as a linear combination of atoms in a dictionary trained only from the panchromatic image; thus, the problem of generating an optimal dictionary is solved. Second, we propose an iterative algorithm to obtain the sparsest coefficients. The sparsest coefficients ensure that the estimated HRI component can be correctly recovered from the panchromatic image. The IKONOS, QuickBird, and WorldView-2 data are used to evaluate the performance of the proposed method. The experimental results demonstrate that the proposed method generates high-quality pan-sharpened multispectral bands quantitatively and perceptually. Morteza Ghahremani, Hassan Ghassemian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Remote Sensing Image Fusion Using Ripplet Transform and Compressed SensingabstractIn this letter, we propose a novel remote sensing image fusion method based on the ripplet transform and the compressed sensing (CS) theory. The ripplet transform generalizes the curvelet transform by adding two parameters, namely, support c and degree d. These parameters provide the ripplet transform with anisotropy capability of representing singularities along arbitrarily shaped curves, and the curvelet transform is just a special case of the ripplet transform with c=1 and d=2. In the proposed method, the spatial details are first extracted from the PAN image by means of ripplets, and then, they are injected into the MS bands by the proposed injection model named CS-based injection. The aim of this model, which is based on the CS theory, is to minimize the spectral distortion in the pan-sharpened MS bands with respect to the original ones. The experimental results carried out on IKONOS and QuickBird data sets demonstrate that the proposed method provides better fused images in terms of the visual and quantitative evaluations. Morteza Ghahremani, Hassan Ghassemian |
IEEE Geosci. Remote. Sens. Lett. | 1 |