EDBT 2026 Demo / reviewers in the wild / expert
Maryam Haghighat
dblp:192/5015
· DBLP profile ↗
16ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-2080-8483ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Effective and Robust Multimodal Medical Image AnalysisabstractMultimodal Fusion Learning ( MFL ), leveraging disparate data from various imaging modalities (e.g., MRI, CT, SPECT ), has shown great potential for addressing medical problems such as skin cancer and brain tumor prediction. However, existing MFL methods face three key limitations: a) they often specialize in specific modalities, and overlooks effective shared complementary information across diverse modalities, hence limiting their generalizability for multi-disease analysis; b) they rely on computationally expensive models, restricting their applicability in resource-limited settings; and c) they lack robustness against adversarial attacks, compromising reliability in medical AI applications. To address these limitations, we propose a novel Multi-Attention Integration Learning ( MAIL ) network, incorporating two key components: a) an efficient residual learning attention block for capturing refined modality-specific multi-scale patterns and b) an efficient multimodal cross-attention module for learning enriched complementary shared representations across diverse modalities. Furthermore, to ensure adversarial robustness, we extend MAIL network to design Robust-MAIL by incorporating random projection filters and modulated attention noise. Extensive evaluations on 20 public datasets show that both MAIL and Robust-MAIL outperform existing methods, achieving performance gains of up to 9.34% while reducing computational costs by up to 78.3%. These results highlight the superiority of our approaches, ensuring more reliable predictions than top competitors. Joy Dhar, Nayyar Abbas Zaidi, Maryam Haghighat |
KDD (1) | 3 |
| 2026 | HyPCA-Net: Advancing Multimodal Fusion in Medical Image AnalysisabstractMultimodal fusion frameworks, which integrate diverse medical imaging modalities (e.g., MRI, CT), have shown great potential in applications such as skin cancer detection, dementia diagnosis, and brain tumor prediction. However, existing multimodal fusion methods face significant challenges. First, they often rely on computationally expensive models, limiting their applicability in low-resource environments. Second, they often employ cascaded attention modules, which potentially increase risk of information loss during inter-module transitions and hinder their capacity to effectively capture robust shared representations across modalities. This restricts their generalization in multi-disease analysis tasks. To address these limitations, we propose a Hybrid Parallel-Fusion Cascaded Attention Network (HyPCA-Net), composed of two core novel blocks: (a) a computationally efficient residual adap tive learning attention block for capturing refined modality-specific representations, and (b) a dual-view cascaded attention block aimed at learning robust shared representations across diverse modalities. Extensive experiments on ten publicly available datasets exhibit that HyPCA-Netsignificantly outperforms existing leading methods, with improvements of up to 5.2% in performance and reductions of up to 73.1% in computational cost. Code: https://github.com/misti1203/HyPCA-Net. Joy Dhar, Manish Kumar Pandey, Debashis Das Chakladar, Maryam Haghighat, Azadeh Alavi, Sajib Mistry, Nayyar Abbas Zaidi |
WACV | 4 |
| 2026 | Intra-Class Probabilistic Embeddings for Uncertainty Estimation in Vision-Language ModelsabstractVision-language models (VLMs), such as CLIP, have gained popularity for their strong open vocabulary classification performance, but they are prone to assigning high confidence scores to misclassifications, limiting their reliability in safety-critical applications. We introduce a training-free, post-hoc uncertainty estimation method for contrastive VLMs that can be used to detect erroneous predictions. The key to our approach is to measure visual feature consistency within a class, using feature projection combined with multivariate Gaussians to create class-specific probabilistic embeddings. Our method is VLM-agnostic, requires no fine-tuning, demonstrates robustness to distribution shift, and works effectively with as few as 10 training images per class. Extensive experiments on ImageNet, Flowers102, Food101, EuroSAT and DTD show state-of-the-art error detection performance, significantly outperforming both deterministic and probabilistic VLM baselines. Code is available at https://github.com/zhenxianglin/ICPE. Zhenxiang Lin, Maryam Haghighat, Will Browne, Dimity Miller |
WACV | 2 |
| 2025 | HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial ViewsabstractWe present HOTFormerLoc, a novel and versatile Hierarchical Octree-based TransFormer, for large-scale 3D place recognition in both ground-to-ground and ground-to-aerial scenarios across urban and forest environments. We propose an octree-based multi-scale attention mechanism that captures spatial and semantic features across granularities. To address the variable density of point distributions from spinning lidar, we present cylindrical octree attention windows to reflect the underlying distribution during attention. We introduce relay tokens to enable efficient global-local interactions and multi-scale representation learning at reduced computational cost. Our pyramid attentional pooling then synthesises a robust global descriptor for end-to-end place recognition in challenging environments. In addition, we introduce CS-WildPlaces, a novel 3D cross-source dataset featuring point cloud data from aerial and ground lidar scans captured in dense forests. Point clouds in CS-Wild-Places contain representational gaps and distinctive attributes such as varying point densities and noise patterns, making it a challenging benchmark for cross-view localisation in the wild. HOTFormerLoc achieves a top-1 average recall improvement of 5.5% – 11.5% on the CS-Wild-Places benchmark. Furthermore, it consistently outperforms SOTA 3D place recognition methods, with an average performance gain of 4.9% on well-established urban and forest datasets. The code and CS-Wild-Places benchmark is available at https://csirorobotics.github.io/HOTFormerLoc. Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Milad Ramezani |
CVPR | 2 |
| 2025 | Neural Network Based Current Harmonic Predictions Using PCC Voltage Measurements in Distribution NetworksabstractConventional current harmonic prediction techniques exhibit significant limitations at a system level in distribution networks due to multivariable, complex, and time-varying models under variable operating conditions. This paper proposes a novel neural network-based technique for predicting low-order current harmonics using voltage harmonics at the point of common coupling, a practical and accurate approach. The proposed model employs voltage harmonic data at the point of common coupling, which are easily measurable in practice, to predict the corresponding amplitude of current harmonics generated by nonlinear loads such as grid-connected inverters. The results demonstrate the model's high accuracy in predicting low-order current harmonics, even under the influence of grid background voltage harmonics, variations in operational power, and grid impedance. The proposed method outperforms other models, such as support vector regression, multiple linear regression, and gaussian process regression, in terms of accuracy and consistency. This work offers a practical and reliable approach for predicting low-order current harmonics at a system level (e.g., distribution networks), contributing to the stable and efficient operation of electrical grids maintaining power quality standards. Rhns Jayathissa, Firuz Zare, Simon Denman, Amir Taghvaie, Maryam Haghighat, Dinesh Kumar 0005 |
IECON | 5 |
| 2025 | Online 6DoF Global Localisation in Forests using Semantically-Guided Re-Localisation and Cross-View Factor-Graph OptimisationabstractThis paper presents FGLoc6D, a novel approach for robust global localisation and online 6DoF pose estimation of ground robots in forest environments by leveraging deep semantically-guided re-localisation and cross-view factor graph optimisation. The proposed method addresses the challenges of aligning aerial and ground data for pose estimation, which is crucial for accurate point-to-point navigation in GPS-degraded environments. By integrating information from both perspectives into a factor graph framework, our approach effectively estimates the robot’s global position and orientation. Additionally, we enhance the repeatability of deep-learned keypoints for metric localisation in forests by incorporating a semantically-guided regression loss. This loss encourages greater attention to wooden structures, e.g., tree trunks, which serve as stable and distinguishable features, thereby improving the consistency of keypoints and increasing the success rate of global registration, a process we refer to as re-localisation. The re-localisation module along with the factor-graph structure, populated by odometry and ground-to-aerial factors over time, allows global localisation under dense canopies. We validate the performance of our method through extensive experiments in three forest scenarios, demonstrating its global localisation capability and superiority over alternative state-of-the-art in terms of accuracy and robustness in these challenging environments. Experimental results show that our proposed method can achieve drift-free localisation with bounded positioning errors, ensuring reliable and safe robot navigation through dense forests. Lucas Carvalho de Lima, Ethan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes, Paulo Vinicius Koerich Borges, Michael Brünig, Milad Ramezani |
IROS | 3 |
| 2025 | Multimodal Fusion Learning with Dual Attention for Medical ImagingabstractMultimodal fusion learning has shown significant promise in classifying various diseases such as skin cancer and brain tumors. However, existing methods face three key limitations. First, they often lack generalizability to other diagnosis tasks due to their focus on a particular disease. Second, they do not fully leverage multiple health records from diverse modalities to learn robust complementary information. And finally, they typically rely on a single attention mechanism, missing the benefits of multiple attention strategies within and across various modalities. To address these issues, this paper proposes a dual robust information fusion attention mechanism (DRIFA) that leverages two attention modules - i.e., multi-branch fusion attention module and the multimodal information fusion attention module. DRIFA can be integrated with any deep neural network, forming a multimodal fusion learning framework denoted as DRIFA-Net. We show that the multi-branch fusion attention of DRIFA learns enhanced representations for each modality, such as dermoscopy, pap smear, MRI, and CT-scan, whereas multimodal information fusion attention module learns more refined multimodal shared representations - improving the network's generalization across multiple tasks and enhancing overall performance. Additionally, to estimate the uncertainty of DRIFA-Net predictions, we have employed an ensemble Monte Carlo dropout strategy. Extensive experiments on five publicly available datasets with diverse modalities demonstrate that our approach consistently outperforms state-of-the-art methods. The code is available at https://github.com/misti1203/DRIFA-Net. Joy Dhar, Nayyar Abbas Zaidi, Maryam Haghighat, Sudipta Roy 0002, Puneet Goyal, Azadeh Alavi |
WACV | 3 |
| 2025 | A Comprehensive Survey on Multi-View Classification: Methods, Applications, and ChallengesabstractMulti-view classification (MVC) has emerged as a promising approach in machine learning, aimed at enhancing classification accuracy by leveraging information from multiple perspectives. As the demand for more robust, interpretable, and effective machine learning models grows, MVC has shown significant progress over the past decade, yet it faces new challenges. Despite extensive literature on this subject, there is a notable absence of a comprehensive synthesis of MVC methods. This article addresses this gap by presenting a thorough overview and classification of MVC methods, categorizing them into seven distinct classes: text, image, time series, hyperspectral, video, signal, and 3D shape. Our meticulous examination within each class highlights advancements and evaluates their applicability in both supervised and semi-supervised learning contexts. Beyond this retrospective analysis, we explore future directions for research and development in this domain. This survey serves as a compendium of existing knowledge and as a guide for future endeavors in MVC, shaping the trajectory of ongoing research and innovation. Kamal Berahmand, Fatemeh Daneshfar, Maryam Rahmaninia, Maryam Haghighat, Mahdi Jalili |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2024 | Pre-training with Random Orthogonal Projection Image ModelingabstractMasked Image Modeling (MIM) is a powerful self-supervised strategy for visual pre-training without the use of labels. MIM applies random crops to input images, processes them with an encoder, and then recovers the masked inputs with a decoder, which encourages the network to capture and learn structural information about objects and scenes. The intermediate feature representations obtained from MIM are suitable for fine-tuning on downstream tasks. In this paper, we propose an Image Modeling framework based on random orthogonal projection instead of binary masking as in MIM. Our proposed Random Orthogonal Projection Image Modeling (ROPIM) reduces spatially-wise token information under guaranteed bound on the noise variance and can be considered as masking entire spatial image area under locally varying masking degrees. Since ROPIM uses a random subspace for the projection that realizes the masking step, the readily available complement of the subspace can be used during unmasking to promote recovery of removed information. In this paper, we show that using random orthogonal projection leads to superior performance compared to crop-based masking. We demonstrate state-of-the-art results on several popular benchmarks. Maryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr Koniusz |
ICLR | 1 |
| 2024 | FactoFormer: Factorized Hyperspectral Transformers With Self-Supervised PretrainingabstractHyperspectral images (HSIs) contain rich spectral and spatial information. Motivated by the success of transformers in the field of natural language processing and computer vision where they have shown the ability to learn long-range dependencies within input data, recent research has focused on using transformers for HSIs. However, current state-of-the-art hyperspectral transformers only tokenize the input HSI sample along the spectral dimension, resulting in the underutilization of spatial information. Moreover, transformers are known to be data-hungry and their performance relies heavily on large-scale pretraining, which is challenging due to limited annotated hyperspectral data. Therefore, the full potential of HSI transformers has not been fully realized. To overcome these limitations, we propose a novel factorized spectral–spatial transformer that incorporates factorized self-supervised pretraining procedures, leading to significant improvements in performance. The factorization of the inputs allows the spectral and spatial transformers to better capture the interactions within the hyperspectral data cubes. Inspired by masked image modeling (MIM) pretraining, we also devise efficient masking strategies for pretraining each of the spectral and spatial transformers. We conduct experiments on six publicly available datasets for the HSI classification task and demonstrate that our model achieves state-of-the-art performance in all the datasets. The code for our model will be made available athttps://github.com/csiro-robotics/FactoFormer. Shaheer Mohamed, Maryam Haghighat, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Peyman Moghadam |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Deep Robust Multi-Robot Re-Localisation in Natural EnvironmentsabstractThe success of re-localisation has crucial implications for the practical deployment of robots operating within a prior map or relative to one another in real-world scenarios. Using single-modality, place recognition and localisation can be compromised in challenging environments such as forests. To address this, we propose a strategy to prevent lidar-based re-localisation failure using lidar-image cross-modality. Our solution relies on self-supervised 2D-3D feature matching to predict alignment and misalignment. Leveraging a deep network for lidar feature extraction and relative pose estimation between point clouds, we train a model to evaluate the estimated transformation. A model predicting the presence of misalignment is learned by analysing image-lidar similarity in the embedding space and the geometric constraints available within the region seen in both modalities in Euclidean space. Experimental results using real datasets (offline and online modes) demonstrate the effectiveness of the proposed pipeline for robust re-localisation in unstructured, natural environments. Milad Ramezani, Ethan Griffiths, Maryam Haghighat, Alex Pitt, Peyman Moghadam |
IROS | 3 |
| 2020 | Rate-Distortion Driven Decomposition of Multiview Imagery to Diffuse and Specular ComponentsabstractIn this work, we propose an overcomplete representation of multiview imagery for the purpose of compression. We present a rate-distortion (R-D) driven approach to decompose multiview datasets into two additive parts which can be interpreted as diffuse and specular content. We choose distinct and different sparsifying transforms for the diffuse and specular components and employ an R-D inspired measure as our optimization cost function to drive the decomposition based solely on compressibility. We first describe a framework which performs data separation in a registered domain to avoid the complexity of warping between views. Then a more comprehensive approach is proposed to separate specular data progressively from coordinates of multiple reference views. Experimental results show a coding gain of up to 0.6 dB for synthetic datasets and up to 0.9 dB for real datasets. Maryam Haghighat, Reji Mathew, David S. Taubman |
IEEE Trans. Image Process. | 1 |
| 2019 | Rate-Distortion Driven Separation of Diffuse and Specular Components in Multiview ImageryabstractIn this work we explore an overcomplete representation of multiview imagery for the purpose of compression. We present a rate-distortion (R-D) driven approach to decompose multiview datasets into two additive parts which can be interpreted as being the diffuse and specular components. We apply different transforms to each component such that the compressibility of input data is improved. We describe a framework which performs the R-D optimized separation in a registered domain to avoid the complexity of warping between views. Experimental results highlight the benefits of the proposed source separation approach in the context of compression. Maryam Haghighat, Reji Mathew, David S. Taubman |
ICIP | 1 |
| 2019 | Illumination Estimation and Compensation of Low Frame Rate Video Sequences for Wavelet-Based Video CompressionabstractIn this paper, we are interested in the compression of image sets or video with considerable changes in illumination. We develop a framework to decompose frames into illumination fields and texture in order to achieve sparser representations of frames which is beneficial for compression. Illumination variations or contrast ratio factors among frames are described by a full resolution multiplicative field. First, we propose a Lifting-based Illumination Adaptive Transform (LIAT) framework which incorporates illumination compensation to temporal wavelet transforms. We estimate a full resolution illumination field, taking heed of its spatial sparsity by a rate-distortion (R-D) driven framework. An affine mesh model is also developed as a point of comparison. We find the operational coding cost of the subband frames by modeling a typical t + 2D wavelet video coding system. While our general findings on R-D optimization are applicable to a range of coding frameworks, in this paper, we report results based on employing JPEG 2000 coding tools. The experimental results highlight the benefits of the proposed R-D driven illumination estimation and compensation in comparison with alternative scalable coding methods and non-scalable coding schemes of AVC and HEVC employing weighted prediction. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 1 |
| 2018 | Rate-Distortion Optimized Illumination Estimation for Wavelet-Based Video CodingabstractWe propose a rate-distortion optimized framework for estimating illumination changes (lighting variations, fade in/out effects) in a highly scalable coding system. Illumination variations are realized using multiplicative factors in the image domain and are estimated considering the coding cost of the illumination field and input frames which are first subject to a temporal Lifting-based Illumination Adaptive Transform (LIAT). The coding cost is modelled by an ℓ1-norm optimization problem which is derived to approximate a quadratic-log function which emerges from rate-distortion considerations. The optimization problem is solved using ADMM. The proposed solution works the same or better than a mesh-based approach proposed in prior work, where sparsity was controlled by explicitly choosing mesh parameters. In the compression-inspired formulation presented here, sparsity is discovered automatically through the solution of a convex program that depends only on a target rate-distortion operating point. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, Sean I. Young, David S. Taubman |
ICASSP | 1 |
| 2017 | Lifting-based Illumination Adaptive Transform (LIAT) using mesh-based illumination modellingabstractState-of-the-art video coding techniques employ block-based illumination compensation to improve coding efficiency. In this work, we propose a Lifting-based Illumination Adaptive Transform (LIAT) to exploit temporal redundancy among frames that have illumination variations, such as the frames of low frame rate video or multi-view video. LIAT employs a mesh-based spatially affine model to represent illumination variations between two frames. In LIAT, transformed frames are jointly compressed, together with illumination information, into a layered rate-distortion optimal codestream, using the JPEG2000 format. We show that the LIAT framework significantly improves compression efficiency of temporal subband transforms for both predictive and more general transforms with predict and update steps. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, David S. Taubman |
ICIP | 1 |