VLDB 2026 Research / reviewers in the wild / expert
Ola Ahmad
dblp:99/10407 · also Ola Suleiman Ahmad
· DBLP profile ↗
14ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-6854-9866ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Network-agnostic distortion-robust projections for wide-angle image understandingabstractDue to their increased field of view, wide-angle lenses are increasingly used in applications such as VR, security, or in autonomous driving. Typically, existing models either ignore wide-angle distortions or "undistort" images to a perspective projection, often resulting in severe stretching. More recent distortion-aware architectures address these issues, yet they impose substantial computational burdens and limit the use of powerful pre-trained vision backbones. In this work, we revisit the undistortion strategy by exploring alternative projection functions beyond the conventional perspective model. Specifically, we investigate square-to-disc mapping functions, most notably, the elliptical grid map (EGM) projection, which minimizes stretching. We show how EGM projection can be combined with the known lens distortion curves to achieve distortion invariance directly in image space. This network-agnostic approach seamlessly integrates with existing deep learning architectures, allowing fine-tuning on pre-trained models trained on large perspective datasets, while adapting to both seen and unseen wide-angle lenses without re-training each time the lens changes during evaluation. We perform experiments on the semantic segmentation task, comparing methods on zero-shot adaptation to unseen lenses from different wide-angle lenses. Our extensive experiments show that using the EGM projection with existing segmentation models significantly outperforms baselines when trained on bounded distortion levels and tested across both seen and out-of-distribution distortions. Furthermore, EGM projection achieves improved performance on real-world datasets, highlighting the robustness and practicality of our approach in real-world applications. The code and models are publicly available at https://lvsn.github.io/NADA/. Akshaya Athwale, Ola Ahmad, Jean-François Lalonde |
WACV | 2 |
| 2025 | Promi: an Efficient Prototype-Mixture Baseline for Few-Shot Segmentation with Bounding-Box AnnotationsabstractIn robotics applications, few-shot segmentation is crucial because it allows robots to perform complex tasks with minimal training data, facilitating their adaptation to diverse, real-world environments. However, pixel-level annotations of even small amount of images is highly time-consuming and costly. In this paper, we present a novel few-shot binary segmentation method based on bounding-box annotations instead of pixel-level labels. We introduce, ProMi, an efficient prototype-mixture-based method that treats the background class as a mixture of distributions. Our approach is simple, training-free, and effective, accommodating coarse annotations with ease. Compared to existing baselines, ProMi achieves the best results across different datasets with significant gains, demonstrating its effectiveness. Furthermore, we present qualitative experiments tailored to real-world mobile robot tasks, demonstrating the applicability of our approach in such scenarios. Our code: https://github.com/ThalesGroup/promi. Florent Chiaroni, Ali Ayub, Ola Ahmad |
ICRA | 3 |
| 2025 | DarSwin-Unet: Distortion Aware ArchitectureabstractWide angle fisheye images are becoming increasingly common for perception tasks in applications such as robotics, security, and mobility (e.g. drones, avionics). However, current models often either ignore the distortions in wide angle images or are not suitable to perform pixel-level tasks. In this paper, we present an encoder-decoder model based on a radial transformer architecture that adapts to distortions in wide angle lenses by leveraging the physical character-istics defined by the radial distortion profile. In contrast to the original model, which only performs classification tasks, we introduce a U-Net architecture, DarSwin-Unet, designed for pixel level tasks. Furthermore, we propose a novel strategy that minimizes sparsity when sampling the image for creating its input tokens. Our approach enhances the model capability to handle pixel-level tasks in wide angle fisheye images, making it more effective for real-world applications. Compared to other baselines, DarSwin-Unet achieves the best results across different datasets, with significant gains when trained on bounded levels of distortions (very low, low, medium, and high) and tested on all, including out-of-distribution distortions. We demonstrate its performance on depth estimation and show through extensive experiments that DarSwin-Unet can perform zero-shot adaptation to unseen distortions of different wide angle lenses. The code and models are publicly available at https://lvsn.github.io/darswin-unet/. Akshaya Athwale, Ichrak Shili, Émile Bergeron, Ola Ahmad, Jean-François Lalonde |
WACV | 4 |
| 2024 | Randomized Confidence Bounds for Stochastic Partial MonitoringabstractThe partial monitoring (PM) framework provides a theoretical formulation of sequential learning problems with incomplete feedback. At each round, a learning agent plays an action while the environment simultaneously chooses an outcome. The agent then observes a feedback signal that is only partially informative about the (unobserved) outcome. The agent leverages the received feedback signals to select actions that minimize the (unobserved) cumulative loss. In contextual PM, the outcomes depend on some side information that is observable by the agent before selecting the action. In this paper, we consider the contextual and non-contextual PM settings with stochastic outcomes. We introduce a new class of PM strategies based on the randomization of deterministic confidence bounds. We also extend regret guarantees to settings where existing stochastic strategies are not applicable. Our experiments show that the proposed RandCBP and RandCBPside* strategies have competitive performance against state-of-the-art baselines in multiple PM games. To illustrate how the PM framework can benefit real world applications, we design a use case on the real-world problem of monitoring the error rate of any deployed classification system. Maxime Heuillet, Ola Ahmad, Audrey Durand |
ICML | 2 |
| 2024 | Neural Active Learning Meets the Partial Monitoring FrameworkabstractWe focus on the online-based active learning (OAL) setting where an agent operates over a stream of observations and trades-off between the costly acquisition of information (labelled observations) and the cost of prediction errors. We propose a novel foundation for OAL tasks based on partial monitoring, a theoretical framework specialized in online learning from partially informative actions. We show that previously studied binary and multi-class OAL tasks are instances of partial monitoring. We expand the real-world potential of OAL by introducing a new class of cost-sensitive OAL tasks. We propose NeuralCBP, the first PM strategy that accounts for predictive uncertainty with deep neural networks. Our extensive empirical evaluation on open source datasets shows that NeuralCBP has competitive performance against state-of-the-art baselines on multiple binary, multi-class and cost-sensitive OAL tasks. Maxime Heuillet, Ola Ahmad, Audrey Durand |
UAI | 2 |
| 2024 | Causal Analysis for Robust Interpretability of Neural NetworksabstractInterpreting the inner function of neural networks is crucial for the trustworthy development and deployment of these black-box models. Prior interpretability methods focus on correlation-based measures to attribute model decisions to individual examples. However, these measures are susceptible to noise and spurious correlations encoded in the model during the training phase (e.g., biased inputs, model overfitting, or misspecification). Moreover, this process has proven to result in noisy and unstable attributions that prevent any transparent understanding of the model’s behavior. In this paper, we develop a robust interventional-based method grounded by causal analysis to capture cause-effect mechanisms in pre-trained neural networks and their relation to the prediction. Our novel approach relies on path interventions to infer the causal mechanisms within hidden layers and isolate relevant and necessary information (to model prediction), avoiding noisy ones. The result is task-specific causal explanatory graphs that can audit model behavior and express the actual causes underlying its performance. We apply our method to vision models trained on classification tasks. On image classification tasks, we provide extensive quantitative experiments to show that our approach can capture more stable and faithful explanations than standard attribution-based methods. Furthermore, the underlying causal graphs express the neural interactions in the model, making it a valuable tool in other applications (e.g., model repair). Ola Ahmad, Nicolas Béreux, Loïc Baret, Vahid Hashemi, Freddy Lécué |
WACV | 1 |
| 2024 | MoP-CLIP: A Mixture of Prompt-Tuned CLIP Models for Domain Incremental LearningabstractDespite the recent progress in incremental learning, addressing catastrophic forgetting under distributional drift is still an open and important problem. Indeed, while state-of-the-art domain incremental learning (DIL) methods perform satisfactorily within known domains, their performance largely degrades in the presence of novel domains. This limitation hampers their generalizability, and restricts their scalability to more realistic settings where train and test data are drawn from different distributions. To address these limitations, we present a novel DIL approach based on a mixture of prompt-tuned CLIP models (MoPCLIP), which generalizes the paradigm of S-Prompting to handle both in-distribution and out-of-distribution data at inference. In particular, at the training stage we model the features distribution of every class in each domain, learning individual text and visual prompts to adapt to a given domain. At inference, the learned distributions allow us to identify whether a given test sample belongs to a known domain, selecting the correct prompt for the classification task, or from an unseen domain, leveraging a mixture of the prompt-tuned CLIP models. Our empirical evaluation reveals the poor performance of existing DIL methods under domain shift, and suggests that the proposed MoP-CLIP performs competitively in the standard DIL settings while outperforming state-of-the-art methods in OOD scenarios. These results demonstrate the superiority of MoP-CLIP, offering a robust and general solution to the problem of domain incremental learning. Julien Nicolas, Florent Chiaroni, Imtiaz Masud Ziko, Ola Ahmad, Christian Desrosiers, Jose Dolz |
WACV | 4 |
| 2023 | Towards Safe Reinforcement Learning via OOD Dynamics Detection in Autonomous Driving System (Student Abstract)abstractDeep reinforcement learning (DRL) has proven effective in training agents to achieve goals in complex environments. However, a trained RL agent may exhibit, during deployment, unexpected behavior when faced with a situation where its state transitions differ even slightly from the training environment. Such a situation can arise for a variety of reasons. Rapid and accurate detection of anomalous behavior appears to be a prerequisite for using DRL in safety-critical systems, such as autonomous driving. We propose a novel OOD detection algorithm based on modeling the transition function of the training environment. Our method captures the bias of model behavior when encountering subtle changes of dynamics while maintaining a low false positive rate. Preliminary evaluations on the realistic simulator CARLA corroborate the relevance of our proposed method. Arnaud Gardille, Ola Ahmad |
AAAI | 2 |
| 2023 | DarSwin: Distortion Aware Radial Swin TransformerabstractWide-angle lenses are commonly used in perception tasks requiring a large field of view. Unfortunately, these lenses produce significant distortions making conventional models that ignore the distortion effects unable to adapt to wide-angle images. In this paper, we present a novel transformer-based model that automatically adapts to the distortion produced by wide-angle lenses. We leverage the physical characteristics of such lenses, which are analytically defined by the radial distortion profile (assumed to be known), to develop a distortion aware radial swin transformer (DarSwin). In contrast to conventional transformer-based architectures, DarSwin comprises a radial patch partitioning, a distortion-based sampling technique for creating token embeddings, and an angular position encoding for radial patch merging. We validate our method on classification tasks using synthetically distorted ImageNet data and show through extensive experiments that DarSwin can perform zero-shot adaptation to unseen distortions of different wide-angle lenses. Compared to other baselines, DarSwin achieves the best results (in terms of Top-1 accuracy) with significant gains when trained on bounded levels of distortions (very-low, low, medium, and high) and tested on all including out-of-distribution distortions. The code and models are publicly available at https://lvsn.github.io/darswin/ Akshaya Athwale, Arman Afrasiyabi, Justin Lagüe, Ichrak Shili, Ola Ahmad, Jean-François Lalonde |
ICCV | 5 |
| 2022 | FisheyeHDK: Hyperbolic Deformable Kernel Learning for Ultra-Wide Field-of-View Image RecognitionabstractConventional convolution neural networks (CNNs) trained on narrow Field-of-View (FoV) images are the state-of-the art approaches for object recognition tasks. Some methods proposed the adaptation of CNNs to ultra-wide FoV images by learning deformable kernels. However, they are limited by the Euclidean geometry and their accuracy degrades under strong distortions caused by fisheye projections. In this work, we demonstrate that learning the shape of convolution kernels in non-Euclidean spaces is better than existing deformable kernel methods. In particular, we propose a new approach that learns deformable kernel parameters (positions) in hyperbolic space. FisheyeHDK is a hybrid CNN architecture combining hyperbolic and Euclidean convolution layers for positions and features learning. First, we provide intuition of hyperbolic space for wide FoV images. Using synthetic distortion profiles, we demonstrate the effectiveness of our approach. We select two datasets - Cityscapes and BDD100K 2020 - of perspective images which we transform to fisheye equivalents at different scaling factors (analogue to focal lengths). Finally, we provide an experiment on data collected by a real fisheye camera. Validations and experiments show that our approach improves existing deformable kernel methods for CNN adaptation on fisheye images. Ola Ahmad, Freddy Lécué |
AAAI | 1 |
| 2018 | Spectral Shape Analysis of Human Torsos: Application to the Evaluation of Scoliosis Surgery OutcomeabstractThis paper aims at evaluating the effect of spinal surgery on the torso shape appearance of adolescent patients. Current methods that assess the surgical outcome on the trunk shape are limited to its global asymmetry or rely on unreliable manual measurements. We introduce a novel framework to evaluate pre- to postoperative local asymmetry changes using a spectral representation of the torso shape, more specifically, the Laplacian spectrum (eigenvalues and eigenvectors) of a graph. We conduct a statistical analysis on the eigenvalues to efficiently select the spectral space and determine the significant components between preop and postop groups. On the selected eigenvectors, we propose a local analysis based on the concept of Euler characteristic to detect their local maxima and minima, which are then used to compute local left-right (L-R) asymmetries of torso shape. On 49 patients with a thoracic spinal deformity, the method captures significant pre- to postoperative changes of asymmetry at the waist, shoulder blades, shoulders, and breasts. We have evaluated average correction rates for L-R asymmetry of the waist height (67%), shoulder-blade height (64%) and depth (67%), lateral offset between shoulder and neck (61%), and breast height (52%). Spectral torso shape analysis provides a novel approach to quantify the surgical correction of the scoliotic trunk from local shape asymmetry. The proposed method could help the surgeon to understand the impact of different spinal surgery strategies on the postoperative appearance and choose the one that should provide better patient's satisfaction. Ola Ahmad, Hervé Lombaert, Stefan Parent, Hubert Labelle, Farida Cheriet |
IEEE J. Biomed. Health Informatics | 1 |
| 2016 | Scale-space spatio-temporal random fields: Application to the detection of growing microbial patterns from surface roughness
Ola Ahmad, Christophe Collet 0001 |
Pattern Recognit. | 1 |
| 2015 | Spatio-spectral Gaussian random field modeling approach for target detection on hyperspectral data obtained in very low SNRabstractRandom field geometry has proven relevant results in the context of statistical hypothesis test for solving detection problems in signal and image processing. This paper emphasizes an unsupervised target detection problem in hyperspectral noisy images with very low signal-to-noise ratio (SNR) conditions. The targets have unknown spectral signatures located at unknown bandwidths and positions. To this aim, a spatio-spectral Gaussian random field (SS-GRF) model is proposed to provide a statistical inference about these targets in the full hyperspectral space by means of the geometric features of the noise, notably the expected Euler-characteristic (EC). The performance of the proposed method is demonstrated by the ROC curve analysis on synthetic examples, and confirms its efficiency and capacity to detect hyperspectral targets (astrophysical objects, remote sensing targets). At the end, we discuss the impact of the spectral dimensions on the method. Ola Ahmad, Christophe Collet 0001, Fabien Salzenstein |
ICIP | 1 |
| 2011 | A geometric-based method for recognizing overlapping polygonal-shaped and semi-transparent particles in gray tone images
Ola Ahmad, Johan Debayle, Jean-Charles Pinoli |
Pattern Recognit. Lett. | 1 |