VLDB 2026 Research / reviewers in the wild / expert
Antoine Manzanera
dblp:73/2951
· DBLP profile ↗
33ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-5718-411XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantitatively Audited Multi-View XAI for Medical Image Classification: Application to MGMT Promoter Methylation
Manel Mili, Abderrahman Ben Abdeljelil, Antoine Manzanera, Asma Ben Abdallah, Jose Javier Otero, Mohamed Bedoui Hedi |
ICAART (2) | 3 |
| 2025 | Improved Monocular Depth Prediction Using Distance Transform Over Pre-semantic Contours with Self-supervised Neural NetworksabstractMonocular depth estimation (MDE) with self-supervised training approaches struggles in low-texture areas, where photometric losses may lead to ambiguous depth predictions. To address this, we propose a novel technique that enhances spatial information by applying a distance transform over pre-semantic contours, augmenting discriminative power in low texture regions. Our approach jointly estimates pre-semantic contours, depth and ego-motion. The pre-semantic contours are leveraged to produce new input images, with variance augmented by the distance transform in uniform areas. This approach results in more effective loss functions, enhancing the training process for depth and ego-motion. We demonstrate theoretically that the distance transform is the optimal variance-augmenting technique in this context. Through extensive experiments on KITTI, Cityscapes, Waymo, NYUv2 and ScanNet our model demonstrates robust performance, surpassing competing self-supervised methods in MDE. Marwane Hariat, Antoine Manzanera, David Filliat |
CVPR | 2 |
| 2025 | Double Descent Meets Out-of-Distribution Detection: Theoretical Insights and Empirical Analysis on the Role of Model Complexityabstract**Out-of-distribution (OOD) detection** is essential for ensuring the reliability and safety of machine learning systems. In recent years, it has received increasing attention, particularly through post-hoc detection and training-based methods. In this paper, we focus on **post-hoc OOD detection**, which enables identifying OOD samples without altering the model's training procedure or objective. Our primary goal is to investigate the relationship between **model capacity** and its OOD detection performance. Specifically, we aim to answer the following question:
*Does the Double Descent phenomenon manifest in post-hoc OOD detection?* This question is crucial, as it can reveal whether overparameterization, which is already known to benefit generalization, can also enhance OOD detection.
Despite the growing interest in these topics by the classic supervised machine learning community, this intersection remains unexplored for OOD detection.
We empirically demonstrate that the Double Descent effect does indeed appear in post-hoc OOD detection. Furthermore, we provide theoretical insights to explain why this phenomenon emerges in such setting. Finally, we show that the overparameterized regime does not yield superior results consistently, and we propose a method to identify the optimal regime for OOD detection based on our observations. Mouïn Ben Ammar, David Brellmann, Arturo Mendoza, Antoine Manzanera, Gianni Franchi |
NeurIPS | 4 |
| 2025 | A multimodal gait and ocular geometric representation to generate a Parkinson progression report
John Archila, Ivan Peña, Luis Fernando Celis, Juan A. Olmos, Antoine Manzanera, Fabio Martínez |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Geometric multimodal learning to support prostate cancer diagnosis on limited and multicentric bi-parametric MRI dataabstractAbstract The classification of clinically significant prostate cancer (csPCa) lesions remains one of the most important challenges in prostate cancer diagnosis. For this, multimodal convolutional neural networks (CNNs) have achieved outstanding results. Nevertheless, the data used in these studies may only partially represent the total burden of csPCa cases. Hence, it is necessary to design reliable models that perform well in limited data scenarios and involving information from different centers (multicentric). A deep Riemannian geometric learning architecture was introduced to capture the intermediate relationships between bi-parametric MRI (bp-MRI) deep representations coded from a 3D multimodal convolutional backbone and considering their geometry. For this, several multimodal bp-MRI fusion strategies were explored to assess their ability to classify csPCa lesions in scenarios where the percentage of available training data was progressively reduced and multicentric data were involved. The proposed method outperformed baseline CNN techniques with an AUC-ROC of 0.96. More remarkably, the method remained stable even only using 10% of the available training data. Additionally, considering multicentric information, this approach also demonstrates generalization ability by losing only 5.4% of the AUC testing data from different acquisition centers, compared to the 10.4% loss of the baseline method. A new deep learning-based method that improves generalization under scenarios with limited data translates to better support for clinicians in accurately classifying csPCa lesions on unseen data. Juan A. Olmos, Antoine Manzanera, Fabio Martínez |
Neural Comput. Appl. | 2 |
| 2025 | Learning a geometric deep representation to classify Parkinson smooth pursuit patterns
Luis Fernando Celis, Juan A. Olmos, Antoine Manzanera, Fabio Martínez |
Pattern Anal. Appl. | 3 |
| 2024 | A Novel Hybrid Grid Search and Tree Parzen Estimator for Deep Learning Hyperparameters OptimizationabstractHyperparameter optimization plays a crucial role in maximizing the performance of Deep Learning (DL) models, particularly in the medical field. In this study, we propose a novel hybrid approach called GS-TPE, which combines Grid Search (GS) and Tree Parzen Estimator (TPE) for optimizing the hyperparameters of DL architectures in order to enhance the vigilance states classification from the EEG signals. Our experiments demonstrate that the GS-TPE approach competes with the state of the art on multiple performance metrics, leading to significantly improved classification results. The obtained accuracy with combined one-Dimensional Convolutional Neural Network and Long Short-Term Memory (1D-CNN-LSTM) and with combined Auto-Encoder and LSTM (AE-LSTM) architectures reach 93.74 and 93.53%, respectively. The proposed GS-TPE approach shows great promise for advancing the field of medical signal analysis and enhancing the accuracy of EEG-based diagnostic systems. Souhir Khessiba, Ahmed Ghazi Blaiech, Asma Ben Abdallah, Antoine Manzanera, Khaled Ben Khalifa, Mohamed Bedoui Hedi |
AICCSA | 4 |
| 2024 | NECO: NEural Collapse Based Out-of-distribution detectionabstractDetecting out-of-distribution (OOD) data is a critical challenge in machine learning due to model overconfidence, often without awareness of their epistemological limits. We hypothesize that "neural collapse", a phenomenon affecting in-distribution data for models trained beyond loss convergence, also influences OOD data. To benefit from this interplay, we introduce NECO, a novel post-hoc method for OOD detection, which leverages the geometric properties of “neural collapse” and of principal component spaces to identify OOD data. Our extensive experiments demonstrate that NECO achieves state-of-the-art results on both small and large-scale OOD detection tasks while exhibiting strong generalization capabilities across different network architectures. Furthermore, we provide a theoretical explanation for the effectiveness of our method in OOD detection. We plan to release the code after the anonymity period. Mouïn Ben Ammar, Nacim Belkhir, Sebastian Popescu, Antoine Manzanera, Gianni Franchi |
ICLR | 4 |
| 2024 | InfraParis: A multi-modal and multi-task autonomous driving datasetabstractCurrent deep neural networks (DNNs) for autonomous driving computer vision are typically trained on specific datasets that only involve a single type of data and urban scenes. Consequently, these models struggle to handle new objects, noise, nighttime conditions, and diverse scenarios, which is essential for safety-critical applications. Despite ongoing efforts to enhance the resilience of computer vision DNNs, progress has been sluggish, partly due to the absence of benchmarks featuring multiple modalities. We introduce a novel and versatile dataset named InfraParis that supports multiple tasks across three modalities: RGB, depth, and infrared. We assess various state-of-the-art baseline techniques, encompassing models for the tasks of semantic segmentation, object detection, and depth estimation. More visualizations and the download link for InfraParis are available at https://enstau2is.github.io/infraParis/. Gianni Franchi, Marwane Hariat, Xuanlong Yu, Nacim Belkhir, Antoine Manzanera, David Filliat |
WACV | 5 |
| 2024 | Riemannian SPD learning to represent and characterize fixational oculomotor Parkinsonian abnormalitiesabstractParkinson’s disease (PD) is the second most common neurodegenerative disorder, mainly characterized by motor alterations. Despite multiple efforts, there is no definitive biomarker to diagnose, quantify, and characterize the disease early. Recently, abnormal fixational oculomotor patterns have emerged as a promising disease biomarker with high sensitivity, even at early stages. Nonetheless, the complex patterns and potential correlations with the disease remain largely unexplored, among others, because of the limitations of standard setups that only analyze coarse measures and poorly exploit the associated PD alterations. This work introduces a new strategy to represent, analyze and characterize fixational patterns from non-invasive video analysis, adjusting a geometric learning strategy. A deep Riemannian framework is proposed to discover potential oculomotor patterns aimed at withstanding data scarcity and geometrically interpreting the latent space. A convolutional representation is first built, then aggregated onto a symmetric positive definite matrix (SPD). The latter encodes second-order statistics of deep convolutional features and feeds a non-linear hierarchical architecture that processes SPD data by maintaining them into their Riemannian manifold. The complete representation discriminates between Parkinson and Healthy (Control) fixational observations, even at PD stages 2.5 and 3. Besides, the proposed geometrical representation exhibit capabilities to statistically differentiate observations among Parkinson’s stages. The developed tool demonstrates coherent results from explainability maps back-propagated from output probabilities. Juan A. Olmos, Antoine Manzanera, Fabio Martínez |
Pattern Recognit. Lett. | 2 |
| 2023 | Improving Knee Osteoarthritis Classification with Markerless Pose Estimation and STGCN ModelabstractKnee osteoarthritis (KOA) is a debilitating disease that greatly impacts the quality of life, particularly among the elderly population. Conventional subjective assessment methods for KOA have limitations in terms of accuracy and objective diagnosis. This paper proposes an innovative approach by integrating advanced technologies, specifically the Spatio-Temporal Graph Convolutional Network (STGCN), applied to gait analysis from markerless videos, for precise and quantitative assessment of KOA. The STGCN network is applied to normalized data obtained from Blazepose, a markerless pose estimation technique. Evaluated on an academic dataset of 80 RGB videos, it provides an accuracy of 93.75%. By leveraging the capabilities of the STGCN network, this study significantly enhances the classification of KOA based on gait patterns, offering promising prospects for improved diagnosis and treatment strategies for individuals with KOA. Souhir Khessiba, Ahmed Ghazi Blaiech, Asma Ben Abdallah, Rim Grassa, Antoine Manzanera, Mohamed Bedoui Hedi |
MMSP | 5 |
| 2023 | Rebalancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictionsabstractWe present CoopNet, an approach that improves the co-operation of co-trained networks by dynamically adapting the apportionment of gradient, to ensure equitable learning progress. It is applied to motion-aware self-supervised prediction of depth maps, by introducing a new hybrid loss, based on a distribution model of photo-metric reconstruction errors made by, on the one hand the depth + odometry paired networks, and on the other hand the optical flow network. This model essentially assumes that the pixels from moving objects (that must be discarded for training depth and odometry), correspond to those where the two reconstructions strongly disagree. We justify this model by theoretical considerations and experimental evidences. A comparative evaluation on KITTI and CityScapes datasets shows that CoopNet improves or is comparable to the state-of-the-art in depth, odometry and optical flow predictions. Our code is available here: https://github.com/mhariat/CoopNet. Marwane Hariat, Antoine Manzanera, David Filliat |
WACV | 2 |
| 2023 | Does it work outside this benchmark? Introducing the rigid depth constructor tool
Clement Pinard, Antoine Manzanera |
Multim. Tools Appl. | 2 |
| 2020 | Naturally Constrained Online Expectation MaximizationabstractWith the rise of big data sets, learning algorithms must be adapted to piece-wise mechanisms to tackle large-scale calculations' time and memory costs. Furthermore, for most learning embedded systems, the input data are fed sequentially and contingently: one by one, and possibly class by class. Thus, learning algorithms should not only run online but cope with time-varying, non-independent, and non-balanced training data for the system's entire life. Online Expectation-Maximization is a well-known algorithm for learning probabilistic models in realtime, due to its simplicity and convergence properties. However, these properties are only valid in the case of large, independent and identically distributed samples. In this paper, we propose to constrain the online Expectation-Maximization on the Fisher distance between the parameters. After presenting the algorithm, we make a thorough study of its use in Probabilistic Principal Components Analysis. First, we derive the update rules, and then we analyze the effect of the constraint on major problems of online and sequential learning: convergence, forgetting and interference. Furthermore, we use several algorithmic protocols: iid vs sequential data, and constraint parameters updated stepwise vs class-wise. Our results show that this constraint increases the convergence rate of online Expectation-Maximization, decreases forgetting and slightly introduces positive transfer learning. Daniela Pamplona, Antoine Manzanera |
ICPR | 2 |
| 2019 | A Fractal-Based Approach to Network Characterization Applied to Texture Analysis
Lucas Correia Ribas, Antoine Manzanera, Odemir Martinez Bruno |
CAIP (1) | 2 |
| 2018 | Bio-Inspired Perception Sensor (BIPS) Concept Analysis for Embedded Applications
Louise Sarrabezolles, Antoine Manzanera, Nicolas Hueber |
CIARP | 2 |
| 2017 | Mixing Hough and Color Histogram Models for Accurate Real-Time Object Tracking
Antoine Tran, Antoine Manzanera |
CAIP (1) | 2 |
| 2017 | Fast Semi Dense Epipolar Flow EstimationabstractOptical flow computation consists in recovering the apparent motion field between two images with overlapping fields of view. This paper focuses on a subset of optical flow problems, called epipolar flow, where the camera moves inside a scene containing no moving objects. Accurate solutions exist but their high computational complexities make them non suitable for a large panel of real-time applications. We propose a new epipolar flow approach with low computational complexity achieving the best error rate on the non dense KITTI optical flow 2012 benchmark and running 1000 faster than the second ranked approach. On a 4core 3GHz processor, our multi-core implementation computes a semi dense optical flow field of a 450k pixels image in 260ms. It is a significant advance in reducing the running time of accurate optical flow computation. To achieve such results we rely on the epipolar constraints and the local coherence of the optical flow not only to increase accuracy but also to reduce computational complexity. Our contribution is twofold. It is first, the acceleration and the accuracy increase of current RANSAC based visual odometry algorithms via the estimation of a robust sparse flow field, well distributed over the image domain. And then, the estimation of a semi dense flow field leveraging epipolar constraints and a propagation scheme to speedup the estimation and reduce error rates. Matthieu Garrigues, Antoine Manzanera |
WACV | 2 |
| 2017 | Spatio-temporal multi-scale motion descriptor from a spatially-constrained decomposition for online action recognitionabstractThis study presents a spatio‐temporal motion descriptor that is computed from a spatially‐constrained decomposition and applied to online classification and recognition of human activities. The method starts by computing a dense optical flow without explicit spatial regularisation. Potential human actions are detected at each frame as spatially consistent moving regions of interest (RoIs). Each of these RoIs is then sequentially partitioned to obtain a spatial representation of small overlapped subregions with different sizes. Each of these region parts is characterised by a set of flow orientation histograms. A particular RoI is then described along the time by a set of recursively calculated statistics that collect information from the temporal history of orientation histograms, to form the action descriptor. At any time, the whole descriptor can be extracted and labelled by a previously trained support vector machine. The method was evaluated using three different public datasets: (i) the ViSOR dataset was used for global classification obtaining an average accuracy of 95% and for recognition in long sequences, achieving an average per‐frame accuracy of 92.3%. (ii) The KTH dataset was used for global classification and (iii) the UT‐datasets were used for recognition task, obtaining an average accuracy of 80% (frame rate). Fabio Martínez, Antoine Manzanera, Eduardo Romero 0001 |
IET Comput. Vis. | 2 |
| 2016 | Statistical binary patterns for rotational invariant texture classification
Thanh Phuong Nguyen 0001, Ngoc-Son Vu, Antoine Manzanera |
Neurocomputing | 3 |
| 2016 | Topological Attribute Patterns for texture recognition
Thanh Phuong Nguyen 0001, Antoine Manzanera, Walter G. Kropatsch, Xuan Son Nguyen |
Pattern Recognit. Lett. | 2 |
| 2014 | Spatial Motion Patterns: Action Models from Semi-Dense TrajectoriesabstractA new action model is proposed, by revisiting local binary patterns (LBP) for dynamic texture models, applied on trajectory beams calculated on the video. The use of semi-dense trajectory field allows to dramatically reduce the computation support to essential motion information, while maintaining a large amount of data to ensure robustness of statistical bag of features action models. A new binary pattern, called Spatial Motion Pattern (SMP) is proposed, which captures self-similarity of velocity around each tracked point (particle), along its trajectory. This operator highlights the geometric shape of rigid parts of moving objects in a video sequence. SMPs are combined with basic velocity information to form the local action primitives. Then, a global representation of a space × time video block is provided by using hierarchical blockwise histograms, which allows to efficiently represent the action as a whole, while preserving a certain level of spatiotemporal relation between the action primitives. Inheriting from the efficiency and the invariance properties of both the semi-dense tracker Video extruder and the LBP-based representations, the method is designed for the fast computation of action descriptors in unconstrained videos. For improving both robustness and computation time in the case of high definition video, we also present an enhanced version of the semi-dense tracker based on the so-called super particles, which reduces the number of trajectories while improving their length, reliability and spatial distribution. Thanh Phuong Nguyen 0001, Antoine Manzanera, Matthieu Garrigues, Ngoc-Son Vu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2013 | Motion Trend Patterns for Action Modelling and Recognition
Thanh Phuong Nguyen 0001, Antoine Manzanera, Matthieu Garrigues |
CAIP (1) | 2 |
| 2013 | Revisiting LBP-Based Texture Models for Human Action Recognition
Thanh Phuong Nguyen 0001, Antoine Manzanera, Ngoc-Son Vu, Matthieu Garrigues |
CIARP (2) | 2 |
| 2013 | Action recognition using bag of features extracted from a beam of trajectoriesabstractA new spatio temporal descriptor is proposed for action recognition. The action is modelled from a beam of trajectories obtained using semi dense point tracking on the video sequence. We detect the dominant points of these trajectories as points of local extremum curvature and extract their corresponding feature vectors, to form a dictionary of atomic action elements. The high density of these informative and invariant elements allows effective statistical action description. Then, human action recognition is performed using a bag of feature model with SVM classifier. Experimentations show promising results on several well-known datasets. Thanh Phuong Nguyen 0001, Antoine Manzanera |
ICIP | 2 |
| 2010 | Ground-plane classification for robot navigation: Combining multiple cues toward a visual-based learning systemabstractThis paper describes a vision-based ground-plane classification system for autonomous indoor mobile-robot that takes advantage of the synergy in combining together multiple visual-cues. A priori knowledge of the environment is important in many biological systems, in parallel with their reactive systems. As such, a learning model approach is taken here for the classification of the ground/object space, initialised through a new Distributed-Fusion (D-Fusion) method that captures colour and textural data using Superpixels. A Markov Random Field (MRF) network is then used to classify, regularise, employ a priori constraints, and merge additional ground/object information provided by other visual cues (such as motion) to improve classification images. The developed system can classify indoor test-set ground-plane surfaces with an average true-positive to false-positive rate of 90.92% to 7.78% respectively on test-set data. The system has been designed in mind to fuse a variety of different visual-cues. Consequently it can be customised to fit different situations and/or sensory architectures accordingly. Tobias Low, Antoine Manzanera |
ICARCV | 2 |
| 2010 | Local Jet Based Similarity for NL-Means FilteringabstractReducing the dimension of local descriptors in images is useful to perform pixels comparison faster. We show here that, for computing the NL-means denoising filter, image patches can be favourably replaced by a vector of spatial derivatives (local jet), to calculate the similarity between pixels. First, we present the basic, limited range implementation, and compare it with the original NL-means. We use a fast estimation of the noise variance to automatically adjust the decay parameter of the filter. Next, we present the unlimited range implementation using nearest neighbours search in the local jet space, based on a binary search tree representation. Antoine Manzanera |
ICPR | 1 |
| 2009 | Image Characterization from Statistical Reduction of Local Patterns
Philippe Guermeur, Antoine Manzanera |
CIARP | 2 |
| 2009 | Motion detection: Fast and robust algorithms for embedded systemsabstractThis article introduces a new hierarchical version of a set of motion detection algorithms called ¿¿. These new algorithms are designed to preserve as much as possible the computational efficiency of the basic ¿¿ estimation, in order to target real-time implementation for low power consumption processors and embedded systems. Lionel Lacassagne, Antoine Manzanera, Antoine Dupret |
ICIP | 2 |
| 2008 | Generalization performance of vision based controllers for mobile robots evolved with genetic programmingabstractWe present a genetic programming system to design automatically vision based obstacle avoidance algorithms adapted to the current context. We use a simulation environment to evaluate the controllers. By restricting the structure of the algorithms to facilitate the compromise between obstacle avoidance and target reaching, we improve the generalization performance of the algorithms. Renaud Barate, Antoine Manzanera |
GECCO | 2 |
| 2007 | Sigma-Delta Background Subtraction and the Zipf Law
Antoine Manzanera |
CIARP | 1 |
| 2007 | A new motion detection algorithm based on Sigma-Delta background estimation
Antoine Manzanera, Julien C. Richefeu |
Pattern Recognit. Lett. | 1 |
| 1999 | Medial Faces from a Concise 3D Thinning AlgorithmabstractWe propose in this paper a new 3D fully parallel thinning algorithm that we believe to be the most concise due to its simple characterization. The algorithm is indeed completely defined by a set of five patterns, three removing conditions and two non-removing conditions. These patterns are designed from the two fundamental and compatible constraints usually expected in skeleta: (1) Topology preservation and (2) Medial surface. From these two constraints, the removing patterns terns (/spl alpha//sub 1/, /spl alpha//sub 2/ and /spl alpha//sub 3/) detect the non-local maxima, whereas the non-removing patterns (/spl beta//sub 1/ and /spl beta//sub 2/) prevent any topology change that the removing conditions could imply. We show that the three mentioned constraints are respected. The logical conciseness of our procedure, called MB-3D, makes it to our knowledge the easiest 3D thinning algorithm to implement. Some results are displayed, that illustrate the relevance of our approach. Antoine Manzanera, Thierry M. Bernard, Françoise J. Prêteux, Bernard Longuet |
ICCV | 1 |