VLDB 2026 Research / reviewers in the wild / expert
Alain Pagani
dblp:59/3540
· DBLP profile ↗
29ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-5136-0837ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: dAIEDGE - A Network of Excellence for Distributed, Trustworthy, Efficient and Scalable AI at the EdgeabstractThe dAIEDGE Network of Excellence (NoE) seeks to strengthen and support the development of a dynamic European cutting-edge Artificial intelligence (AI) ecosystem under the umbrella of the European Lighthouse for AI, and to sustain the development of advanced AI. dAIEDGE fosters the exchange of ideas, concepts, and trends on cutting-edge next generation AI, creating links between ecosystem actors to help both the European Commission (EC) and the European Union (EU) and the peripheral AI constituency identify strategies for future developments in Europe. Our main objective is to advance Europe’s innovation and technology base by developing a comprehensive policy and governance approach to AI in order for the EU to become a world leader in innovation in the data economy and its applications. Alain Pagani, Haralampos-G. D. Stratigopoulos, Aysajan Abidin, Mhd Rashed Al Koutayni, Luca Benini, Angelos Bilas, Alessandro Capotondi, Roberto Cavicchioli, Brian Clerkin, Oscar Déniz-Suárez, Margaux Divernois, Baptiste Dupertuis, Dorvan Favre, Giulio Gambardella, Ander García Gangoiti, Carlo Augusto Grazia, Dominik Günzel, Jude Haris, Klodjan K. Hidri, Maïck Huguenin-Vuillemin, Manal Jammal, Paul Kling, Christos Kozanitis, Xavier Lessage, Srikanth Mandapati, Philippe Massonet, Alfio Di Mauro, Varesh Mishra, Juan Odriozola, Javier Parra 0001, Nuria Pazos, Viviane Potocnik, Miguel de Prado, Rohit Prasad, Spyridon Raptis, Gregoire Rebstein, Ignacio Sanudo Olmedo, Mohamed Selim, Chinmay Satish Shrivastav, Noelia Vállez, Giorgos Vasiliadis, Micaela Verrucchi, Enrico Vincenzi, Damian Vizár, Devendra Vyas, Stefan Wiehle |
DATE | 1 |
| 2026 | TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion ModelabstractRecent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses. Nevertheless, generating temporally coherent long-form content remains challenging. Existing approaches are constrained by computational and memory limitations, as they are typically trained on short video segments, thus performing effectively only over limited frame lengths and hindering their potential for extended coherent generation. To address these constraints, we propose TalkingPose, a novel diffusion-based framework specifically designed for producing long-form, temporally consistent human upper-body animations. TalkingPose leverages driving frames to precisely capture expressive facial and hand movements, transferring these seamlessly to a target actor through a stable diffusion backbone. To ensure continuous motion and enhance temporal coherence, we introduce a feedback-driven mechanism built upon image-based diffusion models. Notably, this mechanism does not incur additional computational costs or require secondary training stages, enabling the generation of animations with unlimited duration. Additionally, we introduce a comprehensive, large-scale dataset to serve as a new benchmark for human upper-body animation. Project page: https://dfki-av.github.io/TalkingPose Alireza Javanmardi, Pragati Jaiswal, Tewodros Habtegebrial, Christen Millerdurai, Shaoxiang Wang, Alain Pagani, Didier Stricker |
WACV | 6 |
| 2026 | Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° ScenesabstractDespite recent advances in single-object front-facing inpainting using NeRF and 3D Gaussian Splatting (3DGS), inpainting in complex 360° scenes remains largely underexplored. This is primarily due to three key challenges: (i) identifying target objects in the 3D field of 360° environments, (ii) dealing with severe occlusions in multi-object scenes, which makes it hard to define regions to inpaint, and (iii) maintaining consistent and high-quality appearance across views effectively.To tackle these challenges, we propose Inpaint360GS, a flexible 360° editing framework based on 3DGS that supports multi-object removal and high-fidelity inpainting in 3D space. By distilling 2D segmentation into 3D and leveraging virtual camera views for contextual guidance, our method enables accurate object-level editing and consistent scene completion. We further introduce a new dataset tailored for 360° inpainting, addressing the lack of ground truth object-free scenes. Experiments demonstrate that Inpaint360GS outperforms existing baselines and achieves state-of-the-art performance. Project page: https://dfki-av.github.io/inpaint360gs/ Shaoxiang Wang, Shihong Zhang, Christen Millerdurai, Rüdiger Westermann, Didier Stricker, Alain Pagani |
WACV | 6 |
| 2026 | Multi-View Face and Gesture Animation with Dynamic GaussiansabstractCreating photorealistic 3D human avatars with realistic upper-body motion remains challenging. Existing approaches either focus on the head and overlook hand gestures, or reconstruct the full body but fail to preserve fine-grained facial fidelity and hand pose accuracy. As a result, current methods struggle to capture the subtle dynamics of facial expressions and hand gestures that are crucial for natural human communication. While methods based on full-body parametric models enable avatar reconstruction from monocular or multi-view inputs, they often lack accurate facial animation and detailed hand articulation. To address these limitations, we propose MVFGA, a novel multi-view-consistent pipeline for generating realistic upper-body avatars. Our approach models the face and hands separately and fuses them with a parametric upper-body mesh model, enabling the capture of fine-grained facial expressions and hand poses for accurate upper-body avatar reconstruction. We then splat 3D Gaussians onto the obtained mesh, enabling high-quality rendering of dynamic avatars from novel viewpoints. Furthermore, we introduce MVFGA-MoCap, a multi-view upper-body motion capture dataset featuring controlled facial expression sequences, diverse hand gestures, and free-form communication. Experiments show that MVFGA generates visually realistic avatars with high-fidelity facial expressions and hand motions, outperforming baselines for upper-body avatar animation. Project page: https://dfki-av.github.io/MVFGA/ Alireza Javanmardi, Vippin Kumar Jeetmal, Christen Millerdurai, Alain Pagani, Didier Stricker |
Comput. Graph. Forum | 4 |
| 2025 | 3D Spatial Understanding in MLLMs: Disambiguation and EvaluationabstractMultimodal Large Language Models (MLLMs) have made significant progress in tasks such as image captioning and question answering. However, while these models can generate realistic captions, they often struggle with providing precise instructions, particularly when it comes to localizing and disambiguating objects in complex 3D environments. This capability is critical as MLLMs become more integrated with collaborative robotic systems. In scenarios where a target object is surrounded by similar objects (distractors), robots must deliver clear, spatially-aware instructions to guide humans effectively. We refer to this challenge as contextual object localization and disambiguation, which imposes stricter constraints than conventional 3D dense captioning, especially regarding ensuring target exclusivity. In response, we propose simple yet effective techniques to enhance the model's ability to localize and disambiguate target objects. Our approach not only achieves state-of-the-art performance on conventional metrics that evaluate sentence similarity, but also demonstrates improved 3D spatial understanding through 3D visual grounding model. https://birdy666.github.io/projects/3d_spatial_understanding_in_mllms/ Chun-Peng Chang, Alain Pagani, Didier Stricker |
ICRA | 2 |
| 2025 | Uni-SLAM: Uncertainty-Aware Neural Implicit SLAM for Real-Time Dense Indoor Scene ReconstructionabstractNeural implicit fields have recently emerged as a powerful representation method for multi-view surface reconstruction due to their simplicity and state-of-the-art performance. However, reconstructing thin structures of indoor scenes while ensuring real-time performance remains a challenge for dense visual SLAM systems. Previous methods do not consider varying quality of input RGB-D data and employ fixed-frequency mapping process to reconstruct the scene, which could result in the loss of valuable information in some frames. In this paper, we propose Uni-SLAM, a decoupled 3D spatial representation based on hash grids for indoor reconstruction. We introduce a novel defined predictive uncertainty to reweight the loss function, along with strategic local-to-global bundle adjustment. Experiments on synthetic and real-world datasets demonstrate that our system achieves state-of-the-art tracking and mapping accuracy while maintaining real-time performance. It significantly improves over current methods with a 25% reduction in depth L1 error and a 66.86% completion rate within 1 cm on the Replica dataset, reflecting a more accurate reconstruction of thin structures. Project page: https://shaoxiang777.github.io/project/uni-slam/ Shaoxiang Wang, Yaxu Xie, Chun-Peng Chang, Christen Millerdurai, Alain Pagani, Didier Stricker |
WACV | 5 |
| 2025 | EventEgo3D++: 3D Human Motion Capture from A Head-Mounted Event CameraabstractMonocular egocentric 3D human motion capture remains a significant challenge, particularly under conditions of low lighting and fast movements, which are common in head-mounted device applications. Existing methods that rely on RGB cameras often fail under these conditions. To address these limitations, we introduce EventEgo3D++, the first approach that leverages a monocular event camera with a fisheye lens for 3D human motion capture. Event cameras excel in high-speed scenarios and varying illumination due to their high temporal resolution, providing reliable cues for accurate 3D human motion capture. EventEgo3D++ leverages the LNES representation of event streams to enable precise 3D reconstructions. We have also developed a mobile head-mounted device (HMD) prototype equipped with an event camera, capturing a comprehensive dataset that includes real event observations from both controlled studio environments and in-the-wild settings, in addition to a synthetic dataset. Additionally, to provide a more holistic dataset, we include allocentric RGB streams that offer different perspectives of the HMD wearer, along with their corresponding SMPL body model. Our experiments demonstrate that EventEgo3D++ achieves superior 3D accuracy and robustness compared to existing solutions, even in challenging conditions. Moreover, our method supports real-time 3D pose updates at a rate of 140Hz. This work is an extension of the EventEgo3D approach (CVPR 2024) and further advances the state of the art in egocentric 3D human motion capture. For more details, visit the project page at https://eventego3d.mpi-inf.mpg.de. Christen Millerdurai, Hiroyasu Akada, Jian Wang 0042, Diogo C. Luvizon, Alain Pagani, Didier Stricker, Christian Theobalt, Vladislav Golyanik |
Int. J. Comput. Vis. | 5 |
| 2024 | G3FA: Geometry-guided GAN for Face Animation
Alireza Javanmardi, Alain Pagani, Didier Stricker |
BMVC | 2 |
| 2024 | MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Groundingabstract3D visual grounding involves matching natural language descriptions with their corresponding objects in 3D spaces. Existing methods often face challenges with accuracy in object recognition and struggle in interpreting complex linguistic queries, particularly with descriptions that involve multiple anchors or are view-dependent. In response, we present the MiKASA (Multi-Key-Anchor Scene-Aware) Transformer. Our novel end-to-end trained model integrates a self-attention-based scene-aware object encoder and an original multi-key-anchor technique, enhancing object recognition accuracy and the understanding of spatial relationships. Furthermore, MiKASA improves the explainability of decision-making, facilitating error diagnosis. Our model achieves the highest overall accuracy in the Referit3D challenge for both the Sr3D and Nr3D datasets, particularly excelling by a large margin in categories that require viewpoint-dependent descriptions. The source code and additional resources for this project are available on GitHub: https://github.com/dfki-av/MiKASA-3DVG Chun-Peng Chang, Shaoxiang Wang, Alain Pagani, Didier Stricker |
CVPR | 3 |
| 2024 | SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and its Downstream TasksabstractScene graphs have been recently introduced into 3D spatial understanding as a comprehensive representation of the scene. The alignment between 3D scene graphs is the first step of many downstream tasks such as scene graph aided point cloud registration, mosaicking, overlap checking, and robot navigation. In this work, we treat 3D scene graph alignment as a partial graph-matching problem and propose to solve it with a graph neural network. We reuse the geometric features learned by a point cloud registration method and associate the clustered point-level geometric features with the node-level semantic feature via our designed feature fusion module. Partial matching is enabled by using a learnable method to select the top-k similar node pairs. Subsequent downstream tasks such as point cloud registration are achieved by running a pre-trained registration network within the matched regions. We further propose a point-matching rescoring method, that uses the node-wise alignment of the 3D scene graph to reweight the matching candidates from a pre-trained point cloud registration method. It reduces the false point correspondences estimated especially in low-overlapping cases. Experiments show that our method improves the alignment accuracy by lOrv20% in low-overlap and random transformation scenarios and outperforms the existing work in multiple downstream tasks. Our code and models are available here. Yaxu Xie, Alain Pagani, Didier Stricker |
CVPR | 2 |
| 2024 | SphereCraft: A Dataset for Spherical Keypoint Detection, Matching and Camera Pose EstimationabstractThis paper introduces SphereCraft, a dataset specifically designed for spherical keypoint detection, matching, and camera pose estimation. The dataset addresses the limitations of existing datasets by providing extracted keypoints from various detectors, along with their ground truth correspondences. Synthetic scenes with photo-realistic rendering and accurate 3D meshes are included, as well as real-world scenes acquired from different spherical cameras. SphereCraft enables the development and evaluation of algorithms targeting multiple camera viewpoints, advancing the state-of-the-art in computer vision tasks involving spherical images. Our dataset is available at https://dfki.github.io/spherecraftweb/. Christiano Couto Gava, Yunmin Cho, Federico Raue, Sebastian Palacio, Alain Pagani, Andreas Dengel 0001 |
WACV | 5 |
| 2023 | Structure PLP-SLAM: Efficient Sparse Mapping and Localization using Point, Line and Plane for Monocular, RGB-D and Stereo CamerasabstractThis paper presents a visual SLAM system that uses both points and lines for robust camera localization, and simultaneously performs a piece-wise planar reconstruction (PPR) of the environment to provide a structural map in real-time. One of the biggest challenges in parallel tracking and mapping with a monocular camera is to keep the scale consistent when reconstructing the geometric primitives. This further introduces difficulties in graph optimization of the bundle adjustment (BA) step. We solve these problems by proposing several run-time optimizations on the reconstructed lines and planes. Our system is able to run with depth and stereo sensors in addition to the monocular setting. Our proposed SLAM tightly incorporates the semantic and geometric features to boost both frontend pose tracking and backend map optimization. We evaluate our system exhaustively on various datasets, and show that we outperform state-of-the-art methods in terms of trajectory precision. The code of PLP-SLAM has been made available in open-source for the research community (https://github.com/PeterFWS/Structure-PLP-SLAM). Fangwen Shu, Alain Pagani, Didier Stricker |
ICRA | 3 |
| 2023 | BoxMask: Revisiting Bounding Box Supervision for Video Object DetectionabstractWe present a new, simple yet effective approach to up- lift video object detection. We observe that prior works operate on instance-level feature aggregation that imminently neglects the refined pixel-level representation, resulting in confusion among objects sharing similar appearance or motion characteristics. To address this limitation, we propose BoxMask, which effectively learns discriminative representations by incorporating class-aware pixel-level information. We simply consider bounding box-level annotations as a coarse mask for each object to supervise our method. The proposed module can be effortlessly integrated into any region-based detector to boost detection. Extensive experiments on ImageNet VID and EPIC KITCHENS datasets demonstrate consistent and significant improvement when we plug our BoxMask module into numerous recent state-of-the-art methods. The code will be available at https://github.com/khurramHashmi/BoxMask. Khurram Azeem Hashmi, Alain Pagani, Didier Stricker, Muhammad Zeshan Afzal |
WACV | 2 |
| 2023 | Learning Attention Propagation for Compositional Zero-Shot LearningabstractCompositional zero-shot learning aims to recognize unseen compositions of seen visual primitives of object classes and their states. While all primitives (states and objects) are observable during training in some combination, their complex interaction makes this task especially hard. For example, wet changes the visual appearance of a dog very differently from a bicycle. Furthermore, we argue that relationships between compositions go beyond shared states or objects. A cluttered office can contain a busy table; even though these compositions don’t share a state or object, the presence of a busy table can guide the presence of a cluttered office. We propose a novel method called Compositional Attention Propagated Embedding (CAPE) as a solution. The key intuition to our method is that a rich dependency structure exists between compositions arising from complex interactions of primitives in addition to other dependencies between compositions. CAPE learns to identify this structure and propagates knowledge between them to learn class embedding for all seen and unseen compositions. In the challenging generalized compositional zero-shot setting, we show that our method outperforms previous baselines to set a new state-of-the-art on three publicly available benchmarks. Muhammad Gul Zain Ali Khan, Muhammad Ferjad Naeem, Luc Van Gool, Alain Pagani, Didier Stricker, Muhammad Zeshan Afzal |
WACV | 4 |
| 2021 | PlaneRecNet: Multi-Task Learning with Cross-Task consistency for Piece-Wise Plane Detection and Reconstruction from a Single RGB Image
Yaxu Xie, Fangwen Shu, Jason R. Rambach, Alain Pagani, Didier Stricker |
BMVC | 4 |
| 2021 | Pose Tracking vs. Pose Estimation of AR Glasses with Convolutional, Recurrent, and Non-local Neural Networks: A Comparison
Ahmet Firintepe, Sarfaraz Habib, Alain Pagani, Didier Stricker |
EuroXR | 3 |
| 2021 | SLAM in the Field: An Evaluation of Monocular Mapping and Localization on Challenging Dynamic Agricultural EnvironmentabstractThis paper demonstrates a system capable of combining a sparse, indirect, monocular visual SLAM, with both offline and real-time Multi-View Stereo (MVS) reconstruction algorithms. This combination overcomes many obstacles encountered by autonomous vehicles or robots employed in agricultural environments, such as overly repetitive patterns, need for very detailed reconstructions, and abrupt movements caused by uneven roads. Furthermore, the use of a monocular SLAM makes our system much easier to integrate with an existing device, as we do not rely on a LiDAR (which is expensive and power consuming), or stereo camera (whose calibration is sensitive to external perturbation e.g. camera being displaced). To the best of our knowledge, this paper presents the first evaluation results for monocular SLAM, and our work further explores unsupervised depth estimation on this specific application scenario by simulating RGB-D SLAM to tackle the scale ambiguity, and shows our approach produces reconstructions that are helpful to various agricultural tasks. Moreover, we highlight that our experiments provide meaningful insight to improve monocular SLAM systems under agricultural settings. Fangwen Shu, Paul Lesur, Yaxu Xie, Alain Pagani, Didier Stricker |
WACV | 4 |
| 2020 | The More, the Merrier? A Study on In-Car IR-based Head Pose EstimationabstractDeep learning methods have proven useful for head pose estimation, but the effect of their depth, type and input resolution based on infrared (IR) images still need to be explored. In this paper, we present a study on in-car head pose estimation on the IR images of the AutoPOSE dataset, where we extract 64 x 64 and 128 x 128 pixel cropped head images. We propose the novel networks Head Orientation Network (HON) and ResNetHG and compare them with state-of-the-art methods like the HPN model from DriveAHead on different input resolutions. In addition, we evaluate multiple depths within our HON and ResNetHG networks and their effect on the accuracy. Our experiments show that higher resolution images lead to lower estimation errors. Furthermore, we show that deep learning methods with fewer layers perform better on head orientation regression based on IR images. Our HON and ResNetHG18 architectures outperform the state-of-the-art on IR images on four different metrics, where we achieve a reduction of the residual error of up to 74%. Ahmet Firintepe, Mohamed Selim, Alain Pagani, Didier Stricker |
IV | 3 |
| 2020 | HMDPose: A large-scale trinocular IR Augmented Reality Glasses Pose DatasetabstractAugmented Reality Glasses usually implement an Inside-Out tracking. In case of a driving scenario or glasses with less computation capabilities, an Outside-In tracking approach is required. However, to the best of our knowledge, no public datasets exist that collects images of users wearing AR glasses. To address this problem, we present HMDPose, an infrared trinocular dataset of four different AR Head-mounted displays captured in a car. It contains sequences of 14 subjects captured by three different cameras running at 60 FPS each, adding up to more than 3,000,000 labeled images in total. We provide a ground truth 6DoF-pose, captured by a submillimeter accurate marker-based tracker. We make HMDPose publicly available for non-profit, academic use and non-commercial benchmarking on ags.cs.uni-kl.de/datasets/hmdpose/. Ahmet Firintepe, Alain Pagani, Didier Stricker |
VRST | 2 |
| 2017 | Sparse-MVRVMs Tree for Fast and Accurate Head Pose Estimation in the Wild
Mohamed Selim, Alain Pagani, Didier Stricker |
CAIP (1) | 2 |
| 2017 | Eyes of ThingsabstractResponsible Research and Innovation (RRI) is an approach that anticipates and assesses potential implications and societal expectations with regard to research and innovation, with the aim to foster the design of inclusive and sustainable research and innovation. While RRI includes many aspects, in certain types of projects ethics and particularly privacy, is arguably the most sensitive topic. The objective in Horizon 2020 innovation project Eyes of Things (EoT) is to build a small high-performance, low-power, computer vision platform (similar to a smart camera) that can work independently and also embedded into all types of artefacts. In this paper, we describe the actions taken within the project related to ethics and privacy. A privacy-by-design approach has been followed, and work continues now in four platform demonstrators. Noelia Vállez, José Luis Espinosa-Aranda, Jose M. Rico-Saavedra, Javier Parra-Patino, Oscar Déniz-Suárez, Alain Pagani, Stephan Krauß, Ruben Reiser, Didier Stricker, David Moloney, Aubrey K. Dunne, Dexmont Peña, Martin Wäny, Matteo Sorci, Tim Llewellynn, Christian Fedorczak, Thierry Larmoire, Elodie Roche, Marco Herbst, Andre Seirafi, Kasra Seirafi |
IC2E | 6 |
| 2016 | Learning to Fuse: A Deep Learning Approach to Visual-Inertial Camera Pose EstimationabstractCamera pose estimation is the cornerstone of Augmented Reality applications. Pose tracking based on camera images exclusively has been shown to be sensitive to motion blur, occlusions, and illumination changes. Thus, a lot of work has been conducted over the last years on visual-inertial pose tracking using acceleration and angular velocity measurements from inertial sensors in order to improve the visual tracking. Most proposed systems use statistical filtering techniques to approach the sensor fusion problem, that require complex system modelling and calibrations in order to perform adequately. In this work we present a novel approach to sensor fusion using a deep learning method to learn the relation between camera poses and inertial sensor measurements. A long short-term memory model (LSTM) is trained to provide an estimate of the current pose based on previous poses and inertial measurements. This estimates then appropriately combined with the output of a visual tracking system using a linear Kalman Filter to provide a robust final pose estimate. Our experimental results confirm the applicability and tracking performance improvement gained from the proposed sensor fusion system. Jason R. Rambach, Aditya Tewari, Alain Pagani, Didier Stricker |
ISMAR | 3 |
| 2015 | Real-Time Head Pose Estimation Using Multi-variate RVM on Faces in the Wild
Mohamed Selim, Alain Pagani, Didier Stricker |
CAIP (2) | 2 |
| 2014 | A Superior Tracking Approach: Building a Strong Tracker through Fusion
Christian Bailer, Alain Pagani, Didier Stricker |
ECCV (7) | 2 |
| 2013 | Real-time modeling and tracking manual workflows from first-person visionabstractRecognizing previously observed actions in video sequences can lead to Augmented Reality manuals that (1) automatically follow the progress of the user and (2) can be created from video examples of the workflow. Modeling is challenging, as the environment is susceptible to change drastically due to user interaction and camera motion may not provide sufficient translation to robustly estimate geometry. We propose a piecewise homographic transform that projects the given video material onto a series of distinct planar subsets of the scene. These subsets are selected by segmenting the largest image region that is consistent with a homographic model and contains a given region of interest. We are then able to model the state of the environment and user actions using simple 2D region descriptors. The model elegantly handles estimation errors due to incomplete observation and is robust towards occlusions, e.g., due to the user's hands. We demonstrate the effectiveness of our approach quantitatively and compare it to the current state of the art. Further, we show how we apply the approach to visualize automatically assessed correctness criteria during run-time. Nils Petersen, Alain Pagani, Didier Stricker |
ISMAR | 2 |
| 2010 | SIFT in perception-based color spaceabstractScale Invariant Feature Transform (SIFT) has been proven to be the most robust local invariant feature descriptor. However, SIFT is designed mainly for grayscale images. Many local features can be misclassified if their color information is ignored. Motivated by perceptual principles, this paper addresses a new color space, called perception-based color space, in which the associated metric approximates perceived distances and color displacements and captures illumination invariant relationship. Instead of using grayscale values to represent the input image, the proposed approach builds the SIFT descriptors in the new color space, resulting in a descriptor that is more robust than the standard SIFT with respect to color and illumination variations. The evaluation results support the potential of the proposed approach. Yan Cui 0011, Alain Pagani, Didier Stricker |
ICIP | 2 |
| 2009 | Learning Local Patch Orientation with a Cascade of Sparse RegressorsabstractWe present a new method for infering the local 3D orientation of keypoints from their appearance. The method is based on the idea that the relation between keypoint appearance and pose can be learnt efficiently with an adequate regressor. Using one reference view of a keypoint, it is possible to train a keypoint-specific regressor that takes the point appearance as input and delivers the local perspective transformation as output. We show that an elegant choice of regressor is a set of sparse regressors applied sequentially in a cascade. In our case, we use a set of parametrized multivariate relevance vector machines (MVRVM) to learn the local 8-dimensional homography from the patch normalized pixel values. We show that using a cascade of regressors, ranging from coarse pose approximation to fine rectifications, considerably speeds up the identification and pose estimation process. Moreover, we show that our method improves the precision of classical points detectors, as the location of the point is rectified together with the homography. The resulting system is able to recover the orientation of patches in real time. Alain Pagani, Didier Stricker |
BMVC | 1 |
| 2009 | Integral P-channels for fast and robust region matchingabstractWe present a new method for matching a region between an input and a query image, based on the P-channel representation of pixel-based image features such as grayscale and color information, local gradient orientation and local spatial coordinates. We introduce the concept of integral P-channels, which conciliates the concepts of P-channel and integral images. Using integral images, the P-channel representation of a given region is extracted with a few arithmetic operations. This enables a fast nearest-neighbor search in all possible target regions. We present extensive experimental results and show that our approach compares favorably to existing methods for region matching such as histograms or region covariance. Alain Pagani, Didier Stricker, Michael Felsberg |
ICIP | 1 |
| 2007 | Feature Management for Efficient Camera Tracking
Harald Wuest, Alain Pagani, Didier Stricker |
ACCV (1) | 2 |