VLDB 2026 Research / reviewers in the wild / expert
Bertrand Luvison
dblp:44/4049
· DBLP profile ↗
11ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HUI360 : A 360° Egocentric Dataset and Baselines for Human-Robot Interaction AnticipationabstractAs robots increasingly operate in human-populated environments, anticipating human intentions is essential for enabling proactive and socially aware behavior. Automatic anticipation of human-robot interactions is thus emerging as a crucial perception challenge for embodied agents. To this end, we introduce HUI360, the largest dataset for human-robot interaction anticipation in the wild and its set of baselines. The dataset was collected from a mobile robot, in the wild, over multiple days within a 3-month period, and in several environments, capturing natural, spontaneous behaviors from both passersby and users, and encompassing a diverse range of individuals. This variety enables evaluating and improving the generalization capabilities of interaction anticipation models. We designed a pipeline and share code for automatic interaction annotation in arbitrary 360-degree equirectangular videos, along with interfaces for manual refinement. Using this pipeline, we release the HUI360 open set of 1M pre-processed annotations, including detailed 2D poses, facial keypoints, and segmentation masks, obtained using state-of-the-art computer vision methods and manually curated to ensure high-quality tracking and interaction annotation. Additionally, we release the raw panoptic 360-degree images captured from the robot's egocentric viewpoint (on demand, for research purpose only in compliance with GDPR). Finally, we establish benchmark baselines for interaction anticipation, including the first cross-dataset evaluations for this task: to this end, we also release 6M annotations for another existing in-the-wild outdoor dataset collected from a mobile robot (SSUP-HRI). Dataset and code can be found at https://hucebot.github.io/hui360. Raphael Lorenzo-Louis, Fabio Amadio, Bertrand Luvison, Serena Ivaldi |
FG | 3 |
| 2025 | Fairer Analysis and Demographically Balanced Face Generation for Fairer Face VerificationabstractFace recognition and verification are two computer vision tasks whose performances have advanced with the introduction of deep representations. However, ethical, legal, and technical challenges due to the sensitive nature of face data and biases in real-world training datasets hinder their development. Generative AI addresses privacy by creating fictitious identities, but fairness problems remain. Using the existing DCFace SOTA framework, we introduce a new controlled generation pipeline that improves fairness. Through classical fairness metrics and a proposed indepth statistical analysis based on logit models and ANOVA, we show that our generation pipeline improves fairness more than other bias mitigation approaches while slightly improving raw performance. Alexandre Fournier-Montgieux, Michaël Soumm, Adrian Popescu 0001, Bertrand Luvison, Hervé Le Borgne |
WACV | 4 |
| 2024 | Uncalibrated Multi-view 3D Human Pose Estimation with Geometry Driven AttentionabstractTo make up for the inherent challenging nature of 3D pose estimation, most multi-view frameworks rely on camera calibration, often leading to impractical or constrained architectures. Accurate human pose estimation is key to en-hancing human-computer interaction, gaming, health, sport and surveillance systems. By capturing precise and reliable body positions, our approach enables efficient and innovative downstream tasks. We leverage monocular 3D pose estimations and a novel geometry driven attention mechanism inside of a transformer lightweight architecture to produce high precision, occlusion aware refined 3D poses, with varying number of uncalibrated cameras. Our method shows competitive results on the in-lab dataset Human3.6M and in the in-the-wild environment of SkiPose PTZ-Camera, both in camera frames or in a disentangled person centric referential allowing practical downstream uses. Our approach matches state-of-the-art performance on Human3.6M, while being at least 3 times lighter. On the SkiPose base acquired under particularly difficult conditions, our results exceed those of the state of the art by being at least 3 times faster. Victor Galizzi, Bertrand Luvison |
FG | 2 |
| 2023 | Feature Space Data Augmentation for Viewpoint-Robust Action Recognition in VideosabstractThe ongoing research on human action recognition models is achieving very promising results, and the existing models reach very high performances. However, they still suffer from one major challenge: their performance decreases on viewpoints not seen in the training step of the model. In this paper, we introduce a new approach based on virtual viewpoint augmentation in the feature space to increase the robustness of the action recognition models to different camera viewpoints. This approach was evaluated on two action recognition datasets: DAHLIA and Toyota SmartHome. Our model shows promising results, with a significant performance increase on both datasets for viewpoints not seen during the training step. Carla Geara, Aleksandr Setkov, Astrid Orcesi, Bertrand Luvison |
ICIP | 4 |
| 2022 | Robot Companion, an intelligent interactive robot coworker for the Industry 5.0abstractTo overcome the limitations of the so-called Industry 4.0 focusing on mass production and full automation, a novel paradigm was recently introduced, namely Industry 5.0, which aims at an increased collaboration between humans and machines, and particularly robots, instead of replacing the former with the latter. This challenge requires novel interactive intelligent robots able to perform complex tasks easily and efficiently and to collaborate on the fly with humans whenever required, be it for training or working. In this work, the Robot Companion, a novel demonstrator of this paradigm, is introduced. It combines robotics, Artificial Intelligence, software engineering and embedded systems technologies, and targets industrial assembly tasks. First tests show that this robot can efficiently assemble a representative gear system autonomously or in collaboration with human operators. Florian Gosselin, Selma Kchir, G. Acher, François Keith, Olivier Lebec, C. Louison, Bertrand Luvison, Fabrice Mayran de Chamisso, Boris Meden, Matteo Morelli, Benoit Perochon, Jaonary Rabarisoa, Caroline Vienne, G. Ameyugo |
IROS | 7 |
| 2021 | Detecting Human-to-Human-or-Object (H2O) Interactions with DIABOLOabstractDetecting human interactions is crucial for human behavior analysis. Many methods have been proposed to deal with Human-to-Object Interaction (HOI) detection, i.e., detecting in an image which person and object interact together and classifying the type of interaction. However, Human-to-Human Interactions, such as social and violent interactions, are generally not considered in available HOI training datasets. As we think these types of interactions cannot be ignored and decorrelated from HOI when analyzing human behavior, we propose a new interaction dataset to deal with both types of human interactions: Human-to-Human-or-Object (H2O). In addition, we introduce a novel taxonomy of verbs, intended to be closer to a description of human body attitude in relation to the surrounding targets of interaction, and more independent of the environment. Unlike some existing datasets, we strive to avoid defining synonymous verbs when their use highly depends on the target type or requires a high level of semantic interpretation. As H2O dataset includes V-COCO images annotated with this new taxonomy, images obviously contain more interactions. This can be an issue for HOI detection methods whose complexity depends on the number of people, targets or interactions. Thus, we propose DIABOLO (Detecting Inter Actions By Only Looking Once), an efficient subject-centric single-shot method to detect all interactions in one forward pass, with constant inference time independent of image content. In addition, this multi-task network simultaneously detects all people and objects. We show how sharing a network for these tasks does not only save computation resource but also improves performance collaboratively. Finally, DIABOLO is a strong baseline for the new proposed challenge of H2O- Interaction detection, as it outperforms all state-of-the-art methods when trained and evaluated on HOI dataset V-COCO. We hope that this new dataset and new baseline will foster future research. H2O is available on https:/lkalisteo.cea.fr/. Astrid Orcesi, Romaric Audigier, Fritz Poka Toukam, Bertrand Luvison |
FG | 4 |
| 2021 | Single-shot 3D multi-person pose estimation in complex images
Abdallah Benzine, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard |
Pattern Recognit. | 2 |
| 2020 | PandaNet: Anchor-Based Single-Shot Multi-Person 3D Pose EstimationabstractRecently, several deep learning models have been proposed for 3D human pose estimation. Nevertheless, most of these approaches only focus on the single-person case or estimate 3D pose of a few people at high resolution. Furthermore, many applications such as autonomous driving or crowd analysis require pose estimation of a large number of people possibly at low-resolution. In this work, we present PandaNet (Pose estimAtioN and Dectection Anchor-based Network), a new single-shot, anchor-based and multi-person 3D pose estimation approach. The proposed model performs bounding box detection and, for each detected person, 2D and 3D pose regression into a single forward pass. It does not need any post-processing to regroup joints since the network predicts a full 3D pose for each bounding box and allows the pose estimation of a possibly large number of people at low resolution. To manage people overlapping, we introduce a Pose-Aware Anchor Selection strategy. Moreover, as imbalance exists between different people sizes in the image, and joints coordinates have different uncertainties depending on these sizes, we propose a method to automatically optimize weights associated to different people scales and joints for efficient training. PandaNet surpasses previous single-shot methods on several challenging datasets: a multi-person urban virtual but very realistic dataset (JTA Dataset), and two real world 3D multi-person datasets (CMU Panoptic and MuPoTS-3D). Abdallah Benzine, Florian Chabot, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard |
CVPR | 3 |
| 2020 | Classifying All Interacting Pairs in a Single ShotabstractIn this paper, we introduce a novel human interaction detection approach, based on CALIPSO (Classifying ALl Interacting Pairs in a Single shOt), a classifier of human-object interactions. This new single-shot interaction classifier estimates interactions simultaneously for all human-object pairs, regardless of their number and class. State-of- the-art approaches adopt a multi-shot strategy based on a pairwise estimate of interactions for a set of human-object candidate pairs, which leads to a complexity depending, at least, on the number of interactions or, at most, on the number of candidate pairs. In contrast, the proposed method estimates the interactions on the whole image. Indeed, it simultaneously estimates all interactions between all human subjects and object targets by performing a single forward pass throughout the image. Consequently, it leads to a constant complexity and computation time independent of the number of subjects, objects or interactions in the image. In detail, interaction classification is achieved on a dense grid of anchors thanks to a joint multi-task network that learns three complementary tasks simultaneously: (i) prediction of the types of interaction, (ii) estimation of the presence of a target and (iii) learning of an embedding which maps interacting subject and target to a same representation, by using a metric learning strategy. In addition, we introduce an object-centric passive-voice verb estimation which significantly improves results. Evaluations on the two well-known Human-Object Interaction image datasets, V- COCO and HICO-DET, demonstrate the competitiveness of the proposed method (2nd place) compared to the state-of- the-art while having constant computation time regardless of the number of objects and interactions in the image. Sanaa Chafik, Astrid Orcesi, Romaric Audigier, Bertrand Luvison |
WACV | 4 |
| 2019 | Deep, Robust and Single Shot 3D Multi-Person Human Pose Estimation from Monocular ImagesabstractIn this paper, we propose a new single shot method for multi-person 3D pose estimation, from monocular RGB images. Our model jointly learns to locate the human joints in the image, to estimate their 3D coordinates and to group these predictions into full human skeletons. Our approach leverages and extends the Stacked Hourglass Network and its multi-scale feature learning to manage multi-person situations. Thus, we exploit the Occlusions Robust Pose Maps (ORPM) to fully describe several 3D human poses even in case of strong occlusions or cropping. Then, joint grouping and human pose estimation for an arbitrary number of people are performed using associative embedding. We evaluate our method on the challenging CMU Panoptic dataset, and demonstrate that it achieves better results than the state of the art. Abdallah Benzine, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard |
ICIP | 2 |
| 2017 | Crowd Behavior Analysis Using Local Mid-Level Visual DescriptorsabstractCrowd behavior analysis has recently emerged as an increasingly important and dedicated problem for crowd monitoring and management in the visual surveillance community. In particular, it is receiving a lot of attention to detect potentially dangerous situations and to prevent overcrowdedness. In this paper, we propose to quantify crowd properties by a rich set of visual descriptors. The calculation of these descriptors is realized through a novel spatio-temporal model of the crowd. It consists of modeling time-varying dynamics of the crowd using local feature tracks. It also involves a Delaunay triangulation to approximate neighborhood interactions. In total, the crowd is represented as an evolving graph, where the nodes correspond to the tracklets. From this graph, various mid-level representations are extracted to determine the ongoing crowd behaviors. In particular, the effectiveness of the proposed visual descriptors is demonstrated within three applications: crowd video classification, anomaly detection, and violence detection in crowds. The obtained results on videos from different data sets prove the relevance of these visual descriptors to crowd behavior analysis. In addition, by means of comparisons to other existing methods, we demonstrate that the proposed descriptors outperform the state-of-the-art methods with a significant margin using the most challenging data sets. Hajer Fradi, Bertrand Luvison, Quoc Cuong Pham |
IEEE Trans. Circuits Syst. Video Technol. | 2 |