EDBT 2026 Demo / reviewers in the wild / expert
Quoc Cuong Pham
dblp:72/5867 · also Quoc-Cuong Pham
· DBLP profile ↗
27ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0003-3032-7090ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVAT: Multi-View Aware Teacher for Weakly Supervised 3D Object DetectionabstractAnnotating 3D data remains a costly bottleneck for 3D object detection, motivating the development of weakly supervised annotation methods that rely on more accessible 2D box annotations. However, relying solely on 2D boxes introduces projection ambiguities since a single 2D box can correspond to multiple valid 3D poses. Furthermore, partial object visibility under a single viewpoint setting makes accurate 3D box estimation difficult. We propose MVAT, a novel framework that leverages temporal multi-view present in sequential data to address these challenges. Our approach aggregates object-centric point clouds across time to build 3D object representations as dense and complete as possible. A Teacher-Student distillation paradigm is employed: The Teacher network learns from single viewpoints but targets are derived from temporally aggregated static objects. Then the Teacher generates high quality pseudo-labels that the Student learns to predict from a single viewpoint for both static and moving objects. The whole framework incorporates a multi-view 2D projection loss to enforce consistency between predicted 3D boxes and all available 2D annotations. Experiments on the nuScenes and Waymo Open datasets demonstrate that MVAT achieves state-of-the-art performance for weakly supervised 3D object detection, significantly narrowing the gap with fully supervised methods without requiring any 3D box annotations. % \footnote{Code available upon acceptance} Our code is available in our public repository (\href{https://github.com/CEA-LIST/MVAT}{code}). Saad Lahlali, Alexandre Fournier-Montgieux, Nicolas Granger 0001, Hervé Le Borgne, Quoc Cuong Pham |
WACV | 5 |
| 2025 | Cross-Modal Distillation for 2D/3D Multi-Object Discovery from 2D MotionabstractObject discovery, which refers to the task of localizing objects without human annotations, has gained significant attention in 2D image analysis. However, despite this growing interest, it remains under-explored in 3D data, where approaches rely exclusively on 3D motion, despite its several challenges. In this paper, we present a novel framework that leverages advances in 2D object discovery which are based on 2D motion to exploit the advantages of such motion cues being more flexible and generalizable and to bridge the gap between 2D and 3D modalities. Our primary contributions are twofold: (i) we introduce DIOD-3D, the first baseline for multi-object discovery in 3D data using 2D motion, incorporating scene completion as an auxiliary task to enable dense object localization from sparse input data; (ii) we develop xMOD, a cross-modal training framework that integrates 2D and 3D data while always using 2D motion cues. xMOD employs a teacher-student training paradigm across the two modalities to mitigate confirmation bias by leveraging the domain gap. During inference, the model supports both RGB-only and point cloud-only inputs. Additionally, we propose a late-fusion technique tailored to our pipeline that further enhances performance when both modalities are available at inference. We evaluate our approach extensively on synthetic (TRIP-PD) and challenging real-world datasets (KITTI and Waymo). Notably, our approach yields a substantial performance improvement compared with the 2D object discovery state-of-the-art on all datasets with gains ranging from +8.7 to +15.1 in F 1@50 score. The code is available at https://github.com/CEA-LIST/xMOD Saad Lahlali, Sandra Kara, Hejer Ammar, Florian Chabot, Nicolas Granger 0001, Hervé Le Borgne, Quoc Cuong Pham |
CVPR | 7 |
| 2025 | Explainable Speech Emotion Recognition Through Attentive Pooling: Insights from Attention-Based Temporal LocalizationabstractInternational audience Tahitoa Leygue, Astrid Sabourin, Christian Bolzmacher, Sylvain Bouchigny, Margarita Anastassova, Quoc Cuong Pham |
INTERSPEECH | 6 |
| 2025 | ALPI: Auto-Labeller with Proxy Injection for 3D Object Detection using 2D Labels Onlyabstract3D object detection plays a crucial role in various applications such as autonomous vehicles, robotics and augmented reality. However, training 3D detectors requires a costly precise annotation, which is a hindrance to scaling annotation to large datasets. To address this challenge, we propose a weakly supervised 3D annotator that relies solely on 2D bounding box annotations from images, along with size priors. One major problem is that supervising a 3D detection model using only 2D boxes is not reliable due to ambiguities between different 3D poses and their identical 2D projection. We introduce a simple yet effective and generic solution: we build 3D proxy objects with annotations by construction and add them to the training dataset. Our method requires only size priors to adapt to new classes. To better align 2D supervision with 3D detection, our method ensures depth invariance with a novel expression of the 2D losses. Finally, to detect more challenging instances, our annotator follows an offline pseudo-labelling scheme which gradually improves its 3D pseudo-labels. Extensive experiments on the KITTI dataset demonstrate that our method not only performs on-par or above previous works on the Car category, but also achieves performance close to fully supervised methods on more challenging classes. We further demonstrate the effectiveness and robustness of our method by being the first to experiment on the more challenging nuScenes dataset. We additionally propose a setting where weak labels are obtained from a 2D detector pre-trained on MS-COCO instead of human annotations. The code is available at https://github.com/CEA-LIST/ALPI Saad Lahlali, Nicolas Granger 0001, Hervé Le Borgne, Quoc Cuong Pham |
WACV | 4 |
| 2024 | DIOD: Self-Distillation Meets Object DiscoveryabstractInstance segmentation demands substantial labeling re-sources. This has prompted increased interest to explore the object discovery task as an unsupervised alternative. In particular, promising results were achieved in localizing instances using motion supervision only. However, the motion signal introduces complexities due to its inherent noise and sparsity, which constrains the effectiveness of current methodologies. In the present paper we propose DIOD (self DIstillation meets Object Discovery), the first method that places the motion-guided object discovery within a framework of continuous improvement through knowledge distillation, providing solutions to existing limitations ($i$) DIOD robustly eliminates the noise present in the exploited motion maps providing accurate motion-supervision (ii) DIOD leverages the discovered objects within an it-erative pseudo-labeling framework, enriching the initial motion-supervision with static objects, which results in a cost-efficient increase in performance. Through experiments on synthetic and real-world datasets, we demonstrate the benefits of bridging the gap between object discovery and distillation, by significantly improving the state-of-the-art. This enhancement is also sustained across other demanding metrics so far reserved for supervised tasks. https://github.com/CEA-LIST/DIOD Sandra Kara, Hejer Ammar, Julien Denize, Florian Chabot, Quoc Cuong Pham |
CVPR | 5 |
| 2024 | 3D-COCO: Extension of MS-COCO Dataset for Scene Understanding and 3D ReconstructionabstractWe introduce 3D-COCO, an extension of the original MS-COCO [1] dataset providing 3D models and 2D-3D alignment annotations. 3D-COCO was designed to achieve computer vision tasks such as 3D reconstruction or image detection configurable with textual, 2D image, and 3D CAD model queries. We complete the existing MS-COCO [1] dataset with 28 K 3D models collected on ShapeNet [2] and Objaverse [3]. By using an IoU-based method, we match each MS-COCO [1] annotation with the best 3D models to provide a 2D-3D alignment. The open-source nature of $3 \mathrm{D}-\mathrm{COCO}$ is a premiere that should pave the way for new research on 3D-related topics. The dataset and its source codes is available at https://kalisteo.cea.fr/index.php/ coco3d-object-detection-and-reconstruction/ Bideaux Maxence, Phe Alice, Mohamed Chaouch, Luvison Bertrand, Quoc Cuong Pham |
ICIP | 5 |
| 2024 | The Background Also Matters: Background-Aware Motion-Guided Objects DiscoveryabstractRecent works have shown that objects discovery can largely benefit from the inherent motion information in video data. However, these methods lack a proper background processing, resulting in an over-segmentation of the non-object regions into random segments. This is a critical limitation given the unsupervised setting, where object segments and noise are not distinguishable. To address this limitation we propose BMOD, a Background-aware Motion-guided Objects Discovery method. Concretely, we leverage masks of moving objects extracted from optical flow and design a learning mechanism to extend them to the true foreground composed of both moving and static objects. The background, a complementary concept of the learned foreground class, is then isolated in the object discovery process. This enables a joint learning of the objects discovery task and the object/non-object separation. The conducted experiments on synthetic and real-world datasets show that integrating our background handling with various cutting-edge methods brings each time a considerable improvement. Specifically, we improve the objects discovery performance with a large margin, while establishing a strong baseline for object/non-object separation. Sandra Kara, Hejer Ammar, Florian Chabot, Quoc Cuong Pham |
WACV | 4 |
| 2023 | Generalized Pseudo-Labeling in Consistency Regularization for Semi-Supervised LearningabstractSemi-Supervised Learning (SSL) reduces annotation cost by exploiting large amounts of unlabeled data. A popular idea in SSL image classification is Pseudo-Labeling (PL), where the predictions of a network are used in order to assign a label to an unlabeled image. However, this practice exposes learning to confirmation bias. In this paper we propose Generalized Pseudo-Labeling (GPL), a simple and generic way to exploit negative pseudo-labels in consistency regularization, entailing minimal additional computational overhead and hyperpameter fine-tuning. GPL makes learning more robust by using the information that an image does not belong to a certain class, which is more abundant and reliable. We showcase GPL in the context of FixMatch. In the benchmark using only 40 labels of the CIFAR-10 dataset, adding GPL on top of FixMatch improves the error rate from 7.93% to 6.58%, and on CIFAR-100 with 2500 labels, from 28.02% to 26.85%. Nikolaos Karaliolios, Florian Chabot, Camille Dupont, Hervé Le Borgne, Quoc Cuong Pham, Romaric Audigier |
ICIP | 5 |
| 2023 | Image Segmentation-based Unsupervised Multiple Objects DiscoveryabstractUnsupervised object discovery aims to localize objects in images, while removing the dependence on annotations required by most deep learning-based methods. To address this problem, we propose a fully unsupervised, bottom-up approach, for multiple objects discovery. The proposed approach is a two-stage framework. First, instances of object parts are segmented by using the intra-image similarity between self-supervised local features. The second step merges and filters the object parts to form complete object instances. The latter is performed by two CNN models that capture semantic information on objects from the entire dataset. We demonstrate that the pseudo-labels generated by our method provide a better precision-recall trade-off than existing single and multiple objects discovery methods. In particular, we provide state-of-the-art results for both unsupervised class-agnostic object detection and unsupervised image segmentation. Sandra Kara, Hejer Ammar, Florian Chabot, Quoc Cuong Pham |
WACV | 4 |
| 2021 | Evaluating Robustness over High Level Driving Instruction for Autonomous DrivingabstractIn recent years, we have witnessed increasingly high performance in the field of autonomous end-to-end driving. In particular, more and more research is being done on driving in urban environments, where the car has to follow high level commands to navigate. However, few evaluations are made on the ability of these agents to react in an unexpected situation. Specifically, no evaluations are conducted on the robustness of driving agents in the event of a bad high-level command. We propose here an evaluation method, namely a benchmark that allows to assess the robustness of an agent, and to appreciate its understanding of the environment through its ability to keep a safe behavior, regardless of the instruction. Florence Carton, David Filliat, Jaonary Rabarisoa, Quoc Cuong Pham |
IV | 4 |
| 2021 | UCP-Net: Unstructured Contour Points for Instance SegmentationabstractThe goal of interactive segmentation is to assist users in producing segmentation masks as fast and as accurately as possible. Interactions have to be simple and intuitive and the number of interactions required to produce a satisfactory segmentation mask should be as low as possible. In this paper, we propose a novel approach to interactive segmentation based on unconstrained contour clicks for initial segmentation and segmentation refinement. Our method is class-agnostic and produces accurate segmentation masks (IoU > 85%) for a lower number of user interactions than state-of-the-art methods on popular segmentation datasets (COCO MVal, SBD and Berkeley). Camille Dupont, Yanis Ouakrim, Quoc Cuong Pham |
SMC | 3 |
| 2021 | Single-shot 3D multi-person pose estimation in complex images
Abdallah Benzine, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard |
Pattern Recognit. | 3 |
| 2020 | PandaNet: Anchor-Based Single-Shot Multi-Person 3D Pose EstimationabstractRecently, several deep learning models have been proposed for 3D human pose estimation. Nevertheless, most of these approaches only focus on the single-person case or estimate 3D pose of a few people at high resolution. Furthermore, many applications such as autonomous driving or crowd analysis require pose estimation of a large number of people possibly at low-resolution. In this work, we present PandaNet (Pose estimAtioN and Dectection Anchor-based Network), a new single-shot, anchor-based and multi-person 3D pose estimation approach. The proposed model performs bounding box detection and, for each detected person, 2D and 3D pose regression into a single forward pass. It does not need any post-processing to regroup joints since the network predicts a full 3D pose for each bounding box and allows the pose estimation of a possibly large number of people at low resolution. To manage people overlapping, we introduce a Pose-Aware Anchor Selection strategy. Moreover, as imbalance exists between different people sizes in the image, and joints coordinates have different uncertainties depending on these sizes, we propose a method to automatically optimize weights associated to different people scales and joints for efficient training. PandaNet surpasses previous single-shot methods on several challenging datasets: a multi-person urban virtual but very realistic dataset (JTA Dataset), and two real world 3D multi-person datasets (CMU Panoptic and MuPoTS-3D). Abdallah Benzine, Florian Chabot, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard |
CVPR | 4 |
| 2019 | Deep, Robust and Single Shot 3D Multi-Person Human Pose Estimation from Monocular ImagesabstractIn this paper, we propose a new single shot method for multi-person 3D pose estimation, from monocular RGB images. Our model jointly learns to locate the human joints in the image, to estimate their 3D coordinates and to group these predictions into full human skeletons. Our approach leverages and extends the Stacked Hourglass Network and its multi-scale feature learning to manage multi-person situations. Thus, we exploit the Occlusions Robust Pose Maps (ORPM) to fully describe several 3D human poses even in case of strong occlusions or cropping. Then, joint grouping and human pose estimation for an arbitrary number of people are performed using associative embedding. We evaluate our method on the challenging CMU Panoptic dataset, and demonstrate that it achieves better results than the state of the art. Abdallah Benzine, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard |
ICIP | 3 |
| 2017 | Crowd Behavior Analysis Using Local Mid-Level Visual DescriptorsabstractCrowd behavior analysis has recently emerged as an increasingly important and dedicated problem for crowd monitoring and management in the visual surveillance community. In particular, it is receiving a lot of attention to detect potentially dangerous situations and to prevent overcrowdedness. In this paper, we propose to quantify crowd properties by a rich set of visual descriptors. The calculation of these descriptors is realized through a novel spatio-temporal model of the crowd. It consists of modeling time-varying dynamics of the crowd using local feature tracks. It also involves a Delaunay triangulation to approximate neighborhood interactions. In total, the crowd is represented as an evolving graph, where the nodes correspond to the tracklets. From this graph, various mid-level representations are extracted to determine the ongoing crowd behaviors. In particular, the effectiveness of the proposed visual descriptors is demonstrated within three applications: crowd video classification, anomaly detection, and violence detection in crowds. The obtained results on videos from different data sets prove the relevance of these visual descriptors to crowd behavior analysis. In addition, by means of comparisons to other existing methods, we demonstrate that the proposed descriptors outperform the state-of-the-art methods with a significant margin using the most challenging data sets. Hajer Fradi, Bertrand Luvison, Quoc Cuong Pham |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Bidirectional sparse representations for multi-shot person re-identificationabstractWith the development of surveillance cameras, person re-identification has gained much interest, however re-identifying people across cameras remains a challenging problem which not only requires a good feature description but also a reliable matching scheme. Our method can be applied with any feature and focuses on the second requirement. We propose a robust bidirectional sparse coding method that improves simple sparse coding performances. Some recent work have already explored sparse representation for the re-identification task but none has considered the problem from both the probe and the gallery perspectives. We propose a bidirectional sparse representations method which searches for the most likely match for the test element in the gallery set and makes sure that the selected gallery match is indeed closely related to the probe. Extensive experiments on two datasets, CUHK03 and iLIDS-VID, show the effectiveness of our approach. Solene Chan-Lang, Quoc Cuong Pham, Catherine Achard |
AVSS | 2 |
| 2016 | RIMOC, a feature to discriminate unstructured motions: Application to violence detection for video-surveillance
Pedro Ribeiro 0006, Romaric Audigier, Quoc Cuong Pham |
Comput. Vis. Image Underst. | 3 |
| 2015 | Collaborative Tracking and Distributed Control for an IP-PTZ Camera Network
Pierrick Paillet, Romaric Audigier, Frédéric Lerasle, Quoc Cuong Pham |
ICPRAM (2) | 4 |
| 2014 | Fast and accurate video annotation using dense motion hypothesesabstractBuilding large video datasets is a crucial task for many applications but is also very expensive in practice. In order to avoid annotating all the frames, the annotations from the labeled frames can be propagated using an offline tracker for each object. Following methods based on dynamic programming and eventually distance transforms, we introduce a new penalization which favors some given displacements between two frames without increasing the complexity of the optimization. In order to speed up this step we also propose to use an exact coarse to fine process. Experimental results show that the proposed energy performs better than previous ones and that our exact coarse to fine optimization leads to a significant speed-up in some scenarios. Loïc Fagot-Bouquet, Jaonary Rabarisoa, Quoc Cuong Pham |
ICIP | 3 |
| 2014 | Joint hierarchical learning for efficient multi-class object detectionabstractIn addition to multi-class classification, the multi-class object detection task consists further in classifying a dominating background label. In this work, we present a novel approach where relevant classes are ranked higher and background labels are rejected. To this end, we arrange the classes into a tree structure where the classifiers are trained in a joint framework combining ranking and classification constraints. Our convex problem formulation naturally allows to apply a tree traversal algorithm that searches for the best class label and progressively rejects background labels. We evaluate our approach on the PASCAL VOC 2007 dataset and show a considerable speed-up of the detection time with increased detection performance. Hamidreza Odabai Fard, Mohamed Chaouch, Quoc Cuong Pham, Antoine Vacavant, Thierry Chateau |
WACV | 3 |
| 2013 | IMM-Based Tracking and Latency Control with Off-the-Shelf IP PTZ Camera
Pierrick Paillet, Romaric Audigier, Frédéric Lerasle, Quoc Cuong Pham |
ACIVS | 4 |
| 2010 | Background subtraction adapted to PTZ cameras by keypoint density estimationabstractConstant Guillot1 [email protected] Maxime Taron1 [email protected] Patrick Sayd1 [email protected] Quoc-Cuong Pham1 [email protected] Christophe Tilmant2 [email protected] Jean-Marc Lavest2 [email protected] 1 CEA LIST Laboratoire Vision et Ingenierie des Contenus, BP 94, Gif-sur-Yvette, F-91191 France 2 LASMEA UMR 6602, PRES Clermont Universite/CNRS, 63177 Aubiere cedex, France Constant Guillot, Maxime Taron, Patrick Sayd, Quoc Cuong Pham, Christophe Tilmant, Jean-Marc Lavest |
BMVC | 4 |
| 2008 | Video Monitoring of Vulnerable People in Home Environment
Quoc Cuong Pham, Yoann Dhome, Laetitia Gond, Patrick Sayd |
ICOST | 1 |
| 2007 | Real-Time Posture Analysis in a Crowd using Thermal ImagingabstractThis article describes a video-surveillance system developed within the ISCAPS project. Thermal imaging provides a robust solution to visibility change (illumination, smoke) and is a relevant technology for discriminating humans in complex scenes. In this article, we demonstrate its efficiency for posture analysis in dense groups of people. The objective is to automatically detect several persons lying down in a very crowded area. The presented method is based on the detection and segmentation of individuals within groups of people using a combination of several weak classifiers. The classification of extracted silhouettes enables to detect abnormal situations. This approach was successfully applied to the detection of terrorist gas attacks on railway platform and experimentally validated in the project. Some of the results are presented here. Quoc Cuong Pham, Laetitia Gond, Julien Begard, Nicolas Allezard, Patrick Sayd |
CVPR | 1 |
| 2005 | A Practical Guide to Marker Based and Hybrid Visual Registration for AR Industrial Applications
Steve Bourgeois, Hanna Martinsson, Quoc Cuong Pham, Sylvie Naudet |
CAIP | 3 |
| 2003 | A 3-D model-based registration approach for the PET, MR and MCG cardiac data fusion
Timo Mäkelä, Quoc Cuong Pham, Patrick Clarysse, Jukka Nenonen, Jyrki Lötjönen, Outi Sipilä, Helena Hänninen, Kirsi Lauerma, Juhani Knuuti, Toivo Katila, Isabelle E. Magnin |
Medical Image Anal. | 2 |
| 2002 | A Review of Cardiac Image Registration MethodsabstractIn this paper, the current status of cardiac image registration methods is reviewed. The combination of information from multiple cardiac image modalities, such as magnetic resonance imaging, computed tomography, positron emission tomography, single-photon emission computed tomography, and ultrasound, is of increasing interest in the medical community for physiologic understanding and diagnostic purposes. Registration of cardiac images is a more complex problem than brain image registration because the heart is a nonrigid moving organ inside a moving body. Moreover, as compared to the registration of brain images, the heart exhibits much fewer accurate anatomical landmarks. In a clinical context, physicians often mentally integrate image information from different modalities. Automatic registration, based on computer programs, might, however, offer better accuracy and repeatability and save time. Timo Mäkelä, Patrick Clarysse, Outi Sipilä, Nicoleta Pauna, Quoc Cuong Pham, Toivo Katila, Isabelle E. Magnin |
IEEE Trans. Medical Imaging | 5 |