Duy-Tho Le

dblp:309/9041 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-2356-4530ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 23% 3D vision · 20% Image recognition and object detection · 18%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
multi-object tracking
1.422024
JRDB-PanoTrack: An Open-World Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments · CVPR 2024
JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking · CVPR 2023
Computer vision › 3D vision
3d scene understanding
1.012026
Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Convex Parametric Shapes · AAAI 2026
Computer vision › Image recognition and object detection › object detection
iou-based loss
1.012026
Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Convex Parametric Shapes · AAAI 2026
Computer vision › 3D vision
3d object detection
0.812024
Diffusion Model for Robust Multi-sensor Fusion in 3D Object Detection and BEV Segmentation · ECCV (68) 2024
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.812024
JRDB-PanoTrack: An Open-World Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments · CVPR 2024
Robotics › Robot navigation and mapping
sensor fusion
0.812024
Diffusion Model for Robust Multi-sensor Fusion in 3D Object Detection and BEV Segmentation · ECCV (68) 2024
Computer vision › Face, body and person analysis
human pose estimation
0.712023
JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking · CVPR 2023
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation
0.712023
JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking · CVPR 2023
Computer vision › Video understanding and tracking › object tracking
person tracking
0.712023
JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking · CVPR 2023
Computer vision › Image recognition and object detection › object detection
bounding box regression
0.312026
Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Convex Parametric Shapes · AAAI 2026
Computer vision › Image recognition and object detection
object detection
0.312026
Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Convex Parametric Shapes · AAAI 2026
Computer vision › Segmentation and scene understanding › image segmentation › scene segmentation
bird's eye view segmentation
0.212024
Diffusion Model for Robust Multi-sensor Fusion in 3D Object Detection and BEV Segmentation · ECCV (68) 2024
Robotics › Autonomous driving › perception
environment perception
0.212024
JRDB-PanoTrack: An Open-World Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments · CVPR 2024
Robotics › Robot navigation and mapping › social navigation
socially-aware navigation
0.212023
JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking · CVPR 2023

Methods — techniques the papers use, named apart from their topics

generalized iou · 1.0differentiable loss · 1.0tracking · 0.8panoptic segmentation · 0.8diffusion model · 0.8tracking benchmark · 0.7pose estimation benchmark · 0.7
YearPublicationVenuePosition
2026 Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Convex Parametric Shapes
abstract
Optimizing the similarity between parametric shapes is crucial for numerous computer vision tasks, where Intersection over Union (IoU) stands as the canonical measure. However, existing optimization methods exhibit significant shortcomings: regression-based losses like L1/L2 lack correlation with IoU, IoU-based losses are unstable and limited to simple shapes, and task-specific methods are computationally intensive and not generalizable across domains. As a result, the current landscape of parametric shape objective functions has become scattered, with each domain proposing distinct IoU approximations. To address this, we unify the parametric shape optimization objective functions by introducing Marginalized Generalized IoU (MGIoU), a novel loss function that overcomes these challenges by projecting structured convex shapes onto their unique shape Normals to compute one-dimensional normalized GIoU. MGIoU offers a simple, efficient, fully differentiable approximation strongly correlated with IoU. We extend MGIoU to MGIoU+ that supports optimizing unstructured convex shapes. Together, MGIoU and MGIoU+ unify parametric shape optimization across diverse applications. Experiments on standard benchmarks demonstrate that MGIoU and MGIoU+ demonstrate higher performance while reducing loss computation latency up to 10-40x. Also, MGIoU and MGIoU+ satisfy metric properties and scale-invariance, ensuring robustness as an objective function. We further propose MGIoU- for minimizing overlaps in tasks like collision-free trajectory prediction.
Duy-Tho Le, Trung Pham, Jianfei Cai 0001, Seyed Hamid Rezatofighi
AAAI1
2024 JRDB-PanoTrack: An Open-World Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
abstract
Autonomous robot systems have attracted increasing research attention in recent years, where environment understanding is a crucial step for robot navigation, human-robot interaction, and decision. Real-world robot systems usually collect visual data from multiple sensors and are required to recognize numerous objects and their movements in complex human-crowded settings. Traditional benchmarks, with their reliance on single sensors and limited object classes and scenarios, fail to provide the comprehensive environmental understanding robots need for accurate navigation, interaction, and decision-making. As an extension of JRDB dataset, we unveil JRDB-PanoTrack, a novel open-world panoptic segmentation and tracking benchmark, towards more comprehensive environmental perception. JRDB-PanoTrack includes (1) various data involving indoor and outdoor crowded scenes, as well as comprehensive 2D and 3D synchronized data modalities; (2) high-quality 2D spatial panoptic segmentation and temporal tracking annotations, with additional 3D label projections for further spatial understanding; (3) diverse object classes for closed- and open-world recognition benchmarks, with OSPA-based metrics for evaluation. Extensive evaluation of leading methods shows significant challenges posed by our dataset.
Duy-Tho Le, Chenhui Gou, Stavya Datta, Hengcan Shi, Ian D. Reid 0001, Jianfei Cai 0001, Seyed Hamid Rezatofighi
CVPR1
2024 Diffusion Model for Robust Multi-sensor Fusion in 3D Object Detection and BEV Segmentation
Duy-Tho Le, Hengcan Shi, Jianfei Cai 0001, Seyed Hamid Rezatofighi
ECCV (68)1
2023 JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and Tracking
abstract
Autonomous robotic systems operating in human environments must understand their surroundings to make accurate and safe decisions. In crowded human scenes with close-up human-robot interaction and robot navigation, a deep understanding of surrounding people requires reasoning about human motion and body dynamics over time with human body pose estimation and tracking. However, existing datasets captured from robot platforms either do not provide pose annotations or do not reflect the scene distribution of social robots. In this paper, we introduce JRDB-Pose, a large-scale dataset and benchmark for multi-person pose estimation and tracking. JRDB-Pose extends the existing JRDB which includes videos captured from a social navigation robot in a university campus environment, containing challenging scenes with crowded indoor and outdoor locations and a diverse range of scales and occlusion types. JRDB-Pose provides human pose annotations with per-keypoint occlusion labels and track IDs consistent across the scene and with existing annotations in JRDB. We conduct a thorough experimental study of state-of-the-art multi-person pose estimation and tracking methods on JRDB-Pose, showing that our dataset imposes new challenges for the existing methods. JRDB-Pose is available at https://jrdb.erc.monash.edu/.
Edward Vendrow, Duy-Tho Le, Jianfei Cai 0001, Seyed Hamid Rezatofighi
CVPR2