Yuchen Yang 0003

dblp:06/7124-3 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0002-2907-6458ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 39% Video understanding and tracking · 24% Face, body and person analysis · 16%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.322023
DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds · ICCV 2023
LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal Fusion · CVPR 2023
Computer vision › Face, body and person analysis › human pose estimation
articulated pose estimation
1.012026
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis · AAAI 2026
Computer vision › Video understanding and tracking › object tracking › single object tracking
ball tracking
1.012026
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis · AAAI 2026
Computer vision › Video understanding and tracking
object tracking
1.012026
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis · AAAI 2026
Computer vision › 3D vision
pose estimation
1.012026
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis · AAAI 2026
Robotics › Autonomous driving
trajectory prediction
1.012026
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis · AAAI 2026
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.812024
Mask as Supervision: Leveraging Unified Mask Information for Unsupervised 3D Pose Estimation · ECCV (44) 2024
Computer vision › 3D vision › multimodal perception
LiDAR-camera fusion
0.712023
LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal Fusion · CVPR 2023
Computer vision › Video understanding and tracking
multi-object tracking
0.712023
DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds · ICCV 2023
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.712023
UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase · ICCV 2023
Computer vision › 3D vision
point cloud segmentation
0.712023
UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase · ICCV 2023
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
0.712023
UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase · ICCV 2023
Visualization and visual analytics › visual analytics
sports analytics
0.312026
RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis · AAAI 2026
Machine learning › Learning paradigms
unsupervised learning
0.212024
Mask as Supervision: Leveraging Unified Mask Information for Unsupervised 3D Pose Estimation · ECCV (44) 2024
Computer vision › Vision and language
multimodal fusion
0.212023
UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase · ICCV 2023
Robotics › Autonomous driving
perception
0.212023
LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal Fusion · CVPR 2023

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 2.0cross-attention · 2.0voxel features · 0.7multi-frame detection · 0.7local-to-global fusion · 0.7feature dynamic aggregation · 0.7decomposed regression · 0.7cross-view association · 0.7cross-modal association · 0.7attention mechanism · 0.7
YearPublicationVenuePosition
2026 RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
abstract
We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research into complex human-object interactions. It is designed to tackle three interconnected tasks: fine-grained ball tracking, articulated racket pose estimation, and predictive ball trajectory forecasting. Our evaluation of established baselines reveals a critical insight for multi-modal fusion: while naively concatenating racket pose features degrades performance, a Cross-Attention mechanism is essential to unlock their value, leading to trajectory prediction results that surpass strong unimodal baselines. RacketVision provides a versatile resource and a strong starting point for future research in dynamic object tracking, conditional motion forecasting, and multi-modal analysis in sports.
Linfeng Dong, Yuchen Yang 0003, Wei Wang 0333, Yuenan Hou, Zhihang Zhong, Xiao Sun 0001
AAAI2
2024 Mask as Supervision: Leveraging Unified Mask Information for Unsupervised 3D Pose Estimation
Yuchen Yang 0003, Yu Qiao 0001, Xiao Sun 0001
ECCV (44)1
2023 LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal Fusion
abstract
LiDAR-camera fusion methods have shown impressive performance in 3D object detection. Recent advanced multi-modal methods mainly perform global fusion, where image features and point cloud features are fused across the whole scene. Such practice lacks fine-grained region-level information, yielding suboptimal fusion performance. In this paper, we present the novel Local-to-Global fusion network (LoGoNet), which performs LiDAR-camerafusion at both local and global levels. Concretely, the Global Fusion (GoF) of LoGoNet is built upon previous literature, while we exclusively use point centroids to more precisely represent the position of voxel features, thus achieving better crossmodal alignment. As to the Local Fusion (LoF), we first divide each proposal into uniform grids and then project these grid centers to the images. The image features around the projected grid points are sampled to be fused with position-decorated point cloud features, maximally uti-lizing the rich contextual information around the proposals. The Feature Dynamic Aggregation (FDA) module is further proposed to achieve information interaction between these locally and globally fused features, thus producing more informative multi-modal features. Extensive experiments on both Waymo Open Dataset (WOD) and KITTI datasets show that LoGoNet outperforms all state-of-the-art 3D detection methods. Notably, LoGoNet ranks 1st on Waymo 3D object detection leaderboard and obtains 81.02 mAPH (L2) detection performance. It is noteworthy that, for the first time, the detection performance on three classes surpasses 80 APH (L2) simultaneously. Code will be available at https://github.com/sankin97/LoGoNet.
Xin Li 0110, Tao Ma 0002, Yuenan Hou, Botian Shi, Yuchen Yang 0003, Youquan Liu, Xingjiao Wu, Qin Chen 0001, Yikang Li 0002, Yu Qiao 0001, Liang He 0001
CVPR5
2023 UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase
abstract
Point-, voxel-, and range-views are three representative forms of point clouds. All of them have accurate 3D measurements but lack color and texture information. RGB images are a natural complement to these point cloud views and fully utilizing the comprehensive information of them benefits more robust perceptions. In this paper, we present a unified multi-modal LiDAR segmentation network, termed UniSeg, which leverages the information of RGB images and three views of the point cloud, and accomplishes semantic segmentation and panoptic segmentation simultaneously. Specifically, we first design the Learnable cross-Modal Association (LMA) module to automatically fuse voxel-view and range-view features with image features, which fully utilize the rich semantic information of images and are robust to calibration errors. Then, the enhanced voxel-view and range-view features are transformed to the point space, where three views of point cloud features are further fused adaptively by the Learnable cross-View Association module (LVA). Notably, UniSeg achieves promising results in three public benchmarks, i.e., SemanticKITTI, nuScenes, and Waymo Open Dataset (WOD); it ranks 1st on two challenges of two benchmarks, including the LiDAR semantic segmentation challenge of nuScenes and panoptic segmentation challenges of SemanticKITTI. Besides, we construct the OpenPCSeg codebase, which is the largest and most comprehensive outdoor LiDAR segmentation codebase. It contains most of the popular outdoor LiDAR segmentation algorithms and provides reproducible implementations. The OpenPCSeg codebase will be made publicly available at https://github.com/PJLab-ADG/PCSeg.
Youquan Liu, Runnan Chen, Xin Li 0110, Lingdong Kong, Yuchen Yang 0003, Zhaoyang Xia, Yeqi Bai, Xinge Zhu, Yuexin Ma, Yikang Li 0002, Yu Qiao 0001, Yuenan Hou
ICCV5
2023 DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds
abstract
Existing offboard 3D detectors always follow a modular pipeline design to take advantage of unlimited sequential point clouds. We have found that the full potential of off-board 3D detectors is not explored mainly due to two reasons: (1) the onboard multi-object tracker cannot generate sufficient complete object trajectories, and (2) the motion state of objects poses an inevitable challenge for the object-centric refining stage in leveraging the long-term temporal context representation. To tackle these problems, we propose a novel paradigm of offboard 3D object detection, named DetZero. Concretely, an offline tracker coupled with a multi-frame detector is proposed to focus on the completeness of generated object tracks. An attention-mechanism refining module is proposed to strengthen contextual information interaction across long-term sequential point clouds for object refining with decomposed regression methods. Extensive experiments on Waymo Open Dataset show our DetZero outperforms all state-of-the-art onboard and offboard 3D detection methods. Notably, DetZero ranks 1st place on Waymo 3D object detection leaderboard1with 85.15 mAPH (L2) detection performance. Further experiments validate the application of taking the place of human labels with such high-quality results. Our empirical study leads to rethinking conventions and interesting findings that can guide future research on offboard 3D object detection.
Tao Ma 0002, Xuemeng Yang, Hongbin Zhou, Xin Li 0110, Botian Shi, Yuchen Yang 0003, Zhizheng Liu, Liang He 0001, Yu Qiao 0001, Yikang Li 0002, Hongsheng Li 0001
ICCV7