Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chunyong Hu

dblp:288/2242 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Autonomous driving · 33% Efficient and distributed learning · 31% 3D vision · 21%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
perception
1.922026
GUIDE: Gaussian Unified Instance Detection for Enhanced Obstacle Perception in Autonomous Driving · AAAI 2026
SAM4D: Segment Anything in Camera and LiDAR Streams · ICCV 2025
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.912025
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning · CVPR 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning · CVPR 2025
Computer vision › 3D vision › point cloud analysis
point cloud learning
0.912025
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning · CVPR 2025
Computer vision › Segmentation and scene understanding
prompt-based segmentation
0.912025
SAM4D: Segment Anything in Camera and LiDAR Streams · ICCV 2025
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.312026
GUIDE: Gaussian Unified Instance Detection for Enhanced Obstacle Perception in Autonomous Driving · AAAI 2026

Methods — techniques the papers use, named apart from their topics

gaussian-to-voxel splatting · 1.03d gaussian splatting · 1.0prompt tuning · 0.9multimodal positional encoding · 0.9multi-scale token selection · 0.9cross-modal memory attention · 0.9automated data engine · 0.9LoRA · 0.9
YearPublicationVenuePosition
2026 GUIDE: Gaussian Unified Instance Detection for Enhanced Obstacle Perception in Autonomous Driving
abstract
In the realm of autonomous driving, accurately detecting surrounding obstacles is crucial for effective decision-making. Traditional methods primarily rely on 3D bounding boxes to represent these obstacles, which often fail to capture the complexity of irregularly shaped, real-world objects. To overcome these limitations, we present GUIDE, a novel framework that utilizes 3D Gaussians for instance detection and occupancy prediction. Unlike conventional occupancy prediction methods, GUIDE also offers robust tracking capabilities. Our framework employs a sparse representation strategy, using Gaussian-to-Voxel Splatting to provide fine-grained, instance-level occupancy data without the computational demands associated with dense voxel grids. Experimental validation on the nuScenes dataset demonstrates GUIDE's performance, with an instance occupancy mAP of 21.61, marking a 50% improvement over existing methods, alongside competitive tracking capabilities. GUIDE establishes a new benchmark in autonomous perception systems, effectively combining precision with computational efficiency to better address the complexities of real-world driving environments.
Chunyong Hu, Jianyun Xu, Song Wang 0019, Sheng Yang 0007
AAAI1
2025 PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning
abstract
Self-supervised representation learning for point cloud has demonstrated effectiveness in improving pre-trained model performance across diverse tasks. However, as pre-trained models grow in complexity, fully fine-tuning them for downstream applications demands substantial computational and storage resources. Parameter-efficient fine-tuning (PEFT) methods offer a promising solution to mitigate these resource requirements, yet most current approaches rely on complex adapter and prompt mechanisms that increase tunable parameters. In this paper, we propose PointLoRA, a simple yet effective method that combines low-rank adaptation (LoRA) with multi-scale token selection to efficiently fine-tune point cloud models. Our approach embeds LoRA layers within the most parameter-intensive components of point cloud transformers, reducing the need for tunable parameters while enhancing global feature capture. Additionally, multi-scale token selection extracts critical local information to serve as prompts for downstream fine-tuning, effectively complementing the global context captured by LoRA. The experimental results across various pre-trained models and three challenging public datasets demonstrate that our approach achieves competitive performance with only 3.43% of the trainable parameters, making it highly effective for resource-constrained applications. Source code is available at: https://github.com/songw-zju/PointLoRA.
Song Wang 0019, Lingdong Kong, Jianyun Xu, Chunyong Hu, Gongfan Fang, Wentong Li 0001, Jianke Zhu, Xinchao Wang
CVPR5
2025 SAM4D: Segment Anything in Camera and LiDAR Streams
abstract
We present SAM4D, a multi-modal and temporal foundation model designed for promptable segmentation across camera and LiDAR streams. Unified Multi-modal Positional Encoding (UMPE) is introduced to align camera and LiDAR features in a shared 3D space, enabling seamless cross-modal prompting and interaction. Additionally, we propose Motion-aware Cross-modal Memory Attention (MCMA), which leverages ego-motion compensation to enhance temporal consistency and long-horizon feature retrieval, ensuring robust segmentation across dynamically changing autonomous driving scenes. To avoid annotation bottlenecks, we develop a multi-modal automated data engine that synergizes VFM-driven video masklets, spatiotemporal 4D reconstruction, and cross-modal masklet fusion. This framework generates camera-LiDAR aligned pseudo-labels at a speed orders of magnitude faster than human annotation while preserving VFM-derived semantic fidelity in point cloud representations. We conduct extensive experiments on the constructed Waymo-4DSeg, which demonstrate the powerful cross-modal segmentation ability and great potential in data annotation of proposed SAM4D.
Jianyun Xu, Song Wang 0019, Ziqian Ni, Chunyong Hu, Sheng Yang 0007, Jianke Zhu
ICCV4