Yujia Zhang 0003

dblp:116/1506-3 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0006-3884-9730ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 38% Autonomous driving · 25% Representation and self-supervised learning · 25%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › geometric deep learning
3d representation learning
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Computer vision › 3D vision
3d scene understanding
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › multimodal self-supervised learning
cross-modal self-supervised learning
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Robotics › Autonomous driving
multimodal large language model for driving
0.912025
DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving · CVPR 2025
Machine learning › Reinforcement learning › imitation learning
online imitation learning
0.912025
DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving · CVPR 2025
Computer vision › 3D vision › 3d scene understanding
point cloud scene understanding
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Machine learning › Representation and self-supervised learning
spatial representation learning
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

self-distillation · 0.9multimodal large language model · 0.9multi-view visual tokenizer · 0.9imitation learning · 0.9contrastive learning · 0.9CLIP · 0.9
YearPublicationVenuePosition
2025 DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving
abstract
Multimodal large language models (MLLMs) possess the ability to comprehend visual images or videos, and show impressive reasoning ability thanks to the vast amounts of pretrained knowledge, making them highly suitable for autonomous driving applications. Unlike the previous work, DriveGPT4-V1, which focused on open-loop tasks, this study explores the capabilities of LLMs in enhancing closed-loop autonomous driving. DriveGPT4-V2 processes camera images and vehicle states as input to generate low-level control signals for end-to-end vehicle operation. A multi-view visual tokenizer (MV-VT) is employed enabling DriveGPT4-V2 to perceive the environment with an extensive range while maintaining critical details. The model architecture has been refined to improve decision prediction and inference speed. To further enhance the performance, an additional expert LLM is trained for online imitation learning. The expert LLM, sharing a similar structure with DriveGPT4-V2, can access privileged information about surrounding objects for more robust and reliable predictions. Experimental results show that DriveGPT4-V2 outperforms all baselines on the challenging CARLA Longest6 benchmark. The code and data of DriveGPT4-V2 will be publicly available.
Zhenhua Xu 0003, Yujia Zhang 0003, Zhuoling Li, Kwan-Yee Kenneth Wong, Hengshuang Zhao
CVPR3
2025 Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
abstract
Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation of human concept learning for spatial cognition, combining 3D intra-modal self-distillation with 2D-3D cross-modal joint embedding. Despite its simplicity, Concerto learns more coherent and informative spatial features, as demonstrated by zero-shot visualizations. It outperforms both standalone SOTA 2D and 3D self-supervised models by 14.2\% and 4.8\%, respectively, as well as their feature concatenation, in linear probing for 3D scene perception. With full fine-tuning, Concerto sets new SOTA results across multiple scene understanding benchmarks (e.g., 80.7\% mIoU on ScanNet). We further present a variant of Concerto tailored for video-lifted point cloud spatial understanding, and a translator that linearly projects Concerto representations into CLIP’s language space, enabling open-world perception. These results highlight that Concerto emerges spatial representations with superior fine-grained geometric and semantic consistency.
Yujia Zhang 0003, Xiaoyang Wu 0002, Yixing Lao, Chengyao Wang, Zhuotao Tian, Naiyan Wang, Hengshuang Zhao
NeurIPS1