Dingyuan Zhang

dblp:302/2373 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0001-5022-8172ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PANDA: Empowering Small Language Models for Proactive Dialogue Through Agent-Based Synthesis (Student Abstract)
abstract
Proactive dialogue systems, which are designed to guide conversations toward predetermined goals. However, contemporary LLMs predominantly function as passive assistants, mechanically executing human instructions. A key challenge contributing to this limitation is the inherent difficulty in acquiring and annotating high-quality training data for proactive dialogue. Consequently, the scarcity of such data results in a notable deficiency in the proactive conversational capabilities of current LLMs.In this paper, we introduce PANDA (Proactive Agent-based Negotiation Dialogue Augmentation), a method designed to generate accurate, complex, and diverse proactive dialogue data for a challenging task—financial dispute mediation—where a LLM acts as the mediator. PANDA leverages a novel self-evolving synthesis process to manage a pool of user profiles and generate dialogues through structured interactions between multiple LLM-driven agents. To ensure data fidelity, we propose a comprehensive evaluation framework and build a two-level validation system combining automated and expert human verification. Our experiments demonstrate that an 8B-parameter model, trained on our synthesized dataset, achieves state-of-the-art results in the task's evaluation framework. Its performance rivals top closed-source models guided by heavily engineered prompts, even when provided with only essential information.
Rongyu Zhang, Dingyuan Zhang
AAAI2
2026 Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
abstract
Large Vision-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing research typically focuses on partial objects within scenes and simple question-answer pair annotations, struggling to achieve comprehensive scene understanding. Meanwhile, existing LVLMs suffer from the lack of mapping relationship between 2D and 3D and insufficient integration of 3D spatial understanding and instruction following. To tackle these limitations, we first introduce NuInteract, a large-scale dataset with over 1.5M multi-view image-language pairs spanning dense scene captions and diverse interactive tasks. Furthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries. The spatial processor, designed as a plug-and-play component, can be initialized with pre-trained 3D detectors to provide structured geometric priors for language-conditioned 3D grounding. Our experiments show that DriveMonkey outperforms general LVLMs, especially achieving a notable 9.86% improvement on the 3D visual grounding task. The dataset and code are released at https://github.com/zc-zhao/DriveMonkey.
Zongchuang Zhao, Haoyu Fu, Dingkang Liang, Xin Zhou 0013, Dingyuan Zhang, Hongwei Xie, Xiang Bai
IEEE Trans. Image Process.5
2025 Orion: A Holistic End-To-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Dingkang Liang, Dingyuan Zhang, Hongwei Xie, Xiang Bai
ICCV7
2025 HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
abstract
Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which involves interpreting and reasoning about the driving environment. In this paper, we present a unified Driving World Model named HERMES. We seamlessly integrate 3D scene understanding and future scene evolution (generation) through a unified framework in driving scenarios. Specifically, HERMES leverages a Bird's-Eye View (BEV) representation to consolidate multi-view spatial information while preserving geometric relationships and interactions. We also introduce world queries, which incorporate world knowledge into BEV features via causal attention in the Large Language Model, enabling contextual enrichment for understanding and generation tasks. We conduct comprehensive studies on nuScenes and OmniDrive-nuScenes datasets to validate the effectiveness of our method. HERMES achieves state-of-the-art performance, reducing generation error by 32.4% and improving understanding metrics such as CIDEr by 8.0%. The model and code will be publicly released at https://github.com/LMD0311/HERMES.
Xin Zhou 0013, Dingkang Liang, Sifan Tu, Xiwu Chen, Yikang Ding, Dingyuan Zhang, Feiyang Tan, Hengshuang Zhao, Xiang Bai
ICCV6
2025 AVS-Net: Point sampling with adaptive voxel size for 3D scene understanding
Hongcheng Yang, Dingkang Liang, Dingyuan Zhang, Zhe Liu 0033, Zhikang Zou, Xingyu Jiang 0005, Yingying Zhu 0005
Neurocomputing3
2024 Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression
Dingyuan Zhang, Dingkang Liang, Zichang Tan, Xiaoqing Ye, Cheng Zhang 0020, Jingdong Wang 0001, Xiang Bai
ECCV (47)1
2024 SAM3D: zero-shot 3D object detection via the segment anything model
Dingyuan Zhang, Dingkang Liang, Hongcheng Yang, Zhikang Zou, Xiaoqing Ye, Zhe Liu 0033, Xiang Bai
Sci. China Inf. Sci.1
2023 A Simple Vision Transformer for Weakly Semi-supervised 3D Object Detection
abstract
Advanced 3D object detection methods usually rely on large-scale, elaborately labeled datasets to achieve good performance. However, labeling the bounding boxes for the 3D objects is difficult and expensive. Although semi-supervised (SS3D) and weakly-supervised 3D object detection (WS3D) methods can effectively reduce the annotation cost, they suffer from two limitations: 1) their performance is far inferior to the fully-supervised counterparts; 2) they are difficult to adapt to different detectors or scenes (e.g, indoor or outdoor). In this paper, we study weakly semi-supervised 3D object detection (WSS3D) with point annotations, where the dataset comprises a small number of fully labeled and massive weakly labeled data with a single point annotated for each 3D object. To fully exploit the point annotations, we employ the plain and non-hierarchical vision transformer to form a point-to-box converter, termed ViT-WSS3D. By modeling global interactions between LiDAR points and corresponding weak labels, our ViT-WSS3D can generate high-quality pseudo-bounding boxes, which are then used to train any 3D detectors without exhaustive tuning. Extensive experiments on indoor and outdoor datasets (SUN RGBD and KITTI) show the effectiveness of our method. In particular, when only using 10% fully labeled and the rest as point labeled data, our ViT-WSS3D can enable most detectors to achieve similar performance with the oracle model using 100% fully labeled data.
Dingyuan Zhang, Dingkang Liang, Zhikang Zou, Xiaoqing Ye, Zhe Liu 0033, Xiao Tan 0001, Xiang Bai
ICCV1