EDBT 2026 Demo / reviewers in the wild / expert
Zhihao Yuan
dblp:49/2982
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-3100-1252ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TAVEN: Task-driven Adaptive Viewpoint Exploration for Training-Free 3D Spatial Reasoning and UnderstandingabstractUnderstanding and reasoning over 3D scenes from natural language queries is essential for intuitive human-machine interaction in robotics, navigation, and mixed reality. While existing VLM-based methods address this task by training on 3D data, they face significant challenges in scalability and adaptability due to the limited availability of such data. Training-free approaches offer a promising alternative but often depend on fixed camera paths, coarse view sampling, or preprocessed inputs, limiting effective reasoning. We propose TAVEN, a training-free framework that leverages MLLM (i.e., GPT) for adaptive, task-driven exploration of 3D scenes. TAVEN features: 1) Chain-of-Thought Query Decomposition with Dual-Focus Reasoning to break down queries into sub-tasks focused on relevant entities and goals; 2) Global Memory-Based Dual-Level Retrieval for retrieving contextually relevant views; 3) Progressive View Adjustment with Tri-Criteria Evaluation to iteratively refine viewpoints; and 4) Trajectory Aware View Rectification to suggest improved (rectified) views based on exploration history. By integrating visual feedback with task semantics, TAVEN enables zero-shot, goal-oriented 3D reasoning, moving beyond passive perception toward intelligent spatial understanding. Shuyi Jiang, Zhihao Yuan, Na Zhao 0004 |
ICMR | 2 |
| 2026 | PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Modelsabstract3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and suboptimal quality of 3D datasets. Current approaches attempt to transfer knowledge from 2D MLLMs to expand 3D instruction data, but still face modality and domain gaps. To this end, we introduce PiSA-Engine (Point-Self-Augmented-Engine), a new framework for generating instruction point-language datasets enriched with 3D spatial semantics. We observe that existing 3D MLLMs offer a comprehensive understanding of point clouds for annotation, while 2D MLLMs excel at cross-validation by providing complementary information. By integrating holistic 2D and 3D insights from off-the-shelf MLLMs, PiSA-Engine enables a continuous cycle of high-quality data generation. We select PointLLM as the baseline and adopt this co-evolution training framework to develop an enhanced 3D MLLM, termed PointLLM-PiSA. Additionally, we identify limitations in previous 3D benchmarks, which often feature coarse language captions and insufficient category diversity, resulting in inaccurate evaluations. To address this gap, we further introduce PiSA-Bench, a comprehensive 3D benchmark covering six key aspects with detailed and diverse labels. Experimental results demonstrate PointLLM-PiSA’s state-of-the-art performance in zero-shot 3D object captioning and generative classification on our PiSA-Bench, achieving significant improvements of 46.45% (+8.33%) and 63.75% (+16.25%), respectively. Project page: CG-ops/PiSA. Zilu Guo, Zhihao Yuan, Chaoda Zheng, Pengshuo Qiu, Dongzhi Jiang, Renrui Zhang, Chun-Mei Feng 0001, Zhen Li 0026 |
WACV | 3 |
| 2026 | Enhancing medical image segmentation with the modification of U-shaped network
Shiren Li, Maksim Davydov, Serestina Viriri, Irsa Talib, Zhihao Yuan, Guangguang Yang |
Vis. Comput. | 6 |
| 2025 | Empowering Large Language Models with 3D Situation AwarenessabstractDriven by the great success of Large Language Models (LLMs) in the 2D image domain, their application in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions (e.g., "left" or "right"). However, current LLM-based methods overlook the egocentric perspective and use datasets from a global viewpoint. To address this issue, we propose a novel approach to automatically generate a situation-aware dataset by leveraging the scanning trajectory during data collection and utilizing Vision-Language Models (VLMs) to produce high-quality captions and question-answer pairs. Furthermore, we introduce a situation grounding module to explicitly predict the position and orientation of the observer’s viewpoint, thereby enabling LLMs to ground situation descriptions in 3D scenes. We evaluate our approach on several benchmarks, demonstrating that our method effectively enhances the 3D situational awareness of LLMs while significantly expanding existing datasets and reducing manual effort. Zhihao Yuan, Yibo Peng, Jinke Ren, Yinghong Liao, Yatong Han, Chun-Mei Feng 0001, Hengshuang Zhao, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
CVPR | 1 |
| 2025 | Toward Fine-Grained 3-D Visual Grounding Through Referring Textual PhrasesabstractRecent progress in 3-D scene understanding has explored visual grounding [3D visual grounding (3DVG)] to localize a target object through a language description. However, existing methods only consider the dependency between the entire sentence and the target object, ignoring fine-grained relationships between contexts and nontarget ones. In this article, we extend 3DVG to a more fine-grained task, called 3D phrase-aware grounding (3DPAG). The 3DPAG task aims to localize the target objects in a 3-D scene by explicitly identifying all phrase-related objects and then conducting the reasoning according to contextual phrases. To tackle this problem, we manually labeled about 227 K phrase-level annotations using a self-developed platform, from 88 K sentences of widely used 3DVG datasets, i.e., Natural Reference in 3-D (Nr3D), Spatial Reference in 3-D (Sr3D), and ScanRefer. By tapping on our datasets, we can extend previous 3DVG methods to the fine-grained phrase-aware scenario. It is achieved through the proposed novel phrase-object alignment (POA) optimization and phrase-specific pretraining (PSP), boosting conventional 3DVG performance as well. Extensive results confirm significant improvements, i.e., previous state-of-the-art method achieves 3.9%, 3.5%, and 4.6% overall accuracy gains on Nr3D, Sr3D, and ScanRefer, respectively. Our datasets and platform are released in https://github.com/CurryYuan/PhraseRefer. Zhihao Yuan, Xu Yan 0005, Xuhao Li, Yao Guo 0002, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | GSmoothFace: Generalized Smooth Talking Face Generation via Fine Grained 3D Face GuidanceabstractAlthough existing speech-driven talking face generation methods achieve significant progress, they are far from real-world application due to the avatar-specific training demand and unstable lip movements. To address the above issues, we propose the GSmoothFace, a novel two-stage generalized talking face generation model guided by a fine-grained 3D face model, which can synthesize smooth lip dynamics while preserving the speaker's identity. Our proposed GSmoothFace model mainly consists of the Audio to Expression Prediction (A2EP) module and the Target Adaptive Face Translation (TAFT) module. Specifically, we first develop the A2EP module to predict expression parameters synchronized with the driven speech. It uses a transformer to capture the long-term audio context and learns the parameters from the fine-grained 3D facial vertices, resulting in accurate and smooth lip-synchronization performance. Afterward, the well-designed TAFT module, empowered by Morphology Augmented Face Blending (MAFB), takes the predicted expression parameters and target video as inputs to modify the facial region of the target video without distorting the background content. The TAFT effectively exploits the identity appearance and background context in the target video, which makes it possible to generalize to different speakers without retraining. Both quantitative and qualitative experiments confirm the superiority of our method in terms of realism, lip-synchronization, and visual quality. Haiming Zhang 0001, Zhihao Yuan, Chaoda Zheng, Xu Yan 0005, Baoyuan Wang, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Visual Programming for Zero-Shot Open-Vocabulary 3D Visual Groundingabstract3D Visual Grounding (3DVG) aims at localizing 3D object based on textual descriptions. Conventional supervised methods for 3DVG often necessitate extensive annotations and a predefined vocabulary, which can be restrictive. To address this issue, we propose a novel visual programming approach for zero-shot open-vocabulary 3DVG, leveraging the capabilities of large language models (LLMs). Our approach begins with a unique dialog-based method, engaging with LLMs to establish a foundational understanding of zero-shot 3DVG. Building on this, we design a visual program that consists of three types of modules, i.e., view-independent, view-dependent, and functional modules. These modules, specifically tailored for 3D scenarios, work collaboratively to perform complex reasoning and inference. Furthermore, we develop an innovative language-object correlation module to extend the scope of existing 3D object detectors into open-vocabulary scenarios. Extensive experiments demonstrate that our zero-shot approach can outperform some supervised baselines, marking a significant stride towards effective 3DVG. Code is available at https://curryyuan.github.io/Z5VG3D. Zhihao Yuan, Jinke Ren, Chun-Mei Feng 0001, Hengshuang Zhao, Shuguang Cui, Zhen Li 0026 |
CVPR | 1 |
| 2024 | Comprehensive Visual Question Answering on Point Clouds through Compositional Scene ManipulationabstractVisual Question Answering on 3D Point Cloud (VQA-3D) is an emerging yet challenging field that aims at answering various types of textual questions given an entire point cloud scene. To tackle this problem, we propose the CLEVR3D, a large-scale VQA-3D dataset consisting of 171K questions from 8,771 3D scenes. Specifically, we develop a question engine leveraging 3D scene graph structures to generate diverse reasoning questions, covering the questions of objects' attributes (i.e., size, color, and material) and their spatial relationships. Through such a manner, we initially generated 44K questions from 1,333 real-world scenes. Moreover, a more challenging setup is proposed to remove the confounding bias and adjust the context from a common-sense layout. Such a setup requires the network to achieve comprehensive visual understanding when the 3D scene is different from the general co-occurrence context (e.g., chairs always exist with tables). To this end, we further introduce the compositional scene manipulation strategy and generate 127K questions from 7,438 augmented 3D scenes, which can improve VQA-3D models for real-world comprehension. Built upon the proposed dataset, we baseline several VQA-3D models, where experimental results verify that the CLEVR3D can significantly boost other 3D scene understanding tasks. Xu Yan 0005, Zhihao Yuan, Yinghong Liao, Yao Guo 0002, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | X -Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Captioningabstract3D dense captioning aims to describe individual objects in 3D scenes by natural language, where 3D scenes are usually represented as RGB-D scans or point clouds. However, only exploiting single modal information, e.g., point cloud, previous approaches fail to produce faithful descriptions. Though aggregating 2D features into point clouds may be beneficial, it introduces an extra computational burden, especially in the inference phase. In this study, we investigate a cross-modal knowledge transfer using Transformer for 3D dense captioning, namely X-Trans2Cap. Our proposed X-Trans2Cap effectively boost the performance of single-modal 3D captioning through the knowledge distillation enabled by a teacher-student framework. In practice, during the training phase, the teacher network exploits auxiliary 2D modality and guides the student network that only takes point clouds as input through the feature consistency constraints. Owing to the well-designed cross-modal feature fusion module and the feature alignment in the training phase, X-Trans2Cap acquires rich appearance information embedded in 2D images with ease. Thus, a more faithful caption can be generated only using point clouds during the inference. Qualitative and quantitative results confirm that X-Trans2Cap outperforms previous state-of-the-art by a large margin, i.e., about +21 and +16 CIDEr points on ScanRefer and Nr3D datasets, respectively. Zhihao Yuan, Xu Yan 0005, Yinghong Liao, Yao Guo 0002, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
CVPR | 1 |
| 2021 | InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringabstractCompared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer1, to achieve a superior 3D visual grounding through the grounding-by-matching strategy. In practice, our model first predicts the target category from the language descriptions using a simple language classification model. Then, based on the category, our model sifts out a small number of instance candidates (usually less than 20) from the panoptic segmentation on point clouds. Thus, the non-trivial 3D visual grounding task has been effectively re-formulated as a simplified instance-matching problem, considering that instance-level candidates are more rational than the redundant 3D object proposals. Subsequently, for each candidate, we perform the multi-level contextual inference, i.e., referring from instance attribute perception, instance-to-instance relation perception, and instance-to-background global localization perception, respectively. Eventually, the most relevant candidate is selected and localized by ranking confidence scores, which are obtained by the cooperative holistic visual-language feature matching. Experiments confirm that our method outperforms previous state-of-the-arts on ScanRefer online benchmark and Nr3D/Sr3D datasets. Zhihao Yuan, Xu Yan 0005, Yinghong Liao, Ruimao Zhang, Sheng Wang 0001, Zhen Li 0026, Shuguang Cui |
ICCV | 1 |
| 2021 | Revisiting Hard Example for Action RecognitionabstractVideo-based action recognition, which needs to handle temporal motion and spatial cues simultaneously, remains a challenging task. In this paper, our motivation is to address this issue by fully utilizing temporal information. Specially, a novel light-weight Voting-based Temporal Correlation (VTC) module is proposed to enhance temporal information. Multiple branches with different temporal sampling intervals are included in this module and they are regarded as voters. The final classification result is “voted” by these branches together. VTC module integrates sparse temporal sampling strategy into feature sequences, so it mitigates the effect of redundant information and focuses more on temporal modeling. Additionally, we propose a simple and intuitive Similarity Loss (SL) to guide the training procedure of the VTC module and the backbone network. When we introduce confusion in the predicted vector intentionally, SL eases intra-class variation by discovering class-specific common motion patterns rather than sample-specific discriminative information. SL neither needs excessive parameter tuning during training nor adds significant computation overhead during test time. By combining VTC module and SL with complementary advances in the field, we clearly outperform state-of-the-art results and achieve 83.0, 98.4, 49.6 and 77.8 accuracy on HMDB51, UCF101, something-something-v1, and Kinetics respectively. Jianguo Hu, Shiren Li, Zhihao Yuan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Rethinking Temporal-Related Sample for Human Action RecognitionabstractTemporal-related samples always have huge intra-class appearance variation, on which lots of existing action recognition algorithms have poor performance. In this paper, our motivation is to address this issue by utilizing temporal information more effectively. A novel light-weight Voting-based Temporal Correlation module (VTC) is proposed to enhance temporal cues. VTC integrates sparse temporal sampling strategy into feature sequences, so it mitigates the effect of redundant information and focuses more on temporal modeling. Furthermore, we propose a simple and intuitive Similarity Loss (SL) to guide the training procedure for VTC. Introducing confusion in the predicted vector intentionally, SL eases intra-class variation by discovering class-specific common motion pattern rather than sample-specific discriminative information. Combining VTC and SL with complementary advances in this field, we clearly outperform state-of-the-art results on HMDB51, UCF101, and Something-something-v1 dataset. The code has been made publicly available on https://github.com/FingerRec/TRS. Shiren Li, Zhikui Duan, Zhihao Yuan |
ICASSP | 4 |
| 2017 | Sciunits: Reusable Research ObjectsabstractScience is conducted collaboratively, often requiring knowledge sharing about computational experiments. When experiments include only datasets, they can be shared using Uniform Resource Identifiers (URIs) or Digital Object Identifiers (DOIs). An experiment, however, seldom includes only datasets, but more often includes software, its past execution, provenance, and associated documentation. The Research Object has recently emerged as a comprehensive and systematic method for aggregation and identification of diverse elements of computational experiments. While a necessary method, mere aggregation is not sufficient for the sharing of computational experiments. Other users must be able to easily recompute on these shared research objects. In this paper, we present the sciunit, a reusable research object in which aggregated content is recomputable. We describe a Git-like client that efficiently creates, stores, and repeats sciunits. We show through analysis that sciunits repeat computational experiments with minimal storage and processing overhead. Finally, we provide an overview of sharing and reproducible cyberinfrastructure based on sciunits gaining adoption in the domain of geosciences. Dai Hai Ton That, Gabriel Fils, Zhihao Yuan, Tanu Malik |
eScience | 3 |
| 2003 | Pseudo Context-Sensitive Models for Parsing Isolating Languages: Classical Chinese - A Case Study
Liang Huang 0001, Yinan Peng, Zhihao Yuan, Hui Liu 0002 |
CICLing | 4 |