VLDB 2026 Research / reviewers in the wild / expert
Zhenyang Liu
dblp:263/5268
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 46% Motion planning and robot control · 17% Reinforcement learning · 17% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d scene understanding
3d visual grounding |
1.7 | 2 | 2025 | A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding · ACM Multimedia 2025 ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning · CVPR 2025 |
Computer vision › 3D vision
3d scene understanding |
1.1 | 2 | 2025 | ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning · CVPR 2025 Spatial-Temporal Aware Visuomotor Diffusion Policy Learning · ICCV 2025 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.9 | 1 | 2025 | ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning · CVPR 2025 |
Computer vision › 3D vision › implicit neural representation
3d language field |
0.9 | 1 | 2025 | A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding · ACM Multimedia 2025 |
Robotics › Robot manipulation
diffusion policy |
0.9 | 1 | 2025 | Spatial-Temporal Aware Visuomotor Diffusion Policy Learning · ICCV 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.9 | 1 | 2025 | Spatial-Temporal Aware Visuomotor Diffusion Policy Learning · ICCV 2025 |
Robotics › Motion planning and robot control
robot learning |
0.9 | 1 | 2025 | Spatial-Temporal Aware Visuomotor Diffusion Policy Learning · ICCV 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning |
0.9 | 1 | 2025 | A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding · ACM Multimedia 2025 |
Machine learning › Reinforcement learning › imitation learning › learning from observation
visual imitation learning |
0.9 | 1 | 2025 | Spatial-Temporal Aware Visuomotor Diffusion Policy Learning · ICCV 2025 |
Robotics › Motion planning and robot control › robot learning › visuomotor learning
visuomotor policy learning |
0.9 | 1 | 2025 | Spatial-Temporal Aware Visuomotor Diffusion Policy Learning · ICCV 2025 |
Computer vision › Vision and language
vision-language model |
0.3 | 1 | 2025 | ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
large vision-language model · 0.9large language model · 0.9gaussian world model · 0.9gaussian splatting · 0.9diffusion model · 0.9behavior cloning · 0.9SAM · 0.9CLIP embedding · 0.9CLIP · 0.93d gaussian splatting · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and ReasoningabstractOpen-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigation and autonomous robotics. However, current methods struggle because they rely heavily on fine-tuning with 3D annotations and mask proposals, which limits their ability to handle diverse semantics and common knowledge required for effective reasoning. In this work, we propose ReasonGrounder, an LVLM-guided framework that uses hierarchical 3D feature Gaussian fields for adaptive grouping based on physical scale, enabling open-vocabulary 3D grounding and reasoning. ReasonGrounder interprets implicit instructions using large vision-language models (LVLM) and localizes occluded objects through 3D Gaussian splatting. By incorporating 2D segmentation masks from the SAM and multi-view CLIP embeddings, ReasonGrounder selects Gaussian groups based on object scale, enabling accurate localization through both explicit and implicit language understanding, even in novel, occluded views. We also contribute ReasoningGD, a new dataset containing over 10K scenes and 2 million annotations for evaluating open-vocabulary 3D grounding and amodal perception under occlusion. Experiments show that ReasonGrounder significantly improves 3D grounding accuracy in real-world scenarios. Zhenyang Liu, Yikai Wang 0002, Sixiao Zheng, Tongying Pan, Longfei Liang, Yanwei Fu 0001, Xiangyang Xue 0001 |
CVPR | 1 |
| 2025 | Spatial-Temporal Aware Visuomotor Diffusion Policy LearningabstractVisual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limiting their 3D spatial and 4D spatiotemporal awareness. Consequently, these methods struggle to capture the 3D structures and 4D spatiotemporal relationships necessary for real-world deployment. In this work, we propose 4D Diffusion Policy (DP4), a novel visual imitation learning method that incorporates spatiotemporal awareness into diffusion-based policies. Unlike traditional approaches that rely on trajectory cloning, DP4 leverages a dynamic Gaussian world model to guide the learning of 3D spatial and 4D spatiotemporal perceptions from interactive environments. Our method constructs the current 3D scene from a single-view RGB-D observation and predicts the future 3D scene, optimizing trajectory generation by explicitly modeling both spatial and temporal dependencies. Extensive experiments across 17 simulation tasks with 173 variants and 3 real-world robotic tasks demonstrate that the 4D Diffusion Policy (DP4) outperforms baseline methods, improving the average simulation task success rate by 16.4% (Adroit), 14% (DexArt), and 6.45% (RLBench), and the average real-world robotic task success rate by 8.6%. Zhenyang Liu, Yikai Wang 0002, Kuanning Wang, Longfei Liang, Xiangyang Xue 0001, Yanwei Fu 0001 |
ICCV | 1 |
| 2025 | A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual GroundingabstractOpen-vocabulary 3D visual grounding aims to localize target objects based on free-form language queries, which is crucial for embodied AI applications such as autonomous navigation, robotics, and augmented reality. Learning 3D language fields through neural representations enables accurate understanding of 3D scenes from limited viewpoints and facilitates the localization of target objects in complex environments. However, existing language field methods struggle to accurately localize instances using spatial relations in language queries, such as ''the book on the chair.'' This limitation mainly arises from inadequate reasoning about spatial relations in both language queries and 3D scenes. In this work, we propose SpatialReasoner, a novel neural representation-based framework with large language model (LLM)-driven spatial reasoning that constructs a visual properties-enhanced hierarchical feature field for open-vocabulary 3D visual grounding. To enable spatial reasoning in language queries, SpatialReasoner fine-tunes an LLM to capture spatial relations and explicitly infer instructions for the target, anchor, and spatial relation. To enable spatial reasoning in 3D scenes, SpatialReasoner incorporates visual properties (opacity and color) to construct a hierarchical feature field. This field represents language and instance features using distilled CLIP features and masks extracted via the Segment Anything Model (SAM). The field is then queried using the inferred instructions in a hierarchical manner to localize the target 3D instance based on the spatial relation in the language query. Notably, SpatialReasoner is not limited to a specific 3D neural representation; it serves as a framework adaptable to various representations, such as Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS). Extensive experiments show that our framework can be seamlessly integrated into different neural representations, outperforming baseline models in 3D visual grounding while empowering their spatial reasoning capability. Project Homepage:ZhenyangLiu.github.io/SpatialReasoner. Zhenyang Liu, Sixiao Zheng, Siyu Chen 0023, Cairong Zhao, Longfei Liang, Xiangyang Xue 0001, Yanwei Fu 0001 |
ACM Multimedia | 1 |
| 2025 | DE-NAF: decoupled neural attenuation fields for sparse-view CBCT reconstruction
Tianning Zhao, Guoping Ding, Zhenyang Liu, Hangping Wei, Min Tan 0005, Jiajun Ding |
Pattern Anal. Appl. | 3 |
| 2025 | Transformer-based material recognition via short-time contact sensing
Zhenyang Liu, Yitian Shao, Qiliang Li, Jingyong Su |
Pattern Recognit. | 1 |
| 2025 | VC-GS: view-consistent deblurring Gaussian splatting via alternating branch optimization
Qida Cao, Jiajun Ding, Zhenyang Liu, Zhenzhong Kuang, Yijie Shao, Yilan Shen |
Vis. Comput. | 3 |
| 2025 | Collaborative neural radiance fields for novel view synthesis
Junqing Yuan, Mengting Fan, Zhenyang Liu, Tongxuan Han, Zhenzhong Kuang, Chihao Pan, Jiajun Ding |
Vis. Comput. | 3 |
| 2024 | Unsupervised Fusion of Misaligned PAT and MRI Images via Mutually Reinforcing Cross-Modality Image Generation and RegistrationabstractPhotoacoustic tomography (PAT) and magnetic resonance imaging (MRI) are two advanced imaging techniques widely used in pre-clinical research. PAT has high optical contrast and deep imaging range but poor soft tissue contrast, whereas MRI provides excellent soft tissue information but poor temporal resolution. Despite recent advances in medical image fusion with pre-aligned multimodal data, PAT-MRI image fusion remains challenging due to misaligned images and spatial distortion. To address these issues, we propose an unsupervised multi-stage deep learning framework called PAMRFuse for misaligned PAT and MRI image fusion. PAMRFuse comprises a multimodal to unimodal registration network to accurately align the input PAT-MRI image pairs and a self-attentive fusion network that selects information-rich features for fusion. We employ an end-to-end mutually reinforcing mode in our registration network, which enables joint optimization of cross-modality image generation and registration. To the best of our knowledge, this is the first attempt at information fusion for misaligned PAT and MRI. Qualitative and quantitative experimental results show the excellent performance of our method in fusing PAT-MRI images of small animals captured from commercial imaging systems. Yutian Zhong, Shuangyang Zhang, Zhenyang Liu, Zongxin Mo, Wufan Chen |
IEEE Trans. Medical Imaging | 3 |
| 2023 | EGRA-NeRF: Edge-Guided Ray Allocation for Neural Radiance Fields
Zhenbiao Gai, Zhenyang Liu, Min Tan 0005, Jiajun Ding, Jun Yu 0002, Mingzhao Tong, Junqing Yuan |
Image Vis. Comput. | 2 |