VLDB 2026 Research / reviewers in the wild / expert
Guoliang You
dblp:344/0637
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-0964-7279ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera FusionabstractWe propose Radar-Camera fusion transformer (RaC-Former) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation-if the depth of pixels is not accurately estimated, the naive combination of BEV features actually integrates unaligned visual content. To avoid this problem, we propose a query-based framework that enables adaptive sampling of instance-relevant features from both the bird’s-eye view (BEV) and the original image view. Furthermore, we enhance system performance by two key designs: optimizing query initialization and strengthening the representational capacity of BEV. For the former, we introduce an adaptive circular distribution in polar coordinates to refine the initialization of object queries, allowing for a distance-based adjustment of query density. For the latter, we initially incorporate a radar-guided depth head to refine the transformation from image view to BEV. Subsequently, we focus on leveraging the Doppler effect of radar and introduce an implicit dynamic catcher to capture the temporal elements within the BEV. Extensive experiments on nuScenes and View-of-Delft (VoD) datasets validate the merits of our design. Remarkably, our method achieves superior results of 64.9% mAP and 70.2% NDS on nuScenes. RaCFormer also secures the state-of-the-art performance on the VoD dataset. Code is available at https://github.com/cxmomo/RaCFormer. Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Houqiang Li, Yanyong Zhang |
CVPR | 3 |
| 2025 | GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping Under Flexible Language InstructionsabstractFlexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities of the large language models (LLMs) to establish mappings between expressions and targets, allowing robots to comprehend users' intentions in the instructions. However, the LLM's knowledge about objects' physical properties remains underexplored despite its tight relevance to grasping. In this work, we propose GraspCoT, a 6-DoF grasp detection framework that integrates a Chain-of-Thought (CoT) reasoning mechanism oriented to physical properties, guided by auxiliary question-answering (QA) tasks. Particularly, we design a set of QA templates to enable hierarchical reasoning that includes three stages: target parsing, physical property analysis, and grasp action selection. Moreover, GraspCoT presents a unified multimodal LLM architecture, which encodes multi-view observations of 3D scenes into 3D-aware visual tokens, and then jointly embeds these visual tokens with CoT-derived textual tokens within LLMs to generate grasp pose predictions. Furthermore, we present IntentGrasp, a large-scale benchmark that fills the gap in public datasets for multi-object grasp detection under diverse and indirect verbal commands. Extensive experiments on IntentGrasp demonstrate the superiority of our method, with additional validation in real-world robotic applications confirming its practicality. The code is available at https://github.com/cxmomo/GraspCoT. Xiaomeng Chu, Jiajun Deng, Guoliang You, Jianmin Ji, Yanyong Zhang |
ICCV | 3 |
| 2025 | CELLmap: Enhancing LiDAR SLAM Through Elastic and Lightweight Spherical Map RepresentationabstractSLAM is a fundamental capability of unmanned systems, with LiDAR-based SLAM gaining widespread adoption due to its high precision. Current SLAM systems can achieve centimeter-level accuracy within a short period. However, there are still several challenges when dealing with largescale mapping tasks including significant storage requirements and difficulty of reusing the constructed maps. To address this, we first design an elastic and lightweight map representation called CELLmap, composed of several CELLS, each representing the local map at the corresponding location. Then, we design a general backend including CELL-based bidirectional registration module and loop closure detection module to improve global map consistency. Our experiments have demonstrated that CELLmap can represent the precise geometric structure of large-scale maps of KITTI dataset using only about 60 MB. Additionally, our general backend achieves up to a 26.88% improvement over various LiDAR odometry methods. Yifan Duan, Yao Li 0016, Guoliang You, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang |
ICRA | 4 |
| 2025 | CalibWorkflow: A General MLLM-Guided Workflow for Centimeter-Level Cross-Sensor CalibrationabstractExtrinsic calibration is a fundamental step in sensor fusion systems. However, existing methods often lack generalization capabilities when facing diverse hardware configurations, sensor poses, and environmental conditions, hindering their large-scale deployment. To address this limitation, we propose a general extrinsic calibration method, CalibWorkflow. Our core innovation lies in positioning multimodal large language models (MLLMs) as ''visual guides'' for the calibration process, leveraging their powerful vision-language understanding capabilities to guide parameter search and refinement. This reliance on visual scene understanding, rather than specific geometric features or sensor characteristics, enables the method to generalize effectively across diverse hardware and environmental conditions. Specifically, CalibWorkflow employs a three-stage calibration pipeline: initial parameter search, coarse optimization, and fine optimization. First, it utilizes the MLLM to assess the visual consistency between the projected point cloud and the image, rapidly determining an initial range for the extrinsic parameters. Next, the MLLM serves as a differential evaluator, giving simple ''better'' or ''worse'' feedback on parameter changes to guide the search through the parameter space. Finally, the method refines the calibration by matching edge features and performing non-linear optimization. Extensive experiments are conducted across six diverse scenarios and four heterogeneous sensor combinations. CalibWorkflow achieves state-of-the-art sub-degree and centimeter-level accuracy on four datasets and demonstrates highly competitive performance on others. These results thoroughly validate the generalization and robustness when facing various scenarios. Codes will be available. Wuyang Zhang, Guoliang You, Xiaomeng Chu, Wenhao Yu 0010, Yifan Duan, Yanyong Zhang |
ACM Multimedia | 3 |
| 2025 | VLMPlanner: Integrating Visual Language Models with Motion PlanningabstractIntegrating large language models (LLMs) into autonomous driving motion planning has recently emerged as a promising direction, offering enhanced interpretability, better controllability, and improved generalization in rare and long-tail scenarios. However, existing methods often rely on abstracted perception or map-based inputs, missing crucial visual context, such as fine-grained road cues, accident aftermath, or unexpected obstacles, which are essential for robust decision-making in complex driving environments. To bridge this gap, we propose VLMPlanner, a hybrid framework that combines a learning-based real-time planner with a vision-language model (VLM) capable of reasoning over raw images. The VLM processes multi-view images to capture rich, detailed visual information and leverages its common-sense reasoning capabilities to guide the real-time planner in generating robust and safe trajectories. Furthermore, we develop the Context-Adaptive Inference Gate (CAI-Gate) mechanism that enables the VLM to mimic human driving behavior by dynamically adjusting its inference frequency based on scene complexity, thereby achieving an optimal balance between planning performance and computational efficiency. We evaluate our approach on the large-scale, challenging nuPlan benchmark, with comprehensive experimental results demonstrating superior planning performance in scenarios with intricate road conditions and dynamic elements. Zhipeng Tang, Sha Zhang 0002, Jiajun Deng, Chenjie Wang, Guoliang You, Xinrui Lin, Yanyong Zhang |
ACM Multimedia | 5 |
| 2024 | RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric StrategiesabstractThe recent advances in query-based multi-camera 3D object detection are featured by initializing object queries in the 3D space, and then sampling features from perspective-view images to perform multi-round query refinement. In such a framework, query points near the same camera ray are likely to sample similar features from very close pixels, resulting in ambiguous query features and degraded detection accuracy. To this end, we introduce RayFormer, a camera-ray-inspired query-based 3D object detector that aligns the initialization and feature extraction of object queries with the optical characteristics of cameras. Specifically, RayFormer transforms perspective-view image features into bird's eye view (BEV) via the lift-splat-shoot method and segments the BEV map to sectors based on the camera rays. Object queries are uniformly and sparsely initialized along each camera ray, facilitating the projection of different queries onto different areas in the image to extract distinct features. Besides, we leverage the instance information of images to supplement the uniformly initialized object queries by further involving additional queries along the ray from 2D object detection boxes. To extract unique object-level features that cater to distinct queries, we design a ray sampling method that suitably organizes the distribution of feature sampling points on both images and bird's eye view. Extensive experiments are conducted on the nuScenes dataset to validate our proposed ray-inspired model design. The proposed RayFormer achieves 55.5% mAP and 63.3% NDS, respectively. Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Yao Li 0016, Yanyong Zhang |
ACM Multimedia | 3 |
| 2023 | P3O: Transferring Visual Representations for Reinforcement Learning via PromptingabstractIt is important for deep reinforcement learning (DRL) algorithms to transfer their learned policies to new environments that have different visual inputs. In this paper, we introduce Prompt based Proximal Policy Optimization (P3O), a three-stage DRL algorithm that transfers visual representations from a target to a source environment by applying prompting. The process of P3O consists of three stages: pre-training, prompting, and predicting. In particular, we specify a prompt-transformer for representation conversion and propose a two-step training process to train the prompt-transformer for the target environment, while the rest of the DRL pipeline remains unchanged. We implement P3O and evaluate it on the OpenAI CarRacing video game. The experimental results show that P3O outperforms the state-of-the-art visual transferring schemes. In particular, P3O allows the learned policies to perform well in environments with different visual inputs, which is much more effective than retraining the policies in these environments. Guoliang You, Xiaomeng Chu, Yifan Duan, Jie Peng 0002, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang |
ICME | 1 |
| 2023 | A coordinate attention enhanced swin transformer for handwriting recognition of Parkinson's diseaseabstractAbstract Diagnosing Parkinson's disease (PD) in its early stages is a significant challenge in medicine. Hand tremors and dysgraphia, which are typical early motor symptoms of PD, can manifest for decades before a formal diagnosis is made. Therefore, handwriting analysis has become an important tool for detecting PD. While many machine learning algorithms have been applied in this area, they struggle to capture the subtle changes in handwriting and must describe features from various perspectives. To address these issues, this paper proposes a Coordinate Attention Enhanced Swin Transformer (CAS Transformer) model for PD handwriting recognition. It establishes the long‐term dependence of features on the joint coordinate attention application, which enables the model to more accurately localize the important features of handwriting data and also extract the fuzzy edge features of handwriting images.These characteristics of the CAS Transformer enable it to outperform current advanced deep learning methods in classification, with an accuracy of 92.68% in experiments conducted on two handwritten datasets. Xuesen Niu, Yiyang Yuan, Yunze Sun, Guoliang You, Aite Zhao |
IET Image Process. | 6 |