EDBT 2026 Demo / reviewers in the wild / expert
Zhujin Liang
dblp:134/1828
· DBLP profile ↗
10ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-1278-5226ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
3D vision · 56% Autonomous driving · 23% Segmentation and scene understanding · 5% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 70% Geometric modeling and processing · 30% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d scene understanding |
2.3 | 3 | 2024 | Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion · IJCAI 2024 Hierarchical Temporal Context Learning for Camera-Based Semantic Scene Completion · ECCV (4) 2024 3DSFLabelling: Boosting 3D Scene Flow Estimation by Pseudo Auto-Labelling · CVPR 2024 |
Computer vision › 3D vision › 3d scene understanding
semantic scene completion |
1.5 | 2 | 2024 | Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion · IJCAI 2024 Hierarchical Temporal Context Learning for Camera-Based Semantic Scene Completion · ECCV (4) 2024 |
Robotics › Autonomous driving
end-to-end driving |
0.9 | 1 | 2025 | GraphAD: Interaction Scene Graph for End-to-end Autonomous Driving · IJCAI 2025 |
Robotics › Autonomous driving
interaction modeling |
0.9 | 1 | 2025 | GraphAD: Interaction Scene Graph for End-to-end Autonomous Driving · IJCAI 2025 |
Robotics › Autonomous driving › perception › vision-based perception
lane detection |
0.9 | 1 | 2025 | Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane Detection · CVPR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation · ACM Multimedia 2025 |
Robotics › Autonomous driving › 3d lane detection
monocular 3d lane detection |
0.9 | 1 | 2025 | Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane Detection · CVPR 2025 |
Robotics › Autonomous driving
perception |
0.9 | 1 | 2025 | GraphAD: Interaction Scene Graph for End-to-end Autonomous Driving · IJCAI 2025 |
Computer vision › Segmentation and scene understanding
scene graph |
0.9 | 1 | 2025 | GraphAD: Interaction Scene Graph for End-to-end Autonomous Driving · IJCAI 2025 |
Visual content generation and editing › vector graphics generation
Text-to-SVG generation |
0.9 | 1 | 2025 | SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation · ACM Multimedia 2025 |
Visual content generation and editing
vector graphics generation |
0.9 | 1 | 2025 | SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation · ACM Multimedia 2025 |
Computer vision › 3D vision
3d object detection |
0.8 | 1 | 2024 | Detecting as Labeling: Rethinking LiDAR-Camera Fusion in 3D Object Detection · ECCV (22) 2024 |
Computer vision › 3D vision › multimodal perception
LiDAR-camera fusion |
0.8 | 1 | 2024 | Detecting as Labeling: Rethinking LiDAR-Camera Fusion in 3D Object Detection · ECCV (22) 2024 |
Computer vision › 3D vision › scene flow estimation
LiDAR scene flow |
0.8 | 1 | 2024 | 3DSFLabelling: Boosting 3D Scene Flow Estimation by Pseudo Auto-Labelling · CVPR 2024 |
Computer vision › 3D vision › 3d scene understanding
multi-modal 3d perception |
0.8 | 1 | 2024 | Detecting as Labeling: Rethinking LiDAR-Camera Fusion in 3D Object Detection · ECCV (22) 2024 |
Computer vision › 3D vision › implicit neural representation
neural field |
0.8 | 1 | 2024 | NeuroGauss4D-PCI: 4D Neural Fields and Gaussian Deformation Fields for Point Cloud Interpolation · NeurIPS 2024 |
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles |
0.8 | 1 | 2024 | Detecting as Labeling: Rethinking LiDAR-Camera Fusion in 3D Object Detection · ECCV (22) 2024 |
Computer vision › 3D vision › point cloud analysis › point cloud learning
point cloud data augmentation |
0.8 | 1 | 2024 | 3DSFLabelling: Boosting 3D Scene Flow Estimation by Pseudo Auto-Labelling · CVPR 2024 |
Computer vision › 3D vision › point cloud processing › point cloud restoration
point cloud interpolation |
0.8 | 1 | 2024 | NeuroGauss4D-PCI: 4D Neural Fields and Gaussian Deformation Fields for Point Cloud Interpolation · NeurIPS 2024 |
Computer vision › 3D vision
point cloud processing |
0.8 | 1 | 2024 | NeuroGauss4D-PCI: 4D Neural Fields and Gaussian Deformation Fields for Point Cloud Interpolation · NeurIPS 2024 |
Machine learning › Learning paradigms › semi-supervised learning
pseudo-labeling |
0.8 | 1 | 2024 | 3DSFLabelling: Boosting 3D Scene Flow Estimation by Pseudo Auto-Labelling · CVPR 2024 |
Computer vision › 3D vision
scene flow estimation |
0.8 | 1 | 2024 | 3DSFLabelling: Boosting 3D Scene Flow Estimation by Pseudo Auto-Labelling · CVPR 2024 |
Computer vision › 3D vision
stereo vision |
0.8 | 1 | 2024 | Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion · IJCAI 2024 |
Computer vision › Video understanding and tracking › temporal modeling
temporal context learning |
0.8 | 1 | 2024 | Hierarchical Temporal Context Learning for Camera-Based Semantic Scene Completion · ECCV (4) 2024 |
Geometric modeling and processing › deformation modeling
deformation field |
0.8 | 1 | 2024 | NeuroGauss4D-PCI: 4D Neural Fields and Gaussian Deformation Fields for Point Cloud Interpolation · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.3 | 1 | 2025 | Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane Detection · CVPR 2025 |
Machine learning › Graph learning
graph neural network |
0.3 | 1 | 2025 | GraphAD: Interaction Scene Graph for End-to-end Autonomous Driving · IJCAI 2025 |
Computer vision › Vision and language › vision-language dataset › vision-language dataset construction
multimodal annotation |
0.3 | 1 | 2025 | SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation · ACM Multimedia 2025 |
Machine learning › Graph learning › graph representation
scene graph representation |
0.3 | 1 | 2025 | GraphAD: Interaction Scene Graph for End-to-end Autonomous Driving · IJCAI 2025 |
Computer vision › 3D vision › point cloud processing › point cloud restoration
point cloud densification |
0.2 | 1 | 2024 | NeuroGauss4D-PCI: 4D Neural Fields and Gaussian Deformation Fields for Point Cloud Interpolation · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.7multimodal annotation · 1.7chain-of-thought reasoning · 1.7attention mechanism · 1.7graph representation · 0.9soft clustering · 0.8rigid body motion decomposition · 0.8radial basis functions · 0.8labeling-based detection · 0.8hierarchical temporal context learning · 0.8gaussian splatting · 0.8anchor box motion augmentation · 0.8LiDAR-camera fusion · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane DetectionabstractMonocular 3D lane detection is a fundamental task in autonomous driving. Although sparse-point methods lower computational load and maintain high accuracy in complex lane geometries, current methods fail to fully leverage the geometric structure of lanes in both lane geometry representations and model design. In lane geometry representations, we present a theoretical analysis alongside experimental validation to verify that current sparse lane representation methods contain inherent flaws, resulting in potential errors of up to 20 m, which raise significant safety concerns for driving. To address this issue, we propose a novel patching strategy to completely represent the full lane structure. To enable existing models to match this strategy, we introduce the EndPoint head (EP-head), which adds a patching distance to endpoints. The EP-head enables the model to predict more complete lane representations even with fewer preset points, effectively addressing existing limitations and paving the way for models that are faster and require fewer parameters in the future. In model design, to enhance the model’s perception of lane structures, we propose the PointLane attention (PL-attention), which incorporates prior geometric knowledge into the attention mechanism. Extensive experiments demonstrate the effectiveness of the proposed methods on various state-of-the-art models. For instance, in terms of the overall F1-score, our methods improve Persformer by 4.4 points, Anchor3DLane by 3.2 points, and LATR by 2.8 points. The code will be available soon. Yifan Chang, Zhujin Liang, Dalong Du |
CVPR | 5 |
| 2025 | GraphAD: Interaction Scene Graph for End-to-end Autonomous DrivingabstractModeling complicated interactions among the ego-vehicle, road agents, and map elements has been a crucial part for safety-critical autonomous driving. Previous work on end-to-end autonomous driving relies on the attention mechanism to handle heterogeneous interactions, which fails to capture geometric priors and is also computationally intensive. In this paper, we propose the Interaction Scene Graph (ISG) as a unified method to model the interactions among the ego-vehicle, road agents, and map elements. With the representation of the ISG, the driving agents aggregate essential information from the most influential elements, including the road agents with potential collisions and the map elements to follow. Since a mass of unnecessary interactions are omitted, the more efficient scene-graph-based framework is able to focus on indispensable connections and leads to better performance. We evaluate the proposed method for end-to-end autonomous driving on the nuScenes dataset. Compared with strong baselines, our method significantly outperforms in full-stack driving tasks. Deheng Qian, Yifeng Pan, Zhenbao Liang, Yingzong Liu, Jianhui Mei, Maolei Fu, Zhujin Liang, Dalong Du |
IJCAI | 12 |
| 2025 | SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG GenerationabstractScalable Vector Graphics (SVG) is a code structure used to represent visual information, and with the powerful capabilities of large language models, it holds significant research potential. Current text-to-SVG generation methods lack generalization capabilities and struggle with accurately adhering to input generation instructions. In this paper, we propose a novel approach for generating SVG using large language models, named SVGThinker, which incorporates a reasoning process to align the generation of SVG code with the visualization process, while supporting all SVG primitives. Through sequential rendering of SVG primitives, we first use a multimodal model to annotate the SVG, followed by sequential updates corresponding to the incremental additions of primitives. We then employ a supervised training framework based on Chain-of-Thought reasoning, which enhances the model's robustness and reduces the risk of errors or hallucinations. Through comparisons with state-of-the-art baseline models, our experiments show that our model generates more stable, high-quality, and editable SVG code. In contrast to image-based methods, our approach preserves the structural advantages of SVG and supports precise, hierarchical editing. We believe our work opens new directions for SVG generation, with potential applications in design, content creation, and automated SVG-based graphic generation. Zhongyin Zhao, Ye Chen 0006, Zhujin Liang, Bingbing Ni |
ACM Multimedia | 4 |
| 2024 | 3DSFLabelling: Boosting 3D Scene Flow Estimation by Pseudo Auto-LabellingabstractLearning 3D scene flow from LiDAR point clouds presents significant difficulties, including poor generalization from synthetic datasets to real scenes, scarcity of real-world 3D labels, and poor performance on real sparse Li-DAR point clouds. We present a novel approach from the perspective of auto-labelling, aiming to generate a large number of 3D scene flow pseudo labels for real-world Li-DAR point clouds. Specifically, we employ the assumption of rigid body motion to simulate potential object-level rigid movements in autonomous driving scenarios. By updating different motion attributes for multiple anchor boxes, the rigid motion decomposition is obtained for the whole scene. Furthermore, we developed a novel 3D scene flow data augmentation method for global and local motion. By perfectly synthesizing target point clouds based on augmented motion parameters, we easily obtain lots of 3D scene flow labels in point clouds highly consistent with real scenarios. On multiple real-world datasets including LiDAR KITTI, nuScenes, and Argoverse, our method outperforms all previous supervised and unsupervised methods without requiring manual labelling. Impressively, our method achieves a tenfold reduction in EPE3D metric on the LiDAR KITTI dataset, reducing it from 0.190m to a mere 0.008m error. Chaokang Jiang, Guangming Wang 0001, Jiuming Liu, Hesheng Wang 0001, Zhenqiang Liu, Zhujin Liang, Dalong Du |
CVPR | 7 |
| 2024 | Detecting as Labeling: Rethinking LiDAR-Camera Fusion in 3D Object Detection
Junjie Huang 0005, Zhujin Liang, Dalong Du |
ECCV (22) | 3 |
| 2024 | Hierarchical Temporal Context Learning for Camera-Based Semantic Scene Completion
Bohan Li 0015, Jiajun Deng, Zhujin Liang, Dalong Du, Xin Jin 0014, Wenjun Zeng 0001 |
ECCV (4) | 4 |
| 2024 | Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion
Bohan Li 0015, Yasheng Sun, Zhujin Liang, Dalong Du, Zhuanghui Zhang, Yunnan Wang, Xin Jin 0014, Wenjun Zeng 0001 |
IJCAI | 3 |
| 2024 | NeuroGauss4D-PCI: 4D Neural Fields and Gaussian Deformation Fields for Point Cloud InterpolationabstractPoint Cloud Interpolation confronts challenges from point sparsity, complex spatiotemporal dynamics, and the difficulty of deriving complete 3D point clouds from sparse temporal information. This paper presents NeuroGauss4D-PCI, which excels at modeling complex non-rigid deformations across varied dynamic scenes. The method begins with an iterative Gaussian cloud soft clustering module, offering structured temporal point cloud representations. The proposed temporal radial basis function Gaussian residual utilizes Gaussian parameter interpolation over time, enabling smooth parameter transitions and capturing temporal residuals of Gaussian distributions. Additionally, a 4D Gaussian deformation field tracks the evolution of these parameters, creating continuous spatiotemporal deformation fields. A 4D neural field transforms low-dimensional spatiotemporal coordinates ($x,y,z,t$) into a high-dimensional latent space. Finally, we adaptively and efficiently fuse the latent features from neural fields and the geometric features from Gaussian deformation fields.
NeuroGauss4D-PCI outperforms existing methods in point cloud frame interpolation, delivering leading performance on both object-level (DHB) and large-scale autonomous driving datasets (NL-Drive), with scalability to auto-labeling and point cloud densification tasks. Chaokang Jiang, Dalong Du, Jiuming Liu, Siting Zhu 0001, Zhenqiang Liu, Zhujin Liang, Jie Zhou 0001 |
NeurIPS | 7 |
| 2014 | An expressive deep model for human action parsing from a single imageabstractThis paper aims at one newly raising task in vision and multimedia research: recognizing human actions from still images. Its main challenges lie in the large variations in human poses and appearances, as well as the lack of temporal motion information. Addressing these problems, we propose to develop an expressive deep model to naturally integrate human layout and surrounding contexts for higher level action understanding from still images. In particular, a Deep Belief Net is trained to fuse information from different noisy sources such as body part detection and object detection. To bridge the semantic gap, we used manually labeled data to greatly improve the effectiveness and efficiency of the pre-training and fine-tuning stages of the DBN training. The resulting framework is shown to be robust to sometimes unreliable inputs (e.g., imprecise detections of human parts and objects), and outperforms the state-of-the-art approaches. Zhujin Liang, Xiaolong Wang 0004, Rui Huang 0001, Liang Lin 0004 |
ICME | 1 |
| 2014 | Deep Joint Task Learning for Generic Object Extraction
Xiaolong Wang 0004, Liliang Zhang, Liang Lin 0004, Zhujin Liang, Wangmeng Zuo |
NIPS | 4 |