Zhenbo Song

dblp:267/1972 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-5020-4277ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 11 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Efficient Active Training for Deep LiDAR Odometry
abstract
Robust and efficient deep LiDAR odometry models are crucial for accurate localization and 3D reconstruction, but typically require extensive and diverse training data to adapt to diverse environments, leading to inefficiencies. To tackle this, we introduce an active training framework designed to selectively extract training data from diverse environments, thereby reducing the training load and enhancing model generalization. Our framework is based on two key strategies: Initial Training Set Selection (ITSS) and Active Incremental Selection (AIS). ITSS begins by breaking down motion sequences from general weather into nodes and edges for detailed trajectory analysis, prioritizing diverse sequences to form a rich initial training dataset for training the base model. For complex sequences that are difficult to analyze, especially under challenging snowy weather conditions, AIS uses scene reconstruction and prediction inconsistency to iteratively select training samples, refining the model to handle a wide range of real-world scenarios. Experiments across datasets and weather conditions validate our approach’s effectiveness. Notably, our method matches the performance of full-dataset training with just 52% of the sequence volume, demonstrating the training efficiency and robustness of our active training paradigm. By optimizing the training process, our approach sets the stage for more agile and reliable LiDAR odometry systems, capable of navigating diverse environmental conditions with greater precision.
Beibei Zhou, Zhiyuan Zhang 0004, Zhenbo Song, Jianhui Guo, Hui Kong 0001
IEEE Trans. Intell. Transp. Syst.3
2025 Controllable Satellite-to-Street-View Synthesis with Precise Pose Alignment and Zero-Shot Environmental Control
abstract
Generating street-view images from satellite imagery is a challenging task, particularly in maintaining accurate pose alignment and incorporating diverse environmental conditions. While diffusion models have shown promise in generative tasks, their ability to maintain strict pose alignment throughout the diffusion process is limited. In this paper, we propose a novel Iterative Homography Adjustment (IHA) scheme applied during the denoising process, which effectively addresses pose misalignment and ensures spatial consistency in the generated street-view images. Additionally, currently, available datasets for satellite-to-street-view generation are limited in their diversity of illumination and weather conditions, thereby restricting the generalizability of the generated outputs. To mitigate this, we introduce a text-guided illumination and weather-controlled sampling strategy that enables fine-grained control over the environmental factors. Extensive quantitative and qualitative evaluations demonstrate that our approach significantly improves pose accuracy and enhances the diversity and realism of generated street-view images, setting a new benchmark for satellite-to-street-view generation tasks.
Xianghui Ze, Zhenbo Song, Jianfeng Lu 0003, Yujiao Shi 0002
ICLR2
2025 Gradient-Based Adversarial Attacks on Deep LiDAR Odometry
abstract
Adversarial attacks have been recently investigated in LiDAR perception problems for autonomous driving, where a small perturbation of source inputs can result in incorrect predictions. However, most previous studies focus on attacks on single-frame perception modules, lacking explorations of attacks on consecutive-frame tasks, i.e. the LiDAR odometry. In this paper, we propose a gradient optimization-based adversarial attack towards deep LiDAR odometry networks. To generate point clouds consistent with real-world scenarios, we constrain adversarial points within the range of a small object, e.g. a traffic cone, and render new points to simulate real LiDAR measurements. By incorporating such adversarial points in consecutive frames, we demonstrate a significant decrease in pose estimation accuracy of current popular LiDAR odometry networks. In addition, we also evaluate traditional geometric odometry approaches and report their robustness against adversarial points. Extensive experiments on the KITTI and Waymo datasets illustrate the effectiveness of the proposed attack method and the vulnerability of deep LiDAR odometry networks against adversarial points.
Zhenbo Song, Xuanzhu Chen, Zhenyuan Zhang 0001, Kaihao Zhang, Jianfeng Lu 0003
ICRA1
2025 Lightweight Yet High-Performance Defect Detector for Uav-Based Large-Scale Infrastructure Real-Time Inspection
abstract
Defect diagnosis in urban infrastructure is crucial for public safety. Traditional manual inspections face significant challenges in terms of accuracy and cost-effectiveness. In this paper, we propose a lightweight and hardware-friendly large-scale infrastructure detector, CUPID, highly suitable for unmanned aerial vehicles (UAVs). Given the significant challenges in automatically detecting defects of varying intensity and size within complex infrastructure, along with the tendency of lightweight models to lose detail and fail to fully capture features during the defect extraction process, we propose the CUPID_Block, a multi-level information fusion block to construct the backbone, featuring the CUPID_Conv module equipped with our proposed CCA (CrissCross Attention). Furthermore, CUPID features an auxiliary training branch that assimilates lower feature maps, helping to recover details lost in deeper convolutional layers. To verify the effectiveness of CUPID and to address the lack of a suitable dataset in the community, we establish a multi-scenario infrastructure defect dataset, CUBIT2024, to conduct extensive experiments. Finally, to assess the efficiency and adaptability of CUPID in UAV for online infrastructure inspection, we design a compact autonomous drone, CU-Astro, where the proposed CUPID is deployed on the Jetson Orin NX computer onboard to evaluate the speed and power consumption of the inference.
Benyun Zhao, Qigeng Duan, Guidong Yang, Jerry Tang, Zhenbo Song, Junjie Wen 0001, Xuchen Liu 0001, Qingxiang Li, Lei Lei 0010, Jihan Zhang, Xi Chen 0104, Mark W. Mueller, Ben M. Chen
ICRA5
2025 UV-GA: UV-Guided Gaussian Avatar Reconstruction from Single Image
Zhenbo Song, Zhenyuan Zhang 0001, Jianfeng Lu 0003
PRCV (10)2
2025 V2DGS:Visual Voxel Map-Based 2D Gaussian Splatting for Accurate Outdoor Reconstruction
Zhenbo Song, Qigeng Duan, Benyun Zhao, Jianfeng Lu 0003
PRCV (10)2
2025 Toward Physically Stable Motion Generation: A New Paradigm of Human Pose Representation
abstract
In machine learning, generating realistic human motion is paramount for a range of applications that require lifelike movements. Traditional methods have often overlooked the adherence to physical principles, leading to motion sequences that exhibit unrealistic behaviors such as foot sliding, penetration, and floating. These issues are particularly pronounced in complex tasks like dance choreography, which demand a higher degree of fidelity and realism. To address these challenges, we introduce RF-Rotation, a novel approach to human pose representation that strategically repositions the root joint of the SMPL model to align with both feet, while representing other joints through recursive bone rotations. It not only aligns more closely with the natural dynamics of human movement but also integrates an advanced contact predictor to ascertain the ground contact status of both feet, thereby preventing physically implausible movements on feet. We note that RF-Rotation is compatible with any motion generation tasks, including dance choreography, text-to-motion synthesis, and motion prediction, and can be seamlessly integrated into existing frameworks without modifications. Extensive experiments across three distinct tasks demonstrate the superior performance of RF-Rotation in enhancing the realism and stability of generated motion sequences. This method can significantly reduce foot sliding, floating, and penetration issues, without affecting computational efficiency, underscores its potential to set new standards in human motion generation.
Qiongjie Cui, Zhenyu Lou, Zhenbo Song, Xiangbo Shu
IEEE Trans. Circuits Syst. Video Technol.3
2025 Multi-Scale Semantic-Guidance Networks: Robust Blind Face Restoration Against Adversarial Attacks
abstract
Image processing networks are known to be vulnerable to adversarial examples, where adding carefully crafted adversarial perturbations to the inputs can mislead the model. This paper addresses the problem of robust blind face restoration (BFR) against adversarial attacks. BFR refers to recovering the HQ images from the LQ images, which suffer from diverse unknown degradation, such as noise, blur, artifact removal, low resolution, etc. Although existing BFR methods exhibit good performance, they experience significant degradation when subtle distortions and perturbations are introduced into the input images. This paper is the first to investigate, improve comprehensively, and evaluate BFR methods towards adversarial attacks. Project Gradient Descent (PGD) is employed to generate adversarial examples, and multiple types of attacks were used to thoroughly assess the robustness of various BFR methods across different objectives, regions, and levels. We evaluate the robustness of multiple BFR methods and analyze the advantages of their structures and modules towards adversarial attacks. Experimental results demonstrate that the method utilizing latent feature encoding and pre-trained discrete HQ codebook achieves better robustness than other methods, with the latter outperforming the former. Similarly, multi-scale semantic guidance information also exhibits superior performance in enhancing robustness. Therefore, we propose a powerful BFR method to mitigate this issue while maintaining better performance. Extensive experiments on three real-world datasets demonstrate our method’s state-of-the-art robustness in different scenarios.
Zhenyuan Zhang 0001, Xingqun Qi, Zhenbo Song, Zhiqin Yang, Jianfeng Lu 0003, Muyi Sun, Man Zhang 0005, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.3
2024 Universal Video Face Restoration Method Based on Vision-Language Model
Yipiao Xu, Zhenbo Song, Jianfeng Lu 0003
ACML2
2024 ACR-Pose: Adversarial Canonical Representation Reconstruction Network for Category Level 6D Object Pose Estimation
abstract
In the realm of category-level 6D object pose estimation, canonical 3D representation reconstruction is pivotal, yet current methods show limitations in reconstruction quality, a key step in current pose estimation pipeline. To address this, we introduce an innovative Adversarial Canonical Representation Reconstruction Network (ACR-Pose) in this paper. In particular, ACR-Pose comprises a Reconstructor, with novel sub-modules: a Pose-Irrelevant Module (PIM) for robustness to rotation and translation, and a Relational Reconstruction Module (RRM) for extracting relational information between input modalities. A Discriminator is incorporated to guide the generation of realistic canonical representations through adversarial optimization. Evaluated on the prevalent NOCS-CAMERA and NOCS-REAL datasets, our method significantly improves the performance of baseline models and achieves comparable performance with existing state-of-the-art methods, representing a promising advancement in the field of category-level 6D object pose estimation.
Zhaoxin Fan, Zhenbo Song, Zhicheng Wang 0007, Jian Xu 0027, Kejian Wu, Hongyan Liu 0002, Jun He 0008
ICMR2
2024 STDG: Semi-Teacher-Student Training Paradigm for Depth-guided One-stage Scene Graph Generation
abstract
Scene Graph Generation is a critical enabler of environmental comprehension for autonomous robotic systems.Most of existing methods, however, are often thwarted by the intricate dynamics of background complexity, which limits their ability to fully decode the inherent topological information of the environment.Additionally, the wealth of contextual information encapsulated within depth cues is often left untapped, rendering existing approaches less effective.To address these shortcomings, we present STDG, an avant-garde Depth-Guided One-Stage Scene Graph Generation methodology.The innovative architecture of STDG is a triad of custom-built modules: The Depth Guided HHA Representation Generation Module, the Depth Guided Semi-Teaching Network Learning Module, and the Depth Guided Scene Graph Generation Module.This trifecta of modules synergistically harnesses depth information, covering all aspects from depth signal generation and depth feature utilization, to the final scene graph prediction.Importantly, this is achieved without imposing additional computational burden during the inference phase.Experimental results confirm that our method significantly enhances the performance of onestage scene graph generation baselines.
Xukun Zhou, Zhenbo Song, Jun He 0008, Hongyan Liu 0002, Zhaoxin Fan
ICMR2
2024 On the Robustness of Deep Face Inpainting: An Adversarial Perspective
Zhenbo Song, Zhenyuan Zhang 0001, Jianfeng Lu 0003
MMAsia2
2024 Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion Prediction
abstract
Diverse human motion prediction (HMP) is a fundamental application in computer vision that has recently attracted considerable interest. Prior methods primarily focus on the stochastic nature of human motion, while neglecting the specific impact of external environment, leading to the pronounced artifacts in prediction when applied to real-world scenarios. To fill this gap, this work introduces a novel task: predicting diverse human motion within real-world 3D scenes. In contrast to prior works, it requires harmonizing the deterministic constraints imposed by the surrounding 3D scenes with the stochastic aspect of human motion. For this purpose, we propose DiMoP3D, a diverse motion prediction framework with 3D scene awareness, which leverages the 3D point cloud and observed sequence to generate diverse and high-fidelity predictions. DiMoP3D is able to comprehend the 3D scene, and determines the probable target objects and their desired interactive pose based on the historical motion. Then, it plans the obstacle-free trajectory towards these interested objects, and generates diverse and physically-consistent future motions. On top of that, DiMoP3D identifies deterministic factors in the scene and integrates them into the stochastic modeling, making the diverse HMP in realistic scenes become a controllable stochastic generation process. On two real-captured benchmarks, DiMoP3D has demonstrated significant improvements over state-of-the-art methods, showcasing its effectiveness in generating diverse and physically-consistent motion predictions within real-world 3D environments.
Zhenyu Lou, Qiongjie Cui, Zhenbo Song, Luoming Zhang, Huaxia Li
NeurIPS4
2024 AS-FIBA: Adaptive Selective Frequency-Injection for Backdoor Attack on Deep Face Restoration
abstract
Deep learning-based face restoration models, increasingly prevalent in smart devices, have become targets for sophisticated backdoor attacks. Through subtle trigger injection into input face images, these attacks can lead to unexpected restoration outcomes. Unlike conventional methods focused on classification tasks, our approach introduces a unique degradation objective tailored for attacking restoration models. Moreover, we propose the Adaptive Selective Frequency Injection Backdoor Attack (AS-FIBA) framework, employing a neural network for input-specific trigger generation in the frequency domain, seamlessly blending triggers with benign images. This results in imperceptible yet effective attacks, guiding restoration predictions towards subtly degraded outputs rather than conspicuous targets. Extensive experiments demonstrate the efficacy of the degradation objective on state-of-the-art face restoration models. Additionally, it is notable that AS-FIBA can insert effective backdoors that are more imperceptible than existing backdoor attack methods, including WaNet, ISSBA, and FIBA.
Zhenbo Song, Zhenyuan Zhang 0001, Jianfeng Lu 0003
TrustCom1
2023 Robust Single Image Reflection Removal Against Adversarial Attacks
abstract
This paper addresses the problem of robust deep single-image reflection removal (SIRR) against adversarial attacks. Current deep learning based SIRR methods have shown significant performance degradation due to unnoticeable distortions and perturbations on input images. For a comprehensive robustness study, we first conduct diverse adversarial attacks specifically for the SIRR problem, i.e. towards different attacking targets and regions. Then we propose a robust SIRR model, which integrates the cross-scale attention module, the multi-scale fusion module, and the adversarial image discriminator. By exploiting the multi-scale mechanism, the model narrows the gap between features from clean and adversarial images. The image discriminator adaptively distinguishes clean or noisy inputs, and thus further gains reliable robustness. Extensive experiments on Nature, SIR2, and Real datasets demonstrate that our model remarkably improves the robustness of SIRR across disparate scenes.
Zhenbo Song, Zhenyuan Zhang 0001, Kaihao Zhang, Wenhan Luo, Zhaoxin Fan, Wenqi Ren, Jianfeng Lu 0003
CVPR1
2023 EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation
abstract
Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content. To address this issue, this paper proposes an end-to-end neural network to disentangle different emotions in speech so as to generate rich 3D facial expressions. Specifically, we introduce the emotion disentangling encoder (EDE) to disentangle the emotion and content in the speech by cross-reconstructed speech signals with different emotion labels. Then an emotion-guided feature fusion decoder is employed to generate a 3D talking face with enhanced emotion. The decoder is driven by the disentangled identity, emotional, and content embeddings so as to generate controllable personal and emotional styles. Finally, considering the scarcity of the 3D emotional talking face data, we resort to the supervision of facial blendshapes, which enables the reconstruction of plausible 3D faces from 2D emotional data, and contribute a large-scale 3D emotional talking face dataset (3D-ETF) to train the network. Our experiments and user studies demonstrate that our approach outperforms state-of-the-art methods and exhibits more diverse facial movements. We recommend watching the supplementary video: https://ziqiaopeng.github.io/emotalk
Ziqiao Peng, Zhenbo Song, Xiangyu Zhu 0001, Jun He 0008, Hongyan Liu 0002, Zhaoxin Fan
ICCV3
2023 Incorporating Global Correlation and Local Aggregation for Efficient Visual Localization
Jianfeng Lu 0003, Zhenbo Song, Xuanzhu Chen
ICIG (2)3
2023 GIDP: Learning a Good Initialization and Inducing Descriptor Post-enhancing for Large-scale Place Recognition
abstract
Large-scale place recognition is a fundamental but challenging task, which plays an increasingly important role in autonomous driving and robotics. Existing methods have achieved acceptable good performance, however, most of them are concentrating on designing elaborate global descriptor learning network structures. The importance of feature generalization and descriptor post-enhancing has long been neglected. In this work, we propose a novel method named GIDP to learn a Good Initialization and Inducing Descriptor Pose-enhancing for Large-scale Place Recognition. In particular, an unsupervised momentum contrast point cloud pretraining module and a reranking-based descriptor post-enhancing module are proposed respectively in GIDP. The former aims at learning a good initialization for the point cloud encoding network before training the place recognition model, while the later aims at post-enhancing the predicted global descriptor through reranking at inference time. Ex-tensive experiments on both indoor and outdoor datasets demonstrate that our method can achieve state-of-the-art performance using simple and general point cloud encoding backbones.
Zhaoxin Fan, Zhenbo Song, Hongyan Liu 0002, Jun He 0008
ICRA2
2023 Learning Dense Flow Field for Highly-accurate Cross-view Camera Localization
abstract
This paper addresses the problem of estimating the 3-DoF camera pose for a ground-level image with respect to a satellite image that encompasses the local surroundings. We propose a novel end-to-end approach that leverages the learning of dense pixel-wise flow fields in pairs of ground and satellite images to calculate the camera pose. Our approach differs from existing methods by constructing the feature metric at the pixel level, enabling full-image supervision for learning distinctive geometric configurations and visual appearances across views. Specifically, our method employs two distinct convolution networks for ground and satellite feature extraction. Then, we project the ground feature map to the bird's eye view (BEV) using a fixed camera height assumption to achieve preliminary geometric alignment. To further establish the content association between the BEV and satellite features, we introduce a residual convolution block to refine the projected BEV feature. Optical flow estimation is performed on the refined BEV feature map and the satellite feature map using flow decoder networks based on RAFT. After obtaining dense flow correspondences, we apply the least square method to filter matching inliers and regress the ground camera pose. Extensive experiments demonstrate significant improvements compared to state-of-the-art methods. Notably, our approach reduces the median localization error by 89\%, 19\%, 80\%, and 35\% on the KITTI, Ford multi-AV, VIGOR, and Oxford RobotCar datasets, respectively.
Zhenbo Song, Xianghui Ze, Jianfeng Lu 0003, Yujiao Shi 0002
NeurIPS1
2023 Deep semantic-aware remote sensing image deblurring
Zhenbo Song, Zhenyuan Zhang 0001, Feiyi Fang, Zhaoxin Fan, Jianfeng Lu 0003
Signal Process.1
2022 SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition
abstract
Simultaneous Localization and Mapping (SLAM) and Autonomous Driving are becoming increasingly more important in recent years. Point cloud-based large scale place recognition is the spine of them. While many models have been proposed and have achieved acceptable performance by learning short-range local features, they always skip long-range contextual properties. Moreover, the model size also becomes a serious shackle for their wide applications. To overcome these challenges, we propose a super light-weight network model termed SVT-Net. On top of the highly efficient 3D Sparse Convolution (SP-Conv), an Atom-based Sparse Voxel Transformer (ASVT) and a Cluster-based Sparse Voxel Transformer (CSVT) are proposed respectively to learn both short-range local features and long-range contextual features. Consisting of ASVT and CSVT, SVT-Net can achieve state-of-the-art performance in terms of both recognition accuracy and running speed with a super-light model size (0.9M parameters). Meanwhile, for the purpose of further boosting efficiency, we introduce two simplified versions, which also achieve state-of-the-art performance and further reduce the model size to 0.8M and 0.4M respectively.
Zhaoxin Fan, Zhenbo Song, Hongyan Liu 0002, Zhiwu Lu 0001, Jun He 0008, Xiaoyong Du 0001
AAAI2
2022 Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation from Monocular RGB Image
Zhaoxin Fan, Zhenbo Song, Jian Xu 0027, Zhicheng Wang 0007, Kejian Wu, Hongyan Liu 0002, Jun He 0008
ECCV (2)2
2022 PilotAttnNet: Multi-modal Attention Network for End-to-End Steering Control
Jincan Zhang, Zhenbo Song, Jianfeng Lu 0003, Xingwei Qu, Zhaoxin Fan
PRCV (3)2
2022 Self-Supervised Depth Completion From Direct Visual-LiDAR Odometry in Autonomous Driving
abstract
In this work, a simple yet effective deep neural network is proposed to generate the dense depth map of the scene by exploiting both LiDAR sparse point cloud and the monocular camera image. Specifically, a feature pyramid network is firstly employed to extract feature maps from images across time. Then the relative pose is calculated by minimizing the feature distance between aligned pixels from inter-frame feature maps. Finally, the feature maps and the relative pose are further applied to compute the feature-metric loss for training the depth completion network. The key novelty of this work lies in that a self-supervised mechanism is presented to train the depth completion network by directly using visual-LiDAR odometry between consecutive frames. Comprehensive experiments and ablation studies on benchmark dataset KITTI demonstrate the superior performance over other state-of-the-art methods in terms of pose estimation and depth completion. The detailed performance of the proposed approach (referred to asSelfCompDVLO) can be found on the KITTI depth completion benchmark. The source code, models, and data have been made available at GitHub.
Zhenbo Song, Jianfeng Lu 0003, Yazhou Yao, Jian Zhang 0002
IEEE Trans. Intell. Transp. Syst.1
2020 Deep Novel View Synthesis from Colored 3D Point Clouds
Zhenbo Song, Wayne Chen, Dylan Campbell, Hongdong Li
ECCV (24)1
2020 End-to-end Learning for Inter-Vehicle Distance and Relative Velocity Estimation in ADAS with a Monocular Camera
abstract
Inter-vehicle distance and relative velocity estimations are two basic functions for any ADAS (Advanced driver-assistance systems). In this paper, we propose a monocular camera based inter-vehicle distance and relative velocity estimation method based on end-to-end training of a deep neural network. The key novelty of our method is the integration of multiple visual clues provided by any two time-consecutive monocular frames, which include deep feature clue, scene geometry clue, as well as temporal optical flow clue. We also propose a vehicle-centric sampling mechanism to alleviate the effect of perspective distortion in the motion field (i.e. optical flow). We implement the method by a light-weight deep neural network. Extensive experiments are conducted which confirm the superior performance of our method over other state-of-the-art methods, in terms of estimation accuracy, computational speed, and memory footprint.
Zhenbo Song, Jianfeng Lu 0003, Tong Zhang 0023, Hongdong Li
ICRA1