Ziying Song

dblp:292/5695 · DBLP profile ↗
← Back
30ranked-venue papers
10as first author
30since 2021 · last 2026
0000-0001-5539-2599ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Global relationship awareness 3-dimensional object detection using 4-dimensional radar
Pianzhang Duan, Li Wang 0092, Ziying Song, Ying Li 0036, Wei Fan 0011, Bin Xu 0003
Eng. Appl. Artif. Intell.4
2026 GADet: Geometry-Aware oriented object detection for remote sensing
Xinyu Zhang 0001, Ziying Song, Lei Yang 0060, Haicheng Qu
Knowl. Based Syst.4
2026 CausalPose: Causal visuo-tactile fusion for robust 6-DoF object pose estimation
Peiliang Wu, Yuanzhi Li, Mingyue Niu, Fengda Zhao, Ziying Song, Yongtao Yang
Pattern Recognit.6
2026 DGFusion: Dual-Guided Fusion for Robust Multi-Modal 3D Object Detection
abstract
As a critical task in autonomous driving perception systems, 3D object detection is used to identify and track key objects, such as vehicles and pedestrians. However, detecting distant, small, or occluded objects (hard instances) remains a challenge, which directly compromises the safety of autonomous driving systems. We observe that existing multi-modal 3D object detection methods often follow a single-guided paradigm, failing to account for the differences in information density of hard instances between modalities. In this work, we propose DGFusion, based on the Dual-guided paradigm, which fully inherits the advantages of the Point-guide-Image paradigm and integrates the Image-guide-Point paradigm to address the limitations of the single paradigms. The core of DGFusion, the Difficulty-aware Instance Pair Matcher (DIPM), performs instance-level feature matching based on difficulty to generate easy and hard instance pairs, while the Dual-guided Modules exploit the advantages of both pair types to enable effective multi-modal feature fusion. Experimental results demonstrate that our DGFusion outperforms the baseline methods, with respective improvements of +1.0% mAP, +0.8% NDS, and +1.3% average recall on nuScenes. Extensive experiments demonstrate consistent robustness gains for hard instance detection across ego-distance, size, visibility, and small-scale training scenarios.
Feiyang Jia, Caiyan Jia, Ailin Liu, Shaoqing Xu, Qiming Xia, Lei Yang 0060, Ziying Song
IEEE Trans. Circuits Syst. Video Technol.9
2026 TiGDistill-BEV: Multi-View BEV 3D Object Detection via Target Inner-Geometry Learning Distillation
abstract
Accurate multi-view 3D object detection is essential for applications such as autonomous driving. Researchers have consistently aimed to leverage LiDAR’s precise spatial information to enhance camera-based detectors through methods like depth supervision and bird-eye-view (BEV) feature distillation. However, existing approaches often face challenges due to the inherent differences between LiDAR and camera data representations. In this paper, we introduce the TiGDistill-BEV, a novel approach that effectively bridges this gap by leveraging the strengths of both sensors. Our method distills knowledge from diverse modalities(e.g., LiDAR) as the teacher model to a camera-based student detector, utilizing the Target Inner-Geometry learning scheme to enhance camera-based BEV detectors through both depth and BEV features by leveraging diverse modalities. Specially, we propose two key modules: an inner-depth supervision module to learn the low-level relative depth relations within objects which equips detectors with a deeper understanding of object-level spatial structures, and an inner-feature BEV distillation module to transfer high-level semantics of different keypoints within foreground targets. To further alleviate the domain gap, we incorporate both inter-channel and inter-keypoint distillation to model feature similarity. Extensive experiments on the nuScenes benchmark demonstrate that TiGDistill-BEV significantly boosts camera-based only detectors achieving a state-of-the-art with 62.8% NDS and surpassing previous methods by a significant margin. The codes is available at: https://github.com/Public-BOTs/TiGDistill-BEV.git.
Shaoqing Xu, Peixiang Huang, Ziying Song, Zhi-Xin Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Implicit Illumination-Aware Representation With Cross-Modal Prefusion Alignment for Universal Multispectral Pedestrian Detection
abstract
Traditional pedestrian detection methods based on red-green-blue (RGB) images struggle in adverse illumination, but a key capability required for pedestrian detection is all-day detection due to its critical role in diverse applications, e.g., security, surveillance, and autonomous driving. To address this issue, multispectral pedestrian detection attempts to introduce thermal images to supplement the RGB images, since they can be captured based on heat radiation difference without relying on external light sources. However, how to fuse the two modalities effectively is still lacking in-depth investigation. To prompt this field, we propose an implicit illumination-aware representation to address the limited availability of specific illumination labels in existing multispectral datasets, coupled with a prefusion feature alignment strategy to reconcile spatial misalignments of identical objects across modalities. We also identify four critical fusion challenges, revealing persistent limitations in existing multispectral detectors' ability to holistically address these issues, particularly regarding underdeveloped cross-modal interactions and suboptimal cross-domain feature fusion. To this end, we propose a universal multispectral pedestrian detection paradigm (UMPDP), which includes a modality alignment module (MAM) for adaptive feature space alignment, a differential modality fusion module (DMFM) to enhance the relationship of different modalities, and a task-conditioned illumination module (TCIM) to dynamically adjust network weights based on illumination condition. Extensive experiments on KAIST and CVC-14 datasets demonstrate the general effectiveness of our proposed method. Code is available at https://github.com/gongyan1/UMPDP.
Hao Liu 0114, Yongsheng Gao 0002, Jie Zhao 0003, Ziying Song, Xiaoxi Hu
IEEE Trans. Neural Networks Learn. Syst.7
2025 Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving
abstract
End-to-end autonomous driving frameworks enable seamless integration of perception and planning but often rely on one-shot trajectory prediction, which may lead to unstable control and vulnerability to occlusions in single-frame perception. To address this, we propose the Momentum-Aware Driving (MomAD) framework, which introduces trajectory momentum and perception momentum to stabilize and refine trajectory predictions. MomAD comprises two core components: (1) Topological Trajectory Matching (TTM) employs Hausdorff Distance to select the optimal planning query that aligns with prior paths to ensure coherence; (2) Momentum Planning Interactor (MPI) cross-attends the selected planning query with historical queries to expand static and dynamic perception files. This enriched query, in turn, helps regenerate long-horizon trajectory and reduce collision risks. To mitigate noise arising from dynamic environments and detection errors, we introduce robust instance denoising during training, enabling the planning model to focus on critical signals and improve its robustness. We also propose a novel Trajectory Prediction Consistency (TPC) metric to quantitatively assess planning stability. Experiments on the nuScenes dataset demonstrate that MomAD achieves superior long-term consistency (≥ 3s) compared to SOTA methods. Moreover, evaluations on the curated Turning-nuScenes shows that MomAD reduces the collision rate by 26% and improves TPC by 0.97m (33.45%) over a 6s prediction horizon, while closed- loop on Bench2Drive demonstrates an up to 16.3% improvement in success rate. The source code is available at https://github.com/adept-thu/MomAD.
Ziying Song, Caiyan Jia, Hongyu Pan, Shaoqing Xu, Lei Yang 0060, Yadan Luo
CVPR1
2025 FGU3R: Fine-Grained Fusion via Unified 3D Representation for Multimodal 3D Object Detection
abstract
Multimodal 3D object detection has garnered considerable interest in autonomous driving. However, multimodal detectors suffer from dimension mismatches that derive from fusing 3D points with 2D pixels coarsely, which leads to suboptimal fusion performance. In this paper, we propose a multimodal framework FGU3R to tackle the issue mentioned above via unified 3D representation and fine-grained fusion, which consists of two important components. First, we propose an efficient feature extractor for raw and pseudo points, termed Pseudo-Raw Convolution (PRConv), which modulates multimodal features synchronously and aggregates the features from different types of points on key points based on multimodal interaction. Second, a Cross-Attention Adaptive Fusion (CAAF) is designed to fuse homogeneous 3D RoI (Region of Interest) features adaptively via a cross-attention variant in a fine-grained manner. Together they make fine-grained fusion on unified 3D representation. The experiments conducted on the KITTI and nuScenes show the effectiveness of our proposed method.
Ziying Song, Zhonghong Ou
ICASSP2
2025 ComDrive: Comfort-Oriented End-to-End Autonomous Driving
abstract
We propose ComDrive: the first comfort-oriented end-to-end autonomous driving system to generate temporally consistent and comfortable trajectories. Recent studies have demonstrated that imitation learning-based planners and learning-based trajectory scorers can effectively generate and select safety trajectories that closely mimic expert demonstrations. However, such trajectory planners and scorers face the challenge of generating temporally inconsistent and uncomfortable trajectories. To address these issues, ComDrive first extracts 3D spatial representations through sparse perception, which then serves as conditional inputs. These inputs are used by a Conditional Denoising Diffusion Probabilistic Model (DDPM)-based motion planner to generate temporally consistent multi-modal trajectories. A dual-stream adaptive trajectory scorer subsequently selects the most comfortable trajectory from these candidates to control the vehicle. Experiments demonstrate that ComDrive achieves state-of-the-art performance in both comfort and safety, outperforming UniAD by 17%in driving comfort and reducing collision rates by 25%compared to SparseDrive. More results are available on our project page: https://jmwang0117.github.io/ComDrive/.
Zebin Xing, Songen Gu, Ziying Song, Qian Zhang 0001, Xiaoxiao Long, Wei Yin 0006
IROS7
2025 Formalization and Online Monitoring of Right-of-way Laws for Autonomous Vehicles at Intersections
abstract
With the rapid advancement of autonomous driving, safety concerns have become the primary barrier to its commercialization. Compliance with traffic laws is crucial for ensuring road safety. However, the current laws, formulated for human drivers, present challenges for autonomous systems due to ambiguous language description, complicating accurate judgment and government monitoring. It is imperative to transform traffic laws into machine-interpretable logical frameworks while simul-taneously resolving ambiguities in legal terminology to ensure clarity and precision. This study focuses on urban intersections, characterized by high traffic complexity and diverse participants. We propose a formalization method for right-of-way laws and develop a threshold analysis framework based on processed data from SIND, which rigorously defines the prioritization of right-of-way. The optimal compliance threshold is determined through sensitivity analysis, evaluated using the proposed Weighted TPN score (WTPNs). Meanwhile, the threshold was implemented in online monitoring at intersections. The dataset is available online via: https://github.comlSOTIF-AVLab/SinD
Lingjun Zhang, Chengxiang Zhao, Lei Yang 0060, Ziying Song, Wenhao Yu 0006, Hong Wang 0014
IV5
2025 V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception
abstract
Modern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In recent years, a series of cooperative perception datasets have emerged; however, these datasets primarily focus on cameras and LiDAR, neglecting 4D Radar—a sensor used in single-vehicle autonomous driving to provide robust perception in adverse weather conditions. In this paper, to bridge the gap created by the absence of 4D Radar datasets in cooperative perception, we present V2X-Radar, the first large-scale, real-world multi-modal dataset featuring 4D Radar. V2X-Radar dataset is collected using a connected vehicle platform and an intelligent roadside unit equipped with 4D Radar, LiDAR, and multi-view cameras. The collected data encompasses sunny and rainy weather conditions, spanning daytime, dusk, and nighttime, as well as various typical challenging scenarios. The dataset consists of 20K LiDAR frames, 40K camera images, and 20K 4D Radar data, including 350K annotated boxes across five categories. To support various research domains, we have established V2X-Radar-C for cooperative perception, V2X-Radar-I for roadside perception, and V2X-Radar-V for single-vehicle perception. Furthermore, we provide comprehensive benchmarks across these three sub-datasets.
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Jiaqi Ma 0003, Zhiying Song, Ziying Song, Li Wang 0092, Yang Shen 0005, Chen Lv 0001
NeurIPS8
2025 DMDet: Dynamic Multi-modal Object Detection Network for UAV Aerial Imagery
Zhe Yang 0005, Ziying Song
PRCV (16)3
2025 LVP: Leverage Virtual Points in Multimodal Early Fusion for 3-D Object Detection
abstract
Due to the sparsity and occlusion of point clouds, pure point cloud detection has limited effectiveness in detecting such samples. Researchers have been actively exploring the fusion of multimodal data, attempting to address the bottleneck issue based on LiDAR. In particular, virtual points, generated through depth completion from front-view RGB image, offer the potential for better integration with point clouds. Nevertheless, recent approaches fuse these two modalities in the region of interest (RoI), which limits the fusion effectiveness due to the inaccurate RoI region issue in the point cloud’s branch, especially in hard samples. To overcome it and unleash the potential of virtual points, while combining late fusion, we present leverage virtual point (LVP), a high-performance 3-D object detector which LVPs in early fusion to enhance the quality of RoI generation. LVP consists of three early fusion modules: virtual points painting (VPP), virtual points auxiliary (VPA), and virtual points completion (VPC) to achieve point-level fusion and global-level fusion. The integration of these modules effectively improves occlusion handling and improves the detection of distant small objects. In the KITTI benchmark, LVP achieves 85.45% 3-D mAP. As for large dataset nuScenes, we could improve the detection accuracy of large objects by compensating for errors in depth estimation. Without whistles and bells, these results establish LVP as an impressive solution for a 3-D outdoor object detection algorithm.
Yidong Chen 0006, Guo-Rong Cai, Ziying Song, Zhaoliang Liu, Binghui Zeng, Jonathan Li 0001, Zongyue Wang
IEEE Trans. Geosci. Remote. Sens.3
2025 DWTFreqNet: Infrared Small Target Detection via Wavelet-Driven Frequency Matching and Saliency-Difference Optimization
abstract
In the field of infrared small target detection, targets generally exhibit dim characteristics, and difficult to distinguish from background clutter. Learning-based methods enhance feature representation through layer-by-layer propagation, but the sparse target information often diminishes. To address this, we propose DWTFreqNet, a network that splits input data to enhance both local saliency and global contextual differences. It incorporates complementary feature extraction modules designed to match the data distribution characteristics. Specifically, it first utilizes the discrete wavelet transform (DWT) to decompose the input data into low- and high-frequency components. For the low-frequency part, which carries key target information, we apply component-differential dense connections and DWT-based downsampling to maintain feature integrity. For the high-frequency part, rich in target-background contrast, an Adaptive Wavelet Guidance Mechanism optimizes multi-component fusion via adaptive weighting, while a Layer-wide Discrepancy Relationship Capture Module enhances target discrimination by linking multi-scale feature maps. Comparative experiments on public datasets demonstrate its superiority over state-of-the-art methods. The code will be available at https://github.com/Kingwin97/DWTFreqNet.
Qianwen Ma, Shangwei Deng, Bincheng Li, Ziying Song, Xiaobo Li 0004, Haofeng Hu
IEEE Trans. Geosci. Remote. Sens.5
2024 GraphBEV: Towards Robust BEV Feature Alignment for Multi-modal 3D Object Detection
Ziying Song, Lei Yang 0060, Shaoqing Xu, Caiyan Jia, Feiyang Jia, Li Wang 0092
ECCV (26)1
2024 RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
Ziying Song, Guoxing Zhang, Lei Yang 0060, Shaoqing Xu, Caiyan Jia, Feiyang Jia, Li Wang 0092
IJCAI1
2024 Simplified GZN (Gradient-Zhang Neurodynamic) Continuous-Model and Discrete-Algorithms Handling Temporally-Varying ODLMVE (Over-Determined Linear Matrix-Vector Equation)
Yunong Zhang, Ziying Song, Binbin Qiu
ISNN2
2024 SparseInteraction: Sparse Semantic Guidance for Radar and Camera 3D Object Detection
abstract
Multi-modal fusion techniques, such as radar and images, enable a complementary and cost-effective perception of the surrounding environment regardless of lighting and weather conditions. However, existing fusion methods for surround-view images and radar are challenged by the inherent noise and positional ambiguity of radar, which leads to significant performance losses. To address this limitation effectively, our paper presents a robust, end-to-end fusion framework dubbed SparseInteraction. First, we introduce the Noisy Radar Filter (NRF) module to extract foreground features by creatively using queried semantic features from the image to filter out noisy radar features. Furthermore, we implement the Sparse Cross-Attention Encoder (SCAE) to effectively blend foreground radar features and image features to address positional ambiguity issues at a sparse level. Ultimately, to facilitate model convergence and performance, the foreground prior queries containing position information of the foreground radar are concatenated with predefined queries and fed into the subsequent transformer-based decoder. The experimental results demonstrate that the proposed fusion strategies markedly enhance detection performance and achieve new state-of-the-art results on the nuScenes benchmark. Source code is available at https://github.com/GG-Bonds/SparseInteraction.
Shaoqing Xu, Shengyin Jiang, Li Liu 0069, Ziying Song, Zhi-Xin Yang 0001
ACM Multimedia5
2024 GraphAlign++: An Accurate Feature Alignment by Graph Matching for Multi-Modal 3D Object Detection
abstract
LiDAR and camera are complementary sensors for 3D object detection in autonomous driving. However, it is challenging to explore the unnatural interaction between point clouds and images, and the critical factor is how to conduct feature alignment of these heterogeneous modalities. Currently, many methods achieve feature alignment through projection calibration, without accounting for the impact of sensors misalignment errors, resulting in sub-optimal performance. In this paper, we present GraphAlign++, a more accurate feature alignment framework for 3D object detection by graph matching. Specifically, we construct the nearest neighbor relationship by calculating Euclidean distances of point cloud features within the subspaces. Through the projection calibration between the image and point cloud pairs, we project the nearest neighbors of point cloud features onto the corresponding image. Then by matching the nearest neighbors of a single point-feature of the point cloud with multiple pixel-features of the image, we search for a more appropriate feature alignment. In addition, we provide a self-attention module to enhance the weights of significant relations to fine-tune the feature alignment between these two heterogeneous modalities. Extensive experiments on nuScenes benchmark demonstrate the effectiveness and efficiency of GraphAlign++. Notably, due to the more accurate feature alignment, which contributes to increase mAP by 3.10% on KITTI test hard level, our method is remarkably beneficial for long-range object detection.
Ziying Song, Caiyan Jia, Lei Yang 0060, Haiyue Wei
IEEE Trans. Circuits Syst. Video Technol.1
2024 SparseDet: A Simple and Effective Framework for Fully Sparse LiDAR-Based 3-D Object Detection
abstract
LiDAR-based sparse 3-D object detection plays a crucial role in autonomous driving applications due to its computational efficiency advantages. Existing methods either use the features of a single central voxel as an object proxy or treat an aggregated cluster of foreground points as an object proxy. However, the former cannot aggregate contextual information, resulting in insufficient information expression in object proxies. The latter relies on multistage pipelines and auxiliary tasks, which reduce the inference speed. To maintain the efficiency of the sparse framework while fully aggregating contextual information, in this work, we propose SparseDet that designs sparse queries as object proxies. It introduces two key modules: the local multiscale feature aggregation (LMFA) module and the global feature aggregation (GFA) module, aiming to fully capture the contextual information, thereby enhancing the ability of the proxies to represent objects. The LMFA module achieves feature fusion across different scales for sparse key voxels via coordinate transformations and using nearest neighbor relationships to capture object-level details and local contextual information, whereas the GFA module uses self-attention mechanisms to selectively aggregate the features of the key voxels across the entire scene for capturing scene-level contextual information. Experiments on nuScenes and KITTI demonstrate the effectiveness of our method. Specifically, SparseDet surpasses the previous best sparse detector VoxelNeXt (a typical method using voxels as object proxies) by 2.2% mean average precision (mAP) with 13.5 frames/s on nuScenes and outperforms VoxelNeXt by 1.12%$\text {AP}_{\text {3-D}}$on hard level tasks with 17.9 frames/s on KITTI. What is more, not only the mAP of SparseDet exceeds that of FSDV2 (a classical method using clusters of foreground points as object proxies) but also its inference speed is 1.3 times faster than FSDV2 on the nuScenes test set. The code has been released inhttps://github.com/liulin813/SparseDet.git.
Ziying Song, Qiming Xia, Feiyang Jia, Caiyan Jia, Lei Yang 0060, Hongyu Pan
IEEE Trans. Geosci. Remote. Sens.2
2024 Multi-Sem Fusion: Multimodal Semantic Fusion for 3-D Object Detection
abstract
LIDAR and camera fusion techniques are promising for achieving 3D object detection in autonomous driving. Most multi-modal 3D object detection frameworks integrate semantic knowledge from 2D images into 3D LiDAR point clouds to enhance detection accuracy. Nevertheless, the restricted resolution of 2D feature maps impedes accurate re-projection and often induces a pronounced boundary-blurring effect, which is primarily attributed to erroneous semantic segmentation. To address these limitations, we present theMulti-Sem Fusion (MSF)framework, a versatile multi-modal fusion approach that employs 2D/3D semantic segmentation methods to generate parsing results for both modalities. Subsequently, the 2D semantic information undergoes re-projection into 3D point clouds utilizing calibration parameters. To tackle misalignment challenges between the 2D and 3D parsing results, we introduce an Adaptive Attention-based Fusion (AAF) module to fuse them by learning an adaptive fusion score. Then the point cloud with the fused semantic label is sent to the following 3D object detectors. Furthermore, we propose a Deep Feature Fusion (DFF) module to aggregate deep features at different levels to boost the final detection performance. The effectiveness of the framework has been verified on two public large-scale 3D object detection benchmarks by comparing them with different baselines. And the experimental results show that the proposed fusion strategies can significantly improve the detection performance compared to the methods using only point clouds and the methods using only 2D semantic information. Moreover, our approach seamlessly integrates as a plug-in within any detection framework.
Shaoqing Xu, Ziying Song, Sifen Wang, Zhi-Xin Yang 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Robustness-Aware 3D Object Detection in Autonomous Driving: A Review and Outlook
abstract
In the realm of modern autonomous driving, the perception system is indispensable for accurately assessing the state of the surrounding environment, thereby enabling informed prediction and planning. The key step to this system is related to 3D object detection that utilizes vehicle-mounted sensors such as LiDAR and cameras to identify the size, the category, and the location of nearby objects. Despite the surge in 3D object detection methods aimed at enhancing detection precision and efficiency, there is a gap in the literature that systematically examines their resilience against environmental variations, noise, and weather changes. This study emphasizes the importance of robustness, alongside accuracy and latency, in evaluating perception systems under practical scenarios. Our work presents an extensive survey of camera-only, LiDAR-only, and multi-modal 3D object detection algorithms, thoroughly evaluating their trade-off between accuracy, latency, and robustness, particularly on datasets like KITTI-C and nuScenes-C to ensure fair comparisons. Among these, multi-modal 3D detection approaches exhibit superior robustness, and a novel taxonomy is introduced to reorganize the literature for enhanced clarity. This survey aims to offer a more practical perspective on the current capabilities and the constraints of 3D object detection algorithms in real-world applications, thus steering future research towards robustness-centric advancements.
Ziying Song, Feiyang Jia, Yadan Luo, Caiyan Jia, Lei Yang 0060, Li Wang 0092
IEEE Trans. Intell. Transp. Syst.1
2023 GraphAlign: Enhancing Accurate Feature Alignment by Graph matching for Multi-Modal 3D Object Detection
abstract
LiDAR and cameras are complementary sensors for 3D object detection in autonomous driving. However, it is challenging to explore the unnatural interaction between point clouds and images, and the critical factor is how to conduct feature alignment of heterogeneous modalities. Currently, many methods achieve feature alignment by projection calibration only, without considering the problem of coordinate conversion accuracy errors between sensors, leading to sub-optimal performance. In this paper, we present GraphAlign, a more accurate feature alignment strategy for 3D object detection by graph matching. Specifically, we fuse image features from a semantic segmentation encoder in the image branch and point cloud features from a 3D Sparse CNN in the LiDAR branch. To save computation, we construct the nearest neighbor relationship by calculating Euclidean distance within the subspaces that are divided into the point cloud features. Through the projection calibration between the image and point cloud, we project the nearest neighbors of point cloud features onto the image features. Then by matching the nearest neighbors with a single point cloud to multiple images, we search for a more appropriate feature alignment. In addition, we provide a self-attention module to enhance the weights of significant relations to fine-tune the feature alignment between heterogeneous modalities. Extensive experiments on nuScenes benchmark demonstrate the effectiveness and efficiency of our GraphAlign.
Ziying Song, Haiyue Wei, Lei Yang 0060, Caiyan Jia
ICCV1
2023 URFormer: Unified Representation LiDAR-Camera 3D Object Detection with Transformer
Jun Xie 0003, Zhepeng Wang 0002, Kuihe Yang, Ziying Song
PRCV (3)6
2023 SAT-GCN: Self-attention graph convolutional network-based 3D object detection for autonomous driving
Li Wang 0092, Ziying Song, Xinyu Zhang 0001, Jun Li 0082, Huaping Liu 0001
Knowl. Based Syst.2
2023 VP-Net: Voxels as Points for 3-D Object Detection
abstract
3D object detection with LiDAR point clouds is a challenging problem which requires 3D scene understanding, yet this task is critical to autonomous driving. Existing voxel-based 3D object detectors are becoming increasingly popular but have several shortcomings. For example, during voxelization, features of distant sparse point clouds are largely discarded, which leads to the missing detection of objects. Additionally, the correlation of points between voxels and the importance of different voxels within a region are not well learned. Therefore, we present a robust network (VP-Net) that views voxels as points to accurately detect 3D objects in LiDAR point clouds and can capture objects’ internal relationships. 3D CNN processing shows the output features of VP-Net as key points. The relationship between key points is then constructed into local graphs to enhance object feature extraction via a self-attention mechanism. Finally, the Euclidean distance between the extracted features guides our model’s weight reassignment for strengthening the importance of neighbor points, thereby enhancing the internal feature aggregation of objects. Experiments on KITTI and nuScenes 3D object detection benchmarks demonstrate the efficiency of enhancing inter-voxel validity within object features and show that the proposed VP-Net can achieve state-of-the-art performance.
Ziying Song, Haiyue Wei, Caiyan Jia, Yongchao Xia, Xiaokun Li
IEEE Trans. Geosci. Remote. Sens.1
2023 VoxelNextFusion: A Simple, Unified, and Effective Voxel Fusion Framework for Multimodal 3-D Object Detection
abstract
LiDAR-camera fusion can enhance the performance of 3D object detection by utilizing complementary information between depth-aware LiDAR points and semantically rich images. Existing voxel-based methods face significant challenges when fusing sparse voxel features with dense image features in a one-to-one manner, resulting in the loss of the advantages of images, including semantic and continuity information, leading to sub-optimal detection performance, especially at long distances. In this paper, we present VoxelNextFusion, a multi-modal 3D object detection framework specifically designed for voxel-based methods, which effectively bridges the gap between sparse point clouds and dense images. In particular, we propose a voxel-based image pipeline that involves projecting point clouds onto images to obtain both pixel- and patch-level features. These features are then fused using a self-attention to obtain a combined representation. Moreover, to address the issue of background features present in patches, we propose a feature importance module that effectively distinguishes between foreground and background features, thus minimizing the impact of the background features. Extensive experiments were conducted on the widely used KITTI and nuScenes 3D object detection benchmarks. Notably, our VoxelNextFusion achieved around +3.20% in [email protected] improvement for car detection in hard level compared to the Voxel R-CNN baseline on the KITTI test dataset.
Ziying Song, Jun Xie 0003, Caiyan Jia, Shaoqing Xu, Zhepeng Wang 0002
IEEE Trans. Geosci. Remote. Sens.1
2022 Fast Detection of Multi-Direction Remote Sensing Ship Object Based on Scale Space Pyramid
abstract
Ships in remote sensing images are usually arranged in arbitrary direction, small in size, and densely arranged. As a result, existing object detection algorithms cannot detect ships quickly and accurately. In order to solve the above problems, a lightweight object detection network for fast detection of ships is proposed. The network is composed of backbone network, four-scale fusion network and rotation branch. First, a lightweight network unit S-LeanNet is designed and used to build a low-computing and accurate backbone network. Then, a four-scale feature fusion module is designed to generate a four-scale feature pyramid, which contains more features such as ship shape and texture, and at the same time is conducive to the detection of small ships. Finally, a novel rotation branch module is designed, using balance L1 loss function and R-NMS for post-processing, to realize the precise positioning and regression of the rotating bounding box in one step. Experimental results show that the detection precision of our method in the DOT A remote sensing data set is compared with the latest SCRDet detection method, the precision is increased by 1.1%, and the operating speed is increased by 8 times, which can meet the fast detection requirements of ships.
Ziying Song, Li Wang 0092, Caiyan Jia, Jiangfeng Bi, Haiyue Wei, Yongchao Xia, Lijun Zhao 0003
MSN1
2021 Model Optimization Method Based on Vertical Federated Learning
abstract
When Vertical Federated Learning is used to classify tasks, a large number of invalid parameters are produced. In view of the above problems, we propose a general method of parameter sharing and gradient compression for both sides of communication, and improve the homomorphic encryption transfer parameters. The experimental results show that the evaluation index of the classification model is greatly improved compared with the traditional longitudinal federated learning logic regression algorithm.
Kuihe Yang, Ziying Song, Jianxuan Wang
ISCAS2
2021 MsfNet: a novel small object detection based on multi-scale feature fusion
abstract
This paper proposes a small object detection algorithm based on multi-scale feature fusion. By learning shallow features at the shallow level and deep features at the deep level, the proposed multi-scale feature learning scheme focuses on the fusion of concrete features and abstract features. It constructs object detector (MsfNet) based on multi-scale deep feature learning network and considers the relationship between a single object and local environment. Combining global information with local information, the feature pyramid is constructed by fusing different depth feature layers in the network. In addition, this paper also proposes a new feature extraction network (CourNet), through the way of feature visualization compared with the mainstream backbone network, the network can better express the small object feature information. The proposed algorithm is evaluated on MS COCO dataset and achieves the leading performance. This study shows that the combination of global information and local information is helpful to detect the expression of small objects in different illumination. MsfNet uses CourNet as the backbone network, which has high efficiency and a good balance between accuracy and speed.
Ziying Song, Peiliang Wu, Kuihe Yang
MSN1