VLDB 2026 Research / reviewers in the wild / expert
Zhiyu Xiang
dblp:63/1751
· DBLP profile ↗
50ranked-venue papers
1as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 1 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 since 2021Systems, architecture and hardware · 8 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SDNet: LiDAR Semantic Scene Completion with Sparse-Dense Fusion and Input-Aware Label RefinementabstractLiDAR Semantic Scene Completion (SSC) in autonomous driving requires predicting both dense occupancy and semantic labels from sparse input point cloud. Existing methods typically adopt cascaded architecture for feature dilation and semantic abstraction, which blurs distinctive geometric patterns and reduces feature discriminability. Moreover, given an input, conventional processing of the ground truth labels overlooks voxel predictability in the target, resulting in ill-posed supervision and discards informative voxels. To address these limitations, we propose Sparse-Dense Net (SDNet), a dual-branch architecture that processes the input points through parallel sparse and dense encoders. The complementary features are aligned and fused using a Sparse Dense Feature Fusion (SDFF) module and further refined by a Feature Propagation (FP) module. Additionally, we introduce an input-aware label refinement strategy, including Sparse-Guided Filtering (SGF) to filter unpredictable targets and Ignored Voxel Recycling (IVR) to leverage informative ignored voxels for auxiliary supervision. These innovations enhance both feature learning and label quality. Extensive experiments on SemanticKITTI and nuScenes OpenOccupancy datasets validate the effectiveness of our approach, with SDNet achieving state-of-the-art performance on both datasets and ranking 1st on the official SemanticKITTI benchmark with 42.1 mIoU, outperforming the previous best by 4.2 (+11.1\%). Tingming Bai, Zhiyu Xiang, Peng Xu 0026, Tianyu Pu, Eryun Liu |
AAAI | 2 |
| 2026 | RaLiFlow: Scene Flow Estimation with 4D Radar and LiDAR Point CloudsabstractRecent multimodal fusion methods, integrating images with LiDAR point clouds, have shown promise in scene flow estimation. However, the fusion of 4D millimeter wave radar and LiDAR remains unexplored. Unlike LiDAR, radar is cheaper, more robust in various weather conditions and can detect point-wise velocity, making it a valuable complement to LiDAR. However, radar inputs pose challenges due to noise, low resolution, and sparsity. Moreover, there is currently no dataset that combines LiDAR and radar data specifically for scene flow estimation. To address this gap, we construct a Radar-LiDAR scene flow dataset based on a public real-world automotive dataset. We propose an effective preprocessing strategy for radar denoising and scene flow label generation, deriving more reliable flow ground truth for radar points out of the object boundaries. Additionally, we introduce RaLiFlow, the first joint scene flow learning framework for 4D radar and LiDAR, which achieves effective radar-LiDAR fusion through a novel Dynamic-aware Bidirectional Cross-modal Fusion (DBCF) module and a carefully designed set of loss functions. The DBCF module integrates dynamic cues from radar into the local cross-attention mechanism, enabling the propagation of contextual information across modalities. Meanwhile, the proposed loss functions mitigate the adverse effects of unreliable radar data during training and enhance the instance-level consistency in scene flow predictions from both modalities, particularly for dynamic foreground areas. Extensive experiments on the repurposed scene flow dataset demonstrate that our method outperforms existing LiDAR-based and radar-based single-modal methods by a significant margin. Jingyun Fu, Zhiyu Xiang, Na Zhao 0004 |
AAAI | 2 |
| 2026 | GanCom: Attention-enhanced feature compression for communication-efficient collaborative perception via adversarial training
Shaohong Wang, Hangguan Shan, Zhiyu Xiang, Zhewei Fu, Eryun Liu |
Neurocomputing | 4 |
| 2025 | Privacy-Preserving V2X Collaborative Perception Integrating Unknown CollaboratorsabstractVehicle-to-everything (V2X) collaborative perception has recently gained increasing attention in autonomous driving due to its ability to enhance scene understanding by integrating information from other collaborators, e.g. vehicles or infrastructure. Existing algorithms usually share deep features to achieve a trade-off between accuracy and bandwidth. However, most of these methods require joint training of all agents, which results in privacy leakage and is impractical and unacceptable in the real world. Sharing prediction results seems to be a direct solution, but its performance is suboptimal and sensitive to localization noise and communication delay. In this paper, we propose a privacy-preserving collaborative perception framework, where each agent is separately trained with its own dataset and the ego vehicle needs to integrate with completely unknown collaborators. Specifically, we propose MSD, a multi-scale feature fusion method combined with deformable attention, to better fuse features of different agents. We also propose a plug-in domain adapter to align the features from unknown collaborators to ego-domain. Extensive experiments on the challenging DAIR-V2X and V2V4Real demonstrate that: 1) MSD achieves remarkable performance, outperforming others by at least 2.8% and 6.7% in AP0.7 on DAIR-V2X and V2V4Real, respectively; 2) After domain adaptation, it significantly outperforms the No Fusion, Late Fusion scenarios and can approach or even surpass the performance of joint training. We truly achieves privacy-preserving collaboration, providing a new paradigm for the study of collaborative perception, which is crucial for practical applications. Xinyu Xiao, Changzhou Zhang, Zhiyu Xiang, Hangguan Shan, Eryun Liu |
AAAI | 5 |
| 2025 | SCKD: Semi-Supervised Cross-Modality Knowledge Distillation for 4D Radar Object Detectionabstract3D object detection is one of the fundamental perception tasks for autonomous vehicles. Fulfilling such a task with a 4D millimeter-wave radar is very attractive since the sensor is able to acquire 3D point clouds similar to Lidar while maintaining robust measurements under adverse weather. However, due to the high sparsity and noise associated with the radar point clouds, the performance of the existing methods is still much lower than expected. In this paper, we propose a novel Semi-supervised Cross-modality Knowledge Distillation (SCKD) method for 4D radar-based 3D object detection. It characterizes the capability of learning the feature from a Lidar-radar-fused teacher network with semi-supervised distillation. We first propose an adaptive fusion module in the teacher network to boost its performance. Then, two feature distillation modules are designed to facilitate the cross-modality knowledge transfer. Finally, a semi-supervised output distillation is proposed to increase the effectiveness and flexibility of the distillation framework. With the same network structure, our radar-only student trained by SCKD boosts the mAP by 10.38% over the baseline and outperforms the state-of-the-art works on the VoD dataset. The experiment on ZJUODset also shows 5.12% mAP improvements on the moderate difficulty level over the baseline when extra unlabeled data are available. Zhiyu Xiang, Hanzhi Zhong, Xijun Zhao, Ruina Dang, Peng Xu 0026, Tianyu Pu, Eryun Liu |
AAAI | 2 |
| 2025 | CVFusion: Cross-View Fusion of 4D Radar and Camera for 3D Object Detectionabstract4D radar has received significant attention in autonomous driving thanks to its robustness under adverse weathers. Due to the sparse points and noisy measurements of the 4D radar, most of the research finish the 3D object detection task by integrating images from camera and perform modality fusion in BEV space. However, the potential of the radar and the fusion mechanism is still largely unexplored, hindering the performance improvement. In this study, we propose a cross-view two-stage fusion network called CVFusion. In the first stage, we design a radar guided iterative (RGIter) BEV fusion module to generate high-recall 3D proposal boxes. In the second stage, we aggregate features from multiple heterogeneous views including points, image, and BEV for each proposal. These comprehensive instance level features greatly help refine the proposals and generate high-quality predictions. Extensive experiments on public datasets show that our method outperforms the previous state-of-the-art methods by a large margin, with 9.10% and 3.68% mAP improvements on View-of-Delft (VoD) and TJ4DRadSet, respectively. Our code will be made publicly available. Hanzhi Zhong, Zhiyu Xiang, Jingyun Fu, Peng Xu 0026, Shaohong Wang, Tianyu Pu, Eryun Liu |
ICCV | 2 |
| 2025 | RayFusion: Ray Fusion Enhanced Collaborative Visual PerceptionabstractCollaborative visual perception methods have gained widespread attention in the autonomous driving community in recent years due to their ability to address sensor limitation problems. However, the absence of explicit depth information often makes it difficult for camera-based perception systems, e.g., 3D object detection, to generate accurate predictions. To alleviate the ambiguity in depth estimation, we propose RayFusion, a ray-based fusion method for collaborative visual perception. Using ray occupancy information from collaborators, RayFusion reduces redundancy and false positive predictions along camera rays, enhancing the detection performance of purely camera-based collaborative perception systems. Comprehensive experiments show that our method consistently outperforms existing state-of-the-art models, substantially advancing the performance of collaborative visual perception. Our code will be made publicly available. Shaohong Wang, Lu Bin, Xinyu Xiao, Hanzhi Zhong, Zhiyu Xiang, Hangguan Shan, Eryun Liu |
NeurIPS | 7 |
| 2025 | LiDAR semantic segmentation with local consistency constrained KPConv LSTM
Tingming Bai, Zhiyu Xiang, Xijun Zhao, Peng Xu 0026, Tianyu Pu, Jingyun Fu |
Neurocomputing | 2 |
| 2025 | Bidirectional Segmentation-Aware Network for One-Shot Object Detection
Zhenghua Chen, Yongyi Su, Zhiyu Xiang, Hangguan Shan, Eryun Liu |
Neurocomputing | 4 |
| 2024 | Adaptive Multi-Modal Cross-Entropy Loss for Stereo MatchingabstractDespite the great success of deep learning in stereo matching, recovering accurate disparity maps is still challenging. Currently, L1 and cross-entropy are the two most widely used losses for stereo network training. Compared with the former, the latter usually performs better thanks to its probability modeling and direct supervision to the cost volume. However, how to accurately model the stereo ground-truth for cross-entropy loss remains largely under-explored. Existing works simply assume that the ground-truth distributions are uni-modal, which ignores the fact that most of the edge pixels can be multimodal. In this paper, a novel adaptive multimodal cross-entropy loss (ADL) is proposed to guide the networks to learn different distribution patterns for each pixel. Moreover, we optimize the disparity estimator to further alleviate the bleeding or mis-alignment artifacts in inference. Extensive experimental results show that our method is generic and can help classic stereo networks regain state-of-the-art performance. In particular, GANet with our method ranks 1ston both the KITTI 2015 and 2012 benchmarks among the published methods. Meanwhile, excellent synthetic-to-realistic generalization performance can be achieved by simply replacing the traditional loss with ours. Code is available at https://github.com/xxxupeng/AdL. Peng Xu 0026, Zhiyu Xiang, Chengyu Qiao, Jingyun Fu, Tianyu Pu |
CVPR | 2 |
| 2024 | Lite-SAM Is Actually What You Need for Segment Everything
Jianhai Fu, Yuanjie Yu, Ningchuan Li, Qichao Chen, Jianping Xiong, Zhiyu Xiang |
ECCV (34) | 8 |
| 2024 | IFTR: An Instance-Level Fusion Transformer for Visual Collaborative Perception
Shaohong Wang, Lu Bin, Xinyu Xiao, Zhiyu Xiang, Hangguan Shan, Eryun Liu |
ECCV (87) | 4 |
| 2024 | RISurConv: Rotation Invariant Surface Attention-Augmented Convolutions for 3D Point Cloud Classification and Segmentation
Zhiyuan Zhang 0004, Licheng Yang 0003, Zhiyu Xiang |
ECCV (28) | 3 |
| 2024 | RWT-SLAM: Robust Visual SLAM for Weakly Textured EnvironmentsabstractAs a fundamental task for intelligent robots, visual SLAM has made significant progress in recent years. However, robust SLAM in weakly textured environments remains a challenging task. In this paper, we present a novel visual Robust SLAM for Weak-Textured environments (RWT-SLAM) to address this problem. Unlike existing methods that use detector-based deep networks for interest point detection, we propose extracting distinctive features from a detector-free based network, namely LoFTR, to avoid the difficulty of manual annotations of feature points in weakly textured images. We generate multi-level feature vectors from LoFTR to form dense descriptors for each pixel in the input image. A keypoint localization component is then proposed to measure the saliency of the descriptors and select the distinctive pixels as keypoints. We integrate this new keypoint into the popular ORB-SLAM framework and compare it with the state-of-the-art methods. Extensive experiments on popular TUM RGB-D, OpenLORIS-Scene, as well as our own dataset are carried out. The results demonstrate the superior performance of our method in weakly textured environments. Qihao Peng, Xijun Zhao, Ruina Dang, Zhiyu Xiang |
IV | 4 |
| 2024 | Wheat-YOLO: A Real-Time and High Precision Object Detection for WheatabstractIn the wheat detection work, the wheat is located in a complex environment, weeds and dried leaves will hinder the detection of wheat, the wheat images of shadows, wheat obscuring each other and other phenomena will lead to a reduction in the accuracy of the detection. At the same time, for the small object detection problem, most of the detection algorithms use the detection speed as a cost to improve the detection accuracy, and cannot do a good job of balancing between the two. To address the above issues, this paper proposes an improved YOLO detection algorithm, which aims to improve the detection accuracy of small targets in complex environments while minimising the cost of detection speed. Firstly, two attention mechanisms for small target detection are added to the backbone of YOLOv8; secondly, the target detection header with unified attention is used in the head part, which improves the expressive power of the detection header without any computational expense; and finally, the Inner-MPDIoU function, which is a combination of MPDIoU and Inner ideas, is used as the localisation loss in loss functions.Extensive experiments on publicly available datasets show that the results of the network in this paper are improved in terms of both speed and accuracy compared to the original YOLOv8, enabling a better balance between detection accuracy and speed. Senyue Zhang, Dongdong Sun, Zhiyu Xiang |
SMC | 5 |
| 2024 | Research on Task Assignment of Firefighting UAVs Based on E-CARGO ModelabstractFirefighting UAV technology has become one of the core tools of modern firefighting operations, and its mission execution is constantly expanding in scale and complexity. In the face of this development, it has become even more critical to find an efficient way to ensure that drones can be assigned to perform their most suitable tasks. In this study, we used the Environment-Class, Agent, Role, Group, and Object (E-CARGO) model to systematically analyze the task assignment (FDTA) problem of firefighting drones, and introduced an enhanced whale optimization algorithm (EWOA) to optimize the path planning in the FDTA problem. Finally, simulation experiments are carried out under diverse terrain conditions to demonstrate the efficiency and fast response ability of the improved algorithm under different workloads and environmental conditions. Zhiyu Xiang, Senyue Zhang, Beihang Gao |
SMC | 2 |
| 2023 | Objects matter: Learning object relation graph for robust absolute pose regression
Chengyu Qiao, Zhiyu Xiang, Xinglu Wang, Shuya Chen, YuanGang Fan, Xijun Zhao |
Neurocomputing | 2 |
| 2023 | Improving lane detection with adaptive homography prediction
Yiman Chen, Zhiyu Xiang, Wentao Du |
Vis. Comput. | 2 |
| 2022 | Homography Loss for Monocular 3D Object DetectionabstractMonocular 3D object detection is an essential task in autonomous driving. However, most current methods consider each 3D object in the scene as an independent training sample, while ignoring their inherent geometric relations, thus inevitably resulting in a lack of leveraging spatial constraints. In this paper, we propose a novel method that takes all the objects into consideration and explores their mutual relationships to help better estimate the 3D boxes. More-over, since 2D detection is more reliable currently, we also investigate how to use the detected 2D boxes as guidance to globally constrain the optimization of the corresponding predicted 3D boxes. To this end, a differentiable loss function, termed as Homography Loss, is proposed to achieve the goal, which exploits both 2D and 3D information, aiming at balancing the positional relationships between different objects by global constraints, so as to obtain more ac-curately predicted 3D boxes. Thanks to the concise design, our loss function is universal and can be plugged into any mature monocular 3D detector, while significantly boosting the performance over their baseline. Experiments demon-strate that our method yields the best performance (Nov. 2021) compared with the other state-of-the-arts by a large margin on KITTI3D datasets. Jiaqi Gu 0004, Bojian Wu, Lubin Fan, Jianqiang Huang 0001, Shen Cao, Zhiyu Xiang, Xian-Sheng Hua 0001 |
CVPR | 6 |
| 2022 | CVFNet: Real-time 3D Object Detection by Learning Cross View FeaturesabstractIn recent years 3D object detection from LiDAR point clouds has made great progress thanks to the development of deep learning technologies. Although voxel or point based methods are popular in 3D object detection, they usually involve time-consuming operations such as 3D convolutions on voxels or ball query among points, making the resulting network inappropriate for time critical applications. On the other hand, 2D view-based methods feature high computing efficiency while usually obtaining inferior performance than the voxel or point based methods. In this work, we present a real-time view-based single stage 3D object detector, namely CVFNet to fulfill this task. To strengthen the cross-view feature learning under the condition of demanding efficiency, our framework extracts the features of different views and fuses them in an efficient progressive way. We first propose a novel Point-Range feature fusion module that deeply integrates point and range view features in multiple stages. Then, a special Slice Pillar is designed to well maintain the 3D geometry when transforming the obtained deep point-view features into bird's eye view. To better balance the ratio of samples, a sparse pillar detection head is presented to focus the detection on the nonempty grids. We conduct experiments on the popular KITTI and NuScenes benchmark, and state-of-the-art performances are achieved in terms of both accuracy and speed. Jiaqi Gu 0004, Zhiyu Xiang, Tingming Bai, Lingxuan Wang, Xijun Zhao, Zhiyuan Zhang 0004 |
IROS | 2 |
| 2021 | Real-time Instance Segmentation with Discriminative Orientation MapsabstractAlthough instance segmentation has made considerable advancement over recent years, it’s still a challenge to design high accuracy algorithms with real-time performance. In this paper, we propose a real-time instance segmentation framework termed OrienMask. Upon the one-stage object detector YOLOv3, a mask head is added to predict some discriminative orientation maps, which are explicitly defined as spatial offset vectors for both foreground and background pixels. Thanks to the discrimination ability of orientation maps, masks can be recovered without the need for extra foreground segmentation. All instances that match with the same anchor size share a common orientation map. This special sharing strategy reduces the amortized memory utilization for mask predictions but without loss of mask granularity. Given the surviving box predictions after NMS, instance masks can be concurrently constructed from the corresponding orientation maps with low complexity. Owing to the concise design for mask representation and its effective integration with the anchor-based object detector, our method is qualified under real-time conditions while maintaining competitive accuracy. Experiments on COCO benchmark show that OrienMask achieves 34.8 mask AP at the speed of 42.7 fps evaluated with a single RTX 2080 Ti. The code is available at https://github.com/duwt/OrienMask. Wentao Du, Zhiyu Xiang, Shuya Chen, Chengyu Qiao, Yiman Chen, Tingming Bai |
ICCV | 2 |
| 2021 | Superline: A Robust Line Segment Feature for Visual SLAMabstractAlong with point features, line features play an important role in achieving robust Simultaneous Localization and Mapping (SLAM) under complex environments. This paper proposes a fast and effective method, namely Superline, to simultaneously detect line segments and generate robust descriptors for matching. The entire model is composed of a convolutional backbone and two task heads, i.e., detection head and description head respectively. A line selecting mechanism and a spatial pyramid Line-of-Interest (LOI) pooling module is specially designed in the description head to aggregate multi-scale information into line feature descriptors. The entire model is implemented end-to-end and can be trained on a dataset with only line annotations and without the need of providing ground truth matching. Comparative experimental results on Wireframe and York Urban datasets as well as applying the Superline features on SLAM applications demonstrate the superior performance of our method. Chengyu Qiao, Tingming Bai, Zhiyu Xiang, Yunfeng Bi |
IROS | 3 |
| 2021 | Smart contract for distributed energy trading in virtual power plants based on blockchainabstractAbstract The energy system is evolving from smart grid to energy Internet. Virtual power plant (VPP), as an important part of the energy Internet, plays an important role in the distributed energy generation and trading. In this article, a blockchain‐based VPP transaction model is established for the future energy Internet driven by real‐time electricity price. Then the smart contracts for distributed energy trading in VPPs using blockchain technology are proposed, and the key technological difficulties are analyzed and the solutions are given. Experiments show that the proposed model can reflect the supply and demand information in real time, so that two‐way selection can be carried out under the condition of information symmetry when distributed energy is connected to the grid. If our method is applied, we can help distributed energy suppliers set electricity prices, reduce the trust cost and improve the energy trading efficiency. Also, we can help the distributed energy voluntarily participate in VPPs and joint maintain the system, then solve the problem of VPP's coordinated control and scheduling of distributed energy resources. Jing Lu 0012, Shihong Wu, Hanlei Cheng, Zhiyu Xiang |
Comput. Intell. | 4 |
| 2021 | PGNet: Panoptic parsing guided deep stereo matching
Shuya Chen, Zhiyu Xiang, Chengyu Qiao, Yiman Chen, Tingming Bai |
Neurocomputing | 2 |
| 2020 | SGNet: Semantics Guided Deep Stereo Matching
Shuya Chen, Zhiyu Xiang, Chengyu Qiao, Yiman Chen, Tingming Bai |
ACCV (1) | 2 |
| 2020 | SDP-Net: Scene Flow Based Real-Time Object Detection and Prediction from Sequential 3D Point Clouds
Yi Zhang 0077, Yuwen Ye, Zhiyu Xiang, Jiaqi Gu 0004 |
ACCV (1) | 3 |
| 2020 | A Permissioned Blockchain-Based Platform for Education Certificate Verification
Hanlei Cheng, Jing Lu 0012, Zhiyu Xiang |
BlockSys | 3 |
| 2020 | Partial Fingerprint Verification via Spatial Transformer NetworksabstractPartial fingerprint verification is a challenging task because of the few features contained in small area as well as the large rotation angle and translation between query images and template images. In this paper, we propose a new framework of partial fingerprint verification based on spatial transformer networks (STN) model, where a transform model, i.e., AlignNet network, is proposed to estimate the alignment parameters, and the verification is modeled as a binary classification task. The experimental results on the simulated datasets created from FVC2004 and the real-world dataset FVC2006 DB1 show that our method is invariant to rotation, and also robust to different kinds of scanners, and dramatically outperforms the rank-1 entry of FVC2006 participants. The EER on FVC2006 DB1 of the proposed algorithm is 3.587% compared to that of 5.564%, the best of FVC2006 DB1 entries. Eryun Liu, Zhiyu Xiang |
IJCB | 3 |
| 2020 | DSSF-net: Dual-Task Segmentation and Self-supervised Fitting Network for End-to-End Lane Mark DetectionabstractLane mark detection is one of the key tasks for autonomous driving systems. Accurate detection of lane marks under complex urban environments remains a challenge. In this paper, an end-to-end lane mark detection network named DSSF-net, which is capable of directly outputting the accurate fitted lane curves, is proposed. First, a dual-task segmentation framework for jointing lane category prediction and spatial partition is presented. An IoU-based loss function is put forward to tackle the severely imbalanced category distribution problem. Then a fully self-supervised curve fitting network is proposed to directly output the parameters of lane line upon the probability map. To achieve better accuracy, the fitting network is trained with two sub-stages: coarse regression and confidence-based optimization. Finally the entire DSSF-net is implemented end-to-end. Comprehensive experiments conducted on challenging CULane dataset show that our model achieves 74.9% in F1-score and outperforms the state-of-the-art models. Wentao Du, Zhiyu Xiang, Yiman Chen, Shuya Chen |
IROS | 2 |
| 2020 | Computation Offloading with Reliability Guarantee in Vehicular Edge Computing SystemsabstractThis paper investigates the reliable computation offloading in vehicular edge computing (VEC) systems. Compared with the traditional task replication method in which task replicas are typically assigned to multiple service vehicles at the same time, in our work, a task vehicle allocates the computation tasks and communication resources to its neighboring service vehicles through the vehicle-to-vehicle (V2V) links, and avoids the degradation of delay and computation efficiency. Specifically, an optimization problem is formulated to minimize the task completion delay and ensure offloading reliability. Then, an algorithm based on the penalty and the concave-convex procedure (CCCP) method is proposed to effectively solve the formulated optimization problem. The simulation results show that the task completion delay of the proposed algorithm is only 30% of that in the traditional task replication method. Zhongjie He, Hangguan Shan, Yuanguo Bi, Zhiyu Xiang, Zhou Su 0001, Weihua Wu, Tom H. Luan |
VTC Fall | 4 |
| 2020 | Intelligent document-filling system on mobile devices by document classification and electronizationabstractAbstract Both of an automatic classification method for original documents based on image feature and a layout analysis method based on rule hypothesis tree are proposed. Then an intelligent document‐filling system by electronizing the original documents, which can be applied to cellphones and pads is designed. When users are filling documents online, information can be automatically input to the financial information system merely by taking photos of the original documents. By this means can not only save time but also ensure the accuracy between the data online and that on the original documents. Experiments show that the accuracy of document classification is 88.38%, the accuracy of document‐filling is 87.22%, and it takes 5.042 seconds dealing with per document. The system can be applied to financial, government, libraries, electric power, enterprises and many other industries, which has high economic and application value. Jing Lu 0012, Shihong Wu, Zhiyu Xiang, Hanlei Cheng |
Comput. Intell. | 3 |
| 2020 | RSDCN: A Road Semantic Guided Sparse Depth Completion Network
Nan Zou, Zhiyu Xiang, Yiman Chen |
Neural Process. Lett. | 2 |
| 2019 | Accurate and Real-Time Object Detection Based on Bird's Eye View on 3D Point CloudsabstractAiming at accurate and real-time object detection on 3D point clouds, we proposed a single-stage deep neural network which includes new solutions in three aspects: network architecture, loss function design and data augmentation. Firstly, the point clouds are directly voxelized to build a binary bird's eye view (BEV) map. The network is specially designed to combine the semantic and position information on point clouds to output a final feature map. When regressing the bounding boxes of objects from the bird's eye view, an extra prediction error regression is considered in the loss function to achieve the convergence with higher precision. In training process, a special data augmentation is adopted by mixing 3D point clouds from different frames to improve generalization performance of the network. Experimental results show that our approach achieves higher performance than the state-of-the-art methods on the KITTI BEV object detection benchmark at a frame rate of 20Hz, using only the position information of LIDAR point clouds. Yi Zhang 0077, Zhiyu Xiang, Chengyu Qiao, Shuya Chen |
3DV | 2 |
| 2019 | 3D Reconstruction by Single Camera Omnidirectional Multi-Stereo SystemabstractOmnidirectional catadioptric systems are popular in robotic applications thanks to their large field of view. For 3D scene reconstruction in a single shot, usually two different catadioptric cameras are needed. More cameras may contribute to better reconstruction while larger mounting space and higher power cost are required. In this paper, a single camera multi-stereo catadioptric system with vertical and horizontal baseline structure is proposed. It features achieving multi-pair of central or non-central omnidirectional stereos in a compact manner. To make the 3D reconstruction process general and adaptive to various types of system configurations, a flexible calibration and reconstruction algorithm pipeline is presented. The algorithm features approximating the system into multiple central sub-cameras and carrying out the stereo matching in a spherical representation. In addition, an effective 3D point cloud fusion algorithm is proposed to optimize the reconstruction results from multiple stereo pairs. The experiment carried out with synthetic and real data verified the feasibility and effectiveness of our system. Shuya Chen, Zhiyu Xiang, Nan Zou, Yiman Chen, Chengyu Qiao |
IROS | 2 |
| 2019 | Self-supervised Homography Prediction CNN for Accurate Lane Marking Fitting
Yiman Chen, Wentao Du, Zhiyu Xiang, Nan Zou, Shuya Chen, Chengyu Qiao |
PRCV (3) | 3 |
| 2018 | A Multi-Position Joint Particle Filtering Method for Vehicle Localization in Urban AreaabstractRobust localization is a prerequisite for autonomous vehicles. Traditional visual localization methods like visual odometry suffer error accumulation on long range navigation. In this paper, a flexible road map based probabilistic filtering method is proposed to tackle this problem. To effectively match the ego-trajectory to various curving roads in map, a new representation based on anchor point (AP) which captures the main curving points on the trajectory is presented. Based on APs of the map and trajectory, a flexible Multi-Position Joint Particle Filtering (MPJPF) framework is proposed to correct the position error. The method features the capability of adaptively estimating a series of APs jointly and only updates the estimation at situations with low uncertainty. It explicitly avoids the drawbacks of obliging to determine the current position at large uncertain situations such as dense parallel road branches. The experiments carried out on KITTI benchmark demonstrate our success. Shuxia Gu, Zhiyu Xiang, Yi Zhang 0077 |
IROS | 2 |
| 2017 | Multi-spectrum superpixel based obstacle detection under vegetation environmentsabstractRobust obstacle detection is an important task for unmanned ground vehide(UGV). Vegetation in off-road environments poses great challenges to this task. Usually, vegetation should not be considered as obstacles for off-road UGVs since they are soft and drivable. On the other hand, there are also possibilities that real obstacles exist in the vegetation, which makes the problem difficult. In this paper, a novel multi-spectrum data fusion based algorithm for partial occluded obstacle detection under complex vegetation environment is proposed. First a RGB and Near-infrared (NIR) multi-spectrum superpixel based segmentation strategy is employed to accurately segment the objects in the image. Obstacle candidate superpixels are then obtained through simple geometric computation in 3D laser data. Finally, the heterogeneous texture and 3D features are extracted from each candidate superpixel and fed in Support Vector Machine (SVM) to distinguish the real obstacles from vegetation. Experimental results on real data acquired from various vegetation environments demonstrate our success. Nan Zou, Zhiyu Xiang |
Intelligent Vehicles Symposium | 2 |
| 2017 | Robust object tracking with RGBD-based sparse learningabstractRobust object tracking has been an important and challenging research area in the field of computer vision for decades. With the increasing popularity of affordable depth sensors, range data is widely used in visual tracking for its ability to provide robustness to varying illumination and occlusions. In this paper, a novel RGBD and sparse learning based tracker is proposed. The range data is integrated into the sparse learning framework in three respects. First, an extra depth view is added to the color image based visual features as an independent view for robust appearance modeling. Then, a special occlusion template set is designed to replenish the existing dictionary for handling various occlusion conditions. Finally, a depth-based occlusion detection method is proposed to efficiently determine an accurate time for the template update. Extensive experiments on both KITTI and Princeton data sets demonstrate that the proposed tracker outperforms the state-of-the-art tracking algorithms, including both sparse learning and RGBD based methods. Zi-ang Ma, Zhiyu Xiang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2017 | High-density crowd behaviors segmentation based on dynamical systems
Zhiyu Xiang |
Multim. Syst. | 4 |
| 2017 | Robust Visual Tracking via Binocular Consistent Sparse Learning
Zi-ang Ma, Zhiyu Xiang |
Neural Process. Lett. | 2 |
| 2016 | Robust localization via Turning Point Filtering with road mapabstractTo deal with the frequent failure of GPS in urban areas, vision-based localization methods such as visual odometry (VO) have been popular in recent years. However, VO still suffers from the problem of drift. In this paper, a novel Turning Point Filtering (TPF) algorithm is proposed to restrain the VO's drift by inducing the constraint from a concise road map. Different from the traditional road map based methods, we believe the simple “edge-node” road map is not enough to well model the true trajectory of the vehicle. Therefore our method does not enforce the corrected trajectory exactly on the edges of the map. A flexible turning point filtering mechanism is designed under a particle filter framework to well balance the information from the VO and the road map. The method features making reliable corrections only on the turning points of the trajectory, which adds little additional computation burden to VO. Experiments on various datasets including KITTI and the data acquired in our campus demonstrate the outperformance of our method. Yidong Jin, Zhiyu Xiang |
Intelligent Vehicles Symposium | 2 |
| 2015 | Design of an enhanced visual odometry by building and matching compressive panoramic landmarks onlineabstractEfficient and precise localization is a prerequisite for the intelligent navigation of mobile robots. Traditional visual localization systems, such as visual odometry (VO) and simultaneous localization and mapping (SLAM), suffer from two shortcomings: a drift problem caused by accumulated localization error, and erroneous motion estimation due to illumination variation and moving objects. In this paper, we propose an enhanced VO by introducing a panoramic camera into the traditional stereo-only VO system. Benefiting from the 360° field of view, the panoramic camera is responsible for three tasks: (1) detecting road junctions and building a landmark library online; (2) correcting the robot’s position when the landmarks are revisited with any orientation; (3) working as a panoramic compass when the stereo VO cannot provide reliable positioning results. To use the large-sized panoramic images efficiently, the concept of compressed sensing is introduced into the solution and an adaptive compressive feature is presented. Combined with our previous two-stage local binocular bundle adjustment (TLBBA) stereo VO, the new system can obtain reliable positioning results in quasi-real time. Experimental results of challenging long-range tests show that our enhanced VO is much more accurate and robust than the traditional VO, thanks to the compressive panoramic landmarks built online. Zhiyu Xiang, Jilin Liu |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2015 | Perception in Disparity: An Efficient Navigation Framework for Autonomous Vehicles With Stereo CamerasabstractStereo cameras are widely used in autonomous vehicles as environmental perception sensors because of their availability and low cost. However, efficiently utilizing the obtained disparity images to generate a desirable local path for the vehicle still remains a challenging problem. In this paper, we present a novel navigation framework for autonomous vehicles equipped with stereo cameras, featuring finishing all of the perception and path planning tasks directly within the disparity space. Comparing with the popular 3-D counterpart, disparity is a projected geometric space that contains more primitive information directly computed from the stereo images. Furthermore, disparity image is a more compact representation for large field of 3-D Cartesian space, which makes perception and path finding in longer distance possible. The proposed framework is composed of three modules, namely, local disparity map building, slope analysis and obstacle detection, and path planning. Two important properties concerning the motion and slope in disparity space are presented for the first time, i.e., the motion model and the slope model in disparity space. With the motion model, the framework first fuses consecutive disparity maps to construct a more reliable and complete local map. Then, a novel slope analysis method called V-Intercept is developed based on the slope model. It can efficiently analyze slopes and obstacles, generating a reasonable cost map directly from the disparity image. Finally, the obstacles in the cost map are expanded properly, and a customized A* search algorithm is performed to find a reasonable path in disparity space. The experimental results show that our framework works well under various kinds of environments. The resulting system can efficiently perceive and plan on a much larger range and react to obstacles further beyond the traditional Cartesian-based method. Teng Cao, Zhiyu Xiang, Jilin Liu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Robust vehicle detection using 3D Lidar under complex urban environmentabstractRobust vehicle detection is one of the key task for autonomous vehicle under the complex urban environment. Using 3D Lidar, the difficulty of the task lies in that the appearance of a vehicle in the sparse range data changes greatly with the distance, the angle of view, as well as occlusions. In this paper we present a new algorithm to detect vehicles using the finely segmented 3D object points. In segmentation, RGLOS (Ring Gradient based Local Optimal Segmentation) algorithm is proposed. Instead of processing in the grid map, the point-wise segmentation method keeps the connection between points and is able to extract object points in far distance. Using the local optimal ground height, it produces much more correct object points with less wrong ground points. In feature extraction stage, three novel features: position-related shape, object height along length and reflective intensity histogram are proposed. Finally the kernel based SVM is used to finish the classification task. Experiments are carried out using the real data acquired from urban environment. The results demonstrate the superior performance comparing with previous methods, thanks to the improved segmentation and new features. Zhiyu Xiang, Teng Cao, Jilin Liu |
ICRA | 2 |
| 2014 | Road scene segmentation via fusing camera and lidar dataabstractThis paper presents an approach for pixel-wise object segmentation for road scenes based on the integration of a color image and an aligned 3D point cloud. In light of the advantage of range information in object discovery, we first produce initial object hypotheses by clustering the sparse 3D point cloud. The image pixels registered to the clustered 3D points are taken as samples to learn each object's prior knowledge. The priors are represented by Gaussian Mixture Models (GMMs) of color and 3D location information only, requiring no high-level features. We further formulate the segmentation problem within a Conditional Random Field (CRF) framework, which incorporates the learned prior models, together with hard constraints placed on the registered pixels and pairwise spatial constraints to achieve final results. Our algorithm is validated on the challenging KITTI dataset which contains diverse complicated road scenarios. Both qualitative and quantitative evaluation results show the superiority of our algorithm. Wenqi Huang 0002, Xiaojin Gong, Zhiyu Xiang |
ICRA | 3 |
| 2013 | High-performance visual odometry with two-stage local binocular BA and GPUabstractVisual odometry becomes an important method to deal the localization work in intelligent vehicle and robotics. A high performance visual odometry needs to achieve two requirements: high accuracy and high frequency. So we propose a two-stage local binocular bundle adjustment algorithm doing the optimization and construct a parallel pipeline using GPU acceleration. Finally, our system can run at about 35~40 frames per second with the maximum RMS 3D localization error less than 1%. Zhiyu Xiang, Jilin Liu |
Intelligent Vehicles Symposium | 2 |
| 2010 | An adaptive fast search algorithm for block motion estimation in H.264abstractMotion estimation is an important issue in H.264 video coding systems because it occupies a large amount of encoding time. In this paper, a novel search algorithm which utilizes an adaptive hexagon and small diamond search (AHSDS) is proposed to enhance search speed. The search pattern is chosen according to the motion strength of the current block. When the block is in active motion, the hexagon search provides an efficient search means; when the block is inactive, the small diamond search is adopted. Simulation results showed that our approach can speed up the search process with little effect on distortion performance compared with other adaptive approaches. Congdao Han, Jilin Liu, Zhiyu Xiang |
J. Zhejiang Univ. Sci. C | 3 |
| 2007 | Polarization-based water hazards detection for autonomous off-road navigationabstractA polarization-based method for water hazards detection is presented. The concept of polarization is introduced to computer vision to detect water hazards for autonomous off-road navigation. This method is based on the physical principle that the light reflected from water surface is partial linearly polarized and the polarization phases of them are more similar than those from the scenes around. Water hazards can be detected by comparison of polarization degree and similarity of the polarization phases. Experiments show that the method has good performance in water detection in complex natural backgrounds, especially when there is vegetation reflected in the water. This method has significant complementary advantages with respect to existing techniques, is computationally efficient, and can be easily implemented with existing imaging technology. Bin Xie 0002, Zhiyu Xiang, Huadong Pan, Jilin Liu |
IROS | 2 |
| 2006 | Enhancing Particle Swarm Optimization Based Particle Filter Tracker
Qicong Wang, Jilin Liu, Zhiyu Xiang |
ICIC (2) | 4 |
| 2004 | Locating and Crossing Doors and Narrow Passages for a Mobile Robot
Zhiyu Xiang, Vítor M. F. Santos, Jilin Liu |
ICINCO (2) | 1 |