Zhiye Wang

dblp:44/8548 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MaestroBot: Generalized Gesture-Driven Hierarchical Coordination for Robotic Formations
abstract
Robotic swarm coordination holds transformative potential for applications such as warehouse automation, search & rescue, and entertainment. However, approaches relying on wearable devices or vision-based systems are often constrained by hardware-intensive, high computational requirements, reliance on line-of-sight, and privacy concerns. Wireless sensing, particularly using Channel State Information (CSI), offers a promising alternative by translating environmental perturbations into CSI variation data. Nevertheless, existing CSI-based systems face significant challenges in domain adaptation, resource limitation, and scalability issues. This paper introduces MaestroBot, a hierarchical motion coordination system that combines distributed CSI-based wireless sensing with domain-adaptive learning to address these limitations. For leader robots, the system features a lightweight hand gesture recognition model, built on a “Hybrid-Single” knowledge distillation framework, achieving up to 95.87% accuracy while maintaining adaptability across diverse domains. For follower robots, the hierarchical motion propagation model leverages localized CSI analysis and dual-layer error correction mechanisms to deliver 97.2% accuracy with a low latency of 0.085 seconds, even in multi-row formations. Additionally, its cost-effective hardware design ensures practical scalability and real-world deployability. These results position MaestroBot as an efficient, robust, and privacy-preserving solution for large-scale robotic swarm coordination in dynamic environments.
Zhiye Wang, Yuhan Xu, Haiming Jin, Linghe Kong, Rui Li 0098, Xi Chen 0009, Qiao Xiang, Guihai Chen
IEEE Trans. Mob. Comput.2
2025 E2MN: human-inspired end-to-end mapless navigation with oscillation suppression and short-term memory
abstract
Robotic navigation in unknown environments is challenging due to the lack of high-definition maps. Building maps in real time requires significant computational resources. Nevertheless, sensor data can provide sufficient environmental context for robots’ navigation. This paper presents an interpretable and mapless navigation method using only two-dimensional (2D) light detection and ranging (LiDAR), mimicking human strategies to escape from dead ends. Unlike traditional planners, which depend on global paths or vision-based and learning-based methods, requiring heavy data and hardware, our approach is lightweight and robust, and it requires no prior map. It effectively suppresses oscillations and enables autonomous recovery from local minimum traps. Experiments across diverse environments and routes, including ablation studies and comparisons with existing frameworks, show that the proposed method achieves map-like performance without a map—reducing the average path length by 50.51% when compared to the classical mapless Bug2 algorithm and increasing it by only 17.57% when compared to map-based navigation.
Zhiye Wang, Xuan Kong, Peng Zhi, Rui Zhou 0005, Qingguo Zhou
Frontiers Inf. Technol. Electron. Eng.2
2025 ATCM: Aerial-Terrestrial LiDAR-Based Collaborative Simultaneous Localization and Mapping
abstract
Multi-robot collaborative simultaneous localization and mapping (C-SLAM) offers precise scene reconstruction over single-robot SLAM and enables the data fusion from heterogeneous robots. However, heterogeneous C-SLAM faces challenges in both accurate inter-robot loop closure detection and globally consistent data fusion due to inherent viewpoint disparities and heterogeneous data characteristics. This paper introduces ATCM, an Aerial-Terrestrial LiDAR-based C-SLAM method designed for heterogeneous robots without priori initial relative position. ATCM comprises three modules: single-robot front-end employing diverse SLAM methods, multi-robot loop closure detection, and global pose graph optimization. A novel LiDAR-based cross-view global loop descriptor is proposed for scan-to-scan heterogeneous inter-robot loop closure detection. By uniformly mapping cross-view information into the height domain and integrating dynamic height, the loop descriptor automatically achieves viewpoint correction. Additionally, we introduce a bidirectional loop detection algorithm that validates inter-robot loop closures through both forward and reverse detections. Finally, the two-stage global pose graph optimization integrates multi-source measurements, ensuring globally consistent mapping and localization with cross-view data. We have validated the effectiveness of ATCM on campus scenario datasets and the KITTI dataset, achieving a remarkable 21.95% improvement in trajectory accuracy and a 17.00% enhancement in map precision compared to high-precision point cloud maps, surpassing state-of-the-art LiDAR-based odometry methods. In the ablation experiments, the proposed loop descriptor achieved 97% accuracy in recognizing heterogeneous inter-robot loop closures. Moreover, compared to the traditional unidirectional method, the bidirectional loop detection method demonstrates up to a 31.2% improvement in loop closure accuracy.
Chi Chen 0002, Bisheng Yang, Weitong Wu 0002, Shangzhe Sun, Zhiye Wang, Liuchun Li, Qin Zou 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 GPR-Former: Detection and Parametric Reconstruction of Hyperbolas in GPR B-Scan Images With Transformers
abstract
Ground Penetrating Radar (GPR) enables the non-invasive detection of various subsurface objects such as pipes, stones, etc. The location and size of the object in the medium could be obtained by fitting the generated hyperbolic signatures within the GPR B-scan and analyzing its parameters. In this paper, GPR-Former is proposed for automatic target detection and hyperbola fitting on GPR B-scan images. We have designed a transformer-based neural network to extract features to directly regress the parameters of hyperbolic signatures in the GPR B-scan data to detect targets beneath the ground automatically. A symmetry-constrained analytical solution for the hyperbolic parameters is proposed to refine the parameters derived from the transformer network, serving the extraction and analysis of buried objects in underground opaque spaces. Experiments are conducted on three datasets for the qualitative and quantitative validation of the GPR-Former, including ground-penetrating radar detection of submarine pipelines and land pipelines. Results show that the proposed method is able to automatically and efficiently extract hyperbolas from GPR B-scan images. True hyperbola-point precision (TP_Pre) and true hyperbola-point recall (TP_Rec) metrics are introduced to evaluate performances in parametric hyperbola extraction and fitting. The results show that the TP_Pre and TP_Rec of the proposed method reach 0.867, 0.402, 0.744 and 0.762, 0.736, 0.723, with an improvement of 6%, 22%, 4% compared with the state-of-the-art methods (C3 algorithm and migration learning-based method proposed by Yang), respectively.
Ang Jin, Chi Chen 0002, Bisheng Yang, Qin Zou 0001, Zhiye Wang, Zhengfei Yan, Shaolong Wu, Jian Zhou 0011
IEEE Trans. Geosci. Remote. Sens.5
2024 SGSR-Net: Structure Semantics Guided LiDAR Super-Resolution Network for Indoor LiDAR SLAM
abstract
Multi-Beam LiDAR (MBL) sensors sample the real-world with discrete 3D point clouds (PC) and have become a major and essential 3D sensing capability for autonomous robots. To ensure an accurate point sampling on surfaces, high-resolution MBL sensors (e.g., Ouster OS0-128) are commonly used to collect dense point clouds for robot tasks, including object detection and tracking, simultaneous localization and mapping (SLAM), in applications such as autonomous driving vehicles (ADVs). However, the high cost and large volume/weight/energy consumption of such sensors limit their usage in broader applications such as UAV/UGV swarms with small-scale agents with limited payload. Existing studies on Super-Resolution (SR) upsampling of the PC from low-resolution MBL have not considered the geometry semantics of the scenes, thus resulting in less optimal SR points for downstream subtasks (e.g., SLAM). Thus, this article proposes SGSR-Net, a structure semantics-guided MBL Super-Resolution network. SGSR-Net takes the low-resolution range images of the MBL sensors as input and produces dense and structure-aware Super-Resolution point cloud from those sparse measurements through a vertical spatial and channel attention-enhanced CNN model coupling with guided Monte Carlo filtering, for indoor LiDAR-SLAM applications. The SGSR-Net is validated using datasets collected by a UGV equipped with multiple MBL sensors. The results demonstrate that the proposed CG-LSR (CASE Attention Guided Encoder-Decoder LiDAR Super-Resolution Network) reduces the MAE of the SR points by 12.4% down to 0.177 m when compared with the state-of-the-art (SOTA) method Shan et al. (2020), Ren et al. (2021), Kwon et al. (2022), Long and Wang (2022). The indoor SLAM results with SR-points produced by SGSR-Net show that the mean and RMSE of the absolute pose error (APE) are decreased by 27% and 30%, down to 0.849 m and 0.902 m, respectively, which significantly improve the indoor-SLAM performance and stability of SOTA LiDAR-SLAM systems (i.e. LeGO-LOAM Shan and Englot (2018), Dellenbach et al. (2022), Vizzo et al. (2023), Zhang and Singh (2014)).
Chi Chen 0002, Ang Jin, Zhiye Wang, Yongwei Zheng, Bisheng Yang, Jian Zhou 0011, Zhigang Tu 0001
IEEE Trans. Multim.3
2023 Revisiting Data Poisoning Attacks on Deep Learning Based Recommender Systems
abstract
Deep learning based recommender systems(DLRS) as one of the up-and-coming recommender systems, and their robustness is crucial for building trustworthy recommender systems. However, recent studies have demonstrated that DLRS are vulnerable to data poisoning attacks. Specifically, an unpopular item can be promoted to regular users by injecting well-crafted fake user profiles into the victim recommender systems. In this paper, we revisit the data poisoning attacks on DLRS and find that state-of-the-art attacks suffer from two issues: user-agnostic and fake-user-unitary or target-item-agnostic, reducing the effectiveness of promotion attacks. To gap these two limitations, we proposed our improved method Generate Targeted Attacks(GTA), to implement targeted attacks on vulnerable users defined by user intent and sensitivity. We initialize the fake users by adding seed items to address the cold start problems of fake users so that we can implement targeted attacks. Our extensive experiments on two real-world datasets demonstrate the effectiveness of GTA.
Zhiye Wang, Baisong Liu, Chennan Lin, Xueyuan Zhang, Ce Hu, Jiangcheng Qin, Linze Luo
ISCC1
2023 Two-level Data Augmentation for Calibrated Multi-view Detection
abstract
Data augmentation has proven its usefulness to improve model generalization and performance. While it is commonly applied in computer vision application when it comes to multi-view systems, it is rarely used. Indeed geometric data augmentation can break the alignment among views. This is problematic since multi-view data tend to be scarce and it is expensive to annotate.In this work we propose to solve this issue by introducing a new multi-view data augmentation pipeline that preserves alignment among views. Additionally to traditional augmentation of the input image we also propose a second level of augmentation applied directly at the scene level. When combined with our simple multi-view detection model, our two-level augmentation pipeline outperforms all existing baselines by a significant margin on the two main multi-view multi-person detection datasets WILD-TRACK and MultiviewX.
Martin Engilberge, Haixin Shi, Zhiye Wang, Pascal Fua
WACV3
2023 Crisscross-Global Vision Transformers Model for Very High Resolution Aerial Image Semantic Segmentation
abstract
Semantic segmentation is a key means for understanding very-high resolution (VHR) aerial imagery. With the explosive development of deep learning, deep learning methods are being applied to the segmentation of VHR images, with convolutional neural networks (CNNs) as the basic framework. However, owing to the highly complex details present in VHR images and the high spatial dependence of geographical objects, CNN-based methods are inadequate. This is because the inherent locality of CNNs limits the size of the receptive field, thus limiting the ability to obtain long-range context information. To solve this problem, in this paper, we propose a transformer-based novel deep learning model called crisscross-global vision transformers (CGVT). CGVT exploits the transformer’s inherent ability to obtain long-range context information to solve the restricted receptive field problem. Specifically, we redesign the self-attention mechanism in the transformer and call it crisscross-global attention. It consists of two parts: crisscross transformer encoder block (CC-TEB) and global squeeze transformer encoder block (GS-TEB). CC-TEB overcomes the limitation of the traditional self-attention design (specifically, difficulty applying it to VHR aerial image segmentation) and further increases the local feature representation ability of the model. GS-TEB increases the global feature representation ability of the model. The results of experiments conducted on the popular ISPRS Vaihingen, IEEE GRSS Data Fusion Contest Zeebrugge, and LoveDA Semantic Segmentation Challenge datasets verify the effectiveness and superiority of our proposed method. Specifically, it achieved state-of-the-art performance on both Zeebrugge and LoveDA datasets, and is currently ranked second in Vaihingen dataset.
Guohui Deng, Zhaocong Wu, Miaozhong Xu, Zhiye Wang, Zhongyuan Lu
IEEE Trans. Geosci. Remote. Sens.5
2022 A Federated Multi-Server Knowledge Graph Embedding Framework For Link Prediction
abstract
The federated framework is actively applied in knowledge graph fusion research to obtain a complete knowledge graph without exposing data privacy. It can help local clients learn the knowledge graph embeddings in other clients without revealing data privacy. However, current federated-based knowledge graph embedding frameworks cannot exploit both entity and relation embeddings and may not prevent partial triples from being reconstructed. This paper proposes a novel framework named Federated Multi-server knowledge graph embedding (FedM), which creatively utilizes uploaded entity and relation embeddings while preventing privacy leakage. Expressly, we first set up two central servers for entity and relation embeddings to aggregate and share client-uploaded embeddings. Secondly, we design a knowledge graph secure aggregation algorithm to address the potential privacy concerns in FedM. We conduct comparative experiments on an empirical dataset (divided into three federated datasets) with four commonly-used knowledge graph embedding methods to evaluate the performance of our proposed framework. In addition, our proposed FedM framework is generally superior to the latest baseline frameworks on both privacy preservation and link prediction tasks.
Ce Hu, Baisong Liu, Xueyuan Zhang, Zhiye Wang, Chennan Lin, Linze Luo
ICTAI4
2022 Privacy-Preserving Recommendation with Debiased Obfuscaiton
abstract
As people enjoy the personalized services recommended by Recommender Systems (RSs), the privacy disclosure risk increases with frequent interactions. Malicious adversary often collects public information online to infer private information for illicit profit. As privacy concerns grew, researchers introduced data obfuscation into recommender systems. However, there still exists several limitations in current work. First, although the existing methods effectively reduce the risk of privacy disclosure, they can be detrimental to the quality of the recommendation service. Second, a range of practical issues under the application of recommendation systems are not considered, e.g., long-tail, density, etc. To address those challenges, we propose a novel framework named Want User Defending Inference (WUDI), a high-performance privacy-preserving debiased framework based on data obfuscation. Unlike the original strategies, i.e., adding or removing user ratings, we introduced some novel strategies to generate an obfuscated matrix. Firstly, we define a new method called Cluster Recommend for alleviating the long-tail skewness and data sparsity in RSs. Then we investigate the gender bias in obfuscation and apply a bias mitigating strategy to RSs. Experiments on public datasets demonstrate that WUDI can outperform the state-of-the-art baselines in obfuscation.
Chennan Lin, Baisong Liu, Xueyuan Zhang, Zhiye Wang, Ce Hu, Linze Luo
TrustCom4
2022 Effectively Clustering Single Cell RNA Sequencing Data by Sparse Representation
abstract
Clustering analysis has been widely used in analyzing single-cell RNA-sequencing (scRNA-seq) data to study various biological problems at cellular level. Although a number of scRNA-seq data clustering methods have been developed, most of them evaluate the similarity of pairwise cells while ignoring the global relationships among cells, which sometimes cannot effectively capture the latent structure of cells. In this paper, we propose a new clustering method SPARC for scRNA-seq data. The most important feature of SPARC is a novel similarity metric that uses the sparse representation coefficients of each cell in terms of the other cells to measure the relationships among cells. In addition, we develop an outlier detection method to help parameter selection in SPARC. We compare SPARC with nine existing scRNA-seq data clustering methods on twelve real datasets. Experimental results show that SPARC achieves the state of the art performance. By further analyzing the cell similarity data derived from sparse representations, we find that SPARC is much more effective in mining high quality clusters of scRNA-seq data than two traditional similarity metrics. In conclusion, this study provides a new way to effectively cluster scRNA-seq data and achieves more accurate clustering results than the state of art methods.
Ruiyi Li, Zhiye Wang, Jihong Guan, Shuigeng Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.2