Xiaoqian Zhu

dblp:50/51 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 7 since 2021Systems, architecture and hardware · 6 · 2 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Underwater acoustic signal denoising with diffusion-based generative models
Boqing Zhu, Yanxin Ma, Zemin Zhou, Xiaoqian Zhu
Signal Process.6
2025 RegionMatch: Pixel-Region Collaboration for Semi-Supervised Semantic Segmentation in Remote Sensing Images
abstract
Semi-supervised semantic segmentation (S4) has shown significant promise in reducing the burden of labor-intensive data annotation. However, existing methods mainly rely on pixel-level information, neglecting the strong region consistency inherent in remote sensing images (RSIs), which limits their effectiveness in handling the complex and diverse backgrounds of RSIs. To address this, we propose RegionMatch, a novel approach that leverages unlabeled data from a fresh object-level perspective, which is more tailored to the nature of semantic segmentation. We design the Pixel-Region Synergy Pseudo-Labeling strategy, which explicitly injects object-level contextual information into the S4 pipeline and promotes knowledge collaboration between pixel and region perspectives for generating high-quality pseudo-labels. In addition, we propose the Region Structure-Aware Correlation Consistency, which models object-level relationships by establishing inter-region correlations across images and pixel correlations within regions, providing more effective supervision signals for unlabeled data. Experimental results demonstrate that RegionMatch outperforms state-of-the-art methods on multiple authoritative remote sensing datasets, highlighting its superiority in the RSIs.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Chaowei Fang, Xu Tang 0004, Licheng Jiao
IJCAI1
2025 Beyond Single Pixel: Context Priors Guided Semi-Supervised Building Footprint Segmentation
abstract
Automated building footprint segmentation is crucial in remote sensing with widespread applications in various fields. Semi-supervised semantic segmentation methods are gaining traction in the remote sensing community as they significantly reduce the need for labor-intensive pixel-level annotations in training segmentation models. However, these methods typically rely on individual pixel-level supervision for unlabeled data, neglecting the contextual relationships between pixels. This limits their potential to exploit unlabeled data. To bridge this gap, this paper proposes a novel Context Priors Guided Semi-Supervised Building Footprint Segmentation method that leverages contextual relationships among numerous unlabeled pixels to build supervisory signals that extend beyond individual pixel-level guidance for learning on unlabeled data. The CPG comprises two main components: Spatial context priors-guided pseudo-label regularization and semantic context priors-guided representation learning. By integrating contextual knowledge from both pixel spatial locations and semantic representation spaces, these components capture comprehensive class semantic attributes and enable the model to learn complete shapes of building footprints. Our approach achieves state-of-the-art performance on three publicly available building footprint segmentation datasets, validating its effectiveness.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xiao Han 0012, Licheng Jiao, Lianchao Zhang
IEEE Trans. Geosci. Remote. Sens.1
2025 Temporal-Feedback Self-Training for Semi-Supervised Object Detection in Remote Sensing Images
abstract
Although modern Remote Sensing Object Detection (RSOD) methods have achieved advanced performance, they heavily rely on a large amount of annotated data. This paper explores semi-supervised RSOD to mitigate annotation costs, leveraging recent extensive research in generic Semi-Supervised Object Detection (SSOD) based on the self-training paradigm. Current SSOD methods encounter challenges in adapting to remote sensing images due to the complexity and variability of RSIs. Two key issues remain underexplored: the noise in pseudo-labels caused by model instability and the difficulty in distinguishing similar categories. This paper introduces the Temporal-Feedback Self-Training (TST) framework, a novel approach to tackle these challenges in semi-supervised RSOD. TST consists of two components: Temporal Consistency Based Pseudo-labels Certainty Estimation (TCE) and Temporal Self-Feedback Feature Refinement (TSF). TCE addresses pseudo-label noise during training by evaluating the stability of pseudo-label classification and localization over time series to assess the quality of pseudo-labels. On the other hand, TSF enhances pseudo-label quality by dynamically identifying the models confusing categories as feedback for feature refinement. Both components facilitate the progression of the self-training-based RSOD during training. We conducted extensive experiments on two challenging public datasets, DOTA and DIOR. The results demonstrate that the proposed TST and TCE components significantly improve the baseline models performance, surpassing the state-of-the-art generic SSOD method. This suggests that our approach is more effective than generic SSOD methods in addressing the challenges posed by remote sensing images.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2025 A Novel Expandable Borderline Smote Over-Sampling Method for Class Imbalance Problem
abstract
The class imbalance problem can cause classifiers to be biased toward the majority class and inclined to generate incorrect predictions. While existing studies have proposed numerous oversampling methods to alleviate class imbalance by generating extra minority class samples, these methods still have some inherent weaknesses and make the generated samples less informative. This study proposes a novel over-sampling method named the Expandable Borderline Smote (EB-Smote), which can address the weaknesses of existing over-sampling methods and generate more informative synthetic samples. In EB-Smote, not only minority class but also majority class is oversampled, and the synthetic samples are generated in the area between the selected minority and majority samples, which are close to the borderlines of their respective classes. EB-Smote can generate more informative samples by expanding the borderlines of minority and majority classes toward the actual decision boundary. Based on 27 imbalanced datasets and commonly used machine learning models, the experimental results demonstrate that EB-Smote significantly outperforms the other 8 existing oversampling methods. This study can provide theoretical guidance and practical recommendations to solve the crucial class imbalance problem in classification tasks.
Jianping Li 0001, Xiaoqian Zhu
IEEE Trans. Knowl. Data Eng.3
2024 Learning incremental audio-visual representation for continual multimodal understanding
abstract
Deep learning methods have demonstrated remarkable success in processing static datasets for various video tasks. However, when confronted with continuous data streams, these approaches often encounter the challenge of catastrophic forgetting. This phenomenon leads to a significant decline in overall performance when learning new classes incrementally. Moreover, existing methods tend to overlook the correlation between audio and visual modalities in video incremental learning, despite their joint significance in scene comprehension. How to continuously learn from new classes while maintaining the knowledge of old videos with limited storage and computing resources is becoming imperative in the field of multimodal learning. In this paper, we introduce CavRL, a pioneering benchmark for audio–visual representation learning under class incremental scenarios. To mitigate catastrophic forgetting, we propose a rehearsal-based training approach that leverages a small exemplar set from previous classes. Our approach constrains the memory buffer within strict storage limits, optimizing exemplar selection by learning correlative audio–visual representations. Additionally, we employ a distillation method to mitigate forgetting in a self-supervised manner. Evaluations on two prevalent multimodal tasks: audio–visual event classification and audio–visual speaker recognition, which demonstrate that CavRL outperforms existing state-of-the-art incremental learning methods across various settings. We anticipate that CavRL will significantly advance research in continual multimodal learning. • Catastrophic Forgetting Challenge in Multimodal Learning . Deep learning grapples with catastrophic forgetting in continuous mulitimodal data. • Consideration Audio–Visual Correlation . Consideration relationship between audio and visual modalities in continual learning. • Self-Supervised Distillation to Alleviate Forgetting . CavRL implements a self-supervised distillation method to reduce forgetting. • CavRL Benchmark . CavRL is a new benchmark for audio–visual learning in class incremental scenarios.
Boqing Zhu, Kele Xu, Zemin Zhou, Xiaoqian Zhu
Knowl. Based Syst.6
2024 Masking Hierarchical Tokens for Underwater Acoustic Target Recognition With Self-Supervised Learning
abstract
Deep learning has made data-driven methods effective in underwater acoustic target recognition (UATR) using passive sonar signals. However, a major current challenge is the limited availability of underwater acoustic data, leading to suboptimal performance without sufficient data. Self-supervised learning (SSL) can help address this problem by learning intrinsic patterns within acoustic data. Nonetheless, applying SSL in UATR systems requires efficient learning of meaningful representations that can provide quick prediction speed for real-time recognition systems. To this end, we propose the masking hierarchical tokens (MHT) method to learn meaningful representations via efficient self-supervised learning for our previously proposed UATR-Transformer, giving rise to the MHT-UATR-Transformer. In particular, the MHT-UATR-Transformer first exploits a new designed token-convolution-based hierarchical tokenization to efficiently obtain rich time–frequency information from the input Mel-spectrogram. Then, most of these tokens are masked with a high masking ratio and subsequently reconstructed by an integrated Encoder–Decoder structure. In this way, the MHT-UATR-Transformer can learn intrinsic representations of underwater acoustic signals to achieve better recognition performance with fewer labeled data, which is expected to alleviate the dependency on expensive acoustic data. Experimental results on two widely studied underwater databases show that our proposed method achieves better performance than supervised learning and state-of-the-art SSL method in both accuracy and speed, especially in few-shot and noisy scenarios, thus enhancing its practicality in real marine applications.
Sheng Feng, Xiaoqian Zhu, Shuqing Ma
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Multistage Enhancement Network for Tiny Object Detection in Remote Sensing Images
abstract
With the rapid advances in deep learning techniques, remote sensing object detection has achieved remarkable achievements in recent years. However, tiny object detection remains unsatisfactory and suffers from two main drawbacks, including (1) the high sensitivity of IoU for location deviation in tiny objects and (2) the poor-quality feature representations of tiny objects. To address the aforementioned problems, we propose a Multi-stage Enhancement Network (MENet) that achieves the instance-level and feature-level enhancement of tiny objects from different stages of the detector. Since the IoU-based label assignment drastically deteriorates the positive samples for tiny objects, we first propose a Central Region-based (CR) label assignment to substitute it in the Region Proposal Network (RPN). The CR label assignment regards the anchors that fall into the central region of ground-truth boxes as positive samples, which provides more positive samples for tiny objects. Then, we design a Gated Context Aggregation (GCA) module that selectively aggregates valuable context information to enhance the feature representation of tiny objects. Additionally, we devise a positive RoI feature (pRoI) generator in the Region Convolutional Neural Network (R-CNN) to generate a rich diversity of high-quality positive RoI features for tiny objects. We conduct extensive experiments on AI-TOD and SODA-A datasets, and the results demonstrate the effectiveness of our proposed method.
Tianyang Zhang 0002, Xiangrong Zhang, Xiaoqian Zhu, Guanchun Wang, Xiao Han 0012, Xu Tang 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.3
2023 A novel authentication and key agreement scheme for Internet of Vehicles
abstract
With the proposal of the intelligent transportation system, vehicular ad-hoc networks have been widely concerned and well-developed. Vehicular ad-hoc network is recognized as a major innovation of the Internet of Things technology. It can collect and analyze traffic data uploaded by vehicles through the network, monitor vehicle state in real-time, provide better driving routes and real-time traffic decisions for vehicles, and improve vehicles’ overall level of intelligent driving. To make vehicles quickly join the vehicular ad-hoc networks through the authentication of identity legitimacy, take into account security and efficiency, and further reduce the calculation and communication consumption in authentication, this paper designs a mutual anonymous authentication and key agreement scheme based on an elliptic curve for the vehicular ad-hoc networks. To complete vehicle authentication and session key establishment, we work with lightweight operations like hash, XOR, and connection in conjunction with the elliptic curve discrete logarithm problem to guarantee the confidentiality of communication data. In the scheme, identity authentication is divided into two types: initial authentication and subsequent authentication. When the vehicle is just on the road, it will use the first roadside unit it encounters for initial authentication. All roadside units that cars on the road come across are then authenticated after the accomplishment of the initial authentication with the first roadside unit. Subsequent authentication is lighter and less computationally complex than initial authentication. Of course, subsequent authentication is based on initial authentication. This paper also analyzes the scheme’s security and uses BAN logic analysis and Proverif simulation to verify the scheme’s security. Additionally, performance analysis is used to demonstrate the scheme’s superiority.
Xiaoqian Zhu, Xiaoliang Wang 0002, Junjie Fu
Future Gener. Comput. Syst.2
2023 Semantics and Contour Based Interactive Learning Network for Building Footprint Extraction
abstract
Building footprint extraction plays an important role in the analysis of remote sensing images and has an extensive range of applications. Obtaining precise boundaries of buildings remains a challenge in existing building extraction methods. Some previous works have made notable efforts to address this concern. However, most of these methods require cumbersome and expensive post-processing steps. Moreover, they ignored the correlation between building semantics and contours, which we believe is crucial for building footprint extraction. To mitigate this issue, our paper presents an intuitive and effective framework that explores semantic and contour cues of buildings and fully excavates their correlation. Specifically, we construct an interactive dual-stream decoder. The Intermediate connections within this decoder interactively transmit features between branches, contributing to learning correlations between semantics and contours. We propose the Semantic Collaboration Module (SCM) to strengthen the connection between the two branches. To further boost performance, we build the Multi-Scale Semantic Context Fusion Module (MSCF) to fuse semantic information from the higher and lower layers of the network, allowing the network to obtain superior feature representations. The experimental results on the WHU, INRIA, and Massachusetts building datasets demonstrate the superior performance of our method.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Xu Tang 0004, Puhua Chen, Huiyu Zhou 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 Semantic-Aware Context Modeling for Road Extraction in Remote Sensing Images
abstract
Road extraction faces the great challenges of occlusion, large span, and complex backgrounds in remote sensing images. Many existing methods receive context from regions near the road non-differently, and the context from irrelevant regions instead harms the semantics of features and leads to the mis-classification of the network. To address the above problem, we propose a Semantic-Aware Context Module (SACM) that encourages the network to model the context of different se-mantics supervised by a soft foreground map. And Strip Pooling Module (SPM) is introduced to match the fact that roads tend to be strip-shaped, contributing to the suppression of contamination information in irrelevant regions. Both SACM and SPM enable the network to obtain more specific semanti-cally relevant context. The experimental results on the Deep-Globe dataset show that the proposed method tremendously improves the performance of the network.
Xiangrong Zhang, Xiaoqian Zhu, Peng Zhu 0004, Xu Tang 0004, Licheng Jiao
IGARSS3
2022 Mask Decoupled Head for Instance Segmentation in Remote Sensing Images
abstract
Instance segmentation predicts the categories of all instances and locates them using pixel-level masks. Although existing methods have shown exemplary performance, the poor boundaries due to the lack of fine-grained information re-mains a challenge for RSIs instance segmentation. In this paper, to address the problem, we propose a novel instance segmentation branch, namely Mask Decoupled Head, which is mainly composed of a Feature Enhance Module (FEM) and a Feature Decoupled Module (FDM). FEM enhances the rep-resentation of the body features through the low-frequency component of images. FDM decouples the segmentation task by supervising body and edge separately and leverages fine-grained information to complement the boundary details. We performed comprehensive experiments on NWPU VHR -10 and HRSID datasets to evaluate the effectiveness of our pro-posed method and achieved good performance.
Xiangrong Zhang, Tianyang Zhang 0002, Xiaoqian Zhu, Xu Tang 0004, Licheng Jiao
IGARSS4
2022 A Transformer-Based Deep Learning Network for Underwater Acoustic Target Recognition
abstract
Underwater acoustic target recognition (UATR) is usually difficult due to the complex and multi-path underwater environment. Currently, deep learning-based UATR methods have proven its effectiveness, and have outperformed the traditional methods by using powerful convolution neural networks (CNNs) to extract discriminative features on acoustic spectrograms. However, CNNs always fail to capture the global information implicated in spectrogram due to the use of small kernel, and thus encounter the performance bottleneck. To this end, we propose the UATR-Transformer based on a convolution-free architecture, referred to as the Transformer, which can perceive both the global and local information from acoustic spectrograms, and thus improve the accuracy. Experiments on two real world data demonstrate that our proposed model has achieved comparative results to the state of art CNNs, and thus can be applied to some certain cases in UATR.
Sheng Feng, Xiaoqian Zhu
IEEE Geosci. Remote. Sens. Lett.2
2021 Parallel optimization of three-dimensional wedge-shaped underwater acoustic propagation based on MPI+OpenMP hybrid programming model
Yongxian Wang, Xiaoqian Zhu, Wei Liu 0140, Qiang Lan, Wenbin Xiao, Xinghua Cheng
J. Supercomput.3
2020 Discriminative Feature Pyramid Network For Object Detection In Remote Sensing Images
abstract
Multi-class geospatial object detection in remote sensing images suffer great challenges, such as large scales variability and complex background. Although feature pyramid network (FPN) can alleviate the problem of scale variation to some extent, it causes the loss of spatial and semantic information which is not conducive to object location. To address the above problem, this paper proposes a discriminative feature pyramid network (DFPN) by introducing a global guidance module (GGM) and a feature aggregation module (FAM). Specifically, the global guidance module delivers the high-level semantic information to lower layers, so as to obtain feature maps with stronger semantic information to eliminate the interference caused by complex background. The feature aggregation module enhances the interflow of information between different layers and better captures the discrimination information at each layer. We validate the effectiveness of our method on the NWPU VHR-10 and RSOD datasets, the results outperform baseline by 2.06 and 3.88 points respectively.
Xiaoqian Zhu, Xiangrong Zhang, Tianyang Zhang 0002, Peng Zhu 0004, Xu Tang 0004, Chen Li 0011
IJCNN1
2018 mSNP: A Massively Parallel Algorithm for Large-Scale SNP Detection
abstract
Single Nucleotide Polymorphism (SNP) detection is a fundamental procedure of whole genome analysis. SOAPsnp, a classic tool for detection, would take more than one week to analyze one typical human genome, which limits the efficiency of downstream analyses. In this paper, we present mSNP, an optimized version of SOAPsnp, which leverages Intel Xeon Phi coprocessors for large-scale SNP detection. Firstly, we redesigned the essential data structures of SOAPsnp, which significantly reduces memory footprint and improves computing efficiency. Then we developed a coordinated parallel framework for a higher hardware utilization of both CPU and Xeon Phi. Also, we tailored the data structures and operations to utilize the wide VPU of Xeon Phi to improve data throughput. Last but not the least, we proposed a read-based window division strategy to improve throughput and obtain better load balance. mSNP is the first SNP detection tool empowered by Xeon Phi. We achieved a 38x single thread speedup on CPU, without any loss in precision. Moreover, mSNP successfully scaled to 4,096 nodes on Tianhe-2. Our experiments demonstrate that mSNP is efficient and scalable for large-scale human genome SNP detection.
Yingbo Cui 0001, Shaoliang Peng, Yutong Lu, Xiaoqian Zhu, Bingqiang Wang, Chengkun Wu, Xiangke Liao
IEEE Trans. Parallel Distributed Syst.4
2017 Unified Access Layer with PostgreSQL FDW for Heterogeneous Databases
Ruohang Feng, Xiaoqian Zhu
NPC4
2015 A Method to Accelerate GROMACS in Offload Mode on Tianhe-2 Supercomputer
abstract
Molecular Dynamics(MD) is a computer simulation of physical movements of atoms and molecules in the context of N-body simulation, and is an important part of pharmaceutical industry. GROMACS, which is the most popular software for MD, could not perform satisfactorily with large-scale for the limit of computing resources. In this paper, we proposed a method to accelerate GROMACS with offload mode. In this mode, GROMACS could be arranged efficiently with CPU and the Intel® Xeon PhiTM Many Integrated Core (MIC) coprocessors at the same time, making the full use of Tianhe-2 supercomputer resources. To promote the efficiency of GROMACS, we proposed a series of methods, such as synchronization, data reassemble and array reuse. As we known, we are the first to accelerate GROMACS in offload mode on MIC.
Haiqiang Wang, Shaoliang Peng, Xiaoqian Zhu, Chengkun Wu, Weiliang Zhu, Jinan Wang, Huaiyu Yang
CCGRID3
2015 MICA: A fast short-read aligner that takes full advantage of Many Integrated Core Architecture (MIC)
abstract
BACKGROUND: Short-read aligners have recently gained a lot of speed by exploiting the massive parallelism of GPU. An uprising alterative to GPU is Intel MIC; supercomputers like Tianhe-2, currently top of TOP500, is built with 48,000 MIC boards to offer ~55 PFLOPS. The CPU-like architecture of MIC allows CPU-based software to be parallelized easily; however, the performance is often inferior to GPU counterparts as an MIC card contains only ~60 cores (while a GPU card typically has over a thousand cores). RESULTS: To better utilize MIC-enabled computers for NGS data analysis, we developed a new short-read aligner MICA that is optimized in view of MIC's limitation and the extra parallelism inside each MIC core. By utilizing the 512-bit vector units in the MIC and implementing a new seeding strategy, experiments on aligning 150 bp paired-end reads show that MICA using one MIC card is 4.9 times faster than BWA-MEM (using 6 cores of a top-end CPU), and slightly faster than SOAP3-dp (using a GPU). Furthermore, MICA's simplicity allows very efficient scale-up when multiple MIC cards are used in a node (3 cards give a 14.1-fold speedup over BWA-MEM). SUMMARY: MICA can be readily used by MIC-enabled supercomputers for production purpose. We have tested MICA on Tianhe-2 with 90 WGS samples (17.47 Tera-bases), which can be aligned in an hour using 400 nodes. MICA has impressive performance even though MIC is only in its initial stage of development. AVAILABILITY AND IMPLEMENTATION: MICA's source code is freely available at http://sourceforge.net/projects/mica-aligner under GPL v3. SUPPLEMENTARY INFORMATION: Supplementary information is available as "Additional File 1". Datasets are available at www.bio8.cs.hku.hk/dataset/mica.
Ruibang Luo, Jeanno Cheung, Edward Wu, Sze-Hang Chan, Wai-Chun Law, Guangzhu He, Chi-Man Liu, Dazong Zhou, Yingrui Li, Ruiqiang Li, Jun Wang 0004, Xiaoqian Zhu, Shaoliang Peng, Tak Wah Lam
BMC Bioinform.14
2014 Enabling and Scaling a Global Shallow-Water Atmospheric Model on Tianhe-2
abstract
This paper presents a hybrid algorithm for the petascale global simulation of atmospheric dynamics on Tianhe-2, the world's current top-ranked supercomputer developed by China's National University of Defense Technology (NUDT). Tianhe-2 is equipped with both Intel Xeon CPUs and Intel Xeon Phi accelerators. A key idea of the hybrid algorithm is to enable flexible domain partition between an arbitrary number of processors and accelerators, so as to achieve a balanced and efficient utilization of the entire system. We also present an asynchronous and concurrent data transfer scheme to reduce the communication overhead between CPU and accelerators. The acceleration of our global atmospheric model is conducted to improve the use of the Intel MIC architecture. For the single-node test on Tianhe-2 against two Intel Ivy Bridge CPUs (24 cores), we can achieve 2.07×, 3.18×, and 4.35× speedups when using one, two, and three Intel Xeon Phi accelerators respectively. The average performance gain from SIMD vectorization on the Intel Xeon Phi processors is around 5× (out of the 8× theoretical case). Based on successful computation-communication overlapping, large-scale tests indicate that a nearly ideal weak-scaling efficiency of 93.5% is obtained when we gradually increase the number of nodes from 6 to 8,664 (nearly 1.7 million cores). In the strong-scaling test, the parallel efficiency is about 77% when the number of nodes increases from 1,536 to 8,664 for a fixed 65,664 × 5,664 × 6 mesh with 77.6 billion unknowns.
Wei Xue 0003, Chao Yang 0002, Haohuan Fu, Yangtong Xu, Lin Gan 0001, Yutong Lu, Xiaoqian Zhu
IPDPS8
2013 Balancing accuracy, complexity and interpretability in consumer credit decision making: A C-TOPSIS classification approach
Xiaoqian Zhu, Jianping Li 0001, Dengsheng Wu, Changzhi Liang
Knowl. Based Syst.1