Xu Qiao

dblp:47/6244 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Systems, architecture and hardware · 5 · 4 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A CT-Based Non-invasive Diagnostic Model for the Grading of Esophageal Precancerous Lesions and Early Cancer
Jingxuan Sun, Yibin Jia, Xu Qiao
ICPR (9)5
2025 Design and Development of a GPR-Equipped Robot for Full-space External Diseases Detection in Drainage Pipelines*
abstract
Soil diseases around drainage pipelines are a major factor in road collapse. Robots designed to detect these diseases face multiple challenges, including harsh internal environments, size limitations, difficulties in achieving full external space coverage, and the impact of pose misalignment on disease localization. To address these challenges, this work presents the design and development of a pipeline robot equipped with Ground-Penetrating Radar (GPR), capable of adapting to a pipe diameter range of 500-1000 millimeters and providing comprehensive detection of external space diseases. A radial offset estimation model is introduced, and by integrating multi-sensor data, the robot achieves full-pose perception, overcoming challenges related to angular and positional misalignment during disease localization. Experimental results demonstrate that the robot can achieve a maximum detection speed of up to 0.5 meters per second and is capable of adapting to various field drainage pipeline scenarios, including full water, rough terrain, pose misalignment, and 90-degree bends. Azimuth errors for external disease localization are controlled within 1 degree, and axial displacement errors are controlled within 2 centimeters.
Yuanjin Fang, Xu Qiao, Maoxuan Xu
IROS3
2025 FlashAttest: Self-Attestation for Low-End Internet of Things via Flash Devices
abstract
Remote Attestation (RA) is an effective security service that allows a trusted party (verifier) to initiate the attestation routine on a potentially untrusted remote device (prover) to verify its correct state. Despite their usefulness, traditional challenge-response remote attestation protocols suffer from certain limitations, such as challenges in scaling attestation collection and the forced suspension of normal operation during attestation. Self-attestation tackles these issues by enabling the prover to measure its own state asynchronously with the verifier’s attestation request. Existing self-attestation methods rely on hybrid architectures to provide the required security properties, which may not be compatible with low-end Internet of Things (IoT) devices due to hardware limitations. In addition, these protocols currently lack formal verification of design correctness. In this paper, we present FlashAttest, a formally verified self-attestation protocol for low-end IoT devices. FlashAttest leverages the flash device to fulfill the security properties required by self-attestation, eliminating the requirement for hardware modifications. In particular, FlashAttest allows the prover to initiate the attestation routine and guarantee the trustworthiness of the results based on the verified software-based security architecture. By collaborating with the flash device during attestation to generate timestamped reports, FlashAttest enables the verifier to collect and verify the legitimacy of the attestation results. More importantly, FlashAttest achieves strong security guarantees supported by a formally verified design using the Tamarin prover. We implement and evaluate FlashAttest on MSP430 architecture, showing a reasonable overhead in terms of memory footprint, communication overhead, runtime and power consumption. Compared with state-of-the-art self-attestation schemes, our approach achieves similar runtime overhead, low energy consumption, and reasonable memory overhead while eliminating the need for hardware modifications. The results confirm the suitability of FlashAttest for low-end devices.
Zheng Zhang 0060, Jingfeng Xue, Weizhi Meng 0001, Xu Qiao, Yuanzhang Li 0001, Yu-an Tan 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Enhanced end-to-end regression algorithm for autonomous road damage detection
Hongjia Xing, Xu Qiao, Fanruo Li, Xinxin Huang
J. Supercomput.3
2024 Automatic Diagnosis Model of Gastrointestinal Diseases Based on Tongue Images
Baochen Fu, Miao Duan, Xiuli Zuo, Xu Qiao
ICPR (4)5
2024 I-ASM: Iterative Acoustic Scene Mapping for Enhanced Robot Auditory Perception in Complex Indoor Environments
abstract
This paper addresses the challenge of acoustic scene mapping (ASM) in complex indoor environments with multiple sound sources. Unlike existing methods that rely on prior data association or SLAM frameworks, we propose a novel particle filter-based iterative framework, termed I-ASM, for ASM using a mobile robot equipped with a microphone array and LiDAR. I-ASM harnesses an innovative "implicit association" to align sound sources with Direction of Arrival (DoA) observations without requiring explicit pairing, thereby streamlining the mapping process. Given inputs including an occupancy map, DoA estimates from various robot positions, and corresponding robot pose data, I-ASM performs multi-source mapping through an iterative cycle of "Filtering-Clustering-Implicit Associating". The proposed framework has been tested in real-world scenarios with up to 10 concurrent sound sources, demonstrating its robustness against missing and false DoA estimates while achieving high-quality ASM results. To benefit the community, we open-source all the codes and data at https://github.com/AISLAB-sustech/Acoustic-Scene-Mapping
Linya Fu, Yuanzheng He, Xu Qiao, He Kong 0001
IROS4
2024 UniverDetect: Universal landmark detection method for multidomain X-ray images
abstract
Landmark detection using X-ray imaging is vital for disease screening, treatment, and prognosis. It provides a framework for subsequent tasks, including segmentation, classification, and target detection. Existing landmark detection methods exhibit limits with the increasing number of X-ray examiners. Not only they serve specific data types, but also their ability to learn the complete semantic information from this dataset is limited, potentially affecting their widespread adoption. We address these challenges by proposing a universal landmark detection model for multidomain X-ray imaging, UniverDetect. By progressively acquiring local semantic information, global semantic insights, and anatomical structure knowledge at various levels, the model ensured that the detected landmarks were closely aligned with the ground-truth labels. UniverDetect consists of three key components. The landmark detection module (LDM) utilizes a U-net network equipped with our innovative pyramid depthwise separable (PDS) convolution module for initial landmark detection. The landmark refinement module (LRM) integrates a fine-tuning module comprising continuously extended convolution blocks. Finally, the landmark correction module (LCM) incorporates a graph convolutional network (GCN) to rectify offset errors during partial landmark detection. An inherent feature of this model is its domain- generalization capability, which enables continuous learning across diverse domains. This model can concurrently learn from eight domains, covering 118 landmarks within a diverse dataset of 5969 images. Extensive experiments on multiple datasets demonstrate that this method consistently outperforms state-of-the-art approaches. This study presents a versatile and effective solution to reduce doctors’ workload while providing precise quantitative analysis.
Chenyang Lu 0009, Guangtong Yang, Xu Qiao, Wei Chen 0039, Qingyun Zeng
Neurocomputing3
2024 Achievable Rate Analysis and Power Optimization for Cell-Free Massive MIMO URLLC Systems Over Aging and Correlated Channels
abstract
In this paper, we consider the cell-free massive multiple-input multiple-output (MIMO) system for supporting ultra-reliable and low-latency communication (URLLC) transmission, where a large number of access points (APs) serve a small number of users in the short packet regime. Assuming channel aging and channel spatial correlation, we derive the closed-form expression of the downlink achievable rate with the normalized conjugate beamforming (NCB). Under the goal of maximizing the minimum user rate, we formulate a max-min power optimization problem with a power constraint at each AP. However, it is challenging to solve this problem because the objective function is a complicated function of power coefficients. To tackle this difficulty, we use a path-following method to approximate the objective function to a logarithmic function and transform the polynomial constraint into a monomial. Thus, we can iteratively solve the original problem by reformulating it as a series of geometric programming problems. Numerical results verify the tightness of the closed-form expression for the downlink achievable rate in the short packet regime. Both channel aging and channel spatial correlation significantly degrade the system performance of CF massive MIMO URLLC systems. Moreover, Using NCB and the proposed max-min power allocation can effectively alleviate this impairment and improve the system performance.
Han Hu 0006, Yao Zhang 0016, Xu Qiao, Longxiang Yang, Hongbo Zhu 0002
IEEE Internet Things J.4
2024 Comparing CNN-based and transformer-based models for identifying lung cancer: which is more effective?
abstract
Abstract Lung cancer constitutes the most severe cause of cancer-related mortality. Recent evidence supports that early detection by means of computed tomography (CT) scans significantly reduces mortality rates. Given the remarkable progress of Vision Transformers (ViTs) in the field of computer vision, we have delved into comparing the performance of ViTs versus Convolutional Neural Networks (CNNs) for the automatic identification of lung cancer based on a dataset of 212 medical images. Importantly, neither ViTs nor CNNs require lung nodule annotations to predict the occurrence of cancer. To address the dataset limitations, we have trained both ViTs and CNNs with three advanced techniques: transfer learning, self-supervised learning, and sharpness-aware minimizer. Remarkably, we have found that CNNs achieve highly accurate prediction of a patient’s cancer status, with an outstanding recall (93.4%) and area under the Receiver Operating Characteristic curve (AUC) of 98.1%, when trained with self-supervised learning. Our study demonstrates that both CNNs and ViTs exhibit substantial potential with the three strategies. However, CNNs are more effective than ViTs with the insufficient quantities of dataset.
Lulu Gai, Mengmeng Xing, Xu Qiao
Multim. Tools Appl.5
2024 A Ground-Penetrating Radar Clutter Suppression Algorithm Integrating Signal Processing and Image Fusion
abstract
This study presents a clutter suppression algorithm for underground mine ground-penetrating radar (GPR) B-scan profiles, integrating the sparrow search algorithm (SSA), variational mode decomposition (VMD), and image fusion via the improved nonsubsampled shearlet transform (INSST), aimed at addressing clutter under strict explosion-proof conditions. Initially, the raw B-scan data undergoes preprocessing to derive a grayscale profile. Using SSA, an envelope entropy fitness function is introduced to automatically optimize VMD with respect to the average trajectory of the grayscale profile. Finally, enhanced non-subsampled shearlet transform (NSST) processing is applied to the synthesized profile using a scale-direction adaptive threshold with an L2 norm square constraint and a flexible threshold function. This approach reduces Gibbs artifacts and oversmoothing while retaining valid signals within the threshold range. In simulated profiles, the proposed algorithm demonstrates strong robustness across different noise standard deviations, significantly improving image quality metrics such as peak signal-to-noise ratio (PSNR) and structure similarity index measure (SSIM), while reducing the root-mean-square error (RMSE). Notably, in measured mining profiles, the proposed algorithm outperforms eight other methods by better preserving edge details. Compared to the suboptimal algorithm, the target-to-noise ratio (TNR) improves by 36% in mining fault profile I and by 17% in coal seam floor profile II. Finally, a sensitivity analysis of the proposed algorithm’s parameters was conducted.
Xu Qiao, Jia-Lin Liu, Fanruo Li, Yong-Liang Wen, Zhi-Hai Fan, Zhen-Hong Qi, Zhi-Hua Yang
IEEE Trans. Geosci. Remote. Sens.3
2023 A Stealth Security Hardening Method Based on SSD Firmware Function Extension
Xiao Yu 0005, Xu Qiao, Yu-an Tan 0001, Yuanzhang Li 0001, Li Zhang 0099
ICONIP (9)3
2023 SLAM-Based Joint Calibration of Differential RSS Sensor Array and Source Localization
abstract
Sensor arrays generating differential received signal strength (DRSS) measurements have found many applications in robotics. However, accurate calibration of these sensor arrays remains a challenge. Most existing methods are impractical in that they assume to know signal source positions or certain parameters (i.e., path loss exponent), and try to estimate the others. In this paper, we adopt graph simultaneous localization and mapping (SLAM) as a general framework for jointly estimating the source positions and parameters of the DRSS sensor array. Our contributions are twofold. On the one hand, by using a Fisher information matrix approach, we conduct a systematic observability analysis of the corresponding SLAM setup for the calibration problem. On the other hand, we propose an effective procedure to select the initial value which is fed to Levenberg-Marquardt iterations for further improving optimization accuracy and convergence. Extensive simulation and hardware experiments show that the proposed method renders high-quality calibration results. All the codes and data are publicly available at https://github.com/SUSTech2022/DRSS-sensor-array-calibration.
Linya Fu, Xu Qiao, Shoudong Huang, Guoqiang Mao, Zhiyun Lin, Youfu Li 0001, He Kong 0001
IECON2
2022 Using Vision Transformers in 3-D Medical Image Classifications
abstract
Convolutional Neural Networks (CNNs) have been the controlling deep learning approach for a decade in automated medical image diagnosis. Recently, vision transformers (ViTs) have appeared as a competitive alternative to CNNs in computer vision, yielding similar levels of performance while possessing several interesting properties that could prove to be beneficial for the explanation of deep neural networks. Since most medical images are grayscale scans of CT, MRI, etc. and in 3-dimensional (3-D) spaces, which are highly different from natural images, we explore whether it is possible to move to transformer-based models or if we should keep working with CNNs in the domain of 3-D medical image classifi-cations. If so, what are the advantages and drawbacks of switching to ViTs for medical image diagnosis? We consider these problems in a series of experiments on three 3-D medical image datasets. Our findings show that, while CNNs perform better when trained from scratch, ViTs gain strong benifit when pre-trained on ImageNet and outperform their CNN counterparts using self-supervised learning and sharpness-aware minimizer optimization method on the large datasets.
Lulu Gai, Xu Qiao
ICIP5
2022 A combinatorial precoding scheme of cell-free massive MIMO with channel aging
abstract
Abstract Here, we investigate the precoding schemes for the downlink data transmission of a time‐division duplex cell‐free massive multiple‐input multiple‐output (MIMO) system with channel aging, which arises from the user mobility. Closed‐form spectral efficiency (SE) expressions of the downlink with the normalized conjugate beamforming (NCB), and the full‐pilot zero‐forcing (FZF), are derived, which are used for the analytical system performance evaluation. Then, a novel combinatorial precoding scheme with enhanced system SE performance, which adopts either NCB or FZF according to each user channel aging condition, is proposed. Moreover, a pilot allocation strategy is proposed to alleviate the extra interference brought by the combinatorial precoding scheme. Also, a statistical channel cooperative power control is employed to further improve the performance for all the above precoding schemes. Numerical results show that the proposed precoding scheme can substantially improve the average downlink SE.
Han Hu 0006, Longxiang Yang, Yao Zhang 0016, Xu Qiao
IET Commun.5
2021 A Teacher-Student Learning Based On Composed Ground-Truth Images For Accurate Cephalometric Landmark Detection
abstract
Computer-aided automatic cephalometric landmark localization has been a hot topic since last century. Recent proposed deep learning-based methods have made great contributions to this research topic. Among them, convolutional neural networks (CNN)-based regression is widely used, where ground-truth (GT) information is mainly used in the calculation of loss function, thus, mimics the difference between the predicted landmarks ' locations and the ground-truth locations through backpropagation. However, considering the limited number of annotated cephalometric data, we believe the performance can be better improved by better utilizing ground-truth information. In this paper, we propose a teacher-student learning method using GT images for accurate cephalometric detection. We first use images composed with GT landmarks as input images to train a detection model, which is treated as a teacher model. Then the teacher model is used to guide a student model, which is trained by original images, by transferring useful features. We believe the features between GT images and original images have similar domain distribution since they both represent same structure. We validate our method on public grand challenge dataset. Our method achieves better performance compared with state-of-the-art methods.
Yu Song 0008, Xu Qiao, Yutaro Iwamoto, Yen-Wei Chen 0001
ICIP2
2021 MAU-Net: Multiple Attention 3D U-Net for Lung Cancer Segmentation on CT Images
abstract
Accurate segmentation of lung cancer from computed tomography (CT) is of great significance to constructing an automatic diagnosis system for lung cancer. This paper presents multiple attention 3D U-Net (MAU-Net), a novel deep learning-based architecture for lung cancer segmentation from CT images. In particular, we first apply a dual attention module at the bottleneck of the U-Net that models the semantic interdependencies in spatial and channel dimensions, respectively. A novel multiple attention gate module is then proposed to adaptively recalibrate and fuse multiscale features from the dual attention module, the previous decoder feature maps, and the corresponding features from the encoder. Extensive ablation studies on a clinical dataset consisting of 322 CT images demonstrate the effectiveness of our proposed method. Our model achieved an average Dice similarity coefficient, 95% Hausdorff distance and relative absolute volume difference of 0.8667, 13.0036, and 0.1552, respectively.
Fengchang Yang, Xianru Zhang, Xu Qiao
KES5
2020 Cell-Free Massive MIMO with Few-bit ADCs/DACs: AQNM versus Bussgang
abstract
In this paper, we consider a downlink cell-free massive multi-input multi-output (mMIMO) system, assuming few-bit analog-digital converters (ADCs) and digital-analog converters (DACs) are implemented at the access points (APs). Leveraging on the linear additive quantization noise model (AQNM), we derive a tight approximate rate expression, which provides insights into the impacts of the imperfect quantization error and channel estimation error. Thanks to the trackable result, we quantitatively compare the performance differences between the two quantization models, namely the AQNM and the Bussgang theorem. In particular, the AQNM can offer analytical tractability for few-bit quantization while the Bussgang theorem only characterizes 1-bit quantization since the multi-bit quantization under the Bussgang theorem is difficult to deal with. Simulation results show that under the same 1-bit quantization, the rate performance with the Bussgang theorem is roughly identical to the case of the AQNM.
Yao Zhang 0016, Haotong Cao, Xu Qiao, Shengchen Wu, Longxiang Yang
VTC Spring4
2020 Cloud-based difference algorithm using big GPR data for roadbed damage detection
abstract
Summary Ground‐penetrating radar (GPR) technology is being widely used for urban roadbed damage detection. In this paper, we present the development of a cloud‐based difference detection method using big GPR data for roadbed damage detection. Unlike many processing methods relying on standalone servers, the proposed method can achieve high accuracy and processing efficiency by distributing the computation over a server cluster. The method includes a difference detection algorithm for GPR image enhancement and damage extraction, the Hadoop Distributed File System, and MapReduce distributed parallel computing framework for big GPR data parallel processing and interpretation. Moreover, to automatize roadbed damage classification, a criterion with seven grades is established based on the wave group shape, amplitude, phase, and attenuation characteristics. We validate the proposed method through a detection experiment that was conducted on the Fourth Ring Road in Beijing, China. The detection revealed 187 damage from the efficient processing of big GPR data, thus suggesting the ability of the proposed method to provide timely support for road safety in cities.
Xianlei Xu, Wenru Gao, Taotao Li, Xu Qiao
Concurr. Comput. Pract. Exp.5
2019 Rate Analysis of Cell-Free Massive MIMO with One-Bit ADCs and DACs
abstract
We investigate the downlink rate performance of cell-free massive multiple-input multiple-output (mMIMO) network with conjugate beamforming precoder when the access points (APs) are equipped with one-bit analog-digital converters (ADCs) and digital-analog converters (DACs). Based on Buss-gang decomposition theory, we derive a rigorous closed-form rate expression, which covers the impact of multi-antenna APs, the imperfect quantization error, and channel estimation error. Then, by exploring this closed-form result, we show that the quantization interferences resulted from one-bit quantization can significantly decrease the downlink rate performance. In addition, we also analyze the performance gain provided by adding the total number of antennas or increasing the total transmitted power. We observe that, adding the total number of antenna arrays is a promising way to compensate for the quantization losses. However, these losses cannot be compensated by infinitely increasing the total transmitted power.
Yao Zhang 0016, Haotong Cao, Xu Qiao, Longxiang Yang
PIMRC4
2017 Multi-dimensional data representation using linear tensor coding
abstract
Linear coding is widely used to concisely represent data sets by discovering basis functions of capturing high‐level features. However, the efficient identification of linear codes for representing multi‐dimensional data remains very challenging. In this study, the authors address the problem by proposing a linear tensor coding algorithm to represent multi‐dimensional data succinctly via a linear combination of tensor‐formed bases without data expansion. Motivated by the amalgamation of linear image coding and multi‐linear algebra, each basis function in the authors’ algorithm captures some specific variabilities. The basis‐associated coefficients can be used for data representation, compression and classification. When the authors apply the algorithm on both simulated phantom data and real facial data, the experimental results demonstrate their algorithm not only preserves the original information of input data, but also produces localised bases with concrete physical meanings.
Xu Qiao, Yen-Wei Chen 0001, Zhi-Ping Liu
IET Image Process.1
2012 Group sparse representation of adaptive sub-domain selection for image classification
Xianhua Han, Xu Qiao, Yen-Wei Chen 0001
ICPR2
2010 Statistical Texture Modeling for Medical Volume Using Generalized N-Dimensional Principal Component Analysis Method and 3D Volume Morphing
abstract
In this paper, a statistical texture modeling method is proposed for medical volumes. As the shapes of the human organ are very different from one case to another, 3D volume morphing is applied to normalize all the volume datasets to a same shape for removing shape variations. In order to deal with the problems of high-dimension and small number of medial samples, we propose an effective image compression method named Generalized N-dimensional Principal Component Analysis (GND-PCA) to construct a statistical model. Experiments applied on liver volumes show good performance on generalization using our method. A simple experiment is employed to show that the features extracted by the statistical texture model have capability of discrimination for different types of data, such as normal and abnormal.
Xu Qiao, Yen-Wei Chen 0001
ICPR1
2010 Tensor-based subspace learning and its applications in multi-pose face synthesis
Xu Qiao, Xianhua Han, Takanori Igarashi, Keisuke Nakao, Yen-Wei Chen 0001
Neurocomputing1