Xueshuang Xiang

dblp:135/1100 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0001-7794-4876ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BSDM: Background Suppression Diffusion Model for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection (HAD) is widely used in Earth observation and deep space exploration. A major challenge for HAD is the complex background of the input hyperspectral images (HSIs), resulting in anomalies confused in the background. On the other hand, most existing HAD methods require training a separate model for each HSI, resulting in poor generalization in practical applications. This paper starts the first attempt to study a new and generalizable background learning problem without labeled samples. We present a novel solution BSDM (background suppression diffusion model) for HAD, which can simultaneously learn latent background distributions and generalize to different datasets for suppressing complex background. It is featured in three aspects: (1) For the complex background of HSIs, we design pseudo-background noise and learn the potential background distribution in it with a diffusion model (DM). (2) For the generalizability problem, we apply a statistical offset module so that the BSDM adapts to datasets of different domains without labeling samples. (3) For achieving background suppression, we innovatively improve the inference process of DM by feeding the original HSIs into the denoising network, which removes the background as noise. Our work paves a new background suppression way for HAD that can improve HAD performance without the prerequisite of manually labeled data. Assessments and generalization experiments of four HAD methods on several real HSI datasets demonstrate the above three unique properties of the proposed method. Our project is available at https://github.com/majitao-xd/BSDM-HAD.
Jitao Ma, Weiying Xie, Xueshuang Xiang, Yunsong Li 0001, Leyuan Fang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Dual-Dynamic Cross-Modal Interaction Network for Multimodal Remote Sensing Object Detection
abstract
Multimodal remote sensing object detection (MM-RSOD) holds great promise for around-the-clock applications. However, it faces challenges in effectively extracting complementary features due to the modality inconsistency and redundancy. Inconsistency can lead to semantic-spatial misalignment, while redundancy introduces uncertainty that is specific to each modality. To overcome these challenges and enhance complementarity exploration and exploitation, this article proposes a dual-dynamic cross-modal interaction network (DDCINet), a novel framework comprising two key modules: a dual-dynamic cross-modal interaction (DDCI) module and a dynamic feature fusion (DFF) module. The DDCI module simultaneously addresses both modality inconsistency and redundancy by employing a collaborative design of channel-gated spatial cross-attention (CSCA) and cross-modal dynamic filters (CMDFs) on evenly segmented multimodal features. The CSCA component enhances the semantic-spatial correlation between modalities by identifying the most relevant channel-spatial features through cross-attention, addressing modality inconsistency. In parallel, the CMDF component achieves cross-modal context interaction through static convolution and further generates dynamic spatial-variant kernels to filter out irrelevant information between modalities, addressing modality redundancy. Following the improved feature extraction, the DFF module dynamically adjusts interchannel dependencies guided by modal-specific global context to fuse features, achieving better complementarity exploitation. Extensive experiments conducted on three MM-RSOD datasets confirm the superiority and generalizability of the DDCINet framework. Notably, our DDCINet, based on the RoI Transformer benchmark and ResNet50 backbone, achieves 78.4% mAP50 on the DroneVehicle test set and outperforms state-of-the-art (SOTA) methods by large margins.
Meiyu Huang, Xueshuang Xiang
IEEE Trans. Geosci. Remote. Sens.4
2024 Incremental Learning Strategy with Multi-dimensional Knowledge Distillation for One-stage Object Detection
abstract
In practical applications, object detectors often encounter unknown classes samples. As object detection datasets undergo continuous development, the need arises for object detectors to adeptly recognize an escalating array of new classes while retaining the original detection capability. This paper introduces a new class incremental object detection learning framework with multi-dimensional knowledge distillation, using Yolov5 as the foundational detection model. Specifically, we utilize local feature distillation, attention distillation and global feature distillation to preserve information in the original model's feature maps. Additionally, we employ response distillation to ensure the model remains responsive to previous classes. In comparison to the state-of-the-art methods, our approach achieves higher overall accuracy on PASCAL VOC 2007. Especially in the “19+1” incremental task, our approach gets the highest [email protected] of 0.697 with a minimum mean AP decline rate of 4.45%. For the “4+3” and “6+1” tasks on KITTI, our approach preserves the model's responsiveness to the original task more effectively while achieving the highest overall accuracy.
Xuejiao Liu 0001, Xueshuang Xiang, Yu-an Tan 0001, Weizhi Meng 0001
INDIN3
2024 Confidence-Weighted Teacher: Semi-Supervised Object Detection Based on Confidence Correction
Xuejiao Liu 0001, Xueshuang Xiang
PRCV (13)3
2024 Attention-based network for passive non-light-of-sight reconstruction in complex scenes
Meiyu Huang, Yunqing Huang, Xueshuang Xiang
Vis. Comput.6
2023 Cross-Modal Attentive Recalibration and Dynamic Fusion for Multispectral Pedestrian Detection
Meiyu Huang, Xueshuang Xiang
PRCV (1)4
2023 FASONet: A Feature Alignment-Based SAR and Optical Image Fusion Network for Land Use Classification
Feng Deng, Meiyu Huang, Xueshuang Xiang
PRCV (10)5
2023 Union label smoothing adversarial training: Recognize small perturbation attacks and reject larger perturbation attacks balanced
abstract
Recently, several adversarial training methods have been proposed for rejecting perturbation-based adversarial examples, which enhance the robustness of deep neural networks to larger perturbations. However, they often perform unsatisfactorily when dealing with examples with lower level perturbations. To address these issues, we introduce a novel adversarial training approach called the union label smoothing adversarial training (ULSAT), which employs a new label smoothing curve and a union strategy for adversarial training. The label smoothing curve assigns soft labels to perturbed examples, allowing for a more reasonable adjustment of calibration. The union strategy adds interpolation examples and combines adversarial examples generated under different perturbation tolerances into the training stage, which improves the rejection ability of the model and balances it with the classification ability. Through theoretical analysis and ablation study, we demonstrate the effectiveness of our proposed approach. Numerical experiments show that ULSAT can accurately classifies less disturbed examples while maintaining a good rejection ability for adversarial examples with higher levels of perturbation. Moreover, we introduce an evaluation index that comprehensively considers the classification ability and rejection ability of the model. Under this index, ULSAT achieves state-of-the-art results.
Jinshu Huang, Haidong Xie, Xueshuang Xiang
Future Gener. Comput. Syst.4
2022 LDGAN: Latent Determined Ensemble Helps Removing IID Data Assumption and Cross-node Sampling in Distributed GANs
abstract
Generative Adversarial Networks (GANs) have received a lot of attention due to their powerful generative ability, and many related studies have been carried out. Among them, deploying GANs in distributed scenarios has become a hot research topic due to the sharp increase in the capacity of training data. However, distributed GANs face many challenges, such as different data distribution of each node, limitation of cross-node data sampling, etc. The previous methods defaulted to some strong assumptions to alleviate these problems, but the performance would be degraded once they encountered a natural scene. In this paper, we propose Latent Determined Generative Adversarial Network (LDGAN), a network introducing latent determined ensemble to guide the training of the generator without specific assumptions. LDGAN measures the difference in the latent space of data distribution at each node and uses this as the weight of the feedback information of each discriminator, and there is no need for any cross-node data sampling and the independent and identical distribution (iid) data assumption. Our experiments show that on the MNIST, Fashion-MNIST, and CIFAR-10 datasets, the images generated by LDGAN are more realistic and diverse, and the achieved Fréchet Inception Distance (FID) is 15.0%, 22.8%, and 15.3% smaller than the state-of-the-art models, respectively.
Wei Wang 0250, Ziwen Wu, Xueshuang Xiang, Yue Li 0013
ICPR3
2022 Attention-Guided Multi-modal and Multi-scale Fusion for Multispectral Pedestrian Detection
Meiyu Huang, Xueshuang Xiang
PRCV (1)4
2022 AviPer: assisting visually impaired people to perceive the world with visual-tactile multimodal attention network
Xinrong Li, Meiyu Huang, Yingze Cao, Yamei Lu, Xueshuang Xiang
CCF Trans. Pervasive Comput. Interact.7
2021 Towards GANs' Approximation Ability
abstract
Generative adversarial networks (GANs) have attracted in-tense interest in the field of generative models. This paper will first theoretically analyze GANs’ approximation property. Similar to the universal approximation property of the fully connected neural networks with one hidden layer, we prove that the generator with the input latent variable in GANs can universally approximate the potential data distribution given the increasing hidden neurons. Furthermore, we propose an approach named stochastic data generation (SDG) to enhance GANs’ approximation ability. Our approach is based on the simple idea of imposing randomness through data generation in GANs by a prior distribution on the conditional probability between the layers. The experimental results on synthetic dataset verify the improved approximation ability obtained by this SDG approach. In the practical dataset, three GANs using SDG can also outperform the corresponding traditional GANs when the model architectures are smaller.
Xuejiao Liu 0001, Xueshuang Xiang
ICME3
2021 Blind Adversarial Pruning: Towards The Comprehensive Robust Models With Gradually Pruning Against Blind Adversarial Attacks
abstract
With the growth of interest in the attack and defense of deep neural networks, researchers have increasingly focused on the robustness of their application to devices with limited memory, in order to deal with unknown-budget (blind) adversarial attacks under different compression ratios. We analyze the existing pruning methods and find that the robustness of the pruned models varies drastically with different pruning processes, and the robustness of the pruned model with adversarial training exhibits a high sensitivity to the budget of the adversarial examples. These methods cannot obtain models that are comprehensively robust when confronting blind adversarial attacks with different compression ratios. To address this problem, we propose an approach called blind adversarial pruning (BAP) that introduces the approach of blind adversarial training into the gradual pruning process, to ultimately obtain pruned models with comprehensive robustness under different compression ratios. The experimental results obtained using BAP for pruning classification models based on several benchmarks demonstrate the competitive performance of this method; the robustness of BAP models is more stable compared to various pruning processes, and BAP exhibits better comprehensive robustness against blind adversarial attacks.
Haidong Xie, Lixin Qian, Xueshuang Xiang, Naijin Liu
ICME3
2021 Legitimate Adversarial Patches: Evading Human Eyes and Detection Models in the Physical World
abstract
It is known that deep neural models are vulnerable to adversarial attacks. Digital attacks can craft imperceptible perturbations but lack of the ability to apply in physical environment. To address this issue, efforts have been investigated to study physical patch attacks in the physical world, especially for object detection models. Previous works mostly focus on evading the detection model itself but ignore the impact of human observers. In this paper, we study legitimate adversarial attacks that evade both human eyes and detection models in the physical world. To this end, we delve into the issue of patch rationality, and propose some indicators for evaluating the rationality of physical adversarial patches. Besides, we propose a novel framework with a two-stage training strategy to generate our legitimate adversarial patches (LAPs). Both in numerical simulations and physical experiments our LAPs have significant attack effects and visual rationality.
Jia Tan, Haidong Xie, Xueshuang Xiang
ACM Multimedia4
2019 Task-Driven Common Representation Learning via Bridge Neural Network
abstract
This paper introduces a novel deep learning based method, named bridge neural network (BNN) to dig the potential relationship between two given data sources task by task. The proposed approach employs two convolutional neural networks that project the two data sources into a feature space to learn the desired common representation required by the specific task. The training objective with artificial negative samples is introduced with the ability of mini-batch training and it’s asymptotically equivalent to maximizing the total correlation of the two data sources, which is verified by the theoretical analysis. The experiments on the tasks, including pair matching, canonical correlation analysis, transfer learning, and reconstruction demonstrate the state-of-the-art performance of BNN, which may provide new insights into the aspect of common representation learning.
Xueshuang Xiang, Meiyu Huang
AAAI2
2019 A Lightweight Neural Network Based Human Depth Recovery Method
abstract
Human depth recovery is an essential task for tele-immersive video interaction systems with limited bandwidth consumption. This paper is motivated by the first human depth recovery method, i.e., the weighted large margin nearest center (WLMNC) distance based method, which compresses each human depth map into several skeletal block structures by learning a WLMNC distance at the remote stage and then recovers it by a rough-to-fine approach at the local stage. Since the WLMNC distance is equivalent to employing a linear transformation on the human pixels, for human postures with complex self-occlusion, the depth recovery performance of WLMNC is limited. To address the problem, this paper proposes to learn a nonlinear WLMNC distance via a lightweight neural network. Different from traditional classification or clustering problems, the neural network is introduced for storing skeletal block structure information instead of predicting new information. The quantitative experimental results on the VGA-sized benchmark dataset demonstrates that such a slight modification over WLMNC can achieve an inspiring performance improvement: the average depth recovery MAD error is reduced from 3.56cm to 2.77cm and the total running cost is reduced from about 150 seconds to 43 seconds on a laptop with just a little loss on the sampling rate, from 64x to 51x.
Meiyu Huang, Xueshuang Xiang, Yiqiang Chen 0001
ICME2
2018 Weighted Large Margin Nearest Center Distance-Based Human Depth Recovery With Limited Bandwidth Consumption
abstract
This paper proposes a weighted large margin nearest center (WLMNC) distance-based human depth recovery method for tele-immersive video interaction systems with limited bandwidth consumption. In the remote stage, the proposed method highly compresses the depth data of the remote human into skeletal block structures by learning the WLMNC distance, which is equivalent to downsampling the human depth map at $64{\times}$ the sampling rate. In the local stage, the method first recovers a rough human depth map based on a WLMNC distance augmented clustering approach and then obtains a fine depth map based on a rough depth-guided autoregressive model to preserve the depth discontinuities and suppress texture copy artifacts. The proposed WLMNC distance is learned by the large margin clustering problem with a weighted hinge loss to balance the clustering accuracy and depth recovery accuracy and is verified to be able to preserve depth discontinuities between skeletal block structures with occlusion. A theoretical analysis is conducted to verify the effectiveness of using the weighted hinge loss. Furthermore, a novel data set containing various types of human postures with self-occlusion is built to benchmark the human depth recovery methods. The quantitative comparison with the state-of-the-art depth recovery methods on the introduced benchmark data set demonstrates the effectiveness of the proposed method for human depth recovery with such a high upsampling rate.
Meiyu Huang, Xueshuang Xiang, Yiqiang Chen 0001, Da Fan
IEEE Trans. Image Process.2