VLDB 2026 Research / reviewers in the wild / expert
Xiaowei Fu
dblp:49/11478
· DBLP profile ↗
24ranked-venue papers
10as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Research on collaborative path planning of UAV swarms for urban logistics distribution in dense building environments
Yuanda Lai, Husheng Wu, Yuanqing Xia, Xiaowei Fu, Dongli Duan, Aiai Wang, Meimei Shi |
Expert Syst. Appl. | 4 |
| 2026 | Unsupervised Robust Domain Adaptation: Paradigm, Theory and Algorithm
Fuxiang Huang, Xiaowei Fu, Shiyu Ye, Wen Li 0001, Xinbo Gao 0001, David Zhang 0001, Lei Zhang 0038 |
Int. J. Comput. Vis. | 2 |
| 2026 | M3C: Resist Agnostic Attacks by Mitigating Consistent Class Confusion PriorabstractAdversarial attack is a major obstacle to the deployment of deep neural networks (DNNs) for security-sensitive applications. To address these adversarial perturbations, various adversarial defense strategies have been developed, with Adversarial Training (AT) being one of the most effective methods to protect neural networks from adversarial attacks. However, existing AT methods struggle against training-agnostic attacks due to their limited generalizability. This suggests that the AT models lack a unified perspective for various attacks to conduct universal defense. This paper sheds light on a generalizable prior under various attacks: consistent class confusion (3C), i.e., an AT classifier often confuses the predictions between correct and ambiguous classes in a highly similar pattern among diverse attacks. Relying on this latent prior as a bridge between seen and agnostic attacks, we propose a more generalized AT model by mitigating consistent class confusion (M3C) to resist training-agnostic attacks. Specifically, we optimize an Adversarial Confusion Loss (ACL), which is weighted by uncertainty, to distinguish the most confused classes and encourage the AT model to focus on these confused samples. To suppress malignant features affecting correct predictions and producing significant class confusion, we propose a Gradient-Aware Attention (GAA) mechanism to enhance the classification confidence of correct classes and eliminate class confusion. Experiments on multiple benchmarks and network frameworks demonstrate that our M3C model significantly improves the generalization of AT robustness against agnostic attacks. The finding of the 3C prior reveals the potential and possibility for defending against a wide range of attacks, and provides a new perspective to overcome such challenge in this field. Xiaowei Fu, Fuxiang Huang, Guoyin Wang 0001, Xinbo Gao 0001, Lei Zhang 0038 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Token Calibration for Transformer-Based Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain by learning domain-invariant representations. Motivated by the recent success of Vision Transformers (ViTs), several UDA approaches have adopted ViT architectures to exploit fine-grained patch-level representations, which are unified as Transformer-based $D$ omain $A$ daptation (TransDA) independent of CNN-based. However, we have a key observation in TransDA: due to inherent domain shifts, patches (tokens) from different semantic categories across domains may exhibit abnormally high similarities, which can mislead the self-attention mechanism and degrade adaptation performance. To solve that, we propose a novel $P$ atch- $A$ daptation Transformer (PATrans), which first identifies similarity-anomalous patches and then adaptively suppresses their negative impact to domain alignment, i.e. token calibration. Specifically, we introduce a $P$ atch- $A$ daptation $A$ ttention (PAA) mechanism to replace the standard self-attention mechanism, which consists of a weight-shared triple-branch mixed attention mechanism and a patch-level domain discriminator. The mixed attention integrates self-attention and cross-attention to enhance intra-domain feature modeling and inter-domain similarity estimation. Meanwhile, the patch-level domain discriminator quantifies the anomaly probability of each patch, enabling dynamic reweighting to mitigate the impact of unreliable patch correspondences. Furthermore, we introduce a contrastive attention regularization strategy, which leverages category-level information in a contrastive learning framework to promote class-consistent attention distributions. Extensive experiments on four benchmark datasets demonstrate that PATrans attains significant improvements over existing state-of-the-art UDA methods (e.g., 89.2% on the VisDA-2017). Code is available at: https://github.com/YSY145/PATrans. Xiaowei Fu, Shiyu Ye, Chenxu Zhang 0001, Fuxiang Huang, Xin Xu 0001, Lei Zhang 0038 |
IEEE Trans. Image Process. | 1 |
| 2026 | Multi-population quantum firefly algorithm via ergodic correction mechanism for continuous optimization problems
Xiaowei Fu |
J. Supercomput. | 4 |
| 2026 | Rectifying Adversarial Sample With Low Entropy Prior for Test-Time DefenseabstractExisting defense methods fail to defend against un known attacks and thus raise generalization issue of adversarial robustness. To remedy this problem, we attempt to delve into some underlying common characteristics among various attacks for generality. In this work, we reveal the commonly overlooked low entropy prior (LE) implied in various adversarial samples, and shed light on the universal robustness against unseen attacks in inference phase. LE prior is elaborated as two properties across various attacks as shown in Fig. 1 and 2: 1) low entropy misclassification for adversarial samples and 2) lower entropy prediction for higher attack intensity. This phenomenon stands in stark contrast to the naturally distributed samples. The LE prior can instruct existing test-time defense methods, thus we propose a two-stage REAL approach: Rectify Adversarial sample based on LE prior for test-time adversarial rectification. Specifically, to align adversarial samples more closely with clean samples, we propose to first rectify adversarial samples misclassified with low entropy by reverse maximizing prediction entropy, thereby eliminating their adversarial nature. To ensure the rectified samples can be correctly classified with low entropy, we carry out secondary rectification by forward minimizing prediction entropy, thus creating a Max-Min entropy optimization scheme. Further, based on the second property, we propose an attack aware weighting mechanism to adaptively adjust the strengths of Max-Min entropy objectives. Experiments on several datasets show that REAL can greatly improve the performance of existing sample rectification models. Xiaowei Fu, Fuxiang Huang, Xinbo Gao 0001, Lei Zhang 0038 |
IEEE Trans. Multim. | 2 |
| 2025 | Elite quantum ant colony algorithm based on double chain encoding for static optimization problems
Xiaowei Fu, Huanyu Li 0015 |
Appl. Intell. | 1 |
| 2025 | Remove to Regenerate: Boosting Adversarial Generalization With Attack InvarianceabstractAdversarial attacks pose a huge challenge to the deployment of deep neural networks (DNNs) in security-sensitive applications. Adversarial defense methods are developed to resist adversarial perturbation. However, most defenses overlook the generalization to various attacks. In medical field, it is known that targeted therapy is a treatment approach at the cellular and molecular levels that targets already identified carcinogenic sites. Inspired by the popular targeted therapies for cancer, we view adversarial attacks as local lesions of natural benign samples. The mechanism behind this assumption implies our key finding that the salient attack components in an adversarial sample dominate the attacking process, while trivial attack components unexpectedly provide trustworthy evidence for obtaining generalizable robustness. Based on this finding, an explainable but efficient Adversarial Surgery and Regeneration (ASR) model following the targeted therapy mechanism is developed to improve the adversarial generalization of DNNs, which has three merits: 1) A score-based Pixel Surgery (PS) module is proposed to remove the salient attack components while retaining the trivial attack components as a kind of attack-invariant information. 2) A Semantic Regeneration module (SR) based on a conditional alignment extrapolator is proposed to restore the discriminative content from the attack-free trivial components, which achieves pixel and semantic consistency for adversarial samples. 3) To further harmonize robustness and accuracy and address such an intractable problem in adversarial defense, a self-augmentation regularizer with adversarial R-drop (ARD) is designed. Experiments on numerous benchmarks show the superiority of the proposed ASR approach. The code can be found inhttps://github.com/fxw13/ASR. Xiaowei Fu, Lei Zhang 0038 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Dynamic Weighted Combiner for Mixed-Modal Image RetrievalabstractMixed-Modal Image Retrieval (MMIR) as a flexible search paradigm has attracted wide attention. However, previous approaches always achieve limited performance, due to two critical factors are seriously overlooked. 1) The contribution of image and text modalities is different, but incorrectly treated equally. 2) There exist inherent labeling noises in describing users' intentions with text in web datasets from diverse real-world scenarios, giving rise to overfitting. We propose a Dynamic Weighted Combiner (DWC) to tackle the above challenges, which includes three merits. First, we propose an Editable Modality De-equalizer (EMD) by taking into account the contribution disparity between modalities, containing two modality feature editors and an adaptive weighted combiner. Second, to alleviate labeling noises and data bias, we propose a dynamic soft-similarity label generator (SSG) to implicitly improve noisy supervision. Finally, to bridge modality gaps and facilitate similarity learning, we propose a CLIP-based mutual enhancement module alternately trained by a mixed-modality contrastive loss. Extensive experiments verify that our proposed model significantly outperforms state-of-the-art methods on real-world datasets. The source code is available at https://github.com/fuxianghuang1/DWC. Fuxiang Huang, Lei Zhang 0038, Xiaowei Fu, Suqi Song |
AAAI | 3 |
| 2024 | An Open-World, Diverse, Cross-Spatial-Temporal Benchmark for Dynamic Wild Person Re-Identification
Lei Zhang 0038, Xiaowei Fu, Fuxiang Huang, Yi Yang 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Attack-defense strategy of UAV swarm based on DEP-SIQ in the active target defense scenario
Xiaowei Fu, Zhe Qiao |
Soft Comput. | 1 |
| 2023 | RGB-T salient object detection via excavating and enhancing CNN features
Hongbo Bi, Ranwan Wu, Yuyu Tong, Xiaowei Fu, Keyong Shao |
Appl. Intell. | 5 |
| 2023 | Bioinspired cooperative control method of a pursuer group vs. a faster evader in a limited area
Xiaowei Fu, Jindong Zhu, Qianglong Wang |
Appl. Intell. | 1 |
| 2022 | Localization and measurement of fetal head in ultrasound image by deep neural networksabstractUltrasound imaging is widely used in prenatal diagnosis of fetuses to monitor fetal development, which is of great clinical significance. The main purpose of prenatal diagnosis is to measure some biological parameters of fetuses, such as Abdominal Circumference (AC), Crown-rump Length (CRL), Femur Length (FL), Biparietal Diameter (BPD), Head Circumference (HC) and so on. However, the artifacts, noise and skull loss in fetal ultrasound images create some challenges to the automatic measurement of fetal HC. In order to improve the accuracy of automatic measurement of fetal HC, this paper proposes Feature Channel Spatial Unet(FCS-Unet) with post processing. Firstly, the fetal head positioning is carried out through Faster R-CNN to find out the coordinates of the fetal head positioning box. Then for fetal head circumference segmentation, Feature Pyramid Networks(FPN) is used to replace the encoder of Unet. And in the skip connections, channel and spatial module are incorporated through sum operation. In the post-processing section, the white pixels outside the head positioning box in the resulting binary images are zeroed and turned to black. Finally, HC is calculated by the least square ellipse fitting method which can obtain the ellipse parameters. Experimental results show that the proposed method is helpful to maintain the details of the fetal head, which has a good robustness to the low signal-to-noise ratio ultrasound images. To a certain extent, our method can improve the accuracy of automatic measurement of fetal HC. Xiaowei Fu |
SMC | 2 |
| 2022 | Cross-Modal Cross-Domain Dual Alignment Network for RGB-Infrared Person Re-IdentificationabstractRGB-Infrared cross-modal person re-identification (Re-ID) has drawn increasing attention due to its application value in practice. Most of the current works rely on a supervised training manner. However, in real-world applications, manual collection of pair-wise RGB-Infrared (IR) person data is labor-intensive and time-consuming. Moreover, when a trained model is directly used in another domain, there is usually a significant performance drop. To overcome the above problems, we make the first attempt to transfer the learned model to a new RGB-IR domain which is unlabeled. The practical problem covers two kinds of challenges, i.e., cross-modal (RGB-Infrared) and cross-domain (different dataset) person Re-ID. Previous works have often considered only one of them either cross-modal or cross-domain. In this work, we propose a dual alignment network (DAN) to solve the RGB-Infrared cross-modal cross-domain person Re-ID problem. This network consists of three parts: Domain Adversarial Alignment component (DAA), Pseudo Label Generation module for target domain (PLG), and Cross-Modal Alignment component (CMA). These three modules complement and promote the model to learn domain-invariant and modality-invariant person representations. Further, we propose a protocol of cross-modal cross-domain person Re-ID by synthesizing target domains by adding random noise, adjusting the lighting intensity, and changing the background color, respectively. Experiments on real and synthetic datasets under the same cross-modalities across domains demonstrate the effectiveness of our method. Xiaowei Fu, Fuxiang Huang, Huimin Ma 0001, Xin Xu 0001, Lei Zhang 0038 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Noise suppressed and bias field corrected image segmentation method for porous Ni-YSZ anode microstructure
Xiaowei Fu, Chengzhen Guo |
Multim. Tools Appl. | 1 |
| 2018 | Task Allocation Method for Multi-UAV Teams with Limited Communication BandwidthabstractMultiple unmanned aerial vehicle (UAV) team often needs to perform some certain tasks cooperatively. In the process of UAVs performing these tasks, the bandwidth of the communication network will affect the task assignment results, then the problem of task allocation for multi-UAV teams under the condition of limited communication bandwidth is considered. For a target, UAVs need to perform reconnaissance, attack and assessment tasks in coordination. The Consensus-Based Bundle Algorithm(CBBA) requires each task is assigned to no more than one UAV, so this paper extends CBBA by duplicating cooperative tasks to modify the task list and adding a judgment mechanism to ensure the uniqueness of the duplicate tasks' allocation in order to achieve the purpose of allocating multiple UAVs to the same task. At the same time, this paper utilizes a bid warping link to improve the algorithm performance and applies CBBA in asynchronous environment to reduce communication burden. Simulation results show that this improved algorithm is feasible and more efficient to improve the task allocation result for multi-UAVs teams with fewer messages transmission and a larger global reward. Xiaowei Fu, Xiaoguang Gao 0001, Jun Chen 0036, Kun Zhang 0029 |
ICARCV | 1 |
| 2016 | Network building and communication strategy designing for multi-UAVs cooperative searchabstractTo improve the coverage of multi-UAVs cooperative search, various communication constraints have to be considered: communication distance, communication delay and communication network structure. A kind of network building method based on tree concept of data structure is presented, and a communication strategy based on information fusion is proposed. These method and strategy could effectively improve the efficiency of multi-UAVs cooperative search. Different situations are designed to verify the rationality and validity of the proposed network building method and communication strategy for improving coverage rate of multi-UAVs cooperative search. Xiaowei Fu, Xiaoguang Gao 0001 |
ICARCV | 1 |
| 2015 | Salient object detection from distinctive features in low contrast imagesabstractSaliency computational model with active environment perception can be useful for many applications including image segmentation, image compression, image retrieval, and etc. Conventional saliency computational models rely on handcrafted low level features, such as color or contrast. These models face great difficulties in low lighting scenarios, due to the lack of well-defined feature to interpret saliency information in low contrast images. In this paper, a new approach is proposed to detect salient object from low contrast images. The proposed approach explores the most distinguishable salient information in low contrast images based on low level features. Extensive experiments have been conducted to evaluate the performance of the proposed method against the state-of-the-art saliency computational models. Xin Xu 0007, Nan Mu, Hong Zhang 0022, Xiaowei Fu |
ICIP | 4 |
| 2014 | Active Learning Methods for Classification of Hyperspectral Remote Sensing Image
Bo Li 0002, Xiaowei Fu |
ICIC (2) | 3 |
| 2013 | DTCWT based medical ultrasound images despeckling using LS parameter optimizationabstractThis paper presents a novel despeckling algorithm that can be used to enhance image quality in medical ultrasound images. Firstly, the log-transformed images are transformed by dual-tree complex wavelet transform (DTCWT). And then, we use a non-Gaussian statistical model with an adaptive smoothing parameter for ideal image signal in the transformed domain. According to Bayesian theory, the MAP estimator is obtained with a proposed adaptive threshold which has better despeckling performance by exploiting the interscale properties of wavelet coefficients. The proposed approach results in significant speckle reduction and preserve details of ultrasound images at the same time while the introduced distortions are not noticeable. Xiaowei Fu, Li Chen 0011, Jing Tian 0002 |
ICIP | 2 |
| 2013 | Homogeneity Based Blind Noisy Image Quality AssessmentabstractBlind noisy image quality assessment aims to evaluate the quality of the degraded noisy image without the need for the ground truth image. To tackle this challenge, this paper proposes an image quality assessment approach using block homogeneity. The contribution of the proposed approach is two-fold. First, a block-based homogeneity measure is proposed to estimate the statistics (e.g., variance) of the noise incurred in the image, based on adaptively selected homogeneous image regions. Second, an image quality assessment approach is proposed by exploiting the above-mentioned estimated noise variance, along with the visual masking effect of the human visual system. Experimental results are provided to demonstrate that the proposed image noise estimation approach yields superior accuracy and stability performance to that of conventional approaches, and the proposed image quality assessment approach achieves consistent performance to that of human subjective evaluation. Xiaotong Huang, Li Chen 0011, Jing Tian 0002, Xiaolong Zhang 0002, Xiaowei Fu |
SMC | 5 |
| 2003 | Virtual face image generation for illumination and pose insensitive face recognitionabstractFace recognition has attracted much attention in the past decades for its wide potential applications. Much progress has been made in the past few years. However, specialized evaluation of the state-of-the-art of both academic algorithms and commercial systems illustrates that the performance of most current recognition technologies degrades significantly due to the variations of illumination and/or pose. To solve these problems, providing multiple training samples to the recognition system is a rational choice. However, enough samples are not always available for many practical applications. It is an alternative to augment the training set by generating virtual views from one single face image, that is, relighting the given face images or synthesize novel views of the given face. Based on this strategy, this paper presents some attempts by presenting a ratio-image based face relighting method and a face re-rotating approach based on linear shape prediction and image warp. To evaluate the effect of the additional virtual face images, primary experiments are conducted using our face specific subspace method as face recognition approach, which shows impressive improvement compared with standard benchmark face recognition methods. Wen Gao 0001, Shiguang Shan, Xiujuan Chai, Xiaowei Fu |
ICASSP (4) | 4 |
| 2003 | Virtual face image generation for illumination and pose insensitive face recognitionabstractFace recognition has attracted much attention in the past decades for its wide potential applications. Much progress has been made in the past few years. However, specialized evaluation of the state-of-the-art in both academic algorithms and commercial systems illustrates that the performance of most current recognition technologies degrades significantly due to the variations of illumination and/or pose. To solve these problems, providing multiple training samples to the recognition system is a rational choice. However, enough samples are not always available for many practical applications. It is an alternative to augment the training set by generating virtual views from one single face image, that is relighting the given face images of synthesize novel views of the given face. Based on this strategy, this paper presents some attempts by presenting a ratio-image based face relighting method and a face re-rotating approach based on linear shape prediction and image warp. To evaluate the effect of the additional virtual face images, primary experiments are conducted using our specific substance method as face recognition approach, which shows impressive improvement compared with standard benchmark face recognition methods. Wen Gao 0001, Shiguang Shan, Xiujuan Chai, Xiaowei Fu |
ICME | 4 |