VLDB 2026 Research / reviewers in the wild / expert
Shengwu Xiong 0001
dblp:96/4134-1
· DBLP profile ↗
241ranked-venue papers
5as first author
154since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 113 · 5 first-author · 65 since 2021Graphics, computer vision, multimedia, augmented reality and games · 75 · 51 since 2021Applied, interdisciplinary, general and emerging computing · 52 · 45 since 2021Databases, data management, data science and information retrieval · 11 · 6 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Computer networks · 6 · 4 since 2021Security and privacy · 6 · 1 since 2021Systems, architecture and hardware · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProPL: Universal Semi-Supervised Ultrasound Image Segmentation via Prompt-Guided Pseudo-LabelingabstractExisting approaches for the problem of ultrasound image segmentation, whether supervised or semi-supervised, are typically specialized for specific anatomical structures or tasks, limiting their practical utility in clinical settings. In this paper, we pioneer the task of universal semi-supervised ultrasound image segmentation and propose ProPL, a framework that can handle multiple organs and segmentation tasks while leveraging both labeled and unlabeled data. At its core, ProPL employs a shared vision encoder coupled with prompt-guided dual decoders, enabling flexible task adaptation through a prompting-upon-decoding mechanism and reliable self-training via an uncertainty-driven pseudo-label calibration (UPLC) module. To facilitate research in this direction, we introduce a comprehensive ultrasound dataset spanning 5 organs and 8 segmentation tasks. Extensive experiments demonstrate that ProPL outperforms state-of-the-art methods across various metrics, establishing a new benchmark for universal ultrasound image segmentation. Yaxiong Chen, Qicong Wang, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
AAAI | 6 |
| 2026 | Driving with Advice: Large Model as Motion Advisor for Joint PlanningabstractWe address the challenge of integrating high-level semantic reasoning with low-level trajectory planning in end-to-end autonomous driving, where most existing frameworks decouple perception, decision-making, and control, leading to limited interpretability and poor instruction compliance. To bridge this gap, we propose Driving with Advice, a novel closed-loop framework that treats a vision-language model (VLM) as a motion advisor to provide interpretable, language-mediated guidance for trajectory generation. Our approach introduces three key innovations: (1) Semantic-Intentional Pretraining (SIP), which injects driving rationale into a compact VLM via machine-generated question-answering pairs; (2) a discrete action space grounded in directional and speed primitives, enabling structured and interpretable policy learning; and (3) an advice-following diffusion policy refined via Group Relative Policy Optimization under a multi-objective reward that ensures safety, comfort, and alignment with semantic intent. We evaluate our method on the NAVSIM benchmark in a closed-loop setting, achieving a state-of-the-art Predictive Driver Model Score (PDMS) of 91.5, outperforming strong baselines in safety (NC: 99.2). The results demonstrate that leveraging language as a cognitive interface between perception and control enhances both generalization and behavioral transparency, advancing the paradigm of language-conditioned driving. Junyin Wang, Jinlei Yu, Huikai Liu, Wenqian Zhu, Shengwu Xiong 0001 |
AAAI | 6 |
| 2026 | Multi-Modal Representation for Spatially Resolved Transcriptomics Based on Global Correlation and Dynamic Cluster DiscoveryabstractConstructing effective representations of spatial resolved transcriptomics (SRT) data, by appropriately characterizing the coherence in gene expression and histology with the spatial information of each sequencing spot, plays an important role in understanding the organization and function of complex tissues. Although much progress has been made, existing SRT representation methods typically establish local associative relationships for each spot only with those located in its surrounding spatial areas, thus failing to capture long-range correlations between distant regions. In addition, the absence of supervision signals on which cluster (with similar biological functions, pathological states or cell types) each spot should belong to also poses a great challenge in deriving an effective representation of SRT data. To this end, we propose a novel Multi-Modal SRT Representation Learning (MMSRL) method based on global spot correlation and dynamic cluster discovery. Specifically, given the gene expression and histological image, MMSRL first builds individual graph convolutional networks (GCNs) for these two modalities and bridges them through the adjacency matrix generated from the spatial locations of different spots. The extracted GCN features are then processed by a correlative self-attention operation to enhance their long-range correlations within each modality. Meanwhile, we also design a multi-modal interaction module (MMIM) to make these two-modal features interact with each other, and align their global correlation information across modalities. After that, we develop an attention-weighted fusion module (AWFM) to adaptively fuse the enhanced intra- and inter-modal features obtained above, so as to effectively integrate the multi-modal information. Finally, a dynamic cluster discovery process, which unsupervisedly assigns cluster labels to each spot, is incorporated to further refine the fused multi-modal representations in a contrastive learning manner. Chuanxiu Li, Shengwu Xiong 0001, Mingxi Sun, Qixiang Zou |
KDD (1) | 2 |
| 2026 | Exploring a double task learning framework for makeup transfer
Zhaoyang Sun, Shengwu Xiong 0001, Yaxiong Chen |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Cross-view visual positioning of unmanned aerial vehicle based on the hierarchical feature guided fusion pyramid network
Bolong Yu, Xuexian Geng, Zhixiang Fang, Shengwu Xiong 0001, Jiaming Qin |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Dual-branch bidirectional attention fusion of temporal and hierarchical representations for robust ECG arrhythmia classification
Basheer A. Hassoon, Shengwu Xiong 0001, Mushtaq A. Hasson, Ridwan Salahudeen, Aminu Onimisi Abdulsalami, Tauqeer Khan |
Expert Syst. Appl. | 2 |
| 2026 | SLOcc: Selective interaction and long-range modelling for occupancy prediction
Junyin Wang, Chenghu Du, Tongao Ge, Shengwu Xiong 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Spatio-temporal synchronous adaptive graph neural network for traffic demand forecasting
Shengwu Xiong 0001 |
Neurocomputing | 2 |
| 2026 | A secure federated feature selection framework for horizontally distributed medical data
Aminu Onimisi Abdulsalami, Farhad Soleimanian Gharehchopogh, Mohammed Abdullahi, Mohamed E. Abd Elaziz, Basheer A. Hassoon, Shengwu Xiong 0001 |
Inf. Process. Manag. | 6 |
| 2026 | A multi-view complementarity deep clustering with self-paced dual attention mechanism
Ridwan Salahudeen, Shengwu Xiong 0001, Adeyemi Abel Ajibesin, Aliyu Garba |
Knowl. Based Syst. | 2 |
| 2026 | WCEDNet: A Weighted Cascaded Encoder-Decoder Network for Hyperspectral Change Detection Based on Spatial-Spectral Difference FeaturesabstractThe core of hyperspectral change detection lies in accurately capturing spectral feature differences across different temporal phases to determine whether surface objects have changed. Since spectral variations of different ground objects often manifest more prominently in specific wavelength bands, we design a Weighted Cascaded Encoder-Decoder Network based on spatial-spectral difference features for hyperspectral change detection. Firstly, unlike conventional change detection frameworks based on siamese networks, our proposed single-branch approach focuses more intensively on extracting spatial-spectral difference features. Secondly, the weighted cascaded structure introduced in the encoder stage enables differential attention to different bands, enhancing focus on spectral bands with high responsiveness. Furthermore, we have developed a spatial-spectral cross-attention module to model intra-feature correlations within spatial and spectral domains. Our method was evaluated on three challenging hyperspectral change detection datasets, and experimental results demonstrate its superior performance compared to competitive models. The detailed code has been open-sourced at https://github.com/WUTCM-Lab/WCEDNet. Bo Zhang 0069, Yaxiong Chen, Ruilin Yao, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2026 | Adversarial discriminant attack on text-to-image diffusion models
Hanxiao Wu, Shengwu Xiong 0001, Dong Yi, Lingxiang Wu, Jianqing Zhu, Guibo Zhu, Jinqiao Wang |
Neural Networks | 2 |
| 2026 | OV-Pro: Enhancing open-vocabulary 3D object detection by prototype contrastive distillation
Tongao Ge, Junyin Wang, Chenghu Du, Hui Li 0010, Huikai Liu, Shengwu Xiong 0001 |
Pattern Recognit. | 6 |
| 2026 | ArtGlyphDiffuser: Text-driven artistic glyph generation via Style-to-CLIP Projection and Multi-Level Controlled diffusion
Xiongbo Lu, Yaxiong Chen, Shengwu Xiong 0001 |
Pattern Recognit. | 4 |
| 2026 | FPMT: Fast and precise high-resolution makeup transfer via Laplacian pyramid
Zhaoyang Sun, Shengwu Xiong 0001 |
Pattern Recognit. | 2 |
| 2026 | D3PD: Dual distillation and dynamic fusion for camera-radar 3D perception
Junyin Wang, Chenghu Du, Tongao Ge, Bingyi Liu, Shengwu Xiong 0001 |
Pattern Recognit. | 5 |
| 2026 | A Novel Approach for Distinguishing Human and AI-Generated TextsabstractThis research introduces innovative methods for differentiating between human-generated and AI-generated text. We developed two new approaches: a hybrid model combining a genetic algorithm with a multilayer perceptron (GA-MLP) and a bidirectional encoder representation from transformers (BERT) approach. Both models show significant improvements over traditional techniques. The article includes a comparative analysis of various conventional methods, such as support vector machines, Naive Bayes, logistic regression, multilayer perceptron, and convolutional neural networks. The experimental results indicate that the GA-MLP approach improves the classification accuracy by approximately 18% compared to traditional models. Furthermore, the BERT approach improves classification accuracy by around 20% over the same conventional methods. This study demonstrates the robustness and effectiveness of the proposed strategies for text classification tasks, which offer superior performance in distinguishing text types compared to traditional techniques. The project can be accessed at https://github.com/zxzxzx77/Human-and-AI-Generated-Texts . Ahmed Abdulhamed, Prabhat Ranjan Singh, Shengwu Xiong 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2026 | RDNet: Rotate-Groundtruth Augmentation and Decoupled Attention HEAD for 3D Object DetectionabstractLiDAR is one of the most important sensors in the field of autonomous driving and allows for better and more accurate perception of the changes in the surrounding environment. Most of the existing 3D object detection methods use data augmentation and feature fusion enhancement to improve the performance of detection, but the majority of the methods ignore the handling of sample imbalance problems during data augmentation. Also, the designed feature fusion and enhancement methods were not well suited to work with the enhancement methods. To this end, we developed a combined method involving data augmentation and feature enhancement. The designed approach has two main objectives: 1) to address the problem of unbalanced sample distribution in detection scenes through data augmentation, and 2) to enhance feature perception using a special feature enhancement module. Our proposed method solves the problem of class imbalance by directly increasing the number of pedestrian samples in the scene through mixed data augmentation, i.e., RG-Aug. In addition, we introduce the Decoupling and Attention Fusion module (DAF), which combines classification headers with high-level features and prediction branches with low-level features. Leverage data features between different layers of features to get a more robust feature representation. Finally, the multi-scale pyramid attention enhancement module is designed to achieve feature enhancement of multi-scale features by means of attention to improve the detection ability of small objects in the scene, especially the detection ability of pedestrians. Our method can achieve 1.57%, 2.16%, and 2.05% performance improvement on the KITTI dataset for Easy, Mod, and Hard samples, respectively. Furthermore, for the detection of pedestrians, our method has a significant competitive advantage over other state-of-the-art techniques with a mAP of 73.42%. Zhenchang Xia, Guanqun Zheng, Shengwu Xiong 0001, Junyin Wang, Jianqun Cui, Yanan Chang, Chenghu Du, Jia Wu 0001 |
IEEE Trans. Big Data | 3 |
| 2026 | DECNet: A Dual-Encoder Contrastive Model for Multilabel Detection of Hate and Abusive Speech in the Sudanese VernacularabstractThe widespread use of social media platforms has drastically transformed communication, enabling users to engage in discussions, share ideas, and express opinions on a global scale. However, this rise in digital interaction has also contributed to the proliferation of harmful content, including hate speech and abusive language, posing significant challenges to online safety and social cohesion. While numerous studies have addressed hate speech detection in the Arabic language, a critical gap remains in handling vernacular variations, particularly in Sudanese vernacular (SV). To address this, we propose a dual encoder contrastive network (DECNet), a novel dual-path encoder framework designed for detecting hate and abusive speech in SV. It integrates a convolutional neural network (CNN) and Transformer backbone with positive-pair contrastive loss (PPCL) and hypersphere projection (HSP) to enhance cross-view alignment and representation diversity. In addition, we introduce the first multilabel dataset specifically curated for SV text. Experimental results on both datasets show that DECNet outperforms all baseline models, achieving a Micro-F1 of 0.92 (a 2–4% improvement over the best transformer baseline) and reducing Hamming Loss by 33%, indicating strong robustness to noise and dialectal variation. Musa Eldow, Shengwu Xiong 0001, Pengfei Duan 0005, Suzanne Hussein |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | ReCoTR: Reducing Semantic Cognitive Shift via Dual-Consensus Token Compression for Remote Sensing Image-Text RetrievalabstractWith the rapid advancement of vision-language models (VLMs) in general-purpose settings, their application to cross-modal retrieval and semantic understanding of large-scale multimodal remote sensing (RS) data is emerging as a key enabler for urban governance, environmental monitoring, and disaster response. However, the pervasive issue of semantic shift in RS image poses a significant challenge to the transferability of pre-trained VLMs. To address this limitation, we propose ReCoTR, an enhanced CLIP-based cross-modal retrieval framework tailored for remote sensing applications. ReCoTR tackles region-level granularity bias and contextual semantic drift through a Dual Consensus Token Evaluation (DCTE) module, which leverages a mixture-of-experts strategy to fuse inter-modal semantic consensus with intra-modal structural consistency, enabling fine-grained estimation of semantic confidence for visual tokens. Moreover, to mitigate representational contamination caused by background noise, we introduce the Semantic Confidence Token Compression (SCTC) module. This module selectively filters and aggregates tokens with high semantic relevance, thus reducing redundancy and alleviating the noise amplification inherent in CLIP's average pooling. Experimental results on three benchmark RS cross-modal retrieval datasets demonstrate that ReCoTR consistently outperforms existing methods on bidirectional image-text retrieval tasks, validating its effectiveness and robustness in remote sensing semantic alignment scenarios. Our source codes are available at: https://github.com/Jerry710/ReCoTR.git. Jirui Huang, Yaxiong Chen, Chuang Du, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Image Process. | 4 |
| 2025 | Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label GenerationabstractEfficiently applying fully supervised learning to virtual try-on tasks is challenging due to the lack of paired ground truth in available training samples. Recent works have achieved virtual try-ons by employing self-supervised learning-based inpainting paradigms. However, this approach is heavily dependent on the constraints of inpainting masks. An incorrect mask can mislead the generated results, while overly large mask areas can lose essential original information, thereby hindering the synthesis of high-quality results. To address these problems, we propose a latent diffusion model-based virtual try-on network that achieves fully supervised learning using the concept of cycle consistency and knowledge distillation. Specifically, we divide our approach into pretext and downstream tasks. In the pretext task, we generate a pseudo-label (pseudo-person image) to form paired training samples, which enables the downstream task to achieve fully supervised learning. To prevent the unreliable pseudo-person image from introducing irresponsible prior knowledge, we propose a noise-covering strategy, which aims at fully optimizing the pseudo-label to eliminate the impact of the incorrect inpainting mask as much as possible. Additionally, we propose a skin refinement loss to further enhance the generation of details in the skin region. Extended experiments demonstrate that our proposed method is superior to state-of-the-art methods. Chenghu Du, Junyin Wang, Feng Yu 0017, Shengwu Xiong 0001 |
AAAI | 4 |
| 2025 | GarFast: Realistic and Fast Garment Transfer with a Simplified Parser-Free ApproachabstractA good garment try-on model should learn the transfer between different types of garments while satisfying: 1) high fidelity and 2) low inference speed. Existing methods address either of these two issues, limited processing speed or low generation quality. We directly use a lightweight encoder-decoder, ensuring faster speeds. To tackle the problem of lower image quality typically generated by lighter models, we present GarFast, a simplified, parser-free framework that optimizes the same lightweight network through a two-stage transformation of real data roles (from input to supervision), thereby greatly promoting model convergence. Specifically, first, we propose a correction strategy to prevent the difficulty of convergence caused by the lack of ground truth in the first stage. Second, we propose a fine-grained domain consistency to ensure that the results generated in the unsupervised first stage are highly realistic clothed human images. Finally, we propose a skin-variant refinement loss and a skinMix regularization to amplify texture differences and enhance the realism of skin-variant regions, thereby improving the quality of the generated skin. Extensive experiments thoroughly demonstrate that our method achieves high resolution, near real-time performance, and superior reconstruction quality compared to state-of-the-art approaches, with processing times of less than 0.03 seconds on an Nvidia A100. Chenghu Du, Junyin Wang, Feng Yu 0017, Shengwu Xiong 0001 |
AAAI | 5 |
| 2025 | RealisID: Scale-Robust and Fine-Controllable Identity Customization via Local and Global ComplementationabstractRecently, the success of text-to-image synthesis has greatly advanced the development of identity customization techniques, whose main goal is to produce realistic identity-specific photographs based on text prompts and reference face images. However, it is difficult for existing identity customization methods to simultaneously meet the various requirements of different real-world applications, including the identity fidelity of small face, the control of face location, pose and expression, as well as the customization of multiple persons. To this end, we propose a scale-robust and fine-controllable method, namely RealisID, which learns different control capabilities through the cooperation between a pair of local and global branches. Specifically, by using cropping and up-sampling operations to filter out face-irrelevant information, the local branch concentrates the fine control of facial details and the scale-robust identity fidelity within the face region. Meanwhile, the global branch manages the overall harmony of the entire image. It also controls the face location by taking the location guidance as input. As a result, RealisID can benefit from the complementarity of these two branches. Finally, by implementing our branches with two different variants of ControlNet, our method can be easily extended to handle multi-person customization, even only trained on single-person datasets. Extensive experiments and ablation studies indicate the effectiveness of RealisID and verify its ability in fulfilling all the requirements mentioned above. Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong 0001 |
AAAI | 7 |
| 2025 | ROME: Radar Sparsity Improvement and Omnimodal Enhancement for 3D Object Detection in Bird's Eye ViewsabstractCombining omnimodal feature interaction using LiDAR, surround-view camera, and Radar to form a network has a great guarantee for the safety of autonomous driving, but most of the current omnimodal fusion methods focus on the interaction enhancement of LiDAR and surround-view camera, ignoring the focus on Radar. Enhancing the contextual representation of Radar can ensure better all-weather capability of the perceptual network. To this end, we design the ROME method based on Radar sparsity improvement to better enhance the performance and robustness of the model in terms of alleviating Radar sparsity shortcomings. Firstly, we design the Autocorrelation Point Enhancement (APE) module to improve Radar sparsity leveraging the point-to-point autocorrelation of Radar. Moreover, for omnimodal Bird’s Eye View (BEV) features, an Omnimodal Adaptive Fusion (OAF) module is designed to improve the robustness of BEV features. With the improved Radar modality, the performance of BEV features for the whole driving scene is further improved. Comprehensive experiments on the nuScenes dataset and comparisons with state-of-the-art methods demonstrate the advantages of our proposed method. Yilong Guo, Junyin Wang, Chenghu Du, Shengwu Xiong 0001, Yaxiong Chen |
ICASSP | 4 |
| 2025 | Enhancing Zero-Shot Relation Extraction through Staged Interaction with Large Language ModelsabstractZero-shot Relation Triplet Extraction (ZeroRTE) is a challenging yet valuable task that extracts relation triplets from unstructured texts for new relation types, significantly reducing the time and effort needed for data labeling. With the advancement of the zero-shot capabilities of large language models, the performance of many zero-shot tasks has been further improved only simply by interacting with large language models (LLMs). In this work, we transform the zero-shot triplet extraction task into a two-stage chat with LLMs. Specifically, in the first stage, we prompt the LLMs to perform Named Entity Recognition (NER). In the second stage, we prompt the LLMs to perform Relation Classification (RC) using the results from the first stage. Experiments on Wiki-ZSL and FewRel datasets show the efficacy of Relation Prompt for the ZeroRTE task. Notably, our method significantly outperforms strong baselines, achieving an impressive 15.89% increase in F1 scores, particularly on WikiZSL with 15 unseen relations. Pengfei Duan 0005, Shengwu Xiong 0001 |
ICASSP | 4 |
| 2025 | All Parts Matter: A Unified Mask-Free Virtual Try-On Framework
Chenghu Du, Shengwu Xiong 0001 |
ICCV | 2 |
| 2025 | Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot SegmentationabstractThis paper studies the few-shot segmentation (FSS) task, which aims to segment objects belonging to unseen categories in a query image by learning a model on a small number of well-annotated support samples. Our analysis of two mainstream FSS paradigms reveals that the predictions made by prototype learning methods are usually conservative, while those of affinity learning methods tend to be more aggressive. This observation motivates us to balance the conservative and aggressive information captured by these two types of FSS frameworks so as to improve the segmentation performance. To achieve this, we propose a **P**rototype-**A**ffinity **H**ybrid **Net**work (PAHNet), which introduces a Prototype-guided Feature Enhancement (PFE) module and an Attention Score Calibration (ASC) module in each attention block of an affinity learning model (called affinity learner). These two modules utilize the predictions generated by a pre-trained prototype learning model (called prototype predictor) to enhance the foreground information in support and query image representations and suppress the mismatched foreground-background (FG-BG) relationships between them, respectively. In this way, the aggressiveness of the affinity learner can be effectively mitigated, thereby eventually increasing the segmentation accuracy of our PAHNet method. Experimental results show that PAHNet outperforms most recently proposed methods across 1-shot and 5-shot settings on both PASCAL-5$^i$ and COCO-20$^i$ datasets, suggesting its effectiveness. The code is available at: [GitHub - tianyu-zou/PAHNet: Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot Segmentation (ICCV'25)](https://github.com/tianyu-zou/PAHNet) Tianyu Zou, Shengwu Xiong 0001, Ruilin Yao |
ICCV | 2 |
| 2025 | Enhancing Virtual Try-On with Text-Image Fusion Guidance
Jingyi Guo, Pengfei Duan 0005, Chenghu Du, Shengwu Xiong 0001 |
ICIC (21) | 4 |
| 2025 | Scene Knowledge Enhanced Multimodal Retrieval Model for Dense Video Captioning
Mingru Huang, Pengfei Duan 0005, Jiawang Peng, Shengwu Xiong 0001 |
ICIC (5) | 6 |
| 2025 | Exploring Flexibility in Incremental Few-Shot Object DetectionabstractIncremental few-shot object detection (iFSD) is critical for real-world applications, enabling rapid adaptation to novel categories with minimal data while mitigating catastrophic forgetting. However, existing methods lack flexibility, particularly in feature representation. The pursuit of a flexible approach to iFSD presents a substantial challenge. To address this, we propose an Attention-Based Feature Aggregation (AFA) that dynamically refines feature representations guided by limited support samples, and a Conditional Classifier (CC) that dynamically refines the generated class prototypes based on the existing knowledge, while conditioning on the limited support images, enhancing flexibility and adaptability. We conducted comprehensive experiments on the MS COCO and LVIS datasets to validate the superiority of our approach. Dongdong Gong, Tengfei Gong, Yaxiong Chen, Jinglin Yuan, Shengwu Xiong 0001 |
ICME | 5 |
| 2025 | MDC: Modality Distribution Consistent Distillation for Multi-View 3D Object DetectionabstractThe purely visual, multi-view perception approach provides a cost-effective solution for autonomous driving perception. However, vision-based systems struggle to achieve the same precision in object localization as LiDAR due to fundamental differences in their sensing mechanisms. To address this, we introduce MDC, a novel method that integrates LiDAR’s superior spatial information into camera-based systems. Our approach includes three key distillation modules: Distribution Consistency Distillation (DCD), Mask Adaptive Distillation (MAD), and Result Distillation (RD). DCD aligns point cloud and multi-view voxel distributions to boost 3D spatial perception. MAD uses adaptive masking to refine BEV feature alignment with LiDAR. RD ensures consistency in the decoding phase. Experiments conducted on the nuScenes benchmark demonstrate that our method achieves a performance improvement of 2.7% to 3.4% over the student network, highlighting its potential to enhance autonomous driving perception capabilities. Huikai Liu, Junyin Wang, Wenqian Zhu, Shengwu Xiong 0001 |
ICME | 5 |
| 2025 | AnyArtisticGlyph: Multilingual Controllable Artistic Glyph GenerationabstractArtistic Glyph Image Generation (AGIG) differs from current creativity-focused generation models by offering finely controllable deterministic generation. It transfers the style of a reference image to a source while preserving its content. Although advanced and promising, current methods may reveal flaws when scrutinizing synthesized image details, often producing blurred or incorrect textures, posing a significant challenge. Hence, we introduce AnyArtisticGlyph, a diffusion-based, multilingual controllable artistic glyph generation model. It includes a font fusion and embedding module, which generates latent features for detailed structure creation, and a vision-text fusion and embedding module that uses the CLIP model to encode references and blends them with transformation caption embeddings for seamless global image generation. Moreover, we incorporate a coarse-grained feature-level loss to enhance generation accuracy. Experiments show that it produces natural, detailed artistic glyph images with state-of-the-art performance. Our project will be open-sourced on https://github.com/jiean001/AnyArtisticGlyph to advance text generation technology. Xiongbo Lu, Yaxiong Chen, Shengwu Xiong 0001 |
ICME | 3 |
| 2025 | Location-Oriented Sound Event Localization and Detection with Spatial Mapping and Regression LocalizationabstractSound Event Localization and Detection (SELD) combines the Sound Event Detection (SED) with the corresponding Direction Of Arrival (DOA). Recently, adopted event-oriented multi-track methods affect the generality in polyphonic environments due to the limitation of the number of tracks. To enhance the generality in polyphonic environments, we propose Spatial Mapping and Regression Localization for SELD (SMRL-SELD). SMRL-SELD segments the 3D spatial space, mapping it to a 2D plane, and a new regression localization loss is proposed to help the results converge toward the location of the corresponding event. SMRL-SELD is location-oriented, allowing the model to learn event features based on orientation. Thus, the method enables the model to process polyphonic sounds regardless of the number of overlapping events. We conducted experiments on STARSS23 and STARSS22 datasets and our proposed SMRL-SELD outperforms the existing SELD methods in overall evaluation and polyphony environments. Xueping Zhang, Yaxiong Chen, Ruilin Yao, Yunfei Zi, Shengwu Xiong 0001 |
ICME | 5 |
| 2025 | Mask Does Not Matter: A Unified Latent Diffusion-Enhanced Framework for Mask-Free Virtual Try-OnabstractA good virtual try-on model should introduce minimal redundant conditional information to avoid instability and increase inference efficiency. Existing methods rely on inpainting masks to guide the generation of the object, but the masks, generated by unstable human parsers, often produce unreliable results with fabric residues due to wrong segmentation. Moreover, large mask regions can lose spatial structure and identity information, requiring extra conditional inputs to compensate, which increases model instability and reduces efficiency. To tackle the problem, we present a novel Mask-Free virtual Try-ON (MFTON) framework. Specifically, we propose a mask-free strategy to eliminate all denoising conditions except for clothing and person images, thereby directly extracting spatial structure and identity information from the person image to improve efficiency and reduce instability. Additionally, to optimize the generated clothing regions, we propose a clothing texture-aware attention mechanism to enable the model to focus on texture generation with significant visual differences. We then introduce a geometric detail capture loss to further enable the model to capture more high-frequency information. Finally, we propose an appearance consistency inference method to reduce the initial randomness of the sampling process significantly. Extensive experiments on popular datasets demonstrate that our method outperforms state-of-the-art virtual try-on methods. Chenghu Du, Junyin Wang, Shengwu Xiong 0001 |
IJCAI | 4 |
| 2025 | Enhancing Semi-Supervised Medical Image Segmentation Through Unbiased Pseudo-Labeling and Student Diversity MaximizationabstractSemi-supervised methodologies effectively utilize limited labeled and abundant unlabeled data to mitigate the scarcity of labeled datasets. Traditional Mean Teacher (MT) approaches, where a student model learns from a teacher’s predictions on unlabeled data, often rely solely on a single model’s predictions, leading to significant confirmation bias and unreliable pseudo-labels. To address these limitations, we propose a Dual Mean-Teacher framework integrated with an unbiased pseudo-labeling strategy and a student diversity maximization strategy. Our approach leverages uncertainties to fuse multiple predictions, reducing confirmation bias and enhancing pseudo-label reliability. The diversity maximization ensures that student models follow distinct learning paths, preventing convergence to similar solutions and improving overall ensemble performance. Comprehensive experiments conducted on three public medical image datasets, encompassing both 2D (MRI) and 3D (CT and MRI) images, demonstrate that our method significantly enhances semi-supervised medical image segmentation, achieving superior results compared to traditional approaches. Guangxing Du, Jinming Xu 0004, Shengwu Xiong 0001 |
IJCNN | 4 |
| 2025 | MAP: Parameter-Efficient Tuning for Referring Expression Comprehension via Multi-Modal Adaptive Positional EncodingabstractThis paper studies the challenging task of Referring Expression Comprehension (REC), which aims at detecting the text-referred target object in an input image. To achieve this, most recent works attempt to adapt powerful pretrained models through integrating additional structures (e.g., low-rank adaptation (LoRA) or adapter modules) to enable efficient parameter tuning. However, all these methods process pretrained features in a position-agnostic manner. This will limit their effectiveness in REC tasks, where the positional information is essential to correctly localize the target object. To this end, we propose a novel parameter-efficient tuning approach, named Multi-Modal Adaptive Positional Encoding (MAP), which addresses the above problem from a new perspective of positional encoding. More specifically, MAP first generates initial positional embeddings for different visual encoder layers from a set of learnable vectors, and then adjusts them adaptively based on spatial-wise visual-linguistic correlations of input data. In this way, the positional information of different image tokens can be appropriately modeled and utilized by MAP, thus making it more applicable to REC tasks. Extensive experiments on five widely-used datasets demonstrate that MAP achieves comparable results to full fine-tuning methods with much fewer extra parameters and outperforms other parameter-efficient tuning approaches. Our source code is available at: https://github.com/Mr-Bigworth/MAP. Ruilin Yao, Tianyu Zou, Bo Zhang 0069, Jian Li 0062, Shengwu Xiong 0001, Shili Xiong |
ACM Multimedia | 6 |
| 2025 | Mitigating Occlusions in Virtual Try-On via A Simple-Yet-Effective Mask-Free FrameworkabstractThis paper investigates the occlusion problems in virtual try-on (VTON) tasks. According to how they affect the try-on results, the occlusion issues of existing VTON methods can be grouped into two categories: (1) Inherent Occlusions, which are the ghosts of the clothing from reference input images that exist in the try-on results. (2) Acquired Occlusions, where the spatial structures of the generated human body parts are disrupted and appear unreasonable. To this end, we analyze the causes of these two types of occlusions, and propose a novel mask-free VTON framework based on our analysis to deal with these occlusions effectively. In this framework, we develop two simple-yet-powerful operations: (1) The background pre-replacement operation prevents the model from confusing the target clothing information with the human body or image background, thereby mitigating inherent occlusions. (2) The covering-and-eliminating operation enhances the model's ability of understanding and modeling human semantic structures, leading to more realistic human body generation and thus reducing acquired occlusions. Moreover, our method is highly generalizable, which can be applied in in-the-wild scenarios, and our proposed operations can also be easily integrated into different generative network architectures (e.g., GANs and diffusion models) in a plug-and-play manner. Extensive experiments on three VTON datasets validate the effectiveness and generalization ability of our method. Both qualitative and quantitative results demonstrate that our method outperforms recently proposed VTON benchmarks. Chenghu Du, Shengwu Xiong 0001, Junyin Wang, Shili Xiong |
NeurIPS | 2 |
| 2025 | Multimodal Co-aware Scale-Spatial Network for Medical Image Segmentation
Enming Huang, Teng Fei Ggong, Yaxiong Chen, Shengwu Xiong 0001 |
PRCV (14) | 4 |
| 2025 | Progressive language-aware encoding and decoding for referring expression comprehension
Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001 |
Sci. China Inf. Sci. | 4 |
| 2025 | Federated clustering with mutual knowledge distillation for traffic flow prediction
Shengwu Xiong 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Artificial Bee Colony Algorithm with Multi-objective in Collaboration Edge ComputingabstractThis work introduces a novel advancement to edge computing by introducing a multi-objective optimization approach. The primary objective of this study is to address the existing research challenges associated with integrating edge computing and Internet of Things (IoT) devices. The utilization of an artificial bee colony technique has led to a decrease in both response time and energy usage within edge computing environments. A formulation of an optimization algorithm based on edges is proposed in order to effectively optimize the trade-off between response time and energy cost. The proposed method exhibits encouraging outcomes through meticulous evaluation in substantially decreasing response time and enhancing energy efficiency. This pioneering approach highlights the potential of the artificial bee colony algorithm as a robust algorithm for enhancing the performance of collaborative edge computing systems. Ahmed Abdulhamed, Prabhat Ranjan Singh, Tanya Shakir Jarad, Shengwu Xiong 0001 |
Int. J. Cooperative Inf. Syst. | 4 |
| 2025 | Make you said that: A motion robust multi-knowledge fusion framework for speaker-agnostic visual dubbing
Shengwu Xiong 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Cross-Domain Density Map-Generated Ship Counting Network for Remote Sensing ImageabstractIn recent years, with the continuous development of remote sensing technology, maritime ship monitoring has become an important research area. Accurately counting the number of ships in remote sensing images is crucial for maritime traffic safety, fisheries management, and marine environmental protection. Existing methods typically use Gaussian kernel functions to generate density maps; however, due to the varied shapes of ships that do not conform to the Gaussian kernel, the resulting density maps fail to accurately reflect the true forms of ships, thereby affecting counting performance. To overcome these limitations, we introduce the cross-domain density map-generated ship counting network (CDDMNet). This network innovatively incorporates a cross-domain feature fusion module (CDFFM), which effectively adapts to ships of varying sizes and shapes. In addition, we have introduced the feature correlation regularization constraint (FCRC) and the integrated loss function, which effectively overcome the disturbances that may arise from variations in ship sizes and enhance the model’s adaptability to changes in ship types and environmental conditions. Experimental results show that the CDDMNet has achieved excellent performance across multiple remote sensing image datasets. Finally, on the RSOC dataset, the mean absolute error (MAE) reached 52.80 and the root mean squared error (RMSE) reached 69.77. Yaxiong Chen, Qijian Li, Kai Yan 0001, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Temporal-Aware Spatial Interaction Transformer for Crop Yield Prediction Based on Multisensor Satellite Image Time SeriesabstractCrop yield prediction is crucial for agricultural decision making. Satellite Image Time Series (SITS) data, which provide continuous temporal observations of vegetation changes, have become a standard for accurate prediction. Recent studies have shown that the integration of multimodal data from different satellite sensors significantly enhances the performance of SITS-based crop yield predictions. However, existing methods often rely on simplistic combinations of multimodal data. Temporal inconsistency between different modals is not considered. In addition, the influence of spatial interaction on crop growth is ignored. To address this issue, we propose the TASI-Transformer (Temporal-Aware Spatial Interaction Transformer) for crop yield prediction using multisensor satellite image time series, which incorporates two innovative modules: a Temporal Enhanced Position Encoding (TEPE) module that incorporates crop growth dates to extract unique temporal information from different satellite time series. A Spatial Enhanced Multimodal Interaction (SEMI) module that learns the impact of spatial relationships between different regions and the interaction between multiple modals. Experimental results on the SICKLE and CROPNET datasets demonstrate that the proposed method achieves state-of-the-art performance, with a Mean Absolute Percentage Error (MAPE) as low as 26.99% in SICKLE using actual season data, and a Root Mean Squared Error (RMSE) of 6.85, Coefficient of Determination(R²) of 0.51, and Pearson Correlation (CORR) of 0.71 in CROPNET, outperforming existing methods. Tengfei Gong, Xinchao Zhu, Yaxiong Chen, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Spatial Invariant Hash Based on Self-Attention Mechanism for Remote Sensing Ship Image RetrievalabstractIn the task of remote sensing ship image retrieval, due to the significant Angle changes and multi-scale characteristics of ships in the image, it is difficult to extract advanced features and the efficiency of feature descriptors is low. Therefore, this paper proposes a spatial invariant hash based on self-attention mechanism for remote sensing ship image retrieval (SIHS). The algorithm consists of two core modules: First, a module based on spatial invariance is designed, which uses deep convolutional neural network to extract continuous real-valued descriptors, and introduces the spatial transformation attention mechanism, and enhances the adaptation ability and learning efficiency of the model to the spatial invariance features through self-learning affine transformation and attention calculation; Secondly, a module based on self-attention hashing is proposed, which improves the efficiency of image representation by multi-scale image embedding, optimizes the attention regularization in the visual encoder, and effectively solves the problem of quantization loss in hash mapping. The experimental results show that the retrieval performance of SIHS algorithm on GGWS, DSCR, FGSC-23 and FGSCR-42 datasets is superior to the existing methods based on deep features. Fuwei Huang, Yaxiong Chen, Kai Yan 0001, Yin Ye, Xuehu Liu, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Entity Naming in NLP: Hybrid Approach GPT Transformer and Multi-level RNNabstractThis research addresses the limitations of existing entity-naming algorithms in natural language processing when the algorithm faces the complexities of polysemy and intricate sentence structures. We propose a novel Transformer-multi-level fusion recurrent Neural Network (T-MFRNN) model that integrates transformer-based layers for contextual understanding with a multi-level fusion recurrent neural network to capture temporal dependencies. Through rigorous experimentation on prominent NLP models like BERT, GPT2, ELECTRA, and XNet, the proposed T-MFRNN demonstrated a significant performance enhancement. Compared to existing methods, such as the boundary assembly model (BAM), the T-MFRNN exhibits a marked improvement of up to 18.66% in the F1-score, highlighting its superior ability to accurately label entities in complex biomedical texts. The T-MFRNN model’s outstanding performance underscores its potential as a robust solution for biomedical entity naming applications. Ahmed Abdulhamed, Prabhat Ranjan Singh, Shengwu Xiong 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2025 | Hybrid Approach for Automatic Text Summarization for Low-resourced Amharic LanguageabstractAutomatic text summarization creates a concise version of the given document while retaining the original content's core ideas, logical structure, and understandability. Despite extensive research on summarization in English and other languages, there remains a shortage of work in Amharic due to limited resources and the challenges posed by the language's complex morphology, syntax, and semantics. Moreover, several feature selection methods have been put forth for major languages. Still, there are no published works on how well they work with the Amharic language in a limited resource context. Furthermore, before putting all of the features together, their individual effects on the hybrid summarization have still not been well investigated for the Amharic language. Our research identifies the best features and addresses the linguistic challenges in Amharic summarization by presenting a hybrid strategy that combines extractive and abstractive methodology and data scarcity issues. The extractive approach utilizes statistical and semantic features such as sentence position, length, and semantics to extract essential sentences. Integrating semantic features with the abstractive approach yielded promising results, surpassing even the combination of statistical and semantic features with the abstractive approach. Moges Ahmed Mehamed, Shengwu Xiong 0001, Awet Fesseha |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2025 | GLV: Geometric Correlation Distillation for Latent Diffusion-Enhanced Parser-Free Virtual Try-OnabstractApplying knowledge distillation to virtual try-on tasks is challenging because current methods fail to fully and efficiently exploit responsible teacher knowledge. In other words, existing approaches merely transfer prior knowledge to the student model via pseudo-labels generated by the teacher model, resulting in shallow knowledge representation and low training efficiency. To address these limitations, we propose a novel teacher-student architecture for parser-free virtual try-on, named GLV, which generates high-quality try-on results with realistic body details. Specifically, we propose a deformation-related prior distillation method to effectively leverage the valuable deformation information contained in the teacher warpage model. This enhances the convergence efficiency of the student warpage model, preventing it from getting stuck in a local minima. Moreover, we are the first to propose a geometric correlation distillation, which models the underlying geometric relationship between clothing and the person and transfers this relationship from the teacher to the student. This enables the student warpage model to reduce the entanglement of deformation-irrelevant features, such as color and texture. Finally, we propose a clothing-body retouching method for try-on result synthesis, which refines the denoising process in the latent space of a well-trained diffusion model, thereby preventing catastrophic forgetting. This method seamlessly transforms the parser-based inpainting synthesis paradigm into a parser-free synthesis paradigm and enables efficient convergence of the diffusion model with only fine-tuning. Extensive experiments demonstrate the generality of our approach and highlight its superiority over previous methods. Chenghu Du, Junyin Wang, Shengwu Xiong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Bow Direction Detection Based on Angular Coding With Heading Intersection Over Union LossabstractAccurate bow direction detection is essential for ship trajectory prediction and port monitoring. Existing ship detection networks typically output angles within 180°, while extending to 360° introduces cyclic issues affecting rotation intersection over union (RIoU) accuracy. This study proposes a novel bow direction detection algorithm that extends network output to 360° and integrates a heading intersection over union (HIoU) loss to enhance detection accuracy and robustness. Additionally, an HIoU loss function is designed to improve bow direction identification and reduce quantization errors in hash codes. The algorithm is evaluated on three datasets: FGSD, OHD-SJTU-S, and OHD-SJTU-L. On FGSD, it achieves mean average precision (mAP) of 91.14%. On OHD-SJTU-S, it attains an$\text {mAP}_{50:95}$of 63.3% and a bow direction prediction accuracy of 90.7%. On OHD-SJTU-L, the$\text {mAP}_{50:95}$is 29.2%, with an accuracy of 80.2%. Yaxiong Chen, Qiangqiang Huang, Hao Sun 0014, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Bilinear Parallel Fourier Transformer for Multimodal Remote Sensing ClassificationabstractVision Transformers (ViTs) have shown promise in multimodal fusion image classification, yet face performance challenges in complex remote sensing scenarios. Single fusion frameworks often fail to fully utilize multimodal diversity, and the uneven distribution of image categories complicates the accurate construction of spatial structures by Transformers. Additionally, traditional cross-entropy tends to favor majority classes, neglecting minority classes, resulting in suboptimal predictions and reduced overall accuracy (OA). To solve these challenges, we propose a novel deep neural network, a bilinear parallel Fourier Transformer (BPFT). We propose a novel dual-fusion feature interaction (DFFI) module that utilizes two distinct types of fused features for learning, namely the spatial-spectral fusion feature and the global fusion feature. Besides, we introduce a dual-feature interaction (DFI) module to improve the utilization of fused feature information. To enable the Transformer to better establish spatial structural relationships, we employ the Fourier transform in place of the self-attention mechanism. To address the focus on minority class labels, we propose an exponential label smoothing cross-entropy loss function. This loss function comprises two components: exponential cross-entropy and label smoothing. The exponential cross-entropy component applies a strong penalty to misclassified samples, thereby increasing attention on minority class labels. To validate the efficacy of our approach, extensive experiments are conducted across two multimodal remote sensing datasets: Augsburg and Berlin, encompassing hyperspectral imaging (HSI) data and synthetic aperture radar (SAR) data. The results of these experiments affirm the superior performance of our proposed BPFT model compared to existing state-of-the-art models in multimodal remote sensing image classification tasks. Yaxiong Chen, Qicong Wang, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Global-Local Fusion With Semantic Information Guidance for Accurate Small Object Detection in UAV Aerial ImagesabstractIn recent years, the rapid development of the unmanned aerial vehicle (UAV) technology has generated a large number of aerial photography images captured by UAV. Consequently, the object detection in UAV aerial images has emerged as a recent research focus. However, due to the flexible flight heights and diverse shooting angles of UAV, two significant challenges have arisen in UAV aerial images: extreme variation in target scale and the presence of numerous small targets. To address these challenges, this article introduces a semantic information-guided fusion module specifically tailored for small targets. This module utilizes high-level semantic information to guide and align the underlying texture information, thereby enhancing the semantic representation of small targets at the feature level and subsequently improving the model’s ability to detect them. In addition, this article introduces a novel global–local fusion detection strategy to strengthen the detection of small targets. We have redesigned the foreground region assembly method to address the drawbacks of previous methods that involved multiple inferences. Extensive experiments conducted on the VisDrone and UAVDT datasets demonstrate that our two self-designed modules can significantly enhance the detection capability of small targets compared with the YOLOX-M model. Our code is publicly available at:https://github.com/LearnYZZ/GLSDet. Yaxiong Chen, Zhengze Ye, Haokai Sun 0001, Tengfei Gong, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | VGRSS: Datasets and Models for Visual Grounding in Remote Sensing Ship ImagesabstractThis paper introduces a task named Visual Grounding of Remote Sensing Ship Images (VGRSS). The goal of VGRSS is to locate ship objects in remote sensing images guided by natural language. Extensive research has been conducted on multimodal processing of remote sensing images and text to retrieve rich information from remote sensing images using natural language. However, due to the unique characteristics of remote sensing ship images, ship localization using natural language remains a challenge. Therefore, in this work, we construct datasets for the VGRSS task and explore deep learning models. Specifically, our contributions can be summarized as follows: First, we construct two remote sensing ship datasets for visual grounding. One is based on the optical remote sensing dataset, named RSSVG, while the other is based on the synthetic aperture radar (SAR) dataset, named SARVG. Second, we propose a Language-Guided Visual Feature Enhancement (LVFE) module. This module enhances visual features through language guidance before Visual-Linguistic Fusion. Third, we propose a Visual-Linguistic Fusion (VLF) module based on multimodal feature stacking. This module inputs the stacked language and visual features, and then performs feature fusion using a Transformer, enabling effective cross-modal interaction and integration. Fourth, we introduce a novel loss calculation method by incorporating Enhanced Intersection over Union (EIoU) into the loss function. Finally, we benchmark extensive state-of-the-art (SOTA) natural image visual grounding methods on the constructed RSSVG and SARVG datasets, then provide insightful analysis based on the results. This work offers valuable insights for developing better VGRSS models. Yaxiong Chen, Liwen Zhan, Yichen Zhao, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | A Dual-Stage Wavelet and Linear Attention Enhancement Network for Agricultural Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification faces unique challenges in agricultural scenario due to spectral-spatial feature similarity caused by complex planting structures and high spectral similarity. Existing spatial-spectral joint feature extraction methods fail to fully exploit the advantages of spatial and spectral information, thus have certain limitations and cannot effectively distinguish similar crops in agricultural scenarios. To address these limitations, we proposed a dual-stage wavelet and linear attention enhancement network (DSW-LAN) for agricultural HSI classification, addressing the challenges of complex spatial-spectral information and high redundancy. We integrates a direction factorized deformable 3D Convolution (DFDWConv3D) module to capture multi-scale spatial-spectral features through adaptive kernel adjustments, while wavelet transform decomposes spatial features into low-frequency (structural) and high-frequency (textural) components for targeted enhancement. Additionally, a spectral probe-guided linear attention mechanism efficiently models long-range spectral dependencies with reduced computational complexity by prioritizing discriminative bands. Experimental results demonstrate superior performance on three challenging agricultural HSI datasets, achieving enhanced classification accuracy with reduced computational complexity. Yaxiong Chen, Bo Zhang 0069, Shili Xiong, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Discover the Unknown Ones in Fine-Grained Ship DetectionabstractRemote sensing image-based ship identification technology has great applications in areas such as national defense and fishery management. However, existing remote sensing ship studies mainly focus on a closed environment and overlook actual sea conditions, while new military ships will be encountered. These unknown categories of ships will be ignored or misclassified by existing models, dramatically affecting the accurate assessment of the maritime situation. Furthermore, existing unknown detection methods for natural images fail to tackle the remote sensing ship detection problem for the property of high similarity in overall appearance. To cope with this problem, this paper proposes a fine-grained unknown ship detection network. Firstly, we explore a class-balanced proposal sampler to avoid inefficient information learning. Secondly, we propose a finegrained memory bank-based contrastive learning strategy to separate different categories. Finally, to further separate unknown classes, we adopt an uncertainty-aware unknown learner with logit to reduce the uncertainty of fine-grained predictions. Experiments conducted in three public ship detection datasets ShipRSImageNet, DOSR, and HRSC2016 show that the method not only achieves good detection on unknown class ships, but also improves the detection accuracy on known classes. The code is available at https://github.com/FoRGEU/DUONet. Tengfei Gong, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Multibranch Fusion-Based Feature Enhance for Remote-Sensing Scene ClassificationabstractRemote-sensing (RS) scene classification is a fundamental and significant task in RS image interpretation, involving the annotation of semantic content. RS scene images are characterized by complex backgrounds, rich content, and multiscale targets, exhibiting both intraclass separation and interclass convergence. Therefore, extracting features that effectively express the intrinsic attributes of images and possess high discriminative is crucial for RS scene classification. Existing global-based methods often lack the ability to capture significant detailed information in similar scenes. Conversely, methods based on local discriminative features tend to overlook the interrelationships of objects within the same scene. To address these issues, this article proposes a unified framework named MBFNet to align and fuse features of different scales and levels for accurate RS scene classification. We utilize a multibranch feature-extracting network structure with parallel convolution and Transformer modules. Simultaneously, a kernel-selected multiscale aggregation (KSMSA) module is designed to efficiently process the diverse scale features emanating from these parallel branches. By selecting different convolution kernels, a dynamic receptive field is established to adaptively process features of different scales, reducing semantic differences to achieve effective aggregation of multiscale features. Moreover, a learnable multilevel aggregation (LMLA) module is designed to integrate shallow features, such as shape information, into deep features for more comprehensive feature fusion. Benefiting from KSMSA and LMLA, the proposed MBFNet improves the discriminability of features, thereby enhancing classification performance. Comprehensive experiments on three benchmark datasets demonstrate that the proposed method outperforms state-of-the-art RS scene classification methods in terms of performance. Xiongbo Lu, Meng Yang 0034, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | SSPNet: Spatial-Spectral Perception Network for Mineral Hyperspectral Image ClassificationabstractUnlike general scenes, mineral hyperspectral images often exhibit similar spatial and spectral characteristics across different mines, making traditional classification methods less effective due to compromised robustness. To address this, we propose a Spatial-Spectral Perception Network for mineral hyperspectral image classification. This approach divides spatial-spectral feature extraction into two stages. In the spatial feature perception stage, we introduce a Spatial Frequency Perceptron that maps three-dimensional spatial features into low-frequency and high-frequency domains. We then apply Triple-Cross-Attention to each frequency domain to better differentiate spatial features of similar mines. In the spectral perception stage, we design a Spectral Linear Perceptron using Absolute Linear Attention, which captures fine-grained spectral differences by establishing internal relationships between spectral features through Absolute Positional Weighting. This enables effective separation of similar spectra for final classification. Extensive experiments on three publicly available mineral hyperspectral image datasets and one agricultural hyperspectral dataset show that our method outperforms popular alternatives in both effectiveness and robustness. The open-source code can be accessed at https://github.com/WUTCM-Lab/SSPNet. Bo Zhang 0069, Yaxiong Chen, Ruilin Yao, Shili Xiong, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Hyperspectral Image Classification via Cascaded Spatial Cross-Attention NetworkabstractIn hyperspectral images (HSIs), different land cover (LC) classes have distinct reflective characteristics at various wavelengths. Therefore, relying on only a few bands to distinguish all LC classes often leads to information loss, resulting in poor average accuracy. To address this problem, we propose a method called Cascaded Spatial Cross-Attention Network (CSCANet) for HSI classification. We design a cascaded spatial cross-attention module, which first performs cross-attention on local and global features in the spatial context, then uses a group cascade structure to sequentially propagate important spatial regions within the different channels, and finally obtains joint attention features to improve the robustness of the network. Moreover, we also design a two-branch feature separation structure based on spatial-spectral features to separate different LC Tokens as much as possible, thereby improving the distinguishability of different LC classes. Extensive experiments demonstrate that our method achieves excellent performance in enhancing classification accuracy and robustness. The source code can be obtained from https://github.com/WUTCM-Lab/CSCANet. Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Image Process. | 3 |
| 2025 | HybridBEV: Hybrid Encode and Distillation for Improved BEV 3D Object DetectionabstractThe development of surround-view cameras is crucial for the advancement of autonomous driving. Utilizing depth information and image features to simulate LiDAR bird’s-eye-view (BEV) features can accomplish efficient 3D object detection tasks. Existing dense BEV generation methods heavily rely on the use of depth features, however, the suboptimal exploitation of these features often results in ambiguity in object location and feature representation during the BEV generation process. To address this, we have designed a hybrid encode and distillation method to enhance 3D object detection performance, termed HybridBEV. Initially, we designed the HybridEncode module, which employs a resampling strategy of depth features in voxel space to obtain BEV features that more accurately reflect the distribution of objects. Subsequently, we introduced multiple distillation methods to supervise the network’s voxel features and BEV feature representations, assisting the student network in learning critical features from the teacher model and ensuring that BEV features can more distinctly represent object distribution. Furthermore, during network training, we loaded pre-trained weights from the teacher network to guide network optimization and accelerate training. Extensive experiments on the nuScenes benchmark demonstrate that HybridBEV can effectively improve the performance of the student network and outperform previous state-of-the-art methods based on surround-view cameras. The code will be published athttps://github.com/wjyxx/HybridBEV Junyin Wang, Chenghu Du, Huikai Liu, Zhenchang Xia, Bingyi Liu, Shengwu Xiong 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Uncertainty Co-Estimator for Improving Semi-Supervised Medical Image SegmentationabstractRecently, combining the strategy of consistency regularization with uncertainty estimation has shown promising performance on semi-supervised medical image segmentation tasks. However, most existing methods estimate the uncertainty solely based on the outputs of a single neural network, which results in imprecise uncertainty estimations and eventually degrades the segmentation performance. In this paper, we propose a novel Uncertainty Co-estimator (UnCo) framework to deal with this problem. Inspired by the co-training technique, UnCo establishes two different mean-teacher modules (i.e., two pairs of teacher and student models), and estimates three types of uncertainty from the multi-source predictions generated by these models. Through combining these uncertainties, their differences will help to filter out incorrect noise in each estimate, thus allowing the final fused uncertainty maps to be more accurate. These resulting maps are then used to enhance a cross-consistency regularization imposed between the two modules. In addition, UnCo also designs an internal consistency regularization within each module, so that the student models can aggregate diverse feature information from both modules, thus promoting the semi-supervised segmentation performance. Finally, an adversarial constraint is introduced to maintain the model diversity. Experimental results on four medical image datasets indicate that UnCo can achieve new state-of-the-art performance on both 2D and 3D semi-supervised segmentation tasks. The source code will be available at https://github.com/z1010x/UnCo. Shengwu Xiong 0001, Jinming Xu 0004, Guangxing Du |
IEEE Trans. Medical Imaging | 2 |
| 2025 | SSAT++: A Semantic-Aware and Versatile Makeup Transfer Network With Local Color Consistency ConstraintabstractThe purpose of makeup transfer (MT) is to transfer makeup from a reference image to a target face while preserving the target's content. Existing methods have made remarkable progress in generating realistic results but do not perform well in terms of semantic correspondence and color fidelity. In addition, the straightforward extension of processing videos frame by frame tends to produce flickering results in most methods. These limitations restrict the applicability of previous methods in real-world scenarios. To address these issues, we propose a symmetric semantic-aware transfer network (SSAT++) to improve makeup similarity and video temporal consistency. For MT, the feature fusion (FF) module first integrates the content and semantic features of the input images, producing multiscale fusion features. Then, the semantic correspondence from the reference to the target is obtained by measuring the correlation of fusion features at each position. According to semantic correspondence, the symmetric mask semantic transfer (SMST) module aligns the reference makeup features with the target content features to generate MT results. Meanwhile, the semantic correspondence from the target to the reference is obtained by transposing the correlation matrix and applied to the makeup removal task. To enhance color fidelity, we propose a novel local color loss that forces the transferred results to have the same color histogram distribution as the reference. Furthermore, a morphing simulation is designed to ensure temporal consistency for video MT without requiring additional video frame input and optical flow estimation. To evaluate the effectiveness of our SSAT++, extensive experiments have been conducted on the MT dataset which has a variety of makeup styles, and on the MT-Wild dataset which contains images with diverse poses and expressions. The experiments show that SSAT++ outperforms existing MT methods through qualitative and quantitative evaluation and provides more flexible makeup control. Code and trained model will be available at https://gitee.com/sunzhaoyang0304/ssat-msp and https://github.com/Snowfallingplum/SSAT. Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | MakeupDiffuse: a double image-controlled diffusion model for exquisite makeup transfer
Xiongbo Lu, Yaxiong Chen, Shengwu Xiong 0001 |
Vis. Comput. | 5 |
| 2025 | Semi-hard constraint augmentation of triplet learning to improve image corruption classification
Shengwu Xiong 0001, Zhaoyang Sun, Jianwen Xiang |
Vis. Comput. | 2 |
| 2024 | CycleVTON: A Cycle Mapping Framework for Parser-Free Virtual Try-OnabstractImage-based virtual try-on aims to transfer a target clothing onto a specific person. A significant challenge is arbitrarily matched clothing and person lack corresponding ground truth to supervised learning. A recent pioneering work leveraged an improved cycleGAN to enable one network to generate the desired image for another network during training. However, there is no difference in the result distribution before and after the clothing changes. Therefore, using two different networks is unnecessary and may even increase the difficulty of convergence. Furthermore, the introduced human parsing used to provide body structure information in the input also have a negative impact on the try-on result. How to employ a single network for supervised learning while eliminating human parsing? To tackle these issues, we present a Cycle mapping Virtual Try-On Network (CycleVTON), which can produce photo-realistic try-on results by using a cycle mapping framework without the parser. In particular, we introduce a flow constraint loss to achieve supervised learning of arbitrarily matched clothing and person as inputs to the deformer, thus naturally mimicking the interaction between clothing and the human body. Additionally, we design a skin generation strategy that can adapt to the shape of the target clothing by dynamically adjusting the skin region, i.e., by first removing and then filling skin areas. Extensive experiments conducted on challenging benchmarks demonstrate that our proposed method exhibits superior performance compared to state-of-the-art methods. Chenghu Du, Junyin Wang, Shuqing Liu, Shengwu Xiong 0001 |
AAAI | 6 |
| 2024 | PViT: Pooling Vision Transformer for Active Trachoma Image Classification
Mulugeta Shitie Zewudie, Shengwu Xiong 0001, Xiaohan Yu 0001, Aminu Onimisi Abdulsalami |
ADMA (4) | 2 |
| 2024 | Content-Style Decoupling for Unsupervised Makeup Transfer without Generating Pseudo Ground TruthabstractThe absence of real targets to guide the model training is one of the main problems with the makeup transfer task. Most existing methods tackle this problem by synthesizing pseudo ground truths (PGTs). However, the generated PGTs are often sub-optimal and their imprecision will eventually lead to performance degradation. To alleviate this issue, in this paper, we propose a novel Content-Style Decoupled Makeup Transfer (CSD-MT) method, which works in a purely unsupervised manner and thus eliminates the negative effects of generating PGTs. Specifically, based on the frequency characteristics analysis, we assume that the low-frequency (LF) component of a face image is more associated with its makeup style information, while the high-frequency (HF) component is more related to its content details. This assumption allows CSD-MT to decouple the content and makeup style information in each face image through the frequency decomposition. After that, CSD-MT realizes makeup transfer by maximizing the consistency of these two types of information between the transferred result and input images, respectively. Two newly designed loss functions are also introduced to further improve the transfer performance. Extensive quantitative and qualitative analyses show the effectiveness of our CSD-MT method. Our code is available at https://github.com/Snowfallingplum/CSD-MT. Zhaoyang Sun, Shengwu Xiong 0001, Yaxiong Chen |
CVPR | 2 |
| 2024 | IFNET: Integrating Data Augmentation and Decoupled Attention Fusion for 3D Object DetectionabstractLiDAR is a key sensor for accurately sensing of the environment in autonomous driving. While existing 3D object detection methods generally rely on data augmentation and feature fusion to improve performance, the challenge of dealing with sample imbalance is often overlooked. We design a novel 3D detection network, IFNet, that tackles these issues by introducing mutually reinforcing data augmentation and feature enhancement strategies. It aims to achieve a dual purpose: 1) correcting the category imbalance by directly enhancing pedestrian samples using mixed data augmentation, i.e., RG-Aug; and 2) enhancing feature perception by introducing the decoupling and attention fusion module (DAF). DAF enables robust feature representations across different layers, improving the detection performance, especially for small objects in the scene. Comprehensive experiments on the KITTI dataset and comparisons with state-of-the-art methods demonstrate the superiority of our proposed approach. Zhenchang Xia, Guanqun Zheng, Shengwu Xiong 0001, Jia Wu 0001, Junyin Wang, Chenghu Du |
ICASSP | 3 |
| 2024 | ST-CLIP: Spatio-Temporal Enhanced CLIP Towards Dense Video Captioning
Pengfei Duan 0005, Mingru Huang, Jingyi Guo, Shengwu Xiong 0001 |
ICIC (11) | 5 |
| 2024 | Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event DetectionabstractAlthough the Prototypical Network (ProtoNet) has demonstrated effectiveness in few-shot biological event detection, two persistent issues remain. Firstly, there is difficulty in constructing a representative negative prototype due to the absence of explicitly annotated negative samples. Secondly, the durations of the target biological vocalisations vary across tasks, making it challenging for the model to consistently yield optimal results across all tasks. To address these issues, we propose a novel adaptive learning framework with an adaptive learning loss to guide classifier updates. Additionally, we propose a negative selection strategy to construct a more representative negative prototype for ProtoNet. All experiments ware performed on the DCASE 2023 TASK5 few-shot bioacoustic event detection dataset. The results show that our proposed method achieves an F-measure of 0.703, an improvement of 12.84%. Yaxiong Chen, Xueping Zhang, Yunfei Zi, Shengwu Xiong 0001 |
ICME | 4 |
| 2024 | Multi-Agent Reinforcement Learning Based Resource Allocation for Efficient Message Dissemination in C-V2X NetworksabstractIn order to support diverse applications in intelligent transportation, intelligent connected vehicles (ICVs) need to send multiple types of messages, such as periodic messages and event-driven messages with different frame specifications. However, existing researches often concentrate on the transmission of single-message types, overlooking hybrid communication scenarios where multiple types of messages coexist, posing challenges in meeting the diverse transmission needs of different message types. To optimize the Quality of Service (QoS) in such scenarios, we take the perspective of ICVs and formulate their decision making as a multi-agent reinforcement learning problem. More specifically, we propose a cooperative individual rewards assisted multi-agent reinforcement learning (CIRA) framework. The transformer structure in CIRA is used to avoid mutual interference during the transmission of different vehicles. Besides, the introduction of individual rewards and the dual-layer architecture of CIRA contribute to providing ICVs with more forward-looking message dissemination scheme. Finally, we set up a simulator to create dynamic traffic scenarios reflecting different real-world conditions. We conduct extensive experiments to evaluate the proposed CIRA framework’s performance. The results show that CIRA can significantly improve the packet reception rates and ensure low communication delays in various scenarios. Bingyi Liu, Jingxiang Hao, Enshu Wang, Dongyao Jia, Weizhen Han, Shengwu Xiong 0001 |
IWQoS | 7 |
| 2024 | CausalCLIPSeg: Unlocking CLIP's Potential in Referring Medical Image Segmentation with Causal Intervention
Yaxiong Chen, Minghong Wei, Zixuan Zheng, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (3) | 6 |
| 2024 | Striving for Simplicity: Simple Yet Effective Prior-Aware Pseudo-labeling for Semi-supervised Ultrasound Image Segmentation
Yaxiong Chen, Zixuan Zheng, Jingliang Hu, Yilei Shi, Shengwu Xiong 0001, Xiao Xiang Zhu 0001, Lichao Mou |
MICCAI (9) | 6 |
| 2024 | Visual Grounding with Multi-modal Conditional AdaptationabstractVisual grounding is the task of locating objects specified by natural language expressions. Existing methods extend generic object detection frameworks to tackle this task. They typically extract visual and textual features separately using independent visual and textual encoders, then fuse these features in a multi-modal decoder for final prediction. However, visual grounding presents unique challenges. It often involves locating objects with different text descriptions within the same image. Existing methods struggle with this task because the independent visual encoder produces identical visual features for the same image, limiting detection performance. Some recently approaches propose various language-guided visual encoders to address this issue, but they mostly rely solely on textual information and require sophisticated designs. In this paper, we introduce Multi-modal Conditional Adaptation (MMCA), which enables the visual encoder to adaptively update weights, directing its focus towards text-relevant regions. Specifically, we first integrate information from different modalities to obtain multi-modal embeddings. Then we utilize a set of weighting coefficients, which generated from the multimodal embeddings, to reorganize the weight update matrices and apply them to the visual encoder of the visual grounding model. Extensive experiments on four widely used datasets demonstrate that MMCA achieves significant improvements and state-of-the-art results. Ablation experiments further demonstrate the lightweight and efficiency of our method. Our source code is available at: https://github.com/Mr-Bigworth/MMCA. Ruilin Yao, Shengwu Xiong 0001, Yichen Zhao |
ACM Multimedia | 2 |
| 2024 | SHMT: Self-supervised Hierarchical Makeup Transfer via Latent Diffusion ModelsabstractThis paper studies the challenging task of makeup transfer, which aims to apply diverse makeup styles precisely and naturally to a given facial image. Due to the absence of paired data, current methods typically synthesize sub-optimal pseudo ground truths to guide the model training, resulting in low makeup fidelity. Additionally, different makeup styles generally have varying effects on the person face, but existing methods struggle to deal with this diversity. To address these issues, we propose a novel Self-supervised Hierarchical Makeup Transfer (SHMT) method via latent diffusion models. Following a "decoupling-and-reconstruction" paradigm, SHMT works in a self-supervised manner, freeing itself from the misguidance of imprecise pseudo-paired data. Furthermore, to accommodate a variety of makeup styles, hierarchical texture details are decomposed via a Laplacian pyramid and selectively introduced to the content representation. Finally, we design a novel Iterative Dual Alignment (IDA) module that dynamically adjusts the injection condition of the diffusion model, allowing the alignment errors caused by the domain gap between content and makeup representations to be corrected. Extensive quantitative and qualitative analyses demonstrate the effectiveness of our method. Our code is available at https://github.com/Snowfallingplum/SHMT. Zhaoyang Sun, Shengwu Xiong 0001, Yaxiong Chen |
NeurIPS | 2 |
| 2024 | A Fine Rendering High-Resolution Makeup Transfer network via inversion-editing strategy
Zhaoyang Sun, Shengwu Xiong 0001, Yaxiong Chen |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Fisher ratio-based multi-domain frame-level feature aggregation for short utterance speaker verificationabstractAs the durations of the short utterances are small, it is difficult to learn sufficient information to distinguish the person, thus, short utterance speaker recognition is highly challenging. In this paper, we propose a multi-domain frame-level feature joint learning method to aggregate the discriminative information from multiple dimensions and domain, which is different domains of the speech, time-domain, frequency-domain, and spectral-domain, represent distinct physical characteristics and provide different dimension information, the time domain captures information about the temporal aspect of the physical signal, the frequency domain represents the signal strength in different frequency ranges, and the spectral domain reflects the overall information of the speech, then, based on the extracted multi-domain frame-level features, using the Multi-Fisher criterion aggregates feature parameters categorically and match the corresponding Multi-Fisher ratio weights to the feature parameters as a way to achieve effective feature aggregation and to preserve more effective information, termed Firm-Domain. Extensive experiments are carried out on short-duration text-independent speaker verification datasets derived from the VoxCeleb, SITW, and NIST SRE corpora, which contain speech samples of varying lengths and scenarios. The results demonstrate that the proposed method outperforms the state-of-the-art deep learning architectures by at least 13%, respectively, in the test set. The results of the ablation experiments demonstrate that our proposed methods can significantly outperform previous approaches. Yunfei Zi, Shengwu Xiong 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Multi-Fisher and Triple-Domain Feature Enhancement-Based Short Utterance Speaker Verification for IoT Smart ServiceabstractSpeech authentication in IoT smart services typically involves short utterances. However, due to the short duration of these utterances (e.g., less than 3 s) and the limited enrolment and/or test data available, it is difficult to learn enough information to accurately distinguish the person. As a result, speaker recognition from short utterances is very challenging. In this article, in the acoustic end, we propose a Multi-Fisher feature enhancement method to enrich short utterance effective information. The method aggregates the effective feature parameters by assigning the corresponding Multi-Fisher ratio weights to these parameters. In the architecture end, we propose a triple-domain feature joint learning method to enhance discriminative information from multiple dimensions. This approach provides different dimensions of information through the different physical meanings of speech in the time domain, the frequency domain, and the spectral domain. Extensive experiments were conducted on VoxCeleb, SITW, and NIST SRE short-duration text-independent speaker verification tasks, containing speech samples of varying lengths and scenarios. The results show that our proposed method outperforms existing acoustic feature extraction approaches and state-of-the-art deep learning architectures by at least 6% and 12%, respectively. The ablation experiments further illustrate that our proposed approaches can achieve substantial improvement over previous methods. Yunfei Zi, Shengwu Xiong 0001 |
IEEE Internet Things J. | 2 |
| 2024 | An Improved Heterogeneous Comprehensive Learning Symbiotic Organism Search for Optimization Problems
Aminu Onimisi Abdulsalami, Mohamed E. Abd Elaziz, Farhad Soleimanian Gharehchopogh, Ahmed Tijani Salawudeen, Shengwu Xiong 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Adaptive feature selection for active trachoma image classificationabstractTrachoma is a neglected tropical eye disease caused by ocular strains of Chlamydia trachomatis, which affects millions of people worldwide. To examine the eye for signs of active trachoma, healthcare providers typically look for clusters of five or more follicles on the conjunctiva of the upper eyelid for the follicular inflammatory trachoma stage. However, it is also possible to find individual follicles scattered throughout the conjunctiva, particularly in mild or early-stage trachoma cases. Additionally, the datasets are photographic images collected in the field that can be high-dimensional and may contain large amounts of redundant information. We propose integrating novel attention-based feature extraction and feature selection techniques to address these challenges. First, we present the Lambda layer within the Convolutional Block Attention Module (L-CBAM) to normalize attention weights and improve the feature extraction process. Second, we introduce an adaptive mechanism, Adaptive Beta Hill Climbing (AβHC) with Social Ski-Driver (SSD), which adjusts the exploration-exploitation trade-off during the search process, allowing for better exploration of the search space and more efficient convergence toward an optimal feature subset. We then use the multilayer perceptron (MLP) classifier to produce final classification results using selected subsets. We evaluated the proposed approach on active trachoma inverted eyelid images and obtained accuracy scores of 93.3% with only 19.7% of the selected features, surpassing many of the algorithms used for comparison. Our proposed method has demonstrated excellent performance compared to recent works utilizing the same datasets. Mulugeta Shitie Zewudie, Shengwu Xiong 0001, Xiaohan Yu 0001, Xiaoyu O. Wu, Moges Ahmed Mehamed |
Knowl. Based Syst. | 2 |
| 2024 | ResCount: A Residual Feature Fusion Network for Ship Counting in Remote Sensing ImagesabstractShip counting is used to count the number of ships in an image. It has a wide range of research backgrounds in areas such as port management and maritime security. In specific areas such as ports, due to their large number of ships, the ships captured by remote sensing images often have problems of uneven distribution and large differences in ship sizes, which will affect the performance of ship counting. To address the above problems, this letter proposes a residual feature fusion network for ship counting (ResCount). The model first uses a feature extraction network to extract the feature map of the image, and then uses a dual-branch structure to further enhance the feature map. One branch uses a visual encoder module to learn the connection between different regions in the image to improve the problem of decreased counting accuracy in scenes with uneven distribution of ships. However, the visual encoder will lose information such as the outline and texture of the ship. Therefore, the other branch uses a regional context feature fusion module (CAF) proposed in this letter to extract local features of different scales and context features of ships to improve the counting accuracy in scenes with large differences in ship size. In addition, this letter proposes a residual feature fusion (RFF) module to enhance the model’s attention to sparse areas and finally regress to obtain a density map. In addition, we conducted a large number of experiments to verify the method. Finally, on the remote sensing object counting dataset (RSOC), the mean absolute error (MAE) index reached 60.08 and the root mean squared error (RMSE) index reached 79.62. Kai Yan 0001, Yaxiong Chen, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Residual Deformable Convolution for better image de-weathering
Huikai Liu, Wenqian Zhu, Bingjian Ding, Shengwu Xiong 0001 |
Pattern Recognit. | 6 |
| 2024 | CTOD: Cross-Attentive Task-Alignment for One-Stage Object DetectionabstractExisting one-stage object detectors are commonly implemented in a multi-task learning based manner, which simultaneously solves two different sub-tasks: object classification and localization. To achieve this, the detection heads with two independent branches are typically utilized to extract specific image features for each task separately. However, due to the lack of interaction between the parallel branches, the difference in learning objectives of classification and localization will lead to spatial misalignment between the predictions of these two tasks. In this work, we propose a novel Cross-attentive Task-aligned Object Detection (CTOD) method to handle this problem by explicitly promoting the prediction consistency for both tasks. Specifically, we first design a Dual Task Interaction (DTI) module, which generates task-interactive embeddings for each branch from task-specific features by using a task cross-attention mechanism. Then based on these embeddings, we propose a Spatial Feature Aggregation (SFA) module that calculates offsets and weights to aggregate information from nearby feature points at each spatial location of the task-specific feature maps. Meanwhile, we also generate adjustment parameters from the task-interactive embeddings to finally align the prediction results of the two tasks obtained from the enhanced task-specific features described above. Extensive experiments are conducted on the MS-COCO dataset. When using ResNeXt-101-$64\times 4$d-DCN as the backbone, our CTOD method achieves a detection result of 51.8 AP with single-model and single-scale testing, outperforming the recently proposed one-stage detectors ATSS, VFNet, LD and TOOD by 4.1, 1.9, 1.3 and 0.7 AP, respectively. The analysis of qualitative results also illustrates the effectiveness and superiority of CTOD in solving the task misalignment problem for object detection. Our code is available athttps://github.com/Mr-Bigworth/CTOD. Ruilin Yao, Qiangqiang Huang, Shengwu Xiong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Scale-Aware Adaptive Refinement and Cross-Interaction for Remote Sensing Audio-Visual Cross-Modal Retrieval
Yaxiong Chen, Chuang Du, Yunfei Zi, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Multiscale Salient Alignment Learning for Remote-Sensing Image-Text RetrievalabstractRemote-sensing image–text (RSIT) retrieval involves the use of either textual descriptions or remote-sensing images (RSI) as queries to retrieve relevant RSIs or corresponding text descriptions. Many traditional cross-modal RSIT retrieval methods tend to overlook the importance of capturing salient information and establishing the prior similarity between RSIs and texts, leading to a decline in cross-modal retrieval performance. In this article, we address these challenges by introducing a novel approach known as multiscale salient image-guided text alignment (MSITA). This approach is designed to learn salient information by aligning text with images for effective cross-modal RSIT retrieval. The MSITA approach first incorporates a multiscale fusion module and a salient learning module to facilitate the extraction of salient information. In addition, it introduces an image-guided text alignment (IGTA) mechanism that uses image information to guide the alignment of texts, enabling the effective capture of fine-grained correspondences between RSI regions and textual descriptions. In addition to these components, a novel loss function is devised to enhance the similarity across different modalities and reinforce the prior similarity between RSIs and texts. Extensive experiments conducted on four widely adopted RSIT datasets affirm that the MSITA approach significantly enhances cross-modal RSIT retrieval performance in comparison to other state-of-the-art methods. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Thread the Needle: Cues-Driven Multiassociation for Remote Sensing Cross-Modal RetrievalabstractRapid advances in Earth observation technologies have yielded numerous remotely sensed images and corresponding text data, enabling cross-modal image–text retrieval to extract valuable clues. However, current methods often focus on learning global semantic information from text and remote sensing (RS) images, while neglecting fine-grained semantic alignment and correlation. In addition, contrastive learning between modalities is often insufficient. To address these issues, we propose an innovative cues-driven multiassociation feature matching network (CDMAN) for cross-modal RS image retrieval. The proposed method primarily involves two key steps: 1) aligning positive samples and enhancing fusion for negative samples based on modal cues. To achieve precise alignment between RS images and text and facilitate the learning process for negative samples in contrastive learning, we have developed a novel fine-grained cues injection module that aligns and guides modalities using fine-grained cues; and 2) establishing multigranularity associative learning. To address the issue of insufficient association between RS images and text, we have implemented multigranularity collaborative associative learning, focusing on general and fine-grained modal associations. By fully leveraging modal cues, our method maintains both detailed associations and overall consistency in global associations. Experiments demonstrate that, compared to baseline methods, this approach achieves more accurate cross-modal retrieval (MCR) by combining fine-grained alignment and multigranularity associations. Yaxiong Chen, Jirui Huang, Zhaoyang Sun, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Integrating Multisubspace Joint Learning With Multilevel Guidance for Cross-Modal Retrieval of Remote Sensing ImagesabstractIn recent years, with the continuous advancement of remote sensing technology and text processing techniques, there has been a growing abundance of remote sensing images and associated textual data. Combining remote sensing images with their corresponding textual data allows for integrated analysis and retrieval, which holds significant practical implications across multiple application domains, including geographic information systems (GIS), environmental monitoring, and agricultural management. Remote sensing images have the characteristics of multi-targets and multi-scales, and the textual descriptions of these targets are not fully utilized, leading to a decrease in retrieval accuracy. Previous methods have struggled to balance inter-modality information interaction and intra-modality feature fusion, and they have paid little attention to the consistency of distribution within modalities. In light of this, this paper proposes a symmetric multi-level guidance network (SMLGN) for cross-modal retrieval in remote sensing. SMLGN first introduces fusion guidance between local and global within modalities and fine-grained bidirectional guidance between modalities, allowing for the learning of a common semantic space. Furthermore, to address the distribution differences of different modalities within the common semantic space, we design an adversarial joint learning framework and a multi-objective loss function to optimize the SMLGN method and achieve consistency in data distribution. The experimental results demonstrate that the SMLGN method performs well in the task of cross-modal retrieval between remote sensing images and textual data. It effectively integrates the information from both modalities, improving the accuracy and reliability of the retrieval process. Yaxiong Chen, Jirui Huang, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Integrating Detailed Features and Global Contexts for Semantic Segmentation in Ultrahigh-Resolution Remote Sensing ImagesabstractSemantic segmentation of ultrahigh-resolution (UHR) remote sensing images is a fundamental task for many downstream applications. Achieving precise pixel-level classification is paramount for obtaining exceptional segmentation results. This challenge becomes even more complex due to the need to address intricate segmentation boundaries and accurately delineate small objects within the remote sensing imagery. To meet these demands effectively, it is critical to integrate two crucial components: global contextual information and spatial detail feature information. In response to this imperative, the multilevel context-aware segmentation network (MCSNet) emerges as a promising solution. MCSNet is engineered to not only model the overarching global context but also extract intricate spatial detail features, thereby optimizing segmentation outcomes. The strength of MCSNet lies in its two pivotal modules, the spatial detail feature extraction (SDFE) module and the refined multiscale feature fusion (RMFF) module. Moreover, to further harness the potential of MCSNet, a multitask learning approach is employed. This approach integrates boundary detection and semantic segmentation, ensuring that the network is well-rounded in its segmentation capabilities. The efficacy of MCSNet is rigorously demonstrated through comprehensive experiments conducted on two established international society for photogrammetry and remote sensing (ISPRS) 2-D semantic labeling datasets: Potsdam and Vaihingen. These experiments unequivocally establish MCSNet stands as a pioneering solution, that delivers state-of-the-art performance, as evidenced by its outstanding mean intersection over union (mIoU) and mean$F1$-score (mF1) metrics. The code is available at:https://github.com/WUTCM-Lab/MCSNet. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | A Joint Saliency Temporal-Spatial-Spectral Information Network for Hyperspectral Image Change DetectionabstractHyperspectral image change detection (HSI-CD) is a fundamental task in the field of remote sensing (RS) observation, which utilizes the rich spectral and spatial information in bitemporal HSIs to detect subtle changes on the Earth’s surface. However, modern deep learning (DL)-based HSI-CD methods mostly rely on patch-based methods, which leads to spectral band redundancy and spatial information noise in limited receiving domains, thus ignoring the extraction and utilization of saliency information and limiting the improvement of CD performance. To address these issues, this article proposes a joint saliency temporal–spatial–spectral information network (STSS-Net) for HSI-CD. The principal contributions of this article can be summarized: 1) we have designed a spatial saliency information extraction (SSIE) module for denoising based on distance from center pixels and spectral similarity of the substance, which increases the attention to spatial differences between similar spectral substances and different spectral substances; 2) we have designed a compact high-level spectral information tokenizer (CHLSIT) for spectral saliency information, where the high-level conceptual information of changes in spectral interest can be represented by nonlinear combinations of spectral bands, and redundancy can be removed by extracting high-level spectral conceptual features; and 3) utilizing the advantages of CNN and transformer architectures to combine temporal–spatial–spectral information. The experimental results on three real HSI-CD datasets show that STSS-Net can improve the accuracy of CD and has a certain improvement in the detection of edge information and complex information. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Learning Starts From Optimizing the Composition of Temporal Information for Hyperspectral Change DetectionabstractHyperspectral image change detection (HSI-CD) is a task that utilizes both spectral and spatial features to more effectively detect changes. The features of HSI captured at different times are often influenced by external factors. The annoying variability is not useful for CD. Current methods extract effective information through complex learning together with it. They did not focus on whether different categories (changed and unchanged) of sample pairs have the same information composition. They lack a fundamental analysis and optimization based on the composition of information pairs to address current issues such as insufficient recognition of changed samples and poor feature fusion. To address these issues, we rethink the information composition of bitemporal sample pairs in different categories for HSI-CD, and design a frequency-domain information exchange and generation network with Siamese U-shaped structure (FDIEG-UNet) for HSI-CD. Our contribution can be summarized as follows: 1) by designing the FDGJLS module for learning and separating global-joint spectral-spatial features, to provide better basic frequency-domain information for subsequent feature-domain optimization; 2) based on the information from FDGJLS, we use the ULDT and SSTDFE modules to remove the influence of temporal-correlated invalid style information in feature fusion. Two modules work together to improve the equality of feature description between changed and unchanged samples in CD; and 3) the experimental results on three real HSI-CD datasets demonstrate the effectiveness of our proposed method in solving the above problems. The code is available athttps://github.com/WUTCM-Lab/FDIEG-UNet. Yaxiong Chen, Jirui Huang, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Cross-Modal Remote Sensing Image-Audio Retrieval With Adaptive Learning for Aligning CorrelationabstractAn important challenge that existing work has yet to address is the relatively small differences in audio representations compared with the rich content provided by remote sensing (RS) images, making it easy to overlook certain details in the images. This imbalance in information between modalities poses a challenge in maintaining consistent representations. In response to this challenge, we propose a novel cross-modal RS image-audio (RSIA) retrieval method called adaptive learning for aligning correlation (ALAC). ALAC integrates region-level learning into image annotation through a region-enhanced learning attention (RELA) module. By collaboratively suppressing features at different region levels, ALAC is able to provide a more comprehensive visual feature representation. In addition, a novel adaptive knowledge transfer (AKT) strategy has been proposed, which guides the learning process of the frontend network using aligned feature vectors. This approach allows the model to adaptively acquire alignment information during the learning process, thereby facilitating better alignment between the two modalities. Finally, to better use mutual information between different modalities, we introduce a plug-and-play result rerank module. This module optimizes the similarity matrix using retrieval mutual information between modalities as weights, significantly improving retrieval accuracy. Experimental results on four RSIA datasets demonstrate that ALAC outperforms other methods in retrieval performance. Compared with state-of-the-art methods, improvements of 1.49%, 2.25%, 4.24%, and 1.33% were, respectively, achieved by ALAC. The codes are accessible athttps://github.com/huangjh98/ALAC. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Visual Contextual Semantic Reasoning for Cross-Modal Drone Image-Text RetrievalabstractThe cross-modal drone image-text (DIT) retrieval task involves using either text or drone images as queries to retrieve relevant drone images or corresponding text. The primary challenge stems from the diverse and intricate nature of drone images, making effective alignment between image and text challenging. In response, we propose an innovative approach called visual contextual semantic reasoning (VCSR), aimed at precisely aligning information across different modalities. VCSR employs textual cues to guide rich semantic reasoning within the visual context, reducing redundancy in visual information. Furthermore, the method captures drone image information relevant to the text, revealing subtle correspondences between drone image regions and textual content. To enhance visual semantic learning, context region learning (CRL) term and consistency semantic alignment (CSA) terms are introduced for stronger guidance, further intensifying the cross-modal interaction between textual and visual data, resulting in more robust feature representation. Extensive experiments conducted on two self-constructed DIT datasets demonstrate that VCSR outperforms alternative methods in terms of DIT retrieval performance. The codes are accessible athttps://github.com/huangjh98/VCSR. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Oriented Object Detector With Gaussian Distribution Cost Label Assignment and Task-Decoupled HeadabstractRecently, oriented object detection in remote sensing images has garnered significant attention due to its broad range of applications. Early oriented object detection adhered to the established general object detection frameworks, utilizing the label assignment strategy based on the horizontal bounding box annotations or rotation-agnostic cost function. Such strategy may not reflect the large aspect ratio and rotation of arbitrary-oriented objects in remota sensing images and require high parameter-tuning efforts in training process, which will eventually harm the detector performance. Furthermore, the localization quality of oriented object depends on precise rotation angle prediction, exacerbating the inconsistency between classification and regression tasks in oriented object detection. To address these issues, we propose the Gaussian Distribution Cost Optimal Transport Assignment (GCOTA) and Decoupled Layer Attention Angle Head (DLAAH). Specifically, GCOTA utilize Gaussian distribution based cost function for the optimal transport label assignment in training process, alleviating the impact of rotation angle and large aspect ratio in remote sensing images. DLAAH predicts rotation angle independently and incorporates layer attention to obtain the task-specific features based on the shared FPN features, enhancing the angle prediction and improving consistency across different tasks. Based on these proposed components, we present an anchor-free oriented detector, namely Gaussian Distribution and Task-Decoupled head oriented Detector(GTDet) and a a multi-class ship detection dataset in real scenarios (CGWX), which provides a benchmark for fine-grained object recognition in remote sensing images. Comprehensive experiments are conducted on CGWX and several public challenging datasets, including DOTAv1.0, HRSC2016, to demonstrate that our method achieves superior performance on oriented object detection task. The code is available at https://github.com/WUTCM-Lab/GTDet. Qiangqiang Huang, Ruilin Yao, Xiaoqiang Lu, Jishuai Zhu, Shengwu Xiong 0001, Yaxiong Chen |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | SANet: A Self-Attention Network for Agricultural Hyperspectral Image ClassificationabstractUnlike conventional hyperspectral image (HSI) classification in general scenes, agricultural HSI classification poses greater challenges due to the increased occurrence of “same spectrum different object” and “different spectrum same object” phenomena caused by class similarities. Furthermore, the dense spatial distribution of land cover categories in agricultural scenes and the mixing of spatial–spectral features at crop boundaries add to the complexity of agricultural HSIs. To tackle these issues, we propose SANet, a network designed to enhance crop classification. SANet integrates spectral and contextual information while emphasizing self-correlation within the HSIs. It combines the spatial–spectral nonlocal block structure and the multiscale spectral self-attention (SSA) structure, allocating more attention resources to spatial and spectral dimensions and modeling the existing correlations within the spectral–spatial domain. Additionally, we introduce a two-branch spatial–spectral semantic extraction and fusion structure that can adaptively learn results from both branches. Experimental results demonstrate the promising performance of SANet in agricultural HSI classification by effectively utilizing spectral data, contextual information, and self-attention mechanisms. Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Global-Group Attention Network With Focal Attention Loss for Aerial Scene ClassificationabstractAerial scene classification, aiming at assigning a specific semantic class to each aerial image, is a fundamental task in the remote sensing community. Aerial scene images have more diverse and complex geological features. While some statistics of images can be well fit using convolution, it limits such models to capturing the global context hidden in aerial scenes. Furthermore, to optimize the feature space, many methods add class information to the feature embedding space. However, they seldom combine model structure with class information to obtain more separable feature representations. In this article, we propose to address these limitations in a unified framework (i.e., CGFNet) from two aspects: focusing on the key information of input images and optimizing the feature space. Specifically, we propose a global-group attention module (GGAM) to adaptively learn and selectively focus on important information from input images. GGAM consists of two parallel branches: the adaptive global attention branch (AGAB) and the region-aware attention branch (RAAB). AGAB utilizes an adaptive pooling operation to better model the global context in aerial scenes. As a supplement to AGAB, RAAB combines grouping features with spatial attention to spatially enhance the semantic distribution of features (i.e., selectively focus on effective regions of features and ignore irrelevant semantic regions). In parallel, a focal attention loss (FA-Loss) is exploited to introduce class information into attention vector space, which can improve intraclass consistency and interclass separability. Experimental results on four publicly available and challenging datasets demonstrate the effectiveness of our method. The source code will be released at:https://github.com/zoecheno/CGFNet. Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Co-Enhanced Global-Part Integration for Remote-Sensing Scene ClassificationabstractRemote sensing (RS) scene classification aims to classify remote sensing images with similar scene characteristics into one category. Plenty of RS images are complex in background, rich in content, and multi-scale in target, exhibiting the characteristics of both intra-class separation and inter-class convergence. Therefore, discriminative feature representations designed to highlight the differences between classes are the key to RS scene classification. Existing methods represent scene images by extracting either global context or discriminative part features from RS images. However, global-based methods often lack salient details in similar RS scenes, while part-based methods tend to ignore the relationships between local ground objects, thus weakening the discriminative feature representation. In this paper, we propose to combine global context and part-level discriminative features within a unified framework called CGINet for accurate RS scene classification. To be specific, we develop a light context-aware attention block (LCAB) to explicitly model the global context to obtain larger receptive fields and contextual information. A co-enhanced loss module (CELM) is also devised to encourage the model to actively locate discriminative parts for feature enhancement. In particular, CELM is only used during training and not activated during inference, which introduces less computational cost. Benefiting from LCAB and CELM, our proposed CGINet improves the discriminability of features, thereby improving classification performance. Comprehensive experiments over four benchmark datasets show that the proposed method achieves consistent performance gains over state-of-the-art RS scene classification methods. Yichen Zhao, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu, Xiao Xiang Zhu 0001, Lichao Mou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | An Efficient Message Dissemination Scheme for Cooperative Drivings via Cooperative Hierarchical Attention Reinforcement LearningabstractA group ofconnected and autonomous vehicleswith common interests can drive in a cooperative manner, namely cooperative driving. In such a networked control system, an efficient message dissemination scheme is critical for cooperative drivings to periodically broadcast their kinetic status, i.e.,beacon. However, most existing researches are designed for a simple or specific scenario, e.g., ignoring the impacts of the complex communication environment and emerging hybrid traffic scenarios. Worse still, the inevitable message transmission interference and the limited interaction among vehicles in harsh communication environments seriously hinder cooperation among cooperative drivings and deteriorate the beaconing performance. In this paper, we formulate the decision-making process of cooperative drivings as a Markov game. Furthermore, we propose acooperative hierarchical attention reinforcement learning (CHA)framework to solve this Markov game. Specifically, the hierarchical structure of CHA leads cooperative drivings to be foresighted. Besides, we integrate each hierarchical level of CHA separately with graph attention networks to incorporate agents' mutual influences in the decision-making process. Moreover, each hierarchical level learns a cooperative reward function to motivate each agent to cooperate with others under harsh communication conditions. Finally, we set up a simulator and conduct extensive experiments to validate the effectiveness of CHA. Bingyi Liu, Weizhen Han, Enshu Wang, Shengwu Xiong 0001, Chunming Qiao, Jianping Wang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Multi-Agent Attention Double Actor-Critic Framework for Intelligent Traffic Light Control in Urban Scenarios With Hybrid TrafficabstractIn real-world urban environments, hybrid and disorder traffic brings new challenges for the intelligent traffic light control system (ITLCS). Apart from coordinating traffic flows around intersections, the ITLCS is responsive to ensuring high priority vehicles pass through intersections quickly. To this end, we formulate the multiple intersections’ decision-making problem as a Semi-Markov game and propose amulti-agent attention double actor-critic (MAADAC)framework to solve this game, integrating theoptions frameworkwithgraph attention networks (GATs). Specifically, the options framework empowers agents to learn to make a long sequence of satisfactory decisions, such as keeping a reasonable phase for a short period to ensure high priority vehicles pass through intersections quickly. Besides, we adopt GATs to capture graph-structure mutual influences among agents. We set up a simulator based on real-world city road networks and conduct extensive experiments to evaluate the performance of MAADAC. The experimental results show that MAADAC can reduce high priority vehicles’ waiting time in the interval of 18.16%-38.14% versus the density of vehicles in real-world urban scenarios over several state-of-the-art approaches. Also, our framework can guarantee the passing efficiency of high priority vehicles under various traffic conditions with the change in the proportion of high priority vehicles. Bingyi Liu, Weizhen Han, Enshu Wang, Shengwu Xiong 0001, Qian Wang 0002, Jianping Wang 0001, Chunming Qiao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Orthogonal integral transform for 3D shape recognition with few examples
Chengde Lin, Shengwu Xiong 0001, Ruyi Chen |
Vis. Comput. | 3 |
| 2024 | Reversible data-hiding exploiting huffman encoding in dual image using weighted matrix and generalized exploiting modification direction (GEMD)
Nada Hussien Abd El Salam, Shengwu Xiong 0001 |
Vis. Comput. | 2 |
| 2023 | ESPT: A Self-Supervised Episodic Spatial Pretext Task for Improving Few-Shot LearningabstractSelf-supervised learning (SSL) techniques have recently been integrated into the few-shot learning (FSL) framework and have shown promising results in improving the few-shot image classification performance. However, existing SSL approaches used in FSL typically seek the supervision signals from the global embedding of every single image. Therefore, during the episodic training of FSL, these methods cannot capture and fully utilize the local visual information in image samples and the data structure information of the whole episode, which are beneficial to FSL. To this end, we propose to augment the few-shot learning objective with a novel self-supervised Episodic Spatial Pretext Task (ESPT). Specifically, for each few-shot episode, we generate its corresponding transformed episode by applying a random geometric transformation to all the images in it. Based on these, our ESPT objective is defined as maximizing the local spatial relationship consistency between the original episode and the transformed one. With this definition, the ESPT-augmented FSL objective promotes learning more transferable feature representations that capture the local spatial features of different images and their inter-relational structural information in each input episode, thus enabling the model to generalize better to new categories with only a few samples. Extensive experiments indicate that our ESPT method achieves new state-of-the-art performance for few-shot image classification on three mainstay benchmark datasets. The source code will be available at: https://github.com/Whut-YiRong/ESPT. Xiongbo Lu, Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong 0001 |
AAAI | 5 |
| 2023 | CF-VTON: Multi-Pose Virtual Try-on with Cross-Domain FusionabstractThe multi-pose virtual try-on technology aims to seamlessly fit an in-shop garment onto a reference person in various poses. This technology has attracted considerable attention from researchers due to its potential commercial and practical applications. Previous works in this field have encountered issues such as unnatural garment alignment and difficulty in preserving the person’s identity, arising from the weak mapping relationship between different feature crosses. To address these challenges, this paper proposes a novel multi-pose virtual try-on network named CF-VTON. Our approach involves predicting the "after-try-on" semantic map to guide garment alignment and try-on synthesis, warping the garment using an improved garment alignment network (GANet) to optimize unnatural alignment, synthesizing a coarse result with our proposed try-on synthesis network (TSN), and refining the output to reconstruct the virtual try-on result with rich facial identity and garment details. Qualitative and quantitative experiments demonstrate the superiority of our approach, outperforming state-of-the-art methods in an efficient manner. Chenghu Du, Shengwu Xiong 0001 |
ICASSP | 2 |
| 2023 | NREE: Nested Relation Extraction in the Economic FieldabstractThe traditional relation extraction (RE) task is to identify whether there is a relation between two entities in a given sentence and determine their relation types. However, especially in professional documents in fields like economics and finance, they often involve nested relations, where relation triples can serve as components of higher-level relations. In this paper, we explore the problem of nested relations within the context of the economic field. To address the task of extracting nested relations in the economic field, we first construct a specialized dataset that is specifically designed for nested relation extraction within the economic domain. Subsequently, we define a representation model for nested relations and introduce the concept of Missing Entity Relations within the dataset. Furthermore, we propose a novel relation extraction model, the BERT-BiLSTM-Transformer network (BBTN), capable of extracting both flat and nested relations. We experimentally evaluate our model on the constructed economy dataset and demonstrate its superior performance compared to the baseline model. Pengfei Duan 0005, Aoni Wu, Wenyan Hu 0003, Shengwu Xiong 0001 |
ICPADS | 5 |
| 2023 | DLFusion: Painting-Depth Augmenting-LiDAR for Multimodal Fusion 3D Object DetectionabstractSurround-view cameras combined with image depth transformation to 3D feature space and fusion with point cloud features are highly regarded. The transformation of 2D features into 3D feature space by means of predefined sampling points and depth distribution happens throughout the scene, and this process generates a large number of redundant features. In addition, multimodal feature fusion unified in 3D space often happens in the previous step of the downstream task, ignoring the interactive fusion between different scales. To this end, we design a new framework, focusing on the design that can give 3D geometric perception information to images and unify them into voxel space to accomplish multi-scale interactive fusion, and we mitigate feature alignment between modal features by geometric relationships between voxel features. The method has two main designs. First, a Segmentation-guided Image View Transformation module is used to accurately transform the pixel region containing the object into a 3D pseudo-point voxel space with the help of a depth distribution. This allows subsequent feature fusion to be performed in a unified voxel feature. Secondly, a Voxel-centric Consistent Fusion module is used to alleviate the errors caused by depth estimation, as well as to achieve better feature fusion between unified modalities. Through extensive experiments on the KITTI and nuScenes datasets, we validate the effectiveness of our camera-LIDAR fusion method. Our proposed approach shows competitive performance on both datasets and outperforms state-of-the-art methods in certain classes of 3D object detection benchmarks. https://github.com/no-Name128/DLFusion [code release] Junyin Wang, Chenghu Du, Hui Li 0010, Shengwu Xiong 0001 |
ACM Multimedia | 4 |
| 2023 | Greatness in Simplicity: Unified Self-Cycle Consistency for Parser-Free Virtual Try-OnabstractImage-based virtual try-on tasks remain challenging, primarily due to inherent complexities associated with non-rigid garment deformation modeling and strong feature entanglement of clothing within human body. Recent groundbreaking formulations, such as in-painting, cycle consistency, and knowledge distillation, have facilitated self-supervised generation of try-on images. However, these paradigms necessitate the disentanglement of garment features within human body features through auxiliary tasks, such as leveraging 'teacher knowledge' and dual generators. The potential presence of irresponsible prior knowledge in the auxiliary task can serve as a significant bottleneck for the main generator (e.g., 'student model') in the downstream task. Moreover, existing garment deformation methods lack the ability to perceive the correlation between the garment and the human body in the real world, leading to unrealistic alignment effects. To tackle these limitations, we present a new parser-free virtual try-on network based on unified self-cycle consistency (USC-PFN), which enables robust translation between different garments using just a single generator, faithfully replicating non-rigid geometric deformation of garments in real-life scenarios. Specifically, we first propose a self-cycle consistency architecture with a circular mode. It utilizes real unpaired garment-person images exclusively as input for training, effectively eliminating the impact of irresponsible prior knowledge at the model input end. Additionally, we formulate a Markov Random Field to simulate a more natural and realistic garment deformation. Furthermore, USC-PFN can leverage a general generator for self-supervised cycle training. Experiments demonstrate that our method achieves state-of-the-art performance on a popular virtual try-on benchmark. Chenghu Du, Junyin Wang, Shuqing Liu, Shengwu Xiong 0001 |
NeurIPS | 4 |
| 2023 | Improving Speaker Recognition by Time-Frequency Domain Feature Enhanced Method
Yunfei Zi, Shengwu Xiong 0001 |
PRICAI (2) | 3 |
| 2023 | LMGFuse: Language Models and Graph reasoning Fuse deeply for question answeringabstractThe combination of pre-trained language models (LM) and knowledge graphs (KG) can enhance the reasoning ability for Question Answering.However, previous methods typically fuse the two modalities in a shallow or knowledgedraining manner, not taking full advantage of the knowledge representation of both.How to effectively fuse the different knowledge representations is still a problem of current research.In our work, a novel model is proposed that fuses LM modal knowledge representations and graph neural network (GNN) modal knowledge representations deeply over multiple layers of modality interaction operations.Specifically, the model includes an information interaction unit, through which KG and LM knowledge can be transferred between modalities to realize knowledge fusion directly, reducing information loss.In addition, we add the context node of implicit knowledge from LM encoding in the construction of the reasoning subgraph in advance for enhancing the reasoning of the GNN.We evaluate our model on two domains in the biomedical benchmark (MedQA-USMLE) and commonsense benchmarks (OpenBookQA and CommonsenseQA).Experimental results show that our model achieves a particular improvement over existing LM and LM+KG models for reasoning over both situational constraints and structured knowledge. Aoxing Wang, Pengfei Duan 0005, Yongbing Li, Wenyan Hu 0003, Shengwu Xiong 0001 |
SEKE | 5 |
| 2023 | Incorporating higher order network structures to improve miRNA-disease association prediction based on functional modularityabstractAs microRNAs (miRNAs) are involved in many essential biological processes, their abnormal expressions can serve as biomarkers and prognostic indicators to prevent the development of complex diseases, thus providing accurate early detection and prognostic evaluation. Although a number of computational methods have been proposed to predict miRNA-disease associations (MDAs) for further experimental verification, their performance is limited primarily by the inadequacy of exploiting lower order patterns characterizing known MDAs to identify missing ones from MDA networks. Hence, in this work, we present a novel prediction model, namely HiSCMDA, by incorporating higher order network structures for improved performance of MDA prediction. To this end, HiSCMDA first integrates miRNA similarity network, disease similarity network and MDA network to preserve the advantages of all these networks. After that, it identifies overlapping functional modules from the integrated network by predefining several higher order connectivity patterns of interest. Last, a path-based scoring function is designed to infer potential MDAs based on network paths across related functional modules. HiSCMDA yields the best performance across all datasets and evaluation metrics in the cross-validation and independent validation experiments. Furthermore, in the case studies, 49 and 50 out of the top 50 miRNAs, respectively, predicted for colon neoplasms and lung neoplasms have been validated by well-established databases. Experimental results show that rich higher order organizational structures exposed in the MDA network gain new insight into the MDA prediction based on higher order connectivity patterns. Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Shengwu Xiong 0001, Lun Hu |
Briefings Bioinform. | 5 |
| 2023 | Federated clustering for recognizing driving styles from private trajectories
Lin Lu 0002, Yuan Wen, Jinxiong Zhu, Shengwu Xiong 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Joint filter combination-based central difference feature extraction and attention-enhanced Dense-Res2Block network for short-utterance speaker recognition
Yunfei Zi, Shengwu Xiong 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Rethinking one-shot face reenactment: A spatial-temporal reconstruction view
Shengwu Xiong 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Exploration of multi-source discriminative acoustic feature for speaker recognition with short-duration audio signal
Yunfei Zi, Shengwu Xiong 0001 |
Multim. Tools Appl. | 2 |
| 2023 | A novel framework for message dissemination with consideration of destination prediction in VFC
Bingyi Liu, Enshu Wang, Shengwu Xiong 0001 |
Neural Comput. Appl. | 7 |
| 2023 | A Lie algebra representation for efficient 2D shape classification
Xiaohan Yu 0001, Yongsheng Gao 0001, Mohammed Bennamoun, Shengwu Xiong 0001 |
Pattern Recognit. | 4 |
| 2023 | Information bottleneck disentanglement based sparse representation for fair classification
Xiongbo Lu, Yaxiong Chen, Shengwu Xiong 0001 |
Pattern Recognit. Lett. | 4 |
| 2023 | Aerial image recognition in discriminative bi-transformer
Yichen Zhao, Yaxiong Chen, Xiongbo Lu, Lei Zhou 0008, Shengwu Xiong 0001 |
Signal Process. | 5 |
| 2023 | BSML: Bidirectional Sampling Aggregation-based Metric Learning for Low-resource Uyghur Few-shot Speaker VerificationabstractIn recent years, text-independent speaker verification has remained a hot research topic, especially for the limited enrollment and/or test data. At the same time, due to the lack of sufficient training data, the study of low-resource few-shot speaker verification makes the models prone to overfitting and low accuracy of recognition. Therefore, a bidirectional sampling aggregation-based meta-metric learning method is proposed to solve the low-accuracy problem of speaker recognition in a low-resource environment with limited data, termed bidirectional sampling multi-scale Fisher feature fusion (BSML). First, the BSML method was used for effective feature enhancement in the feature extraction stage; second, a large number of similar and disjoint tasks were used to train the models to learn how to compare sample similarity; finally, new tasks were used to identify unknown samples by calculating the similarity of the samples. Extensive experiments are conducted on a short-duration text-independent speaker verification dataset generated from the THUYG-20 low-resource Uyghur with limited data, which comprised speech samples of diverse lengths. The experimental result has shown that the metric learning approach is effective in avoiding model overfitting and improving model generalization, with significant results in the identification of short-duration speaker verification in low-resource Uyghur with few-shot. It also demonstrates that BSML outperforms the state-of-the-art deep-embedding speaker recognition architectures and recent metric learning approach by at least 18%–67% in the few-shot test set. The ablation experiments further illustrate that our proposed approaches can achieve substantial improvement over prior methods and achieves better performance and generalization ability. Yunfei Zi, Shengwu Xiong 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Deep Saliency Smoothing Hashing for Drone Image RetrievalabstractDeep hashing algorithms are widely exploited in retrieval tasks due to its low storage and retrieval efficiency. Most of which focus on global feature learning, whilst neglecting local fine-grained features and saliency information for drone images. In this paper, we tackle these dilemmas with a novelDeep Saliency Smoothing Hashing(DSSH) algorithm, which can leverage saliency capture mechanism, distribution smoothing term, global features and local fine-grained features to learn effective hash codes for drone image retrieval. The DSSH algorithm first designs information extraction module to capture global features and local fine-grained features for drone images. Meanwhile, a saliency capture module is proposed to perform information interaction attention and visual enhancement attention, which can capture the saliency area of drone images effectively. On top of the two paths, a novel objective function is designed to preserve the similarity of hash codes, smooth the distribution of drone image datasets and reduce the quantization errors between hash codes and hash-like codes concurrently. Extensive experiments on the Drone Action Dataset and ERA Drone Dataset demonstrate that the DSSH algorithm can further improve the retrieval performance compared to other deep hashing algorithms. Yaxiong Chen, Lichao Mou, Pu Jin, Shengwu Xiong 0001, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Fine Aligned Discriminative Hashing for Remote Sensing Image-Audio RetrievalabstractFor cross-modal remote sensing image-audio retrieval task, hashing technology has attracted much attention in recent works. Most of them focus on mappingRemote Sensing(RS) images and audios into a Hamming space, whilst neglecting discriminative information of RS images and fine alignment for RS images and audios. In this paper, we tackle these dilemmas with a novelFine Aligned Discriminative Hashing(FADH) approach, which can learn hash codes to capture discriminative information of RS images and learn the corresponding detailed information between RS images and audios simultaneously. We first develop a new discriminative information learning module to learn discriminative information of RS images. Meanwhile, a fine alignment module is proposed to unearth the fine correspondence for RS image regions and audios, which can effectively improve the retrieval performance. On top of the two paths, we design a new objective function, which can maintain the similarity of hash codes, preserve the semantic information of RS image features and audio features and eliminate cross-modal differences. The reliability and significance of the designed framework are effectively demonstrated by diverse experiments on three remote sensing image-audio datasets. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Self-Supervision Interactive Alignment for Remote Sensing Image-Audio RetrievalabstractCross-modal remote sensing image-audio retrieval aims to use audio or remote sensing images as queries to retrieve relevant remote sensing images or corresponding audios. Although many approaches leverage labeled samples to achieve good performance, the performance cost of labeled samples is high, because cross-modal remote sensing labeled samples usually requires huge labor resources. Therefore, unsupervised cross-modal learning is very important in real-world applications. In this paper, we propose a novel unsupervised cross-modal remote sensing image-audio retrieval approach, namedSelf-Supervision Interactive Alignment(SSIA), which can take advantage of large amounts of unlabeled samples to learn the salient information, cross-modal alignment and the similarity between remote sensing images and audios. Since self-supervised learning lacks the supervision of label information, we leverage the similarity between the input remote sensing image information and audio information as the supervision information. Besides, to perform cross-modal alignment, a novel interactive alignment module is designed to explore fine correspondence relation for remote sensing images and audios. Moreover, we design an audio guided image de-redundant module to reduce the redundant information of visual information, which can capture salient information of remote sensing images. Extensive experiments on four widely-used remote sensing image-audio datasets testify that the SSIA perform gain better remote sensing image-audio retrieval performance than other compared approaches. Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | MATNet: A Combining Multi-Attention and Transformer Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) has rich spatial-spectral information, high spectral correlation and large redundancy between information. Due to the sparse background distribution of HSI, existing methods generally perform poorly for the classification of class pixels located in the boundary areas of land cover categories. This is largely because the network is vulnerable to surrounding redundant information during the training stage, leading to inaccurate feature extraction and thus poor generalization ability of the model. Based on previous work, we propose a HSI classification network called MATNet which combines multi-attention and Transformer. The network first uses spatial attention and channel attention to pay more attention to the more significant information parts, then uses tokenizer module to make a semantic level representation of different categories of ground objects, and then performs deep semantic feature extraction using the transformer encoder module. Finally, we design a loss function called Lpoly, which adds a polynomial to the label smoothing loss to tune the original first polynomial to accommodate different datasets and tasks. We perform experiments in several well-known HSI datasets as well as for visualization. The results show that our proposed MATNet performs well in extracting spatial-spectral features of HSIs as well as understanding semantic degrees of semantic degrees. Bo Zhang 0069, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Disentangled face editing via individual walk in personalized facial semantic field
Chengde Lin, Shengwu Xiong 0001, Xiongbo Lu |
Vis. Comput. | 2 |
| 2022 | SSAT: A Symmetric Semantic-Aware Transformer Network for Makeup Transfer and RemovalabstractMakeup transfer is not only to extract the makeup style of the reference image, but also to render the makeup style to the semantic corresponding position of the target image. However, most existing methods focus on the former and ignore the latter, resulting in a failure to achieve desired results. To solve the above problems, we propose a unified Symmetric Semantic-Aware Transformer (SSAT) network, which incorporates semantic correspondence learning to realize makeup transfer and removal simultaneously. In SSAT, a novel Symmetric Semantic Corresponding Feature Transfer (SSCFT) module and a weakly supervised semantic loss are proposed to model and facilitate the establishment of accurate semantic correspondence. In the generation process, the extracted makeup features are spatially distorted by SSCFT to achieve semantic alignment with the target image, then the distorted makeup features are combined with unmodified makeup irrelevant features to produce the final result. Experiments show that our method obtains more visually accurate makeup transfer results, and user study in comparison with other state-of-the-art makeup transfer methods reflects the superiority of our method. Besides, we verify the robustness of the proposed method in the difference of expression and pose, object occlusion scenes, and extend it to video makeup transfer. Zhaoyang Sun, Yaxiong Chen, Shengwu Xiong 0001 |
AAAI | 3 |
| 2022 | A Frame Loss of Multiple Instance Learning for Weakly Supervised Sound Event DetectionabstractSound event detection(SED) consists of two subtasks: predicting the classes of sound events within an audio clip (audio tagging) and indicating the onset and offset times for each event (localization). One of the common approaches for SED with weak label is multiple instance learning (MIL) method. However, the general MIL method only optimizes the global loss calculated from the aggregated clip-wise predictions and weak clip labels, lacking a direct constraint on the frame-wise predictions, which leads to a large number of unreasonable prediction values. To address this issue, we explore the deterministic information that can be used to constrain the framewise predictions and based on which we design a frame loss with two terms. Experimental results on the DCASE2017 Task4 dataset demonstrate that the proposed loss can improve the performance of general MIL method. While this article focuses on SED applications, the proposed methods could be applied widely to MIL problems. Code will be available at WSSED. Xiangjinzi Zhang, Yunfei Zi, Shengwu Xiong 0001 |
ICASSP | 4 |
| 2022 | Feature Space Disentangling Based on Spatial Attention for Makeup TransferabstractMakeup transfer aims at rendering the makeup style from a given reference image to a source image. Most existing works have achieved promising progress by disentangled representation. However, these methods do not consider the spatial distribution of makeup style, which inevitably change the makeup-irrelevant regions. To solve the problem, we introduce a novel feature space disentangling framework based on spatial attention mechanism for makeup transfer. In particular, we first utilize a single encoder to extract all the features of the image. Then we propose a learnable spatial semantic classifier to classify the extracted features into makeup-specific and makeup-irrelevant features. Finally, we complete makeup transfer by swapping the classified features. Experiments demonstrate that the makeup-specific features precisely signify the spatial distribution of makeup style. The superiority of our approach is well demonstrated by the experiment that it produces promising visual results and keeps those makeup-irrelevant regions unchanged. Jinli Zhou, Yaxiong Chen, Zhaoyang Sun, Chang Zhan, Shengwu Xiong 0001 |
ICIP | 6 |
| 2022 | Fusing Acoustic and Text Emotional Features for Expressive Speech SynthesisabstractProminent methods based on Tacotron2 and advanced models have improved the quality of synthesized speech. However, most data-driven Text- To-Speech (TTS) synthesis methods only aim to achieve reasonable neutral prosody, so the synthesized speech is less expressive. In this paper, a method was proposed which fuses acoustic and text emotional features to produce more vivid and realistic speech. Specifically, to obtain acoustic features, two acoustic encoders are leveraged to extract utterance-level and phoneme-level vectors from the target speech, respectively. To obtain the objective sentiment features of the text, the sentiment analysis model is exploited to extract the sentiment vector from the text and expand it. The expanded vector is feature- fused with the output vector of the acoustic model. The experimental results on the LJSpeech dataset show that the naturalness and expressiveness of the MOS score are 3.63 and 3.45, respectively, and the similarity of the SMOS score is 4.14. Pengfei Duan 0005, Yunfei Zi, Yaxiong Chen, Shengwu Xiong 0001 |
ICME | 5 |
| 2022 | Deep Semantic Ranking Hashing Based on Self-Attention for Medical Image RetrievalabstractWith the rapid progress of medical image technology, medical image retrieval has attracted wide attention in medical data processing fields. Deep hashing methods have been proven effective for massive medical image retrieval. However, existing medical image retrieval methods ignore lesion context and category-level semantics, so it is difficult to correctly correspond to the context information and category of the lesion, resulting in poor performance. This paper addresses this dilemma with a novel Deep Semantic Ranking Hashing Based on Self-Attention (DSHA) approach. We first divide the medical triplet into smaller patches and send them to the multi-head self-attention module, which can more effectively encode the context information of the lesion and the interaction among patches. Meanwhile, the weight-sharing triplet networks are used to learn hash codes, which are forced to be semantically aligned. The proposed semantic enhancement loss effectively enhances the category-level semantics of hash codes. In addition, we added a semantic ranking penalty loss to optimize the retrieval accuracy. Extensive experiments on diverse medical image datasets prove that our DSHA method achieves remarkable results compared with the state-of-the-art medical image retrieval methods. Yibo Tang, Yaxiong Chen, Shengwu Xiong 0001 |
ICPR | 3 |
| 2022 | Query and Neighbor-Aware Reasoning Based Multi-hop Question Answering over Knowledge Graph
Shengwu Xiong 0001 |
KSEM (1) | 3 |
| 2022 | Text Style Transfer based on Multi-factor Disentanglement and MixtureabstractText style transfer aims to transfer the reference style of one text image to another text image. Previous works have only been able to transfer the style to a binary text image. In this paper, we propose a framework to disentangle the text images into three factors: text content, font, and style features, and then remix the factors of different images to transfer a new style. Both the reference and input text images have no style restrictions. Adversarial training through multi-factor cross recognition is adopted in the network for better feature disentanglement and representation. To decompose the input text images into a disentangled representation with swappable factors, the network is trained using similarity mining within pairs of exemplars. To train our model, we synthesized a new dataset with various text styles in both English and Chinese. Several ablation studies and extensive experiments on our designed and public datasets demonstrate the effectiveness of our approach for text style transfer. Anna Zhu, Zhanhui Yin, Brian Kenji Iwana, Shengwu Xiong 0001 |
ACM Multimedia | 5 |
| 2022 | BSB: Bringing Safe Browsing to Blockchain Platform
Rongwei Yu, Siwei Wu, Shengwu Xiong 0001 |
NSS | 6 |
| 2022 | Few-shot driver identification via meta-learning
Lin Lu 0002, Shengwu Xiong 0001 |
Expert Syst. Appl. | 2 |
| 2022 | Controllable face editing for video reconstruction in human digital twins
Chengde Lin, Shengwu Xiong 0001 |
Image Vis. Comput. | 2 |
| 2022 | Mutual information maximizing GAN inversion for real face with identity preservation
Chengde Lin, Shengwu Xiong 0001, Yaxiong Chen |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | On the form of parsed sentences for relation extraction
Mi Zhang 0006, Shengwu Xiong 0001, Tieyun Qian |
Knowl. Based Syst. | 3 |
| 2022 | Adaptive PVD-MPK encoding method in transform domain with a modified side match method against dynamic and constant attacks
Nada Hussien Abd El Salam, Ahmed Hamdy Ismael, Shengwu Xiong 0001 |
Multim. Tools Appl. | 3 |
| 2022 | The structural weight design method based on the modified grasshopper optimization algorithm
Yin Ye, Shengwu Xiong 0001, Chen Dong 0002 |
Multim. Tools Appl. | 2 |
| 2022 | Simple Extensible Deep Learning Model for Automatic Arabic DiacritizationabstractAutomatic diacritization is an Arabic natural language processing topic based on the sequence labeling task where the labels are the diacritics and the letters are the sequence elements. A letter can have from zero up to two diacritics. The dataset used was a subset of the preprocessed version of the Tashkeela corpus. We developed a deep learning model composed of a stack of four bidirectional long short-term memory hidden layers of the same size and an output layer at every level. The levels correspond to the groups that we classified the diacritics into (short vowels, double case-endings, Shadda, and Sukoon). Before training, the data were divided into input vectors containing letter indexes and outputs vectors containing the indexes of diacritics regarding their groups. Both input and output vectors are concatenated, then a sliding window operation with overlapping is performed to generate continuous and fixed-size data. Such data is used for both training and evaluation. Finally, we realize some tests using the standard metrics with all of their variations and compare our results with two recent state-of-the-art works. Our model achieved 3% diacritization error rate and 8.99% word error rate when including all letters. We have also generated the confusion matrix to show the performances per output and analyzed the mismatches of the first 500 lines to classify the model errors according to their linguistic nature. Hamza Abbad, Shengwu Xiong 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Local R-Symmetry Co-Occurrence: Characterising Leaf Image Patterns for Identifying CultivarsabstractLeaf image recognition techniques have been actively researched for plant species identification. However it remains unclear whether analysing leaf patterns can provide sufficient information for further differentiating cultivars. This paper reports our attempt on cultivar recognition from leaves as a general very fine-grained pattern recognition problem, which is not only a challenging research problem but also important for cultivar evaluation, selection and production in agriculture. We propose a novel local R-symmetry co-occurrence method for characterising discriminative local symmetry patterns to distinguish subtle differences among cultivars. Through scalable and moving R-relation radius pairs, we generate a set of radius symmetry co-occurrence matrices (RsCoM)and their measures for describing the local symmetry properties of interior regions. By varying the size of the radius pair, the RsCoM measures local R-symmetry co-occurrence from global/coarse to fine scales. A new two-phase strategy of analysing the distribution of local RsCoM measures is designed to match the multiple scale appearance symmetry pattern distributions of similar cultivar leaf images. We constructed three leaf image databases, SoyCultivar, CottCultivar, and PeanCultivar, for an extensive experimental evaluation on recognition across soybean, cotton and peanut cultivars. Encouraging experimental results of the proposed method in comparison with the state-of-the-art leaf species recognition methods demonstrate the effectiveness of the proposed method for cultivar identification, which may advance the research in leaf recognition from species to cultivar. Bin Wang 0041, Yongsheng Gao 0001, Shengwu Xiong 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | Deep Quadruple-Based Hashing for Remote Sensing Image-Sound RetrievalabstractWith the rapid progress of earth observation technology, cross-modal remote sensing (RS) image-sound retrieval has attracted much attention from the field of RS data processing. Existing approaches usually learn the pairwise similarity relations between RS images and sounds. However, these approaches ignore relative semantic similarity relationships, which leads to poor performance of cross-modal RS image-sound retrieval. In this article, we address this dilemma with a noveldeep quadruple-based hashing(DQH) approach. We first devise a novel quadruple-based hashing network to learn relative semantic similarity relationships of hash codes. Meanwhile, we propose a quadruple construction hard module, which randomly selects two triplet hard units to directly learn relative semantic similarity relationships. On top of the two paths, we develop a new objective function to perform effective hash codes learning. The new objective function not only captures the relative semantic correlation of hash codes across different modalities and learns the relative semantic correlation of deep features but also enhances category-level semantics of hash codes and reduces the quantization error between hash-like codes and hash codes. The reasonableness and effectiveness of the proposed architecture are well illustrated by comprehensive experiments on diverse RS image-sound datasets. Yaxiong Chen, Shengwu Xiong 0001, Lichao Mou, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Optimizing Living Material Delivery During the COVID-19 OutbreakabstractThe coronavirus disease 2019 (COVID-19) epidemic has spread worldwide, posing a great threat to human beings. The stay-home quarantine is an effective way to reduce physical contacts and the associated COVID-19 transmission risk, which requires the support of efficient living materials (such as meats, vegetables, grain, and oil) delivery. Notably, the presence of potential infected individuals increases the COVID-19 transmission risk during the delivery. The deliveryman may be the medium through which the virus spreads among urban residents. However, traditional delivery route optimization methods don't take the virus transmission risk into account. Here, we propose a novel living material delivery route approach considering the possible COVID-19 transmission during the delivery. A complex network-based virus transmission model is developed to simulate the possible COVID-19 infection between urban residents and the deliverymen. A bi-objective model considering the COVID-19 transmission risk and the total route length is proposed and solved by the hybrid meta-heuristics integrating the adaptive large neighborhood search and simulated annealing. The experiment was conducted in Wuhan, China to assess the performance of the proposed approach. The results demonstrate that 935 vehicles will totally travel 56,424.55 km to deliver necessary living materials to 3,154 neighborhoods, with total risk [Formula: see text]. The presented approach reduces the risk of COVID-19 transmission by 67.55% compared to traditional distance-based optimization methods. The presented approach can facilitate a well response to the COVID-19 in the transportation sector. Tianhong Zhao, Wei Tu 0001, Zhixiang Fang, Zhengdong Huang, Shengwu Xiong 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Foreign Shadow Robust Makeup Transfer via Hierarchical Deep Aggregation and Disentangled RepresentationabstractFacial makeup transfer aims to transfer the makeup style from a reference makeup image to a source image. In daily life, the reference makeup images from various scenes are uneasy about guaranteeing quality. In contrast, the source images photoed by ourselves are likely to control illumination and avoid shadows cast by occlusion. To this end, we propose a novel robust makeup transfer network (SRMT) for facial makeup and de-makeup under the foreign shadow. The reference image is hierarchically aggregated multi-context features for predicting makeup style which is input into a disentangled network with source image for achieving shadow robust makeup transfer. In particular, we first incorporate a shadow manipulation module(SMM) to manipulate the reference makeup, which furtherly restores the original color distribution on the face. Then we propose a makeup transfer module(MTM) based on disentangled representation, producing a high-quality source image from the predicted reference makeup. Extensive experiments show the superiority of our method in terms of visual effects. Moreover, a carefully designed makeup dataset with paired shadow-deshadow makeup images is available at https://github.com/lucuspring/Deshadow-dataset. Chang Zhan, Lin Wang 0019, Zhaoyang Sun, Jinli Zhou, Shengwu Xiong 0001 |
FG | 6 |
| 2021 | Neural Noise Embedding for End-To-End Speech Enhancement with Conditional Layer NormalizationabstractMost of the deep learning based speech enhancement methods focus on the modeling of complicated relationship between the noisy speech and the clean speech without the consideration of noise information. In order to cope with various complex noise scenes, we introduce a novel enhancement architecture that integrates a deep autoencoder with neural noise embedding. In this study, a new normalization method, termed conditional layer normalization (CLN), is introduced to improve the generalization of deep learning based speech enhancement approaches for unseen environments. The noise embedding is passed through the CLN layers to regularize the network for speech enhancement task. The proposed network can be adaptively adjusted according to different noise information extracted from the noisy speech input. The network in overall is trained in an end-to-end manner and the experimental results show that the proposed scheme produces satisfactory enhancement performance comparing the other methods. The visualization shows that our proposed network captures noise information, which is helpful to improve robustness to unseen environments for speech enhancement. Xiaoqi Li 0011, Yaxing Li, Yuanjie Dong, Shengwu Xiong 0001 |
ICASSP | 6 |
| 2021 | Benchmark Platform for Ultra-Fine-Grained Visual Categorization Beyond Human PerformanceabstractDeep learning methods have achieved remarkable success in fine-grained visual categorization. Such successful categorization at sub-ordinate level, e.g., different animal or plant species, however relies heavily on the visual differences that human can observe and the ground-truths are labelled on the basis of such human visual observation. In contrast, few research has been done for visual categorization at the ultra-fine-grained level, i.e., a granularity where even human experts can hardly identify the visual differences or are not yet able to give affirmative labels by inferring observed pattern differences. This paper reports our efforts towards mitigating this research gap. We introduce the ultra-fine-grained (UFG) image dataset, a large collection of 47,114 images from 3,526 categories. All the images in the proposed UFG image dataset are grouped into categories with different confirmed cultivar names. In addition, we perform an extensive evaluation of state-of-the-art fine-grained classification methods on the proposed UFG image dataset as comparative baselines. The proposed UFG image dataset and evaluation protocols is intended to serve as a benchmark platform that can advance research of visual classification from approaching human performance to beyond human ability, via facilitating benchmark data of artificial intelligence (AI) not to be limited by the labels of human intelligence (HI). The dataset is available online at https://githuh.com/XiaohanYu-GU/Ultra-FGVC. Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
ICCV | 5 |
| 2021 | An Efficient Message Dissemination Scheme for Cooperative Drivings via Multi-Agent Hierarchical Attention Reinforcement LearningabstractA group of connected and autonomous vehicles (CAVs) with common interests can drive in a cooperative manner, namely cooperative driving, which has been verified to significantly improve road safety, traffic efficiency, and environmental sustainability. A more general scenario with various types of cooperative driving applications such as truck platooning and vehicle clustering will coexist on roads in the foreseeable future. To support such multiple cooperative drivings, it is critical to design an efficient message dissemination scheduling for vehicles to broadcast their kinetic status, i.e., beacon periodically. Most ongoing researches suggest designing the communication protocols via traffic and communication modeling on top of dedicated short range communications (DSRC) or cellular-based vehicle-to-vehicle (C-V2V) communications as a potential remedy. However, most of the existing researches are designed for a simple or specific traffic scenario, e.g., ignoring the impacts of the complex communication environment and emerging hybrid traffic scenarios. Moreover, some studies design beaconing strategies based on the implication of channel and traffic conditions in the beacons of other vehicles. However, the delayed perception of these information may seriously deteriorate the beaconing performance. In this paper, we take the perspective of cooperative drivings and formulate their decision-making process as a Markov game. Furthermore, we propose a multi-agent hierarchical attention reinforcement learning (MAHA) framework to solve the Markov game. More concretely, the hierarchical structure of the proposed MAHA can lead cooperative drivings to be foresightful. Hence, even without immediate incentives, the well-trained agents can still take favorable actions that benefit their long-term rewards. Besides, we integrate each hierarchical level of MAHA separately with the graph attention network (GAT) to incorporate agents' mutual influences in the decision-making process. Besides, we set up a simulator and adopt this simulator to generate dynamic traffic scenarios, which reflect the different real-world scenarios faced by cooperative drivings. We conduct extensive experiments to evaluate the proposed MAHA framework's performance. The results show that MAHA can significantly improve the beacon reception rate and guarantee low communication delay in all of these scenarios. Bingyi Liu, Weizhen Han, Enshu Wang, Shengwu Xiong 0001, Chunming Qiao, Jianping Wang 0001 |
ICDCS | 5 |
| 2021 | Accurate and Robust Stereo Direct Visual Odometry for Agricultural EnvironmentabstractVision-based localization and mapping in the agricultural environment is challenging due to the unstructured scene with unstable features, illumination variations, bumpy roads, and dynamic environmental objects. To address these challenges, we propose an accurate and robust stereo direct visual odometry system with modifications on Stereo-DSO. We firstly select some well-matched static stereo points in the latest keyframe to improve the accuracy of inverse depth calculation for tracking. The inverse depth can further distinguish close objects from background, which will avoid large and far-away scene objects in keyframe determination. To boost efficiency and accuracy at the tracking stage, we propose a point selection method to sample map points and remove outliers. Furthermore, altitude smoothness verification with a local flat ground assumption and recovery method for tracking failure on bumpy roads are proposed to improve the system’s robustness. Finally, a far-away keyframe is reserved in the sliding window to alleviate the orientation drift since the agricultural robots usually move straightly following the crop row. Our system achieved new state-of-the-art results on Flourish dataset and the recently released Rosario dataset. Junwei Zhou 0002, Liangliang Wang 0007, Shengwu Xiong 0001 |
ICRA | 4 |
| 2021 | Variational Information Bottleneck Based Regularization for Speaker Recognition
Yuanjie Dong, Yaxing Li, Yunfei Zi, Xiaoqi Li 0011, Shengwu Xiong 0001 |
Interspeech | 7 |
| 2021 | An Efficient Approximation for Quantitative Analysis of Dynamic Fault TreesabstractThis paper presents a feasibility and effective ap-proximation method to estimate the failure probability of the top event of a dynamic fault tree. The method is based on a minimal canonical form and uses a quantitative relationship between the smallest cut sequence and the entire sequence. Comparison with discrete-time Bayesian networks and Monte Carlo simulation methods, the validity of this method is assessed on two case studies approximating the probabilities of the top event of a Hypothetical Cardiac Assist System (HCAS) and a fictitious system. The case study results show that our method can achieve similar accuracy with smaller relative error and shorter execution time. Luyao Ye, Erqing Li, Dongdong Zhao 0001, Shengwu Xiong 0001, Jianwen Xiang |
ISSRE | 4 |
| 2021 | Few-Shot Learning with Unlabeled Outlier Exposure
Jieya Lian, Shengwu Xiong 0001 |
MMM (1) | 3 |
| 2021 | Attention Based Reinforcement Learning with Reward Shaping for Knowledge Graph Reasoning
Shengwu Xiong 0001 |
NLPCC (1) | 3 |
| 2021 | Rule Injection-Based Generative Adversarial Imitation Learning for Knowledge Graph Reasoning
Xiaoyin Chen, Shengwu Xiong 0001 |
PAKDD (3) | 3 |
| 2021 | A bi-level distribution mixture framework for unsupervised driving performance evaluation from naturalistic truck driving data
Lin Lu 0002, Shengwu Xiong 0001, Yaxiong Chen |
Eng. Appl. Artif. Intell. | 2 |
| 2021 | Alternating Primal-Dual Algorithm for Minimizing Multiple-Summed Separable Problems with Application to Image RestorationabstractIn order to discover the difference among dual strategies, we propose an alternating primal-dual algorithm (APDA) that can be considered as a general version for minimizing problem which is multiple-summed separable convex but not necessarily smooth. First, the original multiple-summed problem is transformed into two subproblems. Second, one subproblem is solved in the primal space and the other is solved in the dual space. Finally, the alternating direction method is executed between the primal and the dual part. Furthermore, the classical alternating direction method of multipliers (ADMM) is extended to solve the primal subproblem which is also multiple summed, therefore, the extended ADMM can be seen as a parallel method for the original problem. Thanks to the flexibility of APDA, different dual strategies for image restoration are analyzed. Numerical experiments show that the proposed method performs better than some existing algorithms in terms of both speed and accuracy. Peng Wang 0188, Shengwu Xiong 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | A Tibetan Thangka data set and relative tasks
Yanchun Ma, Yongjian Liu, Qing Xie 0002, Shengwu Xiong 0001, Lihua Bai, Anshu Hu |
Image Vis. Comput. | 4 |
| 2021 | Hyperspectral Image Dimensionality Reduction via Graph Embedding in Core Tensor SpaceabstractThis letter tries to effectively reduce the dimension of hyperspectral images (HSIs) by jointly considering both the spectral redundancy and spatial continuity through a multilinear transformation with graph embedding in core tensor space. The whole process is constructed in the framework of Tucker decomposition (TD). Since the distance between intraclass samples should be relatively smaller than that of the interclass samples, the reduced tensor cores should maintain this property. To achieve this goal, a graph is embedded to the core tensor space during TD. Moreover, considering the unstability of solution of the previous works, we constrain the projected matrices by orthogonality so that the results can be more stable and the extracted features can be more discriminative. We further analyze the effect of different constrains to TD methods for HSI dimensionality reduction. Finally, the experimental results show the superiority of this method to many other tensor methods. Peng Wang 0165, Chengyong Zheng, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | MaskCOV: A random mask covariance network for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 4 |
| 2021 | Learning deep part-aware embedding for person retrieval
Yang Zhao 0019, Chunhua Shen, Xiaohan Yu 0001, Hao Chen 0041, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 6 |
| 2020 | Patchy Image Structure Classification Using Multi-Orientation Region TransformabstractExterior contour and interior structure are both vital features for classifying objects. However, most of the existing methods consider exterior contour feature and internal structure feature separately, and thus fail to function when classifying patchy image structures that have similar contours and flexible structures. To address above limitations, this paper proposes a novel Multi-Orientation Region Transform (MORT), which can effectively characterize both contour and structure features simultaneously, for patchy image structure classification. MORT is performed over multiple orientation regions at multiple scales to effectively integrate patchy features, and thus enables a better description of the shape in a coarse-to-fine manner. Moreover, the proposed MORT can be extended to combine with the deep convolutional neural network techniques, for further enhancement of classification accuracy. Very encouraging experimental results on the challenging ultra-fine-grained cultivar recognition task, insect wing recognition task, and large variation butterfly recognition task are obtained, which demonstrate the effectiveness and superiority of the proposed MORT over the state-of-the-art methods in classifying patchy image structures. Our code and three patchy image structure datasets are available at: https://github.com/XiaohanYu-GU/MReT2019. Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001 |
AAAI | 4 |
| 2020 | Local Facial Makeup Transfer via Disentangled Representation
Zhaoyang Sun, Shengwu Xiong 0001, Wenxuan Liu 0008 |
ACCV (4) | 4 |
| 2020 | Wheat Phenotype Extraction via Adaptive Supervoxel SegmentationabstractWheat is one of the main food crops in the world today. In order to increase wheat yield, breeding experts are committed to discovering the connection between its genotype and phenotype. Existing phenotype extraction methods mostly rely on manual methods, and the amount of collected data is limited and inefficient. 3D Computed Tomography (CT) imaging has the advantages of fine imaging, high dynamic range, and nondestructive detection of internal structures, which can help to extract the high-throughput phenotype of wheat non-destructively and finely. However, the 3D images scanned by CT have the characteristics of large amount of data and highly sparse content, which brings great challenges to extract phenotype from them. We propose an adaptive supervoxel algorithm for the high sparsity of 3D CT images, which is used for the pre-segmentation. Then, an octree convolutional network is used to optimize the sparse structure storage, and further improve the efficiency of fine segmentation. At last, phenotype is extracted automatically from the finely segmented wheat grain. Chengde Lin, Shengwu Xiong 0001 |
BIBM | 3 |
| 2020 | Multi-components System for Automatic Arabic Diacritization
Hamza Abbad, Shengwu Xiong 0001 |
ECIR (1) | 2 |
| 2020 | A Time-Frequency Network with Channel Attention and Non-Local Modules for Artificial Bandwidth ExtensionabstractConvolution neural networks (CNNs) have been achieving increasing attention for the artificial bandwidth extension (ABE) task recently. However, these methods use the flipped low-frequency phase to reconstruct speech signals, which may lead to the well-known invalid short-time Fourier Transform (STFT) problem. The convolutional operations only enable networks to construct informative features by fusing both channel-wise and spatial information within local receptive fields at each layer. In this paper, we introduce a Time-Frequency Network (TFNet) with channel attention (CA) and non-local (NL) modules for ABE. The TFNet exploits the information from both time and frequency domain branches concurrently to avoid the invalid STFT problem. To capture the channels and space dependencies, we incorporate the CA and NL modules to construct a proposed fully convolutional neural network for the time and frequency branches of TFNet. Experimental results demonstrate that the proposed method outperforms the competing method. Yuanjie Dong, Yaxing Li, Xiaoqi Li 0011, Shan Xu 0007, Shengwu Xiong 0001 |
ICASSP | 7 |
| 2020 | TA-MAC: A Traffic-Aware TDMA MAC Protocol for Safety Message Dissemination in MEC-assisted VANETsabstractVehicular ad hoc networks (VANETs) have been widely recognized as a promising solution to improve traffic safety and efficiency for the ability to provide situation awareness even though the potential dangers and traffic anomalies are out of the visual range. In VANETs, time-division multiple access (TDMA) based overlay protocols can prevent transmission collisions, and play an important role in providing an efficient communication channel. However, due to high vehicle mobility and time-varying traffic flow, the existing TDMA-based slot allocation approaches cannot fully utilize the channel resources, which may result in high transmission delay and packet collision. To overcome these shortcomings, we propose a traffic-aware TDMA-based MAC (TA-MAC) protocol which utilizes the capability of mobile edge computing (MEC) in this paper. Specifically, based on MEC and vehicle-to-road-side-units (V2R) communications, a traffic-aware mechanism is first proposed to estimate the traffic condition on the road segment. Then, we propose a new slot assignment method that aims at guaranteeing the high channel utilization and low delay of safety message under dynamic traffic conditions. Finally, we conduct extensive experiments to demonstrate the effectiveness of the proposed protocol. Dongxiao Deng, Wenbi Rao, Bingyi Liu, Dongyao Jia, Yang Sheng, Jianping Wang 0001, Shengwu Xiong 0001 |
ICCCN | 7 |
| 2020 | Pose Guided Person Image Generation Based on Pose Skeleton Sequence and 3D ConvolutionabstractThis paper proposes a novel approach for pose transfer, which aims at transferring the pose of a given person to a target pose. Unlike previous works directly simulating the target pose, we emphasize the pose geometric constraint in our approach to tackle this task. To capture the geometric constraints among the condition pose and target pose, we generate a sequence of different intermediate pose skeleton. Moreover, we introduced the 3D convolution into the generator of our Generative Adversarial Network (GAN) in order to learn spatiotemporal features from the generated pose skeleton sequence. The final target pose will be estimated and refined by a pose transfer block and a series of residual blocks based on the spatiotemporal features. Compared with the existing works, our generated person pose images achieved better performance in terms of Inception Score and Structure Similarity. Extensive experiments on dataset Market-1501 demonstrate the effectiveness of the pose skeleton sequence on pose transfer. Wenbin Zhao, Qing Xie 0002, Yanchun Ma, Yongjian Liu, Shengwu Xiong 0001 |
ICIP | 5 |
| 2020 | Learning Class Prototypes Via Anisotropic Combination of Aligned Modalities for Few-Shot LearningabstractPrototypical networks have shown their simplicity and effectiveness in few-shot learning. Recently, some works leverage cross-modal information to enhance the class prototype in few-shot learning. However, they seldom make use of structured information from textual space to optimize class prototype representation. And they either don't align textual modality to visual modality or align them too rigidly. We argue that proper alignment method is important to improve the performance of cross-modal methods, as query data only has visual information in the few-shot learning tasks. In this paper, we propose a cross-modal alignment method to optimize class prototypes with structured information from textual space. We further introduce an anisotropic combination method to enhance class prototypes with information from two modalities. Experiments show that the cross-modal alignment method and anisotropic combination method achieve state-of-the-art results on the miniImageNet and tieredImageNet benchmark in one-shot regime. Jieya Lian, Shengwu Xiong 0001 |
ICME | 3 |
| 2020 | Few-Shot Classification with Transductive Data Clustering Transformation
Jieya Lian, Shengwu Xiong 0001 |
ICONIP (2) | 3 |
| 2020 | RGB-Infrared Person Re-identification via Image Modality ConversionabstractAs a cross modality retrieval task, RGB-infrared person re-identification(Re-ID) is an important and challenging task, because of its important role in video surveillance applications and large cross-modality variations between visible and infrared images. Most previous works addressed the problem of cross-modality gap with feature alignment by original feature representation learning straightly. In this paper, different from existing works, we propose a novel network(CE2L) to tackle the cross-modality gap with feature alignment. CE2L mainly focuses on adding discriminative information and learning robust features by converting modality between visible and infrared images. Its merits are highlighted in two aspects: 1)Using CycleGAN to convert infrared images into color images can not only increase the recognition characteristics of images, but also allow the network to better learn the two modal image features; 2)Our novel method can serve as data augmentation. Specifically, it can increase data diversity and total data against over-fitting by converting labeled training images to another modal images. Extensive experimental results on two datasets demonstrate superior performance compared to the baseline and the state-of-the-art methods. Huangpeng Dai, Qing Xie 0002, Yanchun Ma, Yongjian Liu, Shengwu Xiong 0001 |
ICPR | 5 |
| 2020 | Face Anti-Spoofing Based on Dynamic Color Texture Analysis Using Local Directional Number PatternabstractFace anti-spoofing is becoming increasingly indispensable for face recognition systems, which are vulnerable to various spoofing attacks performed using fake photos and videos. In this paper, a novel “LDN-TOP representation followed by ProCRC classification” pipeline for face anti-spoofing is proposed. We use local directional number pattern (LDN) with the derivative-Gaussian mask to capture detailed appearance information resisting illumination variations and noises, which can influence the texture pattern distribution. To further capture motion information, we extend LDN to a spatial-temporal variant named local directional number pattern from three orthogonal planes (LDN- TOP). The multi-scale LDN- TOP capturing complete information is extracted from color images to generate the feature vector with powerful representation capacity. Finally, the feature vector is fed into the probabilistic collaborative representation based classifier (ProCRC) for face anti-spoofing. Our method is evaluated on three challenging public datasets, namely CASIA FASD, Replay-Attack database, and UVAD database using sequence-based evaluation protocol. The experimental results show that our method can achieve promising performance with 0.37% EER on CASIA and 5.73% HTER on UVAD. The performance on Replay-Attack database is also competitive. Junwei Zhou 0002, Ke Shu, Peng Liu 0005, Jianwen Xiang, Shengwu Xiong 0001 |
ICPR | 5 |
| 2020 | Scene Text Detection with Selected AnchorsabstractObject proposal technique with dense anchoring scheme for scene text detection were applied frequently to achieve high recall. It results in the significant improvement in accuracy but waste of computational searching, regression and classification. In this paper, we propose an anchor selection-based region proposal network (AS-RPN) using effective selected anchors instead of dense anchors to extract text proposals. The center, scales, aspect ratios and orientations of anchors are learnable instead of fixing, which leads to high recall and greatly reduced numbers of anchors. By replacing the anchor-based RPN in Faster RCNN, the AS-RPN-based Faster RCNN can achieve comparable performance with previous state-of-the-art text detecting approaches on standard benchmarks, including COCO-Text, ICDAR2013, ICDAR2015 and MSRA-TD500 when using single-scale and single model (ResNet50) testing only. Anna Zhu, Shengwu Xiong 0001 |
ICPR | 3 |
| 2020 | Bidirectional LSTM Network with Ordered Neurons for Speech Enhancement
Xiaoqi Li 0011, Yaxing Li, Yuanjie Dong, Shan Xu 0007, Shengwu Xiong 0001 |
INTERSPEECH | 7 |
| 2020 | A Novel Safety Message Dissemination for Region of Interest Coverage Using Vehicle TrajectoryabstractVehicular communication networking (VCN) has been widely recognized as a promising solution to support safety-related applications in urban transportation systems. In VCN, efficient message dissemination can let vehicles be better aware of the potential risks and traffic anomalies, which is critical to road safety and traffic efficiency. Substantial studies have focused on the design of inter-vehicle message dissemination protocols. Nonetheless, most existing designs only consider the rapid end-to-end transmission, few of which take into account the broadcast coverage. In this paper, we propose a new message dissemination scheme in the urban traffic scenario by considering both the time constraint and the spatial distribution of data dissemination. Specifically, based on the temporal and spatial correlation of vehicle trajectory, relay vehicles are selected to construct a temporary warning network (TWN) for a rapid safety message dissemination in the regions of interest (ROI). Finally, we conduct extensive numerical experiments to validate the effectiveness of our method in various traffic scenarios. Bingyi Liu, Zhipeng Fang, Dongyao Jia, Shengwu Xiong 0001, Enshu Wang, Jianping Wang 0001 |
VTC Fall | 5 |
| 2020 | A weighted KNN-based automatic image annotation method
Yanchun Ma, Qing Xie 0002, Yongjian Liu, Shengwu Xiong 0001 |
Neural Comput. Appl. | 4 |
| 2020 | MobileFAN: Transferring deep hidden representation for face alignment
Yang Zhao 0019, Yifan Liu 0001, Chunhua Shen, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 5 |
| 2020 | Software aging and rejuvenation in android: new models and metrics
Jianwen Xiang, Caisheng Weng, Dongdong Zhao 0001, Artur Andrzejak 0001, Shengwu Xiong 0001, Lin Li 0001 |
Softw. Qual. J. | 5 |
| 2020 | Double Graph Regularized Double Dictionary Learning for Image ClassificationabstractIn this paper, we present a novel double graph regularized double dictionary learning (DGRDDL) method for image classification. The proposed method jointly constructs a number of class-specific sub-dictionaries to capture the most discriminative features (class-specific information) of each class, and a class-shared dictionary to model the common patterns (class-shared information) shared by the images from different classes. A novel double graph regularization is proposed to correctly represent and differentiate these two types of information. Specifically, an intra-class similarity graph constraint is imposed on the representation coefficients over the class-specific dictionaries, and an inter-class similarity graph constraint is applied on the representation coefficients over the class-shared dictionary. In this way, the representations learned by the proposed DGRDDL method can correctly model the local similarity relationships of the class-specific and the class-shared information in images, respectively. Moreover, due to the differences between the intra-class and inter-class similarity graphs, the two types of information can be appropriately separated and captured by the learned dictionaries. We evaluate the performance of the proposed method on six public datasets and compared against those of seven benchmark methods. The experimental results demonstrate the effectiveness and superiority of the proposed method in image classification over the benchmark dictionary learning methods. Shengwu Xiong 0001, Yongsheng Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Few-Shot Text Style Transfer via Deep Feature SimilarityabstractGenerating text to have a consistent style with only a few observed highly-stylized text samples is a difficult task for image processing. The text style involving the typography, i.e., font, stroke, color, decoration, effects, etc., should be considered for transfer. In this paper, we propose a novel approach to stylize target text by decoding weighted deep features from only a few referenced samples. The deep features, including content and style features of each referenced text, are extracted from a Convolutional Neural Network (CNN) that is optimized for character recognition. Then, we calculate the similarity scores of the target text and the referenced samples by measuring the distance along the corresponding channels from the content features of the CNN when considering only the content, and assign them as the weights for aggregating the deep features. To enforce the stylized text to be realistic, a discriminative network with adversarial loss is employed. We demonstrate the effectiveness of our network by conducting experiments on three different datasets which have various styles, fonts, languages, etc. Additionally, the coefficients for character style transfer, including the character content, the effect of similarity matrix, the number of referenced characters, the similarity between characters, and performance evaluation by a new protocol are analyzed for better understanding our proposed framework. Anna Zhu, Xiongbo Lu, Xiang Bai, Seiichi Uchida, Brian Kenji Iwana, Shengwu Xiong 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | A Variational Bayesian Framework for Cluster Analysis in a Complex NetworkabstractA complex network is a network with non-trivial topological structures. It contains not just topological information but also attribute information available in the rich content of nodes. Concerning the task of cluster analysis in a complex network, model-based algorithms are preferred over distance-based ones, as they avoid designing specific distance measures. However, their models are only applicable to complex networks where the attribute information is composed of attributes in binary form. To overcome this disadvantage, we introduce a three-layer node-attribute-value hierarchical structure to describe the attribute information in a flexible and interpretable manner. Then, a new Bayesian model is proposed to simulate the generative process of a complex network. In this model, the attribute information is generated by following the hierarchical structure while the links between pairwise nodes are generated by a stochastic blockmodel. To solve the corresponding inference problem, we develop a variational Bayesian algorithm called TARA, which allows us to identify functionally meaningful clusters through an iterative procedure. Our extensive experiment results show that TARA can be an effective algorithm for cluster analysis in a complex network. Moreover, the parallelized version of TARA makes it possible to perform efficiently at its tasks when applied to large complex networks. Lun Hu, Keith C. C. Chan, Shengwu Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2019 | Adapting Layer RBERs Variations of 3D Flash Memories via Multi-granularity Progressive LDPC ReadingabstractExisting studies have uncovered that there exist significant Raw Bit Error Rates (RBERs) variations among different layers of 3D flash memories due to manufacture process variation. These RBER variations would cause significantly diversed read latencies when reading data with traditional Low-Density Parity-Check (LDPC) codes designed for planar flash memories, which induces sub-optimal read performance of flash-based Solid-State Drives (SSDs). Yajuan Du, Meng Zhang 0014, Shengwu Xiong 0001 |
DAC | 5 |
| 2019 | Densely Connected Network with Time-frequency Dilated Convolution for Speech EnhancementabstractThe data driven speech enhancement approaches using regression-based deep neural network usually result in enormous number of model parameters, which increase the computational load and the difficulty of model training. In order to improve the model efficiency, we propose a densely connected network with time-frequency (T-F) dilated convolution for speech enhancement. The T-F dilated convolution block is designed to enlarge the receptive field and capture the contextual information in both temporal and frequency domains. Considering the computational efficiency, the 1-D convolution with the bottleneck structure is exploited in the T-F convolution block. Each T-F convolution block is then densely connected to ensure maximum information flow between layers and alleviate the vanishing gradient problem of the network. The experimental results reveal that the proposed scheme not only improves the computational efficiency significantly but also produces satisfactory enhancement performance comparing the competing methods. Yaxing Li, Xiaoqi Li 0011, Yuanjie Dong, Shan Xu 0007, Shengwu Xiong 0001 |
ICASSP | 6 |
| 2019 | Contour Covariance: A Fast Descriptor for ClassificationabstractThis paper presents a novel shape descriptor to effectively and efficiently characterize the local image statistics. The proposed descriptor, termed contour covariance (CC), characterizes covariance features driven by a moving point on the shape contour at multiple scales. To calculate the covariance matrices, three basic features including texture, intensity and distance map, are extracted from the object image. Based on coefficients of the obtained covariance matrices, the proposed CC descriptor is compact yet informative, as well as invariant to rotation, translation and scale. The experimental results on two databases demonstrate the superiority and efficiency of the proposed method among the state-of-the-art methods for shape classification. Xiaohan Yu 0001, Shengwu Xiong 0001, Yongsheng Gao 0001 |
ICIP | 2 |
| 2019 | Character Image Synthesis Based on Selected Content and Referenced Style EmbeddingabstractArbitrary characters synthesis based on a few referenced examples poses a great challenge due to the diversity of characters category and style. We regard this problem as image translation problem and propose a character style transfer network consisting of content selector, style encoder, content encoder, feature embedding and embedded feature decoder to solve it. The content selector is used to select and match the most similar content (i.e., font) from our collected glyph dataset as content references. Then, we apply the style encoder and content encoder to extract the style and content representation separately and mix them for feature embedding. Finally, the embedded features are decoded to generate the target characters. We train them in an end-to-end manner and evaluate the proposed method on MC-GAN dataset and our collected dataset. The experimental results have demonstrated the effectiveness of the proposed model for character synthesis. Anna Zhu, Xiongbo Lu, Shengwu Xiong 0001 |
ICME | 4 |
| 2019 | A Convolutional Neural Network with Non-Local Module for Speech Enhancement
Xiaoqi Li 0011, Yaxing Li, Shan Xu 0007, Yuanjie Dong, Xinrong Sun, Shengwu Xiong 0001 |
INTERSPEECH | 7 |
| 2019 | A correction to: on the linear complexity of the Sidelnikov-Lempel-Cohn-Eastman sequences
Minglong Qi, Shengwu Xiong 0001 |
Des. Codes Cryptogr. | 2 |
| 2019 | Multi-level thresholding-based grey scale image segmentation using multi-objective multi-verse optimizer
Mohamed E. Abd Elaziz, Diego Oliva 0001, Ahmed A. Ewees, Shengwu Xiong 0001 |
Expert Syst. Appl. | 4 |
| 2019 | Coarse-to-fine document localization in natural scene image with regional attention and recursive corner refinement
Anna Zhu, Shengwu Xiong 0001 |
Int. J. Document Anal. Recognit. | 4 |
| 2019 | Task scheduling in cloud computing based on hybrid moth search algorithm and differential evolution
Mohamed E. Abd Elaziz, Shengwu Xiong 0001, K. P. N. Jayasena, Lin Li 0001 |
Knowl. Based Syst. | 2 |
| 2019 | CVSkSA: cross-architecture vulnerability search in firmware based on kNN-SVM and attributed control flow graphabstractTo prevent the same known vulnerabilities from affecting different firmware, searching known vulnerabilities in binary firmware across different architectures is crucial. Because the accuracy of existing cross-architecture vulnerability search methods is not high, we propose a staged approach based on support vector machine (SVM) and attributed control flow graph (ACFG) at the function level to improve the accuracy using prior knowledge. Furthermore, for efficiency, we utilize the k-nearest neighbor (kNN) algorithm to prune and SVM to refine in the function prefilter stage. Although the accuracy of the proposed method using kNN-SVM approach is slightly lower than the accuracy of the method using only SVM, its efficiency is significantly enhanced. We have implemented our approach CVSkSA to search several vulnerabilities in real-world firmware images. The experimental results show that the accuracy of the proposed method using kNN-SVM approach is close to the accuracy of the method using only SVM in most cases, while the former is approximately four times faster than the latter. Dongdong Zhao 0001, Hong Lin 0004, Linjun Ran, Mushuai Han, Shengwu Xiong 0001, Jianwen Xiang |
Softw. Qual. J. | 7 |
| 2019 | Multi-Channel Embedding Convolutional Neural Network Model for Arabic Sentiment ClassificationabstractWith the advent of social network services, Arabs’ opinions on the web have attracted many researchers in recent years toward detecting and classifying sentiments in Arabic tweets and reviews. However, the impact of word embeddings vectors (WEVs) initialization and dataset balance on Arabic sentiment classification using deep learning has not been thoroughly studied. In this article, a multi-channel embedding convolutional neural network (MCE-CNN) is proposed to improve Arabic sentiment classification by learning sentiment features from different text domains, word, and character n-grams levels. MCE-CNN encodes a combination of different pre-trained word embeddings into the embedding block at each embedding channel and trains these channels in parallel. Besides, a separate feature extraction module implemented in a CNN block is used to extract more relevant sentiment features. These channels and blocks help to start training on high-quality WEVs and fine-tuning them. The performance of MCE-CNN is evaluated on several standard balanced and imbalanced datasets to reflect real-world use cases. Experimental results show that MCE-CNN provides a high classification accuracy and benefits from the second embedding channel on both standard Arabic and dialectal Arabic text, which outperforms state-of-the-art methods. Abdelghani Dahou, Shengwu Xiong 0001, Junwei Zhou 0002, Mohamed E. Abd Elaziz |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Efficiently Detecting Protein Complexes from Protein Interaction Networks via Alternating Direction Method of MultipliersabstractProtein complexes are crucial in improving our understanding of the mechanisms employed by proteins. Various computational algorithms have thus been proposed to detect protein complexes from protein interaction networks. However, given massive protein interactome data obtained by high-throughput technologies, existing algorithms, especially those with additionally consideration of biological information of proteins, either have low efficiency in performing their tasks or suffer from limited effectiveness. For addressing this issue, this work proposes to detect protein complexes from a protein interaction network with high efficiency and effectiveness. To do so, the original detection task is first formulated into an optimization problem according to the intuitive properties of protein complexes. After that, the framework of alternating direction method of multipliers is applied to decompose this optimization problem into several subtasks, which can be subsequently solved in a separate and parallel manner. An algorithm for implementing this solution is then developed. Experimental results on five large protein interaction networks demonstrated that compared to state-of-the-art protein complex detection algorithms, our algorithm outperformed them in terms of both effectiveness and efficiency. Moreover, as number of parallel processes increases, one can expect an even higher computational efficiency for the proposed algorithm with no compromise on effectiveness. Lun Hu, Xing Liu 0002, Shengwu Xiong 0001, Xin Luo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | A Novel Multi-scale Invariant Descriptor Based on Contour and Texture for Shape Recognition
Jishan Guo, Yongsheng Gao 0001, Shengwu Xiong 0001 |
ACCV (4) | 5 |
| 2018 | Optimizing the Energy Efficient VM consolidation by a Multi-Objective AlgorithmabstractOptimizing energy efficient Virtual Machine Consolidation (VMC) in a cloud computing environment, which is a non-linear multi-objective NP-hard problem, plays a vital role in decreasing energy consumption, and increasing Quality of Service (QoS). In this paper, VMC is formulated as a multi-objective optimization problem, which has three conflicting objectives, power consumption, Service Level Agreements Violation (SLAV) and Mean Time Before Host Shutdown (MTBHS). We propose a multi-objective optimization algorithm based on Multi-Objective Sine Cosine Algorithm (MOSCA) for the VMC. We evaluate the performance of our model by applying two multi-objective algorithms, namely, Multi-Objective Evolutionary Algorithm based on Decomposition (MOEAD) and Non-dominated Sorting Genetic Algorithm (NSGAII). Our research mainly focus on two tasks, i.e.,evaluating and comparing the multi-objective algorithms to find out the optimal solution and develop a MOSCA based algorithm to solve the proposed VMC model. The simulation results illustrated that the propose multi-objective model meets the optimal solutions amongst the three conflicting objectives, which significantly reduces the power consumption, SLAV and maximize the MTBHS. It got the best performance according to the Multi-objective Optimization Problem (MOP) indicators. K. P. N. Jayasena, Lin Li 0001, Mohamed E. Abd Elaziz, Shengwu Xiong 0001, Jianwen Xiang |
CSCWD | 4 |
| 2018 | Handling Unreasonable Data in Negative Surveys
Jianwen Xiang, Shu Fang, Dongdong Zhao 0001, Shengwu Xiong 0001, Chunhui Yang |
DASFAA (2) | 5 |
| 2018 | A Second-Order Variational Framework for Joint Depth Map Estimation and Image DehazingabstractOutdoor images captured in poor weather conditions (e.g., fog or haze) commonly suffer from reduced contrast and visibility. Increasing attention has recently been paid to single image dehazing, i.e., improving image contrast and visibility. It is generally thought that the dehazing performance highly depends on the accurate depth information. In this work, we first obtain the initial depth map by using the popular dark channel prior. A unified second-order variational framework is then proposed to refine the depth map and restore the haze-free image. The introduced second-order framework has the capacity of preserving important structures in both depth map and haze-free image. Furthermore, the proposed framework performs well for several different types of haze situations. The resulting optimization problems related to depth map estimation and latent image restoration can be effectively handled using the primal-dual algorithm under a two-step numerical framework. The effectiveness of our proposed method has been demonstrated by comparing the imaging performance with several state-of-the-art dehazing methods. Ryan Wen Liu, Shengwu Xiong 0001, Huisi Wu |
ICASSP | 2 |
| 2018 | Multi-Scale Piecewise Line Integral Strategy for Structure Integral TransformabstractStructure Integral transform (SIT) is a mathematical tool for invariant shape recognition. SIT has a superior ability to capture the interior structure information by integrating the shape image function over 2D dissecting structure, but the details on the dissecting structure are discarded. In this paper, for the first time, we propose a novel multi-scale piecewise line integral strategy for SIT which introduces the grayscale information. At a higher scale, the integral line on dissecting structure will be equally divided more times, which brings to a coarse to fine description for each integral line. The proposed multi-scale piecewise integral strategy has been validated to extend SIT from binary shape to grayscale shape successfully through experiments on Leeds' Butterflies dataset and Ponce Group's Butterflies dataset. Yongsheng Gao 0001, Jishan Guo, Shengwu Xiong 0001 |
ICIP | 5 |
| 2018 | Privacy-Preserving K-Means Clustering Upon Negative Databases
Dongdong Zhao 0001, Jianwen Xiang, Xing Liu 0002, Haiying Zhou, Shengwu Xiong 0001 |
ICONIP (4) | 7 |
| 2018 | Two-Phase Transmission Map Estimation for Robust Image Dehazing
Qiaoling Shu, Chuansheng Wu, Ryan Wen Liu, Kwok Tai Chui, Shengwu Xiong 0001 |
ICONIP (6) | 5 |
| 2018 | Multi-frame Quantization of LSF Parameters Using a Deep Autoencoder and Pyramid Vector Quantizer
Yaxing Li, Eshete Derb Emiru, Shengwu Xiong 0001, Anna Zhu, Pengfei Duan 0005, Yichang Li |
INTERSPEECH | 3 |
| 2018 | Multi-frame Coding of LSF Parameters Using Block-Constrained Trellis Coded Vector Quantization
Yaxing Li, Shan Xu 0007, Shengwu Xiong 0001, Anna Zhu, Pengfei Duan 0005, Yueming Ding |
INTERSPEECH | 3 |
| 2018 | Iris Template Protection Based on Randomized Response Technique and Aggregated Block InformationabstractNowadays, biometric recognition has been widely used in real-world applications, but it has also brought potential privacy threats to users. Iris template protection enables an effective iris recognition while protecting personal privacy. In this paper, we propose a method for iris template protection based on randomized response technique and aggregated block information. Specifically, the iris data are first permuted according to an application-specific parameter; next, the permuted data are flipped using the randomized response technique; finally, the result is divided into blocks, and the aggregated information (i.e., the sum of all bits) in each block is calculated and stored instead of original iris data for privacy protection. We demonstrate that the proposed method supports the shifting and masking strategies for enhancing recognition performance. Moreover, the proposed method satisfies the three privacy requirements prescribed in ISO/IEC 24745: irreversibility, revocability and unlinkability. Experimental results show that the proposed method could effectively maintain the recognition performance (w.r.t. the original iris recognition system without privacy protection) on the iris database CASIA-IrisV3-Interval. Dongdong Zhao 0001, Shengwu Xiong 0001, Jianwen Xiang |
ISSRE | 4 |
| 2018 | A gene-phenotype relationship extraction pipeline from the biomedical literature using a representation learning approachabstractMotivation: The fundamental challenge of modern genetic analysis is to establish gene-phenotype correlations that are often found in the large-scale publications. Because lexical features of gene are relatively regular in text, the main challenge of these relation extraction is phenotype recognition. Due to phenotypic descriptions are often study- or author-specific, few lexicon can be used to effectively identify the entire phenotypic expressions in text, especially for plants. Results: We have proposed a pipeline for extracting phenotype, gene and their relations from biomedical literature. Combined with abbreviation revision and sentence template extraction, we improved the unsupervised word-embedding-to-sentence-embedding cascaded approach as representation learning to recognize the various broad phenotypic information in literature. In addition, the dictionary- and rule-based method was applied for gene recognition. Finally, we integrated one of famous information extraction system OLLIE to identify gene-phenotype relations. To demonstrate the applicability of the pipeline, we established two types of comparison experiment using model organism Arabidopsis thaliana. In the comparison of state-of-the-art baselines, our approach obtained the best performance (F1-Measure of 66.83%). We also applied the pipeline to 481 full-articles from TAIR gene-phenotype manual relationship dataset to prove the validity. The results showed that our proposed pipeline can cover 70.94% of the original dataset and add 373 new relations to expand it. Availability and implementation: The source code is available at http://www.wutbiolab.cn: 82/Gene-Phenotype-Relation-Extraction-Pipeline.zip. Supplementary information: Supplementary data are available at Bioinformatics online. Wenhui Xing, Junsheng Qi, Lin Li 0001, Yuhua Fu, Shengwu Xiong 0001, Lun Hu |
Bioinform. | 7 |
| 2018 | Modified Spider Monkey Optimization based on Nelder-Mead method for global optimization
Prabhat Ranjan Singh, Mohamed E. Abd Elaziz, Shengwu Xiong 0001 |
Expert Syst. Appl. | 3 |
| 2018 | Iris Template Protection Based on Local RankingabstractBiometrics have been widely studied in recent years, and they are increasingly employed in real-world applications. Meanwhile, a number of potential threats to the privacy of biometric data arise. Iris template protection demands that the privacy of iris data should be protected when performing iris recognition. According to the international standard ISO/IEC 24745, iris template protection should satisfy the irreversibility, revocability, and unlinkability. However, existing works about iris template protection demonstrate that it is difficult to satisfy the three privacy requirements simultaneously while supporting effective iris recognition. In this paper, we propose an iris template protection method based on local ranking. Specifically, the iris data are first XORed (Exclusive OR operation) with an application-specific string; next, we divide the results into blocks and then partition the blocks into groups. The blocks in each group are ranked according to their decimal values, and original blocks are transformed to their rank values for storage. We also extend the basic method to support the shifting strategy and masking strategy, which are two important strategies for iris recognition. We demonstrate that the proposed method satisfies the irreversibility, revocability, and unlinkability. Experimental results on typical iris datasets (i.e., CASIA-IrisV3-Interval, CASIA-IrisV4-Lamp, UBIRIS-V1-S1, and MMU-V1) show that the proposed method could maintain the recognition performance while protecting the privacy of iris data. Dongdong Zhao 0001, Shu Fang, Jianwen Xiang, Shengwu Xiong 0001 |
Secur. Commun. Networks | 5 |
| 2018 | Energy and Delay Optimization of Heterogeneous Multicore Wireless Multimedia Sensor Nodes by Adaptive Genetic-Simulated Annealing AlgorithmabstractEnergy efficiency and delay optimization are significant for the proliferation of wireless multimedia sensor network (WMSN). In this article, an energy‐efficient, delay‐efficient, hardware and software cooptimization platform is researched to minimize the energy cost while guaranteeing the deadline of the real‐time WMSN tasks. First, a multicore reconfigurable WMSN hardware platform is designed and implemented. This platform uses both the heterogeneous multicore architecture and the dynamic voltage and frequency scaling (DVFS) technique. By this means, the nodes can adjust the hardware characteristics dynamically in terms of the software run‐time contexts. Consequently, the software can be executed more efficiently with less energy cost and shorter execution time. Then, based on this hardware platform, an energy and delay multiobjective optimization algorithm and a DVFS adaption algorithm are investigated. These algorithms aim to search out the global energy optimization solution within the acceptable calculation time and strip the time redundancy in the task executing process. Thus, the energy efficiency of the WMSN node can be improved significantly even under strict constraint of the execution time. Simulation and real‐world experiments proved that the proposed approaches can decrease the energy cost by more than 29% compared to the traditional single‐core WMSN node. Moreover, the node can react quickly to the time‐sensitive events. Xing Liu 0002, Haiying Zhou, Jianwen Xiang, Shengwu Xiong 0001, Kun Mean Hou, Christophe de Vaulx, Tianhui Shen |
Wirel. Commun. Mob. Comput. | 4 |
| 2017 | Bar charts detection and analysis in biomedical literature of PubMed Central
Xiaohan Yu 0001, Yangjing Gan, Tujin Zhu, Shengwu Xiong 0001, Lun Hu |
AMIA | 5 |
| 2017 | Robust Facial Landmark Localization Using LBP Histogram Correlation Based InitializationabstractFacial landmark localization on images with occlusions is an important and challenging task in many visual applications. Recently, the cascaded pose regression has attracted increasing attention, since it achieved superior performance in terms of facial landmark localization under occlusions. However, such approach is sensitive to initialization, where an improper initialization will decrease the performance sharply. In this paper, we propose a novel initialization method to get a robust initial shape by analysing correlation of Local Binary Patterns (LBP) histograms between the estimated face and training faces. The shape of the training face that is most correlated with the estimated face, will be selected as the initialization for the regression. The selected shape is closer to the real shape of the estimated face, which makes the landmark localization more accurate. Besides, in order to make the initial shape more robust to occlusions, we propose a boosted smart restarts technique by checking location and occlusion jointly instead of checking location only. We show that the proposed method significantly improves performance over existing landmark localization methods on the challenging dataset of COFW. The experimental results demonstrate that the proposed method reduces error by 11.9% and failure cases by 20.8% on COFW dataset. Moreover, it detects face occlusions with 85/40% precision/recall. Yiyun Pan, Junwei Zhou 0002, Yongsheng Gao 0001, Jianwen Xiang, Shengwu Xiong 0001, Yanchao Yang 0002 |
FG | 5 |
| 2017 | Identifying overlapping protein complexes in yeast protein interaction network via fuzzy clusteringabstractThe problem of identifying protein complexes is of great significance for studying the protein mechanisms in different cellular systems. It is for this reason that many computational approaches have been proposed to solve the problem. Yet few of them have endeavored to discover overlapping protein complexes, which are crucial to improve the accuracy performance. Hence, in this paper, we explore the feasibility of making use of a fuzzy clustering approach to identify overlapping protein complexes in a natural manner. To do so, we first formulate the identification problem as an optimization problem by following certain intuitions and then develop an algorithm to solve it so that the memberships of each protein to different protein complexes can be optimized to eventually infer the protein complexes of interest. The experimental results on several yeast protein interaction networks show that our algorithm is promising in terms of accuracy. Lun Hu, Shengwu Xiong 0001 |
FUZZ-IEEE | 3 |
| 2017 | Chinese Geographical Knowledge Entity Relation Extraction via Deep Neural Networks
Shengwu Xiong 0001, Jingjing Mao, Pengfei Duan 0005, Shaohao Miao |
ICAART (2) | 1 |
| 2017 | Deep Knowledge Representation based on Compositional Semantics for Chinese Geography
Shengwu Xiong 0001, Pengfei Duan 0005, Abdelghani Dahou |
ICAART (2) | 1 |
| 2017 | Multi-column Deep Neural Network for Offline Arabic Handwriting Recognition
Rolla Almodfer, Shengwu Xiong 0001, MohammedAli Mudhsh, Pengfei Duan 0005 |
ICANN (2) | 2 |
| 2017 | Shape retrieval using multiscale ellipse descriptorabstractIn this paper, a novel multiscale ellipse descriptor (MED) method is proposed for shape description and matching. MED extracts the competitive features of shape contour by measuring the spatial location relationship between contour sample points and topology structure information of segmented multiscale zone. This method not only has the discriminative ability to describe the global and local information, but also is robustness to various linear (rotation, scale and translation transforms) and non-linear (irregular intra-class deformation) transforms. Experimental results on two public available databases consistently demonstrate that, our proposed method is effective and efficient when compared with other state-of-the-art shape retrieval benchmarks (such as 9.72% higher and 64 times faster than popular IDSC method on leaf 100 dataset). Jianwen Xiang, Shengwu Xiong 0001 |
ICIP | 3 |
| 2017 | Very Deep Neural Networks for Hindi/Arabic Offline Handwritten Digit Recognition
Rolla Almodfer, Shengwu Xiong 0001, MohammedAli Mudhsh, Pengfei Duan 0005 |
ICONIP (2) | 2 |
| 2017 | A Hybrid Method of Sine Cosine Algorithm and Differential Evolution for Feature Selection
Mohamed E. Abd Elaziz, Ahmed A. Ewees, Diego Oliva 0001, Pengfei Duan 0005, Shengwu Xiong 0001 |
ICONIP (5) | 5 |
| 2017 | Enhancing AlexNet for Arabic Handwritten words Recognition Using Incremental DropoutabstractCurrently, the growth of mobile technologies, lead to a necessity to develop handwritten recognition applications. While the recognition of handwritten Latin and Chinese has been extensively investigated using various techniques, so little works have been done on Arabic handwritten recognition, and none of the existing techniques is accurate enough for practical application. Over the past few years, deeper convolutional neural networks (CNNs) have widely been employed for improving handwritten recognition performance. In this paper, we enhance the popular AlexNet for Arabic Handwritten Words Recognition (HWR). By adopting a dropout regularization, we prevented our system against overfitting problem and reduced the error recognition rate. We also investigated ReLU and tanH activation functions performance in the fully connected layers. Through several settings of experiments using the benchmarking IFN/ENIT Database, we achieved a new state-of-the-art classification accuracy of 92.13% and 92.55%. Lastly, we compared our best result to those of previous state-of-the-art. Rolla Almodfer, Shengwu Xiong 0001, MohammedAli Mudhsh, Pengfei Duan 0005 |
ICTAI | 2 |
| 2017 | Robust Image Classification via Low-Rank Double Dictionary Learning
Shengwu Xiong 0001, Yongsheng Gao 0001 |
MMM (1) | 2 |
| 2017 | An Ontology-based Knowledge Management System for Software TestingabstractSoftware testing is an important activity in quality assurance and it generates large amount of knowledge.Software testers need to gather domain knowledge to be able to successfully conduct a software testing activity.Not having a proper knowledge base within its own context by software testing environments cause software testers to query limited knowledge available or consult peer software testers, which would greatly impact on their decision-making process.Ontologies emerge as one of the more appropriate knowledge management tools for supporting knowledge representation, processing, storage and retrieval.Given great importance to knowledge for software testing, and the potential benefits of managing software testing knowledge, using semantic web technologies, ontology based knowledge management system is developed.A Software testing knowledge sharing ontology is designed to describe software testing domain knowledge.SPARQL is used as the query language to retrieve software testing knowledge from the semantic storage.Both Ontology experts and non-experts evaluated the developed ontology.We believe our software testing ontology can support other software organizations to improve the sharing of knowledge and learning practices. Shanmuganathan Vasanthapriyan, Dongdong Zhao 0001, Shengwu Xiong 0001, Jianwen Xiang |
SEKE | 4 |
| 2017 | An improved Opposition-Based Sine Cosine Algorithm for global optimization
Mohamed E. Abd Elaziz, Diego Oliva 0001, Shengwu Xiong 0001 |
Expert Syst. Appl. | 3 |
| 2017 | An artificial bee colony-based multi-objective route planning algorithm for use in pedestrian navigation at nightabstractPedestrian navigation at night should differ from daytime navigation due to the psychological safety needs of pedestrians. For example, pedestrians may prefer better-illuminated walking environments, shorter travel distances, and greater numbers of pedestrian companions. Route selection at night is therefore a multi-objective optimization problem. However, multi-objective optimization problems are commonly solved by combining multiple objectives into a single weighted-sum objective function. This study extends the artificial bee colony (ABC) algorithm by modifying several strategies, including the representation of the solutions, the limited neighborhood search, and the Pareto front approximation method. The extended algorithm can be used to generate an optimal route set for pedestrians at night that considers travel distance, the illumination of the walking environment, and the number of pedestrian companions. We compare the proposed algorithm with the well-known Dijkstra shortest-path algorithm and discuss the stability, diversity, and dynamics of the generated solutions. Experiments within a study area confirm the effectiveness of the improved algorithm. This algorithm can also be applied to solving other multi-objective optimization problems. Zhixiang Fang, Qingquan Li 0001, Shengwu Xiong 0001 |
Int. J. Geogr. Inf. Sci. | 6 |
| 2017 | Low-rank double dictionary learning from corrupted data for robust image classification
Shengwu Xiong 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 2 |
| 2016 | Word Embeddings and Convolutional Neural Network for Arabic Sentiment ClassificationabstractWith the development and the advancement of social networks, forums, blogs and online sales, a growing number of Arabs are expressing their opinions on the web. In this paper, a scheme of Arabic sentiment classification, which evaluates and detects the sentiment polarity from Arabic reviews and Arabic social media, is studied. We investigated in several architectures to build a quality neural word embeddings using a 3.4 billion words corpus from a collected 10 billion words web-crawled corpus. Moreover, a convolutional neural network trained on top of pre-trained Arabic word embeddings is used for sentiment classification to evaluate the quality of these word embeddings. The simulation results show that the proposed scheme outperforms the existed methods on 4 out of 5 balanced and unbalanced datasets. Abdelghani Dahou, Shengwu Xiong 0001, Junwei Zhou 0002, Mohamed Houcine Haddoud, Pengfei Duan 0005 |
COLING | 2 |
| 2016 | Discriminative dictionary pair learning from partially labeled dataabstractWhile conventional synthesis dictionary learning approaches have demonstrated tremendous success in various pattern recognition problems, the dictionary pair learning, i.e., jointly learning an analysis dictionary and a synthesis dictionary is still an open problem. Furthermore, the performance of traditional supervised dictionary learning methods is often limited by the amount of labeled training data. In this paper, we propose a novel dictionary pair learning model by utilizing both labeled and unlabeled data for analysis-synthesis dictionary training. In the dictionary learning phase, we integrate the unlabeled samples, whose labels are predicted through an entropy-based method, into their associated classes to increase the amount of the `labeled' data. This strategy promotes the discrimination power of both analysis dictionary and synthesis dictionary. Experimental evaluations on publicly available datasets demonstrate the usefulness of semi-supervised strategy and the effectiveness of the proposed method, especially in the case of limited number of the labeled samples. Shengwu Xiong 0001, Yongsheng Gao 0001 |
ICIP | 2 |
| 2016 | A Game-Theoretic Analysis of Pricing Strategies for Competing Cloud PlatformsabstractIn this paper, we analyse how multiple competing cloud platforms set effective service prices between Web service providers and consumers. We propose a novel economic framework to model this problem. Cloud platforms run double auction mechanisms, where Web service is commodity traded by service providers (sellers) and service consumers (buyers). Multiple cloud platforms compete against each other to attract service providers and consumers. Specifically, we use game theory to analyse the pricing policies of competing cloud platforms, where service providers and consumers can choose to participate in any of the platforms, and bid or ask for the Web service. The platform selection and bidding strategies of service providers and consumers are affected by the pricing policies and vice versa, and so we propose a co-learning algorithm based on fictitious play to analyse this problem. In more detail, we investigate a setting with two competing cloud platforms who can adopt either equilibrium k pricing policy or discriminatory k pricing policy. We find that, when both cloud platforms use the same type of pricing policy, they can co-exist in equilibrium, and they have an extreme bias to service providers or consumers when setting k. When both platforms adopt different types of policies, we find that all service providers and consumers converge to the discriminatory k pricing policy and so the two competing platforms can no longer co-exist. Bing Shi 0002, Yalong Huang, Shengwu Xiong 0001 |
ICPADS | 4 |
| 2016 | Setting an Effective Pricing Policy for Double Auction Marketplaces
Bing Shi 0002, Yalong Huang, Shengwu Xiong 0001, Enrico H. Gerding |
PRICAI | 3 |
| 2016 | InDel marker detection by integration of multiple softwares using machine learning techniquesabstractBACKGROUND: In the biological experiments of soybean species, molecular markers are widely used to verify the soybean genome or construct its genetic map. Among a variety of molecular markers, insertions and deletions (InDels) are preferred with the advantages of wide distribution and high density at the whole-genome level. Hence, the problem of detecting InDels based on next-generation sequencing data is of great importance for the design of InDel markers. To tackle it, this paper integrated machine learning techniques with existing software and developed two algorithms for InDel detection, one is the best F-score method (BF-M) and the other is the Support Vector Machine (SVM) method (SVM-M), which is based on the classical SVM model. RESULTS: The experimental results show that the performance of BF-M was promising as indicated by the high precision and recall scores, whereas SVM-M yielded the best performance in terms of recall and F-score. Moreover, based on the InDel markers detected by SVM-M from soybeans that were collected from 56 different regions, highly polymorphic loci were selected to construct an InDel marker database for soybean. CONCLUSIONS: Compared to existing software tools, the two algorithms proposed in this work produced substantially higher precision and recall scores, and remained stable in various types of genomic regions. Moreover, based on SVM-M, we have constructed a database for soybean InDel markers and published it for academic research. Jianqiu Yang, Xinyi Shi, Lun Hu, Daipeng Luo, Shengwu Xiong 0001, Fanjing Kong, Baohui Liu |
BMC Bioinform. | 6 |
| 2016 | Construction method of concept lattice based on improved variable precision rough set
Ruiling Zhang, Shengwu Xiong 0001, Zhong Chen 0003 |
Neurocomputing | 2 |
| 2016 | Personalized route planning system based on Wardrop Equilibrium model for pedestrian-vehicle mixed evacuation in campus
Pengfei Duan 0005, Shengwu Xiong 0001, Chunhui Yang, Haohao Zhang, Mianfang Liu |
J. Inf. Secur. Appl. | 2 |
| 2016 | Research of a resource-efficient, real-time and fault-tolerant wireless sensor network system
Xing Liu 0002, Haiying Zhou, Shengwu Xiong 0001, Kun Mean Hou, Christophe de Vaulx, Hongling Shi |
J. Inf. Secur. Appl. | 3 |
| 2016 | Research on campus traffic congestion detection using BP neural network and Markov model
Xiaohan Yu 0001, Shengwu Xiong 0001, W. Eric Wong, Yang Zhao 0019 |
J. Inf. Secur. Appl. | 2 |
| 2016 | Enhancing Keyword Suggestion of Web Search by Leveraging Microblog Data
Lin Li 0001, Fang Deng, Shengwu Xiong 0001, Jingling Yuan |
J. Web Eng. | 4 |
| 2015 | An Ontology-Based Approach for Measuring Semantic Similarity Between Words
Ruiling Zhang, Shengwu Xiong 0001, Zhong Chen 0003 |
ICIC (3) | 2 |
| 2015 | Hashtag Biased Ranking for Keyword Extraction from Microblog PostsabstractNowadays, a huge amount of text is being generated for social networking purpose on the Web. Keyword extraction from such text benefit many applications such as advertising, search, and content filtering. Recent studies show that graph based ranking is more effective than traditional term or document frequecy based approaches. However, most work in the literature constructs word to word graph within a document or a collection of documents before applying a kind of random walk. Such a graph does not consider the influence of document importance on keyword extraction. Moreover, social text like a microblog post usually has speical social features such as hashtag and so on, which can help us understand its topic. In this paper, we propose hashtag biased ranking for keyword extraction from a collection of microblog posts. We first build a word-post weighted graph by taking into account the posts themselves. Then, a hashtag biased random walk is applied on this graph, which guides our approach to extract keywords according to the hashtag topic. Last, the final ranking of a word is determined by the stationary probability after a number of interations. We evaluate our proposed method on a real Chinese microblog posts. Experiments show that our method is more effective than the traditional word to word graph based ranking in terms of precision. Lin Li 0001, Yueqing Sun, Shengwu Xiong 0001, Guandong Xu |
KSEM | 4 |
| 2015 | Pedestrian detection algorithm based on video sequences and laser point cloud
Hui Li 0010, Yun Liu 0007, Shengwu Xiong 0001, Lin Wang 0019 |
Frontiers Comput. Sci. | 3 |
| 2014 | Multi-objective optimization model based on steady degree for teaching building evacuationabstractIn this paper, the process of evacuation in teaching building is considered. The concept of steady degree based on cellular automata and potential field is introduced and it can describe the behavior tendency of evacuees during the evacuation process. With the help of steady degree, the model simulates the indoor evacuation behavior. To reduce the congestion and evacuation time, a multi-objective optimization model considering steady degree and evacuation clearance time is proposed. Finally, an experiment in the Teaching Building No.1 of Wuhan University of Technology is carried out. The results show that this model can reduce the clearance time of emergency evacuation in teaching building compared to other models. Pengfei Duan 0005, Shengwu Xiong 0001, Zongbo Hu, Xinlu Zong |
IEEE Congress on Evolutionary Computation | 2 |
| 2014 | Space-time simulation model based on particle swarm optimization algorithm for stadium evacuationabstractIn this paper, a space-time simulation model based on particle swarm optimization algorithm for stadium evacuation is presented. In this new model, the fast evacuation, going with the crowd and the panic behaviors are considered and the corresponding moving rules are defined. The model is applied to a stadium and simulations are carried out to analyze the spacetime evacuation efficiency by different behaviors. The simulation results show that the behaviors of going with the crowd and panic will slow down the evacuation process while quickest evacuation psychology can accelerate the process, and panic is helpful to some extent. The setting of parameters is discussed to obtain best performance. The simulation results can offer effective suggestions for evacuees under emergency situation. Xinlu Zong, Shengwu Xiong 0001, Pengfei Duan 0005 |
IEEE Congress on Evolutionary Computation | 3 |
| 2014 | A kernel support vector machine-based feature selection approach for recognizing Flying Apsaras' streamers in the Dunhuang Grotto Murals, China
Zhong Chen 0003, Shengwu Xiong 0001, Zhixiang Fang, Qingquan Li 0001, Qin Zou 0001 |
Pattern Recognit. Lett. | 2 |
| 2012 | Positive point charge potential field based ACO algorithm for multi-objective evacuation routing optimization problemabstractMulti-objective evacuation routing optimization problem is defined to find out optimal evacuation routes for a group of evacuees according to multiple evacuation objectives. For improving the evacuation efficiency, we abstracted the evacuation zone as a positive-point-charge-potential-field-like model (PPCPF-like model), and we proposed PPCPF-ACO algorithm to solve this problem based on the proposed model. In PPCPF-ACO algorithm, we use non-dominated sorting based roulette wheel routing method (NSRWR) to further improve evacuation efficiency. In Wuhan Sports Center case, we compared PPCPF-ACO with HMERP-ACO (hierarchical multi-objective evacuation routing problem - ant colony optimization) and traditional ACO according to three evacuation objectives, namely, total evacuation time, total evacuation route length and cumulative congestion degree. The experimental results show that PPCPF-ACO has a better performance than HMERP-ACO algorithm and traditional ACO algorithm while solving multi-objective evacuation routing optimization problem. Jialiang Kou, Shengwu Xiong 0001, Zhixiang Fang, Xinlu Zong, Feifei Bian |
IEEE Congress on Evolutionary Computation | 2 |
| 2012 | Prediction based multi-strategy differential evolution algorithm for dynamic environmentsabstractMany real world optimization problems are dynamic optimization problems (DOPs) whose optima change over time. In this paper, we propose new variants of differential evolution (DE) to solve DOPs. A hybrid method that combines population core based multi-population strategy and prediction strategy and new local search scheme is introduced into DE to enhance its performance for solving DOPs. The population core based multi-population strategy is useful to maintain the diversity of population by using the multi-population and population core concept. The prediction strategy is useful to rapidly adapt to the dynamic environment by using the prediction area. The local search scheme is useful to improve the searching accuracy by suing the new chaotic local search method. Experimental results on the moving peaks benchmark show that the proposed schemes enhance the performance of DE in the dynamic environments. Shuzhen Wan, Shengwu Xiong 0001 |
IEEE Congress on Evolutionary Computation | 2 |
| 2010 | Multi-ant colony system for evacuation routing problem with mixed traffic flowabstractEvacuation routing problem with mixed traffic flow is complex due to the interaction among different types of evacuees. The positive feedback mechanism of single ant colony system may lead to congestion on some optimum routes. Like different ant colony systems in nature, different components of traffic flow compete and interact with each other during evacuation process. In this paper, an approach based on multi-ant colony system was proposed to tackle evacuation routing problem with mixed traffic flow. Total evacuation time is minimized and traffic load of the whole road network is balanced by this approach. The experimental results show that this approach based on multi-ant colony system can obtain better solutions than single ant colony system and solve mixed traffic flow evacuation problem with reasonable routing plans. Xinlu Zong, Shengwu Xiong 0001, Zhixiang Fang |
IEEE Congress on Evolutionary Computation | 2 |
| 2009 | Diversity Maintenance Strategy Based on Global Crowding
Shengwu Xiong 0001 |
ISNN (3) | 2 |
| 2009 | Extraction of the Reduced Training Set Based on Rough Set in SVMs
Shengwu Xiong 0001 |
ISNN (2) | 2 |
| 2008 | Self-adaptive Hybrid differential evolution with simulated annealing algorithm for numerical optimizationabstractA self-adaptive hybrid differential evolution with simulated annealing algorithm, termed SaDESA, is proposed. In the novel SaDESA, the choice of learning strategy and several critical control parameters are not required to be pre-specified. During evolution, the suitable learning strategy and parameters setting are gradually self-adapted according to the learning experience. The performance of the SaDESA is evaluated on the set of 25 benchmark functions provided by CEC2005 special session on real parameter optimization. Comparative study exposes the SaDESA algorithm as a competitive algorithm for a global optimization. Zongbo Hu, Qinghua Su, Shengwu Xiong 0001, Fu-gao Hu |
IEEE Congress on Evolutionary Computation | 3 |
| 2006 | Fuzzy Support Vector Machines Based on Spherical Regions
Shengwu Xiong 0001, Xiaoxiao Niu |
ISNN (1) | 2 |
| 2003 | Parallel strength Pareto multi-objective evolutionary algorithm for optimization problemsabstractFinding a good convergence and distribution of solutions near the Pareto-optimal front in a small computational time is an important issue in multiobjective evolutionary optimization. Previous studies have either demonstrated a good distribution with a large computational overhead or a not-so-good distribution quickly, Strength Pareto evolutionary algorithm (SPEA) produces a better distribution with larger computational effort. A Parallel strength Pareto multiobjective evolutionary algorithm (PSPMEA) is proposed. PSPMEA is a parallel computing model designed for solving Pareto-based multiobjective optimization problems by using an evolutionary procedure. In this procedure, both global parallelization and island parallel evolutionary algorithm models are implemented based on Java multi-threaded and distributed computation programmatic technology separately. Each subpopulation evolves separately with different crossover and mutation probability, but they exchange individuals in the elitist archive. The benchmark problems numerical experiment results demonstrate that the proposed method can rapidly converge to the Pareto optimal front and spread widely along the front. Shengwu Xiong 0001 |
IEEE Congress on Evolutionary Computation | 1 |
| 2003 | A new hybrid structure genetic programming in symbolic regressionabstractGenetic programming (GP) has been applied to symbolic regression problem for a long time. The symbolic regression is to discover a function that can fit a finite set of sample data. These sample data can be guided by a simple function, which is continuous and smooth. But in a complex system, they can be produced by a discontinuous or non-smooth function. When conventional GP is applied to this complex system's modelling, it gets poor performance. This paper proposes a new GP representation and algorithm that can be applied to both continuous function's and discontinuous function's regression. Our approach is able to identify both simultaneously the function's structure and the discontinuity points. The numerical experimental results will show that the new GP is able to gain higher success rate, higher convergence rate and better solutions than conventional GP. Shengwu Xiong 0001, Weiwu Wang |
IEEE Congress on Evolutionary Computation | 1 |
| 2003 | A New Genetic Programming Approach in Symbolic RegressionabstractGenetic programming (GP) has been applied to symbolic regression problem for a long time. The symbolic regression is to discover a function that can fit a finite set of sample data. These sample data can be guided by a simple function, which is continuous and smooth, but in a complex system, the sample data can be produced by a discontinuous or non-smooth function. When conventional GP is applied to such complex system's regression, it gets poor performance. This paper proposed a new GP representation and algorithm that can be applied to both continuous function's regression and discontinuous function's regression. The proposed approach is able to identify both the sub-functions and the discontinuity points simultaneously. The numerical experimental results show that the new GP is able to obtain higher success rate, higher convergence rate and better solutions than conventional GP in such complex system's regression. Shengwu Xiong 0001, Weiwu Wang |
ICTAI | 1 |