Feng Xiao 0005

dblp:71/1116-5 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-1969-8616ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Multi-branch perturbation learning with constraint simulation for semi-supervised semantic segmentation
abstract
Current semi-supervised semantic segmentation (SSS) methods improve generalization via weak-to-strong pseudo-supervision with image perturbations. However, many methods are limited by employing a single perturbation mode and a specific weak-to-strong learning strategy, restricting exploration of the perturbation space and hindering performance in fine-grained segmentation. While diverse perturbations are intuitively beneficial, simply combining them can lead to inefficient optimization and instability. In this paper, we propose a multi-branch strong perturbation constraint learning framework for SSS. Our framework introduces a novel multi-branch perturbation learning (MSPL) strategy, employing multiple parallel branches with diverse strong augmentations to expand the perturbation space and capture complex semantic variations. We further design a novel constraint simulation loss (CSSL), based on a hierarchical consistency learning structure (weak-to-strong and strong-to-strong), which enforces strong-to-strong consistency between different perturbation branches. CSSL mitigates instability and enhances robustness to perturbation-induced noise, enabling the network to better generalize and achieve more accurate segmentation, especially for fine object boundaries. Extensive evaluations on benchmark datasets (PASCAL VOC 2012, Cityscapes, COCO) demonstrate that our method achieves state-of-the-art performance. Ablation studies further validate the effectiveness of our proposed MSPL and CSSL components.
Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Shengyong Chen, Houxiang Zhang
Pattern Recognit.2
2026 Adaptive Kernel Selection Module Combined With Feature Enhanced Perception Network for Camouflaged Object Detection
abstract
Camouflaged object detection plays a crucial role in applications such as automatic sorting and defect inspection in industrial production, yet existing methods often struggle to flexibly capture features of diverse shapes, orientations, and scales due to their reliance on fixed receptive fields and rigid windowing schemes. To address these limitations, we propose a dual-branch joint network comprising a reference branch and a segmentation branch. The reference branch learns supplementary cues from salient objects that co-occur with camouflaged targets, guiding the segmentation branch toward more accurate delineation. Within the segmentation branch, we introduce three novel modules: 1) a deformable window interaction mechanism that replaces fixed-size transformer windows with learnable quadrilateral windows to adaptively extract features of arbitrary shape and orientation; 2) a feature enhancement perception module that fuses rich multiscale representations through parallel dilated convolutions at varying rates and channel-/spatial-attention mechanisms; and 3) a receptive field adjustment adaptive module that dynamically adjusts its receptive field size to balance sensitivity to fine details and global context. Comprehensive experiments on COD10 K, NC4K, CAMO, and R2C7K benchmarks demonstrate that our model outperforms the majority of current state-of-the-art approaches, while ablation studies and sensitivity analyses confirm the individual and combined effectiveness of our proposed components.
Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Shengyong Chen
IEEE Trans. Ind. Informatics3
2026 WTCLIP: A Wavelet-Aware CLIP Framework for Boundary-Refined Weakly Supervised Semantic Segmentation
abstract
Some advanced methods have leveraged the zero-shot recognition capability of the contrastive language–image pretraining (CLIP) model and adapted it to weakly supervised semantic segmentation (WSSS), achieving promising performance. However, they primarily use CLIP as an auxiliary feature extractor, leaving the fundamental limitations of class activation mapping unresolved, particularly in preserving fine-grained object boundaries and achieving precise pixelwise localization under sparse supervision. To address these challenges, this article proposes a novel end-to-end WSSS framework WTCLIP, which aims to fully exploit the potential of CLIP for weakly supervised segmentation tasks. Different from traditional methods that use CLIP only as a static feature extractor, we innovatively introduce a learnable wavelet transform decoder to enhance the information extraction capability and significantly improve the model's perception of object boundaries. We dynamically adjust the weight distribution ratio of the CLIP feature layer, capture multiscale edge information, and make full use of the time–frequency localization characteristics of the wavelet transform to significantly improve the quality of pseudolabels and achieve more accurate semantic segmentation. Experimental results show that our method significantly improves the performance of the WSSS task on two public benchmark datasets, notably by4.0%over the state-of-the-art methods, especially in capturing weakly annotated object boundary details.
Feng Xiao 0005, Jianhua Zhang 0002, Peihua Han, Shengyong Chen, Houxiang Zhang
IEEE Trans. Ind. Informatics1
2026 TSCFNet: Temporal Spectral Feature Cross Fusion Network for Imbalanced Sea State Estimation in Autonomous Ships
abstract
Sea state estimation (SSE) is critical to the safety of maritime transport and the reliability of autonomous ships. The frequency of different sea states varies significantly, leading to uneven data distribution. Existing deep learning methods for SSE typically focus on feature extraction, often using simple splicing and fusion, which can result in cross-domain incoherence and degrade model performance. Addressing sea state classification imbalance is often done through distance-based classifiers (e.g., prototype classifiers), but these can be less sensitive to minority classes, and using few prototypes for a class limits the expression of intra-class variations. To overcome these challenges, we propose the Temporal Spectral Cross Fusion Network (TSCFNet), which extracts temporal and spectral features. These are integrated via an innovative temporal spectral cross fusion module to maximize their complementary advantages. Additionally, we introduce a multi-fusion loss function, including temporal, spectral, and fusion losses, to optimize features across different dimensions. This approach improves the performance for minority classes and captures intra-class differences more effectively, solving the problem of category imbalance. Experimental results show that TSCFNet significantly outperforms baseline methods on two imbalanced sea state datasets and multiple multivariate spatio-temporal datasets.
Feng Xiao 0005, Xu Cheng 0003, Xia Xie 0003, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.1
2025 Robust Online Detection of Anomalies in Evolving Data Streams with GAN Imputation
abstract
Online anomaly detection is a critical technique in intelligent industrial systems. Concept drift and missing data that arise during the transmission of real-time data streams can severely impact the performance of anomaly detection. Therefore, ensuring the accuracy of online anomaly detection and adaptability to concept drift in the presence of missing values is a significant challenge. In this paper, we propose an online anomaly detection model that integrates an online imputation module based on Generative Adversarial Networks (GANs) to efficiently impute missing data. Additionally, following the dynamic model pool strategy, the model dynamically selects and adjusts the optimal models in the pool to effectively respond to changes in data distribution, ensuring detection performance under complex conditions. We evaluated the model's performance across multiple datasets with concept drift and under varying missing data rates. The results demonstrate that the proposed model not only adapts flexibly to rapidly changing data streams but also exhibits enhanced robustness in the presence of missing data.
Mengna Liu, Xu Cheng 0003, Jianhua Zhang 0002, Feng Xiao 0005
CSCWD5
2025 Data-Driven Diffusion-Augmented Network for Imbalanced Sea State Estimation with Dynamic Prototypes
abstract
With the advent of Industry 4.0, which emphasizes automation and data-driven decision-making, sea state estimation (SSE) has become a critical component in marine engineering and autonomous vessels. However, traditional SSE methods face significant challenges due to poor real-time performance and high manual costs. Additionally, existing deep learning (DL) approaches struggle to handle imbalanced sea state data effectively, which hinders their generalization ability. To address these issues, this paper proposes a novel DL-based SSE model. The model enhances the expressive capacity of the extracted features through a feature augmentation module and utilizes a diffusion model to extract more robust features from the data. Furthermore, a dynamic prototype update module is designed to effectively address class imbalance and overcome the limitations of traditional prototype classifiers in handling boundary samples. Experimental results show that the proposed method outperforms existing baseline models on two imbalanced sea state datasets, achieving F1-score improvements of 4.5% and 3.5%, respectively, compared to the current state-of-the-art methods. Additionally, the model demonstrates superior performance on several publicly available multivariate time series classification datasets.
Feng Xiao 0005, Xu Cheng 0003, Jianhua Zhang 0002
CSCWD2
2025 Lightweight Self-Supervised Monocular Depth Estimation via Context-Aware Fusion and Separable Depthwise Convolution
abstract
Monocular depth estimation is a critical problem in computer vision, with wide-ranging applications across various domains. However, existing methods often involve high computational costs, making them challenging to deploy efficiently on edge devices. To address the trade-off between computational complexity and inference accuracy, this paper presents an efficient and lightweight model for self-supervised monocular depth estimation. Our model incorporates a Context-Aware Fusion (CAF) module to capture both global and local feature dependencies. In addition, Separable Depthwise Convolution (SDC) module are utilized to reduce computational overhead, and the Multi-Scale Structural Similarity (MS-SSIM) loss function is employed to improve both depth estimation accuracy and visual perception quality. Experimental results show that the proposed model delivers improved accuracy while maintaining a lightweight and efficient architecture, ensuring its compatibility with edge device deployment. A comprehensive analysis of the findings is also presented, along with insights for future optimizations and improvements.
Meina Zhao, Shixin Wang 0014, Feng Xiao 0005, Jianhua Zhang 0002, Xu Cheng 0003, Yunrui Zhu
CSCWD3
2025 Efficiency-Optimized Point Cloud Upsampling with Single-Layer Graph Convolution Network
abstract
Point cloud upsampling is a key technology for improving the density and quality of sparse point clouds, with widespread applications in 3D reconstruction, autonomous driving, and environmental perception. However, traditional point cloud upsampling methods, especially those based on multilayer graph convolution networks (GCNs), typically rely on complex feature extraction modules, which increase computational complexity and model parameters, limiting their use in resource-constrained environments. To overcome these challenges, we propose the EO-PU framework, a lightweight and efficient point cloud upsampling method. This framework combines single-layer GCN and rotation-invariant 4D projection encoding (I4DP) technology, significantly reducing computational load and redundant information, thereby improving upsampling efficiency. Specifically, EO-PU first uses I4DP to map the 3D point cloud data to a rotation-robust 4D feature space, ensuring effective capture of geometric information. A single-layer GCN is then employed to aggregate features, reducing network complexity and computational cost. To further enhance upsampling performance, we introduce the EdgeShuffleNet module, which optimizes feature expansion and rearrangement through efficient local feature aggregation. Experimental results show that EO-PU outperforms or matches existing methods across multiple public datasets while significantly reducing model parameters and computation time, making it highly suitable for deployment in resource-constrained environments.
Yunrui Zhu, Feng Xiao 0005, HaoXiao Wang, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002
IJCNN3
2025 An expert features enhanced temporal and contextual contrasting learning model for detecting wind turbine blade icing
abstract
With global carbon neutrality goals, wind power has rapidly developed, but blade icing remains a major challenge. AI(artificial intelligence) methods show great promise for detecting icing on wind turbine blades. However, early icing data overlap, difficulty obtaining continuous labeled data, and small variations between samples due to short sampling intervals complicate the task. This study proposes an expert feature-enhanced temporal and contextual contrastive learning model for detecting blade icing. This approach efficiently extracts data features and combines self-supervised contrastive learning, maximizing data utilization without requiring extensive labeled data. To validate the effectiveness of this method, extensive experiments were conducted on two public datasets. The results achieved the best performance across multiple metrics, with F1-Score and AUC exceeding 98%, significantly enhancing wind power generation efficiency.
Jiamei Zhou, Feng Xiao 0005, Xu Cheng 0003, Jianhua Zhang 0002
Eng. Appl. Artif. Intell.2
2025 Wavelet-Discrete Cosine Transform Synergy for Ship Motion-Based Sea State Estimation in Autonomous Ships
abstract
Developing a robust autonomous sea state estimation (SSE) model stands as a pivotal challenge in advancing autonomous ships. Presently, deep learning (DL) methodologies have showcased remarkable efficacy in SSE tasks. Nonetheless, the dynamic nature of ship motion introduces temporal variations alongside frequency domain characteristics like periodic swinging, posing challenges for existing DL approaches. Most prevailing DL techniques, predominantly leveraging Convolutional Neural Networks or Long Short-Term Memory Networks, often fail to effectively harness frequency domain information post feature extraction. To tackle these limitations head-on, this paper introduces a pioneering SSE model. Specifically, in order to solve the frequency-domain feature extraction problem, we design a wavelet transform-based frequency domain encoder to extract relevant frequency-domain features from ship motion data by discriminating the contribution of different frequencies in the signal. Subsequently, in order to better integrate the extracted ship motion features, we designed a Feature Perception module based on discrete cosine transform. This module adeptly merges the extracted feature insights while prioritizing crucial frequency domain features. Following rigorous experimentation, our methodology exhibits superior performance compared to existing baseline techniques in SSE, a capability of profound significance for autonomous ships. Moreover, across diverse public multivariate time series classification datasets, our model outperforms current state-of-the-art approaches, underscoring its scalability across distinct domains.
Feng Xiao 0005, Xu Cheng 0003, Sasa Nikolic 0002, Jianhua Zhang 0002, Shengyong Chen
IEEE Trans Autom. Sci. Eng.2
2025 Cascaded State Space and Contrastive Learning for Cross-Domain Few-Shot Segmentation
abstract
Current cross-domain few-shot semantic segmentation (CD-FSS) faces multiple challenges, including inconsistent feature mapping among domains and insufficient utilization of low-level information and background information from the source domain. To address these issues, this article proposes a novel cascade feature enhancement and contrastive learning framework to improve the generalization capability of CD-FSS. Within this framework, we first introduce a cascade feature enhancement module to construct distinctive feature representations, enhancing the model’s transferability across domains. By effectively integrating multilevel feature information from support images, this module strengthens the representation capability of query images. Second, we employ contrastive learning to form positive and negative sample pairs for the foreground and background, capturing rich correlations between them. Finally, the iterative prototype enhancement module we propose gradually refines the correspondence between the support image and the query image through iteration, making full use of the embedded supervisory information in the limited support samples. Experimental results demonstrate that the proposed method outperforms existing approaches on multiple benchmark datasets, achieving up to a 9.7% improvement over state-of-the-art methods.
Feng Xiao 0005, Jianhua Zhang 0002, Peihua Han, Shengyong Chen, Houxiang Zhang
IEEE Trans. Ind. Informatics1
2025 Cross-Scale Denoising Reverse Distillation for Anomaly Detection
abstract
Effective discrepancy representation of anomalies plays a crucial role in visual anomaly detection. Recent advances build upon reverse distillation paradigm that boost the teacher–student model’s discrimination capability on anomalies; however, they are still susceptible to the size variation of unpredictable anomalies. To generalize the anomaly size variation, we propose a new algorithm cross-scale denoising reverse distillation (CDRD), which integrates cross-scale denoising with reverse distillation to exchange multiscale perception and enhance the fine-grained representation of features. Specifically, we introduce a cross-scale anomalous signal suppression procedure in the teacher network to facilitate the interaction of information across different scales, thereby enabling the student network to learn more robust normal data representations. In the knowledge transfer process, a fusion compression module acts as an intermediate transmitter of information, aiming to obtain a compact embedding while abandoning anomaly perturbations. Moreover, we construct a detail supplement module in the student network to prevent the loss of key information in the deconvolution process of the decoder. Experiments on well-known datasets demonstrate that our CDRD brings significant improvements over the next best competitor.
Yanhong Yang, Feng Xiao 0005, Jianhua Zhang 0002, Guodao Zhang, Shengyong Chen
IEEE Trans. Ind. Informatics3
2024 Semi-Supervised Camouflaged Object Detection: Multi Information Fusion Combined with Adaptive Receptive Field Selection Network
Feng Xiao 0005, Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen
PRCV (12)2
2024 KSRB-Net: a continuous sign language recognition deep learning strategy based on motion perception mechanism
Feng Xiao 0005, Yunrui Zhu, Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen
Vis. Comput.1
2021 RCGA-Net: An Improved Multi-hybrid Attention Mechanism Network in Biomedical Image Segmentation
abstract
Drawing support from an effective Medical Image Segmentation (MIS) is conducive to a substantial diagnostic basis for the physicians to identify the focus lesion in the patient body and give the subsequent clinical assessment of the patient status. Although various works have tried the challenging quantitative analysis problem, it is still difficult to conduct precise automatic segmentation, especially the soft tissue organs. In this decade, with the increased amount of available datasets, deep learning-based networks have achieved remarkable performance in image processing. Inspired by the state-of-the-art deep learning works, in this paper, we propose an end-to-end multi-layer network named RCGA-Net. It consists of an encoder-decoder backbone that integrates a coordinate attention mechanism based on space and channel and a global context extraction module to highlight more valuable information. To evaluate the performance of RCGA-Net, we apply it to different kinds of clinical and experimental MIS tasks to testify its generalization ability. Extensive experiments represent that our schema has taken the outperform or compatible results among the comparison methods group. Specifically, the numeric result of RCGA-Net on the pulmonary dataset has achieved a 99.12% optimum F1-score.
Feng Xiao 0005, Shengyong Chen, Zhijun Liao, Jijun Tang
BIBM1
2021 CRB-Net: A Sign Language Recognition Deep Learning Strategy Based on Multi-modal Fusion with Attention Mechanism
abstract
At present, sign language recognition (SLR) researchers are mainly committed to establishing a sign language recognition model based on single-mode data. Nevertheless, this manipulation often leads to a defective understanding of the sign language semantics and ignoring some visual information. In a nutshell, the challenges locate redundancy removing and the alignment of the sign language data with the given tag. To solve the conundrum, this paper proposes a deep learning strategy called CRB-Net, which has used a kind of multimodal fusion attention mechanism. We first extract the features from RGB video and depth video, respectively, then conduct multi-modal fusion. Finally, the fused feature information is fed into an encoder-decoder network to achieve the goal of end-to-end continuous SLR. We verify the effectiveness of our method on three datasets, including the German dataset RWTH-Phoenix-Weather-2014, the Chinese dataset USTC-CSL and the Chinese dataset TJUT-SLRT. As shown by experimental results, the accuracy of 98.5% of our framework CRB-Net has outperformed the state-of-the-art works in the comparison, both in accuracy and algorithm execution efficiency.
Feng Xiao 0005, Tiantian Yuan, Shengyong Chen
SMC1