Xu Cheng 0003

dblp:30/828-3 · DBLP profile ↗
← Back
133ranked-venue papers
18as first author
119since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 5 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 3 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 7 first-author · 29 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 10 since 2021Computer networks · 10 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Beyond Missing Data Imputation: Information-Theoretic Coupling of Missingness and Class Imbalance for Optimal Irregular Time Series Classification
abstract
Irregular time series (IRTS) are prevalent in real-world applications, where uneven sampling and missing data pose fundamental challenges to deep learning-based feature modeling. Although existing methods attempt to retain timestamp information, they often overlook the structured patterns embedded within the missingness itself, and tend to perform poorly when confronted with class imbalance exacerbated by data incompleteness. Specifically, temporal irregularity hinders the modeling of long-range dependencies and local patterns, while sparse observations limit representational capacity, disproportionately impairing minority classes and leading to severe classification bias. To address these deeply coupled challenges, we propose SPECTRA (Structured Pattern and Enriched Context-aware Temporal Representation Architecture), a unified framework for robust IRTS classification. SPECTRA introduces a frequency-guided observation encoder that reconstructs temporal dependencies in a stable manner, mitigating spectral distortion and information corruption. Complementarily, a missingness pattern encoder explicitly captures the dynamic evolution of missing data and leverages it as a discriminative signal. In addition, a prototype-constrained classification paradigm directly optimizes the geometric structure of the feature space, enhancing intra-class compactness and alleviating generalization bottlenecks caused by class imbalance. Extensive experiments on three public IRTS datasets—P12, P19, and PAM—demonstrate the superior performance of SPECTRA under both missing and imbalanced conditions.
Mengna Liu, Xiufeng Liu 0001, Xu Cheng 0003
AAAI7
2026 Radar-APLANC: Unsupervised Radar-based Heartbeat Sensing via Augmented Pseudo-Label and Noise Contrast
abstract
Frequency Modulated Continuous Wave (FMCW) radars can measure subtle chest wall oscillations to enable non-contact heartbeat sensing. However, traditional radar-based heartbeat sensing methods face performance degradation due to noise. Learning-based radar methods achieve better noise robustness but require costly labeled signals for supervised training. To overcome these limitations, we propose the first unsupervised framework for radar-based heartbeat sensing via Augmented Pseudo-Label and Noise Contrast (Radar-APLANC). We propose to use both the heartbeat range and noise range within the radar range matrix to construct the positive and negative samples, respectively, for improved noise robustness. Our Noise-Contrastive Triplet (NCT) loss only utilizes positive samples, negative samples, and pseudo-label signals generated by the traditional radar method, thereby avoiding dependence on expensive ground-truth physiological signals. We further design a pseudo-label augmentation approach featuring adaptive noise-aware label selection to improve pseudo-label signal quality. Extensive experiments on the Equipleth dataset and our collected radar dataset demonstrate that our unsupervised method achieves performance comparable to state-of-the-art supervised methods.
Zhaodong Sun, Xu Cheng 0003, Zuxian He
AAAI3
2026 A2P-Net: Asymmetric Domain-Adaptive Prototype Network for Cross-Domain Multimodal Sensor Retrieval
abstract
Multimodal sensor data from inertial measurement units (IMUs), including accelerometers, gyroscopes, and inclinometers, encode environmental conditions that are difficult to capture through images or text. In maritime settings, classifying sea states from such sensor streams is important for autonomous navigation but faces two interacting difficulties. First, training relies heavily on synthetic simulation data whose idealized physics diverge from real ocean measurements, creating a domain gap. Second, extreme sea states are rare in both domains, and the resulting class imbalance is compounded by the much larger volume of synthetic samples, which together skew gradient updates away from the scarce but operationally important real-world tail classes. We propose A2P-Net, an end-to-end framework that tackles both problems jointly. An adaptive heterogeneous encoder with decoupled channel–temporal attention maps variable-dimension sensor inputs into a shared latent space. Domain-adversarial training aligns synthetic and real feature distributions in that space, and prototype-based metric learning builds per-class retrieval anchors while an asymmetric weighting scheme up-weights real-domain samples to correct the optimization bias. On two custom sea state datasets that mix real and simulated ship motion recordings, A2P-Net reaches 98.7% and 98.9% F1 on real-only evaluation, outperforming the strongest baseline by 1.1–1.3 percentage points. It also ranks first on 15 of 30 UEA multivariate time-series benchmarks with an average accuracy of 74.2%.
Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Shengyong Chen
ICMR4
2026 Target-Enhanced Gated Transformer: A Multi-Behavior Recommendation Framework for Noise Suppression and Target Signal Preservation
abstract
Multi-behavior recommendation systems enhance prediction accuracy for target behavior (e.g., purchase) by integrating auxiliary behaviors such as viewing and adding to cart. However, existing models face two fundamental challenges regarding data characteristics: On one hand, they typically treat all auxiliary behaviors equally or apply simple weighting, lacking effective mechanisms to evaluate and filter inherent noise, leading to contaminated feature representations. On the other hand, when fusing sparse target features with abundant auxiliary features, the critical target signal is easily diluted, losing its dominant role in the final prediction. To address this, we propose a Target-Enhanced Gated Transformer model TEGT. Its core innovations include: a behavior-adaptive gating module that filters source-level noise through nonlinear transformations and hard thresholding while generating importance weights; a target-guided dual-modulation attention mechanism utilizes target behavior as queries to retrieve semantically relevant auxiliary signals, then applies secondary modulation by combining gated weights to ensure both noise resistance and target dominance in the fusion process; a lightweight collaborative semantic enhancement module clusters the fused representations and employs cluster-center contrastive learning to explicitly amplify collaborative signals under sparse target behavior. Extensive experiments on three real-world datasets show that TEGT consistently outperforms state-of-the-art baselines, it achieves remarkable improvements of up to 9.73% in Recall@10 and 9.16% in NDCG@10.
Xu Cheng 0003, Likang Wu, Yingyuan Xiao, Wenguang Zheng
ICMR2
2026 Dynamic Routing-Based Adaptive Multi-LLM Collaboration: A Unified Recommendation Framework with Decision Knowledge Complementation
abstract
Existing LLM-driven recommendation systems (RS) suffer from over-reliance on a single pre-trained model, which limits adaptability across diverse scenarios due to differences in large language models' strengths in semantics, knowledge, and reasoning. To address this issue, we propose AMLrec (Adaptive Multi-LLM Recommendation), a dynamic routing-based adaptive multi-LLM collaboration framework that unifies two dominant paradigms—LLM as Recommender and LLM + Recommender—through decision knowledge complementation. For each user or item, a lightweight encoder generates embeddings that are compared with learnable LLM prototypes using cosine similarity to select the most suitable models. In the first paradigm, selected LLMs generate recommendations via structured prompts, and their outputs are aggregated to form the final recommendation list. In the second paradigm, chosen LLMs produce semantic embeddings, which are fused with learnable embeddings after PCA-based dimensionality reduction and aligned using a lightweight adapter to bridge distribution gaps. Notably, AMLrec does not require fine-tuning of the underlying LLMs, significantly reducing computational overhead. Experiments on real-world datasets demonstrate that the proposed approach consistently outperforms single-LLM baselines across all evaluation metrics, validating its effectiveness. The main contributions of this work are threefold: introducing dynamic routing for multi-LLM recommendation system collaboration, proposing a unified architecture that harmonizes both paradigms, and enabling efficient adaptation without LLM fine-tuning. The code is available at https://github.com/Jiale-12138/AMLrec.
Yingyuan Xiao, Likang Wu, Xu Cheng 0003, Wenguang Zheng, Qingbo Hao, Hongke Zhao
WWW4
2026 Temporal-channel decoupled learning for sea state estimation from ship motion data under class imbalance
Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Shengyong Chen
Eng. Appl. Artif. Intell.4
2026 Learning invariant representation for light field adversarial salient object detection
Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen
Eng. Appl. Artif. Intell.3
2026 Physics-informed dynamic ensemble learning for real-time urban water quality monitoring
abstract
Ensuring high-quality water resources is crucial for sustainable urban development, public health, and resilient city infrastructure, yet traditional anomaly detection methods struggle with the highly variable, non-stationary, and concept-drifting nature of urban water quality data streams. This study proposes a Physics-Informed Dynamic Ensemble Learning (PIDEL) framework, an artificial intelligence approach that combines diverse classical and deep learning models with Physics-Informed Neural Networks (PINNs) embedding convection–diffusion constraints, a Genetic Algorithm (GA) for ensemble optimization, and a Jensen–Shannon Divergence (JSD) based mechanism for dynamic model switching. Applied to a real-world urban water quality dataset, PIDEL achieves an F1-score of 0.95, representing a 59% improvement over the best static ensemble, while reducing false alarms by 73% compared to traditional methods and maintaining F1-scores above 0.9 across all sliding windows. The framework processes each 60-minute window in approximately 2.3 s on standard hardware, demonstrating its suitability for real-time deployment in smart city water systems. These results highlight that integrating physics-informed constraints with dynamic ensemble learning can substantially enhance the reliability, interpretability, and operational value of automated water quality anomaly detection for urban utilities. • Novel LSTM-PINN integrates physics constraints with deep learning for anomaly detection. • Dynamic ensemble adapts to concept drift via Jensen–Shannon divergence-based switching. • Genetic algorithm optimizes ensemble, achieving 95% F1-score and 59% improvement. • Optimal physics loss coefficient ( λ = 0 . 35 ) balances physical and data-driven learning. • Real-time processing (2.3 s per window) enables practical smart city water monitoring.
Renfang Wang, Xiufeng Liu 0001, Xu Cheng 0003, Hong Qiu
Eng. Appl. Artif. Intell.4
2026 Corrigendum to "Physics-informed dynamic ensemble learning for real-time urban water quality monitoring" [Eng. Appl. Artif. Intell. 175 (2026) 114628]
Renfang Wang, Xiufeng Liu 0001, Xu Cheng 0003, Hong Qiu
Eng. Appl. Artif. Intell.4
2026 A progressive layered hybrid expert network with attention routing for multi-task QoS prediction
Xu Cheng 0003, Likang Wu, Wenguang Zheng, Yingyuan Xiao
Expert Syst. Appl.2
2026 BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher's Bias
abstract
Abstract Existing knowledge distillation methods indiscriminately transfer knowledge from teacher networks, including output-level decisional biases, i.e., incorrect final predictions that can mislead student learning and limit student performance. We challenge this paradigm by proposing BTKD++, a framework that systematically filters and rectifies teacher’s output-level biased knowledge into corrective signals. Our approach partitions training data into Easy Tasks (correct teacher predictions) and Hard Tasks (incorrect predictions), then applies bias elimination and rectification modules orchestrated by dynamic learning curriculum. We provide an interpretive information-theoretic abstraction to explain the observed competence-threshold phenomenon, under which bias rectification becomes more effective when teacher errors contain sufficiently structured corrective information. BTKD++ demonstrates broad applicability across classification, detection, and segmentation tasks when task outputs are equipped with suitable probabilistic interfaces, and shows consistent effectiveness across CNNs, Transformers, and State-Space Models. Extensive experiments show consistent student-teacher transcendence, establishing new state-of-the-art results. This work redefines knowledge distillation from blind mimicry to critical learning, proving that students can surpass teachers through principled bias correction. The source code is available at https://github.com/smartyige/BTKD .
Jianhua Zhang 0002, Yu He 0001, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen, Houxiang Zhang, Ruyu Liu
Int. J. Comput. Vis.5
2026 Few-shot video summarization via cross-video temporal invariance
Tinglong Tang, Fanyuan Wu, Shengyong Chen, Xu Cheng 0003
Neurocomputing4
2026 Time-Filtering Graph Learning With Spatial-Temporal Diffusion for Robust Blade Icing Detection
abstract
Accurate detection of wind turbine blade icing is essential for ensuring the safety and efficiency of wind farm operations. Although current machine learning and deep learning approaches are capable of identifying icing states, they still suffer from three major limitations: heavy reliance on manual feature engineering limits the capture of meaningful spatiotemporal dependencies; highly imbalanced data leads to increased false negatives and alarm delays; and sensitivity to data noise often introduces false positives. To address these challenges, this paper proposes an imbalance- and noise-resistant TG-Diff network, which achieves accurate and robust icing state classification by effectively integrating spatiotemporal information from multiple sensors. The TG-Diff comprises three core modules: a Temporal Filtering-based Graph Learning Module (TF-GLM), a Temporal-Spatial Diffusion Graph Convolutional Network (TSD-GCN), and a Distance-Based Classifier (DBC). Specifically, the TF-GLM dynamically infers node relationships to construct robust graph topologies; the TSD-GCN enhances feature representation and suppresses noise through diffusion mechanisms; and the DBC effectively mitigates class imbalance by leveraging distance-based decision boundaries. Experimental results demonstrate that the proposed modules work synergistically to collectively enhance the accuracy and robustness of icing detection in real-world complex environments.
Lingzhu Hu, Xiufeng Liu 0001, Fan Shi 0001, Xu Cheng 0003
IEEE Internet Things J.6
2026 Bridging the gap in cross-domain graph anomaly detection: Enhanced source utilization and label accuracy
Cairui Yan, Xu Cheng 0003, Likang Wu, Yingyuan Xiao, Hongke Zhao, Wenguang Zheng
Inf. Process. Manag.2
2026 Object shape differentiation and texture rendering for neural implicit SLAM
Jiaming Lu, Ruyu Liu, Jianhua Zhang 0002, Xu Cheng 0003
Mach. Vis. Appl.5
2026 ColorSketchNet: Unifying color, sketch and texture for modality-agnostic multi-modal person re-identification
Manman Liu, Xu Cheng 0003, Baowei Wang
Neural Networks2
2026 SaURL-TS: A self-adaptive framework for unsupervised time series representation learning
abstract
Unsupervised time series representation learning, driven by recent advances in contrastive learning-based methods, has become a critical component for downstream tasks like forecasting and classification. However, time series data exhibit complex temporal dependencies and spectral patterns, posing challenges for existing approaches to adapt robustly. Moreover, existing contrastive learning-based approaches overlook frequency-domain information and struggle with selecting effective negative samples, further hindering model performance. To address these issues, we propose S a URL-TS, a novel self-adaptive framework for unsupervised time series representation learning. First, it dynamically learns dataset-specific augmentations to generate high-quality positive samples. Second, an adaptive self-supervised learning module with a multi-domain encoder captures both temporal and spectral patterns without relying on negative samples. Third, a representation-wise attention mechanism assigns dynamic weights to representations across domains. To the best of our knowledge, S a URL-TS is the first self-supervised learning framework to jointly model temporal and spectral patterns across both augmentation and learning stages. Extensive experiments confirm the superior performance of S a URL-TS over state-of-the-art models. Notably, its adaptive data augmentation module is plug-and-play and can be integrated into other contrastive learning frameworks, and its learning stage is capable of adapting to a wide range of time series patterns. Our codebase is available at https://github.com/YusenL/SAURL-TS .
Yusen Liu 0001, Zhichen Lai 0001, Hua Lu 0001, Xu Cheng 0003, Tianqing Zhu, Xiufeng Liu 0001, Huan Huo
Pattern Recognit.4
2026 Multi-branch perturbation learning with constraint simulation for semi-supervised semantic segmentation
abstract
Current semi-supervised semantic segmentation (SSS) methods improve generalization via weak-to-strong pseudo-supervision with image perturbations. However, many methods are limited by employing a single perturbation mode and a specific weak-to-strong learning strategy, restricting exploration of the perturbation space and hindering performance in fine-grained segmentation. While diverse perturbations are intuitively beneficial, simply combining them can lead to inefficient optimization and instability. In this paper, we propose a multi-branch strong perturbation constraint learning framework for SSS. Our framework introduces a novel multi-branch perturbation learning (MSPL) strategy, employing multiple parallel branches with diverse strong augmentations to expand the perturbation space and capture complex semantic variations. We further design a novel constraint simulation loss (CSSL), based on a hierarchical consistency learning structure (weak-to-strong and strong-to-strong), which enforces strong-to-strong consistency between different perturbation branches. CSSL mitigates instability and enhances robustness to perturbation-induced noise, enabling the network to better generalize and achieve more accurate segmentation, especially for fine object boundaries. Extensive evaluations on benchmark datasets (PASCAL VOC 2012, Cityscapes, COCO) demonstrate that our method achieves state-of-the-art performance. Ablation studies further validate the effectiveness of our proposed MSPL and CSSL components.
Ruyu Liu, Feng Xiao 0005, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Shengyong Chen, Houxiang Zhang
Pattern Recognit.5
2026 Light field collaborative perception for visual object tracking
Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001
Pattern Recognit.3
2026 Pioneering Video Semantic Segmentation With Light Field Imaging and Spatial-Angular-Temporal Fusion
Fan Shi 0001, Xu Cheng 0003, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.4
2026 UDA-rPPG: Unsupervised Geometric-Physiological Domain Anchoring for Low-Light rPPG Measurement
abstract
Remote photoplethysmography (rPPG) is a critical technique for non-contact monitoring of human vital signs using facial video data. Most of the existing rPPG approaches, either supervised ones relying on ground-truth physiological signals or less constrained unsupervised ones, primarily address the problem of inaccurate physiological measurements under normal lighting conditions. However, few works focus on handling physiological measurements in extremely low-light scenarios. To this end, we propose an unsupervised geometric-physiological domain anchoring for low-light rPPG measurement (UDA-rPPG). Firstly, we develop a geometric anchoring video enhancement module (GAEM) that can enhance video brightness while preserving rPPG signals, achieving accurate geometric-domain face anchoring. Secondly, we introduce a low-light stable spatial-temporal network (LS-Phys), which focuses on high-frequency information to mitigate noise in low-light scenarios. Finally, a novel highest-peak priority learning strategy is presented to learn physiological-domain rPPG signal anchoring by emphasizing peak information, which enhances the robustness of rPPG measurements in low-light environments. Additionally, we construct a comprehensive low-light rPPG dataset (LRPD) that contains both visible and near-infrared videos under low-light scenarios. Extensive experiments demonstrate the superior performance of our approach over state-of-the-art unsupervised rPPG methods in different light conditions and verify the generalization of UDA-rPPG on cross-dataset testing. Our code and dataset are available at https://github.com/wwenmaositu/LS-rPPG-LRPD.
Xu Cheng 0003, Zhaodong Sun
IEEE Trans. Circuits Syst. Video Technol.2
2026 Time-Frequency Collaborative Learning for Imbalanced Ship Motion Data With Missing Labels in Sea State Estimation
abstract
Semi-supervised learning (SSL) has gained significant attention in the domain of sea state estimation (SSE) due to its capacity to alleviate the reliance of deep learning models on extensive labeled datasets. While existing semi-supervised SSE methodologies leveraging pseudo-labeling have achieved promising results, they often overlook the challenges posed by high class imbalance and the prevalence of missing data in ship motion datasets, which restricts their broader applicability. In this article, we propose a novel SSL approach BalanceSSE based on the class-imbalanced ship motion data for SSE. This approach consists of three main modules: 1) the dynamic imputation (DIT); 2) the imbalance temporal-frequency learning (ITFL); and 3) the ClusterProx classifier (CL). The DIT module dynamically imputes incomplete ship motion data by assigning different weights to various dimensions data. The ITFL module employs time-frequency collaborative learning to generate pseudo-labels and integrate an adaptive confidence strategy to select high confidence pseudo-labels. This process is further enhanced by the CL module to produce better estimates. Experimental tests on UCR datasets and ship motion datasets demonstrate that BalanceSSE outperforms state-of-the-art methods. Ablation studies highlight the critical role of each module in BalanceSSE.
Mengna Liu, Xu Cheng 0003, Junhao Xiao 0001, Shengyong Chen
IEEE Trans. Cybern.3
2026 Learning Domain-Generalizable Discriminative Representations by Mixing Euclidean Dynamics and Hilbert Statistics for Wind Turbine Blade Icing Detection
abstract
As wind energy grows in importance, blade icing threatens turbine efficiency and safety. Existing approaches based solely on convolutional neural networks (CNNs) for Euclidean-space feature extraction struggle with complex dynamics and domain shifts. To address this, we propose the domain-generalizable network for icing turbines via mixed Euclidean and Hilbert representations (DGMEHIT), which runs a CNN deep module and a Hilbert statistical module (HSM) in parallel to extract complementary local sequential and global statistical features from Euclidean-space and Hilbert space. Raw features are mapped into a high-dimensional Hilbert space to enhance global statistical separability and discriminate superficially similar signals. A channel–temporal mixer module further models dynamic dependencies by fusing multichannel and time-domain information. In addition, a maximum mean discrepancy loss is incorporated into the HSM to improve feature consistency across domains, enabling the learning of domain-generalizable representations and enhancing the model’s adaptability to varying geographical and climatic conditions. Experiments on ten public time-series datasets and one real-world icing dataset show DGMEHIT surpasses state-of-the-art methods, improving F1 by 6.4% and MCC by 12.7%, with generalization tests and online evaluations confirming practical robustness.
Mengna Liu, Yunke Li, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen
IEEE Trans. Ind. Informatics4
2026 A Temporal-Spectral Mixer to Class Imbalanced Ship Motion Data-Based Sea State Estimation for Maritime Intelligent Transportation Systems
abstract
Accurate sea state estimation (SSE) is critically important for enabling safe and efficient autonomous maritime transportation systems, yet traditional methods are costly, often exhibit latency, and are less suitable for integration within modern Intelligent Transportation Systems (ITS). While data-driven deep learning offers a promising alternative for real-time SSE within maritime ITS, the inherent class imbalance in naturally occurring sea states poses a significant challenge, hindering robust system development. Deep learning-based SSE methods often underutilize frequency-domain information, struggle with multi-scale wave characteristics, and are biased by class imbalance, limiting ITS effectiveness. To address these limitations and advance deep learning in maritime ITS, this paper proposes a novel class-imbalanced ship motion data-based Temporal-Spectral Mixer model for SSE. This model integrates a Spectral Frequency Adaptive (SFA) module to capture global spectral frequency information, a Temporal Multi-Scale Parallel Convolution (TMSPC) module for extracting local temporal multi-scale features, and an Imbalanced Contrastive Clustering Loss (ICC-Loss) function to mitigate class imbalance and enhance ITS applicability. The TMSPC module captures crucial temporal wave characteristics while the SFA module extracts global spectral context for system-aware SSE. Extensive evaluations demonstrate state-of-the-art performance on diverse datasets, including superior results on both public benchmarks and ship motion data compared to existing methods, including class-imbalance techniques. Ablation and sensitivity studies confirm the effectiveness of each module. The proposed Temporal-Spectral Mixer model offers a robust and promising SSE solution for class-imbalanced scenarios, advancing reliable maritime ITS and holding broader potential for time series classification within transportation systems.
Xu Cheng 0003, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.4
2026 TSCFNet: Temporal Spectral Feature Cross Fusion Network for Imbalanced Sea State Estimation in Autonomous Ships
abstract
Sea state estimation (SSE) is critical to the safety of maritime transport and the reliability of autonomous ships. The frequency of different sea states varies significantly, leading to uneven data distribution. Existing deep learning methods for SSE typically focus on feature extraction, often using simple splicing and fusion, which can result in cross-domain incoherence and degrade model performance. Addressing sea state classification imbalance is often done through distance-based classifiers (e.g., prototype classifiers), but these can be less sensitive to minority classes, and using few prototypes for a class limits the expression of intra-class variations. To overcome these challenges, we propose the Temporal Spectral Cross Fusion Network (TSCFNet), which extracts temporal and spectral features. These are integrated via an innovative temporal spectral cross fusion module to maximize their complementary advantages. Additionally, we introduce a multi-fusion loss function, including temporal, spectral, and fusion losses, to optimize features across different dimensions. This approach improves the performance for minority classes and captures intra-class differences more effectively, solving the problem of category imbalance. Experimental results show that TSCFNet significantly outperforms baseline methods on two imbalanced sea state datasets and multiple multivariate spatio-temporal datasets.
Feng Xiao 0005, Xu Cheng 0003, Xia Xie 0003, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.3
2026 Dynamic Dependency-Aware Collaborative Contrastive Learning for Multi-Behavior Recommendation
abstract
Multi-behavior recommender systems improve prediction accuracy of target behaviors (e.g., purchases) by integrating auxiliary behaviors (e.g., page views). However, existing models face two key limitations: (1) Static propagation mechanisms and inflexible dependency modeling fail to capture dynamic changes in user preferences and cascading relationships between behaviors; (2) Sparse target behavior data usually leads to excessive influence of auxiliary signals, which degrades recommendation quality. To address these challenges, we propose the Dynamic Dependency-Aware Collaborative Contrastive Learning Multi-Behavior Recommendation Model, MBDCC. MBDCC has two dedicated modules: (1) Behavioral gating cascade and cross-attention fusion module, which dynamically models cascading dependencies between behaviors through learnable gate control transfer units controlled by behavioral attributes. This replaces static propagation with adaptive feature flow regulation, capturing evolving user preferences. Meanwhile, it uses a target-guided cross-attention mechanism to selectively fuse semantically relevant auxiliary signals using the target behavior as a query, addressing inflexible cross-behavioral dependency modeling; (2) Collaborative semantic enhancement module, it constructs a user similarity measure matrix based on co-occurrence frequency of interaction items in target behavior, and clusters nodes using a hybrid clustering strategy. By introducing contrastive learning between nodes and their clustering centers, the collaborative semantic information between similar nodes under the target behavior is effectively captured and amplified, alleviating the challenge of sparse target interaction data. Extensive experiments on three real-world datasets show that MBDCC consistently outperforms state-of-the-art baselines, it achieves remarkable improvements of up to 6.84% in Recall@10 and 5.18% in NDCG@50. Moreover, ablation experiments further demonstrate the correctness of our motivation and the necessity of the various modules of the MBDCC model.
Xu Cheng 0003, Likang Wu, Qingbo Hao, Yingyuan Xiao, Wenguang Zheng
ACM Trans. Knowl. Discov. Data2
2026 A Geo-Aware Personalized Network for User and Service Representation and Bilinear Interaction Modeling in QoS Prediction
Xu Cheng 0003, Wenguang Zheng, Yingyuan Xiao
IEEE Trans. Netw. Serv. Manag.2
2026 Hierarchical Spatial-Angular Representation Learning for Point-Supervised Salient Object Detection in Light Fields
abstract
Light Field Salient Object Detection (LFSOD) aims to identify visually distinctive regions by leveraging the complementary spatial–angular information inherent in 4D light field imagery. A major challenge lies in modeling angular dependencies and maintaining spatial coherence under sparse supervision. In this article, we propose a weakly supervised network that consists of three interdependent modules. First, the Light Field Division (LFD) module utilizes epipolar geometry to extract direction-aware boundary features, enhancing the encoding of angular disparities. Second, the Light Field Spatial Association (LFSA) module anchors cross-view feature alignment using central-viewpoint annotations, thereby enforcing spatial consistency and mitigating redundant representations. Third, the Light Field Saliency Local Clustering (LFLC) module introduces a joint boundary-appearance modeling strategy that integrates adaptive clustering with error-aware regularization to refine structural predictions. Experiments on three benchmark datasets show that our method consistently outperforms mainstream weakly supervised approaches. It also achieves superior performance compared to several fully supervised methods.
Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen
ACM Trans. Multim. Comput. Commun. Appl.3
2026 CROMBO: Cross-Modality Bootstrapping for Unified Sketch-Photo Representation Learning
abstract
Sketch–photo recognition refers to matching hand-drawn sketches with their corresponding photos, where the performance essentially depends on how well the representations of the two modalities are aligned in the feature spaces. Existing works bluntly force models to reduce the representation discrepancy between the modalities, making the learning less effective. Besides, the current symmetric feature extraction framework prefers the photo modality for richer information while neglecting the sketch modality. Driven by these observations, we argue that, instead of forcefully wiping out the modality discrepancy, we may utilize the discrepancy to enhance model learning. Thus, we propose a Cross-Modality Bootstrapping learning framework (CROMBO) that utilizes the modality discrepancy to bootstrap cross-modality representation learning via a differentiated interaction manner. Specifically, we first present a Sketch Implicit Bootstrapping (SIB) module to magnify the recognizable elements in the photo modality by utilizing the characteristic of sketches having only contours and key details. Second, a Photo-driven Sketch Refinement (PSR) module is developed to guide the sketch representation in the shared feature extraction process by supplementing rich information from the photo modality. Moreover, we design a second-order alignment strategy to dynamically align the latent distribution of two modalities in a Hilbert space. Also, our CROMBO can learn fewer parameters by freezing the weights of shallow layers in the backbone while making no sacrifice in performance. Extensive experiments on six public datasets verify the superior performance of our CROMBO for sketch–photo-based tasks, such as sketch re-identification (Re-ID), sketch–photo face recognition, and sketch-based image retrieval.
Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 FreeNet: Liberating Depth-Wise Separable Operations for Building Faster Mobile Vision Architectures
abstract
In the pursuit of efficient vision architectures, substantial efforts have been devoted to optimizing operator efficiency. Depth-wise separable operators, such as DWConv, are found cheap in both FLOPs and parameters. As a result, they are increasingly incorporated into efficient backbones, trading for deeper and wider architectures to enhance performance. However, separable operators are not really fast on devices due to the discontinuous memory access requirements. In this paper, we propose FreeNets, a family of simple and efficient backbones that free the separable operation to further accelerate the running speed. We introduce sparse sampling mixers (S2-Mixer) to supersede existing separable token mixers. The S2-Mixer samples multiple segments of partially continuous signals across spatial and channel dimensions for convolutional processing, achieving extremely fast on-device speed. The sparse sampling also enables S2-Mixer to capture long-range pixel relationships from dynamic receptive fields. Furthermore, we introduce a Shift Feed-Forward Network (ShiftFFN) as a faster alternative to existing channel mixers. It utilizes a shift neck architecture that aggregates global information to shift features, enabling faster channel mixing while incorporating global pixel information. Extensive experiments demonstrate that FreeNet offers a superior accuracy-efficiency tradeoff compared to the latest efficient models. On ImageNet-1k, FreeNet-S2 outperforms the StarNet-S4 by 0.4% in top-1 accuracy, while running around 40% faster on desktop GPU and 15% faster on Mobile GPU.
Hao Yu 0015, Haoyu Chen 0001, Wei Peng 0009, Xu Cheng 0003, Guoying Zhao 0001
AAAI4
2025 Can Students Beyond the Teacher? Distilling Knowledge from Teacher's Bias
abstract
Knowledge distillation (KD) is a model compression technique that transfers knowledge from a large teacher model to a smaller student model to enhance its performance. Existing methods often assume that the student model is inherently inferior to the teacher model. However, we identify that the fundamental issue affecting student performance is the bias transferred by the teacher. Current KD frameworks transmit both right and wrong knowledge, introducing bias that misleads the student model. To address this issue, we propose a novel strategy to rectify bias and greatly improve the student model's performance. Our strategy involves three steps: First, we differentiate knowledge and design a bias elimination method to filter out biases, retaining only the right knowledge for the student model to learn. Next, we propose a bias rectification method to rectify the teacher model's wrong predictions, fundamentally addressing bias interference. The student model learns from both the right knowledge and the rectified biases, greatly improving its prediction accuracy. Additionally, we introduce a dynamic learning approach with a loss function that updates weights dynamically, allowing the student model to quickly learn right knowledge-based easy tasks initially and tackle hard tasks corresponding to biases later, greatly enhancing the student model's learning efficiency. To the best of our knowledge, this is the first strategy enabling the student model to surpass the teacher model. Experiments demonstrate that our strategy, as a plug-and-play module, is versatile across various mainstream KD frameworks.
Jianhua Zhang 0002, Ruyu Liu, Xu Cheng 0003, Houxiang Zhang, Shengyong Chen
AAAI4
2025 Robust Online Detection of Anomalies in Evolving Data Streams with GAN Imputation
abstract
Online anomaly detection is a critical technique in intelligent industrial systems. Concept drift and missing data that arise during the transmission of real-time data streams can severely impact the performance of anomaly detection. Therefore, ensuring the accuracy of online anomaly detection and adaptability to concept drift in the presence of missing values is a significant challenge. In this paper, we propose an online anomaly detection model that integrates an online imputation module based on Generative Adversarial Networks (GANs) to efficiently impute missing data. Additionally, following the dynamic model pool strategy, the model dynamically selects and adjusts the optimal models in the pool to effectively respond to changes in data distribution, ensuring detection performance under complex conditions. We evaluated the model's performance across multiple datasets with concept drift and under varying missing data rates. The results demonstrate that the proposed model not only adapts flexibly to rapidly changing data streams but also exhibits enhanced robustness in the presence of missing data.
Mengna Liu, Xu Cheng 0003, Jianhua Zhang 0002, Feng Xiao 0005
CSCWD3
2025 A Correlation-Aware Diffusion Model for Multivariate Time Series Anomaly Detection with Missing Values
abstract
Incomplete time series data is a common problem in real-world application scenarios. Recent research has taken the approach of separating interpolation and anomaly detection, which is not interactive and performs poorly. On the other hand, interpolation using traditional methods relies on a large amount of a priori knowledge, and using deep learning methods takes up a large amount of computational resources and is inefficient. In this study, we propose a correlation-aware diffusion model that successfully bypasses the above problems. Our approach focuses on capturing deep multivariate correlations from limited incomplete data and use low-frequency component to guide generation. Experiments on four realistic scenario datasets covering three domains show that our method achieves better anomaly detection results than existing methods for various missing rates.
Zhanneng Zeng, Renfang Wang, Hong Qiu, Xiufeng Liu 0001, Xu Cheng 0003
CSCWD5
2025 SPDM: Spatiotemporal-Periodic Diffusion Model for Multivariate Time Series Imputation
abstract
This paper presents SPDM, an innovative spatiotemporal-periodic diffusion model for multivariate time series interpolation, which addresses the challenges of spatiotemporal dependency and periodicity inherent in real-world datasets. Unlike prior methods, SPDM integrates conditional features capturing spatiotemporal correlations and geographic relationships, enhancing the model's ability to account for complex interdependencies within time series data. A noise prediction module, leveraging Fast Fourier Transform, decomposes time series into periodic components, thereby enabling the model to capture both intra- and inter-period dynamics and inter-channel correlations effectively. Experimental results on multiple industrial datasets show that SPDM outperforms state-of-the-art methods across various missing data scenarios, highlighting its robustness and effectiveness. This work establishes a new approach to time series interpolation by combining conditional information construction with periodicity-aware diffusion modeling, offering promising insights for further applications in time series analysis.
Qia Zhang, Renfang Wang, Hong Qiu, Xiufeng Liu 0001, Xu Cheng 0003
CSCWD5
2025 Data-Driven Diffusion-Augmented Network for Imbalanced Sea State Estimation with Dynamic Prototypes
abstract
With the advent of Industry 4.0, which emphasizes automation and data-driven decision-making, sea state estimation (SSE) has become a critical component in marine engineering and autonomous vessels. However, traditional SSE methods face significant challenges due to poor real-time performance and high manual costs. Additionally, existing deep learning (DL) approaches struggle to handle imbalanced sea state data effectively, which hinders their generalization ability. To address these issues, this paper proposes a novel DL-based SSE model. The model enhances the expressive capacity of the extracted features through a feature augmentation module and utilizes a diffusion model to extract more robust features from the data. Furthermore, a dynamic prototype update module is designed to effectively address class imbalance and overcome the limitations of traditional prototype classifiers in handling boundary samples. Experimental results show that the proposed method outperforms existing baseline models on two imbalanced sea state datasets, achieving F1-score improvements of 4.5% and 3.5%, respectively, compared to the current state-of-the-art methods. Additionally, the model demonstrates superior performance on several publicly available multivariate time series classification datasets.
Feng Xiao 0005, Xu Cheng 0003, Jianhua Zhang 0002
CSCWD3
2025 Lightweight Self-Supervised Monocular Depth Estimation via Context-Aware Fusion and Separable Depthwise Convolution
abstract
Monocular depth estimation is a critical problem in computer vision, with wide-ranging applications across various domains. However, existing methods often involve high computational costs, making them challenging to deploy efficiently on edge devices. To address the trade-off between computational complexity and inference accuracy, this paper presents an efficient and lightweight model for self-supervised monocular depth estimation. Our model incorporates a Context-Aware Fusion (CAF) module to capture both global and local feature dependencies. In addition, Separable Depthwise Convolution (SDC) module are utilized to reduce computational overhead, and the Multi-Scale Structural Similarity (MS-SSIM) loss function is employed to improve both depth estimation accuracy and visual perception quality. Experimental results show that the proposed model delivers improved accuracy while maintaining a lightweight and efficient architecture, ensuring its compatibility with edge device deployment. A comprehensive analysis of the findings is also presented, along with insights for future optimizations and improvements.
Meina Zhao, Shixin Wang 0014, Feng Xiao 0005, Jianhua Zhang 0002, Xu Cheng 0003, Yunrui Zhu
CSCWD5
2025 QCTKD-PU: Quantum Convolutional Transformer with Knowledge Distillation for Efficient and Robust Point Cloud Upsampling
abstract
Point cloud upsampling is crucial for high-fidelity 3D reconstruction in real-time applications such as autonomous systems. Existing methods based on CNNs or Transformers face three limitations: (1) prohibitive computational complexity hindering real-time deployment, (2) insufficient modeling of multi-scale geometric dependencies in sparse data, (3) sensitivity to noise and outliers. To address these challenges, we propose QCTKD-PU, a framework integrating Quantum Convolutional Transformers (QCT) and Knowledge Distillation (KD) for Point cloud Upsampling. The QCT leverages quantum superposition and self-attention to encode high-dimensional features, enabling efficient multi-scale point interaction learning. Simultaneously, KD transfers knowledge from a teacher model to a lightweight student network, reducing computational costs while maintaining accuracy. Experiments on benchmark datasets demonstrate superior performance in geometric accuracy and noise robustness compared to state-of-the-art methods. This work pioneers the synergy of quantum computing and lightweight learning for resource-constrained 3D vision tasks, while the student model achieves real-time and compact deployment, offering a practical solution for collaborative edge systems.
Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen
CSCWD4
2025 From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identification
abstract
Aiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distributed across multiple devices/entities, raising privacy and ownership concerns that make existing centralized training impractical for VI-ReID. To tackle these challenges, we propose L2RW, a benchmark that brings VI-ReID closer to real-world applications. The rationale of L2RW is that integrating decentralized training into VI-ReID can address privacy concerns in scenarios with limited data-sharing regulation. Specifically, we design protocols and corresponding algorithms for different privacy sensitivity levels. In our new benchmark, we ensure the model training is done in the conditions that: 1) data from each camera remains completely isolated, or 2) different data entities (e.g., data controllers of a certain region) can selectively share the data. In this way, we simulate scenarios with strict privacy constraints which is closer to real-world conditions. Intensive experiments with various server-side federated algorithms are conducted, showing the feasibility of decentralized VI-ReID training. Notably, when evaluated in unseen domains (i.e., new data entities), our L2RW, trained with isolated data (privacy-preserved), achieves performance comparable to SOTAs trained with shared data (privacy-unrestricted). We hope this work offers a novel research entry for deploying VI-ReID that fits real-world scenarios and can benefit the community.
Hao Yu 0015, Xu Cheng 0003, Haoyu Chen 0001, Zhaodong Sun, Guoying Zhao 0001
CVPR3
2025 SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structures
abstract
Pixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morphology and texture, facing challenges in balancing segmentation quality with low computational resource usage. To overcome these limitations, we propose a lightweight Structure-Aware Vision Mamba Network (SCSegamba), capable of generating high-quality pixel-level segmentation maps by leveraging both the morphological information and texture cues of crack pixels with minimal computational cost. Specifically, we developed a StructureAware Visual State Space module (SAVSS), which incorporates a lightweight Gated Bottleneck Convolution (GBC) and a Structure-Aware Scanning Strategy (SASS). The key insight of GBC lies in its effectiveness in modeling the morphological information of cracks, while the SASS enhances the perception of crack topology and texture by strengthening the continuity of semantic information between crack pixels. Experiments on crack benchmark datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods, achieving the highest performance with only 2.8M parameters. On the multi-scenario dataset, our method reached 0.8390 in F1 score and 0.8479 in mIoU. The code is available at https://github.com/Karl1109/SCSegamba.
Fan Shi 0001, Xu Cheng 0003, Shengyong Chen
CVPR4
2025 ICFF-Net: Interlaced Cross-Attention Feature Fusion Network for Music Genre Classification
Shiting Meng, Cairui Yan, Yingyuan Xiao, Wenguang Zheng, Xu Cheng 0003
DASFAA (3)5
2025 Harnessing Light Field Angular Cues and Spatial Geometries for Semantic Segmentation
abstract
4D light field imaging captures rich spatial-angular information, providing essential geometric cues for semantic segmentation tasks. In this paper, we introduce a novel backbone network called the Light Field Extraction Interaction Network (LFEI-Net). LFEI-Net excels in extracting global structures and multi-scale spatial-angular features, capturing feature dependencies through channel modeling and diverse feature interactions. Unlike traditional methods that depend on pyramid and dilated feature extraction, LFEI-Net pioneers an efficient method by integrating large-scale horizontal depth-wise convolution (HDWC) and vertical depth-wise convolution (VDWC) with interactive operations for comprehensive spatial multi-scale feature extraction. Furthermore, we present the Multi-Angular Modeling (MAM) module, which effectively captures scene angle variations from multiple perspectives and precisely delineates object boundaries, thereby improving model adaptability. Our experimental evaluations on two datasets demonstrate that LFEI-Net significantly outperforms state-ofthe-art (SOTA) 2D and 4D light field semantic segmentation methods, achieving mean Intersection over Union (mIoU) of 83.72% and 86.88%, respectively.
Fan Shi 0001, Xu Cheng 0003
ICASSP3
2025 Serial Local Patterns and Irregular Dependencies Extract and Cascaded Fusion Network for Structural Crack Segmentation
abstract
Achieving pixel-level crack segmentation in complex scenarios is a major challenge, as current methods have difficulty effectively integrating both local features and irregular pixel dependencies. In this paper, we introduce a Cascaded Fusion Network (LICFN) specifically designed for crack segmentation, which extracts fine local details and pixel dependencies using a hybrid feature extractor and effectively enhances and fuses them through a cascaded fusion module. To comprehensively evaluate the network, we also created a benchmark dataset, TUT, which includes various scenarios. Experimental results show that our method surpasses others, achieving F1 and mIoU scores of 0.8439 and 0.8509, respectively. The dataset is available at https://github.com/Karl1109/TUT.
Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001
ICASSP3
2025 Collaborative Association Network for Multi-view Multi-Human Association and Tracking using Constraint Optimization and Object Search
abstract
Multi-view multi-human association and tracking (MvMHAT) enhances scene perception using multiple cameras, crucial for applications such as surveillance and crowd analysis. Inherent feature disparities between views complicate similarity calculations. Recent works combine representation and motion information to address this issue. However, existing methods neglect parallax-induced angular issues and inconsistent object counts across views. To address these challenges, we introduce a collaborative association network combining temporal and spatial clues. Our method incorporates multi-scale adaptive alignment, cross-view and cross-frame feature fusion, to obtain comprehensive global feature representations for each object. We also formulate data association as a mixed-constraint optimization problem to enhance the scalability of our method. Additionally, we propose a novel object search loss to improve cross-view and cross-frame data association. Experiments on benchmarks demonstrate the efficiency of our method in MvMHAT task, significantly outperforming state-of-the-art methods.
Fan Shi 0001, Meng Zhao 0001, Xu Cheng 0003
ICASSP5
2025 Efficient Large-Scale Scene Point Cloud Upsampling with Implicit Neural Networks and Spatial Hashing
abstract
Point cloud upsampling is a critical challenge in 3D vision, particularly for large-scale, real-world data. We propose ASFNet, a novel implicit neural network-based approach that uniquely combines adaptive spatial feature representation with efficient spatial hashing. This method significantly improves both upsampling quality and computational efficiency. ASFNet first encodes the point cloud as an implicit surface, employing dynamic search and spatial hashing to optimize query point locations rapidly. This approach creates a uniform, continuous field around surfaces, enabling high-fidelity upsampling. Experiments on benchmark datasets, including Oakland 3D dataset and VMR-Oakland-v2, demonstrate ASFNet’s superiority. Our method achieves state-of-the-art performance with a Chamfer Distance of 5.559 × 10–3on Oakland 3D dataset, while reducing processing time by up to 80% compared to existing methods. On the challenging Oakland 3D dataset, ASFNet completes upsampling in just 150 seconds. These results underscore ASFNet’s potential to advance real-time 3D vision applications in areas such as autonomous navigation and augmented reality.
Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Xiufeng Liu 0001
ICASSP4
2025 FusionPhys: A Flexible Framework for Fusing Complementary Sensing Modalities in Remote Physiological Measurement
Chenhang Ying, Huiyu Yang, Jieyi Ge, Zhaodong Sun, Xu Cheng 0003, Kui Ren 0001
ICCV5
2025 DIFCN: A Hybrid Network for Capturing Dynamic Interests and Feature Co-action in CTR Prediction
Qingbo Hao, Xu Cheng 0003, Yingyuan Xiao
ICIC (8)3
2025 Mamba-SLAM: Enhancing Neural Implicit SLAM with Uncertainty and Mamba
abstract
Current neural implicit SLAM systems struggle with insufficient object shape constraints due to incomplete depth maps and inefficient pixel sampling strategies, leading to inaccuracies in reconstructed scene morphology and texture. To address these limitations, we introduce Mamba-SLAM, a novel framework featuring two key innovations: (1) an uncertainty-based pixel sampling module that enhances object rendering quality in depth-scarce regions by integrating depth map analysis and color image gradients, improving shape and texture consistency by 6.4% on the Replica dataset compared to state-of-the-art methods; and (2) a Mamba-based keyframe selection module that leverages the efficient feature extraction of Mamba to provide rich semantic cues, optimizing pose estimation and enriching reconstruction detail. Experiments on Replica and ScanNet demonstrate that Mamba-SLAM significantly improves scene rendering and object detail, achieving a 13.5% improvement in reconstruction completeness on Replica. The core novelty lies in the synergistic combination of uncertainty-driven pixel selection and Mamba-powered keyframe management for enhanced neural implicit SLAM.
Jiaming Lu, Yunrui Zhu, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Xiufeng Liu 0001
ICME4
2025 Multi-Scale Convolutional Networks with Class-Normalized Logit Clipping for Robust Sea State Estimation from Noisy Ship Motion Data
abstract
Autonomous ships utilize automation systems to achieve unmanned navigation, driving innovation in maritime transportation. However, sea conditions, influenced by dynamic factors such as wave height, wind speed, and ocean currents, present a challenge in accurately assessing these conditions. Traditional classification models often assume accurate labels, but noisy labels are prevalent in real-world applications. Existing methods, such as noise sample filtering or loss function adjustment, have limited applicability and poor generalization when dealing with complex sea condition data. To address this issue, this study proposes an end-to-end neural network model. The model's feature extraction module uses deep representation learning to capture latent patterns in the data, and a loss function is designed to mitigate the impact of outliers. The integration of these components allows the model to perform accurate classification even in the presence of noisy labels. Extensive experiments on public and sea condition datasets validate the effectiveness of this approach, demonstrating that the model exhibits strong generalization capabilities and holds great promise for practical applications.
Mengna Liu, Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001, Jianhua Zhang 0002, Shengyong Chen
ICRA3
2025 FedFAS: Federated Few-shot Abdominal Organs Segmentation across Heterogeneous Clients
abstract
Federated Learning (FL) provides a solution for learning a global model without transferring data, which helps protect privacy in clinical applications. However, existing FL methods often assume that clients have sufficient training samples to generalize the model, and therefore perform poorly in the case of small samples. Furthermore, existing works pay little attention to more challenging medical image segmentation tasks, especially in the case of class-heterogeneous FL. Therefore, in this paper, we construct a framework for federated few-shot medical image segmentation. Specifically, each client obtains local prototypes with limited training samples, which are then uploaded to the server to form a global class prototype library. The clients then select and utilize global class prototypes to calculate global-to-local prototype comparisons to correct local training. In addition, we propose a personalized aggregation strategy for local tasks to enhance the client’s generalization capability for unseen classes and enable the client to learn a discriminative feature space. We establish FL settings using two widely-used datasets and conduct experiments to demonstrate the effectiveness and superiority of our approach.
Yi Zhang 0111, Junpeng Wu, Meng Zhao 0001, Xu Cheng 0003, Yao Zhang 0021, Fan Shi 0001
IJCNN4
2025 Efficiency-Optimized Point Cloud Upsampling with Single-Layer Graph Convolution Network
abstract
Point cloud upsampling is a key technology for improving the density and quality of sparse point clouds, with widespread applications in 3D reconstruction, autonomous driving, and environmental perception. However, traditional point cloud upsampling methods, especially those based on multilayer graph convolution networks (GCNs), typically rely on complex feature extraction modules, which increase computational complexity and model parameters, limiting their use in resource-constrained environments. To overcome these challenges, we propose the EO-PU framework, a lightweight and efficient point cloud upsampling method. This framework combines single-layer GCN and rotation-invariant 4D projection encoding (I4DP) technology, significantly reducing computational load and redundant information, thereby improving upsampling efficiency. Specifically, EO-PU first uses I4DP to map the 3D point cloud data to a rotation-robust 4D feature space, ensuring effective capture of geometric information. A single-layer GCN is then employed to aggregate features, reducing network complexity and computational cost. To further enhance upsampling performance, we introduce the EdgeShuffleNet module, which optimizes feature expansion and rearrangement through efficient local feature aggregation. Experimental results show that EO-PU outperforms or matches existing methods across multiple public datasets while significantly reducing model parameters and computation time, making it highly suitable for deployment in resource-constrained environments.
Yunrui Zhu, Feng Xiao 0005, HaoXiao Wang, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002
IJCNN6
2025 Dual-Path Contrastive Learning For Wind Turbine Icing Detection
abstract
Wind energy, characterized by its clean and replenishable nature, is increasingly used worldwide due to its environmental friendliness and wide distribution of resources. However, ice accretion on turbine blades in cold regions, often resulting from cold weather conditions, significantly impacts both the operational performance and security of wind energy production, which significantly increases maintenance costs, resulting in a significant reduction in the energy output performance of wind turbines. Specifically, blade icing alters the aerodynamic characteristics of the blade surface, increases wind resistance, and reduces wind energy conversion efficiency. Furthermore, ice accretion may result in a non-uniform mass allocation across the blade surfaces. This imbalance can induce vibrations within the turbine system, thereby compromising its operational stability and structural integrity. The key challenges are complex sensor parameter variations, high labeling costs, and data imbalance, making accurate icing prediction difficult. In order to tackle these difficulties, this study introduces a technique based on dual-path contrastive learning. This method balances the dataset using a sliding window technique and utilizes both icing loss features and expert features for dual-path processing to fully exploit the feature sets. Furthermore, ice accretion may result in a non-uniform mass allocation across the blade surfaces. This imbalance can induce vibrations within the turbine system, thereby compromising its operational stability and structural integrity, particularly exhibiting excellent performance in handling data imbalance.
Aili Xu, Jiamei Zhou, Xu Cheng 0003, Fan Shi 0001, Yongming Han, Guoqian Jiang
INDIN3
2025 To Remember, To Adapt, To Preempt: A Stable Continual Test-Time Adaptation Framework for Remote Physiological Measurement in Dynamic Domain Shifts
abstract
Remote photoplethysmography (rPPG) aims to extract non-contact physiological signals from facial videos and has shown great potential. However, existing rPPG approaches struggle to bridge the gap between source and target domains. Recent test-time adaptation (TTA) solutions typically optimize rPPG model for the incoming test videos using self-training loss under an unrealistic assumption that the target domain remains stationary. However, time-varying factors like weather and lighting in dynamic environments often cause continual domain shifts. The erroneous gradients accumulation from these shifts may corrupt the model's key parameters for physiological information, leading to catastrophic forgetting. Therefore, We propose a physiology-related parameters freezing strategy to retain such knowledge. It isolates physiology-related and domain-related parameters by assessing the model's uncertainty to current domain and freezes the physiology-related parameters during adaptation to prevent catastrophic forgetting. Moreover, the dynamic domain shifts with various non-physiological characteristics may lead to conflicting optimization objectives during TTA, which is manifested as the over-adapted model losing its adaptability to future domains. To fix over-adaptation, we propose a preemptive gradient modification strategy. It preemptively adapts to future domains and uses the acquired gradients to modify current adaptation, thereby preserving the model's adaptability. In summary, we propose a stable continual test-time adaptation (CTTA) framework for rPPG measurement, called PhysRAP, which Remembers the past, Adapts to the present, and Preempts the future. Extensive experiments show its state-of-the-art performance, especially in domain shifts. The code is available at https://github.com/xjtucsy/PhysRAP.
Shuyang Chu, Jingang Shi, Xu Cheng 0003, Haoyu Chen 0001, Xin Liu 0012, Guoying Zhao 0001
ACM Multimedia3
2025 LFMamba: Focal Stack-aware State Space Modeling for Light Field Salient Object Detection
abstract
Salient object detection (SOD) in light field data presents unique challenges due to dynamic semantic inconsistencies across focal slices and representation heterogeneity between focal slices and the all-focus image. Existing methods often treat focal slices uniformly or rely on simple fusion strategies, which fail to address focus-induced semantic drift and cross-modal feature misalignment. To tackle these issues, we propose LFMamba, a unified network that jointly models dynamic semantic consistency and adaptive cross-modal fusion. We design the Focal-aware State Space Module (FSSM), which generates focal-aware semantic prompts through low-rank decomposition and adaptively routes them according to focal plane indices, thereby enabling bidirectional semantic propagation across slices through non-causal state transitions. Furthermore, we introduce the Focal-guided Cross-modal Fusion Module (FCFM), which mitigates cross-modal heterogeneity by a two-stage hierarchical strategy, combining structure-aware low-level alignment and gated high-level semantic fusion. Extensive experiments on four public light field SOD benchmarks demonstrate that LFMamba achieves superior performance compared to state-of-the-art methods, with improved robustness and consistency under complex focal variation scenarios.
Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Shengyong Chen
ACM Multimedia3
2025 LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks
abstract
Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interactive fusion of cross-modal features. To address these challenges, we propose a Lightweight Adaptive Cue-Aware Vision Mamba network (LIDAR), which efficiently perceives and integrates morphological and textural cues from different modalities under multimodal crack scenarios, generating clear pixel-level crack segmentation maps. Specifically, LIDAR is composed of a Lightweight Adaptive Cue-Aware Visual State Space module (LacaVSS) and a Lightweight Dual Domain Dynamic Collaborative Fusion module (LD3CF). LacaVSS adaptively models crack cues through the proposed mask-guided Efficient Dynamic Guided Scanning Strategy (EDG-SS), while LD3CF leverages an Adaptive Frequency Domain Perceptron (AFDP) and a dual-pooling fusion strategy to effectively capture spatial and frequency-domain cues across modalities. Moreover, we design a Lightweight Dynamically Modulated Multi-Kernel convolution (LDMK) to perceive complex morphological structures with minimal computational overhead, replacing most convolutional operations in LIDAR. Experiments on three datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods. On the light-field depth dataset, our method achieves 0.8204 in F1 and 0.8465 in mIoU with only 5.35M parameters. Code and datasets are available at https://github.com/Karl1109/LIDAR-Mamba.
Fan Shi 0001, Xu Cheng 0003, Mengfei Shi, Xia Xie 0003, Shengyong Chen
ACM Multimedia4
2025 PolypSense3D: A Multi-Source Benchmark Dataset for Depth-Aware Polyp Size Measurement in Endoscopy
abstract
Accurate polyp sizing during endoscopy is crucial for cancer risk assessment but is hindered by subjective methods and inadequate datasets lacking integrated 2D appearance, 3D structure, and real-world size information. We introduce PolypSense3D, the first multi-source benchmark dataset specifically targeting depth-aware polyp size measurement. It uniquely integrates over 43,000 frames from virtual simulations, physical phantoms, and clinical sequences, providing synchronized RGB, dense/sparse depth, segmentation masks, camera parameters, and millimeter-scale size labels derived via a novel forceps-assisted in-vivo annotation technique. To establish its value, we benchmark state-of-the-art segmentation and depth estimation models. Results quantify significant domain gaps between simulated/phantom and clinical data and reveal substantial error propagation from perception stages to final size estimation, with the best fully automated pipelines achieving an average Mean Absolute Error (MAE) of 0.95 mm on the clinical data subset. Publicly released under CC BY-SA 4.0 with code and evaluation protocols, PolypSense3D offers a standardized platform to accelerate research in robust, clinically relevant quantitative endoscopic vision. The benchmark dataset and code are available at: https://github.com/HNUicda/PolypSense3D and https://doi.org/10.7910/DVN/K13H89.
Ruyu Liu, Mingming Zhou, Jianhua Zhang 0002, Xiufeng Liu 0001, Xu Cheng 0003, Sixian Chan 0001, Yanbin Shen, Sheng Dai, Yuping Yan, Yaochu Jin, Lingjuan Lyu
NeurIPS7
2025 Multi-source Domain Adaptation Image Steganalysis for Cover Source Mismatch
Xiang Zhang 0023, Xinjue Hu, Fan Wang 0024, Xu Cheng 0003, Zhangjie Fu 0001
PRCV (6)5
2025 Enhancing spatiotemporal wind power forecasting with meta-learning in data-scarce environments
abstract
Accurate wind power forecasting is critical for maintaining stable power grids, yet the inherent variability of wind and limited data availability for new wind farms present significant challenges. To address these issues, we present a novel artificial intelligence framework that integrates a self-attention enhanced Spatiotemporal Long Short-Term Memory (ST-LSTM) network with Model-Agnostic Meta-Learning (MAML), termed as the Meta-Learning Spatiotemporal Attention Long Short-Term Memory framework (MAML-STALSTM). This deep learning combination enables the model to effectively capture long-range spatiotemporal dependencies while rapidly adapting to new wind farm configurations or changing wind conditions with minimal training data. By employing rigorous data preprocessing techniques and ensuring temporal separation in data splitting, we mitigate potential data leakage and enhance the model’s generalizability. Extensive experiments conducted on both onshore and offshore wind farm datasets demonstrate that our artificial intelligence approach outperforms established baseline models, particularly excelling in data-scarce environments. Ablation studies highlight the crucial roles of the self-attention mechanism and meta-learning in improving forecasting accuracy, adaptation speed, and model robustness. These results emphasize the practical benefits of our approach in enhancing grid stability and supporting the seamless integration of wind energy, thereby contributing significantly to the advancement of sustainable energy solutions.
Renfang Wang, Jingtong Wu, Xu Cheng 0003, Xiufeng Liu 0001, Hong Qiu
Eng. Appl. Artif. Intell.3
2025 An expert features enhanced temporal and contextual contrasting learning model for detecting wind turbine blade icing
abstract
With global carbon neutrality goals, wind power has rapidly developed, but blade icing remains a major challenge. AI(artificial intelligence) methods show great promise for detecting icing on wind turbine blades. However, early icing data overlap, difficulty obtaining continuous labeled data, and small variations between samples due to short sampling intervals complicate the task. This study proposes an expert feature-enhanced temporal and contextual contrastive learning model for detecting blade icing. This approach efficiently extracts data features and combines self-supervised contrastive learning, maximizing data utilization without requiring extensive labeled data. To validate the effectiveness of this method, extensive experiments were conducted on two public datasets. The results achieved the best performance across multiple metrics, with F1-Score and AUC exceeding 98%, significantly enhancing wind power generation efficiency.
Jiamei Zhou, Feng Xiao 0005, Xu Cheng 0003, Jianhua Zhang 0002
Eng. Appl. Artif. Intell.4
2025 An end-to-end model for time series classification in the presence of missing values
Mengna Liu, Pengshuai Yao, Xu Cheng 0003, Shengyong Chen
Expert Syst. Appl.3
2025 A contrastive clustering loss function increases class-balanced in time series classification
Chaomin Wu, Xu Cheng 0003, Hao Wang 0003
Expert Syst. Appl.2
2025 Hierarchical Spatiotemporal Graph Network for Fault Diagnosis of Industrial Processes
abstract
Intelligent fault diagnosis of industrial processes has received enormous attention in recent years, and deep learning-based methods have excellent performance in accurate health state detection. However, existing methods cannot fully exploit the complex relationship between different subsystems of industrial processes. To address this limit, we convert industrial data into graph structure data with sensors (as nodes) and topological connections between sensors (as edges) to represent complex interactive information. Specifically, we propose a spatiotemporal graph convolutional network with a hierarchical structure (HiSTGCN) for fault diagnosis of industrial processes. First, a local-global graph framework is constructed to explore the correlation between subsystems fully. Particularly, we propose a hierarchical graph structure with a global graph representing the correlation of sensors between subsystems and several local graphs capturing the correlation of sensors within subsystems to enrich the feature extraction. Then, we design a hierarchical spatiotemporal graph neural network to perform a local-global graph framework in both temporal and spatial dimensions. Finally, a synthesized residual health monitor module based on the principal component analysis (PCA) is designed for fault detection and location. Experiment results on an industrial simulation process dataset and a real wind farm dataset show that HiSTGCN has reliable and superior fault diagnosis performance compared to existing methods.
Guoqian Jiang, Kaili Shen, Xiufeng Liu 0001, Xu Cheng 0003
IEEE Internet Things J.4
2025 One Stone, Three Birds: Prototype-Enhanced Federated Learning for Mitigating Data Scarcity, Imbalance, and Heterogeneity in Blade Icing Detection Across Distributed Wind Farms
abstract
Wind turbine blade icing poses a critical challenge to wind power generation in high-latitude regions, necessitating innovative solutions for reliable icing detection. To address this challenge while leveraging the abundance of unlabeled data and preserving data privacy, this study proposes a novel federated semi-supervised prototype learning framework, FedIce. By integrating prototype learning and federated learning, FedIce extracts representative class prototypes at the client level and performs global model updates through federated averaging, significantly enhancing robustness against data heterogeneity. Additionally, it incorporates an advanced separation margin strategy to effectively alleviate the adverse effects of class imbalance. Comprehensive experiments using real-world datasets from 20 wind turbines across two wind farms demonstrate that FedIce outperforms existing methods, achieving a remarkable 95.58% improvement in the$mF_{\beta }$metric and a 33.25% enhancement in the mBA metric compared to FedMatch.
Lele Qi, Mengna Liu, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen
IEEE Internet Things J.3
2025 Alice-SLAM: Accurate and Lite-Communication Collaborative SLAM for Resource-Constrained Multi-Agent
abstract
Multi-agent collaborative simultaneous localization and mapping (Mac-SLAM) facilitates mutual localization among multi-agent and mapping in unknown environments. However, Mac-SLAM faces two main practical challenges in resource-constrained situations: heavy communication load and conflicts among multi-source maps. To address these issues, we propose Alice-SLAM: an accurate and lite-communication client-server collaborative SLAM system, reducing communication load while accuracy-guaranteed. Specifically, regarding high communication demand, we optimize communication load by compressing keyframe data and sharing only key map information instead of full map information. For inconsistency among multi-maps, we combine specific bundle adjustments (BA) and an adaptive strategy for active map optimization to enhance the consistency of the global map. A set of experiments demonstrates the superior accuracy and reduced communication load of the proposed Alice-SLAM on the EuRoC dataset and in multi-user augmented reality (AR) experiments conducted in our lab, highlighting its effectiveness in resource-constrained cases. We plan to open-source our code1to encourage further research and collaboration in this area.
Kaiqi Chen 0001, Ruyu Liu, Xu Cheng 0003, Jianhua Zhang 0002, Shengyong Chen, Houxiang Zhang, Arash Ajoudani
IEEE J. Sel. Areas Commun.4
2025 SpatialIE: Towards adaptive floating waste detection in unpredictable weather
abstract
Accurate detection and subsequent cleanup of floating waste are critical for ecosystem protection. However, existing methods face significant challenges in dealing with unpredictable weather, limiting their generalization capabilities. This paper proposes a novel plug-and-play architecture called Image Enhancer with Spatial Search and Aggregation (SpatialIE) . This architecture includes a novel encoder block (KANsformer) based on KANs and a dynamic decoder module leveraging prompt-based learning. The KANsformer in the encoder enhances the model’s ability to learn complex non-linear relationships in degraded images . At the same time, the Multi-Order-Based Prompt Block in the decoder enables dynamic optimization of feature processing strategies. These components allow SpatialIE to refine detection strategies to accommodate unknown weather conditions adaptively. We integrate SpatialIE with common YOLO detectors in an end-to-end training framework using the FloW-img and Water Surface Object Detection Dataset (WSODD) datasets. Experimental results show that our method delivers outstanding performance across multiple and single degradation scenarios, achieving an optimal balance between detection accuracy and storage efficiency.
Xiufeng Liu 0001, Xu Cheng 0003, Yusen Liu 0001, Tianqing Zhu, Huan Huo
Knowl. Based Syst.3
2025 Multi-scene low-light remote physiological measurement database
Zuxian He, Xu Cheng 0003
Mach. Vis. Appl.5
2025 Adaptive expert fusion model for online wind power prediction
abstract
Wind power prediction is a challenging task due to the high variability and uncertainty of wind generation and weather conditions. Accurate and timely wind power prediction is essential for optimal power system operation and planning. In this paper, we propose a novel Adaptive Expert Fusion Model (EFM+) for online wind power prediction. EFM+ is an innovative ensemble model that integrates the strengths of XGBoost and self-attention LSTM models using dynamic weights. EFM+ can adapt to real-time changes in wind conditions and data distribution by updating the weights based on the performance and error of the models on recent similar samples. EFM+ enables Bayesian inference and real-time uncertainty updates with new data. We conduct extensive experiments on a real-world wind farm dataset to evaluate EFM+. The results show that EFM+ outperforms existing models in prediction accuracy and error, and demonstrates high robustness and stability across various scenarios. We also conduct sensitivity and ablation analyses to assess the effects of different components and parameters on EFM+. EFM+ is a promising technique for online wind power prediction that can handle nonstationarity and uncertainty in wind power generation.
Renfang Wang, Jingtong Wu, Xu Cheng 0003, Xiufeng Liu 0001, Hong Qiu
Neural Networks3
2025 DMANet: Dual-modality alignment network for visible-infrared person re-identification
abstract
Visible–infrared person re-identification (VI-ReID) is a challenging retrieval task, which aims to match the same pedestrian between visible and infrared modalities. Most existing works achieve performance gains by solving the problem of the inherent cross-modality discrepancies. However, they cannot fully mine the modality information and lead to a poor generalization. In addition, the pedestrian images are unable to align well due to the large inter- and intra- class variations. To tackle the above limitations, we propose a novel dual-modality alignment network (DMANet) for VI-ReID. The core idea of our work is to develop multi-granularity features mutual learning (MGFML) for inadequate perception of modalities information, and to solve modality difference by proposing inter- and intra- modality alignment module (IIMA). Specifically, firstly, an effective multi-granularity features mutual learning module is proposed to mine the multi-granularity features, which combines the domain alignment and self-distillation to relieve modality discrepancy. Further, the maximum mean discrepancy loss and mutual learning loss are presented to enhance the identity-aware ability of the DMANet. Secondly, an effective inter- and intra- modality alignment module is presented to explore the potential alignment relation of inter- and intra- modalities. Finally, joint learning mechanism of multi-granularity features and modality alignment is utilized to improve the VI-ReID accuracy. Extensive experiments on mainstream benchmarks demonstrate that our method is superior to the state-of-the-art methods.
Xu Cheng 0003, Shuya Deng, Hao Yu 0015, Guoying Zhao 0001
Pattern Recognit.1
2025 SVD-KD: SVD-based hidden layer feature extraction for Knowledge distillation
Jianhua Zhang 0002, Mian Zhou, Ruyu Liu, Xu Cheng 0003, Sasa Nikolic 0002, Shengyong Chen
Pattern Recognit.5
2025 Recognizing Video Activities in the Wild via View-to-Scene Joint Learning
abstract
Recognizing video actions in the wild is challenging for visual control systems. In-the-wild videos show actions not seen in training data, recorded from various angles and scenes with the same labels. Most existing methods address this challenge by developing complex frameworks to extract spatiotemporal features. To achieve view robustness and scene generalization cost-effectively, we explore view consistency and scene joint understanding. Based on this, we propose a neural network (called Wild-VAR) to learn view and scene information jointly without any 3D pose ground truth labels, a new approach to recognizing video actions in the wild. Unlike most existing methods, first, we propose a Cubing module to self-learn body consistency between views instead of comprehensive image features, boosting the generalization performance of across-view settings. Specifically, we map 3D representations to multiple 2D features and then adopt a self-adaptive scheme to constrain 2D features from different perspectives. Moreover, we propose temporal neural networks (called T-Scene) to develop a recognizing framework, enabling Wild-VAR to flexibly learn scenes across time, including key interactors and context, in video sequences. Extensive experiments show that Wild-VAR consistently outperforms state-of-the-art methods on four benchmarks. Notably, with only half the computation costs, Wild-VAR improves accuracy by 2.2% and 1.3% on the Kinetics-400 and the Something-Somthing V2 datasets, respectively. Note to Practitioners—In human-robot interaction tasks, video action recognition technology is a prerequisite for visual control. In real applications, humans move freely in 3D space, which results in significant changes in the view of video capture and constantly changing scenes. Deep Neural Networks are limited by the perspectives and scenarios contained in the training data, resulting in most existing methods are only effective for identifying actions from 2–4 fixed views, and the background is single. Therefore, existing models are often difficult to generalize to unconstrained application environments. Human view and video scene understanding are often treated separately. Inspired by the human visual system, this paper proposes a view-to-scene video processing method in a cost-efficient way. In real-world applications, this lightweight method can be integrated into robots to help identify human behavior in complex environments. Fewer parameters indicate that the method can be easily migrated to different types of behaviors, and the reduced computational costs represent the ability to achieve real-time performance under limited hardware conditions.
Xuna Wang, Xu Cheng 0003, Zhaojie Ju, Yingke Xu
IEEE Trans Autom. Sci. Eng.4
2025 Wavelet-Discrete Cosine Transform Synergy for Ship Motion-Based Sea State Estimation in Autonomous Ships
abstract
Developing a robust autonomous sea state estimation (SSE) model stands as a pivotal challenge in advancing autonomous ships. Presently, deep learning (DL) methodologies have showcased remarkable efficacy in SSE tasks. Nonetheless, the dynamic nature of ship motion introduces temporal variations alongside frequency domain characteristics like periodic swinging, posing challenges for existing DL approaches. Most prevailing DL techniques, predominantly leveraging Convolutional Neural Networks or Long Short-Term Memory Networks, often fail to effectively harness frequency domain information post feature extraction. To tackle these limitations head-on, this paper introduces a pioneering SSE model. Specifically, in order to solve the frequency-domain feature extraction problem, we design a wavelet transform-based frequency domain encoder to extract relevant frequency-domain features from ship motion data by discriminating the contribution of different frequencies in the signal. Subsequently, in order to better integrate the extracted ship motion features, we designed a Feature Perception module based on discrete cosine transform. This module adeptly merges the extracted feature insights while prioritizing crucial frequency domain features. Following rigorous experimentation, our methodology exhibits superior performance compared to existing baseline techniques in SSE, a capability of profound significance for autonomous ships. Moreover, across diverse public multivariate time series classification datasets, our model outperforms current state-of-the-art approaches, underscoring its scalability across distinct domains.
Feng Xiao 0005, Xu Cheng 0003, Sasa Nikolic 0002, Jianhua Zhang 0002, Shengyong Chen
IEEE Trans Autom. Sci. Eng.3
2025 Advancing Multibehavior Recommendation With Dual-Mode Augmented Contrastive Learning
abstract
Multibehavioral recommender systems improve the prediction accuracy of target behaviors (e.g., purchase) by integrating auxiliary user behaviors (e.g., page view), but often face the challenge of sparse data on target behaviors. Although contrastive learning can effectively alleviate this problem, existing models still have limitations: they reduce behavioral relationships to linear combinations, ignoring users’ unique behavioral combination modes for different items; meanwhile, when constructing contrastive views, they fail to adequately consider the data imbalance between auxiliary and target behaviors, resulting in learned features biased toward auxiliary behaviors. Consequently, we introduced a multibehavioral recommendation model founded on dual mode contrast learning (DMCL). DMCL innovatively defines behavior combination modes and generates interaction graphs from two perspectives: behavior combination and cascading, and designs a specialized graph encoder for each graph to learn node embedding. By introducing adversarial constraints and contrastive learning, DMCL enhances the complementary nature of node embeddings, and at the same time utilizes interaction graphs in different modes to construct a contrast view, which makes full use of the limited interaction information of the target behavior, and effectively mitigates the feature bias problem. Finally, DMCL combines multilayer perceptron to predict recommendation scores. Our comprehensive experiments on three real datasets demonstrate significant improvements in Recall and NDCG metrics, achieving increases of 2.4–10.1% and 2.6–10.3%, respectively, over current state-of-the-art methods.
Xu Cheng 0003, Yingyuan Xiao, Wenguang Zheng
IEEE Trans. Comput. Soc. Syst.2
2025 Learning From Yourself to Others for Unsupervised Visible-Infrared Re-Identification
abstract
Unsupervised visible-infrared person re-identification (US-VI-ReID) aims to match unlabeled pedestrian images captured under varying lighting conditions. The key challenge lies in generating accurate pseudo-labels, alongside alleviating the significant modality gap between visible and infrared modalities. Existing methods mainly focus on mitigating the effects of noisy labels through loss functions during backward propagation. However, these noisy labels already influence the forward propagation, leading to incorrect cross-modality correspondences. To address this issue, we propose a Hierarchical Centrality Collaborative Learning (HCCL) framework for US-VI-ReID, which proactively identifies noisy labels during the forward propagation. The rationale behind HCCL is that intra-modality refinement serves as the foundation for establishing cross-modality correspondences, reflecting the principle of learning from yourself to others. For intra-modality learning, we propose a Closeness Centrality Selection (CCS), quantifying sample confidence using closeness centrality to identify noisy samples. By discarding the noisy samples during forward propagation, CCS mitigates their adverse effects and ensures identity-consistent representation learning. For cross-modality learning, a Hierarchical Consistency Matching (HCM) is proposed to establish local instance-level label associations by leveraging bidirectional consistency with the most reliable samples identified during intra-modality learning. These local associations are then propagated to guide the global cluster-level cross-modality correspondences. Extensive experiments demonstrate that our HCCL achieves competitive performance on mainstream datasets, even surpassing some supervised counterparts. Additionally, outstanding results on corrupted datasets verify its generalizability and robustness.
Wenhui Ji, Xu Cheng 0003, Zhaodong Sun, Guoying Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Prior Knowledge-Driven Hybrid Prompter Learning for RGB-Event Tracking
abstract
Event data can asynchronously capture variations in light intensity, thereby implicitly providing valuable complementary cues for RGB-Event tracking. Existing methods typically employ a direct interaction mechanism to fuse RGB and event data. However, due to differences in imaging mechanisms, the representational disparity between these two data types is not fixed, which can lead to tracking failures in certain challenging scenarios. To address this issue, we propose a novel prior knowledge-driven hybrid prompter learning framework for RGB-Event tracking. Specifically, we develop a frame-event hybrid prompter that leverages prior tracking knowledge from the foundation model as intermediate modal support to mitigate the heterogeneity between RGB and event data. By leveraging its rich prior tracking knowledge, the intermediate modal reduces the gap between the dense RGB and sparse event data interactions, effectively guiding complementary learning between modalities. Meanwhile, to mitigate the internal learning disparities between the lightweight hybrid prompter and the deep transformer model, we introduce a pseudo-prompt learning strategy that lies between full fine-tuning and partial fine-tuning. This strategy adopts a divide-and-conquer approach to assign different learning rates to modules with distinct functions, effectively reducing the dominant influence of RGB information in complex scenarios. Extensive experiments conducted on two public RGB-Event tracking datasets show that the proposed HPL outperforms state-of-the-art tracking methods, achieving exceptional performance.
Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.3
2025 Leveragable Adaptive Multi-Scale Features and Learnable Prototypes for Imbalanced Sea State Estimation Based on Ship Motion Data
abstract
The adoption of autonomous vessels has been accelerated by the flourishing maritime economy and more stringent shipping regulations. Accurate sea state estimation (SSE) is of paramount importance for their safe operation. Traditional SSE methods face several limitations, including subjectivity in manual observation, high radar costs, and the insufficient timeliness of satellite data. In contrast, SSE methods based on ship motion data, particularly deep learning models, offer advantages in capturing complex nonlinear relationships. However, challenges persist, such as imbalanced sea state data, suboptimal performance under extreme conditions, and the complexity of ship motion data, which hampers effective feature extraction. To address these challenges, this paper proposes a novel deep learning model that integrates a dynamic attention mechanism with adaptive selective kernel modules and multi-scale feature fusion to capture spatiotemporal dependencies. Additionally, an enhanced prototype classifier, utilizing cosine similarity and dynamic prototype updating, is introduced to mitigate data imbalance and improve model robustness in extreme conditions. To evaluate the performance of the proposed model, we conducted experiments on 30 benchmark datasets for multivariate time series classification, as well as two sea state datasets. The comparison results demonstrate that our model outperforms most baseline methods in general multivariate time series classification tasks and also surpasses state-of-the-art methods for SSE and class imbalance learning. Furthermore, the experimental results show that the proposed model is robust to data noise and missing values.
Mengna Liu, Xiufeng Liu 0001, Xu Cheng 0003, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.5
2025 GenBEV: Generative Model With Semantic Compensation for Bird's Eye View Segmentation
abstract
Bird’s-Eye View (BEV) semantic segmentation is a key technology for constructing high-precision maps in low-cost visual navigation systems. The main challenge lies in effectively transforming image features into BEV features while preserving rich BEV visual information. Recent works have shown that generative models hold great promise in advancing BEV segmentation. However, these methods primarily focus on producing BEV features using prior knowledge, often overlooking key challenges such as feature shift, confusion, and forgetting during the BEV feature generation process. In this paper, we propose GenBEV, a generative model with semantic compensation that formally addresses inaccuracies and confusion in BEV feature generation. GenBEV leverages the synergistic benefits of data fusion consistency and noise-reduction training to enhance the diversity and reliability of the generated information. This improvement boosts the robustness and generalization of BEV segmentation across diverse scenarios, including those involving complex objects and low-quality images. Specifically, we design an adaptive cross-feature encoder to reduce diffusion variability. During decoding, we integrate the context of BEV features with noisy features to construct semantic embeddings. We show the effectiveness of GenBEV on the nuScenes, KITTI Raw, and KITTI 3D Object datasets. GenBEV achieves segmentation scores of 29.5%, 68.8%, and 39.7%, respectively, surpassing current methods by up to 3.6%, 2.4%, and 2.7%. To the best of our knowledge, GenBEV is the first to address the problem of BEV feature falsification in generative architectures.
Weiming Fan, Yuping Guo, Hong Lyu, Hongwei Gao 0002, Changting Lin, Xu Cheng 0003
IEEE Trans. Intell. Transp. Syst.8
2025 FRAME: Feature Rectification for Class Imbalance Learning
abstract
Class imbalance learning is a challenging task in machine learning applications. To balance training data, traditional class imbalance learning approaches, such as class resampling or reweighting, are commonly applied in the literature. However, these methods can have significant limitations, particularly in the presence of noisy data, missing values, or when applied to advanced learning paradigms like semi-supervised or federated learning. To address these limitations, this paper proposes a novel and theoretically-ensured latentFeatureRectification method for clAss iMbalance lEarning (FRAME). The proposed FRAME can automatically learn multiple centroids for each class in the latent space and then perform class balancing. Unlike data-level methods, FRAME balances feature in the latent space rather than the original space. Compared to algorithm-level methods, FRAME can distinguish different classes based on distance without the need to adjust the learning algorithms. Through latent feature rectification, FRAME can effectively mitigate contaminated noises/missing values without worrying about structural variations in the data. In order to accommodate a wider range of applications, this paper extends FRAME to the following three main learning paradigms: fully-supervised learning, semi-supervised learning, and federated learning. Extensive experiments on 10 binary-class datasets demonstrate that our FRAME can achieve competitive performance than the state-of-the-art methods and its robustness to noises/missing values.
Xu Cheng 0003, Fan Shi 0001, Yao Zhang 0021, Huan Li 0003, Xiufeng Liu 0001, Shengyong Chen
IEEE Trans. Knowl. Data Eng.1
2025 MDANet: Modality-Aware Domain Alignment Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification is a challenging task in video surveillance. Most existing works achieve performance gains by aligning feature distributions or image styles across modalities, whereas the multi-granularity information and domain knowledge are usually neglected. Motivated by these issues, we propose a novel modality-aware domain alignment network (MDANet) for visible-infrared person re-identification (VI-ReID), which utilizes global-local context cues and the generalized domain alignment strategy to solve modal differences and poor generalization. Firstly, modality-aware global-local context attention (MGLCA) is proposed to obtain multi-granularity context features and identity-aware patterns. Secondly, we present a generalized domain alignment learning head (GDALH) to relieve the modality discrepancy and enhance the generalization of MDANet, whose core idea is to enrich feature diversity in the domain alignment procedure. Finally, the entire network model is trained by proposing cross-modality circle, classification, and domain alignment losses in an end-to-end fashion. We conduct comprehensive experiments on two standards and their corrupted VI-ReID datasets to validate the robustness and generalization of our approach. MDANet is obviously superior to the most state-of-the-art methods. Specifically, the proposed method can gain 8.86% and 2.50% in Rank-1 accuracy on SYSU-MM01 (all-search and single-shot mode) and RegDB (infrared to visible mode) datasets, respectively. The source code will be made available soon.
Xu Cheng 0003, Hao Yu 0015, Kevin H. M. Cheng, Zitong Yu, Guoying Zhao 0001
IEEE Trans. Multim.1
2025 DSAF: Dual Space Alignment Framework for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a cross-modality retrieval task that aims to match visible and infrared pedestrian images across non-overlapped cameras. However, we observe that three crucial challenges remain inadequately addressed by existing methods: (i) limited discriminative capacity for modality-shared representation, (ii) modality misalignment, and (iii) neglect of identity consistency knowledge. To solve the above issues, we propose a novel dual space alignment framework (DSAF) to constrain the modality in two specific spaces. Specifically, for (i), we design a lightweight and plug-and-play modality invariant enhancement (MIE) module to capture fine-grained semantic information and render identity discriminative. This facilitates the establishment of correlations between visible and infrared modalities, enabling the model to learn robust modality-shared features. To tackle (ii), a dual space alignment (DSA) is introduced to conduct the pixel-level alignment in both Euclidean space and Hilbert space. DSA establishes an elastic relationship between these two spaces, remaining invariant knowledge across two spaces. To solve (iii), we propose an adaptive identity-consistent learning (AIL) to discover identity-consistent knowledge between visible and infrared modalities in a dynamic manner. Extensive experiments on mainstream VI-ReID benchmarks show the superiority and flexibility of our proposed method, achieving competitive performance on mainstream datasets.
Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001
IEEE Trans. Multim.2
2025 Dual-Path Imbalanced Feature Compensation Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) presents significant challenges on account of the substantial cross-modality gap and intra-class variations. Most existing methods primarily concentrate on aligning cross-modality at the feature or image levels and training with an equal number of samples from different modalities. However, in the real world, there exists an issue of modality imbalance between visible and infrared data. Besides, imbalanced samples between train and test impact the robustness and generalization of the VI-ReID. To alleviate this problem, we propose a dual-path imbalanced feature compensation network (DICNet) for VI-ReID, which provides equal opportunities for each modality to learn inconsistent information from different identities of others, enhancing identity discrimination performance and generalization. First, a modality consistency perception (MCP) module is designed to assist the backbone focus on spatial and channel information, extracting diverse and salient features to enhance feature representation. Second, we propose a cross-modality features re-assignment strategy to simulate modality imbalance by grouping and re-organizing the cross-modality features. Third, we perform bidirectional heterogeneous cooperative compensation with cross-modality imbalanced feature interaction modules (CIFIMs), allowing our network to explore the identity-aware patterns from imbalanced features of multiple groups for cross-modality interaction and fusion. Further, we design a feature re-construction difference loss to reduce cross-modality discrepancy and enrich feature diversity within each modality. Extensive experiments on three mainstream datasets show the superiority of the DICNet. Additionally, competitive results in corrupted scenarios verify its generalization and robustness.
Xu Cheng 0003, Hao Yu 0015, Jingang Shi, Zitong Yu
ACM Trans. Multim. Comput. Commun. Appl.1
2025 A Novel Robustness-Enhancing Adversarial Defense Approach to AI-Powered Sea State Estimation for Autonomous Marine Vessels
abstract
Sea state information is significant for the guide of maritime activities of autonomous vessels. The sea state estimation (SSE) model, powered by artificial intelligence (AI), has shown great effectiveness but is susceptible to malicious data attacks. These attacks can lead to significant declines in the system’s performance and result in incorrect predictions about the sea state. This study introduces SecureSSE, a strategy for protecting SSE models in autonomous marine vessels from adversarial attacks. This approach incorporates three main components: 1) the multiscale feature extraction learning (MFEL) module; 2) the feature convolution aggregation learning (FCAL) module; and 3) the perturbation examples training (PET) module. The PET module is specifically crafted to create perturbation examples that are in line with unaltered data, leveraging the capabilities of both the MFEL and FCAL modules to efficiently extract and integrate detailed features from ship motion data. Our proposed SecureSSE approach is shown to significantly improve the resilience of deep learning models against potential attacks. Through experimental testing, we have validated the effectiveness of this method in enhancing SSE. Additional ablation studies highlight the critical role of each module within the SecureSSE framework. To our knowledge, this is the first study to address adversarial attacks in this context and to propose a comprehensive defense mechanism for SSE systems in autonomous marine vessels.
Xu Cheng 0003, Fan Shi 0001, Hanwei Zhang 0001, Hongning Dai, Houxiang Zhang, Shengyong Chen
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Differentiable Auxiliary Learning for Sketch Re-Identification
abstract
Sketch re-identification (Re-ID) seeks to match pedestrians' photos from surveillance videos with corresponding sketches. However, we observe that existing works still have two critical limitations: (i) cross- and intra-modality discrepancies hinder the extraction of modality-shared features, (ii) standard triplet loss fails to constrain latent feature distribution in each modality with inadequate samples. To overcome the above issues, we propose a differentiable auxiliary learning network (DALNet) to explore a robust auxiliary modality for Sketch Re-ID. Specifically, for (i) we construct an auxiliary modality by using a dynamic auxiliary generator (DAG) to bridge the gap between sketch and photo modalities. The auxiliary modality highlights the described person in photos to mitigate background clutter and learns sketch style through style refinement. Moreover, a modality interactive attention module (MIA) is presented to align the features and learn the invariant patterns of two modalities by auxiliary modality. To address (ii), we propose a multi-modality collaborative learning scheme (MMCL) to align the latent distribution of three modalities. An intra-modality circle loss in MMCL brings learned global and modality-shared features of the same identity closer in the case of insufficient samples within each modality. Extensive experiments verify the superior performance of our DALNet over the state-of-the-art methods for Sketch Re-ID, and the generalization in sketch-based image retrieval and sketch-photo face recognition tasks.
Xu Cheng 0003, Haoyu Chen 0001, Hao Yu 0015, Guoying Zhao 0001
AAAI2
2024 Robust Classification of Incomplete Time Series with Noisy Labels
abstract
Missing data and noisy labeling are common problems in time series analysis. The traditional approach to deal with missing data is to separate interpolation and classification, which is not interactive and provides unsatisfactory performance. While advanced methods can learn features from missing information, feature representation is limited due to the accumulation of interpolation errors. For noisy label interference, a robust loss function is a simpler and more general solution for robust learning. This study proposes an end-to-end neural network that unifies data interpolation and feature learning within a single framework. The focus is placed on extracting useful information from incomplete time series data, and for the computation of classification loss, a robustness loss function is used which effectively reduces the impact of noisy labels. The model is evaluated on 20 univariate time series from the UCR archive after noise processing. The results show that the model outperforms state-of-the-art methods in classifying incomplete time series under noisy labels, especially at high missing rates with high noise rates.
Pengshuai Yao, Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Lili Guo 0001
CSCWD4
2024 Domain Shifting: A Generalized Solution for Heterogeneous Cross-Modality Person Re-Identification
Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001
ECCV (72)2
2024 AdaFSNet: Time Series Classification Based on Convolutional Network with a Adaptive and Effective Kernel Size Configuration
abstract
Time series classification is one of the most critical and challenging problems in data mining, existing widely in various fields and holding significant research importance. Despite extensive research and notable achievements with successful real-world applications, addressing the challenge of capturing the appropriate receptive field (RF) size from one-dimensional or multi-dimensional time series of varying lengths remains a persistent issue, which greatly impacts performance and varies considerably across different datasets. In this paper, we propose an Adaptive and Effective Full-Scope Convolutional Neural Network (AdaFSNet) to enhance the accuracy of time series classification. This network includes two Dense Blocks. Particularly, it can dynamically choose a range of kernel sizes that effectively encompass the optimal RF size for various datasets by incorporating multiple prime numbers corresponding to the time series length. We also design a TargetDrop block, which can reduce redundancy while extracting a more effective RF. To assess the effectiveness of the AdaFSNet network, comprehensive experiments were conducted using the UCR and UEA datasets, which include one-dimensional and multi-dimensional time series data, respectively. Our model surpassed baseline models in terms of classification accuracy, underscoring the AdaFSNet network’s efficiency and effectiveness in handling time series classification tasks.
Haoxiao Wang, Jianhua Zhang 0002, Xu Cheng 0003
IJCNN4
2024 Online Anomaly Detection for Streaming Data in the Presence of Missing Values
abstract
Online anomaly detection is a critical area in data analysis, particularly for handling dynamic data streams and addressing the challenge of concept drift. While current methods for online anomaly detection have achieved significant breakthroughs, creating a system that can continuously and effectively learn in scenarios with missing data remains a formidable challenge. In this paper, we introduce an autoencoder-based online deep anomaly detection model that addresses both missing data and concept drift. The model features a lightweight module specifically designed for efficient missing value processing. Additionally, it incorporates an adaptive model pool to manage the time-varying concept drift commonly observed in dynamic data streams. This flexible and dynamic management mechanism allows the model to adapt to changes in the data stream, maintaining robust anomaly detection performance across various conditions. Empirical validation of our model through ten comparative experiments on high-dimensional datasets affected by concept drift shows that it outperforms existing state-of-the-art methods. These results underscore the effectiveness and practicality of our approach.
Mengna Liu, Xu Cheng 0003, Lei Song 0011, Jianhua Zhang 0002
SMC3
2024 EE-MVSNet: Deep Learning-Based Cascaded High-Precision Multi-View Stereo Network with ECA & EVC
abstract
Multi-view stereo (MVS) has emerged as a pivotal algorithm in 3D reconstruction, garnering significant research attention over the past several decades. While recent coarse-to-fine methods have demonstrated promising results in enhancing the reconstruction quality of traditional algorithms, they often neglect the crucial aspect of feature layer refinement. Additionally, these methods face the challenge of low-cost feature matching. To address these limitations, we propose a novel learning-based MVS framework(EE-MVSNet). Firstly, we propose a novel approach incorporating an explicit visual center (EVC) module within the feature pyramid network (FPN), strengthening the adjustment within feature layers and improving model accuracy. Furthermore, we introduce the ECA+3DCNN module, which utilizes channel attention to alleviate the problem of low-cost feature matching. Finally, our model achieves competitive performance through extensive experimentation on the DTU dataset, showcasing its high-quality 3D reconstruction.
Changfei Kong, Jiafa Mao, Xu Cheng 0003, Sixian Chan 0001
SMC4
2024 A Federated Learning Framework for Cloud-Edge Collaborative Fault Diagnosis of Wind Turbines
abstract
In modern Internet of Things-enhanced wind power systems, most existing data-driven fault diagnosis approaches for wind turbines (WTs) are performed under a centralized paradigm that ignores data privacy. Recently, federated learning (FL) presented a solution to enable edge WTs located at isolated sites to collaboratively learn a shared diagnosis model without accessing local privacy-sensitive data. However, the practical issues of fault label heterogeneity among edge clients and scarcity of labeled data still severely impede the generation of a satisfactory diagnosis model. To address these issues, we propose a diagnostic knowledge-based FL framework (DKFLWT) for collaborative fault diagnosis of distributed edge WTs. In our DKFLWT framework, independently learned diagnostic knowledge from each edge client, rather than model parameters in conventional FL, is uploaded to the cloud server to enrich the client-specific information visible to the server and mitigate the adverse effects on model performance caused by label heterogeneity. To enhance the overall efficiency of the framework, we develop a two-stage, single-round training mechanism, in which the cloud server serves as a universal platform that can accommodate the customized requirements of users, implying the convenient integration of semi-supervised learning to enhance the diagnosis performance in scenarios with limited labeled data. Furthermore, a spatio-temporal memory-enhanced autoencoder is designed to sufficiently exploit essential diagnostic knowledge of different fault patterns from each client. Experimental results demonstrate superior diagnosis performance of our DKFLWT framework with an improvement of more than 22.1% in accuracy and 37.2% in training efficiency against several compared methods in all seriously heterogeneous scenarios.
Guoqian Jiang, Xiufeng Liu 0001, Xu Cheng 0003
IEEE Internet Things J.4
2024 Exploring modality enhancement and compensation spaces for visible-infrared person re-identification
Xu Cheng 0003, Shuya Deng, Hao Yu 0015
Image Vis. Comput.1
2024 Discovering attention-guided cross-modality correlation for visible-infrared person re-identification
Hao Yu 0015, Xu Cheng 0003, Kevin H. M. Cheng, Wei Peng 0009, Zitong Yu, Guoying Zhao 0001
Pattern Recognit.2
2024 Multilevel Signal Decomposition Layer-Specific Residual Network for Blade Icing Prediction
abstract
Wind energy is crucial for sustainable systems but faces reduced productivity due to blade icing. Current detection methods are either costly or heavily reliant on domain-specific knowledge. Data-driven methods show promising performance but encounter challenges such as extracting multi-scale features for blade icing detection from noisy sensor data and addressing the imbalance between icing and non-icing states. To overcome these challenges, we propose an innovative data-driven approach named Multilevel Signal Decomposition Layer-Specific Residual Network (MSD-LRN) for blade icing detection. Our model first employs wavelet decomposition to extract multi-scale features from noisy sensor data and then uses heterogeneous structures to learn the hidden knowledge at each scale. This heterogeneous structure is designed to address the issue of inconsistent information across different scales in wavelet transforms, where information successively decreases. We address data imbalance using resampling techniques. Our approach is validated on three blade icing datasets with varying imbalance ratios, achieving F1 scores of 89.60%, 83.02%, and 78.18%, surpassing existing baselines. Additionally, we introduce random Gaussian noise to test the model's ability to learn robust features from noisy data through wavelet decomposition.
Sizhuo Chen, Mengna Liu, Fan Shi 0001, Xu Cheng 0003
IEEE Signal Process. Lett.5
2024 High-Order Spatial Interactions Enhanced Lightweight Model for Optical Remote Sensing Image-Based Small Ship Detection
abstract
Accurate and reliable optical remote sensing image-based small-ship detection is crucial for maritime surveillance systems, but existing methods often struggle with balancing detection performance and computational complexity. In this article, we propose a novel lightweight framework called HSI-ShipDetectionNet that is based on high-order spatial interactions (HSIs) and is suitable for deployment on resource-limited platforms, such as satellites and unmanned aerial vehicles. HSI-ShipDetectionNet includes a prediction branch specifically for tiny ships and a lightweight hybrid attention block (LHAB) for reduced complexity. In addition, the use of an HSI module improves advanced feature understanding and modeling ability. Our model is evaluated using the public Kaggle and FAIR1M marine ship detection datasets and compared with multiple state-of-the-art models including small object detection models, lightweight detection models, and ship detection models. The results show that HSI-ShipDetectionNet outperforms the other models in terms of detection performance while being lightweight and suitable for deployment on resource-limited platforms.
Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Huan Huo, Shengyong Chen
IEEE Trans. Geosci. Remote. Sens.2
2024 Class-Imbalanced Spatial-Temporal Feature Learning for Blade Icing Recognition of Wind Turbine
abstract
Blade icing detection is vital for wind turbines in cold climates, as it can prevent revenue loss and power degradation. Many machine learning models have been proposed to improve the detection of blade icing; however, earlier studies do not adequately address these issues due to the dynamics of sensor correlations and the imbalance of blade icing data, resulting in low precision and a high false alarm rate. In this study, we aim to address both of these challenges in order to identify blade icing more accurately. On this premise, we develop a spatial–temporal graph convolutional network (SGCN) that leverages the graph convolutional network for adaptively analyzing the dynamics of sensor correlations and a distance-based classifier to improve imbalanced learning. Experiments on the public UEA time series classification datasets and the real-world wind turbine datasets indicate that SGCN is capable of state-of-the-art accuracy, especially in the case of extremely imbalanced data.
Renfang Wang, Hong Qiu, Guoqian Jiang, Xiufeng Liu 0001, Xu Cheng 0003
IEEE Trans. Ind. Informatics5
2024 Bilevel Fusion With Local and Global Cues for Point Cloud Upsampling
abstract
This study focuses on point cloud upsampling, crucial in 3-D data processing but hindered by current 3-D sensor limitations. Point clouds from RGB-D cameras and light detection and ranging (LiDAR) scanners are often sparse, noisy, and irregular, challenging traditional processing methods reliant on prior knowledge and hindering detail preservation. Despite deep learning's transformative impact, issues like hole overfitting and insufficient local-global feature fusion persist. To address these, we introduce the bilevel fusion point cloud upsampling (BiPU) network. It features a parallel extractor for simultaneous local and global feature extraction and a consistency-based feature alignment module employing cross-attention for enhanced multiscale feature transfer. BiPU also incorporates 4-D encoding for rotational invariance and depthwise separable convolutions to reduce complexity and parameters. Tested across multiple datasets, BiPU excels in maintaining hole contours and reducing costs, marking a notable advancement in point cloud processing.
Yunrui Zhu, Xu Cheng 0003, Jianhua Zhang 0002
IEEE Trans. Ind. Informatics3
2024 ST-Phys: Unsupervised Spatio-Temporal Contrastive Remote Physiological Measurement
abstract
Remote photoplethysmography (rPPG) is a non-contact method that employs facial videos for measuring physiological parameters. Existing rPPG methods have achieved remarkable performance. However, the success mainly profits from supervised learning over massive labeled data. On the other hand, existing unsupervised rPPG methods fail to fully utilize spatio-temporal features and encounter challenges in low-light or noise environments. To address these problems, we propose an unsupervised contrast learning approach, ST-Phys. We incorporate a low-light enhancement module, a temporal dilated module, and a spatial enhanced module to better deal with long-term dependencies under the random low-light conditions. In addition, we design a circular margin loss, wherein rPPG signals originating from identical videos are attracted, while those from distinct videos are repelled. Our method is assessed on six openly accessible datasets, including RGB and NIR videos. Extensive experiments reveal the superior performance of our proposed ST-Phys over state-of-the-art unsupervised rPPG methods. Moreover, it offers advantages in parameter reduction and noise robustness.
Mingyue Cao, Xu Cheng 0003, Hao Yu 0015, Jingang Shi
IEEE J. Biomed. Health Informatics2
2024 Selective Feature Fusion and Irregular-Aware Network for Pavement Crack Detection
abstract
Road cracks on highways and main roads are among the most prominent defects. Given the inherent inaccuracy, time-consuming nature, and labor intensiveness of manual road crack detection, there’s a compelling need for automated solutions. The irregular shape of cracks, along with complex background conditions encompassing varying lighting, tree shadows, and dark stains, poses a significant challenge for computer vision-based approaches. Most cracks exhibit irregular edge patterns, which are pivotal features for accurate detection. In response to recent advancements in deep learning within the realm of computer vision, this paper introduces an innovative neural network architecture termed the ‘Selective Feature Fusion and Irregular-Aware Network (SFIAN)’ designed specifically for crack detection on pavements. The proposed network selectively integrates features from multiple levels, enhancing and controlling the flow of valuable information at each stage while effectively modeling irregular crack objects. In an extensive evaluation, this paper conducts experiments on five distinct crack datasets and compares the results with twelve state-of-the-art crack detection methods, including the latest edge detection and semantic segmentation techniques. The experimental findings demonstrate the superior performance of the proposed method, surpassing baseline methods by a notable margin, with an increase of approximately 13.3% in the F1-score, all without introducing additional time complexity. Furthermore, the model achieves real-time processing, achieving a remarkable speed of 35 frames per second (FPS) on images at 320$\times$480 pixels, facilitated by NVIDIA 3090 hardware.
Xu Cheng 0003, Fan Shi 0001, Meng Zhao 0001, Xiufeng Liu 0001, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.1
2024 SAFENESS: A Semi-Supervised Transfer Learning Approach for Sea State Estimation Using Ship Motion Data
abstract
Autonomous vessels have been identified as a promising innovation in advancing marine transportation, providing an effective means to mitigate the risk of accidents, pollution incidents, and carbon dioxide emissions. Accurate sea state estimation (SSE) plays a critical role in facilitating onboard decision-making and optimizing operational efficiency for autonomous ships. Traditional SSE approaches relying on external sensors, such as wave buoys and wave radars, are limited by cost considerations. Model-based methods are highly relying on the understanding of human knowledge to ships. Data-driven models also provide promising solutions, but their generalization is low. To address this challenge, a semi-supervised transfer learning approach for SSE (SAFENESS) is proposed. The model is trained using sufficient data in the source ship and limited data from the target ship and finally applied to the target ship. A data alignment algorithm is utilized to use the limited data of the target ship. To enhance the learning capability of the framework, two attention mechanisms are proposed, and a multi-class adversarial discriminator is introduced that can align the distributions of different domains. The effectiveness of our approach is validated through comprehensive comparisons with eleven established transfer learning methods, demonstrating the superiority of our model. The competitiveness of the proposed attention modules is verified by comparing them with state-of-the-art attention modules. The significance of each component and the influence of key parameters have been thoroughly explored in the ablation and sensitivity analysis. Our method has potential applications in maritime safety, navigation, and operation optimization.
Xu Cheng 0003, Guoyuan Li, Robert Skulstad, Houxiang Zhang
IEEE Trans. Intell. Transp. Syst.1
2024 A Prototype-Empowered Kernel-Varying Convolutional Model for Imbalanced Sea State Estimation in IoT-Enabled Autonomous Ship
abstract
Sea State Estimation (SSE) is essential for Internet of Things (IoT)-enabled autonomous ships, which rely on favorable sea conditions for safe and efficient navigation. Traditional methods, such as wave buoys and radars, are costly, less accurate, and lack real-time capability. Model-driven methods, based on physical models of ship dynamics, are impractical due to wave randomness. Data-driven methods are limited by the data imbalance problem, as some sea states are more frequent and observable than others. To overcome these challenges, we propose a novel data-driven approach for SSE based on ship motion data. Our approach consists of three main components: a data preprocessing module, a parallel convolution feature extractor, and a theoretical-ensured distance-based classifier. The data preprocessing module aims to enhance the data quality and reduce sensor noise. The parallel convolution feature extractor uses a kernel-varying convolutional structure to capture distinctive features. The distance-based classifier learns representative prototypes for each sea state and assigns a sample to the nearest prototype based on a distance metric. The efficiency of our model is validated through experiments on two SSE datasets and the UEA archive, encompassing thirty multivariate time series classification tasks. The results reveal the generalizability and robustness of our approach.
Mengna Liu, Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Hongning Dai, Shengyong Chen
IEEE Trans. Sustain. Comput.2
2023 TOPLight: Lightweight Neural Networks with Task-Oriented Pretraining for Visible-Infrared Recognition
abstract
Visible-infrared recognition (VI recognition) is a challenging task due to the enormous visual difference across heterogeneous images. Most existing works achieve promising results by transfer learning, such as pretraining on the ImageNet, based on advanced neural architectures like ResNet and ViT. However, such methods ignore the neg-ative influence of the pretrained colour prior knowledge, as well as their heavy computational burden makes them hard to deploy in actual scenarios with limited resources. In this paper, we propose a novel task-oriented pretrained lightweight neural network (TOPLight) for VI recognition. Specifically, the TOPLight method simulates the domain conflict and sample variations with the proposed fake do-main loss in the pretraining stage, which guides the network to learn how to handle those difficulties, such that a more general modality-shared feature representation is learned for the heterogeneous images. Moreover, an effective fine-grained dependency reconstruction module (FDR) is developed to discover substantial pattern dependencies shared in two modalities. Extensive experiments on VI person re-identification and VI face recognition datasets demonstrate the superiority of the proposed TOPLight, which signifi-cantly outperforms the current state of the arts while de-manding fewer computational resources.
Hao Yu 0015, Xu Cheng 0003, Wei Peng 0009
CVPR2
2023 Understanding crowd energy consumption behaviors
Xiufeng Liu 0001, Xu Cheng 0003, Yanyan Yang 0002, Huan Huo, Yongping Liu, Per Sieverts Nielsen
EDBT2
2023 Modality Unifying Network for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a challenging task due to large cross-modality discrepancies and intra-class variations. Existing methods mainly focus on learning modality-shared representations by embedding different modalities into the same feature space. As a result, the learned feature emphasizes the common patterns across modalities while suppressing modality-specific and identity-aware information that is valuable for Re-ID. To address these issues, we propose a novel Modality Unifying Network (MUN) to explore a robust auxiliary modality for VI-ReID. First, the auxiliary modality is generated by combining the proposed cross-modality learner and intra-modality learner, which can dynamically model the modality-specific and modality-shared representations to alleviate both cross-modality and intra-modality variations. Second, by aligning identity centres across the three modalities, an identity alignment loss function is proposed to discover the discriminative feature representations. Third, a modality alignment loss is introduced to consistently reduce the distribution distance of visible and infrared images by modality prototype modeling. Extensive experiments on multiple public datasets demonstrate that the proposed method surpasses the current state-of-the-art methods by a significant margin.
Hao Yu 0015, Xu Cheng 0003, Wei Peng 0009, Guoying Zhao 0001
ICCV2
2023 SANet: A novel segmented attention mechanism and multi-level information fusion network for 6D object pose estimation
Xinbo Geng, Fan Shi 0001, Xu Cheng 0003, Mianzhao Wang, Shengyong Chen, Hongning Dai
Comput. Commun.3
2023 POEM: A prototype cross and emphasis network for few-shot semantic segmentation
Xu Cheng 0003, Shuya Deng, Yonghong Peng
Comput. Vis. Image Underst.1
2023 A late-mover genetic algorithm for resource-constrained project-scheduling problems
abstract
The Resource-Constrained Project Scheduling Problem (RCPSP) plays a critical role in various management applications. Despite its importance, research efforts are still ongoing to improve lower bounds and reduce deviation values. This study aims to develop an innovative and straightforward algorithm for RCPSPs by integrating the “1+1” evolution strategy into a genetic algorithm framework. Unlike most existing studies, the proposed algorithm eliminates the need for parameter tuning and utilizes real-valued numbers and path representation as chromosomes. Consequently, it does not require priority rules to construct a feasible schedule. The algorithm's performance is evaluated using the RCPSP benchmark and compared to alternative algorithms, such as cWSA, Hybrid PSO, and EESHHO. The experimental results demonstrate that the proposed algorithm is competitive, while the exploration capability remains a challenge for further investigation.
Yongping Liu, Xiufeng Liu 0001, Guomin Ji, Xu Cheng 0003, Erling Onstein
Inf. Sci.5
2023 A Difference Enhanced Neural Network for Semantic Change Detection of Remote Sensing Images
abstract
Deep learning techniques have been widely used for semantic change detection (SCD) of remote sensing images (RSIs) and have shown encouraging performance. In this paper, we propose a novel neural network by embedding the difference enhancement (DE) module into the adjacent layers of ResNet for SCD of RSIs (DESNet), which can pay more attention to the changes of bi-temporal RSIs. Furthermore, we deploy the module of multi-scale parallel sampling spatial-spectral non-local (SSN) after feature extraction, which can effectively improve the robustness to large-scale changes and the integrity of the changed objects by fusing global features that sampled from the multi-scale feature space. The experimental tests demonstrate that our DESNet can achieve state-of-the-art accuracy on the SECOND dataset and the LandSat-SCD dataset.
Renfang Wang, Hucheng Wu, Hong Qiu, Feng Wang 0031, Xiufeng Liu 0001, Xu Cheng 0003
IEEE Geosci. Remote. Sens. Lett.6
2023 Asymmetric Cascade Fusion Network for Building Extraction
abstract
The U-Net-like model has been widely studied in the field of building extraction. However, most of these models are based on locally sensed Convolutional Neural Networks(CNNs) designed with symmetric structure and single feature processing, which cannot accurately identify buildings with different sizes, shapes, and colors in remote sensing images. To overcome these problems, we propose the asymmetric cascade fusion network(ACFN), based on the Vision Transformer(ViT), to design a novel asymmetric architecture to recognize buildings of different sizes and shapes by processing multi-granularity features by different means. First, the asymmetric architecture obtains multi-granularity features with global contextual information by embedding different types of attention in encoder-decoders of different sizes. This architecture can identify densely distributed and occluded buildings by semantic reasoning in remote sensing images with complex information. Second, we design a multi-branch weighted pyramid pooling module, which sets different branch weights to offset the background noise introduced in introducing global contextual information. Our ACFN significantly improves the Beijing buildings, ISPRS-Vaihingen, and LoveDA datasets.
Sixian Chan 0001, Yuan Wang 0032, Yanjing Lei, Xu Cheng 0003, Wei Wu 0029
IEEE Trans. Geosci. Remote. Sens.4
2023 Visual Object Tracking Based on Light-Field Imaging in the Presence of Similar Distractors
abstract
Visual object tracking is of great importance in the field of computer vision. One of the main challenges is the difficulty of identifying moving targets from nearby similar distractors with a single-view image of the scene. To overcome this challenge, in this article, we acquire multiview images of the scenes by using a light-field camera. The multiview images are able to capture the 4-D structure instead of the 2-D plane of the objects but are more difficult to process. Therefore, we propose a novel representation for multiview images, i.e., the macro-epipolar plane image (macro-EPI), which highlights both spatial topological and angular information of the target and distractors. It is obtained by slicing the original multiview images into pieces and properly restacking these pieces in an ordinal manner. The resulting macro-EPI is mapped into the 2-D space; therefore, we adapt a modified autoencoder network to train a macro-EPI feature extractor. Thereafter, we design a composite framework of two-pattern convolution filters based on a discriminative correlation filter for object tracking, which successfully discriminates the target from the distractors by merging the macro-EPI features and the single-view image features. The experiments also show that our method outperforms the state-of-the-art methods in the presence of similar distractors.
Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Yao Zhang 0021, Shengyong Chen
IEEE Trans. Ind. Informatics3
2023 A Novel Class-Imbalanced Ship Motion Data-Based Cross-Scale Model for Sea State Estimation
abstract
Sea state estimation (SSE) is significant to the development of autonomous ships, which can enhance the sustainable development of maritime transportation. Traditional model-based methods are limited by their drawbacks, such as high costs and inaccurate estimations. The deep learning model shows superior performance, but it requires that the sample quantity for each sea state should be almost the same. Since the occurrence probability of each state is different, the ships mainly work in low sea states, and the collected ship motion data for different sea states are highly imbalanced. This work proposes a novel class-imbalanced ship motion data-based cross-scale model for SSE. The model consists of three major components: a multi-scale feature learning module, a cross-scale feature learning module, and a prototype classifier module. The multi-scale and cross-scale feature learning modules are designed to learn abundant coarse and fine-level features from the ship motion data. The prototype classifier is utilized to overcome the limitation of the conventional softmax classifier to produce better estimates. Our research highlights our model’s remarkable scalability and versatility with 30 publicly available datasets in time series classification, demonstrating superior performance over baseline methods in 21 cases. Notably, it outperformed ShapeNet by 5.72% and EDI by 26.3%. We further validated our model’s proficiency using ship motion datasets, consistently surpassing eight state-of-the-art baselines and five class-imbalanced learning methods. Ablation and sensitivity studies, emphasize the critical role of each model component. Our findings underscore the model’s robustness and its potential to advance time series classification in diverse domains.
Xu Cheng 0003, Xiufeng Liu 0001, Fan Shi 0001, Zhengru Ren, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.1
2023 Coordination and Optimization Control Framework for Vessels Platooning in Inland Waterborne Transportation System
abstract
Vessels sailing in a single platoon could reduce resistance from the perspective of the whole platoon and the individual vessel, and contribute to improving energy benefits. Moreover, transportation energy costs and traffic efficiency are essential indicators for measuring waterborne transportation systems. We attempt to minimize transportation energy costs by coordinating platoon formation using a distributed framework of controllers. A large-scale coordinated vessel platooning program is proposed to minimize transportation energy costs and optimize traffic efficiency while guaranteeing safety. The control framework covers routing, energy consumption-dependent cooperative platooning decision and speed optimization based on graph search algorithm, cluster analysis, optimal control approach and model predictive control. Firstly, a local scheduling strategy combined with the leader vessel selection algorithm is adopted. Furthermore, we used cluster analysis to create a series of mergeable vessel platooning sets. Then, we used the mathematical planning method and a two-step hybrid optimal control approach to calculate the improvement and optimization of each vessel platoon’s path and speed. Finally, the scalability of the scheduling strategy is elucidated. In a simulation of large scale inland waterborne network, savings surpassed 3.5% when six hundreds vessels participated in the system. These simulation results reveal that the scheduling strategy coordinating vessels into vessel platooning, which improves transportation efficiency as well as descends cost, comparing to a fixed origin route in the waterway network.
Man Zhu, Shengyong Chen, Xu Cheng 0003, Yuanqiao Wen, Weidong Zhang 0004, Rudy R. Negenborn, Yusong Pang
IEEE Trans. Intell. Transp. Syst.4
2022 LFBCNet: Light Field Boundary-aware and Cascaded Interaction Network for Salient Object Detection
abstract
In light field imaging techniques, the abundance of stereo spatial information aids in improving the performance of salient object detection. In some complex scenes, however, applying the 4D light field boundary structure to discriminate salient objects from background regions is still under-explored. In this paper, we propose a light field boundary-aware and cascaded interaction network based on light field macro-EPI, named LFBCNet. Firstly, we propose a well-designed light field multi-epipolar-aware learning (LFML) module to learn rich salient boundary cues by perceiving the continuous angle changes from light field macro-EPI. Secondly, to fully excavate the correlation between salient objects and boundaries at different scales, we design multiple light field boundary interactive (LFBI) modules and cascade them to form a light field multi-scale cascade interaction decoder network. Each LFBI is assigned to predict exquisite salient objects and boundaries by interactively transmitting the salient object and boundary features. Meanwhile, the salient boundary features are forced to gradually refine the salient object features during the multi-scale cascade encoding. Furthermore, a light field multi-scale-fusion prediction (LFMP) module is developed to automatically select and integrate multi-scale salient object features for final saliency prediction. The proposed LFBCNet can accurately distinguish tiny differences between salient objects and background regions. Comprehensive experiments on large benchmark datasets prove that the proposed method achieves competitive performance over 2-D, 3-D, and 4-D salient object detection methods.
Mianzhao Wang, Fan Shi 0001, Xu Cheng 0003, Meng Zhao 0001, Yao Zhang 0021, Shengyong Chen
ACM Multimedia3
2022 Dynamic Temporal-Spatial Regularization-Based Channel Weight Correlation Filter for Aerial Object Tracking
abstract
Correlation filter (CF) has drawn extensive interest in aerial object tracking due to its remarkable performance. Recently, the popular CF methods based on temporal–spatial regularization have been proved to be able to effectively improve the tracking results. However, the boundary effect and filter template degradation still influence the speed and accuracy of the trackers. To handle the two problems, a novel dynamic temporal–spatial regularization-based channel weighted tracking (DTSCT) method was proposed in this work. First, we attempted to employ the saliency detection technique to describe object variation for weakening the boundary effect. Then, the filter template was introduced to the temporal regularization to alleviate the template degradation. In addition, an adaptive weighting strategy was utilized to remove data redundancy in the feature channels. Experiments on three benchmark datasets showed the competitive performance of our DTSCT approach compared to the state-of-the-art methods.
Licheng Jiang, Yuhui Zheng, Xu Cheng 0003, Byeungwoo Jeon
IEEE Geosci. Remote. Sens. Lett.3
2022 Multi-Task Convolution Operators With Object Detection for Visual Tracking
abstract
Recently, multi-task correlation filters has drawn much attention in the object tracking field, which utilizes the multi-task learning (MTL) approach to explore the interdependencies among deep features for object tracking. However, the existing multi-task correlation filters based method fails to consider the relations between the correlation filters. To address this problem, a novel correlation filters based visual tracking method is proposed in this paper, with the integration of multi-task convolution operators and object detection. In our method, convolution and correction filters are jointly learnt through using the MTL technique, with the purpose of exploring not only the interdependencies of deep features but also the internal relevance of the convolution filters. In addition, object detection is introduced into our algorithm to handle the problem of object missing to ensure a better performance of our tracking method. Experiments on five benchmark datasets demonstrate that the proposed visual tracking method outperforms existing state-of-the-art approaches.
Yuhui Zheng, Xinyan Liu 0002, Bin Xiao 0002, Xu Cheng 0003, Yi Wu 0001, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.4
2022 A Class-Imbalanced Heterogeneous Federated Learning Model for Detecting Icing on Wind Turbine Blades
abstract
Wind farms are typically located at high latitudes, resulting in a high risk of blade icing. Data-driven approaches offer promising solutions for blade icing detection, but they rely on a considerable amount of data. Data exchange between multiple wind farms would improve the performance of detection models, due to the spatio-temporal dependencies capable of reflecting different meteorological conditions. The traditional centralized approach for icing detection faces many challenges, including the requirement of high storage and computational capacity of the server, vulnerability to cyberattacks, and operators’ reluctance of sharing data for commercial reasons. To address these challenges, this article proposes a heterogeneous federated learning (FL) model for wind turbine blade icing detection. The structures of the server and client models in the presented method are different, in contrast to the traditional FL of sharing the same structure. In addition, this article addresses the class imbalance problem in the training data. Last, this article conducts comprehensive experiments to evaluate the proposed method using real-world data from 20 turbines in two wind farms, and compares it with two state-of-the-art FL models and five well-known class imbalance methods. The experimental results verify the effectiveness and superiority of the proposed method.
Xu Cheng 0003, Fan Shi 0001, Yongping Liu, Jiehan Zhou, Xiufeng Liu 0001
IEEE Trans. Ind. Informatics1
2022 A Blockchain-Empowered Cluster-Based Federated Learning Model for Blade Icing Estimation on IoT-Enabled Wind Turbine
abstract
Wind energy is a fast-growing renewable energy but faces blade icing. Data-driven methods provide talented solutions for blade icing detection, but a considerable amount of Internet of Things data needs to be collected to a central server, which may lead to the leakage of sensitive business data. To address this limitation, this article proposesBLADE, a Blockchain-empowered imbalanced federated learning (FL) model for blade icing detection. With the help of the Blockchain, the conventional FL is improved without worrying about the failure of the single centralized server and boosts the privacy preserving. A validation mechanism is introduced into the Blockchain to enhance the defense against poisoning attacks. In addition, a novel imbalanced learning algorithm is integrated into BLADE to solve the class imbalance problem in the sensor data. BLADE is evaluated on ten wind turbines from two wind farms. The experimental results verify the effectiveness, superiority, and feasibility of the proposed BLADE.
Xu Cheng 0003, Fan Shi 0001, Meng Zhao 0001, Shengyong Chen, Hao Wang 0003
IEEE Trans. Ind. Informatics1
2022 An Uncertainty-Aware Hybrid Approach for Sea State Estimation Using Ship Motion Responses
abstract
Understanding current environmental conditions is essential for autonomous ships, among which real-time estimation of sea conditions is a key aspect. Considering the ship as a large wave buoy, the sea state can be estimated from motion responses without extra sensors installed. This task is challenging since the relationship between the wave and the ship motion is hard to model. Existing methods include a wave buoy analogy (WBA) method, which assumes linearity between wave and ship motion, and a machine learning (ML) approach. Since the data collected from a vessel in the real world are typically limited to a small range of sea states, the ML method might fail when the encountered sea state is not in the training dataset. This article proposes a hybrid approach that combines the above two methods. The ML method is compensated by the WBA method based on the uncertainty of estimation results, and thus, the failure can be avoided. Real-world historical data from the Research Vessel Gunnerus are applied to validate the approach. Results indicate that the hybrid approach improves the estimation accuracy.
Peihua Han, Guoyuan Li, Xu Cheng 0003, Stian Skjong, Houxiang Zhang
IEEE Trans. Ind. Informatics3
2022 Data-Driven Modeling for Transferable Sea State Estimation Between Marine Systems
abstract
Sea state estimation is beneficial for marine systems to enhance on-board decision-making and improve work efficiency. In the era of ship intelligence, artificial intelligence has greatly promoted the technology of sensing environment, such as by using the deep learning. However, it is difficult to collect enough motion data from a marine system to train a deep learning model. In addition, the model for sea state estimation is trained using the data from a specific marine system; applying the model directly to another marine system may result in performance degradation. In this paper, a supervised transfer learning based framework for sea state estimation (STLSSE) is proposed. The STLSSE focuses on knowledge transfer when the collected data for the source marine system is sufficient but the collected data of the target marine system is scarce. In STLSSE, a data pairing algorithm is proposed to determine the relationship of the source and the target marine system. Based on these paired data, a Siamese convolutional neural network, including a new proposed residual fully convolutional network and two novel attention modules, is designed for the semantic alignment. Moreover, the conventional contrastive loss is improved to characterize the distributions when there are only few samples in the target marine system. The extensive comparisons between STLSSE and state-of-the-art transfer learning approaches show its superior performance. The comparisons with state-of-the-art attention modules has verified the competitiveness of the proposed attention modules. The key parameters and each component of STLSSE are emphasized in the ablation and sensitivity studies.
Xu Cheng 0003, Guoyuan Li, Peihua Han, Robert Skulstad, Shengyong Chen, Houxiang Zhang
IEEE Trans. Intell. Transp. Syst.1
2022 An Online Multiobject Tracking Network for Autonomous Driving in Areas Facing Epidemic
abstract
Multi-object tracking is of great importance in autonomous driving. However, with the outbreak of COVID-19, multi-object tracking faces new challenges in areas gripped by epidemics because of complex motion blur, frequent occlusions, and appearance deformations. To reliably improve object trajectory association in epidemic-plagued areas, we propose a temporal-spatial aggregation embedding network (TSAEN) for multi-object tracking. Our embedding network contains a temporal-aware correlation module (TACM) and spatial-aggregate embedding module (SAEM) that can fully obtain and aggregate appearance clues related to moving objects in previous frames. The TACM learns the temporal homogeneity features of the current and previous frames to perceive features with correlated appearance cues. Then, the SAEM adjusts the spatial deformation for each perceived temporal homogeneity feature and aggregates them for re-ID embedding learning. The experimental results demonstrate that our proposed method is able to achieve excellent overall performance.
Mianzhao Wang, Fan Shi 0001, Meng Zhao 0001, Xu Cheng 0003
IEEE Trans. Intell. Transp. Syst.8
2022 A Novel Deep Class-Imbalanced Semisupervised Model for Wind Turbine Blade Icing Detection
abstract
Wind energy is of great importance for future energy development. In order to fully exploit wind energy, wind farms are often located at high latitudes, a practice that is accompanied by a high risk of icing. Traditional blade icing detection methods are usually based on manual inspection or external sensors/tools, but these techniques are limited by human expertise and additional costs. Model-based methods are highly dependent on prior domain knowledge and prone to misinterpretation. Data-driven approaches can offer promising solutions but require a massive amount of labeled training data, which are not generally available. In addition, the data collected for icing detection tend to be imbalanced because, most of the time, wind turbines operate under normal conditions. To address these challenges, this article presents a novel deep class-imbalanced semisupervised (DCISS) model for estimating blade icing conditions. DCISS integrates class-imbalanced and semisupervised learning (SSL) using a prototypical network that can rebalance features and measure the similarities between labeled and unlabeled samples. In addition, a channel calibration attention module is proposed to improve the ability to extract features from raw data. The proposed model has been evaluated using the blade icing datasets of three wind turbines. Compared to the classical anomaly detection and state-of-the-art SSL algorithms, DCISS shows significant advantages in terms of accuracy. Compared to five different class-imbalanced loss functions, the proposed DCISS is competitive. The generalization and practicability of the proposed model are further verified in the use case of online estimation.
Xu Cheng 0003, Fan Shi 0001, Xiufeng Liu 0001, Meng Zhao 0001, Shengyong Chen
IEEE Trans. Neural Networks Learn. Syst.1
2021 ACFIM: Adaptively Cyclic Feature Information-Interaction Model for Object Detection
Xu Cheng 0003, Daqiu Li
PRCV (1)2
2021 Siamese network for object tracking with multi-granularity appearance representations
Zhuoyi Zhang, Yifeng Zhang 0001, Xu Cheng 0003, Guojun Lu
Pattern Recognit.3
2020 SpectralSeaNet: Spectrogram and Convolutional Network-based Sea State Estimation
abstract
Sea State is significant to the operations on the sea. The traditional model-based approaches need lots of knowledge of vessels, which limit the real-world use. This paper proposes a spectrogram-based deep learning model for sea state estimation (SpectralNet). In this model, the ship motion data is converted to spectrogram using short time Fourier transform (STFT). Unlike other methods, the spectrogram of each sensor will be combined to a new image. And then, a 2D convolutional neural network (CNN) is built as the classifier and the sea state can be identified. The experimental results show the proposed approach can achieve higher classification accuracy compared these methods applied directly in raw time series data. Through the comparison results of the proposed approach and the combination of spectrogram of different number of sensors, the proposed approach can achieve highest classification accuracy, and the classification accuracy is growing with the number of combined sensors. The sensitivity analysis finds the classification accuracy is easily influenced by the scale factor of images.
Xu Cheng 0003, Guoyuan Li, Robert Skulstad, Houxiang Zhang, Shengyong Chen
IECON1
2020 Residual Attention SiameseRPN for Visual Tracking
Xu Cheng 0003, Enlu Li, Zhangjie Fu 0001
PRCV (2)1
2020 Multi-Task Deep Dual Correlation Filters for Visual Tracking
abstract
Correlation filters combined with deep features have delivered impressive results in visual tracking task. However, existing approaches treat deep features produced by different network layers independently, limiting their representation power. To address this issue, this paper proposes a multi-task deep dual correlation filters (MDDCF) based method for robust visual tracking. First, a new multi-task learning scheme is designed to take full advantage of the multi-level features of deep networks, where target representation with individual features is regarded as a single task. As such, the interdependencies between different levels of features can be better explored. Second, we reformulate the objective function of the dual correlation filters and propose a new alternating optimization method, allowing joint training of the correlation filters and network parameters. Third, we design an effective object template update scheme which can well capture the target appearance variations. Extensive experimental evaluations on seven benchmark datasets show that the proposed MDDCF tracker performs favorably against state-ofthe-art methods.
Yuhui Zheng, Xinyan Liu 0002, Xu Cheng 0003, Kaihua Zhang 0001, Yi Wu 0001, Shengyong Chen
IEEE Trans. Image Process.3
2019 Modeling and Analysis of Motion Data from Dynamically Positioned Vessels for Sea State Estimation
abstract
Developing a reliable model to identify the sea state is significant for the autonomous ship. This paper introduces a novel deep neural network model (SeaStateNet) to estimate the sea state based on the ship motion data from dynamically positioned vessels. The SeaStateNet mainly consists of three components: an Long-Short-Term Memory (LSTM) recurrent neural network to capture the long dependency in the ship motion data; a convolutional neural network (CNN) to extract time-invariant features; and a Fast Fourier Transform (FFT) block to extract frequency features. A feature fusion layer is designed to learn the degree affected by each component. The proposed model is applied directly to the raw time series data, without needing of any hand-engineered features. A sensitivity analysis (SA) method is applied to assess the influence of data preprocessing. Through benchmark test and experiment on ship motion dataset, SeaStateNet is verified effective for sea state estimation. The investigation on real-time test further shows the practicality of the proposed model.
Xu Cheng 0003, Guoyuan Li, Robert Skulstad, Shengyong Chen, Hans Petter Hildre, Houxiang Zhang
ICRA1
2018 An Attention-Based Approach for Single Image Super Resolution
abstract
The main challenge of single image super resolution (SISR) is the recovery of high frequency details such as tiny textures. However, most of the state-of-the-art methods lack specific modules to identify high frequency areas, causing the output image to be blurred. We propose an attention-based approach to give a discrimination between texture areas and smooth areas. After the positions of high frequency details are located, high frequency compensation is carried out. This approach can incorporate with previously proposed SISR networks. By providing high frequency enhancement, better performance and visual effect are achieved. We also propose our own SISR network composed of DenseRes blocks. The block provides an effective way to combine the low level features and high level features. Extensive benchmark evaluation shows that our proposed method achieves significant improvement over the state-of-the-art works in SISR.
Yuancheng Wang, Nan Li 0064, Xu Cheng 0003, Yifeng Zhang 0001, Yongming Huang 0001, Guojun Lu
ICPR4
2017 Object Tracking via Temporal Consistency Dictionary Learning
abstract
Sparse representation-based methods have been successfully applied to visual tracking. However, complex and inefficient optimization limits their deployment in practical tracking scenarios. In this paper, we propose a temporal consistency dictionary learning tracking algorithm to enable efficient dictionary learning and tracking executive. First, we present an objective function which introduces the fixed dictionary and variance dictionary to reconstruct the object's appearance. In particular, the proposed method takes the temporal consistency into account by adding a regularization term into the objective function to constrain the difference of object appearance at adjacent frames. Then the optimization problem is solved in an iteration way. Moreover, the proposed method can encode the object's local structural information, and the local patches from the same candidate altogether for a global appearance representation. Second, we develop an effective observation likelihood function based on the proposed model. It takes the influence of patches with large reconstruction errors into consideration, thereby, alleviating the drifting of the object. Finally, we present an appearance updating strategy to adapt to the object's appearance variations by the online dictionary learning. Experimental evaluations on the TB50 and TB100 datasets show that the proposed tracking method outperforms sparse representation related visual tracking as well as other state-of-the-art tracking methods.
Xu Cheng 0003, Yifeng Zhang 0001, Jinshi Cui, Lin Zhou 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2016 Learning semantic context feature-tree for action recognition via nearest neighbor fusion
Tongchi Zhou, Nijun Li, Xu Cheng 0003, Qinjun Xu, Lin Zhou 0001, Zhenyang Wu
Neurocomputing3
2016 Recognizing human interactions by genetic algorithm-based random forest spatio-temporal correlation
Nijun Li, Xu Cheng 0003, Zhenyang Wu
Pattern Anal. Appl.2
2014 Tracking deformable parts via dynamic conditional random fields
abstract
Despite the success of many advanced tracking methods in this area, tracking targets with drastic variation of appearance such as deformation, view change and partial occlusion in video sequences is still a challenge in practical applications. In this paper, we take these serious tracking problems into account simultaneously, proposing a dynamic graph based model to track object and its deformable parts at multiple resolutions. The method introduces well learned structural object detection models into object tracking applications as prior knowledge to deal with deformation and view change. Meanwhile, it explicitly formulates partial occlusion by integrating spatial potentials and temporal potentials with an unparameterized occlusion handling mechanism in the dynamic conditional random field framework. Empirical results demonstrate that the method outperforms state-of-the-art trackers on different challenging video sequences.
Suofei Zhang, Xu Cheng 0003, Lin Zhou 0001, Zhenyang Wu
ICIP2
2014 A Hybrid Method for Human Interaction Recognition Using Spatio-temporal Interest Points
abstract
This paper proposes an innovative and effective hybrid way to recognize human interactions, which incorporates the advantages of both global feature (Motion Context, MC) and Spatio-Temporal (S-T) correlation of local Spatio-Temporal Interest Points (STIPs). The MC feature, which also derives from STIPs, is used to train a random forest where Genetic Algorithm (GA) is applied to the training phase to achieve a good compromise between reliability and efficiency. Besides, we design an effective and efficient S-T correlation based match to assist the MC feature, where MC's structure and a biological sequence matching algorithm are employed to calculate the spatial and temporal correlation score, respectively. Experiments on the UT-Interaction dataset show that our GA search based random forest and S-T correlation based match achieve better performance than some other prevalent machine leaning methods, and that a combination of those two methods outperforms most of the state-of-the-art works.
Nijun Li, Xu Cheng 0003, Zhenyang Wu
ICPR2
2014 Realistic human action recognition by Fast HOG3D and self-organization feature map
Nijun Li, Xu Cheng 0003, Suofei Zhang, Zhenyang Wu
Mach. Vis. Appl.2
2013 Recognizing human actions by BP-AdaBoost algorithm under a hierarchical recognition framework
abstract
This paper explores the performance of Neural Network (NN) for human action recognition and proposes a novel hierarchical and boosting-based action recognition system. Specifically, the main contributions of our work are three-fold: (1) A boosted NN based scheme is applied to the human action recognition task for the first time, during which we extend the standard binary AdaBoost algorithm to a multiclass version; (2) A novel hierarchical recognition framework with pre-decision and post-decision modules is proposed, which can significantly enhance the training efficiency as well as the frame-based recognition accuracy; (3) Numerous modified features (both motion and shape features) are utilized and combined in this paper. Experiments on the Weizmann dataset show promising results of our approach in comparison with other state-of-the-art methods.
Nijun Li, Xu Cheng 0003, Suofei Zhang, Zhenyang Wu
ICASSP2
2013 Adaptive object detection by implicit sub-class sharing features
Suofei Zhang, Nijun Li, Xu Cheng 0003, Zhenyang Wu
Signal Process.3
2002 FTA: A File Transfer Agent Using Java
Xu Cheng 0003
CAINE1