VLDB 2026 Research / reviewers in the wild / expert
Zhenghua Chen
dblp:03/7457
· DBLP profile ↗
97ranked-venue papers
9as first author
75since 2021 · last 2026
0000-0002-1719-0328ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 42 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 18 since 2021Computer networks · 9 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Light but Sharp: SlimSTAD for Real-Time Action Detection from Sensor DataabstractSensory Temporal Action Detection (STAD) aims to localize and classify human actions within long, untrimmed sequences captured by non-visual sensors such as WiFi or inertial measurement units (IMUs). Unlike video-based TAD, STAD poses unique challenges due to the low-dimensional, noisy, and heterogeneous nature of sensory data, as well as the real-time and resource constraints on edge devices. While recent STAD models have improved detection performance, their high computational cost hampers practical deployment. In this paper, we propose SlimSTAD, a simple yet effective framework that achieves both high accuracy and low latency for STAD. SlimSTAD features a novel Decoupled Channel Modeling (DCM) encoder, which preserves modality-specific temporal features and enables efficient inter-channel aggregation via lightweight graph attention. An anchor-free cascade predictor then refines action boundaries and class predictions in a two-stage design without dense proposals. Experiments on two real-world datasets demonstrate that SlimSTAD outperforms strong video-derived and sensory baselines by an average of 2.1 mAP, while significantly reducing GFLOPs, parameters, and latency, validating its effectiveness for real-world, edge-aware STAD deployment. Wei Cui 0002, Lukai Fan, Zhenghua Chen, Min Wu 0008, Shili Xiang, Haixia Wang 0003, Bing Li 0002 |
AAAI | 3 |
| 2026 | SPLoc: High-Precision Underwater Tunnel Robot Position Measurement via Structural-Prior ModelingabstractAbstract-High-precision localization of underwater robots in enclosed tunnel environments is challenging due to the absence of satellite signals and the practical difficulty of deploying large-scale acoustic infrastructure. Existing solutions typically fuse inertial navigation with vision, laser, or acoustic sensing, yet their performance is often limited by rapid inertial drift and degraded sensing under turbidity, specular surfaces, and constrained geometries. This paper presents SPLoc, a structural-prior-assisted localization and measurement framework, together with an underwater tunnel pose measurement platform that tightly integrates a blue–green structured-light triangulation sensor and an IMU. The structured-light subsystem acquires dense cross-sectional point clouds with centimeter-scale sampling along both axial and radial directions. From these measurements, tunnel cross-sectional primitives are extracted through feature enhancement, redundancy suppression, and RANSAC-based circle fitting with region-growing verification, and are then parameterized as structural priors. To achieve metrologically consistent fusion, we formulate a least-squares adjustment that explicitly models heading and position states, laser observation residuals, and dominant pose-error terms (including heading/position biases and observation perturbations), and solve it via iterative refinement to provide dynamic pose compensation. Simulation and physical-equivalent experiments on a purpose-built testbed demonstrate that SPLoc achieves < 2 cm localization error and < 0.5° heading error under diverse initial disturbances, while improving robustness against sensing outliers and geometric ambiguities. The proposed platform and estimation model provide a practical measurement route for accurate underwater tunnel robot localization and support engineering deployment where external beacons are unavailable. Minglei Guan, Dejin Zhang, Zhenghua Chen, Chaoyun Song, Yifeng Zeng, Qingquan Li 0001 |
IEEE Internet Things J. | 6 |
| 2026 | CiUAV: Scalable Device-Free Indoor UAV Localization via Multiobjective Optimized Network Using Channel State InformationabstractAccurate and scalable indoor localization for unmanned aerial vehicles (UAVs) is essential for Internet of Things (IoT) applications such as autonomous logistics, infrastructure inspection, and emergency response in GPS-denied environments. However, traditional methods often struggle with cost, deployment complexity, and sensitivity to environmental dynamics, limiting their practicality for large-scale IoT scenarios. This paper presents a method in which Channel State Information (CSI) from low-cost IoT sensors enables robust, device-free 3D UAV localization while optimizing accuracy, sensor adaptability, and data efficiency. We propose CiUAV, leveraging CSI captured by ESP32-S3 sensors, with a Robust CSI Signal Enhancement (RCSE) framework integrating Dynamic AGC Compensation (DAC) and Adaptive Noise Suppression and Outlier Removal (ANSOR), alongside a Sensor-in-Sample (SiS) multi-objective optimization model for adaptive multi-sensor fusion. Experimental evaluations in realistic indoor settings achieve a 3D root mean squared error (RMSE) of 0.2659 meters, outperforming baselines by up to 35% in accuracy and 50% in data efficiency. CiUAV offers a lightweight, scalable, and infrastructure-compatible solution for future IoT-enabled UAV systems. Cunyi Yin, Zhaoke Huang, Hao Jiang 0008, Jing Chen 0022, Xiren Miao, Shaocong Zheng, Jianfei Yang 0001, Zhiwen Chen 0001, Zhenghua Chen, Hong Yan 0001 |
IEEE Internet Things J. | 10 |
| 2026 | Temporal Source Recovery for Time-Series Source-Free Unsupervised Domain AdaptationabstractTime-Series (TS) data has grown in importance with the rise of Internet of Things devices like sensors, but its labeling remains costly and complex. While Unsupervised Domain Adaptation (UDAs) offers an effective solution, growing data privacy concerns have led to the development of Source-Free UDA (SFUDAs), enabling model adaptation to target domains without accessing source data. Despite their potential, applying existing SFUDAs to TS data is challenging due to the difficulty of transferring temporal dependencies-an essential characteristic of TS data-particularly in the absence of source samples. Although prior works attempt to address this by specific source pretraining designs, such requirements are often impractical, as source data owners cannot be expected to adhere to particular pretraining schemes. To address this, we propose Temporal Source Recovery (TemSR), a framework that leverages the intrinsic properties of TS data to generate a source-like domain and recover source temporal dependencies. With this domain, TemSR enables dependency transfer to the target domain without accessing source data or relying on source-specific designs, thereby facilitating effective and practical TS-SFUDA. TemSR features a masking-recovery-optimization process to generate a source-like distribution with restored temporal dependencies. This distribution is further refined through local context-aware regularization to preserve local dependencies, and anchor-based recovery diversity maximization to promote distributional diversity. Together, these components enable effective temporal dependency recovery and facilitate transfer across domains using standard UDA techniques. Extensive experiments across seven TS tasks demonstrate the effectiveness of TemSR, which even surpasses existing TS-SFUDA methods that require source-specific designs. Yucheng Wang 0001, Peiliang Gong, Min Wu 0008, Felix Ott 0001, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Geometric Contrastive Ensemble Distillation with calibration margin induction
Yang Yang 0080, Junyao Hou, Xiang Li 0067, Zhenghua Chen, Yingxue Gao, Chao Wang 0003, Min Wu 0008, Quanjun Yin |
Pattern Recognit. | 6 |
| 2026 | Target-Specific Adaptation and Consistent Degradation Alignment for Cross-Domain Remaining Useful Life PredictionabstractAccurate prediction of the Remaining Useful Life (RUL) in machinery can significantly diminish maintenance costs, enhance equipment up-time, and mitigate adverse outcomes. Data-driven RUL prediction techniques have demonstrated commendable performance. However, their efficacy often relies on the assumption that training and testing data are drawn from the same distribution or domain, which does not hold in real industrial settings. To mitigate this domain discrepancy issue, prior adversarial domain adaptation methods focused on deriving domain-invariant features. Nevertheless, they overlook target-specific information and inconsistency characteristics pertinent to the degradation stages, resulting in suboptimal performance. To tackle these issues, we propose a novel domain adaptation approach for cross-domain RUL prediction named TACDA. Specifically, we propose a target domain reconstruction strategy within the adversarial adaptation process, thereby retaining target-specific information while learning domain-invariant features. Furthermore, we develop a novel clustering and pairing strategy for consistent alignment between similar degradation stages. Through extensive experiments, our results demonstrate the remarkable performance of our proposed TACDA method, surpassing state-of-the-art approaches with regard to two different evaluation metrics. Our code is available at https://github.com/keyplay/TACDA. Yubo Hou, Mohamed Ragab 0002, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001, Zhenghua Chen |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | From Inconsistency to Unity: Benchmarking Deep Learning-Based Unsupervised Domain Adaptation for RULabstractData-driven Remaining Useful Life (RUL) estimation is critical for various industries, yet scarce labeled data poses a significant challenge. Unsupervised Domain Adaptation (UDA) combined with deep learning has emerged as a promising solution by leveraging unlabeled data. However, recent work on deep UDA exhibits notable inconsistencies, including variations in backbone networks and data handling. Such inconsistencies hinder fair comparisons, obscuring if actual advances were made. To address this, we propose CRULE, a comprehensive benchmarking framework that standardizes deep UDA evaluation. CRULE incorporates a unified 1-D CNN backbone architecture, a standardized training scheme with comprehensive hyperparameter tuning, consistent evaluation protocols (both inductive and transductive), and unified performance metrics (RMSE and Score). This ensures fair and reliable comparisons across different UDA methods. Through experiments on three popular RUL datasets, we found, that only one of the evaluated approaches achieves statistically significant improvements over no adaptation. This suggests that deep UDA approaches proposed for RUL estimation may be less reliable under fair evaluation schemes. To catalyze genuine advancements in the field, we open-source CRULE, empowering the research community to develop and consistently benchmark UDA approaches. CRULE is accessible at https://anonymous.4open.science/r/crule-55D1. Tilman Krokotsch, Mohamed Ragab 0002, Min Wu 0008, Xiaoli Li 0001, Zhenghua Chen, Clemens Gühmann |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | MarkingVLM: Vision-Language Model for Few-Shot IC Marking DetectionabstractIntegrated circuit (IC) marking detection faces critical “Triple C” challenges: compact marking scales, complex multidirectional layouts, and costly annotation requirements. Traditional models struggle with data scarcity in dynamic manufacturing environments where fine-grained labeled data are expensive and time-consuming to obtain. This work presents MarkingVLM, a specialized vision-language model designed for IC marking detection with exceptional few-shot learning capabilities. MarkingVLM adapts cross-modal understanding from Contrastive Language-Image Pre-training (CLIP) through unique innovations: 1) dual-granularity detection combining patch-level semantic alignment with pixel-level visual decoding; 2) layout mirror module enabling efficient multidirectional text flow recognition through feature-level augmentation; and 3) enhanced prompt engineering with learnable contexts for effective domain adaptation. Extensive experiments across two IC marking datasets with distinct characteristics substantiate superior performance: 94.2% precision and 96.5% recall on Dataset-1, and 89.6% precision with 92.6% recall on the more challenging Dataset-2. Most significantly, MarkingVLM achieves 92.7% precision with 32 training samples and maintains over 80% recall with merely 8 samples, demonstrating notable data efficiency improvement over conventional methods. Results establish a new paradigm for industrial text detection by bridging open-domain vision-language knowledge with specialized manufacturing requirements. Zhongshu Chen, Zhenghua Chen, Lin Zuo, Yu Liu 0006 |
IEEE Trans. Ind. Informatics | 2 |
| 2026 | A Spatio-Temporal Feature Distribution Network for Device-Free Power Inspection Activity Using WiFi CSIabstractEnsuring personnel safety during power station inspections is a critical yet challenging task due to inherent hazards in such environments. Traditional monitoring methods, including wearable devices and video surveillance, suffer from user discomfort, limited visibility, and high deployment costs. To overcome these limitations, this article proposes PowerHAR, a device-free framework for recognizing power inspection activities based on WiFi channel state information (CSI) acquired from custom-designed ESP32 internet of things (IoT) sensors. PowerHAR introduces a spatio-temporal feature distribution-based power operation recognition network, comprising a transformer-based preprocessing module capable of effectively handling variable-length CSI sequences, and a spatio-temporal extraction module that integrates convolutional operations with multihead self-attention mechanisms for comprehensive feature fusion. By leveraging mutual CSI sensing among distributed sensors, PowerHAR provides robust and accurate recognition of power inspection activities without requiring additional hardware infrastructure. Experimental validation demonstrates that PowerHAR significantly surpasses existing baseline methods, confirming its high reliability and practicality in safety-critical industrial scenarios. Cunyi Yin, Zhaoke Huang, Hao Jiang 0008, Jing Chen 0022, Zhida Wang, Zhenghua Chen, Zhiwen Chen 0001, Hong Yan 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2026 | Evidentially Calibrated Source-Free Time-Series Domain Adaptation With Temporal ImputationabstractSource-free domain adaptation (SFDA) adapts a pre-trained model from a labeled source domain to an unlabeled target domain without source data access, preserving privacy. While SFDA is common in computer vision, it remains largely unexplored in time series analysis, where existing methods struggle to capture temporal dynamics and often produce overconfident predictions on out-of-distribution samples. We propose MAsk And imPUte (MAPU), which tackles temporal consistency through a novel imputation task, where randomly masked time series signals are recovered within the learned embedding space. During adaptation, a dedicated temporal imputer guides the target model to generate features that maintain temporal consistency with source features. However, MAPU relies on standard softmax predictions, leading to overconfident predictions on target samples that fall outside the source domain's support. To address this limitation, we introduce Evidential-MAPU (E-MAPU), which leverages evidential uncertainty estimation to identify these out-of-support samples and adapts the feature extractor to map them closer to the source domain's support, while maintaining the classifier fixed. Extensive experiments on five real-world time series datasets demonstrate significant performance improvements over existing methods. Our approaches effectively handle various time series domain adaptation challenges while maintaining computational efficiency, achieving state-of-the-art performance through its uncertainty-aware adaptation strategy. Mohamed Ragab 0002, Peiliang Gong, Emadeldeen Eldele, Wenyu Zhang 0003, Min Wu 0008, Chuan-Sheng Foo, Daoqiang Zhang, Xiaoli Li 0001, Zhenghua Chen |
IEEE Trans. Knowl. Data Eng. | 9 |
| 2025 | WiFi CSI Based Temporal Activity Detection via Dual Pyramid NetworkabstractWe address the challenge of WiFi-based temporal activity detection and propose an efficient Dual Pyramid Network that integrates Temporal Signal Semantic Encoders and Local Sensitive Response Encoders. The Temporal Signal Semantic Encoder splits feature learning into high and low-frequency components, using a novel Signed Mask-Attention mechanism to emphasize important areas and downplay unimportant ones, with the features fused using ContraNorm. The Local Sensitive Response Encoder captures fluctuations without learning. These feature pyramids are then combined using a new cross-attention fusion mechanism. We also introduce a dataset with over 2,114 activity segments across 553 WiFi CSI samples, each lasting around 85 seconds. Extensive experiments show our method outperforms challenging baselines. Le Zhang 0001, Bing Li 0002, Yingjie Zhou 0001, Zhenghua Chen, Ce Zhu |
AAAI | 5 |
| 2025 | Temporal Restoration and Spatial Rewiring for Source-Free Multivariate Time Series Domain AdaptationabstractSource-Free Domain Adaptation (SFDA) aims to adapt a pre-trained model from an annotated source domain to an unlabelled target domain without accessing the source data, thereby preserving data privacy. While existing SFDA methods have proven effective in reducing reliance on source data, they struggle to perform well on multivariate time series (MTS) due to their failure to consider the intrinsic spatial correlations inherent in MTS data. These spatial correlations are crucial for accurately representing MTS data and preserving invariant information across domains. To address this challenge, we propose Temporal Restoration and Spatial Rewiring (TERSE), a novel and concise SFDA method tailored for MTS data. Specifically, TERSE comprises a customized spatial-temporal feature encoder designed to capture the underlying spatial-temporal characteristics, coupled with both temporal restoration and spatial rewiring tasks to reinstate latent representations of the temporally masked time series and the spatially masked correlated structures. During the target adaptation phase, the target encoder is guided to produce spatially and temporally consistent features with the source domain by leveraging the source pre-trained temporal restoration and spatial rewiring networks. Therefore, TERSE can effectively model and transfer spatial-temporal dependencies across domains, facilitating implicit feature alignment. In addition, as the first approach to simultaneously consider spatial-temporal consistency in MTS-SFDA, TERSE can also be integrated as a versatile plug-and-play module into established SFDA methods. Extensive experiments on three real-world time series datasets demonstrate the effectiveness and versatility of our approach. Our code is available at https://github.com/Tokenmw/TERSE-master. Peiliang Gong, Yucheng Wang 0001, Min Wu 0008, Zhenghua Chen, Xiaoli Li 0001, Daoqiang Zhang |
KDD (2) | 4 |
| 2025 | Augmented Contrastive Clustering with Uncertainty-Aware Prototyping for Time Series Test Time AdaptationabstractTest-time adaptation aims to adapt pre-trained deep neural networks using solely online unlabelled test data during inference. Although TTA has shown promise in visual applications, its potential in time series contexts remains largely unexplored. Existing TTA methods, originally designed for visual tasks, may not effectively handle the complex temporal dynamics of real-world time series data, resulting in suboptimal adaptation performance. To address this gap, we propose Augmented Contrastive Clustering with Uncertainty-aware Prototyping (ACCUP), a straightforward yet effective TTA method for time series data. Initially, our approach employs augmentation ensemble on the time series data to capture diverse temporal information and variations, incorporating uncertainty-aware prototypes to distill essential characteristics. Additionally, we introduce an entropy comparison scheme to selectively acquire more confident predictions, enhancing the reliability of pseudo labels. Furthermore, we utilize augmented contrastive clustering to enhance feature discriminability and mitigate error accumulation from noisy pseudo labels, promoting cohesive clustering within the same class while facilitating clear separation between different classes. Extensive experiments conducted on three real-world time series datasets demonstrate the effectiveness and generalization potential of the proposed method, advancing the underexplored realm of TTA for time series data. Our code is available at https://github.com/Tokenmw/ACCUP-main. Peiliang Gong, Mohamed Ragab 0002, Min Wu 0008, Zhenghua Chen, Yongyi Su, Xiaoli Li 0001, Daoqiang Zhang |
KDD (1) | 4 |
| 2025 | Bidirectional Segmentation-Aware Network for One-Shot Object Detection
Zhenghua Chen, Yongyi Su, Zhiyu Xiang, Hangguan Shan, Eryun Liu |
Neurocomputing | 2 |
| 2025 | Distributed policy evaluation over multi-agent network with communication delays
Yaoyao Zhou, Gang Chen 0014, Changli Pu, Keyu Wu 0002, Zhenghua Chen |
Neurocomputing | 5 |
| 2025 | RMKD: Relaxed matching knowledge distillation for short-length SSVEP-based brain-computer interfaces
Zhen Lan, Zixing Li, Xiaojia Xiang, Dengqing Tang, Min Wu 0008, Zhenghua Chen |
Neural Networks | 7 |
| 2025 | ACCNet: Adaptive cross-frequency coupling graph attention for EEG emotion recognition
Dongyuan Tian, Yucheng Wang 0001, Peiliang Gong, Zhewen Xu, Zhenghua Chen, Min Wu 0008 |
Neural Networks | 5 |
| 2025 | Uncertainty-Aware Self-Knowledge DistillationabstractSelf-knowledge distillation has emerged as a powerful method, notably boosting the prediction accuracy of deep neural networks while being resource-efficient, setting it apart from traditional teacher-student knowledge distillation approaches. However, in safety-critical applications, high accuracy alone is not adequate; conveying uncertainty effectively holds equal importance. Regrettably, existing self-knowledge distillation methods have not met the need to improve both prediction accuracy and uncertainty quantification simultaneously. In response to this gap, we present an uncertainty-aware self-knowledge distillation method named UASKD. UASKD introduces an uncertainty-aware contrastive loss and a prediction synthesis technique within the self-knowledge distillation process, aiming to fully harness the potential of self-knowledge distillation for improving both prediction accuracy and uncertainty quantification. Extensive assessments illustrate that UASKD consistently surpasses other self-knowledge distillation techniques and numerous uncertainty calibration methods in both prediction accuracy and uncertainty quantification metrics across various classification and object detection tasks, highlighting its efficacy and adaptability. Yang Yang 0080, Chao Wang 0003, Lei Gong 0003, Min Wu 0008, Zhenghua Chen, Yingxue Gao, Xuehai Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MTSNet: Convolution-Based Transformer Network With Multi-Scale Temporal-Spectral Feature Fusion for SSVEP Signal DecodingabstractImproving the decoding performance of steady-state visual evoked (SSVEP) signals is crucial for the practical application of SSVEP-based brain-computer interface (BCI) systems. Although numerous methods have achieved impressive results in decoding SSVEP signals, most of them focus only on the temporal or spectral domain information or concatenate them directly, which may ignore the complementary relationship between different features. To address this issue, we propose a dual-branch convolution-based Transformer network with multi-scale temporal-spectral feature fusion, termed MTSNet, to improve the decoding performance of SSVEP signals. Specifically, the temporal branch extracts temporal features from the SSVEP signals using the multi-level convolution- based Transformer (Convformer) that can adapt to the dynamic fluctuations of SSVEP signals. In parallel, the spectral branch takes the complex spectrum converted from temporal signals by the zero-padding fast Fourier transform as input and uses the Convformer to extract spectral features. These extracted temporal and spectral features are then integrated by the multi-scale feature fusion module to obtain comprehensive features with different scale information, thereby enhancing the interactions between the features and improving the effectiveness and robustness. Extensive experimental results on two widely used public SSVEP datasets, Benchmark and BETA, show that the proposed MTSNet significantly outperforms the state-of-the-art calibration-free methods in terms of accuracy and ITR. The superior performance demonstrates the effectiveness of our method in decoding SSVEP signals, which may facilitate the practical application of SSVEP-based BCI systems. Zhen Lan, Zixing Li, Xiaojia Xiang, Dengqing Tang, Min Wu 0008, Zhenghua Chen |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Multi-Scale CNN-Transformer Hybrid Network for Rail Fastener Defect DetectionabstractDefect detection in rail fasteners is crucial for train safety, as defective fasteners can cause derailments and severe safety incidents. However, Existing algorithms often struggle in various real-world scenarios due to challenges such as obscured fasteners, motion blur in images, varying camera angles, and fasteners submerged in water. To address these challenges, we propose a Multi-scale CNN-Transformer Hybrid Network for Rail Fastener Defect Detection (MCHNet-RF2D), specifically designed to identify fastener defects in complex environments. Our approach constructs an efficient CNN block and a multi-scale Vision Transformer block to alternately extract local detail features and global semantic features of the fasteners. These features are seamlessly integrated through multi-scale fusion to enhance defect recognition robustness. By combining comprehensive global recognition with detailed local defect detection, MCHNet-RF2D outperforms existing CNN-Transformer hybrid networks by 2.8% and surpasses current fastener defect detection algorithms by 2.9%. In practical deployment on over 40 trains, our model successfully detected more than 2,000 fastener defects, demonstrating its effectiveness in diverse and challenging conditions. Wei Wang 0278, Fengmao Lv, Haonan Luo 0002, Gexiang Zhang, Zhenghua Chen |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | From Algorithm to Hardware: A Survey on Efficient and Safe Deployment of Deep Neural NetworksabstractDeep neural networks (DNNs) have been widely used in many artificial intelligence (AI) tasks. However, deploying them brings significant challenges due to the huge cost of memory, energy, and computation. To address these challenges, researchers have developed various model compression techniques such as model quantization and model pruning. Recently, there has been a surge in research on compression methods to achieve model efficiency while retaining performance. Furthermore, more and more works focus on customizing the DNN hardware accelerators to better leverage the model compression techniques. In addition to efficiency, preserving security and privacy is critical for deploying DNNs. However, the vast and diverse body of related works can be overwhelming. This inspires us to conduct a comprehensive survey on recent research toward the goal of high-performance, cost-efficient, and safe deployment of DNNs. Our survey first covers the mainstream model compression techniques, such as model quantization, model pruning, knowledge distillation, and optimizations of nonlinear operations. We then introduce recent advances in designing hardware accelerators that can adapt to efficient model compression approaches. In addition, we discuss how homomorphic encryption can be integrated to secure DNN deployment. Finally, we discuss several issues, such as hardware evaluation, generalization, and integration of various compression approaches. Overall, we aim to provide a big picture of efficient DNNs from algorithm to hardware accelerators and security perspectives. Xue Geng, Zhe Wang 0019, Chunyun Chen, Qing Xu 0015, Kaixin Xu, Jin Chao, Manas Gupta, Xulei Yang, Zhenghua Chen, Mohamed M. Sabry, Jie Lin 0001, Min Wu 0008, Xiaoli Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | Local-Global Correlation Fusion-Based Graph Neural Network for Remaining Useful Life PredictionabstractRemaining useful life (RUL) prediction is an essential component for prognostics and health management of a system. Due to the powerful ability of nonlinear modeling, deep learning (DL) models have emerged as leading solutions by capturing temporal dependencies within time series sensory data. However, in RUL prediction tasks, data are typically collected from multiple sensors, introducing spatial dependencies in the form of sensor correlations. Existing methods are limited in effectively modeling and capturing the spatial dependencies, restricting their performance to learn representative features for RUL prediction. To overcome the limitations, we propose a novel LOcal-GlObal correlation fusion-based framework (LOGO). Our approach combines both local and global information to model sensor correlations effectively. From a local perspective, we account for local correlations that represent dynamic changes of sensor relationships in local ranges. Simultaneously, from a global perspective, we capture global correlations that depict relatively stable relations between sensors. An adaptive fusion mechanism is proposed to automatically fuse the correlations from different perspectives. Subsequently, we define sequential micrographs for each sample to effectively capture the fused correlations. Graph neural network (GNN) is introduced to capture the spatial dependencies within each micrograph, and the temporal dependencies between these sequential micrographs are then captured. This approach allows us to effectively model and capture the dependency information within the data for accurate RUL prediction. Extensive experiments have been conducted, verifying the effectiveness of our method. Yucheng Wang 0001, Min Wu 0008, Ruibing Jin, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Temporal and Heterogeneous Graph Neural Network for Remaining Useful Life PredictionabstractPredicting remaining useful life (RUL) plays a crucial role in the prognostics and health management of industrial systems that involve a variety of interrelated sensors. Given a constant stream of time-series sensory data from such systems, deep learning (DL) models have risen to prominence at identifying complex, nonlinear temporal dependencies in these data. In addition to the temporal dependencies of individual sensors, spatial dependencies emerge as important correlations among these sensors, which can be naturally modeled by a temporal graph that describes time-varying spatial relationships. However, the majority of existing studies have relied on capturing discrete snapshots of this temporal graph, a coarse-grained approach that leads to a loss of temporal information. Moreover, given the variety of heterogeneous sensors, it becomes vital that such inherent heterogeneity is leveraged for RUL prediction in temporal sensor graphs. To capture the nuances of the temporal and spatial relationships and heterogeneous characteristics in an interconnected graph of sensors, we introduce a novel model named temporal and heterogeneous graph neural networks (THGNNs). Specifically, THGNN aggregates historical data from neighboring nodes to accurately capture the temporal dynamics and spatial correlations within the stream of sensor data in a fine-grained manner. Moreover, the model leverages feature-wise linear modulation (FiLM) to address the diversity of sensor types, significantly improving the model's capacity to learn the heterogeneity in the data sources. Finally, we have validated the effectiveness of our approach through comprehensive experiments. Our empirical findings demonstrate significant advancements on the N-CMAPSS dataset, achieving improvements of up to 19.2% and 31.6% in terms of two different evaluation metrics over state-of-the-art methods. Zhihao Wen, Yuan Fang 0001, PengCheng Wei, Fayao Liu, Zhenghua Chen, Min Wu 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | A Virtual-Label-Based Hierarchical Domain Adaptation Method for Time-Series ClassificationabstractUnsupervised domain adaptation (UDA) is becoming a prominent solution for the domain-shift problem in many time-series classification tasks. With sequence properties, time-series data contain both local and sequential features, and the domain shift exists in both features. However, conventional UDA methods usually cannot distinguish those two features but mix them into one variable for direct alignment, which harms the performance. To address this problem, we propose a novel virtual-label-based hierarchical domain adaptation (VLH-DA) approach for time-series classification. Specifically, we first slice the original time-series data and introduce virtual labels to represent the type of each slice (called local patterns). With the help of virtual labels, we decompose the end-to-end (i.e., signal to time-series label) time-series task into two parts, i.e., signal sequence to local pattern sequence and local pattern sequence to time-series label. By decomposing the complex time-series UDA task into two simpler subtasks, the local features and sequential features can be aligned separately, making it easier to mitigate distribution discrepancies. Experiments on four public time-series datasets demonstrate that our VLH-DA outperforms all state-of-the-art (SOTA) methods. Wenmian Yang, Lizhi Cheng, Mohamed Ragab 0002, Min Wu 0008, Sinno Jialin Pan, Zhenghua Chen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Fully-Connected Spatial-Temporal Graph for Multivariate Time-Series DataabstractMultivariate Time-Series (MTS) data is crucial in various application fields. With its sequential and multi-source (multiple sensors) properties, MTS data inherently exhibits Spatial-Temporal (ST) dependencies, involving temporal correlations between timestamps and spatial correlations between sensors in each timestamp. To effectively leverage this information, Graph Neural Network-based methods (GNNs) have been widely adopted. However, existing approaches separately capture spatial dependency and temporal dependency and fail to capture the correlations between Different sEnsors at Different Timestamps (DEDT). Overlooking such correlations hinders the comprehensive modelling of ST dependencies within MTS data, thus restricting existing GNNs from learning effective representations. To address this limitation, we propose a novel method called Fully-Connected Spatial-Temporal Graph Neural Network (FC-STGNN), including two key components namely FC graph construction and FC graph convolution. For graph construction, we design a decay graph to connect sensors across all timestamps based on their temporal distances, enabling us to fully model the ST dependencies by considering the correlations between DEDT. Further, we devise FC graph convolution with a moving-pooling GNN layer to effectively capture the ST dependencies for learning effective representations. Extensive experiments show the effectiveness of FC-STGNN on multiple MTS datasets compared to SOTA methods. The code is available at https://github.com/Frank-Wang-oss/FCSTGNN. Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen |
AAAI | 7 |
| 2024 | Graph-Aware Contrasting for Multivariate Time-Series ClassificationabstractContrastive learning, as a self-supervised learning paradigm, becomes popular for Multivariate Time-Series (MTS) classification. It ensures the consistency across different views of unlabeled samples and then learns effective representations for these samples. Existing contrastive learning methods mainly focus on achieving temporal consistency with temporal augmentation and contrasting techniques, aiming to preserve temporal patterns against perturbations for MTS data. However, they overlook spatial consistency that requires the stability of individual sensors and their correlations. As MTS data typically originate from multiple sensors, ensuring spatial consistency becomes essential for the overall performance of contrastive learning on MTS data. Thus, we propose Graph-Aware Contrasting for spatial consistency across MTS data. Specifically, we propose graph augmentations including node and edge augmentations to preserve the stability of sensors and their correlations, followed by graph contrasting with both node- and graph-level contrasting to extract robust sensor- and global-level features. We further introduce multi-window temporal contrasting to ensure temporal consistency in the data for each sensor. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on various MTS classification tasks. The code is available at https://github.com/Frank-Wang-oss/TS-GAC. Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen |
AAAI | 7 |
| 2024 | TSLANet: Rethinking Transformers for Time Series Representation LearningabstractTime series data, characterized by its intrinsic long and short-range dependencies, poses a unique challenge across analytical applications. While Transformer-based models excel at capturing long-range dependencies, they face limitations in noise sensitivity, computational efficiency, and overfitting with smaller datasets. In response, we introduce a novel **T**ime **S**eries **L**ightweight **A**daptive **Net**work (**TSLANet**), as a universal convolutional model for diverse time series tasks. Specifically, we propose an Adaptive Spectral Block, harnessing Fourier analysis to enhance feature representation and to capture both long-term and short-term interactions while mitigating noise via adaptive thresholding. Additionally, we introduce an Interactive Convolution Block and leverage self-supervised learning to refine the capacity of TSLANet for decoding complex temporal patterns and improve its robustness on different datasets. Our comprehensive experiments demonstrate that TSLANet outperforms state-of-the-art models in various tasks spanning classification, forecasting, and anomaly detection, showcasing its resilience and adaptability across a spectrum of noise levels and data sizes. The code is available at https://github.com/emadeldeen24/TSLANet. Emadeldeen Eldele, Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Xiaoli Li 0001 |
ICML | 3 |
| 2024 | Reinforced Cross-Domain Knowledge Distillation on Time Series DataabstractUnsupervised domain adaptation methods have demonstrated superior capabilities in handling the domain shift issue which widely exists in various time series tasks. However, their prominent adaptation performances heavily rely on complex model architectures, posing an unprecedented challenge in deploying them on resource-limited devices for real-time monitoring. Existing approaches, which integrates knowledge distillation into domain adaptation frameworks to simultaneously address domain shift and model complexity, often neglect network capacity gap between teacher and student and just coarsely align their outputs over all source and target samples, resulting in poor distillation efficiency. Thus, in this paper, we propose an innovative framework named Reinforced Cross-Domain Knowledge Distillation (RCD-KD) which can effectively adapt to student's network capability via dynamically selecting suitable target domain samples for knowledge transferring. Particularly, a reinforcement learning-based module with a novel reward function is proposed to learn optimal target sample selection policy based on student's capacity. Meanwhile, a domain discriminator is designed to transfer the domain invariant knowledge. Empirical experimental results and analyses on four public time series datasets demonstrate the effectiveness of our proposed method over other state-of-the-art benchmarks. Qing Xu 0015, Min Wu 0008, Xiaoli Li 0001, Kezhi Mao, Zhenghua Chen |
NeurIPS | 5 |
| 2024 | Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot LearningabstractMeta-learning offers a promising avenue for few-shot learning (FSL), enabling models to glean a generalizable feature embedding through episodic training on synthetic FSL tasks in a source domain. Yet, in practical scenarios where the target task diverges from that in the source domain, meta-learning based method is susceptible to over-fitting. To overcome this, we introduce a novel framework, Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning, which is crafted to comprehensively exploit the cross-domain transferable image prior that each image can be decomposed into complementary low-frequency content details and high-frequency robust structural characteristics. Motivated by this insight, we propose to decompose each query image into its high-frequency and low-frequency components, and parallel incorporate them into the feature embedding network to enhance the final category prediction. More importantly, we introduce a feature reconstruction prior and a prediction consistency prior to separately encourage the consistency of the intermediate feature as well as the final category prediction between the original query image and its decomposed frequency components. This allows for collectively guiding the network's meta-learning process with the aim of learning generalizable image feature embeddings, while not introducing any extra computational cost in the inference phase. Our framework establishes new state-of-the-art results on multiple cross-domain few-shot learning benchmarks. Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Zhenghua Chen, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001 |
NeurIPS | 4 |
| 2024 | Going Deeper into Recognizing Actions in Dark Environments: A Comprehensive Benchmark Study
Yuecong Xu, Haozhi Cao, Jianxiong Yin, Zhenghua Chen, Xiaoli Li 0001, Zhengguo Li, Qianwen Xu 0001, Jianfei Yang 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | PowerSkel: A Device-Free Framework Using CSI Signal for Human Skeleton Estimation in Power StationabstractSafety monitoring of power operations in power stations is crucial for preventing accidents and ensuring stable power supply. However, conventional methods such as wearable devices and video surveillance have limitations such as high cost, dependence on light, and visual blind spots. WiFi-based human pose estimation is a suitable method for monitoring power operations due to its low cost, device-free, and robustness to various illumination conditions. In this paper, a novel Channel State Information (CSI)-based pose estimation framework, namely PowerSkel, is developed to address these challenges. PowerSkel utilizes self-developed CSI sensors to form a mutual sensing network and constructs a CSI acquisition scheme specialized for power scenarios. It significantly reduces the deployment cost and complexity compared to the existing solutions. To reduce interference with CSI in the electricity scenario, a sparse adaptive filtering algorithm is designed to preprocess the CSI. CKDformer, a knowledge distillation network based on collaborative learning and self-attention, is proposed to extract the features from CSI and establish the mapping relationship between CSI and keypoints. The experiments are conducted in a real-world power station, and the results show that the PowerSkel achieves high performance with a PCK@50 of 96.27%, and realizes a significant visualization on pose estimation, even in dark environments. Our work provides a novel low-cost and high-precision pose estimation solution for power operation. Cunyi Yin, Xiren Miao, Jing Chen 0022, Hao Jiang 0008, Jianfei Yang 0001, Yunjiao Zhou, Min Wu 0008, Zhenghua Chen |
IEEE Internet Things J. | 8 |
| 2024 | SEA++: Multi-Graph-Based Higher-Order Sensor Alignment for Multivariate Time-Series Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) methods have been successful in reducing label dependency by minimizing the domain discrepancy between labeled source domains and unlabeled target domains. However, these methods face challenges when dealing with Multivariate Time-Series (MTS) data. MTS data typically originates from multiple sensors, each with its unique distribution. This property poses difficulties in adapting existing UDA techniques, which mainly focus on aligning global features while overlooking the distribution discrepancies at the sensor level, thus limiting their effectiveness for MTS data. To address this issue, a practical domain adaptation scenario is formulated as Multivariate Time-Series Unsupervised Domain Adaptation (MTS-UDA). In this paper, we propose SEnsor Alignment (SEA) for MTS-UDA, aiming to address domain discrepancy at both local and global sensor levels. At the local sensor level, we design endo-feature alignment, which aligns sensor features and their correlations across domains. To reduce domain discrepancy at the global sensor level, we design exo-feature alignment that enforces restrictions on global sensor features. We further extend SEA to SEA++ by enhancing the endo-feature alignment. Particularly, we incorporate multi-graph-based higher-order alignment for both sensor features and their correlations. Extensive empirical results have demonstrated the state-of-the-art performance of our SEA and SEA++ on six public MTS datasets for MTS-UDA. Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001, Zhenghua Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | FedCov: Enhanced Trustworthy Federated Learning for Machine RUL Prediction With Continuous-to-Discrete ConversionabstractNumerous approaches have been proposed for predicting machine remaining useful life (RUL), which helps prevent unnecessary downtime and reduces the maintenance cost in industrial systems. Most existing RUL methods rely on centralized learning and require large-scale datasets with manual labels, which are infeasible to collect. As a decentralized learning paradigm, federated learning (FL) has recently been integrated into these approaches, which aims to utilize the distributed data from local users for model training, while preserving their data privacy. However, data heterogeneity in the industry poses a critical challenge for FL, leading to model drifting issue and degraded global model performance. A straightforward method to tackle this problem is to estimate the data distribution of clients. However, it is difficult to apply this method to a regression task, since the prediction space is continuous, which increases the difficulty in estimating the data distribution. Motivated by the digital-to-analog converter in electronics, we propose a novel approach called FedCov, which involves a converter module that transforms continuous RUL values into discrete categories. Subsequently, a generator is trained to aggregate user information based on the discrete label distribution, and it is broadcasted to users as a data enhancement tool to address the data heterogeneity problem. Furthermore, to improve the performance and reliability of our FedCov, an uncertainty estimation module is proposed, which utilizes the confidence level of model predictions to adjust the training direction. Extensive experiments are conducted on theC-MAPSSbenchmark, which demonstrates that our proposed FedCov effectively solves the model drift issues and improves the performances on the RUL task, achieving state-of-the-arts performances. Yuming Fang 0001, Weide Liu, Ruibing Jin, Jun Cheng 0003, Zhenghua Chen |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | An Adaptive and Dynamical Neural Network for Machine Remaining Useful Life PredictionabstractRecently, many neural networks have been proposed for machine remaining useful life (RUL) prediction. However, most network architectures of the existing approaches are fixed. Since the sequential information depends on the input data and distributes differently, these fixed networks that cannot be dynamically adjusted according to the input data may not be able to capture this sequential information well, resulting in suboptimal performances. To mitigate this issue, we propose an adaptive and dynamical neural network (AdaNet), which can dynamically adjust its architecture according to the input data. A neural network is generally determined by kernel size, depth, and channel size. In this article, we aim to enable our proposed AdaNet to adjust its kernel size and channel size dynamically. First, we explore to adapt the deformable convolution to time-series data, which allows the convolutional kernel to change according to the feature map. With this deformable convolution, the convolutional kernels in the AdaNet become adjustable, which is beneficial to fully exploit the sequential information in time-series data, leading to accurate RUL prediction. In addition, a channel selection module is devised, which can selectively activate the feature channel according to the input, further improving the performance of our AdaNet. Extensive experiments have been carried out on the C-MAPSS dataset, demonstrating that our proposed AdaNet achieves state-of-the-art performances. Ruibing Jin, Duo Zhou, Min Wu 0008, Xiaoli Li 0001, Zhenghua Chen |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Tower Masking MIM: A Self-Supervised Pretraining Method for Power Line InspectionabstractFor intelligent inspection of power lines, a core task is to detect components in aerial images. Currently, deep supervised learning, a data-hungry paradigm, has attracted great attention. However, considering real-world scenarios, labeled data are usually limited, and the utilization of abundant unlabeled data is rarely investigated in this field. This study deploys a pretrained model for power line component detection based on a self-supervised pretraining approach, which exploits useful information from unannotated data. Concretely, we design a new masking strategy based on the structural characteristic of power lines to guide the pretraining process with meaningful semantic content. Meanwhile, a Siamese architecture is proposed to extract complete global features by using dual reconstruction with semantic targets provided by the proposed masking strategy. Then, the knowledge distillation is utilized to enable the pretrained model to learn both domain-specific and general representations. Moreover, a feature pyramid mechanism is adopted to capture multiscale features, which can benefit the detection task. Experimental results show that the proposed approach can successfully improve the performance of a variety of detection frameworks for power line components, and outperforms other self-supervised pretraining methods. Xinyu Liu 0006, Xiren Miao, Hao Jiang 0008, Jing Chen 0022, Min Wu 0008, Zhenghua Chen |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Curriculum-Based Federated Learning for Machine Fault Diagnosis With Noisy LabelsabstractFederated learning (FL) has emerged as an effective machine-learning paradigm for collaborative machine fault diagnosis in a privacy-preserving scheme. However, due to the perception limitation and different annotation criteria of annotators, the data in clients may have noisy labels with varied noise levels, leading to degraded FL performances. Most existing methods in FL for tackling the label noise issue, assume that there is label noise in all clients and treat all clients with the same denoising training. However, these methods may result in sub-optimization and even training instability of local models, so that they cannot perform well on heterogeneous label noise across clients in FL. To address this issue, we propose a curriculum-based federated learning (called FedCNL) method to combat the heterogeneous label noise in FL settings. First, our proposed FedCNL exploits a noise modeling module to adaptively estimate the clean clients and noisy clients, and identify the clean samples and noisy samples in noisy clients in an unsupervised manner. Then, a multi-stage curriculum learning is designed by regarding the noise level as learning complexity, where the model learns from clean to noisy samples, gradually improving the performance of the global model. Moreover, a mixed loss correction method is explored in the curriculum stage to maximize the utilization of data with noisy labels. Experiments performed on fault datasets in non-identically and independently distributed settings indicate that our proposed method addresses the label noise issue for machine fault diagnosis in heterogeneous FL with favorable effectiveness, achieving state-of-the-art performances. Ruqiang Yan 0001, Ruibing Jin, Rui Zhao 0004, Zhenghua Chen |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Graph Convolutional Network With Connectivity Uncertainty for EEG-Based Emotion RecognitionabstractAutomatic emotion recognition based on multichannel Electroencephalography (EEG) holds great potential in advancing human-computer interaction. However, several significant challenges persist in existing research on algorithmic emotion recognition. These challenges include the need for a robust model to effectively learn discriminative node attributes over long paths, the exploration of ambiguous topological information in EEG channels and effective frequency bands, and the mapping between intrinsic data qualities and provided labels. To address these challenges, this study introduces the distribution-based uncertainty method to represent spatial dependencies and temporal-spectral relativeness in EEG signals based on Graph Convolutional Network (GCN) architecture that adaptively assigns weights to functional aggregate node features, enabling effective long-path capturing while mitigating over-smoothing phenomena. Moreover, the graph mixup technique is employed to enhance latent connected edges and mitigate noisy label issues. Furthermore, we integrate the uncertainty learning method with deep GCN weights in a one-way learning fashion, termed Connectivity Uncertainty GCN (CU-GCN). We evaluate our approach on two widely used datasets, namely SEED and SEEDIV, for emotion recognition tasks. The experimental results demonstrate the superiority of our methodology over previous methods, yielding positive and significant improvements. Ablation studies confirm the substantial contributions of each component to the overall performance. Hongxiang Gao, Xingyao Wang 0001, Zhenghua Chen, Min Wu 0008, Zhipeng Cai 0002, Jianqing Li 0002, Chengyu Liu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | WiSR: Wireless Domain Generalization Based on Style RandomizationabstractCurrent wireless cross-domain solutions are limited to cross-one-factor tasks, requiring target domain data participation for training or position-independent feature extraction using multiple transceivers. Therefore, this paper aims to demonstrate cross-domain wireless sensing in a more challenging domain generalization (DG) setting without multiple transceivers or target domain data. Specifically, we propose a style-randomized cross-domain wireless sensing model called WiSR, which extracts domain-invariant features from multiple source domains. It quantifies Channel State Information (CSI) differences in the subcarrier dimensions as subcarrier-domain styles and instructs the feature extractor to gradually bias the gesture signals by randomizing the subcarrier-domain styles at the feature level. Meanwhile, a domain classifier that shares the same feature extractor is instructed to gradually bias the domain signals by randomizing the gesture features. Then, the adversarial training framework enables the domain classifier to reduce the influence of domain signals on the feature extractor. Extensive experiments have been performed on three gesture datasets with varying amounts of subcarriers from devices with different NICs, including cross-one-factor (such as room, user, location, and orientation) and cross-multi-factor sensing tasks. The results demonstrate that our method considerably increases performance on wireless DG tasks. Our code is available at:https://github.com/LiuSjia/WiSR. Shijia Liu, Zhenghua Chen, Min Wu 0008, Chang Liu 0087, Liangyin Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Generalizing Wireless Cross-Multiple-Factor Gesture Recognition to Unseen DomainsabstractCross-domain wireless sensing has always been challenging due to the sensitivity of wireless signals to various environmental factors, which we refer to as subdomains. However, current efforts are limited to cross-one-subdomain tasks requiring target domain data for model training or multiple receivers for data collection. Taking common gesture recognition as an application example, we attempt to demonstrate the feasibility of cross-multiple-subdomain wireless sensing in the more challenging domain generalization (DG) setting. It is possible to extract domain-invariant features from one or several source domain(s), thereby avoiding the need for multiple receivers or target domain data. We also propose an intelligent wireless data augmentation technique based on subdomain-guided perturbations, named WiSGP. Specifically, the independent domain model generates perturbations in the direction of the largest subdomain variations. Then, these subdomain-guided perturbations augment the gesture model's input to enable better domain-invariant feature extraction, even when various subdomains interact. Similarly, gesture-guided perturbations augment the domain model's input, resulting in more accurate subdomain-guided perturbations and minimal gesture label changes. Extensive experiments have been conducted on three datasets collected from various NICs. In terms of room, location, orientation, and user subdomains, WiSGP exhibits excellent accuracy, generalizability, and portability for both cross-one-subdomain and cross-multiple-subdomain tasks. Shijia Liu, Zhenghua Chen, Min Wu 0008, Hao Wang 0034, Liangyin Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Bayesian Uncertainty Calibration for Federated Time Series AnalysisabstractDeep learning models for time series analysis often require large-scale labeled datasets for training. However, acquiring such datasets is cost-intensive and challenging, particularly for individual institutions. To overcome this challenge and concern about data confidentiality among different institutions, federated learning (FL) servers as a viable solution to this dilemma by offering a decentralized learning framework. However, the datasets collected by each institution often suffer from imbalance and may not adhere to uniform protocols, leading to diverse data distributions. To address this problem, we design a global model to approximate the global data distribution of all participant clients, then transfer it to local clients as an induction in the training phase. While discrepancies between the approximate distribution and the actual distribution result in uncertainty in the predicted results. Moreover, the diverse data distributions among various clients within the FL framework, combined with the inherent lack of reliability and interpretability in deep learning models, further amplify the uncertainty of the prediction results. To address these issues, we propose an uncertainty calibration method based on Bayesian deep learning techniques, which captures uncertainty by learning a fidelity transformation to reconstruct the output of time series regression and classification tasks, utilizing deterministic pre-trained models. Extensive experiments on the regression dataset (C-MAPSS) and classification datasets (ESR, Sleep-EDF, HAR, and FD) in the Independent and Identically Distributed (IID) and non-IID settings show that our approach effectively calibrates uncertainty within the FL framework and facilitates better generalization performance in both the regression and classification tasks, achieving state-of-the-art performance. Weide Liu, Xue Xia 0005, Zhenghua Chen, Yuming Fang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Self-Supervised Autoregressive Domain Adaptation for Time Series DataabstractUnsupervised domain adaptation (UDA) has successfully addressed the domain shift problem for visual applications. Yet, these approaches may have limited performance for time series data due to the following reasons. First, they mainly rely on the large-scale dataset (i.e., ImageNet) for source pretraining, which is not applicable for time series data. Second, they ignore the temporal dimension on the feature space of the source and target domains during the domain alignment step. Finally, most of the prior UDA methods can only align the global features without considering the fine-grained class distribution of the target domain. To address these limitations, we propose a SeLf-supervised AutoRegressive Domain Adaptation (SLARDA) framework. In particular, we first design a self-supervised (SL) learning module that uses forecasting as an auxiliary task to improve the transferability of source features. Second, we propose a novel autoregressive domain adaptation technique that incorporates temporal dependence of both source and target features during domain alignment. Finally, we develop an ensemble teacher model to align class-wise distribution in the target domain via a confident pseudo labeling approach. Extensive experiments have been conducted on three real-world time series applications with 30 cross-domain scenarios. The results demonstrate that our proposed SLARDA method significantly outperforms the state-of-the-art approaches for time series domain adaptation. Our source code is available at: https://github.com/mohamedr002/SLARDA. Mohamed Ragab 0002, Emadeldeen Eldele, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Aligning Correlation Information for Domain Adaptation in Action RecognitionabstractDomain adaptation (DA) approaches address domain shift and enable networks to be applied to different scenarios. Although various image DA approaches have been proposed in recent years, there is limited research toward video DA. This is partly due to the complexity in adapting the different modalities of features in videos, which includes the correlation features extracted as long-range dependencies of pixels across spatiotemporal dimensions. The correlation features are highly associated with action classes and proven their effectiveness in accurate video feature extraction through the supervised action recognition task. Yet correlation features of the same action would differ across domains due to domain shift. Therefore, we propose a novel adversarial correlation adaptation network (ACAN) to align action videos by aligning pixel correlations. ACAN aims to minimize the distribution of correlation information, termed as pixel correlation discrepancy (PCD). Additionally, video DA research is also limited by the lack of cross-domain video datasets with larger domain shifts. We, therefore, introduce a novel HMDB-ARID dataset with a larger domain shift caused by a larger statistical difference between domains. This dataset is built in an effort to leverage current datasets for dark video classification. Empirical results demonstrate the state-of-the-art performance of our proposed ACAN for both existing and the new video DA datasets. Yuecong Xu, Haozhi Cao, Kezhi Mao, Zhenghua Chen, Lihua Xie 0001, Jianfei Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | LiteFormer: A Lightweight and Efficient Transformer for Rotating Machine Fault DiagnosisabstractTransformer has shown impressive performance on global feature modeling in many applications. However, two drawbacks induced by its intrinsic architecture limit its application, especially in fault diagnosis. First, the quadratic complexity of its self-attention scheme extremely increases the computation cost, which poses a challenge to apply Transformer to a computationally limited platform like an industry system. In addition, the sequence-based modeling in the Transformer increases the training difficulty and requires a large-scale training dataset. This drawback becomes serious when Transformer is applied in fault diagnosis where only limited data is available. To mitigate these issues, we rethink this common approach and propose a new Transformer, which is more suitable for fault diagnosis. In this article, we first show that the attention module can be actually replaced with or even surpassed by a convolution layer under some conditions in mathematics and experiments. Then, we adopt the convolutions into the Transformer, where the computation burden issue is alleviated and the fault classification accuracy is significantly improved. Furthermore, to increase the computation efficiency, a lightweight Transformer called LiteFormer, is developed by utilizing the depth-wise convolutional layer. Extensive experiments are carried out on four datasets: Case Western Reserve University dataset; Paderborn University dataset; and two gearbox datasets of drivetrain dynamic simulator. Through our experiments, our LiteFormer not only reduces the computation cost in model training, but also sets new state-of-the-art results, surpassing other counterparts in both fault classification accuracy and model robustness. Ruqiang Yan 0001, Ruibing Jin, Jiawen Xu 0002, Yuan Yang 0005, Zhenghua Chen |
IEEE Trans. Reliab. | 6 |
| 2024 | Lighter Sequential Recommendation Algorithm With Time Interval Awareness AugmentationabstractSequential recommendation models analyze users’ historical interactions to predict the next item they will en gage with. In order to better capture users’ dynamic interest preferences, most existing sequential recommendation models that introduce heterogeneous time intervals lead to increased model complexity, which raises computational costs and training difficulty. This is particularly evident in long sequential data, where the model need to handle a large variety of different time intervals. Additionally, accurately modeling the impact of long time intervals on user behavior remains a significant challenge. To address these issues, we propose a lightweight sequential recommendation algorithm with time interval awareness augmen tation (TALSAN). This model introduces a novel uniform data augmentation operator to improve the distribution of original data samples and employs a time-aware self-attention layer to model user interactions, maintaining the continuity of the original sequence. By integrating temporal context with posi tional features, TALSAN constructs a streamlined self-attention network for predicting user behavior. Comparative testing on datasets such as ML-100K, ML-1M, Amazon Beauty, Amazon Toys, and Amazon Fashion demonstrates the model’s superiority over existing baselines. Our results confirm that TALSAN not only mitigates cold start issues but also enhances the ability to learn user preferences, leading to improved prediction accuracy. Xiaoyao Zheng, Shengfei Jiang, Zhenghua Chen, Qingying Yu, Liangmin Guo, Yonglong Luo |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | SEnsor Alignment for Multivariate Time-Series Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) methods can reduce label dependency by mitigating the feature discrepancy between labeled samples in a source domain and unlabeled samples in a similar yet shifted target domain. Though achieving good performance, these methods are inapplicable for Multivariate Time-Series (MTS) data. MTS data are collected from multiple sensors, each of which follows various distributions. However, most UDA methods solely focus on aligning global features but cannot consider the distinct distributions of each sensor. To cope with such concerns, a practical domain adaptation scenario is formulated as Multivariate Time-Series Unsupervised Domain Adaptation (MTS-UDA). In this paper, we propose SEnsor Alignment (SEA) for MTS-UDA to reduce the domain discrepancy at both the local and global sensor levels. At the local sensor level, we design the endo-feature alignment to align sensor features and their correlations across domains, whose information represents the features of each sensor and the interactions between sensors. Further, to reduce domain discrepancy at the global sensor level, we design the exo-feature alignment to enforce restrictions on the global sensor features. Meanwhile, MTS also incorporates the essential spatial-temporal dependencies information between sensors, which cannot be transferred by existing UDA methods. Therefore, we model the spatial-temporal information of MTS with a multi-branch self-attention mechanism for simple and effective transfer across domains. Empirical results demonstrate the state-of-the-art performance of our proposed SEA on two public MTS datasets for MTS-UDA. The code is available at https://github.com/Frank-Wang-oss/SEA Yucheng Wang 0001, Yuecong Xu, Jianfei Yang 0001, Zhenghua Chen, Min Wu 0008, Xiaoli Li 0001, Lihua Xie 0001 |
AAAI | 4 |
| 2023 | Augmenting and Aligning Snippets for Few-Shot Video Domain AdaptationabstractFor video models to be transferred and applied seamlessly across video tasks in varied environments, Video Unsupervised Domain Adaptation (VUDA) has been introduced to improve the robustness and transferability of video models. However, current VUDA methods rely on a vast amount of high-quality unlabeled target data, which may not be available in real-world cases. We thus consider a more realistic Few-Shot Video-based Domain Adaptation (FSVDA) scenario where we adapt video models with only a few target video samples. While a few methods have touched upon Few-Shot Domain Adaptation (FSDA) in images and in FSVDA, they rely primarily on spatial augmentation for target domain expansion with alignment performed statistically at the instance level. However, videos contain more knowledge in terms of rich temporal and semantic information, which should be fully considered while augmenting target domains and performing alignment in FSVDA. We propose a novel SSA2lign to address FSVDA at the snippet level, where the target domain is expanded through a simple snippet-level augmentation followed by the attentive alignment of snippets both semantically and statistically, where semantic alignment of snippets is conducted through multiple perspectives. Empirical results demonstrate state-of-the-art performance of SSA2lign across multiple cross-domain action recognition benchmarks. Code will be provided at: https://github.com/xuyu0010/SSA2lign. Yuecong Xu, Jianfei Yang 0001, Yunjiao Zhou, Zhenghua Chen, Min Wu 0008, Xiaoli Li 0001 |
ICCV | 4 |
| 2023 | Distilling Universal and Joint Knowledge for Cross-Domain Model Compression on Time Series DataabstractFor many real-world time series tasks, the computational complexity of prevalent deep leaning models often hinders the deployment on resource limited environments (e.g., smartphones). Moreover, due to the inevitable domain shift between model training (source) and deploying (target) stages, compressing those deep models under cross-domain scenarios becomes more challenging. Although some of existing works have already explored cross-domain knowledge distillation for model compression, they are either biased to source data or heavily tangled between source and target data. To this end, we design a novel end-to-end framework called UNiversal and joInt Knowledge Distillation (UNI-KD) for cross-domain model compression. In particular, we propose to transfer both the universal feature-level knowledge across source and target domains and the joint logit-level knowledge shared by both domains from the teacher to the student model via an adversarial learning scheme. More specifically, a feature-domain discriminator is employed to align teacher’s and student’s representations for universal knowledge transfer. A data-domain discriminator is utilized to prioritize the domain-shared samples for joint knowledge transfer. Extensive experimental results on four time series datasets demonstrate the superiority of our proposed method over state-of-the-art (SOTA) benchmarks. The source code is available at https://github.com/ijcai2023/UNI KD. Qing Xu 0015, Min Wu 0008, Xiaoli Li 0001, Kezhi Mao, Zhenghua Chen |
IJCAI | 5 |
| 2023 | Source-Free Domain Adaptation with Temporal Imputation for Time Series DataabstractSource-free domain adaptation (SFDA) aims to adapt a pretrained model from a labeled source domain to an unlabeled target domain without access to the source domain data, preserving source domain privacy. Despite its prevalence in visual applications, SFDA is largely unexplored in time series applications. The existing SFDA methods that are mainly designed for visual applications may fail to handle the temporal dynamics in time series, leading to impaired adaptation performance. To address this challenge, this paper presents a simple yet effective approach for source-free domain adaptation on time series data, namely MAsk and imPUte (MAPU). First, to capture temporal information of the source domain, our method performs random masking on the time series signals while leveraging a novel temporal imputer to recover the original signal from a masked version in the embedding space. Second, in the adaptation step, the imputer network is leveraged to guide the target model to produce target features that are temporally consistent with the source features. To this end, our MAPU can explicitly account for temporal dependency during the adaptation while avoiding the imputation in the noisy input space. Our method is the first to handle temporal consistency in SFDA for time series data and can be seamlessly equipped with other existing SFDA methods. Extensive experiments conducted on three real-world time series datasets demonstrate that our MAPU achieves significant performance gain over existing methods. Our code is available at: https://github.com/mohamedr002/MAPU_SFDA_TS. Mohamed Ragab 0002, Emadeldeen Eldele, Min Wu 0008, Chuan-Sheng Foo, Xiaoli Li 0001, Zhenghua Chen |
KDD | 6 |
| 2023 | DDUC: an erasure-coded system with decoupled data updating and codingabstractIn distributed storage systems, replication and erasure code (EC) are common methods for data redundancy. Compared with replication, EC has better storage efficiency, but suffers higher overhead in update. Moreover, consistency and reliability problems caused by concurrent updates bring new challenges to applications of EC. Many works focus on optimizing the EC solution, including algorithm optimization, novel data update method, and so on, but lack the solutions for consistency and reliability problems. In this paper, we introduce a storage system that decouples data updating and EC encoding, namely, decoupled data updating and coding (DDUC), and propose a data placement policy that combines replication and parity blocks. For the ( N, M ) EC system, the data are placed as N groups of M +1 replicas, and redundant data blocks of the same stripe are placed in the parity nodes, so that the parity nodes can autonomously perform local EC encoding. Based on the above policy, a two-phase data update method is implemented in which data are updated in replica mode in phase 1, and the EC encoding is done independently by parity nodes in phase 2. This solves the problem of data reliability degradation caused by concurrent updates while ensuring high concurrency performance. It also uses persistent memory (PMem) hardware features of the byte addressing and eight-byte atomic write to implement a lightweight logging mechanism that improves performance while ensuring data consistency. Experimental results show that the concurrent access performance of the proposed storage system is 1.70–3.73 times that of the state-of-the-art storage system Ceph, and the latency is only 3.4%–5.9% that of Ceph. Yaofeng Tu, Yinjun Han, Zhenghua Chen, Xuecheng Qi, Xinyuan Sun |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2023 | SSA-ICL: Multi-domain adaptive attention with intra-dataset continual learning for Facial expression recognition
Hongxiang Gao, Min Wu 0008, Zhenghua Chen, Yuwen Li 0002, Xingyao Wang 0001, Shan An, Jianqing Li 0002, Chengyu Liu 0001 |
Neural Networks | 3 |
| 2023 | Self-Supervised Contrastive Representation Learning for Semi-Supervised Time-Series ClassificationabstractLearning time-series representations when only unlabeled data or few labeled samples are available can be a challenging task. Recently, contrastive self-supervised learning has shown great improvement in extracting useful representations from unlabeled data via contrasting different augmented views of data. In this work, we propose a novel Time-Series representation learning framework via Temporal and Contextual Contrasting (TS-TCC) that learns representations from unlabeled data with contrastive learning. Specifically, we propose time-series-specific weak and strong augmentations and use their views to learn robust temporal relations in the proposed temporal contrasting module, besides learning discriminative representations by our proposed contextual contrasting module. Additionally, we conduct a systematic study of time-series data augmentation selection, which is a key part of contrastive learning. We also extend TS-TCC to the semi-supervised learning settings and propose a Class-Aware TS-TCC (CA-TCC) that benefits from the available few labeled data to further improve representations learned by TS-TCC. Specifically, we leverage the robust pseudo labels produced by TS-TCC to realize a class-aware contrastive loss. Extensive experiments show that the linear evaluation of the features learned by our proposed framework performs comparably with the fully supervised training. Additionally, our framework shows high efficiency in few labeled data and transfer learning scenarios. Emadeldeen Eldele, Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001, Cuntai Guan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Reinforced Adaptation Network for Partial Domain AdaptationabstractDomain adaptation enables generalized learning in new environments by transferring knowledge from label-rich source domains to label-scarce target domains. As a more realistic extension, partial domain adaptation (PDA) relaxes the assumption of fully shared label space, and instead deals with the scenario where the target label space is a subset of the source label space. In this paper, we propose a Reinforced Adaptation Network (RAN) to address the challenging PDA problem. Specifically, a deep reinforcement learning model is proposed to learn source data selection policies. Meanwhile, a domain adaptation model is presented to simultaneously determine rewards and learn domain-invariant feature representations. By combining reinforcement learning and domain adaptation techniques, the proposed network alleviates negative transfer by automatically filtering out less relevant source data and promotes positive transfer by minimizing the distribution discrepancy across domains. Experiments on three benchmark datasets demonstrate that RAN consistently outperforms seventeen existing state-of-the-art methods by a large margin. Keyu Wu 0002, Min Wu 0008, Zhenghua Chen, Ruibing Jin, Wei Cui 0002, Zhiguang Cao, Xiaoli Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Multi-Source Video Domain Adaptation With Temporal Attentive Moment Alignment NetworkabstractMulti-Source Domain Adaptation (MSDA) is a more practical domain adaptation scenario in real-world scenarios, which relaxes the assumption in conventional Unsupervised Domain Adaptation (UDA) that source data are sampled from a single domain and match a uniform data distribution. The MSDA is more challenging due to the existence of different domain shifts between distinct domain pairs. When considering videos, the negative transfer would be provoked by spatial-temporal features and can be formulated into a more challenging Multi-Source Video Domain Adaptation (MSVDA) problem. In this paper, we address the MSVDA problem by proposing a novel Temporal Attentive Moment Alignment Network (TAMAN) which aims for effective feature transfer by dynamically aligning both spatial and temporal feature moments. The TAMAN further constructs robust global temporal features by attending to dominant domain-invariant local temporal features with high local classification confidence and low disparity between global and local feature discrepancies. To facilitate future research on the MSVDA problem, we introduce comprehensive benchmarks, covering extensive MSVDA scenarios. Empirical results demonstrate a superior performance of the proposed TAMAN across multiple MSVDA benchmarks. Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Keyu Wu 0002, Min Wu 0008, Zhengguo Li, Zhenghua Chen |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Privacy-Preserving Cross-Environment Human Activity RecognitionabstractRecent studies have demonstrated the success of using the channel state information (CSI) from the WiFi signal to analyze human activities in a fixed and well-controlled environment. Those systems usually degrade when being deployed in new environments. A straightforward solution to solve this limitation is to collect and annotate data samples from different environments with advanced learning strategies. Although workable as reported, those methods are often privacy sensitive because the training algorithms need to access the data from different environments, which may be owned by different organizations. We present a practical method for the WiFi-based privacy-preserving cross-environment human activity recognition (HAR). It collects and shares information from different environments, while maintaining the privacy of individual person being involved. At the core of our approach is the utilization of the Johnson-Lindenstrauss transform, which is theoretically shown to be differentially private. Based on that, we further design an adversarial learning strategy to generate environment-invariant representations for HAR. We demonstrate the effectiveness of the proposed method with different data modalities from two real-life environments. More specifically, on the raw CSI dataset, it shows 2.18% and 1.24% improvements over challenging baselines for two environments, respectively. Moreover, with the discrete wavelet transform features, it further yields 5.71% and 1.55% improvements, respectively. Le Zhang 0001, Wei Cui 0002, Bing Li 0002, Zhenghua Chen, Min Wu 0008, Sin G. Teo |
IEEE Trans. Cybern. | 4 |
| 2023 | Component Detection for Power Line Inspection Using a Graph-Based Relation Guiding NetworkabstractDetecting the components in aerial images is a crucial task in automatic visual inspection for power lines. Currently, deep learning models guided by external knowledge have achieved promising performances compared to directly applying the benchmark detectors. However, the component relationship, as human commonsense knowledge for object reasoning, is rarely investigated in this field. This study presents a graph-based relation guided network for power line component detection, which exploits correlations of regions, images, and categories. The visual relation module is employed to learn region-to-region relationship and enhance the visual features of each proposal that may contain components. Meanwhile, two guidance modules are proposed to capture image-to-region correlation and distinctively facilitate the category classification and position regression, which has not been considered in previous methods. Moreover, the category graphs built in these two modules are able to explore category-to-category dependencies that can further promote the network ability. Experimental results demonstrate that the proposed method can achieve more accurate and reasonable component detection compared to previous methods, which verifies the effectiveness of the proposed model incorporated with relation knowledge. Xinyu Liu 0006, Xiren Miao, Hao Jiang 0008, Jing Chen 0022, Min Wu 0008, Zhenghua Chen |
IEEE Trans. Ind. Informatics | 6 |
| 2023 | ECG-CL: A Comprehensive Electrocardiogram Interpretation Method Based on Continual LearningabstractThe value of Electrocardiogram (ECG) monitoring in early cardiovascular disease (CVD) detection is undeniable, especially with the aid of intelligent wearable devices. Despite this, the requirement for expert interpretation significantly limits public accessibility, underscoring the need for advanced diagnosis algorithms. Deep learning-based methods represent a leap beyond traditional rule-based algorithms, but they are not without challenges such as small databases, inefficient use of local and global ECG information, high memory requirements for deploying multiple models, and the absence of task-to-task knowledge transfer. In response to these challenges, we propose a multi-resolution model adept at integrating local morphological characteristics and global rhythm patterns seamlessly. We also introduce an innovative ECG continual learning (ECG-CL) approach based on parameter isolation, designed to enhance data usage effectiveness and facilitate inter-task knowledge transfer. Our experiments, conducted on four publicly available databases, provide evidence of our proposed continual learning method's ability to perform incremental learning across domains, classes, and tasks. The outcome showcases our method's capability in extracting pertinent morphological and rhythmic features from ECG segmentation, resulting in a substantial enhancement of classification accuracy. This research not only confirms the potential for developing comprehensive ECG interpretation algorithms based on single-lead ECGs but also fosters progress in intelligent wearable applications. By leveraging advanced diagnosis algorithms, we aspire to increase the accessibility of ECG monitoring, thereby contributing to early CVD detection and ultimately improving healthcare outcomes. Hongxiang Gao, Xingyao Wang 0001, Zhenghua Chen, Min Wu 0008, Jianqing Li 0002, Chengyu Liu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | ADATIME: A Benchmarking Suite for Domain Adaptation on Time Series DataabstractUnsupervised domain adaptation methods aim at generalizing well on unlabeled test data that may have a different (shifted) distribution from the training data. Such methods are typically developed on image data, and their application to time series data is less explored. Existing works on time series domain adaptation suffer from inconsistencies in evaluation schemes, datasets, and backbone neural network architectures. Moreover, labeled target data are often used for model selection, which violates the fundamental assumption of unsupervised domain adaptation. To address these issues, we develop a benchmarking evaluation suite ( AdaTime ) to systematically and fairly evaluate different domain adaptation methods on time series data. Specifically, we standardize the backbone neural network architectures and benchmarking datasets, while also exploring more realistic model selection approaches that can work with no labeled data or just a few labeled samples. Our evaluation includes adapting state-of-the-art visual domain adaptation methods to time series data as well as the recent methods specifically developed for time series data. We conduct extensive experiments to evaluate 11 state-of-the-art methods on five representative datasets spanning 50 cross-domain scenarios. Our results suggest that with careful selection of hyper-parameters, visual domain adaptation methods are competitive with methods proposed for time series domain adaptation. In addition, we find that hyper-parameters could be selected based on realistic model selection approaches. Our work unveils practical insights for applying domain adaptation methods on time series data and builds a solid foundation for future works in the field. The code is available at github.com/emadeldeen24/AdaTime . Mohamed Ragab 0002, Emadeldeen Eldele, Wee Ling Tan, Chuan-Sheng Foo, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2023 | Learning Large Neighborhood Search for Vehicle Routing in Airport Ground HandlingabstractDispatching vehicle fleets to serve flights is a key task in airport ground handling (AGH). Due to the notable growth of flights, it is challenging to simultaneously schedule multiple types of operations (services) for a large number of flights, where each type of operation is performed by one specific vehicle fleet. To tackle this issue, we first represent the operation scheduling as a complex vehicle routing problem and formulate it as a mixed integer linear programming (MILP) model. Then given the graph representation of the MILP model, we propose a learning assisted large neighborhood search (LNS) method using data generated based on real scenarios, where we integrate imitation learning and graph convolutional network (GCN) to learn a destroy operator to automatically select variables, and employ an off-the-shelf solver as the repair operator to reoptimize the selected variables. Experimental results based on a real airport show that the proposed method allows for handling up to 200 flights with 10 types of operations simultaneously, and outperforms state-of-the-art methods. Moreover, the learned method performs consistently accompanying different solvers, and generalizes well on larger instances, verifying the versatility and scalability of our method. Jianan Zhou 0002, Yaoxin Wu, Zhiguang Cao, Wen Song 0004, Jie Zhang 0002, Zhenghua Chen |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | A Hybrid Accuracy- and Energy-Aware Human Activity Recognition Model in IoT EnvironmentabstractPersonalised health and fitness provide users with information regarding their wellbeing and an opportunity to inform healthcare services for better patient outcomes. Underpinning this industry sector is the need to establish human activity recognition (HAR) in a ubiquitous manner. For example, through the use of smartwatches and/or mobile phones gathering information such as heart rates, movement, and steps of a user. The engineering challenge is providing accurate, informative, and timely data without rapidly depleting the mobile device's battery life. This problem is compounded as a number of algorithms used to process such data require substantial, cloud-based resources, to achieve higher accuracy. Therefore, a balance is required between battery depletion, accuracy of data, and timely delivery of results through a mixture of cloud and local algorithmic execution. In this article, we proposeAE-HAR (Accuracy and Energy Aware-HAR)model that delivers engineered solutions which approach optimal combinations in the consideration of energy consumption, accuracy, and timeliness of results.AE-HARintroduces a “light-weight”machine learningon-device component identifying the probabilistic accuracy of data together with energy consumption identification requirements. A heuristic is then adopted to determine if cloud-enabled calculations are required while including possible performance costs related to the analysis of networking infrastructures. Our model is validated in a real-world environment through experimentation that demonstrates accuracy in excess of 93% and energy consumption savings in excess of 94%. Devki Nandan Jha, Zhenghua Chen, Shudong Liu 0003, Min Wu 0008, Jiahan Zhang, Graham Morgan, Rajiv Ranjan 0001, Xiaoli Li 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2022 | Generalizing Reinforcement Learning through Fusing Self-Supervised Learning into Intrinsic MotivationabstractDespite the great potential of reinforcement learning (RL) in solving complex decision-making problems, generalization remains one of its key challenges, leading to difficulty in deploying learned RL policies to new environments. In this paper, we propose to improve the generalization of RL algorithms through fusing Self-supervised learning into Intrinsic Motivation (SIM). Specifically, SIM boosts representation learning through driving the cross-correlation matrix between the embeddings of augmented and non-augmented samples close to the identity matrix. This aims to increase the similarity between the embedding vectors of a sample and its augmented version while minimizing the redundancy between the components of these vectors. Meanwhile, the redundancy reduction based self-supervised loss is converted to an intrinsic reward to further improve generalization in RL via an auxiliary objective. As a general paradigm, SIM can be implemented on top of any RL algorithm. Extensive evaluations have been performed on a diversity of tasks. Experimental results demonstrate that SIM consistently outperforms the state-of-the-art methods and exhibits superior generalization capability and sample efficiency. Keyu Wu 0002, Min Wu 0008, Zhenghua Chen, Yuecong Xu, Xiaoli Li 0001 |
AAAI | 3 |
| 2022 | Source-Free Video Domain Adaptation by Learning Temporal Consistency for Action Recognition
Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Keyu Wu 0002, Min Wu 0008, Zhenghua Chen |
ECCV (34) | 6 |
| 2022 | Contrastive adversarial knowledge distillation for deep model compression in time-series regression tasks
Qing Xu 0015, Zhenghua Chen, Mohamed Ragab 0002, Chao Wang 0003, Min Wu 0008, Xiaoli Li 0001 |
Neurocomputing | 2 |
| 2021 | Two-Stream Convolution Augmented Transformer for Human Activity RecognitionabstractRecognition of human activities is an important task due to its far-reaching applications such as healthcare system, context-aware applications, and security monitoring. Recently, WiFi based human activity recognition (HAR) is becoming ubiquitous due to its non-invasiveness. Existing WiFi-based HAR methods regard WiFi signals as a temporal sequence of channel state information (CSI), and employ deep sequential models (e.g., RNN, LSTM) to automatically capture channel-over-time features. Although being remarkably effective, they suffer from two major drawbacks. Firstly, the granularity of a single temporal point is blindly elementary for representing meaningful CSI patterns. Secondly, the time-over-channel features are also important, and could be a natural data augmentation. To address the drawbacks, we propose a novel Two-stream Convolution Augmented Human Activity Transformer (THAT) model. Our model proposes to utilize a two-stream structure to capture both time-over-channel and channel-over-time features, and use the multi-scale convolution augmented transformer to capture range-based patterns. Extensive experiments on four real experiment datasets demonstrate that our model outperforms state-of-the-art models in terms of both effectiveness and efficiency. Bing Li 0002, Wei Cui 0002, Wei Wang 0011, Le Zhang 0001, Zhenghua Chen, Min Wu 0008 |
AAAI | 5 |
| 2021 | A Multi-Stage Progressive Learning Strategy for Covid-19 Diagnosis Using Chest Computed Tomography with Imbalanced DataabstractIn this paper, a multi-stage progressive learning strategy is investigated to train classifiers for COVID-19 Diagnosis using imbalanced Chest Computed Tomography Data acquired from patients infected with COVID-19 Pneumonia, Community Acquired Pneumonia (CAP) and from normal healthy subjects. In the first learning stage, pre-processed volumetric CT data together with the segmented lung masks are fed into a 3D ResNet module, and an initial classification result can be obtained. However, due to categorical data imbalance, we observe large differences in sensitivity between COVID-19 and CAP cases. In the second stage, five learning models are independently trained over data with only COVID-19 and CAP cases, and are then ensembled to further discriminate the two classes. The final classification results are obtained by combining the predictions from both stages. Based on the validation dataset, we have evaluated our method and compared it with up-to-date methods in terms of overall accuracy and sensitivity for each class. The validation results validate the accuracy of the proposed multi-stage learning strategy. The overall accuracy of the validation dataset is 88.8%, and the sensitivities are 0.873, 0.789 and 1 for COVID-19, CAP and normal cases, respectively. Zaifeng Yang, Yubo Hou, Zhenghua Chen, Le Zhang 0001, Jie Chen 0026 |
ICASSP | 3 |
| 2021 | Partial Video Domain Adaptation with Partial Adversarial Temporal Attentive NetworkabstractPartial Domain Adaptation (PDA) is a practical and general domain adaptation scenario, which relaxes the fully shared label space assumption such that the source label space subsumes the target one. The key challenge of PDA is the issue of negative transfer caused by source-only classes. For videos, such negative transfer could be triggered by both spatial and temporal features, which leads to a more challenging Partial Video Domain Adaptation (PVDA) problem. In this paper, we propose a novel Partial Adversarial Temporal Attentive Network (PATAN) to address the PVDA problem by utilizing both spatial and temporal features for filtering source-only classes. Besides, PATAN constructs effective overall temporal features by attending to local temporal features that contribute more toward the class filtration process. We further introduce new benchmarks to facilitate research on PVDA problems, covering a wide range of PVDA scenarios. Empirical results demonstrate the state-of-the-art performance of our proposed PATAN across the multiple PVDA benchmarks. Code will be provided at: https://github.com/xuyu0010/PATAN. Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Zhenghua Chen, Kezhi Mao |
ICCV | 4 |
| 2021 | Time-Series Representation Learning via Temporal and Contextual ContrastingabstractLearning decent representations from unlabeled time-series data with temporal dynamics is a very challenging task. In this paper, we propose an unsupervised Time-Series representation learning framework via Temporal and Contextual Contrasting (TS-TCC), to learn time-series representation from unlabeled data. First, the raw time-series data are transformed into two different yet correlated views by using weak and strong augmentations. Second, we propose a novel temporal contrasting module to learn robust temporal representations by designing a tough cross-view prediction task. Last, to further learn discriminative representations, we propose a contextual contrasting module built upon the contexts from the temporal contrasting module. It attempts to maximize the similarity among different contexts of the same sample while minimizing similarity among contexts of different samples. Experiments have been carried out on three real-world time-series datasets. The results manifest that training a linear classifier on top of the features learned by our proposed TS-TCC performs comparably with the supervised training. Additionally, our proposed TS-TCC shows high efficiency in few-labeled data and transfer learning scenarios. The code is publicly available at https://github.com/emadeldeen24/TS-TCC. Emadeldeen Eldele, Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001, Cuntai Guan |
IJCAI | 3 |
| 2021 | Deep Reinforcement Learning Boosted Partial Domain AdaptationabstractDomain adaptation is critical for learning transferable features that effectively reduce the distribution difference among domains. In the era of big data, the availability of large-scale labeled datasets motivates partial domain adaptation (PDA) which deals with adaptation from large source domains to small target domains with less number of classes. In the PDA setting, it is crucial to transfer relevant source samples and eliminate irrelevant ones to mitigate negative transfer. In this paper, we propose a deep reinforcement learning based source data selector for PDA, which is capable of eliminating less relevant source samples automatically to boost existing adaptation methods. It determines to either keep or discard the source instances based on their feature representations so that more effective knowledge transfer across domains can be achieved via filtering out irrelevant samples. As a general module, the proposed DRL-based data selector can be integrated into any existing domain adaptation or partial domain adaptation models. Extensive experiments on several benchmark datasets demonstrate the superiority of the proposed DRL-based data selector which leads to state-of-the-art performance for various PDA tasks. Keyu Wu 0002, Min Wu 0008, Jianfei Yang 0001, Zhenghua Chen, Zhengguo Li, Xiaoli Li 0001 |
IJCAI | 4 |
| 2021 | Learning to Iteratively Solve Routing Problems with Dual-Aspect Collaborative TransformerabstractRecently, Transformer has become a prevailing deep architecture for solving vehicle routing problems (VRPs). However, it is less effective in learning improvement models for VRP because its positional encoding (PE) method is not suitable in representing VRP solutions. This paper presents a novel Dual-Aspect Collaborative Transformer (DACT) to learn embeddings for the node and positional features separately, instead of fusing them together as done in existing ones, so as to avoid potential noises and incompatible correlations. Moreover, the positional features are embedded through a novel cyclic positional encoding (CPE) method to allow Transformer to effectively capture the circularity and symmetry of VRP solutions (i.e., cyclic sequences). We train DACT using Proximal Policy Optimization and design a curriculum learning strategy for better sample efficiency. We apply DACT to solve the traveling salesman problem (TSP) and capacitated vehicle routing problem (CVRP). Results show that our DACT outperforms existing Transformer based improvement models, and exhibits much better generalization performance across different problem sizes on synthetic and benchmark instances, respectively. Yining Ma 0001, Zhiguang Cao, Wen Song 0004, Le Zhang 0001, Zhenghua Chen, Jing Tang 0004 |
NeurIPS | 6 |
| 2021 | RDMA Based Performance Optimization on Distributed Database Systems: A Case Study with GoldenX
Yaofeng Tu, Yinjun Han, Zhenghua Chen, Yanchao Zhao |
WASA (2) | 4 |
| 2021 | Detecting the shuttlecock for a badminton robot: A YOLO based approach
Zhiguang Cao, Tingbo Liao, Wen Song 0004, Zhenghua Chen, Chongshou Li |
Expert Syst. Appl. | 4 |
| 2021 | Deep learning for human activity recognition
Xiaoli Li 0001, Peilin Zhao, Min Wu 0008, Zhenghua Chen, Le Zhang 0001 |
Neurocomputing | 4 |
| 2021 | Attention-based sequence to sequence model for machine remaining useful life prediction
Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Ruqiang Yan 0001, Xiaoli Li 0001 |
Neurocomputing | 2 |
| 2021 | Key Nodes Selection in Controlling Complex Networks via Convex OptimizationabstractKey nodes are the nodes connected with a given number of external source controllers that result in minimal control cost. Finding such a subset of nodes is a challenging task since it impossible to list and evaluate all possible solutions unless the network is small. In this paper, we approximately solve this problem by proposing three algorithms step by step. By relaxing the Boolean constraints in the original optimization model, a convex problem is obtained. Then inexact alternating direction method of multipliers (IADMMs) is proposed and convergence property is theoretically established. Based on the degree distribution, an extension method named degree-based IADMM (D-IADMM) is proposed such that key nodes are pinpointed. In addition, with the technique of local optimization employed on the results of D-IADMM, we also develop LD-IADMM and the performance is greatly improved. The effectiveness of the proposed algorithms is validated on different networks ranging from Erdős-Rényi networks and scale-free networks to some real-life networks. Jie Ding 0007, Changyun Wen, Guoqi Li 0002, Zhenghua Chen |
IEEE Trans. Cybern. | 4 |
| 2021 | Contrastive Adversarial Domain Adaptation for Machine Remaining Useful Life PredictionabstractEnabling precise forecasting of the remaining useful life (RUL) for machines can reduce maintenance cost, increase availability, and prevent catastrophic consequences. Data-driven RUL prediction methods have already achieved acclaimed performance. However, they usually assume that the training and testing data are collected from the same condition (same distribution or domain), which is generally not valid in real industry. Conventional approaches to address domain shift problems attempt to derive domain-invariant features, but fail to consider target-specific information, leading to limited performance. To tackle this issue, in this article, we propose a contrastive adversarial domain adaptation (CADA) method for cross-domain RUL prediction. The proposed CADA approach is built upon an adversarial domain adaptation architecture with a contrastive loss, such that it is able to take target-specific information into consideration when learning domain-invariant features. To validate the superiority of the proposed approach, comprehensive experiments have been conducted to predict the RULs of aeroengines across 12 cross-domain scenarios. The experimental results show that the proposed method significantly outperforms state-of-the-arts with over 21% and 38% improvements in terms of two different evaluation metrics. Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chuan-Sheng Foo, Chee Keong Kwoh 0001, Ruqiang Yan 0001, Xiaoli Li 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | An Attention Based CNN-LSTM Approach for Sleep-Wake Detection With Heterogeneous SensorsabstractIn this article, we propose an attention based convolutional neural network long short-term memory (CNN-LSTM) approach for sleep-wake detection with heterogeneous sensor data, i.e., acceleration and heart rate variability (HRV). Since the three-dimensional acceleration data was sampled with a high frequency, we firstly design a CNN-LSTM structure to effectively learn latent features from the acceleration. Meanwhile, considering the unique format of the HRV data, some effective features are extracted based on domain knowledge. Next, we design a unified architecture to efficiently merge the features learned by CNN-LSTM approach from the acceleration and the extracted features from the HRV, which enables us to make full use of all the available information from these two heterogeneous sources. Taking into consideration that these two heterogeneous sources may have distinct contributions for the sleep and wake states, we propose an attention network to dynamically adjust the importance of features from the two sources. Real-world experiments have been conducted to verify the effectiveness of the proposed approach for sleep-wake detection. The results demonstrate that the proposed method outperforms all existing approaches for sleep-wake classification. In the evaluation of leave-one-subject-out (LOSO) cross-validation which is more challenging and practical, the proposed method achieves remarkable improvements ranging from 5% to 46% over the benchmark approaches. Zhenghua Chen, Min Wu 0008, Wei Cui 0002, Chengyu Liu 0001, Xiaoli Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Mahalanobis Distance Based Adversarial Network for Anomaly DetectionabstractAnomaly detection techniques are very crucial in multiple business applications, such as cyber security, manufacturing and finance. However, developing anomaly detection methods for high-dimensional data with high speed and good performance is still a challenge. Generative Adversarial Networks (GANs) are able to model the complex high-dimensional data, but they still require large computation in inference stage. This paper proposes an efficient method, known as Mahalanobis Distance-based Adversarial Network (MDAN), for anomaly detection. The proposed MDAN models the data using generative adversarial network (GAN) and detects anomalies by using the Mahalanobis distance. The proposed MDAN outperforms conventional GAN-based methods considerably and has a higher inference speed, when applied to several tabular and image datasets. Yubo Hou, Zhenghua Chen, Min Wu 0008, Chuan-Sheng Foo, Xiaoli Li 0001, Raed M. Shubair |
ICASSP | 2 |
| 2020 | URFS: A User-space Raw File System based on NVMe SSDabstractNVMe (Non-Volatile Memory Express) is a protocol designed specifically for SSD (Solid State Drive), which has significantly improved the performance of SSD storage devices. However, the traditional kernel-space IO path hinders the performance of NVMe SSD devices. In this paper, a user-space raw file system (URFS) based on NVMe SSD is proposed. Through the design of the user-space multi-process shared cache, multiple applications can share access to SSD to reduce the amount of SSD access; NVMe-oriented log-free data layout and Multi-granularity IO queue elastic separation technology are used to improve system performance and throughput. Experiments show that, compared to traditional file systems, URFS performance is improved by more than 23% in CDN (Content Delivery Network) scenarios, and URFS performance is improved more in small file scenarios and read-intensive scenarios. Yaofeng Tu, Yinjun Han, Zhenghua Chen, Zhengguang Chen, Bing Chen 0002 |
ICPADS | 3 |
| 2020 | Learning From Paired and Unpaired Data: Alternately Trained CycleGAN for Near Infrared Image ColorizationabstractThis paper presents a novel near infrared (NIR) image colorization approach for the Grand Challenge held by 2020 IEEE International Conference on Visual Communications and Image Processing (VCIP). A Cycle-Consistent Generative Adversarial Network (CycleGAN) with cross-scale dense connections is developed to learn the color translation from the NIR domain to the RGB domain based on both paired and unpaired data. Due to the limited number of paired NIR-RGB images, data augmentation via cropping, scaling, contrast and mirroring operations have been adopted to increase the variations of the NIR domain. An alternating training strategy has been designed, such that CycleGAN can efficiently and alternately learn the explicit pixel-level mappings from the paired NIR-RGB data, as well as the implicit domain mappings from the unpaired ones. Based on the validation data, we have evaluated our method and compared it with conventional CycleGAN method in terms of peak signal-to-noise ratio (PSNR), structural similarity (SSIM) and angular error (AE). The experimental results validate the proposed colorization framework. Zaifeng Yang, Zhenghua Chen |
VCIP | 2 |
| 2020 | Semi-Supervised Deep Learning Based Wireless Interference Identification for IIoT NetworksabstractAccurate wireless interference identification (WII) is vital for wireless industrial internet of things (IIoT) network to coexist with other technologies in the crowded 2.4 GHz unlicensed band. Deep learning (DL) based methods have emerged as a promising candidate for such type of task. However, to achieve good accuracy, DL methods require large amount of labeled training data, which comes from tedious annotation work by domain expert. In contrast, unlabeled data is easier to obtain. In this paper we present a semi-supervised DL based WII algorithm which combines temporal ensembling technique with CNN network to exploit unlabeled data to improve the performance. The proposed algorithm is able to differentiate interference from multiple wireless standards accurately with reduced number of labels, such as IEEE 802.11, IEEE 802.15.4 and IEEE 802.15.1. Specifically, the proposed algorithm achieves 90% accuracy with less than 2% of labeled data with medium to high signal SNR. Extensive simulation results show that the proposed algorithm achieves a better classification accuracy than benchmark algorithms under various SNR conditions and with different number of labeled data. Jiajia Huang 0004, Min Li Huang, Peng Hui Tan, Zhenghua Chen, Sumei Sun |
VTC Fall | 4 |
| 2020 | MobileDA: Toward Edge-Domain AdaptationabstractDeep neural networks (DNNs) have made significant advances in computer vision and sensor-based smart sensing. DNNs achieve prominent results based on standard data sets and powerful servers, whereas, in real applications with domain-shift data and resource-constrained environments such as Internet-of-Things (IoT) devices in the edge computing, DNNs are likely to have degraded performance in terms of accuracy and efficiency. To this end, we develop the MobileDA framework that learns transferable features while keeping the simple structure of the deep model. Our method allows a novel teacher network trained in the server to distill the knowledge for a student network running in the edge device, which is achieved by a cross-domain distillation. Leveraging unlabeled data in the new environment, our student model amends the feature learning to be domain invariant, then being our objective model running in the edge device. Our approach is evaluated on a challenging IoT-based WiFi gesture recognition scenario, and three classic visual adaptation benchmarks. The empirical studies corroborate the effectiveness of distillation for domain transfer, and the overall results show that our model achieves state-of-the-art performance merely using a simple network. Jianfei Yang 0001, Han Zou, Shuxin Cao, Zhenghua Chen, Lihua Xie 0001 |
IEEE Internet Things J. | 4 |
| 2020 | WiFi-Based Indoor Robot Positioning Using Deep Fuzzy ForestsabstractAddressing the positioning problem of a mobile robot remains challenging to date despite many years of research. Indoor robot positioning strategies developed in the literature either rely on sophisticated computer vision techniques to handle visual inputs or require strong domain knowledge for nonvisual sensors. Although some systems have been deployed, the former may be lacking due to the intrinsic limitation of cameras (such as calibration, data association, system initialization, etc.) and the latter usually only works under certain environment layouts and additional equipment. To cope with those issues, we design a lightweight indoor robot positioning system which operates on cost-effective WiFi-based received signal strength (RSS) and could be readily pluggable into any existing WiFi network infrastructures. Moreover, a novel deep fuzzy forest is proposed to inherit the merits of decision trees and deep neural networks within an end-to-end trainable architecture. Real-world indoor localization experiments are conducted and results demonstrate the superiority of the proposed method over the existing approaches. Le Zhang 0001, Zhenghua Chen, Wei Cui 0002, Bing Li 0002, Cen Chen 0002, Zhiguang Cao, Kai-Zhou Gao |
IEEE Internet Things J. | 2 |
| 2019 | A Novel Ensemble ELM for Human Activity Recognition Using Smartphone SensorsabstractHuman activity recognition plays a unique role in many important applications, including ubiquitous computing, health-care services, and smart buildings. Due to the nonintrusive property of smartphones, smartphone sensors are widely used for the identification of human activities. Since the signals of smartphone sensors are quite noisy, feature engineering will be performed to extract more discriminant representations. Then, various machine learning algorithms can be employed to recognize different human activities. Extreme learning machine (ELM) has been shown to be effective in classification tasks with extremely fast learning speed. Due to its randomness property, it is naturally suitable for ensemble learning. In this paper, we propose a novel ensemble ELM algorithm for human activity recognition using smartphone sensors. Gaussian random projection is employed to initialize the input weights of base ELMs. By doing this, more diversities can be generated to boost the performance of ensemble learning. Real experimental data has been applied to evaluate the performance of our proposed approach. We also conduct a comparison of the proposed approach with some state-of-the-art approaches in the literature. The experimental results indicate that our proposed ensemble ELM approach outperforms these approaches and can achieve recognition accuracies of$\text{97.35}\%$and$\text{98.88}\%$on two datasets. Zhenghua Chen, Chaoyang Jiang, Lihua Xie 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | A Novel Semisupervised Deep Learning Method for Human Activity RecognitionabstractHuman activity recognition (HAR) based on inertial sensors has been investigated for many industrial informatics applications, such as healthcare and ubiquitous computing. Existing methods mainly rely on supervised learning schemes, which require large labeled training data. However, labeled data are sometimes difficult to acquire, while unlabeled data are readily available. Thus, we intend to make use of both labeled and unlabeled data with semisupervised learning for accurate HAR. In this paper, we propose a semisupervised deep learning approach, using temporal ensembling of deep long short-term memory, to recognize human activities with smartphone inertial sensors. With the deep neural network processing, features are extracted for local dependencies in the recurrent framework. Besides, with an ensemble approach based on both labeled and unlabeled data, we can combine together the supervised and unsupervised losses, so as to make good use of unlabeled data that the supervised learning method cannot leverage. Experimental results indicate the effectiveness of our proposed semisupervised learning scheme, when compared to several state-of-the-art semisupervised learning approaches. Qingchang Zhu, Zhenghua Chen, Yeng Chai Soh |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | WiFi CSI Based Passive Human Activity Recognition Using Attention Based BLSTMabstractHuman activity recognition can benefit various applications including healthcare services and context awareness. Since human actions will influence WiFi signals, which can be captured by the channel state information (CSI) of WiFi, WiFi CSI based human activity recognition has gained more and more attention. Due to the complex relationship between human activities and WiFi CSI measurements, the accuracies of current recognition systems are far from satisfactory. In this paper, we propose a new deep learning based approach, i.e., attention based bi-directional long short-term memory (ABLSTM), for passive human activity recognition using WiFi CSI signals. The BLSTM is employed to learn representative features in two directions from raw sequential CSI measurements. Since the learned features may have different contributions for final activity recognition, we leverage on an attention mechanism to assign different weights for all the learned features. Real experiments have been carried out to evaluate the performance of the proposed ABLSTM for human activity recognition. The experimental results show that our proposed ABLSTM is able to achieve the best recognition performance for all activities when compared with some benchmark approaches. Zhenghua Chen, Le Zhang 0001, Chaoyang Jiang, Zhiguang Cao, Wei Cui 0002 |
IEEE Trans. Mob. Comput. | 1 |
| 2018 | Building Occupancy Detection from Carbon-dioxide and Motion SensorsabstractOccupant detection using carbon-dioxide sensors is prevalent but its accuracy is restricted by the inherent sensing delays. This paper proposes an indoor occupant detection method using real-time carbon-dioxide and Pyroelectric Infrared (PIR) sensor measurements overcoming the sensing delays. The occupancy detection problem is formulated as a classification problem wherein the classifier learns from offline carbon-dioxide data and the actual occupancy measurements of the room. While the classifier can provide realtime occupancy detection, the delays in carbon-dioxide sensors influence their accuracy. To overcome the delays, observations from PIR sensors are combined with the results of the single-layer feedforward neural network (SLFN) based classifier. The classifier works in four steps: (i) data-preprocessing, (ii) feature-selection, (iii) learning, and (iv) validation. The data is preprocessed by smoothing and several features are selected as input to the SLFN. Then, the classifier is validated with realtime experiments. Our results demonstrate that the proposed approach provides accuracy up to 99.79% and also overcomes the delays found in carbon-dioxide sensors. Chaoyang Jiang, Zhenghua Chen, Lih Chieh Png, Korkut Bekiroglu, Seshadhri Srinivasan, Rong Su 0001 |
ICARCV | 2 |
| 2018 | Distilling the Knowledge From Handcrafted Features for Human Activity RecognitionabstractHuman activity recognition is a core problem in intelligent automation systems due to its far-reaching applications including ubiquitous computing, health-care services, and smart living. Due to the nonintrusive property of smartphones, smartphone sensors are widely used for the identification of human activities. However, unlike applications in vision or data mining domain, feature embedding from deep neural networks performs much worse in terms of recognition accuracy than properly designed handcrafted features. In this paper, we posit that feature embedding from deep neural networks may convey complementary information and propose a novel knowledge distilling strategy to improve its performance. More specifically, an efficient shallow network, i.e., single-layer feedforward neural network (SLFN), with handcrafted features is utilized to assist a deep long short-term memory (LSTM) network. On the one hand, the deep LSTM network is able to learn features from raw sensory data to encode temporal dependencies. On the other hand, the deep LSTM network can also learn from SLFN to mimic how it generalizes. Experimental results demonstrate the superiority of the proposed method in terms of recognition accuracy against several state-of-the-art methods in the literature. Zhenghua Chen, Le Zhang 0001, Zhiguang Cao, Jing Guo 0007 |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | Environmental Sensors-Based Occupancy Estimation in Buildings via IHMM-MLRabstractOccupancy estimation in buildings can benefit various applications such as heating, ventilation, and air-conditioning control, space monitoring, and emergency evacuation. Due to the consideration of temporal dependency in occupancy data, hidden Markov model (HMM) has been shown to be effective in occupancy estimation. However, the conventional HMM that assumes invariant temporal dependency of occupancy dynamics for different time instances is unrealistic. Moreover, the performance of the conventional HMM that utilizes mixture of Gaussian for emission probability in terms of continuous observations can be easily affected by the noise in sensory data. To address these problems, in this paper, we propose a new architecture, i.e., inhomogeneous hidden Markov model with multinomial logistic regression (IHMM-MLR), for building occupancy estimation using nonintrusive environmental sensors. Instead of using the time-invariant transition probability matrix, we apply a time-dependent (inhomogeneous) transition probability matrix which can capture the temporal dependency for different time instances. Meanwhile, we employ an efficient probabilistic model, i.e., MLR, for emission probability. Online and offline occupancy estimation schemes are presented for real-time and accurate long-term applications respectively. Real experiments have indicated the effectiveness of our proposed approach. Zhenghua Chen, Qingchang Zhu, Mustafa K. Masood, Yeng Chai Soh |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | Robust Human Activity Recognition Using Smartphone Sensors via CT-PCA and Online SVMabstractHuman activity recognition using either wearable devices or smartphones can benefit various applications including healthcare, fitness, smart home, etc. Instead of using wearable devices which are intrusive and require extra cost, we shall leverage on modern smartphones embedded with a variety of sensors. Due to the flexibility of using smartphones, the recognition accuracy will degrade with orientation, placement, and subject variations. In this paper, we propose a robust human activity recognition system in terms of orientation, placement, and subject variations based on coordinate transformation and principal component analysis (CT-PCA) and online support vector machine (OSVM). The proposed CT-PCA scheme is utilized to eliminate the effect of orientation variations. Experiments show that the proposed scheme significantly improves the activity recognition accuracy and outperforms the state-of-the-art methods on leave one orientation out experiments, which demonstrates the generalization ability of the proposed scheme on the data from unseen orientations. We also show the effectiveness of this scheme on placement and subject variations. However, the inherent difference of signal properties for different placement and subject dramatically reduces the recognition accuracy, especially for different placement. Thus, we present an efficient OSVM algorithm, that is, online-independent support vector machine (OISVM), which utilizes a small portion of data from the unseen placement or subject to online update the parameters of the SVM algorithm. The experimental results demonstrate the effectiveness of this OISVM algorithm on placement and subject variations. Zhenghua Chen, Qingchang Zhu, Yeng Chai Soh, Le Zhang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2016 | Smartphone Inertial Sensor-Based Indoor Localization and Tracking With iBeacon CorrectionsabstractThe Global Positioning System (GPS) can be readily used for outdoor localization, but GPS signals are degraded in indoor environments. How to develop a robust and accurate indoor localization system is an emergent task. In this paper, we propose a smartphone inertial sensor-based indoor localization and tracking system with occasional iBeacon corrections. Some important issues in a smartphone-based pedestrian dead reckoning (PDR) approach, i.e., step detection, walking direction estimation, and initial point estimation, are studied. One problem of the PDR approach is the drift with walking distance. We apply a recent technology, iBeacon, to occasionally calibrate the drift of the PDR approach. By analyzing iBeacon measurements, we define an efficient calibration range where an extended Kalman filter is utilized. The proposed localization and tracking system can be implemented in resource-limited smartphones. To evaluate the performance of the proposed approach, real experiments under two different environments have been conducted. The experimental results demonstrated the effectiveness of the proposed approach. We also tested the localization accuracy with respect to the number of iBeacons. Zhenghua Chen, Qingchang Zhu, Yeng Chai Soh |
IEEE Trans. Ind. Informatics | 1 |
| 2010 | Ecosystem health assessment by using remote sensing derived data: A case study of terrestrial region along the coast in Zhejiang provinceabstract10-day maximum composite SPOT/VEGETATION NDVI from 1998 to 2007 were collected to assess terrestrial region ecosystem health along the coast in Zhejiang province. Vigor, organization and resilience were the main three characterization of ecosystem health, in this paper the average annual NDVI represented vigor, vegetation percentage and centriod movement represented organization, SLOPE represented resilience. Then integrated health value was calculated by using multiplying of those indicators. Then integrated ecosystem health results were acquired. The results were followings: (1) overall vigor remained steady in recent 10 years, but spatial heterogeneity existed; (2) the vegetation percentage decreased and DNVI centroid moved to inland obviously; (3) the lower SLOPE distributed around cities and coastline, vegetation resilience decreased; (4) the best status of ecosystem health were in Fenghua, Pingyang, the worst were in Shaoxing, Ningbo and Cixi. Zhenghua Chen, Qiu Yin, Li Li 0017 |
IGARSS | 1 |
| 2010 | A comparison of two stream approximation for the discrete ordinate method and the SOS methodabstractThe two-stream discrete ordinates method and the two-stream successive orders of scattering method are compared, and the key features of two methods are discussed. Based on the convergence characteristics of successive scattering, we use a semi-empirical model to improve the computing efficiency in the SOS method. Using the delta-M method, we investigate the effect of the two two-stream method for non-absorbing and absorbing case respectively. The 32-stream DISORT is used as the benchmark for assessments of the relative accuracy of the two methods investigated. With the comparisons for the accuracy of flux, the results of the two two-stream methods are almost the same in general and the absorbing media lead to larger errors of flux compared with that of the non-absorbing case. Weizhen Hou, Qiu Yin, Li Li 0017, Zhenghua Chen |
IGARSS | 5 |
| 2010 | Estimating chlorophyll a concentration in lake water using space-borne hyperspectral dataabstractChlorophyll fluorescence properties are effective in detecting chlorophyll a concentration in lake water, which provided new optional sensitive bands to retrieval chl-a concentration in complex water. Hyper Spectral Imager (HSI) on HJ-1A satellite with great spectral resolution can be used to detect spectral characteristics of chlorophyll fluorescence peak. The effects of empirical algorithms of lake Taihu and the algorithm based on height of fluorescence peak are analyzed by using MODIS data and HSI data respectively. The results show that compared with empirical MODIS algorithms of lake Taihu, chl-a retrieval based on HSI is more similar to the in-situ reference data. Li Li 0017, Qiu Yin, Cailan Gong, Zhenghua Chen |
IGARSS | 5 |
| 2010 | Spectral data analysis of ground objects in Chao Lake basinabstractBased on a great deal of spectral data for different kinds of ground objects which were denoised by wavelet transform method, the spectral characteristics and changing rules of water, paddy field, wheat, cole and vegetable greenhouse in Chao Lake basin were analyzed with ASD portable spectrum analyzer by field investigation and plot survey. The spectral data of typical ground objects in Chao Lake basin were also processed by using mathematics technology such as de-noising technology of derivative spectrum technology and normalization processing technology. It was shows that wavelet transform had advantage in de-noising because it could remove the noise from signal as well as preserve the detail information; derivative spectrum and normalization processing technology had better practicability in suppressing background effect and emphasizing the signal of objects. Qiu Yin, Li Li 0017, Zhenghua Chen, Yuhuan Ren, Weizhen Hou, Pengfei Yin |
IGARSS | 5 |
| 2010 | An atmospheric correction algorithm for hyperspectral imagery of lake water by Chinese satellite HJ-1AabstractThis paper demonstrates the Ruddick's algorithm to utilize atmospheric correction with the hyper-spectral imagery over Chinese turbid lake water obtained by China first hyperspectral imager (HSI) onboard HJ-1A. The paper studies on the sensor characteristics and analyzes the optical properties of turbid lake water in Taihu Lake. Based on consideration about real circumstances, the paper recalibrates parameterαof value taken as 1.43. Results indicate that the recalibrated parameter could enhance the algorithm performance and improve the accuracy through comparison with the in situ measurements. Xingfa Gu, Qiu Yin, Li Li 0017, Zhenghua Chen, Yuhuan Ren, Weizhen Hou, Pengfei Yin |
IGARSS | 5 |
| 2009 | A general framework for automatic on-line replay detection in sports videoabstractReplay detection is a pivotal step for sports video highlight extraction, which is a very promising application of multimedia analysis. In this paper, a general framework, which is based on a Bayesian network, is proposed to make full use of the multiple clues, including shot structure, gradual transition pattern, slow-motion, and sports scene. A novel algorithm based on motion vector reliability classification is proposed to analyze the gradual transition patterns, so that the replay detector can meet the requirements of automatic on-line applications. This is the first integrated general replay detection framework proposed in the literature. Extensive experiments on diversified sports games have proven the scheme efficient, accurate and robust. Zhenghua Chen, Chang Liu 0087, Weiguo Wu |
ACM Multimedia | 3 |
| 2007 | Assessing value of grassland ecosystem services in Gansu Province, northwest of ChinaabstractThere are three characteristics for ecosystem services, zonal characteristics, complexity and conventional civilization, which affect ecological valuation directly. Considering regional environment and anthropogenic influence, grassland's ecosystem services in Gansu Province are concluded 11 types, including gas regulation, climate regulation, disturbance regulation, wind erosion control and sand fixture, water resource provision, soil erosion control, waste treatment, gene resistance, food production, raw materials, recreation and culture. Landsat TM data are collected for manual photointerpretation. The grassland type distribution map is obtained. In this paper, we use stockbreeding income as benchmark, the other ecosystem services values which do not enter market to exchange are decided by experts questionary in this field. The relative weights can help to calculate the non-market ecological service value. The results' spatial distribution are displayed and analyzed by using Geographic Information Systems software. The results indicate: the total Gansu province grassland ecosystem value reaches 283.67times108yuan (RMB) in 1980s. The grassland in the low zone along Hexi corridor and Longnan have the highest ecosystem service value, the grassland in Gannan take second place, and the western of Qingyang and counties around Lanzhou take third place, the other places value in the province are low. The Gansu Province is a severe erosion region in arid and semiarid area, therefore wind erosion control and sand fixture, soil erosion control are the most valuable services. Zhenghua Chen, Jian Wang 0032, Qingyuan Ma |
IGARSS | 1 |
| 2007 | Management decision-making support system of percision agriculture based on CNCSabstractPrecision agriculture needs a mass of data, therefore data availability is a vital obstacle factor for precision agriculture technology widely use. The methods of acquisition data for precision agriculture change from field survey to sensor acquisition gradually. Remote Sensing plays an important role in the large-scale, real-time, multi-spectral data acquisition. But there are still some problems when Remote Sensing data are applied in the actual farm production and management, which result in that precision agriculture's advantages are limited, so finding the solution of data acquisition for precision agriculture based on remote sensing becomes very necessary. In this paper, the NPP, NDVI, LAI etc are retrieved from remote sensing data, the field investigation data are also obtained. We set up a decision-making system based on CNCS (central neural control system), which is composed of several module such as data updating, data editing, decision-making supporting. The system is made for the farmers to decide cultivation strategies on when, where and how much irrigation, fertilizer and pesticide should be applied on each plot. Qingyuan Ma, Zhenghua Chen |
IGARSS | 3 |