Yang Gu 0001

dblp:01/5858-1 · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 State Mamba: Spatiotemporal EEG State-Space Model with Dynamic Brain Alignment for Cross-Subject Representation
abstract
Cross-subject EEG decoding remains a fundamental challenge due to substantial inter-subject variability in brain activity, which hinders the development of subject-independent EEG models. Despite progress in extracting cross-subject invariant features, existing studies neglect the shared neural responses that arise under similar cognitive or emotional states across individuals, limiting their ability to learn generalized and consistent EEG representations. To address the challenges, we propose State Mamba, a novel spatiotemporal EEG state-space model that explicitly models and aligns neural responses and their spatiotemporal state transitions to learn consistent and generalizable representations across subjects. Innovatively, State Mamba theoretically formulates a multi-channel Mamba architecture that jointly models spatial and temporal brain state transitions, supporting principled analysis of neural responses. To enhance spatiotemporal feature coupling, we introduce the LGANN module, which adopts global-local attention to integrate long- and short-term brain activity into a compact EEG representation. Furthermore, we design two self-supervised pretext tasks to extract consistent neural patterns across subjects: (1) representation alignment to align EEG representation, and (2) pattern alignment to align their transition rules under identical conditions, jointly promoting subject-invariant EEG representations. Extensive experiments on three benchmark datasets, FACED, DEAP, and ISRUC, demonstrate the superior performance of State Mamba in cross-subject emotion and sleep recognition tasks, validating its robust generalization capability.
Weining Weng, Yang Gu 0001, Yingwei Zhang 0002, Yiqiang Chen 0001
AAAI2
2026 Multidimensional Haptic Perception and Quantification Method for Ophthalmic Surgery Training
abstract
Virtual ophthalmic surgical training is a cost-effective and time-efficient paradigm. However, the absence of haptic feedback in virtual ophthalmic surgery systems limits the realism and effectiveness of training. To address this challenge, this study proposes a task-driven method tailored to quantify haptic information in ophthalmic cataract surgery. This work designs a virtual cataract surgery interaction scenario and simulates three key haptic interactions: damping when moving the surgical tool within the eyeball, stiffness during corneal incision and phacoemulsification, and pull during continuous curvilinear capsulorhexis. Using a task-driven and subjective-objective integrated data analysis approach, this work quantifies the just noticeable threshold (JNT) and just noticeable difference (JND). Experimental results demonstrate that the proposed method effectively quantifies the perceptual thresholds and perceptual difference for varied individuals. These findings provide empirical evidence for establishing a more realistic and effective haptic feedback in virtual ophthalmic surgical training.
Yang Gu 0001, He Yuan, Maoyan Li, Yiqiang Chen 0001, Weiwei Dai
Int. J. Hum. Comput. Interact.1
2025 Semantic-oriented Visual Prompt Learning for Class Incremental Learning
abstract
Class-incremental learning (CIL) enables models to continuously learn new classes while addressing catastrophic forgetting. With the introduction of pre-trained models, new tuning paradigms have emerged for CIL. This paper revisits parameter-efficient fine-tuning (PEFT) methods in the context of incremental learning. Prior studies reveal that PEFT methods’ extended parameters do not directly contribute to semantic perception, limiting performance with significant category and domain gaps. To address this, we propose semantic-oriented visual prompt learning (SVPL), which enhances semantic perception and improves task-specific knowledge extraction. SVPL assigns learnable prompts to each class, using a contrastive group alignment to align prompts to task-specific semantic spaces, thus preserving relationships between old and new knowledge. Additionally, hierarchical semantic delivery allows the semantic transformation of prompt groups from shallow to deep layers to facilitate efficient knowledge mining and enable effective learning of new knowledge. Extensive experimental results on five benchmarks demonstrate the superior performance of our methods.
Shuai Guo 0001, Yang Gu 0001, Yingwei Zhang 0002, Weining Weng, Weiwei Dai, Yiqiang Chen 0001
ICASSP2
2025 Contrast Memory for Unsupervised Anomaly Detection
abstract
Real-world multivariate time series unsupervised anomaly detection is a challenging problem due to intricate temporal correlations. Recently, impressive progress have been made in tackling this issue through the design of large-scale models, facilitated by the growing model parameters. However, in resource-constrained scenarios such as ubiquitous computing and edge computing, the large-scale models suffer from issues like high parameter complexity and expensive training overheads. Existing methods can only strive for a direct tradeoff between model size and performance. To address this challenge, we propose DiMER (Diminutive Memory-Enhanced Reconstruction), a model with parameters of the order of 0.1M. In DiMER, we introduce a novel contrast memory mechanism to learn normal patterns with diminutive network and propose a temporal reconstruction loss nto add the autocorrelation information. In addition, we introduce a multi-space composite detection criterion, an anomaly score calculation that takes into account both memory space and data space. Extensive experiments on real-world datasets across various domains demonstrate that the proposed model achieves comparable or even superior performance to large-scale models while maintaining lightweight.
Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan
ICASSP4
2025 HYMAN: Hybrid Memory and Attention Network for Unsupervised Anomaly Detection
abstract
Detecting anomalies in unsupervised multivariate time series is challenging due to the intricate temporal patterns present in both local short-term and global long-term dependencies. Long short-term memory has achieved impressive results in this domain, yet it is gradually being supplemented by Transformers, due to limitations such as non-parallelization, gradient vanishing, and difficulty in focusing on local information. Leveraging the attention mechanism, Transformers can attend to all time steps simultaneously, effectively capturing local dependencies. However, they may often face challenges in efficiently modeling long-term dependencies, particularly in real-world scenarios. To address these issues, we propose the HYbrid Memory and Attention Network (HYMAN), which integrates memory and attention mechanisms together to model both global and local information. The attention captures short-term dependencies by focusing on temporal autocorrelation, while the memory stores and updates key historical patterns in global information, facilitating the learning of long-term dependencies. In contrast to previous approaches, HYMAN eliminates the need for auxiliary loss, simplifying the training by reducing the effort for coefficients tuning. In the inference phase, HYMAN introduces a novel anomaly scoring method that fuses features from both the temporal and latent spaces, offering high-performance detection compared to traditional methods that rely solely on reconstruction. Extensive experiments on real-world benchmark datasets demonstrate that HYMAN achieves state-of-the-art performance by leveraging the complementary strengths of memory and attention mechanisms.
Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan
ICASSP4
2025 DTAD: A Distribution-Transformed Supervised Anomaly Detection Method
abstract
Most anomaly detection (AD) methods adopt an unsupervised approach, relying exclusively on normal samples during training, which limits the model’s discriminative ability. In real world scenarios, only small amounts of anomaly data are typically available, which can still provide valuable insights for model learning. However, supervised anomaly detection methods may be impacted by the scarcity of anomaly samples, leading to significant overlap in the feature distributions of normal and anomaly samples, thereby degrading overall performance. To address this, we propose a distribution-transformed supervised anomaly detection method (DTAD). This method employs a two-stage distribution transformation to progressively reduce the overlap between normal and anomaly distributions, thereby enhancing the model’s discriminative performance. In the first stage, residual calculation is used to initially separate normal and abnormal distributions, with an attention network highlighting critical feature contributions. In the second stage, a contrastive loss function (semi-push-pull loss) is employed to expand the decision boundary between normal and abnormal samples, further improving distribution separation. On the benchmark datasets MVTecAD, AITEX, ELPV, and BrainMRI, our method outperforms recent state-of-the-art approaches, demonstrating its effectiveness.
Lingxing Chen, Yang Gu 0001, Jianqi Chen, Yingting Zhu, Yehong Zhuo, Dongmei Jiang, Yiqiang Chen 0001
ICME2
2025 Empowering Healthcare Time-Series Reasoning by Reinforcement Learning-Enhanced Large Language Models
Yang Gu 0001
ICONIP (5)2
2025 CARE: Contextual Residual and soft Encoding Relation Extraction for clinical medicine
abstract
Clinical event extraction involves extracting event attributes from clinical medical records. However, arguments involved in clinical events exhibit specificity, diversity and ambiguity, posing substantial challenges for existing models. The scarcity of Chinese clinical datasets further impedes research on clinical event extraction. Existing models commonly experience issues of contextual information decay during multi-task processes. Furthermore, unlike entities in general domains, medical entities are expressed in complex ways, resulting in low recall. To address these challenges, we propose a Clinical Event Extraction model based on Contextual ResiduAl and Soft Encoding Relation Extraction(CARE), which consists of an encoding module, a relation detection module, and an entity recognition module. The relation detection module identifies potential relations within a medical record. To prevent error propagation in relation extraction, a soft encoding strategy is proposed to discern target relations from candidate ones. The entity recognition module employs the contextual residual connection mechanism to concatenate the text with relation between semantic templates before feeding them into the entity recognition module. On CHIP-CDEE and CEMRs, CARE achieves F1 scores of 73.22% and 94.83%, respectively, which outperforms all baseline models, including LLMs, demonstrating its effectiveness for this task.
Kaifei Li, Qian Chen 0023, Suge Wang, Jian Liao 0005, Yang Gu 0001
IJCNN6
2025 An Effective GANs-Based Method for High-Information-Value Multimodal Wearable Sensor Data Synthesis
abstract
In the field of Human Activity Recognition (HAR), current GANs-based approaches for sensor data synthesis frequently fail to account for the varying informational significance of data samples, often treating them with uniform importance. This leads to the generation of a substantial number of data samples with limited informational utility by the trained generator, which ultimately fail to contribute meaningfully to HAR tasks. This paper introduces an effective GANs-based sensor data synthesis method for generating more data samples with high information value. Firstly, we incorporate active learning mechanisms into the adversarial training architecture to guide a more comprehensive learning of the data distribution. Secondly, we enhance the model’s ability to learn temporal, spatial, and spatial-temporal features by combining convolutional, recurrent, and self-attention modules. Thirdly, we validate the proposed method on the public dataset using both quantitative and qualitative metrics. The experimental results demonstrate that the proposed method can synthesize high-information-value multimodal sensor data, which are more valuable for training downstream HAR models.
Yang Gu 0001, Shuwang Zhou
IJCNN2
2025 A Geometric Constraints based Bayesian Neural Network for Virtual Surgery Assessment
abstract
Virtual surgical assessment evaluates a surgeon’s technical skills by analyzing instrument motion data collected during virtual surgery. Although existing methods classify skills accurately from motion data, they struggle to generalize to unfamiliar surgical patterns because surgeons use varied techniques. To address these limitations, this paper proposes a Geometric Constraints based Bayesian Neural Network (GeomBNN) model for surgical assessment. The geometric constraints enforce orthogonality among nonlinear components of the motion data, enhancing the model’s capacity to capture complex surgical patterns. The Bayesian neural network’s probabilistic framework quantifies predictive uncertainty, enabling adaptation to complex procedures and improving assessment robustness. Experiments on four datasets show that GeomBNN achieves higher classification accuracy than existing methods. What is more, the model exhibits progressive feature attention, dynamically focusing on the most discriminative surgical skill features during training, further validating its effectiveness.
He Yuan, Weiwei Dai, Danmin Cao, Yang Gu 0001
IJCNN6
2025 K-Space Bispectrum Steganography for Robust Unlearnable Data
Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan
ACM Multimedia4
2025 DDC: Dynamic distribution calibration for few-shot learning under multi-scale representation
Lingxing Chen, Yang Gu 0001, Dongmei Jiang, Yiqiang Chen 0001
Knowl. Based Syst.2
2025 DDIR: Domain-Disentangled Invariant Representation learning for tailored predictions
Yang Gu 0001, Shuai Guo 0001, Feiyi Fan, Yiqiang Chen 0001
Knowl. Based Syst.2
2025 PhysCL: Knowledge-Aware Contrastive Learning of Physiological Signal Models for Cuff-Less Blood Pressure Estimation
abstract
Training deep learning models for photoplethysmography(PPG)-based cuff-less blood pressure estimation often requires a substantial amount of labeled data collected through sophisticated medical instruments, posing significant challenges in practical applications. To address this issue, we propose Physiological Knowledge-Aware Contrastive Learning (PhysCL), a novel approach designed to reduce the dependence on labeled PPG data while improving blood pressure estimation accuracy. Specifically, PhysCL tackles the semantic consistency problem in contrastive learning by introducing a knowledge-aware augmentation bank, which generates positive physiological signal pairs using knowledge-based constraints during the contrastive pair generation. Additionally, we propose a contrastive feature reconstruction method to enhance feature diversity and prevent model collapse through feature re-sampling and re-weighting. We evaluate PhysCL on data from 106 subjects across the MIMIC III, MIMIC IV, and UQVS datasets under cross-dataset validation settings, comparing it against state-of-the-art contrastive learning methods and blood pressure estimation models. PhysCL achieves an average mean absolute error of 9.5/5.9 mmHg (systolic/diastolic) across the three datasets, using only 2% labeled data combined with 98% unlabeled data for pre-training and 5 samples for personalization, which represents a 6.2% /4.3% improvement, respectively, over the current best supervised methods. The ablation study provides further convincing evidence that the unlabeled data can be utilized to improve the existing cuff-less blood pressure estimation models and shed light on unsupervised contrastive learning for physiological signals.
Renju Liu, Jianfei Shen, Yang Gu 0001, Yiqiang Chen 0001, Jiling Zhang, Qingyu Wu, Chenyang Xu 0007, Feiyi Fan
IEEE J. Biomed. Health Informatics3
2025 Ten Challenging Problems in Federated Foundation Models
abstract
Federated Foundation Models (FedFMs) represent a distributed learning paradigm that fuses general competences of foundation models as well as privacy-preserving capabilities of federated learning. This combination allows the large foundation models and the small local domain models at the remote clients to learn from each other in a teacher-student learning setting. This paper provides a comprehensive summary of the ten challenging problems inherent in FedFMs, encompassing foundational theory, utilization of private data, continual learning, unlearning, Non-IID and graph data, bidirectional knowledge transfer, incentive mechanism design, game mechanism design, model watermarking, and efficiency. The ten challenging problems manifest in five pivotal aspects: “Foundational Theory,” which aims to establish a coherent and unifying theoretical framework for FedFMs. “Data,” addressing the difficulties in leveraging domain-specific knowledge from private data while maintaining privacy; “Heterogeneity,” examining variations in data, model, and computational resources across clients; “Security and Privacy,” focusing on defenses against malicious attacks and model theft; and “Efficiency,” highlighting the need for improvements in training, communication, and parameter efficiency. For each problem, we offer a clear mathematical definition on the objective function, analyze existing methods, and discuss the key challenges and potential solutions. This in-depth exploration aims to advance the theoretical foundations of FedFMs, guide practical implementations, and inspire future research to overcome these obstacles, thereby enabling the robust, efficient, and privacy-preserving FedFMs in various real-world applications.
Tao Fan 0002, Hanlin Gu, Xuemei Cao 0001, Chee Seng Chan, Qian Chen 0023, Yiqiang Chen 0001, Yihui Feng, Yang Gu 0001, Jiaxiang Geng, Bing Luo 0002, Shuoling Liu, WinKent Ong, Chao Ren 0006, Jiaqi Shao, Xiaoli Tang 0001, Hong Xi Tae, Yongxin Tong, Shuyue Wei 0001, Fan Wu 0006, Wei Xi 0003, Mingcong Xu, Xin Yang 0012, Jiangpeng Yan, Hao Yu 0023, Han Yu 0001, Xiaojin Zhang 0002, Zhenzhe Zheng 0001, Lixin Fan, Qiang Yang 0001
IEEE Trans. Knowl. Data Eng.8
2025 Grade-Skewed Domain Adaptation via Asymmetric Bi-Classifier Discrepancy Minimization for Diabetic Retinopathy Grading
abstract
Diabetic retinopathy (DR) is a leading cause of preventable low vision worldwide. Deep learning has exhibited promising performance in the grading of DR. Certain deep learning strategies have facilitated convenient regular eye check-ups, which are crucial for managing DR and preventing severe visual impairment. However, the generalization performance on cross-center, cross-vendor, and cross-user test datasets is compromised due to domain shift. Furthermore, the presence of small lesions and the imbalanced grade distribution, resulting from the characteristics of DR grading (e.g., the progressive nature of DR disease and the design of grading standards), complicates image-level domain adaptation for DR grading. The general predictions of the models trained on grade-skewed source domains will be significantly biased toward the majority grades, which further increases the adaptation difficulty. We formulate this problem as a grade-skewed domain adaptation challenge. Under the grade-skewed domain adaptation problem, we propose a novel method for image-level supervised DR grading via Asymmetric Bi-Classifier Discrepancy Minimization (ABiD). First, we propose optimizing the feature extractor by minimizing the discrepancy between the predictions of the asymmetric bi-classifier based on two classification criteria to encourage the exploration of crucial features in adjacent grades and stretch the distribution of adjacent grades in the latent space. Moreover, the classifier difference is maximized by using the forward and inverse distribution compensation mechanism to locate easily confused instances, which avoids pseudo-label bias on the target domain. The experimental results on two public DR datasets and one private DR dataset demonstrate that our method outperforms state-of-the-art methods significantly.
Yang Gu 0001, Shuai Guo 0001, Shijie Wen, Nianfeng Shi, Weiwei Dai, Yiqiang Chen 0001
IEEE Trans. Medical Imaging2
2024 Utilizing Attention-based Ensemble Mechanism to Identify Discriminative Feature Combinations for Sleep Staging
abstract
Deep learning significantly diminishes reliance on human experts for sleep staging, paving the way for fully automated sleep quality assessment in clinical and daily environments. However, existing methods mostly focus on the integration of novel networks through "black boxes" way to boost classification accuracy, often neglecting model interpretability and the identification of critical sleep features. To address these issues, we develop an innovative architecture utilizing Attention-based ensemble Mechanism (namely AeM) for effective sleep staging and further finding out discriminative feature combinations. AeM consists of three main parts, i.e., modality division and combination (MDC) module, individual feature extraction (IFE) module and attention-based weight and ensemble classification (AWEC) module. MDC module divides the raw input data by modality and recombines them along the channel dimension into new inputs. The IFE module employs multiple naive convolutional neural networks to extract features efficiently from each modality combination. Finally, the AWEC module integrates these features, effectively calculates the weight of each modal combination using the attention mechanism, and outputs the classification result. We evaluated AeM on three public-available datasets: ISRUC-S1, ISRUC-S3, and Sleep-EDF78. Experimental results demonstrate that AeM outperforms all comparative methods and can identify effective digital biomarkers in sleep staging tasks.
Yuwei Dai, Yingwei Zhang 0002, Yiqiang Chen 0001, Yang Gu 0001, Wei Zhang 0082
BIBM4
2024 Class-imbalanced Time Series Adaptation via Multi-Expert Consistency Entropy Minimization
abstract
Wearable-based time series recognition aims to infer behavioral classes based on time series signals collected by sensors. Domain Adaptation (DA) methods have significantly improved the performance on out-of-domain test classification tasks. However, highly imbalanced label distribution in real-world tasks hinders the practicality of DA methods. In this paper, we propose a Multi-Expert Consistency Entropy Minimization (MECEM) method for class-imbalanced domain adaptation. Through the incorporation of a multi-expert committee mechanism, we recalibrate decision boundaries in response to imbalanced source distributions, thereby mitigating bias towards minority classes and promoting fairness in committee decisions. Furthermore, we introduce a time-frequency dynamic augmentation mechanism, which simultaneously addresses domain shift on the target domain and dynamically adapts the augmentation library for the classification task. The MECEM method facilitates the propagation of consistency entropy during backpropagation, conveying both the reliability of pseudo-labels and the classification task relevance of augmentation techniques. Experiments across various tasks on activity recognition and Parkinson’s tremor grading demonstrate the superiority of our method.
Yang Gu 0001, Shuai Guo 0001, Weining Weng, Shijie Wen, Yiqiang Chen 0001
BIBM2
2024 Wafformer: Wave-Frequency Cross-Prediction and Alignment for Knowledge-Enhanced Sleep Representation
abstract
Electroencephalogram (EEG)-based sleep stage classification (SSC) identifies human sleep phases from EEG signals. With the integration of deep learning frameworks, automatic SSC has achieved notable accuracy. However, current SSC approaches encounter significant challenges: 1) Data challenge. The high costs associated with collecting and annotating sleep EEG signals result in a limited number of training samples, complicating supervised training for deep models. 2) Knowledge challenge. Incorporating clinical and human insights related to sleep into models for SSC remains challenging. To address these challenges, we propose Wafformer, a novel self-supervised framework to extract knowledge-enhanced sleep representation from EEG signals. Specifically, Wafformer leverages the Transformer architecture, employing self-attention and cross-attention encoders to extract wave-frequency features. We introduce innovative pretext tasks, such as cross-domain wave peaks and frequency-energy predictions, to integrate essential knowledge about salient waveforms associated with sleep states, complemented by temporal contrastive learning to align wave-frequency features. Extensive experiments on ISRUC, SleepEDF-SC, and SleepEDF-ST datasets demonstrate the superior performance of Wafformer on the sleep staging task. Besides, Wafformer can adaptively capture dynamic waveforms related to sleep stages, offering a comprehensive understanding of sleep phenomena.
Weining Weng, Yang Gu 0001, Shuai Guo 0001, Fulong Xiao, Yiqiang Chen 0001
BIBM2
2024 Information Retrieval Optimization for Non-Exemplar Class Incremental Learning
abstract
Existing non-example class-incremental learning (NECIL) methods usually utilize a combination strategy of replay mechanism and knowledge distillation. However, this combination strategy only focuses on the preservation of old information quantitatively, ignoring the preservation quality. When the old knowledge has wrong redundant information, catastrophic forgetting is more likely to occur. Therefore, obtaining adequate information without impurities as much as possible and removing invalid or even harmful information has become an effective solution to improve the performance of NECIL. This process is consistent with the information bottleneck (IB) theory. Thus, we propose a new NECIL method based on the IB framework. By using the different information obtained from the new and old class samples and the implicit knowledge in the teacher model training process, the error of harmful redundant information learned is eliminated. Specifically, we propose two optimization strategies that align with the two optimization processes of the information bottleneck. Firstly, we employ a pseudo-prototype selection mechanism that selectively incorporates pseudo-samples into the learning process of new and old categories, thus enhancing the distinction between new and old categories and diminishing the mutual information between the input and intermediate features. Secondly, we introduce an attention-based feature distillation method that regulates the distillation strength between feature pairs based on their similarity, thereby augmenting the mutual information between intermediate features and output prediction. Extensive experiments on three benchmarks demonstrate that the proposed method exhibits significant incremental performance improvements over existing methods.
Shuai Guo 0001, Yang Gu 0001, Yingwei Zhang 0002, Weining Weng, Weiwei Dai, Yiqiang Chen 0001
CIKM2
2024 Cascade Memory for Unsupervised Anomaly Detection
abstract
Unsupervised anomaly detection is to detect previously unseen rare samples without any prior knowledge about them. With the emergence of deep learning, many methods employ normal data reconstruction to train detection models, which is expected to yield relatively large errors when reconstructing anomalies. However, recent studies find that anomalies can be overgeneralized, resulting in reconstruction errors as small as normal samples. In this paper, we examine the anomaly overgeneralization problem and propose global semantic information learning. Normal and anomalous samples may share the same local feature such as textures, edges, and corners, but have separability at the global semantic level. To address this, we propose a novel cascade memory architecture designed to capture global semantic information in the latent space and introduce a configurable sparsification and random forgetting mechanism. Our proposed method achieves state-of-the-art experimental results on different public benchmarks, without the introduction of any additional auxiliary loss terms. The code is available at https://github.com/LiJiahao-Alex/Cascade-Memory.
Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan
ECAI4
2024 Learning Critically: Selective Self-Distillation in Federated Learning on Non-IID Data
abstract
Federated learning (FL) enables multiple clients to collaboratively train a global model while keeping local data decentralized. Data heterogeneity (non-IID) across clients has imposed significant challenges to FL, which makes local models re-optimize towards their own local optima and forget the global knowledge, resulting in performance degradation and convergence slowdown. Many existing works have attempted to address the non-IID issue by adding an extra global-model-based regularizing item to the local training but without an adaption scheme, which is not efficient enough to achieve high performance with deep learning models. In this paper, we propose a Selective Self-Distillation method for Federated learning (FedSSD), which imposes adaptive constraints on the local updates by self-distilling the global model’s knowledge and selectively weighting it by evaluating the credibility at both the class and sample level. The convergence guarantee of FedSSD is theoretically analyzed and extensive experiments are conducted on three public benchmark datasets, which demonstrates that FedSSD achieves better generalization and robustness in fewer communication rounds, compared with other state-of-the-art FL methods.
Yuting He 0008, Yiqiang Chen 0001, Xiaodong Yang 0005, Hanchao Yu, Yihua Huang 0002, Yang Gu 0001
IEEE Trans. Big Data6
2023 Auto-Augmentation Contrastive Learning for Wearable-based Human Activity Recognition
abstract
For low-semantic sensor signals from human activity recognition (HAR), contrastive learning (CL) is essential to implement novel applications or generic models without manual annotation, which is a high-performance self-supervised learning (SSL) method. However, CL relies heavily on data augmentation for pairwise comparisons. Especially for low semantic data in the HAR area, conducting good performance augmentation strategies in pretext tasks still rely on manual attempts lacking generalizability and flexibility. To reduce the augmentation burden, we propose an end-to-end auto-augmentation contrastive learning (AutoCL) method for wearable-based HAR. AutoCL is based on a Siamese network architecture that shares the parameters of the backbone and with a generator embedded to learn auto-augmentation. AutoCL trains the generator based on the representation in the latent space to overcome the disturbances caused by noise and redundant information in raw sensor data. The architecture empirical study indicates the effectiveness of this design. Furthermore, we propose a stop-gradient design and correlation reduction strategy in AutoCL to enhance encoder representation learning. Extensive experiments based on four wide-used HAR datasets demonstrate that the proposed AutoCL method significantly improves recognition accuracy compared with other SOTA methods.
Qingyu Wu, Jianfei Shen, Feiyi Fan, Yang Gu 0001, Chenyang Xu 0007, Yiqiang Chen 0001
BIBM4
2023 Letting Go of Self-Domain Awareness: Multi-Source Domain-Adversarial Generalization via Dynamic Domain-Weighted Contrastive Transfer Learning
abstract
Domain generalization (DG), which aims to learn a model that can generalize to an unseen target domain, has recently attracted increasing research interest. A major approach is to learn domain invariant representations to avoid greedily capturing all the correlations found in source domains caused by empirical risk minimization. Nevertheless, overly emphasizing learning of domain invariant representations might lead to learning overly-compressed domain invariant representations, causing confusion of different classes in a same domain. To address this limitation, we introduce a novel dynamic domain-weighted contrastive loss, which maximizes the subdomain differences between different classes especially those belonging to the same domain, while minimizing the average distance between the points of the convex hull of the aligned source domains. We propose Multi-source domain-adversarial generalization via dynamic domain-weighted Contrastive transfer learning (MsCtrl), a novel domain-adversarial generalization framework, which optimizes the distribution alignment of source and potential target subdomains in an adversarial manner under the “control” of the aforementioned contrastive loss. Extensive experiments based on real-world datasets demonstrate significant advantages of MsCtrl over existing state-of-the-art methods.
Yiqiang Chen 0001, Han Yu 0001, Yang Gu 0001, Shijie Wen, Shuai Guo 0001
ECAI4
2023 Pre-training A Prompt Pool for Vision-Language Model
abstract
Pre-trained vision-language(VL) model improves the performance in vision-language tasks, but requires a large amount of labeled data to apply the model to downstream tasks. It is challenging to obtain good results with limited training data. Prompt Tuning(PT), which freezes pre-train language models(PLMs) and only tunes soft prompts, provides an effective solution for adapting PLMs to downstream tasks. However, PT performs comparably with Full-model Tuning(FT) when the data are sufficient and performs much worse in few-shot settings, primarily due to the initialization of soft prompts. In this paper, we propose a new training framework Pre-train a Prompt Pool for Vision-Language Models called “P3VLM”, to give better initialization to PT. The objective of P3VLM is to optimize prompts selected by the visual feature, allowing the prompt to learn relevant knowledge in different data domains and store it in the prompt pool. Then we use diverse prompts in the prompt pool as the initialization of PT instead of a single pre-trained prompt. In the pre-train stage, the training samples are used to train a prompt pool by selecting prompts that best match the visual features. In the downstream datasets, the model selects the best matching prompts as the initialization for PT, freezes the PLM, and only tunes the prompts. Extensive experiments show that our method obtains strong performance on two image caption datasets in both zero-shot and few-shot scenarios.
Yang Gu 0001, Zhaohua Yang, Shuai Guo 0001, Huaqiu Liu, Yiqiang Chen 0001
IJCNN2
2023 Discriminative Domain Adaptation Network for Fine-grained Disease Severity Classification
abstract
Unsupervised Domain Adaptation (UDA) has shown promise in improving medical diagnosis tasks on the unlabeled target domain by utilizing rich labels on the source domain. However, in real medical scenarios, it is crucial to obtain a fine-grained classification of the disease, in order to support physicians in making accurate diagnoses and treatment plans for patients. Unfortunately, accurately labeling medical data at a fine-grained level is challenging because of the diversity of patients, and the variety of diseases. This leads to a difference between the given label and the true label, referred to as label bias. The existing UDA methods are based on the premise that the given label on the source domain is the true label, so it is easy to transfer the biased knowledge to the target domain in Fine-grained Unsupervised Domain Adaptation (FUDA), which leads to the poor performance of the model on the target domain. We find the key factor of FUDA is the sample with a large label bias (bias-sample) which is located near the decision boundary of adjacent fine-grained classes. To solve this problem, we propose Discriminative domain adaptation Network for Fine-grained classification (DNF) with Discriminative Cross Entropy (DCE) and Discriminative Local multi-kernel Maximum Mean Discrepancy (DLMMD). DNF employs two classifiers with different parameters to discriminate bias-samples and reduce their weight for classification, and use the outputs of the classifiers to modify the expectation of each class made by the given label to make it closer to the expectation of the true label. Therefore, DNF transfers non-bias knowledge from the source domain to the target domain. Experiments on Hand Tremor (HT) and Gait Freezing (GF) show that our approach outperforms SOTA.
Shijie Wen, Yiqiang Chen 0001, Shuai Guo 0001, Yang Gu 0001, Piu Chan
IJCNN5
2023 Time Series Adaptation Network for Sensor-Based Cross Domain Human Activity Recognition
abstract
Domain adaptation can apply knowledge learned from the source domain to the target domain by reducing data distribution discrepancy inter domains. However, existing domain adaptation algorithms do not do as well on sensor datasets as on image datasets because of the neglect of intra data distribution discrepancy. The long time of collecting a raw data segment on sensors will lead to a shift with time, and the shift distribution will change with the variety of sensors and wearing positions, causing the time series distribution discrepancy intra and inter domains. To solve this problem, we design a new model, Time Series Adaptation Network (TSAN), and a new loss, Time series Contrastive Loss (TCL). TSAN uses a siamese network and “pack” the samples divided from the same segment into the network. Furthermore, TCL is defined as the similarity of “unpack” network output, which leads the model to learn time-independent features. In particular, TSAN can be used as a plug-in to combine with existing domain adaptation algorithms, so the intra and inter distribution discrepancies can be considered simultaneously. We conduct extensive experiments with eight existing domain adaptation algorithms on sensor-based cross domain human activity recognition (HAR) tasks, including three Routine Activity Recognition (RAR) datasets and four Parkinson's tremor Detection (PD) datasets. The results show that all the existing algorithms are improved by an average of 5.4% (RAR), 2.2% (PD) with TSAN.
Shijie Wen, Yiqiang Chen 0001, Shuai Guo 0001, Yang Gu 0001, Piu Chan
IJCNN5
2022 KiCi: A Knowledge Importance Based Class Incremental Learning Method for Wearable Activity Recognition
abstract
Wearable-based human activity recognition (HAR) is commonly employed in real-world scenarios such as health monitoring, auxiliary diagnosis, etc. As implementing activity recognition is a daunting challenge in an open dynamic environment, incremental learning has become a common method to adapt to variable behavior patterns of users and create dynamic modeling in activity recognition. However, catastrophic forgetting is a significant challenge with incremental learning. This is contrary to our expectations of identifying new activity classes while remembering existing ones. To address this problem, we propose a knowledge importance-based class incremental learning method called KiCi and construct an incremental learning model based on the framework of self-iterative knowledge distillation for dynamic activity recognition. To eliminate the prediction bias of the teacher model on the old knowledge, we utilize the trained weights of previous incremental steps generated by the teacher model as the prior knowledge to obtain knowledge importance. Then use it to make the student model have a reasonable trade-off between old and new knowledge and mitigate catastrophic forgetting by avoiding negative transfer. We conduct extensive experiments on four public HAR datasets and our method consistently outperforms the existing state-of-the-art methods by a large margin.
Shuai Guo 0001, Yang Gu 0001, Shijie Wen, Yiqiang Chen 0001, Chunyu Hu 0001
CIKM2
2022 A wearable-HAR oriented sensory data generation method based on spatio-temporal reinforced conditional GANs
Yiqiang Chen 0001, Yang Gu 0001
Neurocomputing3
2021 A Collaborative Multi-modal Fusion Method Based on Random Variational Information Bottleneck for Gesture Recognition
Yang Gu 0001, Yiqiang Chen 0001, Jianfei Shen
MMM (1)1
2020 A Mobile Cloud Collaboration Fall Detection System Based on Ensemble Learning
abstract
Falls are one of the major causes of accidental or unintentional injury death worldwide. Therefore, this paper proposes a reliable fall detection algorithm and a mobile cloud collaboration system for fall detection. The algorithm is an ensemble learning method based on decision tree, named Fall-detection Ensemble Decision Tree (FEDT). The mobile cloud collaboration system is composed of three stages: 1) mobile stage: a light-weighted threshold method is used to filter out activities of daily livings (ADLs), 2) collaboration stage: TCP protocol is used to transmit data to cloud and meanwhile features are extracted in the cloud, 3) cloud stage: the model trained by FEDT is deployed to give the final detection result with the extracted features. Experiments show that the proposed FEDT outperforms the others' over 1-3% both on sensitivity and specificity and has superior robustness on different devices.
Tong Wu 0009, Yang Gu 0001, Yiqiang Chen 0001
ASSETS2
2020 A Lightweight Fully Convolutional Neural Network of High Accuracy Surface Defect Detection
Yiqiang Chen 0001, Yang Gu 0001, Jianquan Ouyang 0001, Ni Zeng
ICANN (2)3
2020 Multi-Layer Cross Loss Model for Zero-Shot Human Activity Recognition
Tong Wu 0009, Yiqiang Chen 0001, Yang Gu 0001, Zhanghu Zhechen
PAKDD (1)3
2020 Highly Fluent Sign Language Synthesis Based on Variable Motion Frame Interpolation
abstract
Sign Language Synthesis (SLS) is a domain-specific problem where multiple sign language words are stitched to generate a whole sentence in video, which serves to facilitate communications between the hearing-impaired people and healthy population. This paper presents a Variable Motion Frame Interpolation (VMFI) method for highly fluent SLS in scattered videos. Existing approaches for SLS mainly focus on mechanical virtual human technology, lacking high flexibility and natural effect. Also, the representative solutions to interpolate frames usually assume that the motion object moves at a constant speed which is not suitable for predicting the complex hand motion in frames of scattered sign language videos. To address the above issues, the proposed VMFI adopts acceleration to predict more accurate interpolated frames based on an end-to-end convolutional neural network. The framework of VMFI consists of variable optical flow estimation network and high-quality frame synthesis network that can approximate and fuse the intermediate optical flow to generate interpolated frames for synthesis. Experimental results on our realistic collected Chinese sign language dataset demonstrate that the proposed VMFI model achieves efficiency by performing better in PSNR (Peak Signal to Noise Ratio), SSIM (Structural Similarity) and MA (Motion Activity) and gets higher score in MOS (Mean Opinion Score) than other two representative methods.
Ni Zeng, Yiqiang Chen 0001, Yang Gu 0001, Yunbing Xing
SMC3
2018 SensoryGANs: An Effective Generative Adversarial Framework for Sensor-based Human Activity Recognition
abstract
This study focuses on improving the performance of human activity recognition when a small number of sensor data are available under some special practical scenarios and resource-limited environments, such as some high-risk projects, anomaly monitoring and actual tactical scenarios. The Human Activity Recognition (HAR) based on wearable sensors is an attractive research topic in machine learning and ubiquitous computing over the last few decades, and has extremely practicality in health surveillance, medical assistance, personalized services, etc. However, with the limitation of sensor sampling rate, sustainability, deployment, and other restricted conditions, it is difficult to collect enough and resultful sensor data anywhere, anytime. Therefore, the HAR based on wearable sensors always faces the challenges of the low-data regime under some practical scenarios, which leads to a low accuracy of activity recognition and needs to be solved urgently. Currently, the Generative Adversarial Networks (GANs) provide a powerful method for training resultful generative models that could generate very convincing verisimilar images. The framework of GANs and its variants shed many lights on improving the performance of HAR. In this paper, we propose a new generative adversarial networks framework called SensoryGANs that can effectively generate available sensor data used for HAR. To the best of our knowledge, SensoryGANs is the first unbroken generative adversarial networks applied in generating sensor data in the HAR research field. Firstly, we tried exploring and devising three activity-special GANs models for three human daily activities. Secondly, these specific models are trained with the guidance of unbroken vanilla GANs. Thirdly, the trained generators from adversarial optimization process are used to generate synthetic sensor data. Finally, the synthetic sensor data from SensoryGANs are used to enrich the original authentic sensor datasets, which can improve the performance of target activity recognition model. Meanwhile, we propose three visual evaluation methods for assessing synthetic sensor data produced by the trained generators in SensoryGANs models. Experimental results show that SensoryGANs models have the capability of capturing the implicit distribution of real sensor data of human activity, and then the synthetic sensor data generated by SensoryGANs models Have a potential for improving human activity recognition.
Yiqiang Chen 0001, Yang Gu 0001, Yunlong Xiao, Haonan Pan
IJCNN3
2018 FSELM: fusion semi-supervised extreme learning machine for indoor localization with Wi-Fi and Bluetooth fingerprints
Xinlong Jiang, Yiqiang Chen 0001, Junfa Liu, Yang Gu 0001, Lisha Hu
Soft Comput.4
2016 Feature Adaptive Online Sequential Extreme Learning Machine for lifelong indoor localization
Xinlong Jiang, Junfa Liu, Yiqiang Chen 0001, Dingjun Liu, Yang Gu 0001, Zhenyu Chen 0003
Neural Comput. Appl.5
2015 Semi-supervised deep extreme learning machine for Wi-Fi based localization
Yang Gu 0001, Yiqiang Chen 0001, Junfa Liu, Xinlong Jiang
Neurocomputing1
2014 Constraint Online Sequential Extreme Learning Machine for lifelong indoor localization system
abstract
As an important technology in LBS (Location Based Services) field, Wi-Fi based indoor localization suffers signal fluctuation problem which prevents lifelong and high performance running. With the fluctuation of wireless signal over time, fingerprints collected at the same location become different; therefore existing model cannot fit the new collected data well, which decreases the localization accuracy. In this paper, a novel indoor localization method COSELM (Constraint Online Sequential Extreme Learning Machine) is proposed, utilizing incremental data to update the old model and overcome the fluctuation problem. The performance of COSELM is validated in real Wi-Fi indoor environment. Compared with OSELM, it can improve more than 5% localization accuracy on average; and in contrast to batch learning, COSELM can save more than 50% time consumption.
Yang Gu 0001, Junfa Liu, Yiqiang Chen 0001, Xinlong Jiang
IJCNN1
2014 TOSELM: Timeliness Online Sequential Extreme Learning Machine
Yang Gu 0001, Junfa Liu, Yiqiang Chen 0001, Xinlong Jiang, Hanchao Yu
Neurocomputing1