EDBT 2026 Demo / reviewers in the wild / expert
Tong Zhang 0015
dblp:07/4227-15
· DBLP profile ↗
139ranked-venue papers
18as first author
120since 2021 · last 2026
0000-0002-7025-6365ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 8 first-author · 65 since 2021Applied, interdisciplinary, general and emerging computing · 46 · 5 first-author · 35 since 2021Human-computer interaction and ubiquitous computing · 25 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 13 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An accurate and efficient online broad learning system for data stream classification
Chunyu Lei, Guang-Ze Chen, C. L. Philip Chen, Tong Zhang 0015 |
Sci. China Inf. Sci. | 4 |
| 2026 | A dynamic graph attention network for traffic flow prediction based on multi-domain features fusion
Nan Ma 0001, Qinfen Wang, Shi-Yuan Han, Jie Liu 0002, Jin Zhou 0003, Tong Zhang 0015, C. L. Philip Chen |
Expert Syst. Appl. | 7 |
| 2026 | Broad fractional order fuzzy system for noise and outlier resistant data classification
Tong Zhang 0015, Tao Zhang 0103, C. L. Philip Chen, Yuzhong Sun |
Expert Syst. Appl. | 2 |
| 2026 | Simplified implementation and universal approximation of multi-input single-output hierarchical fuzzy systems with correction factors
Linlin Guo, Changle Sun, Shi-Yuan Han, Jin Zhou 0003, Yalu Li, Tong Zhang 0015, C. L. Philip Chen |
Fuzzy Sets Syst. | 6 |
| 2026 | Collaborative multi-view fuzzy clustering based on Gaussian mixture model
Shi-Yuan Han, Jin Zhou 0003, C. L. Philip Chen, Tong Zhang 0015, Yuehui Chen, Lin Wang 0004, Tao Du 0002 |
Neurocomputing | 5 |
| 2026 | Unsupervised representation learning for anomaly detection in power grid communications via spatio-temporal graph neural networks
Tong Zhang 0015, Zongyan Zhang, Weijie Qiu, Tingwen Yu, C. L. Philip Chen |
Neurocomputing | 2 |
| 2026 | Broad learning system based on mixture correntropy criterion
Jiajun Liu 0012, Tao Zhang 0103, Tong Zhang 0015, C. L. Philip Chen |
Neurocomputing | 4 |
| 2026 | IDFG: Information-Regularized Diversity-Fidelity Graph for EEG emotion recognition
Bianna Chen, C. L. Philip Chen, Tong Zhang 0015 |
Knowl. Based Syst. | 3 |
| 2026 | QuAC: Quality-aware chinese pre-training proactive defense framework
Zhongyi Deng, C. L. Philip Chen, Yile Chen 0004, Tong Zhang 0015 |
Pattern Recognit. | 4 |
| 2026 | Contrastive adversarial tuning: Enhancing discriminability and robustness of LLMs for emotion recognition in conversation
Kankan Lan, C. L. Philip Chen, Zongyan Zhang, Tong Zhang 0015 |
Pattern Recognit. | 4 |
| 2026 | CRIA: A cross-view interaction and instance-adapted pre-training framework for generalizable EEG representations
Puchun Liu, C. L. Philip Chen, Tong Zhang 0015 |
Pattern Recognit. | 4 |
| 2026 | Test-time Adaptive Hierarchical Co-enhanced Denoising Network for reliable multimodal classification
Shu Shen, C. L. Philip Chen, Tong Zhang 0015 |
Pattern Recognit. | 3 |
| 2026 | F2FNet: An Efficient Affective State Analysis Network Reconstructing EEG From Few-Channel to Full-ChannelabstractDue to the prohibitive costs and lack of portability associated with full-channel devices, portable miniature EEG devices with only a few electrodes (few-channel) are more suitable for widespread deployment in consumer applications. However, affective state analysis with few-channel EEG presents a significant challenge, as these devices only capture signals from partial brain regions. Existing approaches designed for few-channel devices do not adequately account for the critical role of whole-brain connectivity and suffer from the impact of noise in affective state analysis. To address the challenges, we propose an efficient affective state analysis Network which reconstructs EEG from Few-channel To Full-channel (F2FNet). This method aims to leverage knowledge from full-channel EEG and exclude irrelevant information to enhance the affective state analysis performance of few-channel EEG based on whole-brain connectivity patterns. To ensure information validity and consistency during the reconstruction process, we introduce mechanisms for information compression and control. Additionally, we perform alignment of compressed features within the VAD emotional space to ensure that the model's understanding and extraction of emotional features conform to the definitions of affective theories. Extensive experiments on basic emotion recognition and depression detection demonstrate the efficacy of the proposed method. Zihua Xu, Tianjun Li, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 4 |
| 2026 | Multimodal Affect Perception With Large Language Model Enhancement NetworkabstractMultimodal Sentiment Analysis (MSA) plays a vital role in understanding emotional content from social media and multimedia data. However, existing methods often rely on large-scale labeled datasets, leading to high annotation costs and poor adaptability. They also suffer from modality imbalance and suboptimal feature fusion. To address these issues, we propose MapleNet—a Multimodal Affect Perception framework enhanced by Large Language Models. MapleNet integrates a prototype-guided fusion strategy and a dynamic modality balancing mechanism to improve alignment and collaboration between text and image features. Specifically, a shared-space encoder combined with prompt optimization ensures semantic consistency across modalities. Within the prototype learning framework, the model dynamically adjusts modalityspecific learning by aligning features with class prototypes, thus mitigating imbalance and uncovering complementary affective cues. In addition, MapleNet employs a similaritybased sample retrieval module to construct contextual prompts, enriching sentiment understanding in few-shot settings. Experiments on six benchmark datasets show that MapleNet consistently outperforms state-of-the-art methods, especially under few-shot conditions, achieving superior accuracy and generalization. The relevant code is available at https://github.com/YFanLuo/MapleNet. Kaixiang Yang 0001, Zongyan Zhang, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 5 |
| 2026 | MCGC-Net: Multi-Scale Controllable Graph Convolutional Network on Music Emotion Recognition
Tong Zhang 0015, Xueyue Yang, C. L. Philip Chen |
IEEE Trans. Affect. Comput. | 1 |
| 2026 | DCAL: Dual-Temporal Contrastive Adversarial Learning for Cross-Subject EEG Sleep Analysis
Yuxin Yin, Bianna Chen, Zihua Xu, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | MEDS-Net: Meta-Evidential Dual-Stream Network for Multimodal Few-Shot Driver Fatigue DetectionabstractFatigued driving is a major societal risk factor in the road safety system, affecting the efficiency of the entire transportation network and public safety. Fatigue detection methods based on multimodal physiological signals can objectively reflect a driver’s neurocognitive state, offering higher monitoring reliability. However, existing methods face two main challenges: 1) the high cost of annotating high-quality physiological signals leads to challenges in learning with limited data; and 2) the signal-to-noise ratio of multimodal physiological signals is variable, leading to imbalances in the confidence in the decision between modalities. To address these issues, this article proposes the meta-evidential dual-stream network (MEDS-Net), which integrates a meta-learning framework with evidential deep learning to establish a solution for driver fatigue detection for the first time. Specifically, MEDS-Net employs a dual-stream evidence generation network to extract heterogeneous features from multimodal physiological signals and uses a Dirichlet distribution to quantify cognitive uncertainty for each modality. Additionally, MEDS-Net introduces a meta-learning framework to enhance the model’s ability to handle data scarcity through the design of task simulation mechanisms and a two-layer optimization strategy. Extensive experiments on the SEED-VIG dataset demonstrate that MEDS-Net outperforms current state-of-the-art methods in terms of effectiveness and generalizability. Furthermore, MEDS-Net can precisely quantify the importance of global features in explaining modal contributions and average activation patterns. Yuankang Fu, Xin-Rong Gong, Kaixiang Yang 0001, Tong Zhang 0015, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Social Power Evolution of Multiple DeGroot Individuals With Centralized MediaabstractIn this article, the social power evolution problem is investigated for a social network with a centralized media and multiple DeGroot individuals, where the centralized media impacts the evolution of individuals’ opinions through the centralization parameter at the broadcast moment, while the network topology among the individuals is influenced by the relative interaction matrix. Then, the convergence of the corresponding opinion dynamics on the time scale is derived by discussing three distinct initial social powers. By integrating the reflected appraisal mechanism, a social power evolution model with the centralized media is established, which is essentially a nonlinear mapping. Based on the Jacobian matrix of this nonlinear mapping, it is proved that both the social powers of the centralized media and DeGroot individuals can converge provided that the centralization parameter exceeds a certain threshold; furthermore, a lower bound for the centralized media’s final social power is estimated, which helps to demonstrate that the centralized media possesses the greatest social power within the whole network. Additionally, concerning individuals’ final social powers, all individuals are first divided into three categories, and a sufficient condition is presented to ensure that the balanced individual has the greatest social power except for the centralized media. Finally, the obtained results are illustrated by a numerical example. Hong-xiang Hu, Jialing Zhou, Yun Chen 0008, Tong Zhang 0015, Guanghui Wen |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Unified Multimodal Representation Learning With Prototype Calibration for Cross-Subject Few-Shot Emotion RecognitionabstractMultimodal cross-subject emotion recognition has gained increasing attention due to the complementary nature of physiological and behavioral signals, such as electroencephalography (EEG) and eye movements (EM). However, modality heterogeneity and subject variability remain major obstacles, especially in few-shot scenarios. Existing approaches often combine increasingly complex architectures with unsupervised domain adaptation to address these issues, but they typically require large amounts of unlabeled target data and struggle to generalize when samples are scarce. To overcome these limitations, this article proposes a unified multimodal representation learning (UMRL) framework with prototype calibration under the few-shot paradigm. Specifically, the UMRL strategy is designed from an optimization perspective rather than architectural complexity to learn a compact and discriminative joint representation. It encourages compact representations by jointly minimizing cross-modal inconsistency and penalizing modality dominance, ensuring that complementary cues are effectively integrated in a unified embedding space. Nevertheless, emotional prototypes constructed from a few labeled samples of novel subjects are often unstable and sensitive to subject-dependent variability. To further alleviate this issue, this article introduces an emotional prototype calibration mechanism that aligns novel emotional prototypes with semantic priors learned from seen subjects to suppress noise and improve cross-subject generalization. Extensive experiments on multiple multimodal emotion recognition benchmarks with EEG and EM demonstrate that the proposed framework achieves state-of-the-art performance and delivers consistent robustness under few-shot conditions. Haiqi Liu, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2026 | SAIL: Seeking Attribute in Language to Improve Chinese PretrainingabstractImbalanced exposure of Chinese characters during pretraining often leads to inconsistent performance across downstream tasks. This problem stems primarily from the randomness of the masking strategy and is further aggravated by the semantic granularity mismatch between Chinese and English. To address these challenges, we propose the seeking attribute in language (SAIL) scheme to enhance Chinese language modeling. SAIL constructs a hierarchical attribute pyramid grounded in the rich semantics of Chinese and comprises three core tasks: masked language modeling (MLM), masked character restructuring (MCR), and masked quality evaluation (MQE). MQE employs a Chinese association graph to dynamically assess the quality of randomly masked samples and uses this assessment as a supervisory signal during pretraining. MCR integrates character-level structural information to mitigate the imbalance caused by varying word frequencies. To facilitate multilevel semantic interaction, MLM not only models word-level semantics but also bridges the character-level structural semantics from MCR and the sentence-level semantics from MQE. In turn, MCR and MQE reevaluate and enhance the performance of MLM from the lower and upper layers of the hierarchical pyramid, respectively. Experimental results across a range of Chinese language understanding tasks, including multidomain comprehension, fine-grained entity recognition, and diverse text classification, demonstrate that our model consistently achieves superior average performance. These findings show the generalization capability and transferability of the SAIL framework in modeling complex Chinese semantics. Tong Zhang 0015, Zhongyi Deng, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2026 | ACM-GNN: Adaptive Cluster-Oriented Modularity Graph Neural Network for EEG Depression DetectionabstractMajor depressive disorder (MDD) is typically accompanied by varying topological dynamics across brain network modularity due to the influence of time-variant and subject-specific factors. Current works primarily characterize the electroencephalograms (EEG) topological relationships based on prior predefined brain regions. However, this predefined strategy cannot dynamically fit to different individuals, which may affect the adaptability of unseen individual for depression detection. This article proposes an adaptive cluster-oriented modularity graph neural network (ACM-GNN) to enhance the adaptability of individual topological interaction for depression detection. Specifically, a cluster-oriented modularity construction (CMC) module dynamically clusters EEG channels into different brain regions based on channel-pairs contrastive learning. It can adaptively construct brain modularity to fit different individual instances. Furthermore, a modularity graph interaction learning (MGIL) module performs multilayer graph information interaction between EEG globality and modularity levels. In this way, more powerful hierarchical information can be integrated by further aggregating representations at different levels. Experiments on two public datasets, MODMA and PRED+CT, demonstrate that the proposed method outperforms the state-of-the-art EEG depression detection methods. Finally, investigations on brain activities reveal the importance of dynamic modular relations for depression detection. Tong Zhang 0015, Zihua Xu, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2026 | DPS-Net: Direction-Aware Pseudo-Stereo Network for Accurate Road Surface ReconstructionabstractThe geometry of road surfaces plays a critical role in the performance of autonomous driving systems. Consequently, achieving accurate and efficient road surface reconstruction (RSR) is of paramount importance. However, due to the inherent effects of perspective projection, distant regions often exhibit geometric distortions and a long-tailed distribution, which pose significant challenges to existing reconstruction methods. To address these issues, we propose a novel framework, termed Direction-aware Pseudo-Stereo Road Reconstruction Network (DPS-Net), which incorporates two lightweight and plug-and-play modules: Direction-Aware Feature Enhancement (DFE) module and Pseudo-Stereo Fusion (PSF) module. The DFE module is designed to enhance the perception of sparse and geometry-invariant features by integrating directional context, while the PSF module captures global dependencies across spatial and channel dimensions through pseudo-stereo fusion. Both modules are constructed with an emphasis on maintaining low computational complexity. We conducted extensive experiments on the public RSRD dataset to evaluate the effectiveness and superiority of our proposed method. The code is available at https://github.com/yidanyi/DPS-Net. Shi-Yuan Han, Yidan Pei, Rui Wang 0199, Tong Zhang 0015, C. L. Philip Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Prompts Libra: Enhanced Image Outpainting Diffusion Model With Balanced Bimodal GuidanceabstractImage outpainting, a challenging generative task, has advanced significantly with the introduction of text-to-image diffusion models (DM). Despite these advances, DM-based methods frequently encounter the phenomenon in which one modal takes precedence over another, causing the image to be over-guided. Current research relies on manual hyperparameters to achieve bimodal balance. To reduce reliance, Prompt Libra is proposed to automatically balance bimodal prompts during inference and enhance extrapolated images. Given the variation of bimodal cross-attention during DM denoising, we create an adaptive bimodal attention module via attention maps. Furthermore, we design a classifier-free guidance computation based on masked images to improve the semantic control of the masked part and enhance the quality of images. Finally, we propose a semantic transformer to address the problem of quality degradation caused by incomplete prompts. It extracts limited semantics from the source images, which is suitable for scenarios lacking text prompts. Experimental results demonstrate that our method generates images that achieve the state-of-the-art effect on several image quality evaluation metrics while maintaining the image and text prompts in balance. Zongyan Zhang, C. L. Philip Chen, Zepeng Su, Tong Zhang 0015 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Dynamic Extraction of Subdialogs for Dialog Emotion Recognition
Zhenyu Yang 0002, Zhibo Zhang 0009, Yuhu Cheng 0001, Tong Zhang 0015, Xuesong Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient DeploymentsabstractLarge Language Models (LLMs) have advanced rapidly but face significant memory demands.While quantization could alleviate the memory-bound issue, current methods typically require lengthy training to recover accuracy under low bit width.In that circumstance, deployment across scenarios with different resource constraints necessitates repeated training, amplifying the issue of protracted training.It is beneficial to train a once-for-all (OFA) supernet capable of offering optimal subnets for downstream applications.To extend the oncefor-all setting to LLMs, we decouple the shared weights to mitigate the interference and integrate Low-Rank adapters to enhance training efficiency.Furthermore, it is observed that there is an imbalance in the allocation of training resources due to traditional uniform sampling.A non-parametric scheduler is introduced to adjust the sampling rate for each quantization configuration, thereby achieving a more balanced allocation among subnets with varying demands.We validate the approach on LLaMA families and Mistral on downstream evaluation, demonstrating high performance while significantly reducing deployment time faced with multiple scenarios.1 Ke Yi 0003, Heng Chang, Tong Zhang 0015, Jia Li 0009 |
ACL (1) | 5 |
| 2025 | A Parameter-Efficient and Fine-Grained Prompt Learning for Vision-Language ModelsabstractCurrent vision-language models (VLMs) understand complex vision-text tasks by extracting overall semantic information from largescale cross-modal associations.However, extracting from large-scale cross-modal associations often smooths out semantic details and requires large computations, limiting multimodal fine-grained understanding performance and efficiency.To address this issue, this paper proposes a detail-oriented prompt learning (DoPL) method for vision-language models to implement fine-grained multi-modal semantic alignment with merely 0.25M trainable parameters.According to the low-entropy information concentration theory, DoPL explores shared interest tokens from text-vision correlations and transforms them into alignment weights to enhance text prompt and vision prompt via detail-oriented prompt generation.It effectively guides the current frozen layer to extract fine-grained text-vision alignment cues.Furthermore, DoPL constructs detail-oriented prompt generation for each frozen layer to implement layer-by-layer localization of finegrained semantic alignment, achieving precise understanding in complex vision-text tasks.DoPL performs well in parameter-efficient finegrained semantic alignment with only 0.12% tunable parameters for vision-language models.The state-of-the-art results over the previous parameter-efficient fine-tuning methods and full fine-tuning approaches on six benchmarks demonstrate the effectiveness and efficiency of DoPL in complex multi-modal tasks. Yongbin Guo, Shuzhen Li, Zhulin Liu, Tong Zhang 0015, C. L. Philip Chen |
ACL (1) | 4 |
| 2025 | Incongruity-aware Tension Field Network for Multi-modal Sarcasm DetectionabstractMulti-modal sarcasm detection (MSD) identifies sarcasm and accurately understands users' real attitudes from text-image pairs.Most MSD researches explore the incongruity of textimage pairs as sarcasm information through consistency preference methods.However, these methods prioritize consistency over incongruity and blur incongruity information under their global feature aggregation mechanisms, leading to incongruity distortions and model misinterpretations.To address the above issues, this paper proposes a pioneering inconsistency preference method called incongruityaware tension field network (ITFNet) for multimodal sarcasm detection tasks.Specifically, ITFNet extracts effective text-image feature pairs in fact and sentiment perspectives.It then constructs a fact/sentiment tension field with discrepancy metrics to capture the contextual tone and polarized incongruity after the iterative learning of tension intensity, effectively highlighting incongruity information during such inconsistency preference learning.It further standardizes the polarized incongruity with reference to contextual tone to obtain standardized incongruity, effectively implementing instance standardization for unbiased decision-making in MSD.ITFNet performs well in extracting salient and standardized incongruity through an incongruity-aware tension field, significantly tackling incongruity distortions and cross-instance variance.Moreover, ITFNet achieves state-of-the-art performance surpassing LLaVA1.5-7B with only 17.3M trainable parameters, demonstrating its optimal performance-efficiency in multi-modal sarcasm detection tasks. Jiecheng Zhang, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
ACL (1) | 4 |
| 2025 | Scaling Mesh Generation via Compressive TokenizationabstractWe propose a compressive yet effective mesh tokenization, Blocked and Patchified Tokenization (BPT), facilitating the generation of meshes exceeding 8k faces. BPT compresses mesh sequences by employing block-wise indexing and patch aggregation, reducing their length by approximately 75% compared to the vanilla coordinate sequences. This compression milestone unlocks the potential to utilize mesh data with significantly more faces, thereby enhancing detail richness and improving generation robustness. Empowered with the BPT, we have built a foundation mesh generative model training on scaled mesh data to support flexible control for point clouds and images. Our model demonstrates the capability to generate meshes with intricate details and accurate topology, achieving SoTA performance on mesh generation and reaching the level for direct product usage. Haohan Weng, Zibo Zhao 0001, Biwen Lei, Xianghui Yang, Jian Liu 0036, Zeqiang Lai, Zhuo Chen 0054, Jie Jiang 0015, Chunchao Guo, Tong Zhang 0015, Shenghua Gao, C. L. Philip Chen |
CVPR | 11 |
| 2025 | An Orthogonal High-Rank Adaptation for Large Language ModelsabstractLow-rank adaptation (LoRA) efficiently adapts LLMs to downstream tasks by decomposing LLMs' weight update into trainable low-rank matrices for fine-tuning.However, the random low-rank matrices may introduce massive taskirrelevant information, while their recomposed form suffers from limited representation spaces under low-rank operations.Such dense and choked adaptation in LoRA impairs the adaptation performance of LLMs on downstream tasks.To address these challenges, this paper proposes OHoRA, an orthogonal high-rank adaptation for parameter-efficient fine-tuning on LLMs.According to the information bottleneck theory, OHoRA decomposes LLMs' pre-trained weight matrices into orthogonal basis vectors via QR decomposition and splits them into two low-redundancy high-rank components to suppress task-irrelevant information.It then performs dynamic rank-elevated recomposition through Kronecker product to generate expansive task-tailored representation spaces, enabling precise LLM adaptation and enhanced generalization.OHoRA effectively operationalizes the information bottleneck theory to decompose LLMs' weight matrices into low-redundancy high-rank components and recompose them in rank-elevated manner for more task-tailored representation spaces and precise LLM adaptation.Empirical evaluation shows OHoRA's effectiveness by outperforming LoRA and its variants and achieving comparable performance to full fine-tuning with only 0.0371% trainable parameters. Xin Zhang 0100, Guang-Ze Chen, Shuzhen Li, Zhulin Liu, C. L. Philip Chen, Tong Zhang 0015 |
EMNLP | 6 |
| 2025 | Enhancing Generalized EEG Classification with Decomposed Statistics-diverse Feature AugmentationabstractLearning a generalized EEG representation under limited data and subject variability is a long-standing challenge. Most studies utilized data augmentation to extend the distribution of training data, which may hinder the diversity of augmented samples to cover more subject variability. In this paper, we propose a decomposed statistics-diverse feature augmentation (DSFA) framework for generalized EEG learning. The wavelet-based decomposed representation module decomposes signals into approximation and detail features, thus deriving semantics from original signals into approximation features to avoid over-transformation. The statistics-diverse feature augmentation module augments features to extend beyond the original feature space by manipulating the statistics of features. Extensive experiments on three datasets demonstrate that our approach achieves state-of-the-art performance in different classification tasks. Our repository is public at https://github.com/1940653868/DSFA. C. L. Philip Chen, Bianna Chen, Tong Zhang 0015 |
ICASSP | 4 |
| 2025 | TimeBooth: Disentangled Facial Invariant Representation for Diverse and Personalized Face Aging
Zepeng Su, Zhulin Liu, Zongyan Zhang, Tong Zhang 0015, C. L. Philip Chen |
ICCV | 4 |
| 2025 | Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inferenceabstractLarge language models have demonstrated promising capabilities upon scaling up parameters. However, serving large language models incurs substantial computation and memory movement costs due to their large scale. Quantization methods have been employed to reduce service costs and latency. Nevertheless, outliers in activations hinder the development of INT4 weight-activation quantization. Existing approaches separate outliers and normal values into two matrices or migrate outliers from activations to weights, suffering from high latency or accuracy degradation. Based on observing activations from large language models, outliers can be classified into channel-wise and spike outliers.
In this work, we propose Rotated Runtime Smooth (**RRS**), a plug-and-play activation smoother for quantization, consisting of Runtime Smooth and the Rotation operation. Runtime Smooth (**RS**) is introduced to eliminate **channel-wise outliers** by smoothing activations with channel-wise maximums during runtime. The Rotation operation can narrow the gap between **spike outliers** and normal values, alleviating the effect of victims caused by channel-wise smoothing.
The proposed method outperforms the state-of-the-art method in the LLaMA and Qwen families and improves WikiText-2 perplexity from 57.33 to 6.66 for INT4 inference. Ke Yi 0003, Zengke Liu, Jianwei Zhang 0012, Tong Zhang 0015, Junyang Lin, Jingren Zhou 0001 |
ICLR | 5 |
| 2025 | PivotMesh: Generic 3D Mesh Generation via Pivot Vertices GuidanceabstractGenerating compact and sharply detailed 3D meshes poses a significant challenge for current 3D generative models. Different from extracting dense meshes from neural representation, some recent works try to model the native mesh distribution (i.e., a set of triangles), which generates more compact results as humans crafted. However, due to the complexity and variety of mesh topology, most of these methods are typically limited to generating meshes with simple geometry. In this paper, we introduce a generic and scalable mesh generation framework PivotMesh, which makes an initial attempt to extend the native mesh generation to large-scale datasets. We employ a transformer-based autoencoder to encode meshes into discrete tokens and decode them from face level to vertex level hierarchically. Subsequently, to model the complex typology, our model first learns to generate pivot vertices as coarse mesh representation and then generate the complete mesh tokens with the same auto-regressive Transformer. This reduces the difficulty compared with directly modeling the mesh distribution and further improves the model controllability. PivotMesh demonstrates its versatility by effectively learning from both small datasets like Shapenet, and large-scale datasets like Objaverse and Objaverse-xl. Extensive experiments indicate that PivotMesh can generate compact and sharp 3D meshes across various categories, highlighting its great potential for native mesh modeling. Haohan Weng, Tong Zhang 0015, C. L. Philip Chen |
ICLR | 3 |
| 2025 | DMDM: Photorealistic Face Age Transformation by Dual-Modal Collaborative Attention using Diffusion ModelsabstractIn this work, we focus on enhancing the realism of face age transformation. Previous methods often relied on style transfer strategies or text-attention manipulation, which frequently result in undesirable artifacts or distorted facial defects. We propose DMDM, an age transformation method by Dual-Modal collaborative attention using Diffusion Models. Specifically, we introduce a collaborative text-image attention based editing method that balances age semantic control and visual harmony of generated face. To further refine fidelity, we propose the Softer Image Attention Injection Mechanism, which dynamically integrates image guidance. Additionally, we design a query image set to retrieve relevant images accroding to source face and other attributes for image guidance. Finally, Age-aware Face Restoration module is proposed to enhance high-frequency details according to age through a cascaded refinement pipeline. Extensive experiment demonstrates that DMDM achieves state-of-the-art performance, especially in visual quality. Zepeng Su, Zhulin Liu, Zongyan Zhang, Tong Zhang 0015, C. L. Philip Chen |
ICME | 4 |
| 2025 | DiBAN: Dual-Drive Broad Attentive Network for Speech Emotion RecognitionabstractData-Knowledge dual-driven fashion can enhance model performance by complementing data-driven basis with expert knowledge. However, cutting-edge works in speech emotion recognition (SER) primarily evolve in data-driven training, failing to incorporate prior knowledge to form a closed loop and posing an obstacle to capture task-specific details when used independently. In this paper, we propose a Dual-Drive Broad Attentive Network (DiBAN) to achieve comprehensive emotional learning for SER. Specifically, the Dual-Drive Emotional Modeling module incorporates handcrafted extractor, pre-trained model and tailored base models, to conduct integral emotional modeling. Subsequently, the Multi-Model Attention-Aware Learning module is designed to refine the data-knowledge emotional disparities based on the attention-enhanced entropy loss. Finally, the Broad Adaptive Decision Fusion module performs adaptive fusion of emotional decisions from different drives. Extensive experiments on seven SER corpora demonstrate that DiBAN achieves significant improvements over the base models and outperforms comparative methods, fully showcasing its superiority. Gongli Zhang, C. L. Philip Chen, Tong Zhang 0015, Zhulin Liu, Xiaoman Hu, Bianna Chen |
ICME | 3 |
| 2025 | DFMU: Distribution-based Framework for Modeling Aleatoric Uncertainty in Multimodal Sentiment AnalysisabstractIn Multimodal Sentiment Analysis (MSA), data noise arising from various sources can lead to uncertainty in Aleatoric Uncertainty (AU), significantly impacting model performance. Current efforts to address AU have insufficiently explored its sources. They primarily focus on modeling noise rather than implementing targeted modeling based on its origin. Consequently, these approaches struggle to effectively mitigate the influence of AU, resulting in sustained limitations in model performance. Our research identifies that the AU primarily stems from two problems: subjective bias in the annotation process and the complex set relationships of sentiment features. To specifically address them, we propose DFMU, a Distribution-based Framework for Modeling Aleatoric Uncertainty, which incorporates an uncertainty modeling block capable of encoding uncertainty distributions and adaptively adjusting optimization objectives. Furthermore, we introduce distribution-based contrastive learning with sentiment words replacement to better capture the complex relationships among features. Extensive experiments on three public MSA datasets, i.e., MOSI, MOSEI, and SIMS, demonstrate that the proposed model maintains robust performance even under high noise conditions and achieves state-of-the-art results on these popular datasets. Tingrui Shen, Xin-Rong Gong, Tong Zhang 0015 |
IJCAI | 5 |
| 2025 | InfoGA: Enhancing Generalized EEG Emotion Recognition via Information-Aware Graph AugmentationabstractLimited EEG data and subject variability pose significant challenges to the generalization of EEG-based emotion recognition. Most existing approaches augment EEG data using deterministic methods, often neglecting to ensure both diversity and fidelity in the generated samples. This oversight leads to insufficient domain diversity and emotional semantic information for a generalized model independent of individuals. This paper proposes an Information-Aware Graph Augmentation (InfoGA) framework for generalized EEG emotion recognition. The graph uncertainty augmentation module augments both the connectivity and features of EEG graphs by modeling statistical uncertainty, enabling the model to simulate domain shifts and improve generalizability against subject variability. Additionally, two information-aware constraints are introduced to ensure diversity and fidelity in the augmented EEG graphs. The graph diversity constraint enriches the emotional knowledge of the augmented graphs, while the graph fidelity constraint preserves their emotional semantic fidelity by integrating consistency learning with supervised learning. Extensive experiments on three public EEG emotion datasets, i.e., SEED, SEED-IV, and SEED-V, demonstrate that InfoGA achieves superior generalizability compared to baseline methods. Bianna Chen, C. L. Philip Chen, Tong Zhang 0015 |
SMC | 3 |
| 2025 | MoRa: Multi-Graph Orthogonal Representation Adaptation Network for Cross-Subject EEG Emotion RecognitionabstractThe redundancy of emotion-agnostic features and inter-subject variability significantly undermine the adaptability of EEG-based emotion recognition models. Previous studies have largely overlooked the influence of emotional responses at both intra- and inter-region levels across individuals. Consequently, the emotion-aware representations exhibit substantial redundancy and limited adaptability to cross-subject variability. To address these issues, this paper proposes a multi-graph orthogonal adaptation network (MoRa) that enhances the robustness of emotion-aware representations and mitigates inter-subject variability in cross-subject EEG emotion recognition. Specifically, the multi-graph module performs intra-region orthogonal decoupling of emotion-aware and agnostic information, reducing redundancy and enhancing consistency across dual emotional spaces. To reduce inter-region redundancy, the cross-region emotion refinement module integrates soft orthogonality and knowledge distillation to enhance collaborative information exchange across regions. Moreover, cross-subject experiments on the SEED and SEED-IV datasets demonstrate that the MoRa achieves state-of-the-art performance in EEG emotion recognition. C. L. Philip Chen, Tong Zhang 0015 |
SMC | 3 |
| 2025 | CoDD: Convenient-Oriented Depression Detection from Few-Channel EEG with Channel ReconstructionabstractMajor depressive disorder (MDD) is characterized by imbalanced connectivity between brain regions, rather than simply increased or decreased activity in a specific region. Current studies on depression detection using EEG predominantly utilize whole-brain signal (full-channel). However, due to the cost and convenience limitations of full-channel devices, portable miniaturized EEG devices with only a few electrodes (few-channel) are more suitable for widespread use in daily scenarios. Depression detection from few-channel EEG data is challenging because these devices can only capture EEG signals from a portion of the brain regions. To address the challenge, we propose a novel Convenience-Oriented Depression Detection model (CoDD). The model aims to enhance the performance of MDD detection using few-channel EEG by leveraging knowledge from full-channel EEG. Specifically, we capture prior whole-brain connectivity patterns from full-channel data to reconstruct missing channels in few-channel EEG data, supplementing the few-channel data with critical encodings and cooperative relationships related to depression. In addition, we use full-channel data to guide the process of information mining and depression detection to ensure that the training process conforms to the original data distribution. Experiments conducted on the MODMA and PRED+CT datasets demonstrate that the model achieves SOTA performance in depression detection under few-channel device conditions, validating the effectiveness of the proposed technique. Zihua Xu, C. L. Philip Chen, Tong Zhang 0015 |
SMC | 3 |
| 2025 | Multi-agent reinforcement learning for vibration control of regenerative active suspension
Xiaotian Gao, Yu Du 0009, Shi-Yuan Han, Wenxiu Zhao, Jin Zhou 0003, Tong Zhang 0015, C. L. Philip Chen |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Dynamic Spatial-Temporal Imputation Network With Missing Features for Traffic Data ImputationabstractMissing traffic data caused by sensor failures or communication errors significantly hinders the efficiency of downstream tasks in Intelligent Transportation Systems (ITS), such as the critical functions of traffic monitoring and decision-making. Since missing data contains important information, it is essential to extract dynamic spatial-temporal correlations in traffic processes by incorporating these missing features. Motivated by these concerns, a novel Dynamic Spatial-Temporal Imputation Network with Missing Features (DSTMIN) is proposed to accurately impute traffic data. DSTMIN comprises an embedding layer, a Mask Attention module (MA), and a Fusion Graph Convolution module (FGC). Specifically, an embedding layer is designed to accurately represent the distribution of missing data, thereby capturing both temporal features and missing features. Furthermore, in order to effectively capture the temporal correlations, MA integrates the missing features to emphasize the significance of observed data and reduce the adverse effects caused using incomplete data. To capture spatial correlations, FGC constructs the spatial graphs and model dynamic spatial correlations from traffic subsequences and the missing-data graph in the presence of missing features. The proposed DSTMIN is adequately evaluated to demonstrate its superior performance on two datasets, which achieves a remarkable 20% reduction in imputation error compared to state-of-the-art methods. Hao Li 0100, Shi-Yuan Han, Jie Liu 0002, Jin Zhou 0003, Tong Zhang 0015, C. L. Philip Chen |
IEEE Internet Things J. | 6 |
| 2025 | Broad learning system based on fractional order optimization
Tong Zhang 0015, Zhang Tao, C. L. Philip Chen |
Neural Networks | 2 |
| 2025 | Label Feature Co-Learning for Facial and EEG Emotion RecognitionabstractRecognizing human emotions through behavioral and physiological signals is fundamental to overall health. However, since emotion occurs transiently, a semantics mismatch exists between the uniformly annotated label and multimodal temporal signals, leading to emotional ambiguity. Previous studies used the annotated label to guide feature learning, which makes it intractable to accurately identify emotional elicitation moments within each signal. Moreover, the inconsistency of specific elicitation moments across different signals complicates emotion recognition. The model hardly learns discriminative features due to emotional ambiguity, which weakens its ability to differentiate between emotions. To tackle the above challenges, this paper proposes a novel label feature co-learning model (LFCL) for emotion recognition through multimodal signals. Specifically, the LFCL leverages unimodal and multimodal information and adaptively generates instance-level emotion labels, thus precisely locating emotion elicitation moments within each signal. To promote emotion consistency across different signals, the LFCL incorporates a dynamic label calibration mechanism to balance the label generation process with historical information. Furthermore, to enhance the deep interaction between signals, the LFCL conducts multimodal interactive fusion to integrate multi-level multimodal features and extract global emotional information. The LFCL performs precise label-to-feature alignment to capture discriminative features of each signal, effectively alleviating emotional ambiguity and improving the ability to distinguish different emotions. Extensive experiments on three publicly available datasets demonstrate the effectiveness and generalization of the proposed model. Mingchen Cai, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | AM-ConvBLS: Adaptive Manifold Convolutional Broad Learning System for Cross-Session and Cross-Subject Emotion RecognitionabstractEmotion recognition based on electroencephalography (EEG) data has gained rapid development because of its believability and accuracy. However, most existing EEG emotion recognition methods suffer from three issues: 1) underutilization of multi-scale emotion representations, 2) underexploitation of emotion labels and intrinsic geometric structures, and 3) inherent non-stationarity characteristic and individual variability of EEG signals. To this end, we propose an Adaptive Manifold Convolutional Broad Learning System (AM-ConvBLS) to capture the multi-scale distribution aligned emotion patterns in the geometric structure preserved emotion submanifold space. To begin with, a Multi-Scale Representation Learning (MSRL) module is developed to learn diverse multi-scale emotion representations. The Domain Discrepancy Elimination (DDE) module is then developed to further align feature distributions in the source and target domains. To further utilize EEG emotion labels, we devise a Label Manifold Information Exploration (LMIE) module to retain the sample label consistency. In addition, a global feature importance analysis method for AM-ConvBLS based on Shapley Additive Global importancE (SAGE) is designed to investigate the EEG frequency band importance and brain neural activation patterns. Experiments on SEED, SEED-IV, and SEED-V datasets demonstrate the effectiveness and superiority of our AM-ConvBLS. Chunyu Lei, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | DGC-Link: Dual-Gate Chebyshev Linkage Network on EEG Emotion RecognitionabstractEEG emotion recognition presents several challenges, including region correlation, local and long-range node connectivity, and multi-channel patterns, necessitating advanced methods capable of effectively capturing and utilising complex EEG signal information. This paper introduces a novel method, the dual-gate Chebyshev Linkage network (DGC-Link), which comprises three main components: the Chebyshev Linkage (CL) module for extracting regional correlation features, the dual-gate module for regulating the flow of different-order information, and the deep network for extracting multi-channel features and enhancing representation capabilities. Validated on three datasets (SEED, DREAMER, and MPED) with ablation experiments demonstrating each component's effectiveness, DGC-Link achieves superior recognition performance compared to state-of-the-art methods. Notably, it achieves 96.43% accuracy on differential entropy on the SEED dataset, and 98.58%, 97.62%, and 98.01% for valence, arousal, and dominance classifications on the DREAMER dataset, along with 78.48% and 44.93% for 3-class and 7-class classifications on the MPED dataset. These results highlight DGC-Link's potential for improved performance in EEG emotion recognition. Tong Zhang 0015, C. L. Philip Chen, Xiaowei Zhang 0001, Bin Hu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Grop: Graph Orthogonal Purification Network for EEG Emotion RecognitionabstractThe existence of emotion-irrelevant representations and individual variability impedes the extraction of robust emotional representations, limiting the adaptability of EEG emotion recognition. Massive studies focus on the mining of emotion-aware information, overlooking emotion-agnostic information, which is insufficient for the extraction of emotion-relevant features against redundancy and variation. In this paper, Graph Orthogonal Purification Network (Grop) is proposed to enhance individual adaptability through improvements in the orthogonality and transferability between emotion-relevant and emotion-irrelevant features. Specifically, the proposed Grop utilized a graph representation extraction module to capture both emotion-relevant and emotion-irrelevant features by the dual graph. The representation orthogonal purification module is developed to eliminate redundant information through feature projection and feature purification. Moreover, the dual emotional space alignment module is imposed to align distribution discrepancies in different emotion feature spaces. To assess the effectiveness of the proposed Grop, various experiments are conducted on two public EEG emotion datasets, i.e., SEED and SEED-IV. The results achieve state-of-the-art performance, demonstrating the capability of the Grop to capture robust emotion features and alleviate the intra- and inter-subject discrepancies. C. L. Philip Chen, Bianna Chen, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | TFAGL: A Novel Agent Graph Learning Method Using Time-Frequency EEG for Major Depressive Disorder DetectionabstractThe abnormality in depression exhibits reciprocal imbalanced connectivity between brain regions rather than increased or decreased activity of one particular area. Current works primarily align the distributions of EEG electrodes with insufficient simulation of neurophysiological structures. Moreover, they neglect significant collaborative relationships among diverse brain regions, which limits the performance of MDD detection. Considering the comprehensive information across brain regions and domains, we propose a novel EEG-based MDD detection model named Time-Frequency Agent Graph Learning (TFAGL), to capture the specific whole-brain level collaborative mechanism of MDD. Specifically, we generate agent nodes adaptively to perform global interactions among regions to sufficiently simulate the function of principal neurons, thereby forming a dynamic local-global connectivity graph to capture connectivity patterns for intra- and inter-regions. Furthermore, interactive learning across different receptive fields through multi-scale graph convolution is applied for each domain and connectivity. Besides, we construct feature extractors for both time and frequency domains and apply intra- and inter-domain constraints to remove redundancy and enhance the discriminability, thus obtaining comprehensive information representations. Extensive experiments on the public EEG MDD detection datasets demonstrate the superiority of TFAGL compared with the state-of-the-art methods. Zihua Xu, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Improving the Interpretability Through Maximizing Mutual Information for EEG Emotion RecognitionabstractTrustworthy Graph Neural Networks (GNNs) for EEG emotion recognition should identify emotions accurately and elucidate corresponding rationales. Current GNNs have achieved notable performance by dynamically modeling emotional connections between EEG channels. However, these GNNs lack interpretability due to the absence of explicit rationale behind their predictions. This paper conducts a comprehensive identification of important EEG channels to enhance the interpretability of EEG emotion recognition from the perspective of mutual information. Specifically, an Adjacency-Explainable Graph Neural Network (AEG) for ante-hoc interpretability is proposed to capture genuine EEG emotional connections, which gives a theoretical guarantee to remove spurious connections. Moreover, a Channel-wise Adaptive Class Activation Mapping Explainer (CACA) for post-hoc interpretability is developed to locate the EEG channels that contribute most to predictions. Experimental results on three datasets, i.e., SEED, SEED-IV, and DREAMER, prove that imbuing training processes with enhanced interpretability ensures significant performance improvements in emotion recognition. Quantitative comparisons of post-hoc interpretability also demonstrate the superiority of CACA. Furthermore, this paper illustrates two potential applications of the proposed methodologies, showing their broader utility and significance. C. L. Philip Chen, Bianna Chen, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Semantic and Emotional Dual Channel for Emotion Recognition in ConversationabstractEmotion recognition in conversation (ERC) aims at accurately identifying emotional states expressed in conversational content. Existing ERC methods, although relying on semantic understanding, often encounter challenges when confronted with incomplete or misleading semantic information. In addition, when dealing with the interaction between emotional and semantic information, existing methods are often difficult to effectively distinguish the complex relationship between the two, which affects the accuracy of emotion recognition. To address the problems of semantic misdirection and emotional cross-talk encountered by traditional models when confronted with complex conversational data, we propose a semantic and emotional dual channel (SEDC) strategy for emotion recognition in conversations to process emotional and semantic information independently. Under this strategy, emotion information provides an auxiliary recognition function when the semantics are unclear or lacking, enhancing the accuracy of the model. Our model consists of two modules: the emotion processing module accurately captures the emotional features of each utterance through contrastive learning, and then constructs a dialogue emotion propagation map to simulate the emotional information conveyed in the dialogue; the semantic processing module combines an external knowledge base to enhance the semantic expression of the dialogue through knowledge enhancement strategies. This divide-and-conquer approach allows us to more deeply analyze the emotional and semantic dimensions of complex dialogues. Experimental results on the IEMOCAP, EmoryNLP, MELD, and DailyDialog datasets show that our approach significantly outperforms existing techniques and effectively improves the accuracy of dialogue emotion recognition. Zhenyu Yang 0002, Zhibo Zhang 0009, Yuhu Cheng 0001, Tong Zhang 0015, Xuesong Wang 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Ugan: Uncertainty-Guided Graph Augmentation Network for EEG Emotion RecognitionabstractThe underlying time-variant and subject-specific brain dynamics lead to statistical uncertainty in electroencephalogram (EEG) representations and connectivities under diverse individual biases. Current works primarily augment statisticallike EEG data based on deterministic modes without comprehensively considering uncertain statistical discrepancies in representations and connectivities. This results in insufficient domain diversity to cover more domain variations for a generalized model independent of individuals. This article proposes an uncertainty-guided graph augmentation network (Ugan) to generalize EEG emotion recognition across subjects by comprehensively mimicking and constraining the uncertain statistical shifts across individuals. Specifically, an uncertainty-guided graph augmentation module is employed to augment both connectivities and features of EEG graph by manipulating domain statistical characteristics. With the original and augmented EEG graph covering diverse domain variations, the model can mimic the uncertain domain shifts to achieve better generalizability against potential subject variability. To extract discriminative characteristics and preserve emotional semantics after augmentation, a graph coteaching learning module is designed to facilitate coteaching knowledge learning between the original and augmented views. Moreover, a coteaching regularization module is developed to constrain semantic domain invariance and consistency, thereby rendering the model invariant to uncertain statistical shifts. Extensive experiments on three public EEG emotion datasets, i.e., Shanghai Jiao Tong University emotion EEG dataset (SEED), SEED-IV, and SEED-V, validate the superior generalizability of Ugan compared to the state-of-the-art methods. Bianna Chen, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Debiased Sequential Recommendation by Separating Long-Term and Short-Term InterestsabstractA significant problem in sequential recommendation (SR) is the over-recommendation of popular items, leading to popularity bias, as users often follow these items due to conformity. Existing methods measure users’ conformity factors to reduce the impact of popularity bias. However, these methods do not consider the differences in users’ conformity behavior in the long-term and short-term. To address this, we propose LSDRec, a novel debiased SR method structured around three key tasks: degree-centrality conformity awareness, dual-scale interest encoding, and adaptive conformity information fusing. The degree-centrality conformity awareness task constructs a multiuser interaction graph, employs a graph convolutional network (GCN) to obtain global user conformity representations, and uses the degree centrality algorithm to compute users’ long-term and short-term conformity factors. The dual-scale interest encoding task models users’ long-term and short-term interests separately, obtaining corresponding interest representations and further enhancing them through the adaptive conformity information fusing task. The adaptive conformity information fusing task contrasts global conformity representations with long-term and short-term interest representations, adaptively integrating conformity factors and dynamically adjusting the degree of conformity information transfer. Together, these three tasks effectively mitigate popularity bias and improve the accuracy of user interest modeling. Our extensive evaluations of four diverse datasets demonstrate LSDRec's superior performance over current state-of-the-art methods. Zhenyu Yang 0002, Wenyue Hu, Tong Zhang 0015, Yuhu Cheng 0001, Xuesong Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | LLM40FD: Unlocking the Potential of LLM for Anonymous Zero-Shot Fraud DetectionabstractCredit card (CC) fraud detection within the realm of financial security faces challenges such as data imbalance, large-scale anonymized transaction datasets, and the need for system-specific model training. Past methods often fail to address these aforementioned issues simultaneously. Utilizing a single model results in a lack of zero-shot capability without adaptation for real-world scenarios. This article introduces LLM40FD, a novel framework that leverages a large language model (LLM) to overcome these obstacles in anonymous zero-shot fraud detection. LLM40FD addresses the aforementioned challenges in CC fraud detection by employing a distribution-based one-class function and the walking embedding, without reliance on labeled data or fine-tuning in downstream. Additionally, LLM40FD enhances the model’s ability to detect fraudulent patterns and define robust decision boundaries. This is achieved through a dual-augmentation strategy and implicit contrastive learning, which generate enriched positive and negative samples. Our experiments demonstrate that LLM40FD not only achieves state-of-the-art (SOTA) performance in the full-shot setting but also exhibits strong zero-shot capability even with limited training data. Furthermore, we conduct additional experiments to validate the effectiveness and working mechanism of LLM40FD. Kaixiang Yang 0001, Zhiwen Yu 0002, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | AE-AMT: Attribute-Enhanced Affective Music Generation With Compound Word RepresentationabstractAffective music generation is a challenge for symbolic music generation. Existing methods face the problem that the perceived emotion of the generated music is not evident because music datasets containing emotional labels are relatively small in quantity and scale. To address this issue, an attribute-enhanced affective music transformer (AE-AMT) model is proposed to generate perceived affective music with attribute enhancement. In addition, a multiquantile-based attribute discretization (MQAD) strategy is designed, enabling the model to generate intensity-controllable affective music pieces. Furthermore, A replication-expanded compound representation of the control signals (RECR) method is designed for control signals to improve the controllability of the model. In objective experiments, the AE-AMT model demonstrated a 29.25% and 19.5% improvement in overall emotion accuracy, along with a 30% and 32% improvement in arousal accuracy on the datasets EMOPIA and VGMIDI. These improvements are achieved without significant difference in objective music quality, while also providing ample novelty and diversity compared to the current state-of-the-art approach. Moreover, subjective experiments revealed that the AE-AMT model outperformed comparison models, especially in low valence and arousal based on the Wilcoxon signed ranks test. Additionally, the soft variant model of AE-AMT exhibited a significant advantage in valence, low arousal, and overall music quality. These experiments showcase the AE-AMT model's ability to significantly enhance arousal performance and strike a balance between emotional intensity and musical quality through adaptable strategies. Weiyi Yao, C. L. Philip Chen, Zongyan Zhang, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Self-Prompt Guided Image Outpainting Model for Captions Absence in Social ScenesabstractThe limitations of acquisition equipment often result in scene image data of limited size, posing a challenge for comprehensive analysis of social image datasets. Advances in generative models have introduced image outpainting techniques that expand the size of acquired social scene images, thereby enhancing the value of social image data. Stable diffusion (SD), which benefits from the guidance of caption prompts, shows excellent performance in image outpainting. However, its heavy reliance on manual prompts leads to a significant drawback: a decrease in the quality of generated images without prompts. To overcome this challenge, we propose a novel self-prompt diffusion model for image outpainting that extrapolates images based on the semantics of the source image, thereby removing the dependence on manual prompts. Specifically, we design a prompt autoencoder that uses an autoregressive transformer to map prompt embeddings into their semantic space, facilitating the construction of a semantic decoder. The semantic decoder and prompt embeddings are then cooptimized within the proposed prompt embedding network, allowing the mapping of image features to the stable diffusion prompt embeddings. Furthermore, by exploiting the inherent generative capabilities of diffusion models, we introduce a seam line regeneration mechanism to address the common problem of seam lines when splicing input and generated images. Comparative experiments on the Places2 and COCO datasets show that our method outperforms current state-of-the-art approaches on visual quality metrics and is adaptable to the stable diffusion model without additional fine-tuning. Zongyan Zhang, C. L. Philip Chen, Haohan Weng, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | From Disagreement to Unity: A Cascade Perspective Network for Subjective TasksabstractSubjective tasks involve annotating instances according to personal opinions, emotions, and feelings. The neural network model must learn about human thought and expression complexity. Handling instances with different annotation opinions is the main challenge in subjective tasks. Existing methods focus on majority voting or integrating opinions from a few assigned annotators, leading to biased decisions and limited performance. To address this issue, this article proposes a cascade perspective network (CPNet) to uncover reliable disagreement for subjective tasks. Specifically, CPNet learns each annotator’s personalized knowledge from annotation disagreement and stores them in an annotator bank through the personal perspective module for abundant disagreement information. Then, CPNet obtains consistent opinions by referring to all annotators’ opinions from the annotator bank through the comprehensive perspective module to reduce bias caused by noise. CPNet improves decision-making by considering the diversity and comprehensiveness of all annotators’ opinions. Moreover, it performs well in subjective tasks with limited or numerous annotators. The state-of-the-art (SOTA) results on subjective datasets from different domains demonstrate the effectiveness and generalizability of CPNet. Tong Zhang 0015, Canhui Zhang, Shuzhen Li, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | AdamGraph: Adaptive Attention-Modulated Graph Network for EEG Emotion RecognitionabstractThe underlying time-variant and subject-specific brain dynamics lead to inconsistent distributions in electroencephalogram (EEG) topology and representations within and between individuals. However, current works primarily align the distributions of EEG representations, overlooking the topology variability in capturing the dependencies between channels, which may limit the performance of EEG emotion recognition. To tackle this issue, this article proposes an adaptive attention-modulated graph network (AdamGraph) to enhance the subject adaptability of EEG emotion recognition against connection variability and representation variability. Specifically, an attention-modulated graph connection module is proposed to explicitly capture the individual important relationships among channels adaptively. Through modulating the attention matrix of individual functional connections using spatial connections based on prior knowledge, the attention-modulated weights can be learned to construct individual connections adaptively, thereby mitigating individual differences. Besides, a deep node-graph representation learning module is designed to extract long-range interaction characteristics among channels and alleviate the over-smoothing problem of representations. Furthermore, a graph domain co-regularized learning module is imposed to tackle the individual distribution discrepancies in connection and representations across different domains. Extensive experiments on three public EEG emotion datasets, i.e., SEED, DREAMER, and MPED, validate the superior performance of AdamGraph compared with state-of-the-art methods. C. L. Philip Chen, Bianna Chen, Tong Zhang 0015 |
IEEE Trans. Cybern. | 3 |
| 2025 | Broad Metric Learning: A Fast and Efficient Discriminative Metric Learning ModelabstractMetric learning aims to learn a discriminative metric space, where samples of the same class stay close, and those of different classes far apart. Existing classical metric learning methods based on linear transformation have limited learning performance due to the low representation capability. Although deep metric learning learns nonlinear mappings, the training may come across convergence issues and be unstable. Additionally, many classical metric learning algorithms suffer from long computational time for iterative optimization especially when data dimension is high. Deep metric learning also requires high training cost. To learn a metric space more efficiently and effectively, this article proposes a novel broad metric learning (BML) model, which learns the data transformation by training a broad network. BML maps input data to a broad feature space by fast and convenient nonlinear feature mapping based on random weights, and learns a linear transformation to a discriminative output space. Intraclass distance is reduced by minimizing the distance between data and their class-specific reference points in the target space. The hard-triplet distance learning (HDL) is proposed to learn the distance of hard positive and negative sample pairs, which enhances the intraclass compactness and interclass separation. Closed-form solutions are adopted to solve the optimization problems efficiently when learning the linear transformation. Experiments are conducted on nine datasets to verify the efficiency and effectiveness of BML. BML learns fast and achieves high classification and clustering accuracies in the learned data space. Xiaoman Hu, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Cybern. | 3 |
| 2025 | Co-Training Broad Siamese-Like Network for Coupled-View Semi-Supervised LearningabstractMultiview semi-supervised learning is a popular research area in which people utilize cross-view knowledge to overcome the limitation of labeled data in semi-supervised learning. Existing methods mainly utilize deep neural network, which is relatively time-consuming due to the complex network structure and back propagation iterations. In this article, co-training broad Siamese-like network (Co-BSLN) is proposed for coupled-view semi-supervised classification. Co-BSLN learns knowledge from two-view data and can be used for multiview data with the help of feature concatenation. Different from existing deep learning methods, Co-BSLN utilizes a simple shallow network based on broad learning system (BLS) to simplify the network structure and reduce training time. It replaces back propagation iterations with a direct pseudo inverse calculation to further reduce time consumption. In Co-BSLN, different views of the same instance are considered as positive pairs due to cross-view consistency. Predictions of views in positive pairs are used to guide the training of each other through a direct logit vector mapping. Such a design is fast and effectively utilizes cross-view consistency to improve the accuracy of semi-supervised learning. Evaluation results demonstrate that Co-BSLN is able to improve accuracy and reduce training time on popular datasets. Yikai Li 0001, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Cybern. | 3 |
| 2025 | Cyclic Data Distillation Semi-Supervised Learning for Multi-Modal Emotion RecognitionabstractMulti-modal emotion recognition (MER) integrates multi-modal signals to help computers comprehensively understand human emotions, which is a crucial technology in human-computer interactions. However, the amount of labeled multi-modal emotion data is small and limits MER performance due to its expensive manual annotations. Meanwhile, semi-supervised learning (SSL) methods improving MER models with enormous unlabeled data suffer from confirmation bias, resulting in biased data distribution. To tackle these challenges, this paper proposes a cyclic data distillation semi-supervised learning (CDD-SSL) for MER tasks. CDD-SSL leverages multiple pre-trained unimodal teacher models and confidence-boosting pseudo-labelling (CBPL) to boost the confidence of multi-modal ensemble outputs and distill reliable and class-representative data from numerous unlabeled data. It then utilizes reliable and less-biased data to train a multi-modal student model and provides feedback to update all unimodal teacher models. CDD-SSL is a cyclic teacher-student framework with a feedback mechanism that gradually mitigates confirmation bias and obtains an effective MER model. Experimental results on four benchmark datasets demonstrate that CDD-SSL achieves superior performance over both the semi-supervised methods and the state-of-the-art fully-supervised models in MER tasks. Shuzhen Li, Tong Zhang 0015, C. L. Philip Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | SGB-Net: Scalable Graph Broad NetworkabstractDue to the complexity and self-evolutionary property of graph data in reality, graph learning methods require both validity to represent unstructured data and scalability to adapt to evolving graphs. However, current works have representation learning limitations on optimizable graph feature space due to the bottleneck of the structure depth. Moreover, they encounter a complete retraining process when graphs evolve, especially in the case without the assistance of new labels. To address the above issues, we propose a scalable graph broad network (SGB-Net), which contains three proposed modules: the graph feature broad transformation layer (GFBT layer) for enhancing graph embedding and two update algorithms (SGB-Net-U, SGB-Net-S) for endowing scalability. The GFBT layer aims to explicitly expand the graph feature space and broadly build the model. It constructs two expandable feature spaces in various graph scales to embed graphs discriminatively. SGB-Net-U is an exploratory method designed to tackle the label-free graph incremental learning (GIL) problem by leveraging unsupervised incremental knowledge to expand graph representation. SGB-Net-S endows scalability in classical incremental learning scenarios involving labels. Benefiting from its broad construction framework, SGB-Net not only enhances graph embeddings but also seamlessly adapts and improves performance in response to graph expansion without requiring retraining. In the experiments conducted on 15 benchmark datasets, SGB-Net outperforms state-of-the-art GNNs in terms of both effectiveness and scalability. Yuebin Xu, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Hierarchical Dynamic Graph Convolutional Network With Interpretability for EEG-Based Emotion RecognitionabstractGraph convolutional networks (GCNs) have shown great prowess in learning topological relationships among electroencephalogram (EEG) channels for EEG-based emotion recognition. However, most existing GCN-only methods are designed with a single spatial pattern, lacking connectivity enhancement within local functional regions and ignoring the data dependencies of EEG original data. In this article, hierarchical dynamic GCN (HD-GCN) is proposed to explore dynamic multilevel spatial information among EEG channels, with discriminative features of EEG signals as auxiliary information. Specifically, representation learning in topological space consists of two branches: one for extracting global dynamic information and one for exploring augmentation information in local functional regions. In each branch, a layerwise adjacency matrix is utilized to enrich the expressive power of GCN. Furthermore, a data-dependent auxiliary information module (AIM) is developed to capture multidimensional fusion features. Extensive experiments on two public datasets, SJTU emotion EEG dataset (SEED) and DREAMER, demonstrate that the proposed method consistently exceeds state-of-the-art methods. Interpretability analysis of the proposed model is performed, discovering the active brain regions and important electrode pairs related to emotion. Mengqing Ye, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Discovery of Shared Latent Nonlinear Effective Connectivity for EEG-Based Depression DetectionabstractGranger causality (GC) effective connectivity (EC) calculated from electroencephalogram (EEG) signals has been widely used in mental disorder detection. However, the existing methods only take into account linear dynamics or nonlinear dynamics within a single sample, ignoring the nonlinear dynamics shared by the same class of subjects. In this article, a model combining graph neural networks (GNNs) and variational autoencoders (VAEs) is proposed to construct shared latent nonlinear EC from raw EEG signals for depression detection. Several convolution modules and fully connected layers are used in the graph encoding network to learn the embeddings of the connectivity connected by every two EEG channels. In the graph decoding network, a class-specific Gaussian mixture model (GMM) is introduced in the VAEs to model shared dynamics in EC of the same class of subjects, and the shared dynamics combine the encoded embeddings of the EC and the past time series to restore raw EEG signals. Through a node-to-edge encoding process and an edge-to-node decoding process, the shared latent nonlinear EC in EEG signals can ultimately be learned by gradually optimizing the model's loss function. The performance of the proposed method is verified on several open-accessed datasets. The excellent results prove that the proposed neural networks can learn more generalized nonlinear EC representations, and shared latent dynamics discovery can also help to identify depression better. The code is available at https://github.com/william-yuan2012/DSLNEC-tscausality. Wenjie Yuan 0001, Xiaowei Zhang 0001, Xuejuan Zhang, Shuangyan Wang, Tianzhi Wang, Tong Zhang 0015, Qinglin Zhao, Bin Hu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Broad Learning System Based on Fractional Feature OptimizationabstractBroad learning system (BLS) have demonstrated excellent performance in terms of both speed and accuracy in tasks such as image classification. In BLS, the feature nodes predominantly utilize linear features, and sparse representation is mainly employed in the feature optimization component. The robustness of these features to different data needs to be improved. Although there are many improved algorithms for BLS in feature optimization, there is no improvement based on fractional calculus at present. This article proposes BLS-FC, a novel data classification and regression method that can seamlessly combine BLS and fractional calculation. Fractional calculus describes the properties of data between integer orders and has memory properties. Fractional Fourier transform (Frft) also has time domain and frequency domain information. First, Frft is added to the broad learning feature node extraction to enrich the node features, which is called BLS-Frft. Second, fractional calculus is integrated into the BLS-Frft sparse representation feature optimization, and the feature representation capability is enhanced by fractional differential memory. This part is called BLS-FS. Finally, in order to solve the problem of unstable features of random fractional order subspaces, a fractional order multiscale feature interaction based on BLS-Frft is proposed, which is called BLS-MF. Experimental results across various classification and regression datasets demonstrate the superior performance of the proposed method. Tong Zhang 0015, C. L. Philip Chen, Tao Zhang 0103 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Robust Incremental Broad Learning System for Data Streams of Uncertain ScaleabstractDue to its marvelous performance and remarkable scalability, a broad learning system (BLS) has aroused a wide range of attention. However, its incremental learning suffers from low accuracy and long training time, especially when dealing with unstable data streams, making it difficult to apply in real-world scenarios. To overcome these issues and enrich its relevant research, a robust incremental BLS (RI-BLS) is proposed. In this method, the proposed weight update strategy introduces two memory matrices to store the learned information, thus the computational procedure of ridge regression is decomposed, resulting in precomputed ridge regression. During incremental learning, RI-BLS updates two memory matrices and renews weights via precomputed ridge regression efficiently. In addition, this update strategy is theoretically analyzed in error, time complexity, and space complexity compared with existing incremental BLSs. Different from Greville's method used in the original incremental BLS, its results are closer to the solution of one-shot calculation. Compared with the existing incremental BLSs, the proposed method exhibits more stable time complexity and superior space complexity. The experiments prove that RI-BLS outperforms other incremental BLSs when handling both stable and unstable data streams. Furthermore, experiments demonstrate that the proposed weight update strategy applies to other random neural networks as well. Linjun Zhong, C. L. Philip Chen, Jifeng Guo 0002, Tong Zhang 0015 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Temporal Group Attention Network with Affective Complementary Learning for Gait Emotion RecognitionabstractSkeleton-based methods in Gait Emotion Recognition (GER) that abstract gait as a spatial-temporal graph in a non-Euclidean space have achieved remarkable success. However, existing studies neglect capturing crucial spatial-temporal implicit dependencies and make insufficient use of manually crafted affective complementary information. In this paper, we propose a novel Temporal Group Attention Network with Affective Complementary Learning (T2A). Specifically, we propose a Temporal Group Attention (TGA) module, which captures spatial-temporal implicit dependencies and crucial features across space-time. Moreover, we design an affective complementary learning strategy, which first introduces an Affective Denoising (AD) module to reduce the impact of noise in weak emotion-related information and learn augmented affective features. Then, we adopt a decision-level fusion of the denoised affective features with the deep skeleton features to compensate for the discrepancies between different feature spaces. Experimental results on Emotion-Gait and EMOGAIT demonstrate that our proposed method significantly outperforms state-of-the-art methods. C. L. Philip Chen, Haiqi Liu, Tong Zhang 0015 |
BIBM | 4 |
| 2024 | Desigen: A Pipeline for Controllable Design Template GenerationabstractTemplates serve as a good starting point to implement a design (e.g., banner, slide) but it takes great effort from designers to manually create. In this paper, we present Desigen, an automatic template creation pipeline which generates background images as well as harmonious layout elements over the background. Different from natural images, a background image should preserve enough non-salient space for the overlaying layout elements. To equip existing advanced diffusion-based models with stronger spatial control, we propose two simple but effective techniques to constrain the saliency distribution and reduce the attention weight in desired regions during the background generation process. Then conditioned on the background, we synthesize the layout with a Transformer-based autoregressive generator. To achieve a more harmonious composition, we propose an iterative inference strategy to adjust the synthesized background and layout in multiple rounds. We constructed a design dataset with more than 40k advertisement banners to verify our approach. Extensive experiments demonstrate that the proposed pipeline generates high-quality templates comparable to human designers. More than a single-page design, we further show an application of presentation generation that outputs a set of theme-consistent slides. The data and code are available at https://whaohan.github.io/desigen. Haohan Weng, Danqing Huang, Chin-Yew Lin, Tong Zhang 0015, C. L. Philip Chen |
CVPR | 6 |
| 2024 | Spatial and Frequency-Based Feature Reconstruction for Cross-Database Micro-Expression RecognitionabstractThe vulnerability of individual-database learned model for Micro-Expression Recognition (MER) has significantly hindered their performance in real-world scenarios. Cross-Database Micro-Expression Recognition (CDMER) aims to enhance the robustness and generalization performance for more complicated situations. Most existing CDMER methods typically attempt to leverage spatial information to eliminate the inevitable domain shift between source and target domains. In this paper, we propose a novel feature space reconstruction model that provides feature space with better generalization for CDMER. Specifically, we introduce the Spatial and Frequency Domain Co-learning (SFDC), a three-branch module, that adaptively exploits the spatial and frequency characteristics of intermediate feature representations, capturing the global and local features of micro-expression adequately. Furthermore, to facilitate the synchronization of two domains, Source and Target Domain Synchronization (STDS) module is employed to guide the alignment of different subspaces simultaneously. Extensive experimental results on the SMIC and CASME II databases demonstrate the effectiveness of the reconstruction and the superiority of our proposed method over state-of-the-art (SOTA) methods. Zhi Feng, C. L. Philip Chen, Tong Zhang 0015 |
ECAI | 4 |
| 2024 | Disentanglement Network: Disentangle the Emotional Features from Acoustic Features for Speech Emotion RecognitionabstractSpeech emotion recognition plays a crucial role in human-computer interaction. However, data distribution of speech signals varies among individuals for emotion recognition. It may guide models to focus more on identity information rather than emotional information, which impairs the generalization ability of models. To address this issue, this paper proposes a novel Disentanglement Network (DTNet) to disentangle emotional features from acoustic features. Specifically, DTNet first captures hidden identity features from acoustic features through an identity-aware module. Then, we design a disentanglement module to disentangle emotional features from acoustic features within the constraints of a reconstruction module and the hidden identity features. These modules enable the DTNet to extract more discriminative emotional features for emotion recognition. Experimental results on both speaker-independent and speaker-dependent settings have proven the effectiveness of DTNet, and this method achieves an unweighted accuracy (UA) of 74.8% on the IEMOCAP dataset and UA of 95.5% on the Emo-DB dataset, outperforming the state-of-the-art methods on both datasets. Zhichen Yuan, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
ICASSP | 4 |
| 2024 | Multi-Scale Prompt Memory-Augmented Model for Black-Box ScenariosabstractXiaojun Kuang, C. L. Philip Chen, Shuzhen Li, Tong Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xiaojun Kuang, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
NAACL-HLT | 4 |
| 2024 | Adaptive Domain-Enhanced Transfer Learning for Welding Defect ClassificationabstractThe integration of Intelligent Welding Systems (IWS) in smart manufacturing leverages advancements in sensors, robotics, and artificial intelligence to optimize welding processes. However, in industry practice, we still face challenges such as sufficient data is not available for every manufacturing task, the costs associated with welding data annotation quality, and the risk of knowledge forgetting during the continual welding process. To tackle these issues, we developed an Adaptive Domain-Enhanced Transfer Learning (ADETL) framework that integrates self-supervised and continual learning strategies. This framework is adept at using incremental and unlabeled data for pre-training, in which we analyze the parameter space, loss landscape, and make the model understand the behaviour of knowledge transfer from diverse source domains. The ADETL framework improves the performance of defect classification, offering a promising solution to the challenges inherent in automatic, continuous welding operations. Dan Dai, Pasquale Franciosa, Tong Zhang 0015, C. L. Philip Chen, Dariusz Ceglarek |
SMC | 4 |
| 2024 | DPCA: Dynamic Probability Calibration Algorithm in LMaaSabstractProbability calibration is a method to improve the reliability of models by linking the predicted probability to accuracy. Most research follow a static strategy of full fine-tuning. These studies do not consider dynamic data and sparse parameters in Language Models as a Service(LMaaS), leading to limited effectiveness of probability calibration. To address above issues, we propose a dynamic probability calibration algorithm (DPCA) to consider both data flow and parameter freezing. DPCA consists of streaming annotation (SA) task and dynamic calibration (DC) task. The SA task takes a specified number of samples from the training data stream. The sampled data is automatically annotated according to the deviation between predicted probability and true label. The DC task injects the probability deviation of the SA task into next training epoch through adapter-tuning. DPCA achieves data augmentation in LMaaS through joint learning of sample labels and their predicted probability deviation. This work validates DPCA through the implementation on BERT architecutre. The proposed model achieves overall performance improvement on both Chinese and English NLP tasks. Experimental results demonstrate a 2.17% decrease in average ECE without de-creasing in accuracy. Experimental analysis demonstrate the effectiveness and generalizability of DPCA in LMaaS. Zhongyi Deng, C. L. Philip Chen, Tong Zhang 0015 |
SMC | 3 |
| 2024 | DHFusion: Deep Hypergraph Convolutional Network for Multi-Modal Physiological Signals FusionabstractMulti-modal physiological signals fusion integrates multiple heterogeneous physiological signals to characterize human physiological activities, which is basic research on biomedical signal processing. Most researches on multi-modal physiological signal fusion are based on heterogeneous graph fusion networks (HGFNs) while they ignore the high-order correlations between multi-modal physiological signals. Moreover, most HGFNs suffer from the over-smoothing problem, causing failures in deeper networks. To address these issues, this paper proposes a deep hypergraph convolutional network called DHFusion for multi-modal physiological signal fusion. Specially, DHFusion designs deep hypergraph convolution (DHGCN) layer to effectively integrate multiple physiological signals and obtain multi-modal features. DHFusion then introduces JK readout layer to enable multi-layer DHGCN to capture deep multi-modal features based on high-order correlations. DHFusion effectively solves the over-smoothing of HGFNs and performs well in representing deep multi-modal features. Experimental results over the state-of-the-art methods on two benchmark datasets demonstrate the effectiveness of the proposed method. Yuanhang Shen, C. L. Philip Chen, Tong Zhang 0015 |
SMC | 3 |
| 2024 | OneDConv: Generalized Convolution for Transform-Invariant RepresentationabstractConvolutional Neural Networks (CNNs) have ex-hibited great power in various vision tasks. However, the lack of transform-invariant property limits their further applications in complicated real-world scenarios. In this work, we pro-posed a novel generalized one-dimension convolutional operator (OneDConv), which dynamically transforms the convolution kernels based on the input features in a computationally and parametrically efficient manner. The proposed operator can extract the transform-invariant features naturally. It improves the robustness and generalization of convolution without sac-rificing the performance of common images. The proposed OneDConv operator can substitute the vanilla convolution. Thus, it can readily be incorporated into popular convolutional architectures, supporting end-to-end training. Empirical evaluations on popular benchmarks reveal OneDConv's superior performance over the standard convolution and competitive models in handling canonical and distorted images. Haohan Weng, C. L. Philip Chen, Ke Yi 0003, Haiqi Liu, Tong Zhang 0015 |
SMC | 5 |
| 2024 | Dual-Domain Attention Based Adaptive Graph Convolutional Network for EEG Emotion RecognitionabstractThe asymmetry of emotional responses is observed in electroencephalogram (EEG) of different frequency bands across various spatial brain regions in neuroscience research. Many prior works have primarily emphasized the dependencies among channels in the spatial domain, neglecting the dynamic interaction of EEG in both spatial and frequency domains, which may limit the performance of EEG emotion recognition. To address these issues, we propose the dual-domain attention based adaptive graph convolutional network (DDA-AGCN) for EEG emotion recognition. Specifically, we propose the lightweight dual-domain attention mechanism (DDA) based on random vector similarity measurement and the squeezeexcitation technique to capture important characteristics in the channel and frequency domain respectively. Furthermore, the adaptive graph convolutional network (AGCN) is utilized to adaptively filter and refine low signal-to-noise ratio EEG data, while also learning the dynamic connectivity patterns among important EEG channels and extracting higher-level abstract features for emotion recognition tasks. To validate the effectiveness of the proposed method, experimental comparisons were conducted on SEED, SEED-IV, and MPED. The experimental results show that our method achieves highly competitive classification performance compared to existing methods. Moreover, under fair comparison, the DDA demonstrates better performance and computational efficiency than self-attention. Tie Xu, Tong Zhang 0015, Bianna Chen, C. L. Philip Chen |
SMC | 2 |
| 2024 | Multi-Granularity Temporal-Spectral Representation Learning for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) captures emotional information from speech signals to recognize users' emotional states, which plays a crucial role in conversational human-computer interaction. Most SER researches focus on ex-ploiting emotional information from global temporal or spectral features, but it may neglect detailed emotion-related information such as phonemes and syllables. To address this problem, this paper proposes a multi-granularity temporal-spectral representation learning (MG-TSRL) network for speech emotion recognition tasks. Specifically, MG-TSRL extracts different temporal features in phonetic, syllabic, and sentential granular-ity from spectrograms to retain more detailed emotional-related information. It then designs multilayer emotion-aware units to capture emotion-related frequency patterns and obtain deep spectrum features at each temporal granularity feature. MG-TSRL further introduces a fast broad learning system and feeds deep temporal-spectral features to it to obtain more accurate emotions. MG-TSRL gradually achieves effective temporal-spectral representation learning through multi-granularity temporal features and multilayer frequency pattern learning. The state-of-the-art results on the CASIA, RAVDESS, and SAVEE datasets are respectively 95.17%, 92.78%, and 87.50% in unweighted accuracy, demonstrating the effectiveness of MG-TSRL in speech emotion recognition. Zhichen Yuan, C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
SMC | 4 |
| 2024 | STFGCN: Spatial-temporal fusion graph convolutional network for traffic prediction
Hao Li 0100, Jie Liu 0002, Shi-Yuan Han, Jin Zhou 0003, Tong Zhang 0015, C. L. Philip Chen |
Expert Syst. Appl. | 5 |
| 2024 | Deuce: Dual-diversity Enhancement and Uncertainty-awareness for Cold-start Active LearningabstractAbstract Cold-start active learning (CSAL) selects valuable instances from an unlabeled dataset for manual annotation. It provides high-quality data at a low annotation cost for label-scarce text classification. However, existing CSAL methods overlook weak classes and hard representative examples, resulting in biased learning. To address these issues, this paper proposes a novel dual-diversity enhancing and uncertainty-aware (Deuce) framework for CSAL. Specifically, Deuce leverages a pretrained language model (PLM) to efficiently extract textual representations, class predictions, and predictive uncertainty. Then, it constructs a Dual-Neighbor Graph (DNG) to combine information on both textual diversity and class diversity, ensuring a balanced data distribution. It further propagates uncertainty information via density-based clustering to select hard representative instances. Deuce performs well in selecting class-balanced and hard representative data by dual-diversity and informativeness. Experiments on six NLP datasets demonstrate the superiority and efficiency of Deuce. C. L. Philip Chen, Shuzhen Li, Tong Zhang 0015 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | GDDN: Graph Domain Disentanglement Network for Generalizable EEG Emotion RecognitionabstractCross-subject EEG emotion recognition suffers a major setback due to high inter-subject variability in emotional responses. Many prior studies have endeavored to alleviate the inter-subject discrepancies of EEG feature distributions, ignoring the variable EEG connectivity and prediction deviation caused by individual differences, which may cause poor generalization to the unseen subject. This paper proposes a graph domain disentanglement network (GDDN) to generalize EEG emotion recognition across subjects in terms of EEG connectivity, representation, and prediction. More specifically, a graph domain disentanglement module is proposed to extract common-specific characteristics on both EEG graph connectivity and graph representation, enabling a more comprehensive network transferability to the unseen individual. Meanwhile, to strengthen stable emotion prediction capability, a domain-adaptive classifier aggregation module is developed to facilitate adaptive emotional prediction for the unseen individual conditioned on the domain weights of the input individuals. Finally, an auxiliary supervision module is imposed to alleviate the domain discrepancy and reduce information loss during the disentanglement learning. Extensive experiments on three public EEG emotion datasets, i.e., SEED, SEED-IV, and MPED, validate the superior generalizability of GDDN compared with the state-of-the-art methods. Bianna Chen, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | CiABL: Completeness-Induced Adaptative Broad Learning for Cross-Subject Emotion Recognition With EEG and Eye Movement SignalsabstractAlthough multimodal physiological data from the central and peripheral nervous systems can objectively respond to human emotional states, the individual differences caused by non-stationary and low signal-to-noise properties bring several challenges to cross-subject emotion recognition tasks. Many previous studies usually focused on learning high correlation information between different modalities, which easily leads to incomplete descriptions of different physiological signals and difficulties in aligning critical emotional information. To tackle these challenges, this paper proposes a novel multimodal emotion recognition model for improving the generalization performance to unseen target domain subjects, termed Completeness-induced Adaptative Broad Learning (CiABL). The proposed CiABL can gradely explore the completeness modality representation that encompasses both modality-relevant and modality-independent information, avoiding the loss of performance due to spurious correlations from different modalities. Subsequently, a well-designed weighted representation distribution alignment mechanism of CiABL can appropriately align the marginal and conditional distributions to reduce the influences of individual differences greatly. Extensive experiments on the SEED and SEED-FRA datasets demonstrate the effectiveness and generalization of the proposed CiABL, which outperforms current state-of-the-art methods. In addition, CiABL can precisely quantify the importance of global features to properly explain the modality contribution and averaged activation patterns of the brain under cross-subject emotion recognition tasks. Xin-Rong Gong, C. L. Philip Chen, Bin Hu 0001, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Gusa: Graph-Based Unsupervised Subdomain Adaptation for Cross-Subject EEG Emotion RecognitionabstractEEG emotion recognition has been hampered by the clear individual differences in the electroencephalogram (EEG). Nowadays, domain adaptation is a good way to deal with this issue because it aligns the distribution of data across subjects. However, the performance for EEG emotion recognition is limited by the existing research, which mainly focuses on the global alignment between the source domain and the target domain and ignores much fine-grained information. In this study, we propose a method called Graph-based Unsupervised Subdomain Adaptation (Gusa), which simultaneously aligns the distribution between the source and target domains in a fine-grained way from both the channel and emotion subdomains. Gusa employs three modules, such as the Node-wise Domain Constraints Module to align each EEG channel and obtain a domain-variant representation, the Class-level Distribution Constraints Module, and the Emotion-wise Domain Constraints Module, to collect more fine-grained information, create more discriminative representations for each emotion, and lessen the impact of noisy emotion labels. The studies on the SEED, SEED-IV, and MPED datasets demonstrate that Gusa significantly improves the ability of EEG to recognize emotions and can extract more granular and discriminative representations for EEG. C. L. Philip Chen, Bianna Chen, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Fine-Grained Interpretability for EEG Emotion Recognition: Concat-Aided Grad-CAM and Systematic Brain Functional NetworkabstractEEG emotion recognition plays a significant role in various mental health services. Deep learning-based methods perform excellently, but still suffer from interpretability. Although methods such as Gradient-weighted Class Activation Mapping(Grad-CAM) can cope with the above problem, their coarse granularity cannot accurately reveal the mechanism to promote emotional intelligence. In this paper, fine-grained interpretability is proposed, called Concat-aided Grad-CAM. Specifically, the multi-level feature mapping before the fully connected layer is concatenated to obtain the gradients of the target concept so that the discriminant information can be directly located in the high-precision area. Unlike coarse-grained interpretability methods applied in EEG emotion recognition, it can accurately highlight the EEG channels related to emotion rather than an obscure area. In addition, a systematic brain functional network is proposed to reveal the relationship between those channels and to further improve emotion recognition performance. The channels with greater contributions are connected, and those connections are learned by dynamic graph convolutional networks, while the others are independent to eliminate interference. Experiments on two EEG emotion recognition datasets manifest that Concat-aided Grad-CAM can be interpreted by the fine-grained. In addition, it has been shown that the learned brain functional network can improve the performance of the baselines. Significantly, the experiment results achieve state-of-the-art performance in subject-dependent experiments. Bingxiu Liu, Jifeng Guo 0002, C. L. Philip Chen, Xia Wu 0001, Tong Zhang 0015 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Emotion Recognition in Conversation Based on a Dynamic Complementary Graph Convolutional NetworkabstractEmotion recognition in conversation (ERC) is a widely used technology in both affective dialogue bots and dialogue recommendation scenarios, where motivating a system to correctly recognize human emotions is crucial. Uncovering as much contextual information as possible with a limited amount of dialogue information is essential for eventually identifying the correct emotion of each sentence. The integration of contextual information using the existing approaches often results in inadequate access to information or information redundancy. Deeply integrating the different knowledge behind utterances is also difficult. Therefore, to address these problems, we propose a dynamic complementary graph convolutional network (DCGCN) for conversational emotion recognition. Our approach uses commonsense knowledge to complement the contextual information contained in utterances and enrich the extracted conversation information. We creatively propose the concept of utterance density to prevent redundancy and the loss of utterance information in context-dependent contextual information modeling cases. An utterance dependency structure is dynamically determined by the utterance density, and the contextual information is fully integrated into each sentence representation. We evaluate our proposed model in extensive experiments conducted on four public benchmark datasets that are commonly used for ERC. The results demonstrate the effectiveness of the DCGCN, which achieves competitive results in terms of well-known evaluation metrics. Our code is available athttps://github.com/Tars-is-a-robot/Conversational-emotion-recognition.git. Zhenyu Yang 0002, Yuhu Cheng 0001, Tong Zhang 0015, Xuesong Wang 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Cross-Cultural Emotion Recognition With EEG and Eye Movement Signals Based on Multiple Stacked Broad Learning SystemabstractWith increasing social globalization, interaction between people from different cultures has become more frequent. However, there are significant differences in the expression and comprehension of emotions across cultures. Therefore, developing computational models that can accurately identify emotions among different cultures has become a significant research problem. This study aims to investigate the similarities and differences in emotion cognition processes in different cultural groups by employing a fusion of electroencephalography (EEG) and eye movement (EM) signals. Specifically, an effective adaptive region selection method is proposed to investigate the most emotion-related activated brain regions in different groups. By selecting these commonly activated regions, we can eliminate redundant features and facilitate the development of portable acquisition devices. Subsequently, the multiple stacked broad learning system (MSBLS) is designed to explore the complementary information of EEG and EM features and the effective emotional information still contained in the residual value. The intracultural subject-dependent (ICSD), intracultural subject-independent, and cross-cultural subject-independent (CCSI) experiments have been conducted on the SEED-CHN, SEED-GER, and SEED-FRA datasets. Extensive experiments manifest that MSBLS achieves superior performance compared with current state-of-the-art methods. Moreover, we discover that some brain regions (the anterior frontal, temporal, and middle parieto-occipital lobes) and Gamma frequency bands show greater activation during emotion cognition in diverse cultural groups. Xin-Rong Gong, C. L. Philip Chen, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | SIA-Net: Sparse Interactive Attention Network for Multimodal Emotion RecognitionabstractMultimodal emotion recognition (MER) integrates multiple modalities to identify the user's emotional state, which is the core technology of natural and friendly human–computer interaction systems. Currently, many researchers have explored comprehensive multimodal information for MER, but few consider that comprehensive multimodal features may contain noisy, useless, or redundant information, which interferes with emotional feature representation. To tackle this challenge, this article proposes a sparse interactive attention network (SIA-Net) for MER. In SIA-Net, the sparse interactive attention (SIA) module mainly consists of intramodal sparsity and intermodal sparsity. The intramodal sparsity provides sparse but effective unimodal features for multimodal fusion. The intermodal sparsity adaptively sparses intramodal and intermodal interactive relations and encodes them into sparse interactive attention. The sparse interactive attention with a small number of nonzero weights then act on multimodal features to highlight a few but important features and suppress numerous redundant features. Furthermore, the intramodal sparsity and intermodal sparsity are deep sparse representations that make unimodal features and multimodal interactions sparse without complicated optimization. The extensive experimental results show that SIA-Net achieves superior performance on three widely used datasets. Shuzhen Li, Tong Zhang 0015, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Adaptive 3DCNN-Based Interpretable Ensemble Model for Early Diagnosis of Alzheimer's DiseaseabstractAdaptive interpretable ensemble model based on three-dimensional Convolutional Neural Network (3DCNN) and Genetic Algorithm (GA), i.e., 3DCNN+EL+GA, was proposed to differentiate the subjects with Alzheimer's Disease (AD) or Mild Cognitive Impairment (MCI) and further identify the discriminative brain regions significantly contributing to the classifications in a data-driven way. Plus, the discriminative brain sub-regions at a voxel level were further located in these achieved brain regions, with a gradient-based attribution method designed for CNN. Besides disclosing the discriminative brain sub-regions, the testing results on the datasets from the Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Open Access Series of Imaging Studies (OASIS) indicated that 3DCNN+EL+GA outperformed other state-of-the-art deep learning algorithms and that the achieved discriminative brain regions (e.g., the rostral hippocampus, caudal hippocampus, and medial amygdala) were linked to emotion, memory, language, and other essential brain functions impaired early in the AD process. Future research is needed to examine the generalizability of the proposed method and ideas to discern discriminative brain regions for other brain disorders, such as severe depression, schizophrenia, autism, and cerebrovascular diseases, using neuroimaging. Dan Pan 0001, Genqiang Luo, An Zeng, Chao Zou, Haolin Liang, Tong Zhang 0015, Baoyao Yang |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2024 | Adaptive Dual-Space Network With Multigraph Fusion for EEG-Based Emotion RecognitionabstractMost of the work on electroencephalogram (EEG)-based emotion recognition aims to extract the distinguishing features from high-dimensional EEG signals, ignoring the complementarity of information between EEG latent space and graph space. Furthermore, the influence of brain connectivity on emotions encompasses both physical structure and functional connectivity, which may have varying degrees of importance for different individuals. To address these issues, this article introduces an adaptive dual-space network (ADS-Net) with multigraph fusion aimed at capturing more comprehensive information by integrating dual-space representations. Specifically, ADS-Net models the spatial correlation of EEG channels in graph topological space, while exploring long-range dependencies and frequency relationships from EEG data in latent space. Subsequently, these representations are adaptively combined through an innovative gated fusion approach to extract complementary corepresentations. Moreover, drawing on the principles of brain connectivity theory, the proposed method constructs a multigraph to indicate the associativity of EEG channels. To further capture individual differences, an adaptive multigraph fusion mechanism is developed for the dynamic integration of physical and functional connectivity graphs. When compared to state-of-the-art methods, the superior experimental results underscore the effectiveness and broad applicability of the proposed method. Mengqing Ye, C. L. Philip Chen, Wenming Zheng, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | TT-GCN: Temporal-Tightly Graph Convolutional Network for Emotion Recognition From GaitsabstractThe human gait reflects substantial information about individual emotions. Current gait emotion recognition methods focus on capturing gait topology information and ignore the importance of fine-grained temporal features. This article proposes the temporal-tightly graph convolutional network (TT-GCN) to extract temporal features. TT-GCN comprises three significant mechanisms: the causal temporal convolution network (casual-TCN), the walking direction recognition auxiliary task, and the feature mapping layer. To obtain tight temporal dependencies and enhance the relevance among gait periods, the causal-TCN is introduced. Based on the assumption of emotional consistency in the walking directions, the auxiliary task is proposed to enhance the ability of fine-grained feature extraction. Through the feature mapping layer, affective features can be mapped into the appropriate representation and fused with deep learning features. TT-GCN shows the best performance across five comprehensive metrics. All experimental results verify the necessity and feasibility of exploring fine-grained temporal feature extraction. Tong Zhang 0015, Yelin Chen, Shuzhen Li, Xiping Hu, C. L. Philip Chen |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Robust Saliency-Aware Distillation for Few-Shot Fine-Grained Visual RecognitionabstractRecognizing novel sub-categories with scarce samples is an essential and challenging research topic in computer vision. Existing literature addresses this challenge by employing local-based representation approaches, which may not sufficiently facilitate meaningful object-specific semantic understanding, leading to a reliance on apparent background correlations. Moreover, they primarily rely on high-dimensional local descriptors to construct complex embedding space, potentially limiting the generalization. To address the above challenges, this article proposes a novel model, Robust Saliency-aware Distillation (RSaD), for few-shot fine-grained visual recognition. RSaD introduces additional saliency-aware supervision via saliency detection to guide the model toward focusing on the intrinsic discriminative regions. Specifically, RSaD utilizes the saliency detection model to emphasize the critical regions of each sub-category, providing additional object-specific information for fine-grained prediction. RSaD transfers such information with two symmetric branches in a mutual learning paradigm. Furthermore, RSaD exploits inter-regional relationships to enhance the informativeness of the representation and subsequently summarize the highlighted details into contextual embeddings to facilitate the effective transfer, enabling quick generalization to novel sub-categories. The proposed approach is empirically evaluated on three widely used benchmarks, demonstrating its superior performance. Haiqi Liu, C. L. Philip Chen, Xin-Rong Gong, Tong Zhang 0015 |
IEEE Trans. Multim. | 4 |
| 2024 | A Broad Generative Network for Two-Stage Image OutpaintingabstractImage outpainting is a challenge for image processing since it needs to produce a big scenery image from a few patches. In general, two-stage frameworks are utilized to unpack complex tasks and complete them step-by-step. However, the time consumption caused by training two networks will hinder the method from adequately optimizing the parameters of networks with limited iterations. In this article, a broad generative network (BG-Net) for two-stage image outpainting is proposed. As a reconstruction network in the first stage, it can be quickly trained by utilizing ridge regression optimization. In the second stage, a seam line discriminator (SLD) is designed for transition smoothing, which greatly improves the quality of images. Compared with state-of-the-art image outpainting methods, the experimental results on the Wiki-Art and Place365 datasets show that the proposed method achieves the best results under evaluation metrics: the Fréchet inception distance (FID) and the kernel inception distance (KID). The proposed BG-Net has good reconstructive ability with faster training speed than those of deep learning-based networks. It reduces the overall training duration of the two-stage framework to the same level as the one-stage framework. Furthermore, the proposed method is adapted to image recurrent outpainting, demonstrating the powerful associative drawing capability of the model. Zongyan Zhang, Haohan Weng, Tong Zhang 0015, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | FACExplainer: Generating Model-faithful Explanations for Graph Neural Networks Guided by Spatial InformationabstractGraph neural networks (GNNs) have been widely applied in various decision-crucial fields, where accurate predictions with high interpretability are desired. Thus, numerous post-hoc explainers for GNNs have been proposed. However, some prioritize human-intelligible explanations through graph rules, such as the connection rule, which undermines the explanation’s faithfulness to the model. This paper proposes an innovative method, FACExplainer, that re-examines the role of spatial information within GNNs for generating model-faithful explanations. FACExplainer employs activation maps from the last graph convolution to narrow down a compact search space. Our approach further identifies the subgraph that maximizes mutual information as the explanation, eliminating the need for domain-specific knowledge about the downstream task. Empirical analysis of FACExplainer on seven benchmark datasets with three classical GNNs reveals significantly improved explanation quality while consuming less time when compared to leading explainers. The source code of FACExplainer is freely available at https://github.com/HuaYangttt/Facexplainer/. C. L. Philip Chen, Bianna Chen, Tong Zhang 0015 |
BIBM | 4 |
| 2023 | Learn and Sample Together: Collaborative Generation for Graphic Design LayoutabstractIn the process of graphic layout generation, user specifications including element attributes and their relationships are commonly used to constrain the layouts (e.g.,"put the image above the button''). It is natural to encode spatial constraints between elements using a graph. This paper presents a two-stage generation framework: a spatial graph generator and a subsequent layout decoder which is conditioned on the previous output graph. Training the two highly dependent networks separately as in previous work, we observe that the graph generator generates out-of-distribution graphs with a high frequency, which are unseen to the layout decoder during training and thus leads to huge performance drop in inference. To coordinate the two networks more effectively, we propose a novel collaborative generation strategy to perform round-way knowledge transfer between the networks in both training and inference. Experiment results on three public datasets show that our model greatly benefits from the collaborative generation and has achieved the state-of-the-art performance. Furthermore, we conduct an in-depth analysis to better understand the effectiveness of graph condition modeling. Haohan Weng, Danqing Huang, Tong Zhang 0015, Chin-Yew Lin |
IJCAI | 3 |
| 2023 | Siamese labels auxiliary learning
Wenrui Gan, Zhulin Liu, C. L. Philip Chen, Tong Zhang 0015 |
Inf. Sci. | 4 |
| 2023 | MIA-Net: Multi-Modal Interactive Attention Network for Multi-Modal Affective AnalysisabstractWhen a multi-modal affective analysis model generalizes from a bimodal task to a trimodal or multi-modal task, it is usually transformed into a hierarchical fusion model based on every two pairwise modalities, similar to a binary tree structure. This easily leads to large growth in model parameters and computation as the number of modalities increases, which limits the model's generalization. Moreover, many multi-modal fusion methods ignore that different modalities contribute differently to affective analysis. To tackle these challenges, this article proposes a general multi-modal fusion model that supports trimodal or multi-modal affective analysis tasks, called Multi-modal Interactive Attention Network (MIA-Net). Instead of treating different modalities equally, MIA-Net takes the modality that contributes the most to emotion as the main modality and the others as auxiliary modalities. MIA-Net introduces multi-modal interactive attention modules to adaptively select the important information of each auxiliary modality one by one to improve the main-modal representation. Moreover, MIA-Net enables quick generalization to trimodal or multi-modal tasks through stacking multiple MIA modules, which maintains efficient training and only requires linear computation and stable parameter counts. Experimental results of the transfer, generalization, and efficiency experiments on the widely-used datasets demonstrate the effectiveness and generalization of the proposed method. Shuzhen Li, Tong Zhang 0015, Bianna Chen, C. L. Philip Chen |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Semantic Learning for Facial Action Unit DetectionabstractThis article proposes semantic embedding for image transformers (SEiTs) to explore semantic features of facial morphology in the action unit (AU) detection task. The conventional approaches typically rely on external information (e.g., facial landmarks) to obtain the location of facial components, whereas the SEiT can learn morphological features intrinsically from the face image. The pre-training task, namely semantic masked facial image modeling (SMFIM), aims to actively obtain facial morphological information. The pixels of the input facial image are randomly erased with semantic masks (e.g., nose, eyes, eyebrows, mouth, and lip). The embedding model tries to predict the presence of facial components for the input image that can learn semantic representations of the face simultaneously. The learned semantic embeddings are fed to transformer blocks, which enable global interaction between semantic elements. The SEiT integrates facial morphological information and global interaction characters, appropriate for AU detection. The experiments are conducted on the Binghamton-Pittsburgh 4D (BP4D) dataset and Denver intensity of spontaneous facial action (DISFA) dataset, and the results demonstrate the effectiveness of the proposed SEiT. Xuehan Wang, C. L. Philip Chen, Haozhang Yuan, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | MIAR: Interest-Activated News Recommendation by Fusing Multichannel InformationabstractThe different news clicked by users reflects the diverse interests of users. Most of the existing news recommendation methods do not consider the interaction with candidate news in the process of modeling user interest representation. This method makes it challenging to precisely match candidate news to specific user interests. We propose a user interest activation recommendation method that fuses multichannel information—MIAR. It utilizes the word embedding of the user’s historical clicked news and the news title embedding generated by aggregation and interacts with the candidate news, respectively, to better match the candidate news with the user’s interests. Our proposed method contains two frameworks (interactive framework and distributed framework). In the interactive framework, we propose a user multichannel interest modeling framework MIF from the word embedding level of news headlines to capture more semantic cues related to user interests. In the distributed framework, we design a candidate-aware interest activation module TAR from the news embedding representation level obtained by attention aggregation. It uses different candidate news vectors to adjust the user representations learned from the user’s historical reading records. This allows the model to build candidate-guided user representations to accurately match candidate news to parts of user interests that are relevant to the candidate news. Finally, we effectively assign the weights of the two frame scores so that the models can fuse better. Extensive experiments on the MIND news recommendation dataset demonstrate the effectiveness of our method. Zhenyu Yang 0002, Laiping Cui, Xuesong Wang 0001, Tong Zhang 0015, Yuhu Cheng 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | Recommendation Model Based on Enhanced Graph Convolution That Fuses Review PropertiesabstractIn rating prediction research, how to capture user and item features from review text is a key to improving model prediction accuracy. The sparsity of review text and the accuracy of the description of items in the review text make it difficult to obtain accurate feature representations of users and items by modeling the text content alone. Therefore, it is important to evaluate the usefulness of the reviews at first because not all review texts are valuable. How to analyze the usefulness of a review is a key to modeling the review text. The way previous models use attention to inscribe semantic weights on the review text is not sufficient to indicate the degree of usefulness of a review, so we suggest adding property information to model reviews. Based on this, we propose an interaction recommendation model that is based on enhanced graph convolution and fuses review properties (PGIR), which incorporates property information into text modeling by different activations and matches useful property feature interaction pairs for review text in a self-supervised manner. This allows the model to obtain an accurate feature representation of the review text. In addition, we analyze the high-order connectivity among user–item pairs. Then, we design an enhanced graph convolution method to capture the collaborative signals between users and items and model the dynamic features of users and items on this basis. After extensive experiments conducted on five standard datasets based on Amazon, the results show that the PGIR model achieves a substantial improvement over existing state-of-the-art models in terms of rating prediction. In addition, we experimentally demonstrate the superiority of our proposed property activation method, which further improves the rating prediction performance of the PGIR model. Zhenyu Yang 0002, Yuhu Cheng 0001, Tong Zhang 0015, Xuesong Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | AIA-Net: Adaptive Interactive Attention Network for Text-Audio Emotion RecognitionabstractEmotion recognition based on text-audio modalities is the core technology for transforming a graphical user interface into a voice user interface, and it plays a vital role in natural human-computer interaction systems. Currently, mainstream multimodal learning research has designed various fusion strategies to learn intermodality interactions but hardly considers that not all modalities play equal roles in emotion recognition. Therefore, the main challenge in multimodal emotion recognition is how to implement effective fusion algorithms based on the auxiliary structure. To address this problem, this article proposes an adaptive interactive attention network (AIA-Net). In AIA-Net, text is treated as a primary modality, and audio is an auxiliary modality. AIA-Net adapts to textual and acoustic features with different dimensions and learns their dynamic interactive relations in a more flexible way. The interactive relations are encoded as interactive attention weights to focus on the acoustic features that are effective for textual emotional representations. AIA-Net performs well in adaptively assisting the textual emotional representation with the acoustic emotional information. Moreover, multiple collaborative learning (co-learning) layers of AIA-Net achieve multiple multimodal interactions and the deep bottom-up evolution of emotional representations. Experimental results on three benchmark datasets demonstrate the great effectiveness of the proposed method over the state-of-the-art methods. Tong Zhang 0015, Shuzhen Li, Bianna Chen, Haozhang Yuan, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2023 | Transfer Learning-Based Collaborative Multiview ClusteringabstractCollaborative multiview clustering methods can efficiently realize the view fusion by exploring complementary and consistent information among multiple views. However, these studies ignore all the differences between multiple views in fusion. In fact, in the multiview clustering, the data are diverse from view to view. The larger the difference between any two views is, the more the fusion of these views is required. Moreover, a global tradeoff parameter is generally adopted to restrain the penalty related to the disagreement of all views, which is often defined empirically. Inspired by the idea of transfer learning, a series of novel collaborative multiview clustering algorithms are proposed to tackle these challenges. In the most basic one, each view performs clustering independently and learns from others to improve its own clustering performance, in which a global learning factor is defined to control the interaction between multiple views. The fuzzy memberships are regarded as the important knowledge to provide guidance between views, and the consensus constraint is defined to ensure the consistent partitions of all views. In addition, the local adaptive learning factors between any two views instead of a global fixed one are adopted in an improved version to emphasize the difference between views, and the adjustment strategy for the learning factor is further designed to guarantee the stability of multiview clustering without the influence of initial values. Finally, to identify the significance of different views to the clustering, the extended versions are excavated with the assignment of view weights and the maximum entropy regularization technique is employed to optimize the weights. Experiments on various real-world multiview datasets verify the superiority of the presented approaches. Xiangdao Liu, Jin Zhou 0003, C. L. Philip Chen, Tong Zhang 0015, Yuehui Chen, Shi-Yuan Han, Tao Du 0002, Ke Ji, Kun Zhang 0013 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2023 | Random Feature-Based Collaborative Kernel Fuzzy Clustering for Distributed Peer-to-Peer NetworksabstractKernel clustering has the ability to get the inherent nonlinear structure of the data. But the high computational complexity and the unknown representation of the kernel space make it unavailable for the data clustering in distributed peer-to-peer (P2P) networks. To solve this issue, we propose a new series of random feature-based collaborative kernel clustering algorithms in this article. In the most basic algorithm, each node in a distributed P2P network first maps its data into a low-dimensional random feature space with the approximation of the given kernel by using the random Fourier feature mapping method. Then, each node independently searches the clusters with its local data and the collaborative knowledge from its neighbor nodes, and the distributed clustering is performed among all network nodes until reaching the global consensus result, i.e., all nodes have the same cluster centers. In addition, an improved version is designed with assignment of feature weights, which is optimized by the maximum-entropy technique to extract important features for the cluster identification. What’s more, to relief the impact of different kernel functions and related parameters on clustering results, the combination of multiple kernels rather than a single kernel is adopted for the low-dimensional approximation, and the optimized weights are assigned to provide the guidance on the choice of the kernels and their parameters and discover significant features at the same time. Experiments on synthetic and real-world datasets show that the proposed methods achieve similar and even better results than the traditional kernel clustering methods on various performance metrics, including the average classification rate, the average normalized mutual information, and the average adjusted rand index. More importantly, the low-dimensional random features approximated to kernels and the distributed clustering mechanism adopted in these methods bring the greatly lower temporal complexity. Yingxu Wang 0002, Shi-Yuan Han, Jin Zhou 0003, Long Chen 0001, C. L. Philip Chen, Tong Zhang 0015, Zhulin Liu, Lin Wang 0004, Yuehui Chen |
IEEE Trans. Fuzzy Syst. | 6 |
| 2023 | A VAE-Based User Preference Learning and Transfer Framework for Cross-Domain RecommendationabstractThe core idea of cross-domain recommendation is to alleviate the problem of data scarcity. Previous methods have made brilliant successes. However, many of them mainly focus on learning an ideal mapping function across-domains, ignoring the user preferences within a specific domain, which leads to suboptimal results. In this paper, we propose a Cross-Domain Recommendation Variational AutoEncoder framework (CDRVAE), a novel extension of a variational autoencoder on cross-domain recommendations for user behaviour distribution modeling. It applies a new hybrid architecture of VAE as the backbone and simultaneously constructs two information flows, within-domain and cross-domain modeling. For the former, an asymmetric codec structure is designed to reconstruct preference distribution from domain-specific latent factors. To relieve the posterior collapse dilemma, a combined prior is employed to increase the distribution complexity. The equivalent transition by a transformation matrix and the unobserved interaction generation by cross-domain reconstruction contribute to the latter. We combine all the above components for the more accurate and reliable user features. Extensive experiments are conducted on three public benchmark datasets to validate the effectiveness of the proposed CDRVAE. Experimental results demonstrate that CDRVAE is consistently superior to other state-of-the-art alternative baseline models. Tong Zhang 0015, Chen Chen 0128, Dan Wang 0002, Jie Guo 0008, Bin Song 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Constructing Microstructural Evolution System for Cement Hydration From Observed Data Using Deep LearningabstractCement has been widely used in civil engineering directly and plays a critical role in cement-based materials, e.g., concrete. As the microstructural evolution of cement hydration predominates the final physical properties, an accurate simulation of hydration is highly required to enable scientists to evaluate the performance and help design new cementitious materials. However, despite significant effort and progress, a satisfactory model to realistically and accurately simulate the evolution of three-dimensional (3-D) microstructure has not yet to be constructed, mainly because cement hydration is one of the most complex phenomena in material science. In this work, a novel near-realistic microstructural model is proposed to simulate the cement hydration system using deep learning and cellular automata. It is designed to break through the bottleneck of fidelity to real microstructural evolution. The dynamical system is constructed based on a 3-D cellular automaton, in which behavior is controlled by deep neural networks distilled from microstructural images. In addition, a dynamic stratified sampling method with variable capacity is proposed to ensure the representativeness of samples for reducing the computation cost of training. Experiments manifest that the simulated hydration is in accordance with the actual development in different aspects, such as near-realistic microstructure and approximate process. Furthermore, the constructed system also demonstrates promising generalization capability even under various conditions. Jifeng Guo 0002, C. L. Philip Chen, Lin Wang 0004, Bo Yang 0001, Tong Zhang 0015 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | Multi-Person Pose Estimation in the Wild: Using Adversarial Method to Train a Top-Down Pose Estimation NetworkabstractRecent studies estimate human anatomical key points through the single monocular image, in which multichannel heatmaps are the key factor in determining the quality of human pose estimation. Multichannel heatmaps can efficiently handle the image-to-coordinate mapping task and the processing of semantic features. Most methods ignore physical constraints and internal relationships of human body parts, which easily misclassify left and right symmetrical parts as similar features. Some studies use RNNs on the top to incorporate priors about the structure of pose components and body configuration. Therefore, a novel top-down convolutional network is proposed to consider these priors during training, which can improve the robustness under complex field conditions in the wild. In order to learn the prior knowledge of human pose configuration, the hierarchy of fully convolutional networks (discriminator) is used to distinguish real poses from fake ones. Consequently, the pose network is inclined to make a pose estimation that the discriminator misjudges as true, which is reasonable in complex situations. The performance of the method is experimentally validated by pose estimation on the MS COCO human key point detection task. The proposed approach outperforms the original method and generates robust pose predictions, demonstrating efficiency by using adversarial learning. Tong Zhang 0015, Jingxiang Lian, Jingtao Wen, C. L. Philip Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Adaptive Broad Learning Neural Network for Fault-Tolerant Control of 2-DOF Helicopter SystemsabstractThis study is aimed to design a fault-tolerant control using a broad learning neural network (BLNN) for a two-degree-of-freedom (2-DOF) nonlinear helicopter system. Compared with the conventional radial basis function neural network, the BLNN can approximate uncertainties and unknown functions with smaller tracking errors by adding incremental and enhancement nodes. Considering possible actuator faults in during actual application, an adaptive auxiliary parameter is established to prevent their effects on control. Through direct Lyapunov method, the stability and convergence of the closed-loop system are analyzed. The results from simulations and experiments conducted on a 2-DOF helicopter laboratory platform of Quanser demonstrate the validity and feasibility of the proposed control method. Zhijia Zhao 0002, Weitian He, Tao Zou 0001, Tong Zhang 0015, C. L. Philip Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Self-labeling with feature transfer for speech emotion recognition
Guihua Wen, Huiqiang Liao, Pengcheng Wen, Tong Zhang 0015, Sande Gao |
Knowl. Based Syst. | 5 |
| 2022 | GCB-Net: Graph Convolutional Broad Network and Its Application in Emotion RecognitionabstractIn recent years, emotion recognition has become a research focus in the area of artificial intelligence. Due to its irregular structure, EEG data can be analyzed by applying graphical based algorithms or models much more efficiently. In this work, a Graph Convolutional Broad Network (GCB-net) was designed for exploring the deeper-level information of graph-structured data. It used the graph convolutional layer to extract features of graph-structured input and stacks multiple regular convolutional layers to extract relatively abstract features. The final concatenation utilized the broad concept, which preserves the outputs of all hierarchical layers, allowing the model to search features in broad spaces. To improve the performance of the proposed GCB-net, the broad learning system (BLS) was applied to enhance its features. For comparison, two individual experiments were conducted to examine the efficiency of the proposed GCB-net based on the SJTU emotion EEG dataset (SEED) and DREAMER dataset respectively. In SEED, compared with other state-of-art methods, the GCB-net could better promote the accuracy (reaching 94.24 percent) on the DE feature of the all-frequency band. In DREAMER dataset, GCB-net performed better than other models with the same setting. Furthermore, the GCB-net reached high accuracies of 86.99, 89.32 and 89.20 percent on dimensions of Valence, Arousal and Dominance respectively. The experimental results showed the robust classifying ability of the GCB-net and BLS in EEG emotion recognition. Tong Zhang 0015, Xuehan Wang, Xiangmin Xu 0001, C. L. Philip Chen |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | BMT-Net: Broad Multitask Transformer Network for Sentiment AnalysisabstractSentiment analysis uses a series of automated cognitive methods to determine the author's or speaker's attitudes toward an expressed object or text's overall emotional tendencies. In recent years, the growing scale of opinionated text from social networks has brought significant challenges to humans' sentimental tendency mining. The pretrained language model designed to learn contextual representation achieves better performance than traditional learning word vectors. However, the existing two basic approaches for applying pretrained language models to downstream tasks, feature-based and fine-tuning methods, are usually considered separately. What is more, different sentiment analysis tasks cannot be handled by the single task-specific contextual representation. In light of these pros and cons, we strive to propose a broad multitask transformer network (BMT-Net) to address these problems. BMT-Net takes advantage of both feature-based and fine-tuning methods. It was designed to explore the high-level information of robust and contextual representation. Primarily, our proposed structure can make the learned representations universal across tasks via multitask transformers. In addition, BMT-Net can roundly learn the robust contextual representation utilized by the broad learning system due to its powerful capacity to search for suitable features in deep and broad ways. The experiments were conducted on two popular datasets of binary Stanford Sentiment Treebank (SST-2) and SemEval Sentiment Analysis in Twitter (Twitter). Compared with other state-of-the-art methods, the improved representation with both deep and broad ways is shown to achieve a better F1 -score of 0.778 in Twitter and accuracy of 94.0% in the SST-2 dataset, respectively. These experimental results demonstrate the abilities of recognition in sentiment analysis and highlight the significance of previously overlooked design decisions about searching contextual features in deep and broad spaces. Tong Zhang 0015, Xin-Rong Gong, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2022 | A Structure Constraint Matrix Factorization Framework for Human Behavior SegmentationabstractThis article presents a structure constraint matrix factorization framework for different behavior segmentation of the human behavior sequential data. This framework is based on the structural information of the behavior continuity and the high similarity between neighboring frames. Due to the high similarity and high dimensionality of human behavior data, the high-precision segmentation of human behavior is hard to achieve from the perspective of application and academia. By making the behavior continuity hypothesis, first, the effective constraint regular terms are constructed. Subsequently, the clustering framework based on constrained non-negative matrix factorization is established. Finally, the segmentation result can be obtained by using the spectral clustering and graph segmentation algorithm. For illustration, the proposed framework is applied to the Weiz dataset, Keck dataset, mo_86 dataset, and mo_86_9 dataset. Empirical experiments on several public human behavior datasets demonstrate that the structure constraint matrix factorization framework can automatically segment human behavior sequences. Compared to the classical algorithm, the proposed framework can ensure consistent segmentation of sequential points within behavior actions and provide better performance in accuracy. Hongbo Gao 0001, Chen Lv 0001, Tong Zhang 0015, Hongfei Zhao, Yi Huang 0038 |
IEEE Trans. Cybern. | 3 |
| 2022 | Research Review for Broad Learning System: Algorithms, Theory, and ApplicationsabstractIn recent years, the appearance of the broad learning system (BLS) is poised to revolutionize conventional artificial intelligence methods. It represents a step toward building more efficient and effective machine-learning methods that can be extended to a broader range of necessary research fields. In this survey, we provide a comprehensive overview of the BLS in data mining and neural networks for the first time, focusing on summarizing various BLS methods from the aspects of its algorithms, theories, applications, and future open research questions. First, we introduce the basic pattern of BLS manifestation, the universal approximation capability, and essence from the theoretical perspective. Furthermore, we focus on BLS's various improvements based on the current state of the theoretical research, which further improves its flexibility, stability, and accuracy under general or specific conditions, including classification, regression, semisupervised, and unsupervised tasks. Due to its remarkable efficiency, impressive generalization performance, and easy extendibility, BLS has been applied in different domains. Next, we illustrate BLS's practical advances, such as computer vision, biomedical engineering, control, and natural language processing. Finally, the future open research problems and promising directions for BLSs are pointed out. Xin-Rong Gong, Tong Zhang 0015, C. L. Philip Chen, Zhulin Liu |
IEEE Trans. Cybern. | 2 |
| 2022 | Transfer Collaborative Fuzzy Clustering in Distributed Peer-to-Peer NetworksabstractThe traditional collaborative fuzzy clustering can effectively perform data clustering in distributed peer-to-peer networks, which is an impossible task to complete for the centralized clustering methods due to privacy and security requirements or network transmission technology constraints. But it will increase the number of clustering iterations and lead to lower efficiency of the clustering. Moreover, the collaborative mechanism hidden in the iterative process of clustering cannot be well revealed and explained. In this article, a novel series of transfer collaborative fuzzy clustering algorithms are proposed to solve these issues. In the first basic algorithm, the transfer learning among neighbor nodes vividly expresses the collaborative mechanism and enhances the information collaboration to accelerate the convergence of fuzzy clustering. Meanwhile, neighbor nodes can learn the knowledge from each other to further promote their respective clustering performance. Then, an improved version, with the learning-rate-adjustable strategy instead of fixed values, is designed to highlight the different influence between neighbor nodes, and the appropriate learning rates between neighbor nodes are achieved to ensure the stable clustering accuracy. Finally, two extended versions with the attribute-weight-entropy regularization technique are presented for the clustering of high dimensional sparse data and the extraction of important subspace features. Experiments show the efficiency of the proposed algorithms compared with the related prototype-based clustering methods. Bozhan Dang, Yingxu Wang 0002, Jin Zhou 0003, Long Chen 0001, C. L. Philip Chen, Tong Zhang 0015, Shi-Yuan Han, Lin Wang 0004, Yuehui Chen |
IEEE Trans. Fuzzy Syst. | 7 |
| 2022 | An Efficient Inspection System Based on Broad Learning: Nondestructively Estimating Cement Compressive Strength With Internal FactorsabstractCement has been widely used in civil engineering, whose quality directly affects the safety of buildings. Cement compressive strength, as an important quality indicator, its accurate estimation is of great significance in quality inspections and the design of high-performance products. However, existing measurement technology remains traditional and destructive. Except for high time-consuming and the waste of various resources, it requires significant improvement since the unprofessional operations will give rise to large errors. In this article, an efficient system is proposed to estimate the cement compressive strength based on the broad learning and internal factors, in which the index system describes the internal factors affecting the compressive strength, and the broad learning system distills the potential correlation between the compressive strength and those factors. It can nondestructively estimate the strength directly with the internal factors, e.g., clinker composition and physical properties. In addition, to verify its practicability and to assist the formula optimization in the application, the robustness test and factorial analysis are designed. The experimental results prove that this model can accurately estimate the strength with excellent generalization ability, which saves labor power and material, avoids large errors caused by unprofessional operations, and aids high-performance cement production. Especially, its ability to rapidly build an accurate estimation model is beneficial for the production of various cement in industry. Jifeng Guo 0002, Zhulin Liu, C. L. Philip Chen, Tong Zhang 0015, Lin Wang 0004, Kaipeng Fan |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Compensating for Local Ambiguity With Encoder-Decoder in Urban Scene SegmentationabstractSemantic segmentation plays a critical role in scene understanding for self-driving vehicles. A line of efforts has proven that global context matters in urban scene segmentation due to massive scale changes. However, we find that existing methods suffer from local ambiguities when dissipating continuous local context, i.e. scrambling to a huge receptive field of global cues by coarse pooling. To this end, this paper proposes a new Context Aggregation Module (CAM) that consists of two primary components: context encoding using no coarse pooling but encoder-decoders with appropriate sampling scales and gated fusion that extends gate attention mechanism to balance different-scale context during feature fusion. Weeding out coarse pooling and applying the encoder-decoder inherits the merits of exploring global context while avoiding the drawback of losing local contextual continuity. We then construct a Context Aggregation Network (CANet) and conduct extensive evaluations on challenging autonomous driving benchmarks of Cityscapes, CamVid and BDD100K. Consistently improved results evidence the effectiveness. Notably, we attain competitive mIoU 82.7% on Cityscapes and optimal mIoU 80.5% on CamVid. Quan Tang 0001, Fagui Liu, Tong Zhang 0015, Jun Jiang 0003, Yu Zhang 0144, Boyuan Zhu, Xuhao Tang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problem in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from three aspects. First, we establish a CDMER experimental evaluation protocol aiming to allow the researchers to conveniently work on this topic and evaluate their proposed methods under the same standard. Second, we conduct benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating CDMER problem from two different perspectives. Third, we propose a novel DA method called region selective transfer regression (RSTR) to deal with the CDMER task. The overall superior performance of RSTR over the state-of-the-art DA methods demonstrates that taking into consideration the facial local region information used in RSTR contributes to developing effective DA methods for dealing with CDMER problem. Tong Zhang 0015, Yuan Zong, Wenming Zheng, C. L. Philip Chen, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Situational Assessment for Intelligent Vehicles Based on Stochastic Model and Gaussian Distributions in Typical Traffic ScenariosabstractIn intelligent driving, situational assessment (SA) is an important technology, which helps to improve the cognitive ability of intelligent vehicles in the environment. Uncertainty analysis is very significant in situation assessment. This article proposes an SA method based on uncertainty risk analysis. Under uncertain conditions, according to the random environment model and Gaussian distribution model, the collision probability between multiple vehicles is estimated by comprehensive trajectory prediction. The proposed method considers collision probabilities of different prediction points within and outside the prediction range and obtains long-term accurate prediction results. The method is suitable for the situation risk assessment of sensor systems in the presence of unexpected dynamic obstacles, sensor failures or communication losses in traffic, and different environmental sensing accuracy. The experimental results show that in the dynamic traffic environment, the proposed scenario assessment method can not only accurately predict and assess the situation risks within the prediction range, but also provide accurate scenario risk assessment outside the prediction range. Hongbo Gao 0001, Juping Zhu, Tong Zhang 0015, Guotao Xie, Zhen Kan, Zhengyuan Hao, Kang Liu 0023 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Cross-Subject EEG Emotion Recognition Using Domain Adaptive Few-Shot Learning NetworksabstractDue to the individual differences and nonstationary of EEG signals, it is difficult to classify EEG emotions with traditional machine methods, which assume that the training and testing set come from the same data distribution, but this assumption is usually not true in the EEG field, therefore the accuracy of emotion recognition is very poor. In this paper, a Single-Source Domain Adaptive Few-Shot Learning Networks (SDA-FSL) was proposed for cross-subject EEG emotion recognition. This is the first time that domain adaptation method with few-shot learning has been used in the field of EEG emotion recognition. A CBAM-based feature mapping module was designed to extract the common features of the two domains, and the domain adaptation module was used to align the data distribution of two domains. In addition, Prototypical Networks with instance-attention mechanism is introduced to preserve domain-specific information. The proposed method was evaluated on DEAP and SEED datasets in within-dataset and cross-dataset experiments under various N-way k-shot settings. Experimental results show that the performance of SDA-FSL outperforms other comparison methods and has superior generalization performance on cross-dataset experiments. Run Ning, C. L. Philip Chen, Tong Zhang 0015 |
BIBM | 3 |
| 2021 | Edge computing and its role in Industrial Internet: Methodologies, applications, and future directions
Tong Zhang 0015, Yikai Li 0001, C. L. Philip Chen |
Inf. Sci. | 1 |
| 2021 | Attention-guided chained context aggregation for semantic segmentation
Quan Tang 0001, Fagui Liu, Tong Zhang 0015, Jun Jiang 0003, Yu Zhang 0144 |
Image Vis. Comput. | 3 |
| 2021 | Emotion Recognition From Multimodal Physiological Signals Using a Regularized Deep Fusion of Kernel MachineabstractThese days, physiological signals have been studied more broadly for emotion recognition to realize emotional intelligence in human-computer interaction. However, due to the complexity of emotions and individual differences in physiological responses, how to design reliable and effective models has become an important issue. In this article, we propose a regularized deep fusion framework for emotion recognition based on multimodal physiological signals. After extracting the effective features from different types of physiological signals, we construct ensemble dense embeddings of multimodal features using kernel matrices, and then utilize a deep network architecture to learn task-specific representations for each kind of physiological signal from these ensemble dense embeddings. Finally, a global fusion layer with a regularization term, which can efficiently explore the correlation and diversity among all of the representations in a synchronous optimization process, is designed to fuse generated representations. Experiments on two benchmark datasets show that this framework can improve the performance of subject-independent emotion recognition compared to single-modal classifiers or other fusion methods. Data visualization also demonstrates that the final fusion representation exhibits higher class-separability power for emotion recognition. Xiaowei Zhang 0001, Jinyong Liu, Jian Shen 0004, Kechen Hou, Tong Zhang 0015, Bin Hu 0001 |
IEEE Trans. Cybern. | 8 |
| 2021 | AS-NAS: Adaptive Scalable Neural Architecture Search With Reinforced Evolutionary Algorithm for Deep LearningabstractNeural architecture search (NAS) is a challenging problem in the design of deep learning due to its nonconvexity. To address this problem, an adaptive scalable NAS method (AS-NAS) is proposed based on the reinforced I-Ching divination evolutionary algorithm (IDEA) and variable-architecture encoding strategy. First, unlike the typical reinforcement learning (RL)-based and evolutionary algorithm (EA)-based NAS methods, a simplified RL algorithm is developed and used as the reinforced operator controller to adaptively select the efficient operators of IDEA. Without the complex actor–critic parts, the reinforced IDEA based on simplified RL can enhance the search efficiency of the original EA with lower computational cost. Second, a variable-architecture encoding strategy is proposed to encode neural architecture as a fixed-length binary string. By simultaneously considering variable layers, channels, and connections between different convolution layers, the deep neural architecture can be scalable. Through the integration with the reinforced IDEA and variable-architecture encoding strategy, the design of the deep neural architecture can be adaptively scalable. Finally, the proposed AS-NAS are integrated with the${L}_{1/2}$regularization to increase the sparsity of the optimized neural architecture. Experiments and comparisons demonstrate the effectiveness and superiority of the proposed method. Tong Zhang 0015, Chunyu Lei, Zongyan Zhang, Xianbing Meng, C. L. Philip Chen |
IEEE Trans. Evol. Comput. | 1 |
| 2021 | Stacked Broad Learning System: From Incremental Flatted Structure to Deep ModelabstractThe broad learning system (BLS) has been proved to be effective and efficient lately. In this article, several deep variants of BLS are reviewed, and a new adaptive incremental structure, Stacked BLS, is proposed. The proposed model is a novel incremental stacking of BLS. This invariant inherits the efficiency and effectiveness of BLS that the structure and weights of lower layers of BLS are fixed when the new blocks are added. The incremental stacking algorithm computes not only the connection weights between the newly stacking blocks but also the connection weights of the enhancement nodes within the BLS block. The Stacked BLS is considered as the increment of “layers” and “neurons” dynamically during the training for multilayer neural networks. The proposed architecture along with the training algorithms that utilizes the residual characteristic is very versatile in comparison with traditional fixed architecture. Finally, experimental results on UCI datasets, MNIST dataset, NORB dataset, CIFAR-10 dataset, SVHN dataset, and CIFAR-100 dataset indicate that the proposed method outperforms the selected state-of-the-art methods on both accuracy and training speed, such as deep residual networks. The results also imply that the proposed structure could highly reduce the number of nodes and the training time of the original BLS in the classification task of some datasets. Zhulin Liu, C. L. Philip Chen, Feng Shuang 0001, Qiying Feng, Tong Zhang 0015 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2021 | Hierarchical Lifelong Learning by Sharing Representations and Integrating HypothesisabstractIn lifelong machine learning (LML) systems, consecutive new tasks from changing circumstances are learned and added to the system. However, sufficiently labeled data are indispensable for extracting intertask relationships before transferring knowledge in classical supervised LML systems. Inadequate labels may deteriorate the performance due to the poor initial approximation. In order to extend the typical LML system, we propose a novel hierarchical lifelong learning algorithm (HLLA) consisting of two following layers: 1) the knowledge layer consisted of shared representations and integrated knowledge basis at the bottom and 2) parameterized hypothesis functions with features at the top. Unlabeled data is leveraged in HLLA for pretraining of the shared representations. We also have considered a selective inherited updating method to deal with intertask distribution shifting. Experiments show that our HLLA method outperforms many other recent LML algorithms, especially when dealing with higher dimensional, lower correlation, and fewer labeled data problems. Tong Zhang 0015, Guoxi Su, Chunmei Qing, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Multi-Channel EEG Based Emotion Recognition Using Temporal Convolutional Network and Broad Learning SystemabstractAutomatic real-time emotion recognition based on multi-channel EEG signals is a significant and challenging task in neurology and psychiatry. In recent years, deep learning has been used in EEG emotion recognition. However, many existing deep learning based methods still require complex pre-processing or additional feature extraction, which make it difficult to achieve real-time emotion recognition. In this paper, an end-to-end model named Temporal Convolutional Broad Learning System (TCBLS) was designed for multi-channel EEG based emotion recognition. The TCBLS takes one-dimensional EEG signals as input, then extracts emotion-related features of EEG automatically. In this model, the Temporal Convolutional Network (TCN) is designed to extract EEG temporal features and deep abstract features simultaneously, then Broad Learning System (BLS) is used to map the features to a more discriminative space and further enhance the features. We evaluated our method on DEAP database, performing 10-fold cross-validation on each subject to obtain the classification accuracy. Experimental results indicate that the performance of TCBLS is better than other comparison methods, and the mean accuracy of TCBLS is 99.5755% and 99.5781% on valence and arousal classification task respectively. The results demonstrate the effectiveness and robustness of TCBLS in EEG emotion recognition. Tong Zhang 0015, C. L. Philip Chen, Zhulin Liu, Long Chen 0001, Guihua Wen, Bin Hu 0001 |
SMC | 2 |
| 2020 | Multiple Spatial Information Weighted Fuzzy Clustering for Image SegmentationabstractFor image segmentation, fuzzy clustering methods with single spatial information cannot ensure robustness to the image corrupted by different noises. In this paper, to figure out this problem, we propose a multiple spatial information weighted fuzzy clustering method, in which the original pixel intensity and its two spatial information, the mean and median of neighbors within a local window, are combined with different weights to obtain precise segmentation results of noise images. And the entropy-regularized method is employed to optimize the weight of each term to handle the images with different noise. What's more, the kernelization of the proposed method is presented to relief the impact of outliers. It is worth noting that our methods can be further extended by combining with other spatial information. Experiments on synthetic images and natural images show the superiority and efficiency of the proposed methods. Xiangdao Liu, Jin Zhou 0003, C. L. Philip Chen, Tong Zhang 0015, Lin Wang 0004, Shi-Yuan Han, Yuehui Chen |
SMC | 5 |
| 2020 | Exploring privileged information from simple actions for complex action recognition
Fang Liu 0030, Xiangmin Xu 0001, Tong Zhang 0015, Kailing Guo, Lin Wang 0004 |
Neurocomputing | 3 |
| 2020 | Underwater Internet of Things in Smart Ocean: System Architecture and Open IssuesabstractThe development of the smart ocean requires that various features of the ocean be explored and understood. The Underwater Internet of Things (UIoT), an extension of the Internet of Things (IoT) to the underwater environment, constitutes powerful technology for achieving the smart ocean. This article provides an overview of the UIoT with emphasis on current advances, future system architecture, applications, challenges, and open issues. The UIoT is enabled by the most recent developments in autonomous underwater vehicles, smart sensors, underwater communication technologies, and underwater routing protocols. In the coming years, the UIoT is expected to bridge diverse technologies for sensing the ocean, allowing it to become a smart network of interconnected underwater objects that has self-learning and intelligent computing capabilities. This article first provides a horizontal overview of the UIoT. Then, we present a five-layer system architecture for the future UIoT, which consists of a sensing, communication, networking, fusion, and application layer. Finally, we suggest the current challenges and the future UIoT research trends, in which cloud computing, fog computing, and artificial intelligence are combined. Tie Qiu 0001, Zhao Zhao 0002, Tong Zhang 0015, Chen Chen 0006, C. L. Philip Chen |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Brain-Robot Interface-Based Navigation Control of a Mobile Robot in Corridor EnvironmentsabstractThis paper proposes a brain-robot interface (BRI)-based control strategy in combination with the simultaneous localization and mapping (SLAM) to achieve the navigation and control of a mobile robot in uncertain environments. The BRI is based on steady state visually evoked potentials, utilizing the multivariate synchronization index classification algorithm to analyze the human electroencephalograph (EEG) signals in such a manner that human intentions can be recognized and motion commands can be produced for the brain controlled robot. The entire system is semi-autonomous since the navigation of mobile robot is commanded by the BRI, and the low-level motion of the mobile robot is autonomous with a designed kinematic controller. By utilizing vanishing points and door plates as the environmental features, a global metric map of the environment has been built by a sequential SLAM algorithm. The main contribution of this paper is the combination of an artificial potential field (APF) and the brain signals, which builds up the relationship between the strength of EEG signals and the intensity of the potential field. Through the proposed EEG-APF method, motion commands that would plan an obstacle-free trajectory in un-structured environments, can be obtained. The entire system has been tested with eight volunteer subjects, and all subjects are able to successfully fulfill manipulating mobile robot in the experiments. Yiliang Liu, Zhijun Li 0001, Tong Zhang 0015, Suna Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | Multi-Kernel Broad Learning systems Based on Random Features: A Novel Expansion for Nonlinear Feature NodesabstractThe Broad Learning System has been proved to be effective and efficient. However, the associated feature nodes in the system are mainly based on linear mappings. Although such kind of features has been successful in various datasets and applications, more general features (especially for the nonlinear features) are necessary for specific applications. Motivated by the powerful capability of the kernel methods, a novel expansion of broad learning system based on multiple kernels is proposed in this paper. Firstly, the nonlinear feature mappings in the form of multiple kernels are merged into the feature nodes of broad learning system. After that, the resulted features are further enhanced through nonlinear activation functions. The experimental results on UCI datasets indicate that the proposed method outperforms the other methods. Zhulin Liu, C. L. Philip Chen, Tong Zhang 0015, Jin Zhou 0003 |
SMC | 3 |
| 2018 | EEG Emotion Recognition Using Dynamical Graph Convolutional Neural Networks and Broad Learning System
Xuehan Wang, Tong Zhang 0015, Xiangmin Xu 0001, Long Chen 0001, Xiao-Fen Xing, C. L. Philip Chen |
BIBM | 2 |
| 2018 | Improved Quantification of 18O Labeled LC-MS Based on I-Ching Divination Evolutionary AlgorithmabstractAn innovative quantification method for 18O labeled LC-MS data is proposed based on I-Ching divination evolutionary algorithm(IDEA). Considering label efficiency for calculating the least squares regression function, traditional methods based on genetic algorithm(GA) or other optimized algorithms will bring high level of computation complexity. The proposed method applies very flexible I-Ching operators(ICOs)— intrication operator, turnover operator, and mutual operator. The objective is the function of determining coefficients, which include the 18O/16O ratio r, the label efficiency f, and the abundance a of 16O. Comparing with GA, the proposed algorithm can significantly improve the accuracy and precision of peptide ratio measurements and better performs in the evolution procedure over mathematically calculating the function. Simultaneously we run the experiment with mix peptide raw data of predefined ratio. The result shows that our proposal algorithm is superior to the conventional GA in exploring optimum solution for better quantification accuracy. Tianjun Li, C. L. Philip Chen, Long Chen 0001, Tong Zhang 0015, Bianna Chen, Xiangmin Xu 0001 |
SMC | 4 |
| 2018 | Facial Expression Recognition via Broad Learning SystemabstractIn recent years, research on facial expression recognition (FER) has become an increasingly active research topic. Deep learning is a new area, which gives a new way to classify images of human faces into emotion categories. However, it faces many difficulties caused by poor robustness and real-time performance. This paper designs a new architecture network based on Broad Learning System (BLS) for facial expressions recognition. It is established as a flat network. The original inputs are transferred and placed as mapped features in feature nodes, while the structure is expanded in wide sense in the enhancement nodes. To evaluate our architecture we tested the proposed method with the Extended Cohn-Kanade Dataset (CK+). The experimental results show that the BLS approach is very effective in facial expression recognition to compare with convolutional neural networks. Tong Zhang 0015, Zhulin Liu, Xuehan Wang, Xiao-Fen Xing, C. L. Philip Chen, Enhong Chen |
SMC | 1 |
| 2018 | Design of Highly Nonlinear Substitution Boxes Based on I-Ching OperatorsabstractThis paper is to design substitution boxes (S-Boxes) using innovative I-Ching operators (ICOs) that have evolved from ancient Chinese I-Ching philosophy. These three operators-intrication, turnover, and mutual- inherited from I-Ching are specifically designed to generate S-Boxes in cryptography. In order to analyze these three operators, identity, compositionality, and periodicity measures are developed. All three operators are only applied to change the output positions of Boolean functions. Therefore, the bijection property of S-Box is satisfied automatically. It means that our approach can avoid singular values, which is very important to generate S-Boxes. Based on the periodicity property of the ICOs, a new network is constructed, thus to be applied in the algorithm for designing S-Boxes. To examine the efficiency of our proposed approach, some commonly used criteria are adopted, such as nonlinearity, strict avalanche criterion, differential approximation probability, and linear approximation probability. The comparison results show that S-Boxes designed by applying ICOs have a higher security and better performance compared with other schemes. Furthermore, the proposed approach can also be used to other practice problems in a similar way. Tong Zhang 0015, C. L. Philip Chen, Long Chen 0001, Xiangmin Xu 0001, Bin Hu 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Emotion classification using deep neural networks and emotional patchesabstractEmotion is closely related to healthy and abnormal mood is the alarm of our body. This paper is concentrated on the objective and accurate emotion classification using EEG signal. We propose emotional patches and combine it with the deep belief network(DBN) to achieve high-precision emotion classification. DBN is able to fit the distribution of the EEG signal and mapping the extracted feature to the higher-level characteristics space where we can easily perform high-precision classification. Compared with the other method, our method uses the emotional patches which have considered the temporal information of emotion and reduce the influence of noise. In addition, our model doesn't need to be trained twice to complete higher classification accuracy. We divide the EEG signal and choose the vital β frequency band where we perform feature extraction. Based on the SJTU Emotion EEG Dataset(SEED), we perform the emotion classification experiment and compare our method with the commonly used classifiers such as SVM, LR and CCA ect. The experimental result demonstrates that our method achieves the highest classification accuracy and outperform the state-of-theart emotion classification approaches based on EEG. Jungming Huang, Xiangmin Xu 0001, Tong Zhang 0015 |
BIBM | 3 |
| 2017 | Multi-scale convolutional neural networks for crowd countingabstractCrowd counting on static images is a challenging problem due to scale variations. Recently deep neural networks have been shown to be effective in this task. However, existing neural-networks-based methods often use the multi-column or multi-network model to extract the scale-relevant features, which is more complicated for optimization and computation wasting. To this end, we propose a novel multi-scale convolutional neural network (MSCNN) for single image crowd counting. Based on the multi-scale blobs, the network is able to generate scale-relevant features for higher crowd counting performances in a single-column architecture, which is both accuracy and cost effective for practical applications. Complemental results show that our method outperforms the state-of-the-art methods on both accuracy and robustness with far less number of parameters. Lingke Zeng, Xiangmin Xu 0001, Bolun Cai, Suo Qiu, Tong Zhang 0015 |
ICIP | 5 |
| 2017 | Spectral clustering based on JS-divergence for uncertain dataabstractSpectral clustering is one of the most effective methods of data mining, in which the adjacency matrix is constructed by using the similarity matrix. In this paper, to extend spectral clustering method for uncertain data clustering, we propose a new spectral clustering method based on JS-divergence. In the proposed method, the JS-divergence is used to construct the adjacency matrix in the spectral clustering, which is more suitable to calculate the similarity between uncertain data objects as a symmetrical measurement compared to the KL-divergence. Yingxu Wang 0002, Jiwen Dong, Jin Zhou 0003, Lin Wang 0004, Shi-Yuan Han, Tong Zhang 0015, C. L. Philip Chen |
SMC | 6 |
| 2017 | I-Ching Divination Evolutionary Algorithm and its Convergence AnalysisabstractAn innovative simulated evolutionary algorithm (EA), called I-Ching divination EA (IDEA), and its convergence analysis are proposed and investigated in this paper. Inherited from ancient Chinese culture, I-Ching divination has always been used as a divination system in traditional and modern China. There are three operators evolved from I-Ching transformations in this new optimization algorithm, intrication operator, turnover operator, and mutual operator. These new operators are very flexible in the evolution procedure. Additionally, two new spaces are defined in this paper, which are denoted as hexagram space and state space. In order to analyze the convergence property of I-Ching divination algorithm, Markov model was adopted to analyze the characters of the operators. Meanwhile, the proposed algorithm is proved to be a homogeneous Markov chain with the positive transition matrix. After giving some basic concepts of necessary theorems, definition of admissible functions and I-Ching map, a precise proof of the states converge to the global optimum is presented. Compared with the genetic algorithm, particle swarm optimization, and differential evolution algorithm, our proposed IDEA is much faster in reaching the global optimum. C. L. Philip Chen, Tong Zhang 0015, Long Chen 0001, Sik Chung Tam |
IEEE Trans. Cybern. | 2 |
| 2015 | Matrix Factorization with Scale-Invariant Parameters
Guangxiang Zeng, Hengshu Zhu, Qi Liu 0003, Ping Luo 0001, Enhong Chen, Tong Zhang 0015 |
IJCAI | 6 |
| 2014 | Impact of ratio k on two-layer neural networks with dynamic optimal learning rateabstractLearning process is an important part in two-layer networks. It is imperative to search for an optimal learning rate to get a maximum error reduction in each learning step. Related literature has proposed various kinds of methods to find such an optimal learning rate in the past decades. In this paper, we proposed an improved dynamic optimal learning rate by adding an optimal ratio k. It is found that our improved dynamic optimal learning rate can generate a better result in learning processes. Meanwhile, we have proved the existence of the ratio kby giving it a proper math expression. Furthermore, we also applied the improved learning rate to solve inverse problem and compared the difference of the improved learning rate with the previous approach. It is observed that our proposed method performs better. Therefore, it can be concluded that our new method to search for dynamic optimal learning rate is valuable in the intelligence learning applications of neural networks, or it is effective in the aspect of tested problem at least. Tong Zhang 0015, C. L. Philip Chen, Jin Zhou 0003 |
IJCNN | 1 |
| 2014 | A novel evolutionary algorithm solving optimization problemsabstractThis paper develops an novel evolutionary algorithm, I Ching algorithm (ICA) for solving optimization problems. The new algorithm employs an novel method by implying new operators from I Ching, which comes from ancient Chinese culture. There are some transformation methods such as a penalty method and a multiplier method. The penalty method is often used to solve optimization problems, because the solutions are often near the boundary of the feasible set and the method is used easily for its simplicity. In design the ICA, three operators - mutation operator, turnover operator, and mutual operator were developed by the authors based on the concept of I Ching transformations. These new operators are very flexible and search on the designed I Ching network in the evolution procedure. The proposed algorithm was applied to solving two optimization benchmark functions, Booth function and Hump function. Then, we compare the performance of ICA with genetic algorithm. The experimental results show that our proposed I Ching algorithm performs better than genetic algorithm in reaching the global optimum. It is much faster than those of genetic algorithms. Additionally, the ICA is also a universal method, which is suitable to different optimization problems. C. L. Philip Chen, Tong Zhang 0015, Sik Chung Tam |
SMC | 2 |
| 2012 | Image encryption algorithm based on a new combined chaotic systemabstractChaotic theory has been applied to image encryption as an effective and robust technique due to its unique properties. In this paper, we introduce a new combined chaotic system, which shows better chaotic behaviors than the traditional ones. Applying this chaotic system to image processing, a new image encryption algorithm is introduced based on the confusion and diffusion in encryption procedure. Experimental results show that the proposed algorithm has a higher security level and excellent performance in image encryption. C. L. Philip Chen, Tong Zhang 0015, Yicong Zhou |
SMC | 2 |
| 2011 | On the classification of cancer cell gene via Expressive Value Distance (EVD) algorithm and its comparison to the optimally trained ANN methodabstractIn recent years, cancer can be detected and recognized by analyzing the sample's expression profile. The cancer gene expression data are high dimensional, high variable dependent, and very noisy. The dimension reduction method is often used for processing the high dimensional data. In this study, a new statistical dimension reduction method called Expressive Value Distance (EVD) is developed and proposed for the practical high-dimensional gene expression cancer data. The feature genes data extracted by EVD are arranged for training the optimally trained Artificial Neural Network (ANN). The trained ANN is then used to classify whether the unseen gene data is cancer or not. In comparison of ANN classification with and without EVD, it is found that both of the ANN can classify the cancer data in good accuracy. With the EVD method, the great amount of data (2000 genes) can be effectively reduced to 16 genes. Therefore, EVD is an effective dimension reduction method. Even the EVD method is not used, the optimally trained ANN is also an advanced method for classifying the high dimensional and complicated cancer data. Briefly, it proves that optimally trained ANN is a very robust classification technique. Tong Zhang 0015, Chi-Hsu Wang, Sik Chung Tam, C. L. Philip Chen |
FUZZ-IEEE | 1 |