Chong Zhang 0003

dblp:74/3128-3 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-2162-4344ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 19 since 2021Artificial intelligence and machine learning · 15 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 MDSF-YOLO: Advancing Object Detection With a Multiscale Dilated Sequence Fusion Network
abstract
Accurate and fast detection of traffic signs is critical for autonomous driving, particularly in complex environments with diverse sign scales and varying detection distances. Existing approaches, incorporating attention modules or modifying detection heads, frequently encounter high rates of false positives and omissions due to the increased sampling depth. To address these limitations, we propose MDSF-you only look once (YOLO), a novel detection framework that integrates multiscale sequence fusion (MSF) for synergistic feature integration across granularities, enhancing the precision of both localization and semantic information fusion. Additionally, our dilated-wise residual (DWR) module leverages dilated convolutions and channel-wise reparameterization to improve fine-grained feature extraction. The architecture further introduces a $P_{2}$ detection head for shallow features and fully decouples all detection heads, optimizing target localization and category identification. Extensive experiments on the TT100K and CCTSDB2021 datasets demonstrate the superiority of MDSF-YOLO over benchmark models, including YOLOv11s, with significant improvements in mAP by 8.8% and 2.4% on respective datasets while substantially reducing false positives and leakage rate. Besides, the marked improvement of MDSF-YOLO on the VisDrone2019 dataset verifies its enhanced capability to address drone-based object detection. These advances underscore the efficiency and robustness of the proposed model, providing a promising solution for autonomous driving and similar object detection scenarios.
Chong Zhang 0003, Xuyang Jing, Qing-Guo Wang
IEEE Trans. Neural Networks Learn. Syst.2
2025 UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
abstract
Yidi Jiang, Qian Chen, Shengpeng Ji, Yu Xi, Wen Wang, Chong Zhang, Xianghu Yue, ShiLiang Zhang, Haizhou Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yidi Jiang, Qian Chen 0003, Shengpeng Ji, Yu Xi, Wen Wang 0001, Chong Zhang 0003, Xianghu Yue, Shiliang Zhang, Haizhou Li 0001
ACL (1)6
2025 HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
abstract
The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently trained and concatenated networks may lead to inconsistent representations and poor speech quality, especially in out-of-domain scenarios. In this work, we propose HiFi-SR, a unified network that leverages end-to-end adversarial training to achieve high-fidelity speech super-resolution. Our model features a unified transformer-convolutional generator designed to seamlessly handle both the prediction of latent representations and their conversion into time-domain waveforms. The transformer network serves as a powerful encoder, converting low-resolution mel-spectrograms into latent space representations, while the convolutional network upscales these representations into high-resolution waveforms. To enhance high-frequency fidelity, we incorporate a multi-band, multi-scale time-frequency discriminator, along with a multi-scale mel-reconstruction loss in the adversarial training process. HiFi-SR is versatile, capable of upscaling any input speech signal between 4 kHz and 32 kHz to a 48 kHz sampling rate. Experimental results demonstrate that HiFi-SR significantly outperforms existing speech SR methods across both objective metrics and ABX preference tests, for both in-domain and out-of-domain scenarios.
Shengkui Zhao, Kun Zhou 0003, Zexu Pan, Chong Zhang 0003, Bin Ma 0001
ICASSP5
2025 Conditional Latent Diffusion-Based Speech Enhancement via Dual Context Learning
abstract
Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or spectral domains, leading to increased generation complexity and slower inference speeds. Additionally, these methods have primarily modelled clean speech distributions, with limited exploration of noise distributions, thereby constraining the discriminative capability of diffusion models for speech enhancement. To address these issues, we propose a novel approach that integrates a conditional latent diffusion model (cLDM) with dual-context learning (DCL). Our method utilizes a variational autoencoder (VAE) to compress mel-spectrograms into a low-dimensional latent space. We then apply cLDM to transform the latent representations of both clean speech and background noise into Gaussian noise by the DCL process, and a parameterized model is trained to reverse this process, conditioned on noisy latent representations and text embeddings. By operating in a lower-dimensional space, the latent representations reduce the complexity of the generation process, while the DCL process enhances the model’s ability to handle diverse and unseen noise environments. Our experiments demonstrate the strong performance of the proposed approach compared to existing diffusion-based methods, even with fewer iterative steps, and highlight the superior generalization capability of our models to out-of-domain noise datasets.
Shengkui Zhao, Zexu Pan, Kun Zhou 0003, Chong Zhang 0003, Bin Ma 0001
ICASSP5
2025 Online Audio-Visual Autoregressive Speaker Extraction
Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang 0003, Kun Zhou 0003, Bin Ma 0001
INTERSPEECH4
2025 Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
Zexu Pan, Shengkui Zhao, Kun Zhou 0003, Chong Zhang 0003, Bin Ma 0001
INTERSPEECH6
2024 Loss Masking Is Not Needed In Decoder-Only Transformer For Discrete-Token-Based ASR
abstract
Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a single decoder-only Transformer on a mixture of speech tasks. However, these models rely on the Loss Masking strategy for the ASR task, which ignores the dependency among speech tokens. In this paper, we propose to model speech tokens in an autoregressive way, similar to text. We find that applying the conventional cross-entropy loss on input speech tokens does not consistently improve the ASR performance over the Loss Masking approach. To address this issue, we propose a novel approach denoted Smoothed Label Distillation (SLD), which applies a KL divergence loss with smoothed labels on speech tokens. Our experiments show that SLD effectively models speech tokens and outperforms Loss Masking for decoder-only Transformers in ASR tasks with different speech discretization methods1.
Qian Chen 0003, Wen Wang 0001, Shiliang Zhang, Chong Deng, Jiaqing Liu, Chong Zhang 0003
ICASSP10
2024 Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?
abstract
Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. However, not many people understand how and why this is so. In this study, we aim to deepen our understanding of this emerging method by investigating the role of soft prompts in automatic speech recognition (ASR). Our findings highlight their role as zero-shot learners in improving ASR performance while also exposing them to the risk of malicious modifications. Soft prompts aid generalization but are not obligatory for inference. We also identify two primary roles of soft prompts: content refinement and noise information enhancement, which enhances robustness against background noise. Additionally, we propose an effective modification on noise prompts to show that they are capable of zero-shot learning on adapting to out-of-distribution noise environments.
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Fabian Ritter Gutierrez, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001
ICASSP2
2024 SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance
abstract
Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, which comprise half a dual-path model’s parameters, contribute minimally to performance. Thus, we propose the Single-Path Global Modulation (SPGM) block to replace inter-blocks. SPGM is named after its structure consisting of a parameter-free global pooling module followed by a modulation module comprising only 2% of the model’s total parameters. The SPGM block allows all transformer layers in the model to be dedicated to local feature modelling, making the overall model single-path. SPGM achieves 22.1 dB SI-SDRi on WSJ0-2Mix and 20.4 dB SI-SDRi on Libri2Mix, exceeding the performance of Sepformer by 0.5 dB and 0.3 dB respectively and matches the performance of recent SOTA models with up to 8 times fewer parameters. Model and weights are available at huggingface.co/yipjiaqi/spgm
Jia Qi Yip, Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Dianwen Ng, Chng Eng Siong, Bin Ma 0001
ICASSP5
2024 MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
abstract
Our previously proposed MossFormer has achieved promising performance in monaural speech separation. However, it predominantly adopts a self-attention-based MossFormer module, which tends to emphasize longer-range, coarser-scale dependencies, with a deficiency in effectively modelling finer-scale recurrent patterns. In this paper, we introduce a novel hybrid model that provides the capabilities to model both long-range, coarse-scale dependencies and fine-scale recurrent patterns by integrating a recurrent module into the MossFormer framework. Instead of applying the recurrent neural networks (RNNs) that use traditional recurrent connections, we present a recurrent module based on a feedforward sequential memory network (FSMN), which is considered "RNN-free" recurrent network due to the ability to capture recurrent patterns without using recurrent connections. Our recurrent module mainly comprises an enhanced dilated FSMN block by using gated convolutional units (GCU) and dense connections. In addition, a bottleneck layer and an output layer are also added for controlling information flow. The recurrent module relies on linear projections and convolutions for seamless, parallel processing of the entire sequence. The integrated MossFormer2 hybrid model demonstrates remarkable enhancements over MossFormer and surpasses other state-of-the-art methods in WSJ0-2/3mix, Libri2Mix, and WHAM!/WHAMR! benchmarks.
Shengkui Zhao, Chongjia Ni, Chong Zhang 0003, Hao Wang 0199, Trung Hieu Nguyen 0001, Kun Zhou 0003, Jia Qi Yip, Dianwen Ng, Bin Ma 0001
ICASSP4
2024 Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
Kun Zhou 0003, Shengkui Zhao, Chong Zhang 0003, Hao Wang 0199, Dianwen Ng, Chongjia Ni, Trung Hieu Nguyen 0001, Jia Qi Yip, Bin Ma 0001
INTERSPEECH4
2024 Fine-Tuning Channel-Pruned Deep Model via Knowledge Distillation
Chong Zhang 0003, Hongzhi Wang 0001, Hongwei Liu 0001
J. Comput. Sci. Technol.1
2024 Tuning Large Language Model for Speech Recognition With Mixed-Scale Re-Tokenization
abstract
Large Language Models (LLMs) have proven successful across a spectrum of speech-related tasks, such as speech recognition, text-to-speech, and spoken language understanding. Recently, the use of discretized speech features has gained attention as an efficient and compatible alternative to continuous features for LLMs. This is mainly due to their reduced storage requirements and better alignment of these features with LLM's input space. However, the typical practice of freezing the speech encoder during training poses challenges in bridging the modality gap between speech and text. To address this, we propose to use a mixed-scale re-tokenization layer, integrating multiple granularities in discretized speech features directly within the LLM's input module. Our experimental results demonstrated that the proposed method can effectively enhance the performance of ASR in the setting of continuous learning of an LLM, highlighting the importance of a meticulously designed input module for the integration of discretized speech features with an LLM.
Chong Zhang 0003, Qian Chen 0003, Wen Wang 0001, Bin Ma 0001
IEEE Signal Process. Lett.2
2023 Ditto: A Simple and Efficient Approach to Improve Sentence Embeddings
abstract
Prior studies diagnose the anisotropy problem in sentence representations from pre-trained language models, e.g., BERT, without finetuning.Our analysis reveals that the sentence embeddings from BERT suffer from a bias towards uninformative words, limiting the performance in semantic textual similarity (STS) tasks.To address this bias, we propose a simple and efficient unsupervised approach, Diagonal Attention Pooling (Ditto), which weights words with model-based importance estimations and computes the weighted average of word representations from pre-trained models as sentence embeddings.Ditto can be easily applied to any pre-trained language model as a postprocessing operation.Compared to prior sentence embedding approaches, Ditto does not add parameters nor requires any learning.Empirical evaluations demonstrate that our proposed Ditto can alleviate the anisotropy problem and improve various pre-trained models on the STS benchmarks. 1
Qian Chen 0003, Wen Wang 0001, Chong Deng, Jiaqing Liu, Chong Zhang 0003
EMNLP9
2023 Auxiliary Pooling Layer For Spoken Language Understanding
abstract
End-to-end spoken language understanding requires speech data annotated with semantic information and may suffer from the shortage of annotated data. Recent progresses leverage unlabelled speech data to pre-train a speech encoder. However, it remains a challenge for the pre-trained speech encoder to encode semantic information. Existing works explore transferring knowledge from a pre-trained text model with different alignment losses at a fixed granularity. In this paper, we address the variable granularity in transferring knowledge from texts to speech representation via APLY, an auxiliary pooling layer, that fuses the global information with the adaptively encoded local context. We demonstrate the effectiveness of APLY on three benchmarks of spoken language understanding.
Trung Hieu Nguyen 0001, Jinjie Ni, Wen Wang 0001, Qian Chen 0003, Chong Zhang 0003, Bin Ma 0001
ICASSP6
2023 De'hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech Recognition
abstract
Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a single source, which tends to do poorly when noise is present during testing. Nonetheless, it is crucial to overcome the adverse influence of noise for real-world applications. In this work, we propose a novel training framework, called deHuBERT, for noise reduction encoding inspired by H. Barlow’s redundancy-reduction principle. The new framework improves the HuBERT training algorithm by introducing auxiliary losses that drive the self- and cross-correlation matrix between pairwise noise-distorted embeddings towards identity matrix. This encourages the model to produce noise- agnostic speech representations. With this method, we report improved robustness in noisy environments, including unseen noises, without impairing the performance on the clean set.
Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Jinjie Ni, Chong Zhang 0003, Chongjia Ni, Chng Eng Siong, Bin Ma 0001
ICASSP6
2023 Contrastive Speech Mixup for Low-Resource Keyword Spotting
abstract
Most of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the rising demand for smart devices to become more person-alized, KWS models need to adapt quickly to smaller user samples. To tackle this challenge, we propose a contrastive speech mixup (CosMix) learning algorithm for low-resource KWS. CosMix introduces an auxiliary contrastive loss to the existing mixup augmentation technique to maximize the relative similarity between the original pre-mixed samples and the augmented samples. The goal is to inject enhancing constraints to guide the model towards simpler but richer content-based speech representations from two augmented views (i.e. noisy mixed and clean pre-mixed utterances). We conduct our experiments on the Google Speech Command dataset, where we trim the size of the training set to as small as 2.5 mins per keyword to simulate a low-resource condition. Our experimental results show a consistent improvement in the performance of multiple models, which exhibits the effectiveness of our method.
Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Chng Eng Siong, Bin Ma 0001
ICASSP4
2023 Adaptive Knowledge Distillation Between Text and Speech Pre-Trained Models
abstract
Learning on a massive amount of speech corpus leads to the recent success of many self-supervised speech models. With knowledge distillation, these models may also benefit from the knowledge encoded by language models that are pre-trained on rich sources of texts. The distillation process, however, is challenging due to the modal disparity between textual and speech embedding spaces. This paper studies metric-based distillation to align the embedding space of text and speech with only a small amount of data without modifying the model structure. Since the semantic and granularity gap between text and speech has been omitted in literature, which impairs the distillation, we propose the Prior-informed Adaptive knowledge Distillation (PAD) that adaptively leverages text/speech units of variable granularity and prior distributions to achieve better global and local alignments between text and speech pre-trained models. We evaluate on three spoken language understanding benchmarks to show that PAD is more effective in transferring linguistic knowledge than other metric-based distillation approaches.
Jinjie Ni, Wen Wang 0001, Qian Chen 0033, Dianwen Ng, Han Lei, Trung Hieu Nguyen 0001, Chong Zhang 0003, Bin Ma 0001, Erik Cambria
ICASSP8
2023 Adapter-tuning with Effective Token-dependent Representation Shift for Automatic Speech Recognition
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Qian Chen 0003, Wen Wang 0001, Chng Eng Siong, Bin Ma 0001
INTERSPEECH2
2023 Dual Acoustic Linguistic Self-supervised Representation Learning for Cross-Domain Speech Recognition
Dianwen Ng, Chong Zhang 0003, Xiao Fu 0001, Wei Xi 0003, Chongjia Ni, Chng Eng Siong, Bin Ma 0001, Jizhong Zhao
INTERSPEECH3
2023 A Unified Recognition and Correction Model under Noisy and Accent Speech Conditions
Dianwen Ng, Chong Zhang 0003, Wei Xi 0003, Chongjia Ni, Jizhong Zhao, Bin Ma 0001, Chng Eng Siong
INTERSPEECH3
2023 Dual-Memory Multi-Modal Learning for Continual Spoken Keyword Spotting with Confidence Selection and Diversity Enhancement
Dianwen Ng, Xizhe Li, Chong Zhang 0003, Wei Xi 0003, Chongjia Ni, Jizhong Zhao, Bin Ma 0001, Chng Eng Siong
INTERSPEECH4
2023 ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention
Jia Qi Yip, Duc-Tuan Truong, Dianwen Ng, Chong Zhang 0003, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Chng Eng Siong, Bin Ma 0001
INTERSPEECH4
2019 A Cost-Sensitive Deep Belief Network for Imbalanced Classification
abstract
Imbalanced data with a skewed class distribution are common in many real-world applications. Deep Belief Network (DBN) is a machine learning technique that is effective in classification tasks. However, conventional DBN does not work well for imbalanced data classification because it assumes equal costs for each class. To deal with this problem, cost-sensitive approaches assign different misclassification costs for different classes without disrupting the true data sample distributions. However, due to lack of prior knowledge, the misclassification costs are usually unknown and hard to choose in practice. Moreover, it has not been well studied as to how cost-sensitive learning could improve DBN performance on imbalanced data problems. This paper proposes an evolutionary cost-sensitive deep belief network (ECS-DBN) for imbalanced classification. ECS-DBN uses adaptive differential evolution to optimize the misclassification costs based on the training data that presents an effective approach to incorporating the evaluation measure (i.e., G-mean) into the objective function. We first optimize the misclassification costs, and then apply them to DBN. Adaptive differential evolution optimization is implemented as the optimization algorithm that automatically updates its corresponding parameters without the need of prior domain knowledge. The experiments have shown that the proposed approach consistently outperforms the state of the art on both benchmark data sets and real-world data set for fault diagnosis in tool condition monitoring.
Chong Zhang 0003, Kay Chen Tan, Haizhou Li 0001, Geok Soon Hong
IEEE Trans. Neural Networks Learn. Syst.1
2018 Gated Recurrent Units Based Neural Network For Tool Condition Monitoring
abstract
Tool condition monitoring (TCM) is a prerequisite to ensure high finishing quality of workpiece in manufacturing automation. One of the most important components in TCM system is tool wear estimation. How to achieve estimation with high accuracy is still an open question. In the past few decades, recurrent neural network (RNN) has shown a great success in learning long-term dependence of the sequential data. However, traditional RNNs (e.g., vanilla RNN, etc.) suffer gradient vanishing or exploding problem as well as long computational training time when the model is trained through back propagation through time (BPTT). To address these issues, we propose a gated recurrent units (GRU) based neural network to estimate the tool wear for tool condition monitoring. The GRU neural network can analyze time-series data on multiple time scales and can avoid gradient vanishing during training. A real-world gun drilling experimental dataset is used as a case study for tool condition monitoring in this paper. The performance of the proposed GRU based TCM approach is compared with other well-known models including support vector regression (SVR) and multi-layer perceptron (MLP). The experimental results show that the proposed GRU based TCM approach outperforms other competing models on this real-world gun drilling dataset.
Chong Zhang 0003, Geok Soon Hong, Junhong Zhou, Jihoon Hong, Keng Soon Woon
IJCNN2
2017 A data-driven prognostics framework for tool remaining useful life estimation in tool condition monitoring
abstract
Tool Condition Monitoring (TCM) is an important topic in manufacturing industry, which improves product quality, production efficiency, reduces costs and downtime. This paper develops a new data-driven framework for estimating tool remaining useful life (RUL) in TCM. The framework includes the following modular components: data preprocessing with a proposed adaptive Baysian change point detection (ABCPD) for automatic data alignment, time window process, feature extraction, feature selection and a multi-layer neural network as the main machine learning algorithm. The proposed framework is evaluated on a real-world gun drilling experimental dataset with multiple sensor measurements (i.e. thrust force, torque, 12 vibration signals). Different model selection, sensor selection, feature selection methods have been investigated in this paper. The simulation performance of the proposed framework is studied with the gun drilling dataset and it has been shown that the proposed framework has good performance.
Chong Zhang 0003, Geok Soon Hong, Kay Chen Tan, Junhong Zhou, Hian-Leng Chan, Haizhou Li 0001
ETFA1
2017 Multiobjective Deep Belief Networks Ensemble for Remaining Useful Life Estimation in Prognostics
abstract
In numerous industrial applications where safety, efficiency, and reliability are among primary concerns, condition-based maintenance (CBM) is often the most effective and reliable maintenance policy. Prognostics, as one of the key enablers of CBM, involves the core task of estimating the remaining useful life (RUL) of the system. Neural networks-based approaches have produced promising results on RUL estimation, although their performances are influenced by handcrafted features and manually specified parameters. In this paper, we propose a multiobjective deep belief networks ensemble (MODBNE) method. MODBNE employs a multiobjective evolutionary algorithm integrated with the traditional DBN training technique to evolve multiple DBNs simultaneously subject to accuracy and diversity as two conflicting objectives. The eventually evolved DBNs are combined to establish an ensemble model used for RUL estimation, where combination weights are optimized via a single-objective differential evolution algorithm using a task-oriented objective function. We evaluate the proposed method on several prognostic benchmarking data sets and also compare it with some existing approaches. Experimental results demonstrate the superiority of our proposed method.
Chong Zhang 0003, Pin Lim, A. K. Qin 0001, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.1
2016 Training cost-sensitive Deep Belief Networks on imbalance data problems
abstract
Many real-world problems are usually unbalanced, where datasets present skewed class distributions, such as failure diagnosis, spam detection, anomaly detection, fraud detection, oil spillage detection and medical diagnosis, etc. Deep Belief Network (DBN) is a competitive machine learning technique with good performance in many applications. However, some machine learning methods are likely to give poor performance with imbalanced data between classes since they assume equal costs for each class intrinsically. To deal with this problem, existing researches only focus on sampling based approaches and lack of studies about cost-sensitive based approaches. This paper proposes cost-sensitive Deep Belief Networks for such imbalanced classification problems. The proposed approach is extended to multi-class scenario. Unequalized misclassification costs between classes have been applied to DBN. Extensive comparison with extreme learning machines is provided as a proof of the ability of the proposed approach to perform competitively on imbalanced datasets. An evolutionary algorithm is also implemented to optimize the misclassification costs for each class in cost matrix.
Chong Zhang 0003, Kay Chen Tan, Ruoxu Ren
IJCNN1
2015 Deep Belief Networks Ensemble with Multi-objective Optimization for Failure Diagnosis
abstract
Early diagnosis that can detect faults from some symptoms accurately is critical, because it provides the potential benefits such as reducing maintenance costs, improving productivity and avoiding serious damages. Degradation pattern classification for early diagnosis has not been explored in many researches yet. This paper will use hybrid ensemble model for degradation pattern classification. Supervised training of deep models (e.g. Many-layered Neural Nets) is difficult for optimization problem with unlabeled datasets or insufficient data sample. Shallow models (SVMs, Neural Networks, etc...) are unlikely candidates for learning high-level abstractions, since they are affected by the curse of dimensionality. Therefore, deep learning network (DBN), an unsupervised learning model, in diagnosis problem has been investigated to do classification. Few researches have been done for exploring the effects of DBN in diagnosis. In this paper, an ensemble of DBNs with MOEA/D has been applied for diagnosis to handle failure degradation with multivariate sensory data. Turbofan engine degradation dataset is employed to demonstrate the efficacy of the proposed model. We believe that deep learning with multi-objective ensemble for degradation pattern classification can shed new light on failure diagnosis, and our work presented the applicability of this method to diagnosis as well as prognostics.
Chong Zhang 0003, Jia Hui Sun, Kay Chen Tan
SMC1