Jia-Ching Wang

dblp:41/2001 · DBLP profile ↗
← Back
98ranked-venue papers
16as first author
30since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 61 · 4 first-author · 24 since 2021Artificial intelligence and machine learning · 21 · 4 first-author · 7 since 2021Systems, architecture and hardware · 10 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 1 since 2021Security and privacy · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Mixture of Ordered Scoring Experts for Cross-prompt Essay Trait Scoring
abstract
Po-Kai Chen, Bo-Wei Tsai, Shao Kuan Wei, Chien-Yao Wang, Jia-Ching Wang, Yi-Ting Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Po-Kai Chen, Bo-Wei Tsai, Shao-Kuan Wei, Chien-Yao Wang, Jia-Ching Wang, Yi-Ting Huang
ACL (1)5
2025 HistoFS: Non-IID Histopathologic Whole Slide Image Classification via Federated Style Transfer with RoI-Preserving
abstract
Federated learning for pathological whole slide image (WSI) classification allows multiple clients to train a global multiple instance learning (MIL) model without sharing their privacy-sensitive WSIs. To accommodate the non-independent and identically distributed (non-i.i.d.) feature shifts, cross-client style transfer has been popularly used but is subject to two fundamental issues: (1) WSI contains multiple morphological structures, each corresponding to a distinct style. (2) Performing style transfer may potentially shift the region of interests (RoIs) in the augmented WSIs. To address these challenges, we propose HistoFS, a federated learning framework for computational pathology on non-i.i.d. feature shifts in WSI classification. Specifically, we introduce pseudo bag styles that capture multiple style variations within a single WSI. In addition, an authenticity module is introduced to ensure that RoIs are preserved, allowing local models to learn WSIs with diverse styles while maintaining essential RoIs. Extensive experiments validate the superiority of HistoFS over state-of-the-art methods on three clinical datasets. Our code is available at https://lalakitchen.github.io/HistoFS/.
Farchan Hakim Raswa, Chun-Shien Lu, Jia-Ching Wang
CVPR3
2025 Impact of Glyph Information on Latent Space Diffusion Models for Accurate Handwritten Text Generation
abstract
The generation of high-quality stylized handwritten text images is a challenging task in computer vision and artificial intelligence. While advanced approaches using Latent Diffusion Models (LDMs) for generating stylized handwritten text have shown effectiveness, they often struggle with maintaining the structural integrity of certain characters, resulting in issues such as missing or extraneous strokes. In this work, we propose GlyphLDM, an innovative model that integrates glyph image information into both the diffusion and denoising processes in the latent space, enhancing the structural accuracy of generated text images. In the early training stages, our method demonstrated a significant improvement in the structural accuracy of the generated text images, with the Average Confidence Score increasing by approximately 40% compared to the baseline method. These experimental results indicate that incorporating glyph image information has promising potential to enhance the structural accuracy and overall quality of generated text images. This approach provides an effective solution for generating more accurate and diverse handwritten text images.
Ying-Li Lin, Chung-I Huang, Chien-Yao Wang, Jia-Ching Wang
ICASSP5
2025 A Key to Effective Multi-task Learning: Separate Query Selection for Task-Synergized Handling and Node Utilization
abstract
In the realm of computer vision, effectively handling multi-tasks simultaneously presents a challenge that necessitates innovative solutions. To better address multiple vision problems, we introduce SeTano, an integrated Graph Neural Network (GNN)-based framework. This framework comprises a Dynamic Edge-Sensing GNN (DES-GNN) backbone, which can dynamically adjust edges to extract more pivotal features, and a downstream design which includes a node reduction and a separate query selection strategy. To validate our approach, we perform multi-task experiments on the ImageNet and MS COCO datasets. The results indicate that the integrated design of SeTano leads to enhanced performance in various vision multi-tasks.
Shan-Ya Yang, Chien-Yao Wang, Jia-Ching Wang, Chun-Yi Lee
ICASSP4
2025 User-Customizable Voice Anonymization Through Personalized Style Transfer
abstract
The growing collection of personal voice data online has heightened the demand for effective privacy protection through speaker de-identification. While existing anonymization methods successfully obscure speaker identity, they fail to simultaneously achieve robust identity protection, natural speech preservation, and flexible user customization. We address this limitation through three key innovations: (1) a dynamic neural style-transfer framework that generates perceptually natural yet anonymized speech via reference-guided interpolation; (2) a privacy-preserving disentanglement technique using a triple-encoder architecture to suppress speaker identity while preserving linguistic content and transferable prosodic feature; and (3) a user-customizable design that supports intentional voice persona modulation, catering to emerging existing methods in anonymization effectiveness while maintaining superior speech quality and adaptability. Experimental results and comparisons demonstrate the effectiveness of our method, effectively bridging the gap between privacy protection and speech utility in real-world applications.
Wenny Ramadha Putri, Chun-Shien Lu, Jia-Ching Wang
IJCB3
2025 Diffusion to Confusion: Naturalistic Adversarial Patch Generation Based on Diffusion Model for Object Detector
abstract
Many physical adversarial patch generation methods are widely proposed to protect personal privacy from malicious monitoring using object detectors. However, they usually fail to generate satisfactory patch images in terms of both stealthiness and attack performance without making huge efforts on careful hyperparameter tuning. To address this issue, we propose a novel naturalistic diffusion model-based (DM) adversarial patch generation method. Through sampling the optimal image from the pretrained DM model upon natural images, it allows us to stably craft high-quality physical adversarial patches without suffering serious mode collapse problems as other deep generative models. Moreover, the generated patches are not only visually pleasing but also robust against the state-of-the-art adversarial patch removal algorithm due to different image statistics from the traditional ones. In addition, to resolve the huge memory requirement of the diffusion model during backpropagation, we also utilize adjoint method for patch generation. With extensive experiments, the results demonstrate the effectiveness of the proposed approach to generate better-quality and more stealthy adversarial patches than other approaches against the adversarial patch removal algorithm while achieving comparable attack performance than other state-of-the-art patch generation methods.
Shuo-Yen Lin, Ernie Chu, Po-Hung Yeh, Jun-Cheng Chen, Jia-Ching Wang
ICIP5
2025 Defense Against Backdoor Attacks on Image Retrieval Models Through Strategic Manipulations
Hung-Lei Lee, Chun-Shien Lu, Jia-Ching Wang
ICISSP (2)3
2025 MLSS: Mandarin English Code-Switching Speech Recognition via Mutual Learning-Based Semi-Supervised Method
abstract
Code-switching is a phenomenon of alternating use of two or more languages within or between utterances in communication that often occurs in multilingual communities. Recently, code-switching natural language processing and automatic speech recognition (ASR) have attracted numerous studies. However, a major obstacle affecting the results of these studies is the lack of transcribed data. In this letter, we propose a novel semi-supervised learning (SSL) approach to deal with this problem, namely Mutual Learning-Based Semi-Supervised Method (MLSS). The MLSS method involves the utilization of two networks for interleaved fine-tuning on a combination of transcribed dataset and pseudo-labeled data generated from another network. This iterative fine-tuning process repeats until all unlabeled data are selected for training or reaches a certain number of iterations. By incorporating mutual learning between the two networks, our approach effectively leverages the knowledge acquired from previous iterations during the training stage and combines the knowledge from both networks during the decoding process, resulting in a more robust and effective approach. To evaluate the effectiveness of our proposed method, we conduct experiments on the SEAME Mandarin-English code-switching corpus. The experimental results clearly illustrate that our approach outperforms other state-of-the-art methods, as evidenced by achieving a Mixed Error Rate (MER) of 15.6% /21.1% on test$_{man}$/test$_{sge}$sets.
Cao Hong Nga, Duc-Quang Vu, Phuong Le Thi, Huong Hoang Luong, Jia-Ching Wang
IEEE Signal Process. Lett.5
2024 Enhancing Breast Cancer Detection: A Novel Training Strategy and Batch Scheduler Method
abstract
Recent advancements in deep learning and computational power have opened new possibilities that were once unattainable. Researchers are now eager to transfer their expertise in deep learning to various domains, including medical diagnostics, autonomous systems, and remote sensing. In the medical field, the application of deep learning promises to reduce the workload of healthcare professionals, streamline screening processes, and improve time efficiency. Breast cancer remains a significant challenge for many women, with mass-type cancers having increased difficulty due to their heterogeneity and anomalies. Addressing this issue requires a more detailed investigation. In this study, we propose an innovative training strategy aimed at improving the precision and F1 score of breast cancer detection models. Furthermore, we introduce a novel method, named the Batch Scheduler, which dynamically adjusts batch sizes during the training phase, rather than maintaining a constant size throughout. This approach has been shown to improve the performance of the existing system by 0.4%. For our training and testing, we used ‘ConvNext’ equipped with pre-trained weights, which further contributed to the robustness of our model.
Akumalla Brahma Reddy, Bach-Tung Pham, Jia-Ching Wang
AVSS3
2024 Knowledge Sharing via Mimicking Attention Guided-Discriminative Features in Whole Slide Image Classification
abstract
The difficulty of collecting histopathology whole slide images (WSIs) and lack of disease-positives within slide image is a major obstacle to the development of computer-aided diagnosis. Existing works suggest sharing knowledge learned by mimicking the discriminative features, in which a student model with lack-features is trained to mimic a teacher model with rich-features. However, most feature mimicking methods, designed for natural image tasks, might be failed in the case of whole slide images. We propose a new method to mimic features for knowledge sharing in WSI classification. On the one hand, attention guided feature selection and normalization is proposed to extract discriminative features from a learning model and use attention scores to quantify feature contributions so as to identify the diseases-positive regions (a.k.a Region of Interests). On the other hand, we propose to learn by mimicking the high-discriminative features based on disease-positive regions. Our method is evaluated on two datasets with rich-features and two datasets with lack-features. Results demonstrate that our proposed method can boost performance compared to competitive MIL and knowledge sharing methods at the WSI level.
Farchan Hakim Raswa, Chun-Shien Lu, Jia-Ching Wang
HealthCom3
2024 Scene Text Recognition Using Progressive Rectification Network And Spelling Error Correction Language Model
abstract
Scene text recognition has gained popularity in deep neural network research. Compared to document text recognition, scene text recognition faces challenges such as complex backgrounds, diverse fonts, and blurred characters. Visual and semantic information must be considered in text recognition. While recent research has focused on improving semantic information, most studies have used English text datasets. Directly applying these methods to Chinese text datasets may not be effective. This work proposes a vision model with progressive rectification network and a Chinese scene text recognition method that uses a robust error-correcting language model to correct errors predicted by vision models. Firstly, the proposed progressive rectification network is of more effectiveness on the multi-oriented scene text images compared to present rectification method. On the other hand, the designed language model is used to correct errors predicted by vision models. The language model can handle low-quality images, including blurred, occluded, or nonsensical text. Experiments demonstrate that our method outperforms recent classic and state-of-the-art methods, making it a more powerful and suitable option for Chinese scene text recognition.
Ming-Zheng Peng, Phuong-Thi Le, Cheng-Chun Wang, Chien-Yao Wang, Jia-Ching Wang
ICIP6
2024 mmAlphabet: Air Writing Alphabet Recognition System Based on mmWave FMCW Radar and Convolutional Neural Network
Chao-Wang Huang, Chien-Yao Wang, Jia-Ching Wang
ICPR (28)3
2024 Attention-Guided Prototype Mixing: Diversifying Minority Context on Imbalanced Whole Slide Images Classification Learning
abstract
Real-world medical datasets often suffer from class imbalance, which can lead to degraded performance due to limited samples of the minority class. In another line of research, Transformer-based multiple instance learning (Transformer-MIL) has shown promise in addressing the pairwise correlation between instances in medical whole slide images (WSIs) with gigapixel resolution and non-uniform sizes. However, these characteristics pose challenges for state-of-the-art (SOTA) oversampling methods aiming at diversifying the minority context in imbalanced WSIs.In this paper, we propose an Attention-Guided Prototype Mixing scheme at the WSI level. We leverage Transformer-MIL training to determine the distribution of semantic instances and identify relevant instances for cutting and pasting across different WSI (bag of instances). To our knowledge, applying Transformer is often limited by memory requirements and time complexity, particularly when dealing with gigabyte-sized WSIs. We introduce the concept of prototype instances that have smaller representations while preserving the uniform size and intrinsic features of the WSI.We demonstrate that our proposed method can boost performance compared to competitive SOTA oversampling and augmentation methods at an imbalanced WSI level.
Farchan Hakim Raswa, Chun-Shien Lu, Jia-Ching Wang
WACV3
2024 Multi-view and multi-augmentation for self-supervised visual representation learning
Van-Nhiem Tran, Chi-En Huang, Shen-Hsuan Liu, Muhammad Saqlain Aslam, Kai-Lin Yang, Yung-Hui Li, Jia-Ching Wang
Appl. Intell.7
2024 LCSL: Long-Tailed Classification via Self-Labeling
abstract
During the last decades, deep learning (DL) has been proven to be a very powerful and successful technique in many real-world applications, e.g., video surveillance or object detection. However, when class label distributions are highly skewed, DL classifiers tend to be biased towards majority classes during training phases. This leads to poor generalization of minority classes and consequently reduces the overall accuracy. How to effectively deal with this long-tailed class distribution in DL, i.e., deep long-tailed classification (DLC), remains a challenging problem despite many research efforts. Among various approaches, data augmentation, which aims at generating more samples for reducing label imbalance, is the most common and practical one. However, simply relying on existing class-agnostic augmentation strategies without properly considering the label differences would worsen the problem since more head-class samples can be inevitably augmented than tail-class ones. Moreover, none of the existing works consider the quality and suitability of augmented samples during the training process. Our proposed approach, called Long-tailed Classification via Self-Labeling (LCSL), is specifically designed to address these limitations. LCSL fundamentally differs from existing works by the way it iteratively exploits the preceding network during the training process to re-label the labeled augmented samples and uses the output confidence to decide whether new samples belong to minority classes before adding them to the data. Not only does this help to reduce imbalance ratios among classes, but this also helps to reduce the uncertainty of class prediction problems for minority classes by selecting more confident samples to the data. This incremental learning and generating scheme thus provide a new robust approach for decreasing model over-fitting, thus enhancing the overall accuracy, especially for minority classes. Extensive experiments have demonstrated that LCSL acquires better performance than state-of-the-art long-tailed learning techniques on various standard benchmark datasets. More specifically, our LCSL obtains 85.8%, 54.4%, and 56.2% in terms of accuracy on CIFAR10-LT, CIFAR100-LT, and ImageNet-LT (with moderate to extreme imbalance ratios), respectively. The source code is available athttps://github.com/vdquang1991/lcsl/.
Duc-Quang Vu, Trang T. T. Phung, Jia-Ching Wang, Son T. Mai
IEEE Trans. Circuits Syst. Video Technol.3
2024 HAPiCLR: heuristic attention pixel-level contrastive loss representation learning for self-supervised pretraining
Van-Nhiem Tran, Shen-Hsuan Liu, Chi-En Huang, Muhammad Saqlain Aslam, Kai-Lin Yang, Yung-Hui Li, Jia-Ching Wang
Vis. Comput.7
2023 EMIX: A Data Augmentation Method for Speech Emotion Recognition
abstract
In the last few years, many deep learning (DL) models have been developed to improve the accuracy of speech emotion recognition (SER). However, as SER datasets are generally small and insufficient due to their difficult and expensive collection, the DL models are prone to overfitting, so their performance is limited. In this paper, we introduce a novel data augmentation (DA) method for the SER problem, namely EMix, which is simple but effective. The method creates new data by mixing pairs of selective samples from the original data. The generated mixtures will be noisier or less ambiguous than their constructive ones. To verify the effectiveness of the proposed DA, we develop a transformer-based network for the SER task, and experiment with the two public datasets including IEMOCAP and Crema-D. The experimental results demonstrate the superiority of EMix over other DA methods. In comparison with state-of-the-art methods, our approach shows competitive performance.
An Dang, Toan H. Vu, Le Dinh Nguyen, Jia-Ching Wang
ICASSP4
2023 Code-Switching Speech Synthesis Based on Self-Supervised Learning and Domain Adaptive Speaker Encoder
abstract
Recently, end-to-end speech synthesis models based on deep learning have made great progress in speech quality, and gradually replaced traditional speech synthesis methods into the mainstream. However, these methods are still challenging to synthesize highly natural speech. In order to solve the above problems, we introduce self-supervised learning and frame-level domain adversarial training into the speaker encoder based on the speaker verification task, so that the speaker vectors of different languages keep a consistent distribution in the speaker space, and the performance of speech synthesis is improved. In addition, we use a non-autoregressive speech synthesis model in the selection of speech synthesis model, so as to solve the problem of unnatural speech rate caused by cross-language speech synthesis. We first demonstrate that in the mixed language dataset of LibriTTS and AISHELL3, the speaker encoder trained with self-supervised representation has a 4.968% absolute EER reduction compared to the traditional MFCC on the speaker verification task, indicating that self-supervised representation has better generalization for domain-complex datasets. Then we obtain MOS scores of 3.635 and 3.675 for speech naturalness and speaker similarity in the code-switching speech synthesis task, respectively. Our approach simplifies the need to use multiple monolingual encoders to model linguistic information in the past literature, and adds frame-level domain adversarial training to optimize the speaker vectors in the speaker feature space to facilitate the code-switching speech synthesis task.
Yi-Xing Lin, Cheng-Hsun Pai, Phuong Le Thi, Bima Prihasto, Chien-Ling Huang, Jia-Ching Wang
ICASSP6
2023 Dense Adversarial Transfer Learning Based On Class-Invariance
abstract
This work proposes the dense adversarial transfer learning based on class-invariance, which is a novel, unsupervised, conditional adversarial domain adaptation approach. The proposed framework concatenates feature maps from the last layer of each backbone’s block to improve transfer learning; these features are weighted and densely connected to the features of each block along with the gradient-reversal layer. Classifiers are also added to the domain discriminators so that the network not only retains the classifying abilities when learning the domain-invariant features, but also has its domain adaptation abilities improved. In the experiment, the benchmark dataset Office-31 is used to compare the performance of similar existing frameworks. In three transfer tasks, the proposed method enhances the accuracy by approximately 3% to 5%, demonstrating the improvement provided by the proposed network towards unsupervised domain adaptation.
Bach-Tung Pham, Ting-Yu Wang, Phuong Le Thi, Khai-Thinh Nguyen, Yuan-Shan Lee, Tzu-Chiang Tai, Jia-Ching Wang
ICASSP7
2023 CNEG-VC: Contrastive Learning Using Hard Negative Example In Non-Parallel Voice Conversion
abstract
Contrastive learning has advantages for non-parallel voice conversion, but the previous conversion results could be better and more preserved. In previous techniques, negative samples were randomly selected in the features vector from different locations. A positive example could not be effectively pushed toward the query examples. We present contrastive learning in non-parallel voice conversion to solve this problem using hard negative examples. We named it CNEG-VC. Specifically, we teach the generator to generate negative examples. Our proposed generator has specific features. First, Instance-wise negative examples are generated based on voice input. Second, when taught with an adversarial loss, it can produce hard negative examples. The generator significantly improves non-parallel voice conversion performance. Our CNEG-VC achieved state-of-the-art results by outperforming previous techniques.
Bima Prihasto, Yi-Xing Lin, Phuong Le Thi, Chien-Lin Huang, Jia-Ching Wang
ICASSP5
2023 Discriminative Vector Learning with Application to Single Channel Speech Separation
abstract
In this paper, we introduce a discriminative vector learning method and apply it to single-channel speech separation. First, speech samples are transformed into discriminative vectors using two backbone networks. These vectors are easily separated by simple clustering algorithms. Among them, vectors with lower similarity are separated into different clusters, while vectors in the same cluster have higher similarity. This property is very important in image segmentation, audio separation, and data clustering problems. In our work, we design the network architecture to improve the discriminativeness of vectors through learning, taking this task as spectrogram segmentation. Experiments show that our method significantly improves performance compared to other deep clustering methods for speech separation.
Ha Minh Tan, Kai-Wen Liang, Jia-Ching Wang
ICASSP3
2023 Selinet: A Lightweight Model for Single Channel Speech Separation
abstract
The time-domain speech separation methods adopting deep learning have obtained impressive performance. However, the computational complexity, model size, and performance are still the challenges for the implementation on real-time low-resource devices. In this paper, we introduce a lightweight yet effective network for speech separation, namely SeliNet. The SeliNet is the one-dimensional convolutional architecture that employs bottleneck modules, and atrous temporal pyramid pooling. In bottleneck modules, the depth-wise separable convolution significantly decreases the model size and computational cost meanwhile the squeeze excitation uses a context vector to interact with the entire hidden state vector. Specifically, the atrous temporal pyramid pooling recognizes long-time sequences of various lengths and extracts context at different field-of-views. This helps SeliNet to obtain impressive performance while still maintaining the small computational cost and model size.
Ha Minh Tan, Duc-Quang Vu, Jia-Ching Wang
ICASSP3
2023 3D Face Reconstruction Based on Weakly-Supervised Learning Morphable Face Model
abstract
In this paper, we propose a system for 3D face model reconstruction. Earlier studies on reconstruction methods included the software modeling methods or the instrument scanning modeling methods. But both of the above methods require a lot of development resources and time costs. Therefore, we develop a reconstruction system using a weakly supervised approach combining Convolutional Neural Networks (CNN) and 3D Morphable Face Models (3DMM). Given a sufficient number of 2D face images to train and learn the main features of the face, our system is capable of rapidly constructing 3D face models. The proposed method enhances the efficiency of preprocessing and improves the performance of loss function through image depth feature extraction and regression coefficients. Using two datasets for model evaluation and analysis, this study efficiently reconstructs faces without ground-truth labels.
Kai-Wen Liang, Pin-Hsuan Li, Chung-Hsun Lo, Chien-Yao Wang, Yung-Fang Chen, Jia-Ching Wang, Pao-Chi Chang
ICIP6
2023 Cyclic Transfer Learning for Mandarin-English Code-Switching Speech Recognition
abstract
Transfer learning is a common method to improve the performance of the model on a target task via pre-training the model on pretext tasks. Different from the methods using monolingual corpora for pre-training, in this study, we propose a Cyclic Transfer Learning method (CTL) that utilizes both code-switching (CS) and monolingual speech resources as the pretext tasks. Moreover, the model in our approach is always alternately learned among these tasks. This helps our model can improve its performance via maintaining CS features during transferring knowledge. The experiment results on the standard SEAME Mandarin-English CS corpus have shown that our proposed CTL approach achieves the best performance with Mixed Error Rate (MER) of 16.3% on test$_{man}$, 24.1% on test$_{sge}$. In comparison to the baseline model that was pre-trained with monolingual data, our CTL method achieves 11.4% and 8.7% relative MER reduction on the test$_{man}$and test$_{sge}$sets, respectively. Besides, the CTL approach also outperforms compared to other state-of-the-art methods. The source code of the CTL method can be found athttps://github.com/caohongnga/CTL-CSSR.
Cao Hong Nga, Duc-Quang Vu, Huong Hoang Luong, Chien-Lin Huang, Jia-Ching Wang
IEEE Signal Process. Lett.5
2023 Anti-aliasing convolution neural network of finger vein recognition for virtual reality (VR) human-robot equipment of metaverse
Nghi C. Tran, Toan H. Vu, Tzu-Chiang Tai, Jia-Ching Wang
J. Supercomput.5
2022 Selective Mutual Learning: An Efficient Approach for Single Channel Speech Separation
abstract
Mutual learning, the related idea to knowledge distillation, is a group of untrained lightweight networks, which simultaneously learn and share knowledge to perform tasks together during training. In this paper, we propose a novel mutual learning approach, namely selective mutual learning. This is the simple yet effective approach to boost the performance of the networks for speech separation. There are two networks in the selective mutual learning method, they are like a pair of friends learning and sharing knowledge with each other. Especially, the high-confidence predictions are used to guide the remaining network while the low-confidence predictions are ignored. This helps to remove poor predictions of the two networks during sharing knowledge. The experimental results have shown that our proposed selective mutual learning method significantly improves the separation performance compared to existing training strategies including independently training, knowledge distillation, and mutual learning with the same network architecture.
Ha Minh Tan, Duc-Quang Vu, Chung-Ting Lee, Yung-Hui Li, Jia-Ching Wang
ICASSP5
2022 (2+1)D Distilled ShuffleNet: A Lightweight Unsupervised Distillation Network for Human Action Recognition
abstract
While most existing deep neural networks (DNN) architectures are proposed for increasing performance, they also raise overall model complexity. However, practical applications require lightweight DNN models, that are able to run real-time in edge computing devices. In this work, we present a simple and elegant unsupervised distillation learning paradigm to train a lightweight network to human action recognition called (2+1)D Distilled ShuffleNet. Leveraging the distilling technique, the proposed method allows us to create a lightweight DNN model that achieves high accuracy and real-time speed. Our lightweight (2+1)D Distilled ShuffleNet is designed as an unsupervised paradigm; it does not require labelled data during distilling knowledge from the teacher to the student. Furthermore, to help the student be more "intelligent", we propose to distill the knowledge from two different teachers, i.e., 2D teacher and 3D teacher. The experimental results have shown that our lightweight (2+1)D Distilled ShuffleNet outperforms other state-of-the-art distillation networks with 86.4% and 59.9% top-1 accuracy on UCF101 and HMDB51 datasets, respectively, whereas the inference running time is at 47.16 FPS on CPU with only 17.1M parameters and 12.07 GFLOPs.
Duc-Quang Vu, T. Hoang Ngan Le, Jia-Ching Wang
ICPR3
2022 A Comparative Study of Cross-Model Universal Adversarial Perturbation for Face Forgery
abstract
Although the rapid development of deep generative models (DGM) enables diverse applications of content creation, increasing illegal uses of the technologies also severely threaten the privacy and security of personal information, especially for faces. Several previous works have been proposed to leverage adversarial attacks to fight against these malicious manipulations by adding an imperceptible perturbation to each input image to disrupt the output. In addition, to improve its scalability, a sequential cross-model universal perturbation attack has been proposed to learn a common adversarial perturbation to defend the images from the manipulation of multiple DGMs. However, we find that the order of DGMs for the adversarial perturbation generation does matter and influence the final defense performance. To address this issue, we propose to generate the universal perturbation through joint optimization of multiple DGMs. From the extensive experimental results, we find that the universal perturbation generated by the proposed method can successfully disrupt the output faces of multiple DGMs at the same time and achieves higher attack success rates than the previous state-of-the-art method based on the sequential generation, even under the situations where the model robustness of DGMs are enhanced by random perturbations.
Shuo-Yen Lin, Jun-Cheng Chen, Jia-Ching Wang
VCIP3
2022 Spectral-Temporal Receptive Field-Based Descriptors and Hierarchical Cascade Deep Belief Network for Guitar Playing Technique Classification
abstract
Music information retrieval is of great interest in audio signal processing. However, relatively little attention has been paid to the playing techniques of musical instruments. This work proposes an automatic system for classifying guitar playing techniques (GPTs). Automatic classification for GPTs is challenging because some playing techniques differ only slightly from others. This work presents a new framework for GPT classification: it uses a new feature extraction method based on spectral-temporal receptive fields (STRFs) to extract features from guitar sounds. This work applies a supervised deep learning approach to classify GPTs. Specifically, a new deep learning model, called the hierarchical cascade deep belief network (HCDBN), is proposed to perform automatic GPT classification. Several simulations were performed and the datasets of: 1) data on onsets of signals; 2) complete audio signals; and 3) audio signals in a real-world environment are adopted to compare the performance. The proposed system improves upon the F-score by approximately 11.47% in setup 1) and yields an F-score of 96.82% in setup 2). The results in setup 3) demonstrate that the proposed system also works well in a real-world environment. These results show that the proposed system is robust and has very high accuracy in automatic GPT classification.
Chien-Yao Wang, Pao-Chi Chang, Jian-Jiun Ding, Tzu-Chiang Tai, Andri Santoso, Yu-Ting Liu, Jia-Ching Wang
IEEE Trans. Cybern.7
2021 A Novel Self-Knowledge Distillation Approach with Siamese Representation Learning for Action Recognition
abstract
Knowledge distillation is an effective transfer of knowledge from a heavy network (teacher) to a small network (student) to boost students' performance. Self-knowledge dis-tillation, the special case of knowledge distillation, has been proposed to remove the large teacher network training process while preserving the student's performance. This paper intro-duces a novel Self-knowledge distillation approach via Siamese representation learning, which minimizes the difference between two representation vectors of the two different views from a given sample. Our proposed method, SKD-SRL, utilizes both soft label distillation and the similarity of representation vectors. Therefore, SKD-SRL can generate more consistent predictions and representations in various views of the same data point. Our benchmark has been evaluated on various standard datasets. The experimental results have shown that SKD-SRL significantly improves the accuracy compared to existing supervised learning and knowledge distillation methods regardless of the networks.
Duc-Quang Vu, Thi-Thu-Trang Phung, Jia-Ching Wang
VCIP3
2020 Encoder-Recurrent Decoder Network for Single Image Dehazing
abstract
This paper develops a deep learning model, called Encoder-Recurrent Decoder Network (ERDN), which recovers the clear image from a degrade hazy image without using the atmospheric scattering model. The proposed model consists of two key components- an encoder and a decoder. The encoder is constructed by a residual efficient spatial pyramid (rESP) module such that it can effectively process hazy images at any resolution to extract relevant features at multiple contextual levels. The decoder has a recurrent module which sequentially aggregates encoded features from high levels to low levels to generate haze-free images. The network is trained end-to-end given pairs of hazy-clear images. Experimental results on the RESIDE-Standard dataset demonstrate that the proposed model achieves a competitive dehazing performance compared to the state-of-the-art methods in term of PSNR and SSIM.
An Dang, Toan H. Vu, Jia-Ching Wang
ICASSP3
2020 Learning to Remember Beauty Products
abstract
This paper develops a deep learning model for the beauty product image retrieval problem. The proposed model has two main components- an encoder and a memory. The encoder extracts and aggregates features from a deep convolutional neural network at multiple scales to get feature embeddings. With the use of an attention mechanism and a data augmentation method, it learns to focus on foreground objects and neglect background on images, so can it extract more relevant features. The memory consists of representative states of all database images as its stacks, and it can be updated during training process. Based on the memory, we introduce a distance loss to regularize embedding vectors from the encoder to be more discriminative. Our model is fully end-to-end, requires no manual feature aggregation and post-processing. Experimental results on the Perfect-500K dataset demonstrate the effectiveness of the proposed model with a significant retrieval accuracy.
Toan H. Vu, An Dang, Jia-Ching Wang
ACM Multimedia3
2020 Sound Events Recognition and Retrieval Using Multi-Convolutional-Channel Sparse Coding Convolutional Neural Networks
abstract
This article proposes two novel deep convolutional neural networks (CNN), which are called the sparse coding convolutional neural network (SC-CNN) and the multi-convolutional-channel SC-CNN (MSC-CNN), to address the sound event recognition and retrieval problem. Unlike the general framework of a CNN, in which the feature learning process is performed hierarchically, the proposed framework models the whole memorization process in the human brain, including encoding, storage, and recollection. In particular, the MSC-CNN is designed to recognize multiple sound events that occur simultaneously. The experimental results indicate that the proposed SC-CNN and MSC-CNN outperforms the state-of-the-art systems in sound event recognition and retrieval.
Chien-Yao Wang, Tzu-Chiang Tai, Jia-Ching Wang, Andri Santoso, Seksan Mathulaprangsan, Chin-Chin Chiang, Chung-Hsien Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Speaker Characterization Using TDNN-LSTM Based Speaker Embedding
abstract
In this paper we propose speaker characterization using time delay neural networks and long short-term memory neural networks (TDNN-LSTM) speaker embedding. Three types of front-end feature extraction are investigated to find good features for speaker embedding. Three kinds of data augmentation are used to increase the amount and diversity of the training data. The proposed methods are evaluated with the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) tasks. Experimental results show that the proposed methods achieve a decision cost of 0.400 with the pooled SRE 2018 development set with a single system. In addition, by applying simple average score combination on the outputs of 12 systems, the proposed methods achieve an equal error rate (EER) of 5.56% and a minimum decision cost function of 0.423 with the SRE 2016 evaluation set.
Chia-Ping Chen, Su-Yu Zhang, Chih-Ting Yeh, Jia-Ching Wang, Tenghui Wang, Chien-Lin Huang
ICASSP4
2019 Object Bounding Transformed Network for End-to-End Semantic Segmentation
abstract
In recent years, numerous studies of the use of a Fully Convolutional Network (FCN) for image semantic segmentation have been published. This work introduces an end-to-end Object Bounding Transformed Network (OBTNet) which combines the advantages of the Object Boundary Guided (OBG) and Doman Transform (DT). OBG is an object boundary based approach that increases the integrity of object shape. Based on OBG, we propose an Object Boundary Network (OBN) as the object region and object boundary generator. In addition, our system achieves object region preserving and object boundary preserving by employing DT. The proposed system uses the pretrained multi-scale ResNet101 as the base network and uses multi-scale atrous convolution to preserve the dimensions of the feature map, increasing the accuracy of semantic segmentation. Experiments show that our system yielded a mean IOU of 77.74% and outperformed the baseline model on the VOC2012 test set.
Kuan-Chung Wang, Chien-Yao Wang, Tzu-Chiang Tai, Jia-Ching Wang
ICIP4
2019 Deep Learning Based Vietnamese Diacritics Restoration
abstract
Diacritics are very important in diacritical languages, because the meaning of sentences can be changed in accordance to diacritics. Writing without diacritics makes the sentences ambiguous; however, there are several reasons make people do not write words with diacritics, such as fast typing, convenience, or texting on unsupported diacritics devices. As a result, these texts are very difficult to process on further natural language processing (NLP) tasks like machine translation, sentiment analysis, or question answering system. Therefore, diacritics restoration is critical for further usage or processing in NLP related tasks. In this study, we propose a method which combines convolutional neural network (CNN) and bidirectional gated recurrent unit (Bi-GRU) to restore diacritics. In addition, we use residual block to resolve vanishing gradient problem of recurrent neural networks. We applied the model for restoring diacritics of Vietnamese language that has the highest ratio of diacritics in words. This approach has character accuracy at 98.63% and word accuracy at 94.77%.
Cao Hong Nga, Khai-Thinh Nguyen, Pao-Chi Chang, Jia-Ching Wang
ISM4
2019 Sentiment Analysis Using Residual Learning with Simplified CNN Extractor
abstract
Sentiment analysis has an important role in social media monitoring as it extracts public opinions, emotions, and feelings about certain products or services. There are several publications in building a system to identify opinions from text using rule-based approach, lexicon-based approach, or machine learning. In this paper, we propose and compare several deep learning models to solve sentiment analysis problem of the Internet Movie Database (IMDb) review sentiment dataset. The feature extractor consists of a convolutional layer, followed by a max pooling layer and a batch normalization layer. To solve the vanishing gradient problem, we use a residual connection to concatenate the input values with the extracted features before feeding the output into a recurrent layer. Our best model has an accuracy of 90.02%.
Khai-Thinh Nguyen, Cao Hong Nga, Yuan-Shan Lee, Meng-Lun Wu, Pao-Chi Chang, Jia-Ching Wang
ISM6
2018 Image Representation Using Supervised and Unsupervised Learning Methods on Complex Domain
abstract
Matrix factorization (MF) and its extensions have been intensively studied in computer vision and machine learning. In this paper, unsupervised and supervised learning methods based on MF technique on complex domain are introduced. Projective complex matrix factorization (PCMF) and discriminant projective complex matrix factorization (DPCMF) present two frameworks of projecting complex data to a lower dimension space. The optimization problems are formulated as the minimization of the real-valued functions of complex variables. Motivated by independence among extracted features, Fisher linear discriminant is used as hard constraint on supervised model. Experimental results on facial expression recognition (FER) show improved classification performance in comparison to real-valued features of both unsupervised and supervised NMFs.
Manh-Quan Bui, Viet-Hang Duong, Yung-Hui Li, Tzu-Chiang Tai, Jia-Ching Wang
ICASSP5
2018 Locality-Preserving Complex-Valued Gaussian Process Latent Variable Model for Robust Face Recognition
abstract
Learning a low-dimensional image representation yields effective and efficient face recognition. The use of such a representation helps to weaken the curse of dimensionality. However, the traditional facial representation method is not robust against partial occlusions or variations of expression. To solve this problem, this paper proposes a more reliable, complex-valued representation of facial image. The robust representation is based on the proposed locality-preserving complex-valued Gaussian process latent variable model (LP-CGPLVM). In the LP-CGPLVM, the Euler formula is utilized to transform original facial images into the complex domain. A proper complex GP is employed to model the mapping between the complex-valued high-dimensional data and the corresponding low-dimensional representation. Moreover, the locality-preserving constraint is taken into consideration to preserve the neighborhood data structure. The experimental results indicate that our proposed method is robust against partial occlusions and various facial expressions.
Sih-Huei Chen, Yuan-Shan Lee, Yu-Sheng Hsu, Chung-Hsien Wu 0001, Jia-Ching Wang
ICASSP5
2018 Complex-Valued Gaussian Process Latent Variable Model for Phase-Incorporating Speech Enhancement
abstract
Traditional speech enhancement techniques modify the magnitude of a speech in time-frequency domain, and use the phase of a noisy speech to resynthesize a time domain speech. This work proposes a complex-valued Gaussian process latent variable model (CGPLVM) to enhance directly the complex-valued noisy spectrum, modifying not only the magnitude but also the phase. The main idea that underlies the developed method is the modeling of short-time Fourier transform (STFT) coefficients across the time frames of a speech as a proper complex Gaussian process (GP) with noise added. The proposed method is based on projecting the spectrum into a low-dimensional subspace. Experiments were carried out on the CHTTL database, which contains the digits zero to nine in Mandarin. Several standard measures are used to demonstrate that the proposed method outperforms baselines with various types of noise and SNR levels.
Sih-Huei Chen, Yuan-Shan Lee, Jia-Ching Wang
ICASSP3
2018 Depth Human Action Recognition Based on Convolution Neural Networks and Principal Component Analysis
abstract
In this work, we address human action recognition problem under viewpoint variation. The proposed model is formulated by wisely combining convolution neural network (CNN) model with principle component analysis (PCA). In this context, we pass real depth videos through a CNN model in a frame-wise manner. The view invariant features are extracted by employing convolution layers as mid-outputs and considered as 3D nonnegative tensors. The PCA algorithm is separately imposed on view-invariant high-level space of image and video groups to seek both local and holistic hidden dynamics information. To deal with noisy data and temporal misalignment, we utilize the Fourier temporal pyramid to encode temporal and obtain the final descriptors. Our proposed framework supplies a robust discriminative representation with low dimension and computational requirements. We evaluate the proposed method on two standard multiview depth video datasets. The experimental results show that our method has superior performance compared to other approaches.
Manh-Quan Bui, Viet-Hang Duong, Tzu-Chiang Tai, Jia-Ching Wang
ICIP4
2018 Playing Technique Classification Based on Deep Collaborative Learning of Variational Auto-Encoder and Gaussian Process
abstract
Modeling musical timbre is critical for various music information retrieval (MIR) tasks. This work addresses the task of classifying playing techniques, which involves extremely subtle variations of timbre among different categories. A deep collaborative learning framework is proposed to represent a music with greater discriminative power than previously achieved. Firstly, a novel variational autoencoder (VAE) is developed to eliminate the variation of acoustic features within a class. Secondly, a Gaussian process classifier is jointly learned to distinguish the variations of timbres between classes, which increases the discriminative power of the learned representations. We derive a new lower bound that guides a VAE-based representation. Experiments were conducted on a database of seven classes of guitar playing techniques. The experimental results demonstrated that the proposed method outperforms baselines in terms of the Fl-score and accuracy.
Sih-Huei Chen, Yuan-Shan Lee, Min-Che Hsieh, Jia-Ching Wang
ICME4
2018 Locality Preserving Discriminative Complex-Valued Latent Variable Model
abstract
Techniques for analyzing complex-valued data are required in numerous fields, such as signal processing. This work develops a novel complex-valued latent variable model, named locality-preserving discriminative complex-valued Gaussian process latent variable model (LPD-CGPLVM), for discovering a compressed complex-valued representation of data. The developed LPD-CGPLVM operates on the complex-valued domain. Additionally, we attempt to preserve both global and local data structures while promoting discrimination. A new objective function that imposes a locality-preserving and a discriminative term for complex-valued data is presented. Complex-valued gradient descent is then utilized to obtain a complex-valued representation of high-dimensional data and the hyperparameters in the LPD-CGPLVM. The proposed method was evaluated using two pattern recognition applications - face recognition with occlusion and music emotion recognition. The experimental results thus obtained demonstrated the superior accuracy of the proposed method, especially for situations with only a small number of training data.
Sih-Huei Chen, Yuan-Shan Lee, Jia-Ching Wang
ICPR3
2018 Learning a Hierarchical Latent Semantic Model for Multimedia Data
abstract
This paper develops a hierarchical feature representation that is based on a Bayesian non-parametric method. Feature learning is an important issue in classification and data analysis. It can improve the classification performance and increase the convenience of data processing and analysis. Popular methods of representation learning include methods that are based on mixture models or dictionary learning methods. However, current methods have some disadvantages. The use of a traditional mixture model, such as the Gaussian mixture model (GMM), involves the model selection problem and suffers a lack of hierarchy between components. Inspired by h-LDA, distance-based Gaussian hierarchical Dirichlet allocation (distance-based GhLDA) is proposed herein. This method can automatically determine the number of components and construct a hierarchical representation. The distance function between data is used in the prior distribution. The learnt representation in the proposed model has the advantage of hLDA, which can handle shared components and distinct components. The quantization loss problem, which commonly arises when a topic model is used to deal with continuous data, can be solved by assuming that the distribution of words follows a Gaussian rather than a Dirichlet distribution. The performance of the proposed model in solving audio and image classification problems is evaluated. Experimental results indicate that the distance-based GhLDA outperforms baseline methods.
Shao Hui Wu, Yuan-Shan Lee, Sih-Huei Chen, Jia-Ching Wang
ICPR4
2018 Single-Channel Speech Separation Based on Gaussian Process Regression
abstract
Gaussian process (GP) is a flexible kernel-based learning method which has found widespread applications in signal processing. In this paper, a supervised approach is proposed to handle single-channel speech separation (SCSS) problem. We focus on modeling a nonlinear mapping between mixed and clean speeches based on GP regression, in which reconstructed audio signal is estimated by the predictive mean of GP model. The nonlinear conjugate gradient method was utilized to perform the hyper-parameter optimization. The experiment on a subset of TIMIT speech dataset is carried out to confirm the validity of the proposed approach.
Nguyen Le, Sih-Huei Chen, Tzu-Chiang Tai, Jia-Ching Wang
ISM4
2018 Predicting the Probability Density Function of Music Emotion Using Emotion Space Mapping
abstract
Computationally modeling the affective content of music has been intensively studied in recent years because of its wide applications in music retrieval and recommendation. Although significant progress has been made, this task remains challenging due to the difficulty in properly characterizing the emotion of a music piece. Music emotion perceived by people is subjective by nature and thus complicates the process of collecting the emotion annotations as well as developing the predictive model. Instead of assuming people can reach a consensus on the emotion of music, in this work we propose a novel machine learning approach that characterizes the music emotion as a probability distribution in the valence-arousal (VA) emotion space, not only tackling the subjectivity but also precisely describing the emotions of a music piece. Specifically, we represent the emotion of a music piece as a probability density function (PDF) in the VA space via kernel density estimation from human annotations. To associate emotion with the audio features extracted from music pieces, we learn the combination coefficients by optimizing some objective functions of audio features, and then predict the emotion of an unseen piece by linearly combining the PDFs of the training pieces with the coefficients. Several algorithms for learning the coefficients are studied. Evaluations on the NTUMIR and MediaEval 2013 datasets validate the effectiveness of the proposed methods in predicting the probability distributions of emotion from audio features. We also demonstrate how to use the proposed approach in emotion-based music retrieval.
Yu-Hao Chin, Jia-Ching Wang, Ju-Chiang Wang, Yi-Hsuan Yang
IEEE Trans. Affect. Comput.2
2018 Sound Event Recognition Using Auditory-Receptive-Field Binary Pattern and Hierarchical-Diving Deep Belief Network
abstract
Automatic sound event recognition (SER) has recently attracted renewed interest. Although practical SER system has many useful applications in everyday life, SER is challenging owing to the variations among sounds and noises in the real-world environment. This paper presents a novel feature extraction and classification method to solve the problem of SER. An audio-visual descriptor, called the auditory-receptive-field binary pattern, is designed based on the spectrogram image feature, the cepstral features, and the human auditory receptive field model. The extracted features are then fed into a classifier to perform event classification. The proposed classifier, called the hierarchical-diving deep belief network, is a deep neural network system that hierarchically learns the discriminative characteristics from physical feature representation to the abstract concept. The performance of our proposed system was verified using several experiments under various conditions. Using the RWCP dataset, the proposed system achieved a recognition rate of 99.27% for real-world sound data in 105 categories. Under noisy conditions, the developed system is very robust, with which it achieved 95.06% recognition rate with 0 dB signal-to-noise ratio. Using the TUT sound event dataset, the proposed system achieves error rates of 0.81 and 0.73 in sound event detection in home and residential area scenes. The experimental results reveal that the proposed system outperformed the other systems in this field.
Chien-Yao Wang, Jia-Ching Wang, Andri Santoso, Chin-Chin Chiang, Chung-Hsien Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Kernel weighted Fisher sparse analysis on multiple maps for audio event recognition
abstract
This work presents a novel approach for audio event recognition. The approach develops a weighted kernel fisher sparse analysis method based on multiple maps. The proposed method consists of maps extraction and kernel weighted Fisher sparse analysis. Two maps are firstly extracted from each audio file, i.e. scale-frequency map and damping-frequency map. The scale and frequency of the Gabor atoms are extracted to construct a scale-frequency map. On the other hand, the damping-frequency map is generated according to the frequency and damping factor of damped atoms. Gabor atoms can be utilized to model human auditory perception, and the damped atoms can be used to model commonly observed damped oscillations in natural signals. This work fuses the advantages of these two dictionaries to improve the performance of the system. During the recognition stage, this work constructs a kernel sparse representation-based classifier via the proposed kernel weighted Fisher sparse analysis to enhance separability. The proposed kernel weighted Fisher sparse analysis combines sparse representation with heteroscedastic kernel weighted discriminant analysis (HKWDA), which is useful for providing a discriminative recognition of audio events because a weighted pairwise Chernoff criterion is utilized in the kernel space. Experiments on a 20-class audio event database indicate that the proposed approach can achieve an accuracy rate of 82.70%. Also, integrating the scale-frequency map with MFCCs increases the accuracy rate to 87.70%.
Yu-Hao Chin, Bo-Wei Chen, Jia-Ching Wang
ICASSP3
2017 Exemplar-embed complex matrix factorization for facial expression recognition
abstract
This paper presents an image representation approach which is based on matrix factorization in the complex domain and called exemplar-embed complex matrix factorization (EE-CMF). The proposed EE-CMF approach can very effectively improve the performance of facial expression recognition. Moreover, Wirtinger's calculus was employed to determine derivatives. The gradient descent method was utilized to solve the complex optimization problem. Experiments on facial expression recognition verified the effectiveness of the proposed EE-CMF. It provides consistently better recognition results than standard NMFs.
Viet-Hang Duong, Yuan-Shan Lee, Jian-Jiun Ding, Bach-Tung Pham, Manh-Quan Bui, Jia-Ching Wang
ICASSP7
2017 Multi-pitch streaming of interwoven streams
abstract
In this paper, we discuss the multipitch streaming (MPS) problem for a multi-source audio signal having interweaving pitch contours. We propose two approaches to tackle this challenge, one relates to a feature extracted from the energy levels distributed in multi-channel recordings for better characterization of the source, and the other uses particle swarm optimization (PSO) to enlarge the search space and alleviate the initialization problem in constrained clustering of the features representing different sources. Experiments on music and speech samples having highly interweaving pitch contours are presented to assess its effectiveness.
Chih Yi Kuan, Li Su 0004, Yu-Hao Chin, Jia-Ching Wang
ICASSP4
2017 Fully complex deep neural network for phase-incorporating monaural source separation
abstract
Deep neural network (DNN) have become a popular means of separating a target source from a mixed signal. Most of DNN-based methods modify only the magnitude spectrum of the mixture. The phase spectrum is left unchanged, which is inherent in the short-time Fourier transform (STFT) coefficients of the input signal. However, recent studies have revealed that incorporating phase information can improve the quality of separated sources. To estimate simultaneously the magnitude and the phase of STFT coefficients, this work paper developed a fully complex-valued deep neural network (FCDNN) that learns the nonlinear mapping from complex-valued STFT coefficients of a mixture to sources. In addition, to reinforce the sparsity of the estimated spectra, a sparse penalty term is incorporated into the objective function of the FCDNN. Finally, the proposed method is applied to singing source separation. Experimental results indicate that the proposed method outperforms the state-of-the-art DNN-based methods.
Yuan-Shan Lee, Chien-Yao Wang, Shu-Fan Wang, Jia-Ching Wang, Chung-Hsien Wu 0001
ICASSP4
2017 Dynamic tracking attention model for action recognition
abstract
This paper proposes a dynamic tracking attention model (DTAM), which mainly comprises a motion attention mechanism, a convolutional neural network (CNN) and long short-term memory (LSTM), to recognize human action in a video sequence. In the motion attention mechanism, the local dynamic tracking is used to track moving objects in feature domain and global dynamic tracking corrects the motion in the spectral domain. The CNN is utilized to perform feature extraction, while the LSTM is applied to handle sequential information about actions that is extracted from videos. It effectively fetches information between consecutive frames in a video sequence and has an even higher recognition rate than does the CNN-LSTM. Combining the DTAM with the visual attention model, the proposed algorithm has a recognition rate that is 3.6% and 4.5% higher than that of the CNN-LSTMs with and without the visual attention model, respectively.
Chien-Yao Wang, Chin-Chin Chiang, Jian-Jiun Ding, Jia-Ching Wang
ICASSP4
2017 Hierarchical joint-guided networks for semantic image segmentation
abstract
Semantic image segmentation is now an exciting area of research owing to its various useful applications in daily life. This paper introduces a hierarchical joint-guided network (HJGN) which is mainly composed of proposed hierarchical joint learning convolutional networks (HJLCNs) and proposed joint-guided and making networks (JGMNs). HJLCNs exhibit high robustness in the segmentation of unseen objects that are not contained in training categories. JGMNs are very effective in filling holes and preventing incorrect segmentation predictions. The proposed HJGNs outperform the state-of-the-art methods on the PASCAL VOC 2012 testing set, reaching a mean IU of 80.4%.
Chien-Yao Wang, Jyun-Hong Li, Seksan Mathulaprangsan, Chin-Chin Chiang, Jia-Ching Wang
ICASSP5
2017 Hand gesture recognition based on Bayesian sensing hidden Markov models and Bhattacharyya divergence
abstract
This work develops a system for recognizing common hand gestures. The main idea that underlies the developed system is the incorporation of Bhattacharyya divergence into Bayesian sensing hidden Markov models (BS-HMM). The system consists of two stages. First, a sequence of depth images is captured by Microsoft Kinect. The hand region is identified from the depth images by tracking the position of the hand using information about the skeleton, yielding the segmented depth images. A histogram of the oriented normal 4D (HON4D) and a histogram of oriented gradient (HOG) are then extracted from the segmented depth images to represent the motion patterns. Second, all training feature vectors are transformed by combining every k consecutive feature vectors into a sequence of distributions. The proposed Bhattacharyya divergence based BS-HMM (BDBS-HMM) is trained using the sequence of distributions. The proposed system is compared to the standard HMM and the BS-HMM using MSRGesture3D database and our database. Experimental results indicated that the proposed method outperforms the baseline methods.
Sih-Huei Chen, Ari Hernawan, Yuan-Shan Lee, Jia-Ching Wang
ICIP4
2017 Source separation using dictionary learning and deep recurrent neural network with locality preserving constraint
abstract
Deep learning is a popular method for monaural source separation, and especially for extracting a singing voice from a single-channel song. However, deep learning-based source separation ignores the geometrical structure of the input data. This work develops a novel approach to source separation that is based on non-negative matrix factorization (NMF) and deep recurrent neural networks (DRNN) with a locality-preserving constraint. First, NMF was used to learn patterns from training data. The learned patterns are linearly combined with the output of DRNN. Second, a locality-preserving constraint is developed to exploit the inner-structure of the input data in the DRNN learning process. Experimental results obtained using the MIR-1K dataset reveal that the proposed algorithm outperforms the baselines.
Pham Tuan, Yuan-Shan Lee, Seksan Mathulaprangsan, Jia-Ching Wang
ICME4
2017 Recognition and retrieval of sound events using sparse coding convolutional neural network
abstract
This paper proposes a novel deep convolutional neural network (CNN), called sparse coding convolutional neural network (SC-CNN), to address the problem of sound event recognition and retrieval task. Unlike the general framework of a CNN, in which feature learning process is performed hierarchically, the proposed framework models the whole memorizing procedures in the human brain, including encoding, storage, and recollection. Sound data from the RWCP sound scene dataset with added noise from NOISEX-92 noise dataset are used to compare the performance of the proposed system with the state-of-the-art baselines. The experimental results indicated that the proposed SC-CNN outperformed the state-of-the-art systems in sound event recognition and retrieval. In the sound event recognition task, the proposed system achieved an accuracy of 94.6%, 100% and 100% under 0db, 10db and clean noise conditions, respectively. In the retrieval task, the proposed system improves the mAP rate of the general CNN by approximately 6%.
Chien-Yao Wang, Andri Santoso, Seksan Mathulaprangsan, Chin-Chin Chiang, Chung-Hsien Wu 0001, Jia-Ching Wang
ICME6
2017 Discriminative Training of Complex-valued Deep Recurrent Neural Network for Singing Voice Separation
abstract
Deep neural networks (DNN) have performed impressively in the processing of multimedia signals. Most DNN-based approaches were developed to handle real-valued data; very few have been designed for complex-valued data, despite their being essential for processing various types of multimedia signal. Accordingly, this work presents a complex-valued deep recurrent neural network (C-DRNN) for singing voice separation. The C-DRNN operates on the complex-valued short-time discrete Fourier transform (STFT) domain. A key aspect of the C-DRNN is that the activations and weights are complex-valued. The goal herein is to reconstruct the singing voice and the background music from a mixed signal. For error back-propagation, CR-calculus is utilized to calculate the complex-valued gradients of the objective function. To reinforce model regularity, two constraints are incorporated into the objective function of the C-DRNN. The first is an additional masking layer that ensures the sum of separated sources equals the input mixture. The second is a discriminative term that preserves the mutual difference between two separated sources. Finally, the proposed method is evaluated using the MIR-1K dataset and a singing voice separation task. Experimental results demonstrate that the proposed method outperforms the state-of-the-art DNN-based methods.
Yuan-Shan Lee, Kuo Yu, Sih-Huei Chen, Jia-Ching Wang
ACM Multimedia4
2017 Single channel source separation using graph sparse NMF and adaptive dictionary learning
abstract
The aim of single channel source separation is to accurately recover signals from mixtures. Non-negative matrix factorization (NMF) is a popular method to separate mixed signals using learned dictionaries. These dictionaries can be produced efficiently by sparse NMF to approximate the input signal as closely as possible. However, the literature does not consider the structure of the data in terms of the similarity among vertices of the input signal. Furthermore, state-of-art variants of NMF that are more efficient than conventional ones have not been utilized, and the learned dictionary is typically fixed in the separating phase. This strategy is not favorable because the training data and the testing data totally differ. To deal with these issues, our work proposes a method that incorporates the graph regularization into group sparsity β-NMF to improve the performance of source separation. The proposed algorithms differ from those in the literature by using an adaptive dictionary in which particular characteristics of the testing data are updated to produce newer dictionaries. Experimental results demonstrate that our proposed method is outstandingly effective in speech separation in various scenarios, relative to the baseline.
Yuan-Shan Lee, Yan-Bo Lin, Yung-Hui Li, Tzu-Chiang Tai, Jia-Ching Wang
Intell. Data Anal.6
2017 Music emotion recognition using PSO-based fuzzy hyper-rectangular composite neural networks
abstract
This study proposed a novel system for recognising emotional content in music, and the proposed system is based on particle swarm optimisation (PSO)‐based fuzzy hyper‐rectangular composite neural networks (PFHRCNNs), which integrates three computational intelligence tools, i.e. hyper‐rectangular composite neural networks (HRCNNs), fuzzy systems, and PSO. PFHRCNN is flexible to the complex data due to the fuzzy membership estimation, and an optimisation of the parameters is provided by PSO. First, raw features are extracted from each music clips. After feature extraction, a HRCNN is separately constructed for each class. Each trained HRCNN will result in a set of crisp rules. A problem associated with these generated crisp rules is that some of them are ineffective; therefore, a crisp rule is transformed into a fuzzy rule incorporated with a confidence factor. Next, PSO is adopted to simultaneously trim the rules, search a set of good confidence factors, and fine‐tune the locations of the selected hyper‐rectangles to increase their effectiveness. Finally, a PFHRCNN consisted of a set of fuzzy rules can be generated to recognise the emotion state of music. The experimental result shows that the proposed system has a good performance.
Yu-Hao Chin, Yi-Zeng Hsieh, Mu-Chun Su, Shu-Fang Lee, Miao-Wen Chen, Jia-Ching Wang
IET Signal Process.6
2017 Program Guardian: screening system with a novel speaker recognition approach for smart TV
Yu-Hao Chin, Tzu-Chiang Tai, Jia-Hao Zhao, Kuang-Yao Wang, Chao-Tse Hong, Jia-Ching Wang
Multim. Tools Appl.6
2017 Spectral-temporal receptive fields and MFCC balanced feature extraction for robust speaker recognition
Jia-Ching Wang, Chien-Yao Wang, Yu-Hao Chin, Yu-Ting Liu, En-Ting Chen, Pao-Chi Chang
Multim. Tools Appl.1
2017 Speaker Identification Using Discriminative Features and Sparse Representation
abstract
Speaker identification is an important topic with relevance to various disciplines. This paper proposes a novel speaker identification system, which consists of two major components-feature extraction and sparse representation classifier (SRC). Although SRC has been utilized for many classification purposes, few studies have provided insight into the link between the commonly used speaker identification feature, i-vector, and SRC. To combine i-vector and SRC sufficiently, we use probabilistic principal component analysis and Bartlett test to extract high-quality i-vector to construct a discriminative dictionary in SRC, supporting effective speaker identification. Besides improving dictionary from the i-vector aspect, we also utilize dictionary learning to further enhance the content of the dictionary. Two learning methods are proposed-robust principal component analysis dictionary and SVD-dictionary. Furthermore, we propose constructing a noise dictionary and combine it with the original dictionary to absorb and suppress noise when implementing the sparse coding. Various coding methods are utilized and analyzed. A comparison to the methods for speaker identification reveals that the proposed method outperforms the baselines and confirms its feasibility.
Yu-Hao Chin, Jia-Ching Wang, Chien-Lin Huang, Kuang-Yao Wang, Chung-Hsien Wu 0001
IEEE Trans. Inf. Forensics Secur.2
2016 Locality-preserving K-SVD Based Joint Dictionary and Classifier Learning for Object Recognition
abstract
This paper concerns the development of locality-preserving methods for object recognition. The major purpose is consideration of both descriptor-level locality and image-level locality throughout the recognition process. Two dual-layer locality-preserving methods are developed, in which locality-constrained linear coding (LLC) is used to represent an image. In the learning phase, the discriminative locality-preserving K-SVD (DLP-KSVD) in which the label information is incorporated into the locality-preserving term is proposed. In addition to using class labels to learn a linear classifier, the label-consistent LP-KSVD (LCLP-KSVD) is proposed to enhance the discriminability of the learned dictionary. In LCLP-KSVD, the objective function includes a label-consistent term that penalizes sparse codes from different classes. For testing, additional information about the locality of query samples is obtained by treating the locality-preserving matrix as a feature. The recognition results that were obtained in experiments with the Caltech101 database indicate that the proposed method outperforms existing sparse coding based approaches.
Yuan-Shan Lee, Chien-Yao Wang, Seksan Mathulaprangsan, Jia Hao Zhao, Jia-Ching Wang
ACM Multimedia5
2016 Transportation Mode Detection on Mobile Devices Using Recurrent Nets
abstract
We present an approach to the use of Recurrent Neural Networks (RNN) for transportation mode detection (TMD) on mobile devices. The proposed model, called Control Gate-based Recurrent Neural Network (CGRNN), is an end-to-end model that works directly with raw signals from an embedded accelerometer. As mobile devices have limited computational resources, we evaluate the model in terms of accuracy, computational cost, and memory usage. Experiments on the HTC transportation mode dataset demonstrate that our proposed model not only exhibits remarkable accuracy, but also is efficient with low resource consumption.
Toan H. Vu, Le Dung, Jia-Ching Wang
ACM Multimedia3
2016 Compressive Sensing-Based Speech Enhancement
abstract
This study proposes a speech enhancement method based on compressive sensing. The main procedures involved in the proposed method are performed in the frequency domain. First, an overcomplete dictionary is constructed from the trained speech frames. The atoms of this redundant dictionary are spectrum vectors that are trained by the K-SVD algorithm to ensure the sparsity of the dictionary. For a noisy speech spectrum, formant detection and a quasi-SNR criterion are first utilized to determine whether a frequency bin in the spectrogram is reliable, and a corresponding mask is designed. The mask-extracted reliable components in a speech spectrum are regarded as partial observations and a measurement matrix is constructed. The problem can therefore be treated as a compressive sensing problem. The K atoms of a K-sparsity speech spectrum are found using an orthogonal matching pursuit algorithm. Because the K atoms form the speech signal subspace, the removal of the noise projected onto these K atoms is achieved by multiplying the noisy spectrum with the optimized gain that corresponds to each selected atom. The proposed method is experimentally compared with the baseline methods and demonstrates its superiority.
Jia-Ching Wang, Yuan-Shan Lee, Chang-Hong Lin, Shu-Fan Wang, Chih-Hao Shih, Chung-Hsien Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Genre based emotion annotation for music in noisy environment
abstract
The music listened by human is sometimes exposed to noise. For example, background noise usually exists when listening to music in broadcasts or lives. The noise will worsen the performance in various music emotion recognition systems. To solve the problem, this work constructs a robust system for music emotion classification in a noisy environment. Furthermore, the genre is considered when determining the emotional label for the song. The proposed system consists of three major parts, i.e. subspace based noise suppression, genre index computation, and support vector machine (SVM). Firstly, the system uses noise suppression to remove the noise content in the signal. After that, acoustical features are extracted from each music clip. Next, a dictionary is constructed by using songs that cover a wide range of genres, and it is adopted to implement sparse coding. Via sparse coding, data can be transformed to sparse coefficient vectors, and this paper computes genre indexes for the music genres based on the sparse coefficient vector. The genre indexes are regarded as combination weights in the latter phase. At the training stage of the SVM, this paper train emotional models for each genre. At the prediction stage, the predictions that obtained by emotional models in each genre are weighted combined across all genres using the genre indexes. Finally, the proposed system annotates multiple emotional labels for a song based on the combined prediction. The experimental result shows that the system can achieve a good performance in both normal and noisy environments.
Yu-Hao Chin, Po-Chuan Lin, Tzu-Chiang Tai, Jia-Ching Wang
ACII4
2015 Kernel Sparse Representation Classifier with Center Enhanced SPM for Vehicle Classification
abstract
In this paper, we proposes a visual-based vehicle classification system, in which it involves visual feature representation and classification step. In the feature representation step, we present a center enhanced spatial pyramid matching (CE-SPM) to extract the feature from images. In this work, we defined additional region in the center of each images to calculate the histograms of visual words and then pool them together with some weights to construct the feature representation vector of an image. In the classification step, kernel sparse representation classifier is used to address the problem of visual-based vehicle classification. The kernel function maps the features from original space into higher space dimension. The modified active-set algorithm for l1 non-negative least square problem is adopted to solve the optimization problem. The experimental results show the improvement of proposed method over the original SPM. The proposed method can achieve the performance of 93.7% using particular vehicle image dataset.
Andri Santoso, Chien-Yao Wang, Tzu-Chiang Tai, Jia-Ching Wang
COMPSAC4
2015 Hierarchical Dirichlet Process Mixture Model for Music Emotion Recognition
abstract
This study proposes a novel multi-label music emotion recognition (MER) system. An emotion cannot be defined clearly in the real world because the classes of emotions are usually considered overlapping. Accordingly, this study proposes an MER system that is based on hierarchical Dirichlet process mixture model (HPDMM), whose components can be shared between models of each emotion. Moreover, the HDPMM is improved by adding a discriminant factor to the proposed system based on the concept of linear discriminant analysis. The proposed system represents an emotion using weighting coefficients that are related to a global set of components. Moreover, three methods are proposed to compute the weighting coefficients of testing data, and the weighting coefficients are used to determine whether or not the testing data contain certain emotional content. In the tasks of music emotion annotation and retrieval, experimental results show that the proposed MER system outperforms state-of-the-art systems in terms of F-score and mean average precision.
Jia-Ching Wang, Yuan-Shan Lee, Yu-Hao Chin, Ying-Ren Chen, Wen-Chi Hsieh
IEEE Trans. Affect. Comput.1
2015 Speaker Identification With Whispered Speech for the Access Control System
abstract
This work presents an access control system, which is a speaker identification system based on whispered speech. Speaker identification is a main function of an access control system. Hence, a novel speaker identification system using instantaneous frequencies is proposed. The input speech signals pass through both signal independent and signal dependent filters firstly. Then, we derive the signal's instantaneous frequencies by applying the Hilbert transform. The analyzed instantaneous frequencies are proceeded to be modeled as probability density models. We use these probability density models as the feature in the proposed speaker identification system. In this work, we compare the use of parametric and nonparametric probability density estimation for instantaneous frequency modeling. Furthermore, we propose an approximated probability product kernel support vector machine (APPKSVM). In the APPKSVM, Riemann sum is applied in approximating the probability product kernel. The whisper sounds from the CHAIN speech corpus were used in the experiments. Results of the experiments show the superiority of the proposed speaker identification system.
Jia-Ching Wang, Yu-Hao Chin, Wen-Chi Hsieh, Chang-Hong Lin, Ying-Ren Chen, Ernestasia Siahaan
IEEE Trans Autom. Sci. Eng.1
2015 Robust Environmental Sound Recognition With Fast Noise Suppression for Home Automation
abstract
This paper proposes a robust environmental sound recognition system using a fast noise suppression approach for home automation applications. The system comprises a fast subspace-based noise suppression module and a sound classification module. For the noise suppression module, we propose a noise suppression method that applies fast subspace approximations in the wavelet domain. We show that this method offers a lower computational cost than conventional methods. In the sound classification module, we use a feature extraction method that is also based on the wavelet subspace, derived from seventeen critical bands in a signal's wavelet packet transform. Furthermore, we create a multiclass support vector machine by employing probability product kernels. The experimental results for ten classes of various environmental sounds show that the proposed system offers robust performance in environmental sound recognition tasks.
Jia-Ching Wang, Yuan-Shan Lee, Chang-Hong Lin, Ernestasia Siahaan, Chung-Hsien Yang
IEEE Trans Autom. Sci. Eng.1
2015 Speech Emotion Verification Using Emotion Variance Modeling and Discriminant Scale-Frequency Maps
abstract
This paper develops an approach to speech-based emotion verification based on emotion variance modeling and discriminant scale-frequency maps. The proposed system consists of two parts-feature extraction and emotion verification. In the first part, for each sound frame, important atoms from the Gabor dictionary are selected by using the matching pursuit algorithm. The scale, frequency, and magnitude of the atoms are extracted to construct a nonuniform scale-frequency map, which supports auditory discriminability by the analysis of critical bands. Next, sparse representation is used to transform scale-frequency maps into sparse coefficients to enhance the robustness against emotion variance and achieve error-tolerance improvement. In the second part, emotion verification, two scores are calculated. A novel sparse representation verification approach based on Gaussian-modeled residual errors is proposed to generate the first score from the sparse coefficients. Such a classifier can minimize emotion variance and improve recognition accuracy. The second score is calculated by using the emotional agreement index (EAI) from the same coefficients. These two scores are combined to obtain the final detection result. Experiments on an emotional database of spoken speech were conducted and indicate that the proposed approach can achieve an average equal error rate (EER) of as low as 6.61%. A comparison among different approaches reveals that the proposed method is superior to the others and confirms its feasibility.
Jia-Ching Wang, Yu-Hao Chin, Bo-Wei Chen, Chang-Hong Lin, Chung-Hsien Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 VLSI Design for SVM-Based Speaker Verification System
abstract
This brief presents the chip implementation of a support vector machine (SVM)-based speaker verification system. The proposed chip comprises a speaker feature extraction (SFE) module, an SVM module, and a decision module. The SFE module performs autocorrelation analysis, linear predictive coefficient (LPC) extraction, and LPC-to-cepstrum conversion. The SVM module includes a Gaussian kernel unit and a scaling unit. The purpose of the Gaussian kernel unit is first to evaluate the kernel value of a test vector and a support vector. Four Gaussian kernel processing elements (GK-PEs) are designed to process four support vectors simultaneously. Each GK-PE is designed in the pipeline fashion and is capable of performing 2-norm and exponential operations. An enhanced CORDIC architecture is proposed to calculate the exponential value. As well as the Gaussian kernel unit, a scaling unit is also developed for use in the SVM module. The scaling unit is used to perform scaling multiplications and the remaining operations of SVM decision value evaluation. Finally, the decision module accumulates the frame scores that are generated by all of the test frames, and then compare it with a threshold to see if the test utterance is spoken by the claimed speaker. This designed chip is characterized by its high speed and its ability to handle a large number of support vectors in the SVM. The prototype chip is a semicustom chip that is fabricated using Taiwan Semiconductor Manufacturing Company 0.90-nm CMOS technology on a die with a size of ~7.9 × 7.9 mm2.
Jia-Ching Wang, Li-Xun Lian, Yan-Yu Lin, Jia Hao Zhao
IEEE Trans. Very Large Scale Integr. Syst.1
2014 Gabor-Based Nonuniform Scale-Frequency Map for Environmental Sound Classification in Home Automation
abstract
This work presents a novel feature extraction approach called nonuniform scale-frequency map for environmental sound classification in home automation. For each audio frame, important atoms from the Gabor dictionary are selected by using the Matching Pursuit algorithm. After the system disregards phase and position information, the scale and frequency of the atoms are extracted to construct a scale-frequency map. Principal Component Analysis (PCA) and Linear Discriminate Analysis (LDA) are then applied to the scale-frequency map, subsequently generating the proposed feature. During the classification phase, a segment-level multiclass Support Vector Machine (SVM) is operated. Experiments on a 17-class sound database indicate that the proposed approach can achieve an accuracy rate of 86.21%. Furthermore, a comparison reveals that the proposed approach is superior to the other time-frequency methods.
Jia-Ching Wang, Chang-Hong Lin, Bo-Wei Chen, Min-Kang Tsai
IEEE Trans Autom. Sci. Eng.1
2014 Mixed Sound Event Verification on Wireless Sensor Network for Home Automation
abstract
In this paper, we present the problem of mixed sound event verification in a wireless sensor network for home automation systems. In home automation systems, the sound recognized by the system becomes the basis for performing certain tasks. However, if a target source is mixed with another sound due to simultaneous occurrence, the system would generate poor recognition results, subsequently leading to inappropriate responses. To handle such problems, this study proposes a framework, which consists of sound separation and sound verification techniques based on a wireless sensor network (WSN), to realize sound-triggered automation. In the sound separation phase, we present a convolutive blind source separation system with source number estimation using time-frequency clustering. An accurate mixing matrix can be estimated by the proposed phase compensation technique and used for reconstructing the separated sound sources. In the verification phase, Mel frequency cepstral coefficients and Fisher scores that are derived from the wavelet packet decomposition of signals are used as features for support vector machines. Finally, a sound of interest can be selected for triggering automated services according to the verification result. The experimental results demonstrate the robustness and feasibility of the proposed system for mixed sound verification in WSN-based home environments.
Jia-Ching Wang, Chang-Hong Lin, Ernestasia Siahaan, Bo-Wei Chen, Hsiang-Lung Chuang
IEEE Trans. Ind. Informatics1
2012 A new hybrid and dynamic fusion of multiple experts for intelligent porch system
Ta-Wen Kuan, Hsin-Chun Tsai, Jhing-Fa Wang, Jia-Ching Wang, Bo-Wei Chen, Zong-You Lin
Expert Syst. Appl.4
2012 Fast Mode Decision for H.264/AVC Based on Rate-Distortion Clustering
abstract
Although H.264/AVC is a promising video coding standard that achieves excellent coding performance in terms of visual quality and bitrate, its extremely high encoding complexity raises concerns about its computational burden on real-time applications. This work presents a multi-phase classification (MPC) scheme that builds a mode decision tree according to the clustering of rate-distortion costs. A nearest cluster mean criterion is used to examine candidate modes phase by phase, and a performance control mechanism is incorporated to maintain coding performance. Experimental results confirm that the proposed MPC algorithm reduces encoding time by an average of 68% with only negligible performance degradation.
Yu-Huan Sung, Jia-Ching Wang
IEEE Trans. Multim.2
2012 VLSI Design of an SVM Learning Core on Sequential Minimal Optimization Algorithm
abstract
The sequential minimal optimization (SMO) algorithm has been extensively employed to train the support vector machine (SVM). This work presents an efficient application specific integrated circuit chip design for sequential minimal optimization. This chip is implemented as an intellectual property core, suitable for use in an SVM-based recognition system on a chip. The proposed SMO chip was tested and found to be fully functional, using a prototype system based on the Altera DE2 board with a Cyclone II 2C70 field-programmable gate array.
Ta-Wen Kuan, Jhing-Fa Wang, Jia-Ching Wang, Po-Chuan Lin, Gaung-Hui Gu
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Hardware/software co-design for fast-trainable speaker identification system based on SMO
abstract
Embedded speaker identification system is a popular research, but most of current systems can not provide fast training ability. Because of the low computational ability in the embedded environment, a large amount of waiting time usually makes the human-machine interface not friendly. This paper presents a hardware and software (HW/SW) co-design solution for fast-trainable speaker identification system. Fast training ability makes this embedded speaker identification system possess high flexibility and enhances the convenience to a wide range of real-world applications. The proposed system consists of a training phase and a multiclass identification phase. The sequential minimal optimization (SMO) training algorithm occupies the heaviest computational load and is realized as a dedicated VLSI module, i.e., the hardware component. The other processes such as speech preprocess, speech feature extraction, and SVM voting strategy are implemented by software. Moreover, a data-packed mechanism is presented to improve the bandwidth utilization. Compared with the embedded C code based on ARM processor, our system reduces 90% of the training time and achieves 89.9% identification rate with the NIST 2010 speaker recognition database. The proposed system was tested and found to be fully functional working on a Socle CDK prototype system with an AMBA based Xilinx FPGA and an ARM926EJ processor.
Jhing-Fa Wang, Jr-Shiang Peng, Jia-Ching Wang, Po-Chuan Lin, Ta-Wen Kuan
SMC3
2010 Dynamic Fixed-Point Arithmetic Design of Embedded SVM-Based Speaker Identification System
Jhing-Fa Wang, Ta-Wen Kuan, Jia-Ching Wang, Ta-Wei Sun
ISNN (2)3
2009 Video Knowledge Augmentation based on Summarized Contents and Online Media
abstract
Exploration techniques of video knowledge have been proposed for years to help people discover the details about videos. However, existing systems still yield limited information for users. In this paper, we present a video knowledge browsing system, which can establish the framework of a video based on its summarized contents and expand them by using online correlated media. Thus, users can not only browse key points of a video efficiently but also focus on what they are interested in. In order to construct the fundamental system, we make use of our previous proposed approaches to transforming a video into a graph. After the relational graph is built up, the social network analysis is then performed to explore online relevant resources. We also apply the Markov clustering algorithm to enhance the results of the network analysis. The experiments demonstrate that our system can achieve better performance than the traditional systems.
Bo-Wei Chen, Jhing-Fa Wang, Jia-Ching Wang
ISCAS3
2009 VLSI Design of Sequential Minimal Optimization Algorithm for SVM Learning
abstract
The sequential minimal optimization (SMO) algorithm has been widely used for training the support vector machine (SVM). In this paper, we present the first chip design for sequential minimal optimization. This chip is implemented as an intellectual property (IP) core, suitable to be utilized in an SVM-based recognition system on a chip. The proposed SMO chip has been tested to be fully functional, using a prototype system based on the Altera DE2 board with Cyclone II 2C70 FPGA (field-programmable gate array).
Ta-Wen Kuan, Jhing-Fa Wang, Jia-Ching Wang, Gaung-Hui Gu
ISCAS3
2009 A Novel Video Summarization Based on Mining the Story-Structure and Semantic Relations Among Concept Entities
abstract
Video summarization techniques have been proposed for years to offer people comprehensive understanding of the whole story in the video. Roughly speaking, existing approaches can be classified into the two types: one is static storyboard, and the other is dynamic skimming. However, despite that these traditional methods give brief summaries for users, they still do not provide with a concept-organized and systematic view. In this paper, we present a structural video content browsing system and a novel summarization method by utilizing the four kinds of entities: who, what, where, and when to establish the framework of the video contents. With the assistance of the above-mentioned indexed information, the structure of the story can be built up according to the characters, the things, the places, and the time. Therefore, users can not only browse the video efficiently but also focus on what they are interested in via the browsing interface. In order to construct the fundamental system, we employ maximum entropy criterion to integrate visual and text features extracted from video frames and speech transcripts, generating high-level concept entities. A novel concept expansion method is introduced to explore the associations among these entities. After constructing the relational graph, we exploit graph entropy model to detect meaningful shots and relations, which serve as the indices for users. The results demonstrate that our system can achieve better performance and information coverage.
Bo-Wei Chen, Jia-Ching Wang, Jhing-Fa Wang
IEEE Trans. Multim.2
2008 Robust Environmental Sound Recognition for Home Automation
abstract
This work presents a robust environmental sound recognition system for home automation. Specific home automation services can be activated based on identified sound classes. Additionally, when the sound category is human speech, such speech can be recognized for detecting human intentions as in conventional research on home automation. To attain this ambitious goal, this study uses two key techniques: signal-to-noise ratio-aware subspace-based signal enhancement and sound recognition with independent component analysis mel-frequency cepstral coefficients and a frame-based multiclass support vector machines, respectively. Simulations and an experiment in a real-world environment are given to illustrate the performance of the proposed robust sound recognition system.
Jia-Ching Wang, Hsiao Ping Lee, Jhing-Fa Wang, Cai-Bei Lin
IEEE Trans Autom. Sci. Eng.1
2008 Intensity Gradient Technique for Efficient Intra-Prediction in H.264/AVC
abstract
This study presents an intensity gradient approach for intra-prediction in H.264 encoding system, which enhances the performance and efficiency of previous fast algorithms. We propose a preprocessing stage in which eight orientation features are extracted from a macro block by the intensity gradient filters. The orientation features are utilized to select a subset of prediction modes to be involved in the rate-distortion calculation so that the encoding time can be reduced. The simulation results indicate that the intensity gradient based algorithm for intra-prediction contributes better tradeoff between rate-distorion performance and encoding complexity than the previous algorithms. Compared to H.264 reference software, the proposed algorithm introduces slight PSNR degradation and bit rate increase but saves around 76% of the total encoding time with all intra-frame coding.
An-Chao Tsai, Anand Paul 0001, Jia-Ching Wang, Jhing-Fa Wang
IEEE Trans. Circuits Syst. Video Technol.3
2007 Efficient Intra Prediction in H.264 Based on Intensity Gradient Approach
abstract
This study presents an intensity gradient approach to intra prediction in H.264 encoding system, which enhances the performance and efficiency by means of edge orientation of the gradient filter. We propose a pre-processing stage in which eight-orientation feature are extracted from a macro block that selects four modes to be applied to the block among a set of predefined modes. It is shown that by choosing number of modes used in rate-distortion calculation lead to significant enhancement in the performance of intra prediction. The validity of this algorithm is confirmed experimentally. The simulation results indicate that the intensity gradient-based algorithm for intra prediction contributes to the bit-rate reduction compared to that of previous algorithms and saves around 53% of the total encoding time in H.264.
An-Chao Tsai, Anand Paul 0001, Jia-Ching Wang, Jhing-Fa Wang
ISCAS3
2007 Event-Based Segmentation of Sports Video Using Motion Entropy
abstract
An event-based segmentation method for sports videos is presented. A motion entropy criterion is employed to characterize the level of intensity of relevant object motion in individual frames of a video sequence. The resulting motion entropy curve then is approximated with a piece-wise linear model using a homoscedastic error model based time series change point detection algorithm. It is observed that interesting sports events are correlated with specific patterns of the piece-wise linear model. A set of empirically derived classification rules then is derived based on these observations. Application of these rules to the motion entropy curve leads to this motion entropy curve, one is able to segment the corresponding video sequence into individual sections, each consisting of a semantically relevant event. The proposed method is tested on six hours of sports videos including basketball, soccer and tennis. Excellent experimental results are observed.
Chen-Yu Chen, Jia-Ching Wang, Jhing-Fa Wang, Yu Hen Hu
ISM2
2007 Unsupervised Speaker Change Detection Using SVM Training Misclassification Rate
abstract
This work presents an unsupervised speaker change detection algorithm based on support vector machines (SVM) to detect speaker change (SC) in a speech stream. The proposed algorithm is called the SVM training misclassification rate (STMR). The STMR can identify SCs with less speech data collection, making it capable of detecting speaker segments with short duration. According to experiments on the NIST Rich Transcription 2005 Spring Evaluation (RT-05S) corpus, the STMR has a missed detection rate of only 19.67 percent.
Po-Chuan Lin, Jia-Ching Wang, Jhing-Fa Wang, Hao-Ching Sung
IEEE Trans. Computers2
2007 A Fast Mode Decision Algorithm and Its VLSI Design for H.264/AVC Intra-Prediction
abstract
In this paper, we present a fast mode decision algorithm and design its VLSI architecture for H.264 intra-prediction. A regular spatial domain filtering technique is proposed to compute the dominant edge strength (DES) to reduce the possible predictive modes. Experimental results revealed that the proposed fast intra-algorithm reduces 40% computation with slight peak signal-to-noise ratio (PSNR) degradation. The designed DES VLSI engine comprises a zigzag converter, a DES finite-state machine (FSM), and a DES core. The former two units handle memory allocation and control flow while the last performs pseudoblock computation, edge filtering, and dominant edge strength extraction. With semicustom design fabricated by 0.18 mum CMOS single-poly-six-metal technology, the realized die size is roughly 0.15 times 0.15 mm2and can be operated at 66 MHz.
Jia-Ching Wang, Jhing-Fa Wang, Jar-Ferr Yang, Jang-Ting Chen
IEEE Trans. Circuits Syst. Video Technol.1
2006 An ARM-Based Embedded System Design for Speech-to-Speech Translation
Shun-Chieh Lin, Jhing-Fa Wang, Jia-Ching Wang, Hsueh-Wei Yang
EUC3
2006 Robust Speaker Recognition using SNR-Aware Subspace-Based Enhancement and Probabilistic SVMs
abstract
In this paper, we present a robust text-independent speaker recognition system. The proposed system mainly includes an SNR-aware subspace-based enhancement technique and probabilistic support vector machines (SVMs). First, we construct a perceptual filterbank from psycho-acoustic model and incorporate it with the subspace-based enhancement approach. The prior SNR of each subband within the perceptual filterbank is taken to decide the estimator's gain to effectively suppress environmental background noises. Next, this study uses probabilistic SVMs to identify the speaker from the enhanced speech. The superiority of the proposed system has been demonstrated by twenty speaker recognition from AURORA-2 database with in-car noises
Jia-Ching Wang, Jhing-Fa Wang, Wai-He Kuok, Hsiao Ping Lee, Chung-Hsien Yang
ICME1
2006 Environmental Sound Classification using Hybrid SVM/KNN Classifier and MPEG-7 Audio Low-Level Descriptor
abstract
In this paper, we present a new environmental sound classification architecture. The proposed sound classifier is performed in frame level and fuses the support vector machine (SVM) and the k nearest neighbor rule (KNN). In feature selection, three MPEG-7 audio low-level descriptors, spectrum centroid, spectrum spread, and spectrum flatness are used as the sound features to exploit their ability in sound classification. Experiments carried out on 12-class sound database can achieve an 85.1% accuracy rate. The The The performance comparison between the HMM sound classifier using audio spectrum projection features demonstrates the superiority of the proposed scheme.
Jia-Ching Wang, Jhing-Fa Wang, Wai-He Kuok, Cheng-Shu Hsu
IJCNN1
2006 A novel fast algorithm for intra mode decision in H.264/AVC encoders
abstract
This paper presents a fast mode decision algorithm for H.264 intra prediction based on dominant edge strength (DES). In H.264 intra prediction, the computation-extensive rate distortion optimization (RDO) technique with full intra mode search is used to select the best mode for each macroblock. To reduce the computational load in mode decision, the DES which is corresponding to a decision mode is detected first. In accordance with the detected dominant edge, a subset of the prediction modes is then chosen for RDO calculation. The proposed algorithm only searches 4 modes instead of 9 for the 4/spl times/4 luma blocks. As for the 16/spl times/16 or 8/spl times/8 chroma blocks, instead of 4 modes, only 2 modes are required to be searched. Experimental results revealed that the computation time of the proposed fast intra prediction algorithm is averagely reduced to 40% of the full search method with slight PSNR degradation.
Jhing-Fa Wang, Jia-Ching Wang, Jang-Ting Chen, An-Chao Tsai, Anand Paul 0001
ISCAS2
2006 Efficient news video querying and browsing based on distributed news video servers
abstract
This study presents an efficient news video querying and browsing system based on distributed news video servers. The proposed architecture includes distributed news video preprocessing (NVP) server and visualized querying/browsing (VQB) server. The distributed NVP server receives the news video story from news video web (NVW) server and generates the story abstract, i.e., key frames and key sentences. These story abstract are then combined with the news script of a news story, then sent to the VQB server for news video querying and browsing. The VQB server performs semantic clustering using news script to categorize all the news video stories, it also provides a visualized interface displaying the story abstract so that users can fast grasp the main idea of a news story. The superiority of the proposed system has been demonstrated using news video obtained from NVW servers of EraNews, ETtoday and Formosa TV stations in Taiwan.
Chen-Yu Chen, Jia-Ching Wang, Jhing-Fa Wang
IEEE Trans. Multim.2
2002 Chip design of MFCC extraction for speech recognition
Jia-Ching Wang, Jhing-Fa Wang, Yu-Sheng Weng
Integr.1
2002 Chip design of portable speech memopad suitable for persons with visual disabilities
abstract
This paper presents the design of a speech recognition and compression chip for portable memopad devices, especially suitable for use by the visually impaired. The proposed chip design is based on several cores of which they can be regarded as intellectual property (IP) cores to be used for a variety of speech-related application systems. A cepstrum extraction core and a dynamic warping core are designed for mapping the speech recognition algorithms. In the cepstrum extraction core, a novel architecture computes the autocorrelation between the overlapping frames using two pairs of shift registers and an intelligent accumulation procedure. The architecture of the dynamic time warping core uses only a single processing element, and is based on our extensive study of the relationship among the nodes in the dynamic time warping lattice. Bit rate is the key factor affecting the memory size for speech compression; therefore, a very low bit-rate speech coder is used. The speech coder exploits a line-spectrum-based interpolation method, which yields fine quality synthesized speech despite the low 1.6 kbps bit rate. The 1.6 kbps vocoder core is cost-effective, and it integrates both encoder and decoder algorithms. The proposed design has been tested via hardware simulations on Xilinx Virtex series FPGAs and a semi-custom chip fabricated by 0.35 /spl mu/m CMOS single-poly-four-metal technology on a die size approximately 4.46/spl times/4.46 mm/sup 2/.
Jhing-Fa Wang, Jia-Ching Wang, Han-Chiang Chen, Tai-Lung Chen, Chin-Chan Chang, Ming-Chi Shih
IEEE Trans. Speech Audio Process.2
2001 A voicing-driven packet loss recovery algorithm for analysis-by-synthesis predictive speech coders over Internet
abstract
In this paper, a novel voice-driven adaptive packet loss recovery algorithm is proposed to lessen the possible voice degradation and error propagation for analysis-by-synthesis speech coders in Internet applications. After voicing classification, we adaptively adopt random noise generation, multiresolution excitation generation, or pulse tracking procedure to recover the lost packets, By applying the algorithm to the G.723.1 coder, simulation results show that the proposed algorithm is superior to the recovery algorithm embedded in the G.723.1 standard through the subjective evaluation.
Jhing-Fa Wang, Jia-Ching Wang, Jar-Ferr Yang, Jian-Jia Wang
IEEE Trans. Multim.2
2000 Chip design of mel frequency cepstral coefficients for speech recognition
abstract
The mel frequency cepstral coefficients (MFCC) is one of the mast important features, which is required among various kinds of speech applications. The chip for speech features extraction based on the MFCC algorithm is first proposed. The chip is designed with area efficient consideration and can achieve the following: (1) the reduction of table size and multiplication complexity by means of the symmetric property of the cosine function, (2) the decrease of the multiplication load and required constant memory in the calculation of the weighted energy spectrum by applying the mapping relationship between the mel scale and the frequency scale, (3) the minimization of the look-up table size for logarithm operations by modifying the partitioned table look-up method. The chip is fabricated with 0.6 /spl mu/m double-metal CMOS technology. It contains approximately 10,000 gates occupying 3.2/spl times/3.3 mm/sup 2/ area and the maximum clock rate is 50 MHz.
Jia-Ching Wang, Jhing-Fa Wang, Yu-Sheng Weng
ICASSP1
2000 Single chip implementation of the 1.6 kbps speech vocoder
abstract
In this paper, we propose a low bit rate speech vocoder and its corresponding VLSI implementation. The vocoder exploits the interpolation property so that the fine quality in synthesized speech is obtained even though the bit rate is as low as 1.6 kbps. Two novel methods including pitch detection and LSP decoding which are suitable for VLSI implementation are also proposed. The heuristic pitch detection algorithm avoids the heavy computational load introduced by the traditional normalized autocorrelation method. The memory storing triangular function value is no longer needed after adopting the new LSP decoding process. The chip is designed with area effective feature and is suitable for stand alone application.
Jia-Ching Wang, Jhing-Fa Wang, Han-Chiang Chen
ISCAS1