Chia-Ping Chen

dblp:01/3763 · DBLP profile ↗
← Back
49ranked-venue papers
15as first author
10since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 12 first-author · 8 since 2021Artificial intelligence and machine learning · 27 · 7 first-author · 8 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Enhancing ECAPA-TDNN with Feature Processing Module and Attention Mechanism for Speaker Verification
Shiu-Hsiang Liou, Po-Cheng Chan, Chia-Ping Chen, Tzu-Chieh Lin, Chung-Li Lu, Yu-Han Cheng, Hsiang-Feng Chuang, Wei-Yu Chen
INTERSPEECH3
2024 Bilingual and Code-switching TTS Enhanced with Denoising Diffusion Model and GAN
Huai-Zhe Yang, Chia-Ping Chen, Shan-Yun He, Cheng-Ruei Li
INTERSPEECH2
2024 ReF-LDM: A Latent Diffusion Model for Reference-based Face Image Restoration
abstract
While recent works on blind face image restoration have successfully produced impressive high-quality (HQ) images with abundant details from low-quality (LQ) input images, the generated content may not accurately reflect the real appearance of a person. To address this problem, incorporating well-shot personal images as additional reference inputs may be a promising strategy. Inspired by the recent success of the Latent Diffusion Model (LDM) in image generation, we propose ReF-LDM—an adaptation of LDM designed to generate HQ face images conditioned on one LQ image and multiple HQ reference images. Our LDM-based model incorporates an effective and efficient mechanism, CacheKV, for conditioning on reference images. Additionally, we design a timestep-scaled identity loss, enabling LDM to focus on learning the discriminating features of human faces. Lastly, we construct FFHQ-ref, a dataset consisting of 20,406 high-quality (HQ) face images with corresponding reference images, which can serve as both training and evaluation data for reference-based face restoration models.
Chi-Wei Hsiao, Yu-Lun Liu 0001, Cheng-Kun Yang, Sheng-Po Kuo, Kevin Jou, Chia-Ping Chen
NeurIPS6
2023 Personalized Lightweight Text-to-Speech: Voice Cloning with Adaptive Structured Pruning
abstract
Personalized TTS is an exciting and highly desired application that allows users to train their TTS voice using only a few recordings. However, TTS training typically requires many hours of recording and a large model, making it unsuitable for deployment on mobile devices. To overcome this limitation, related works typically require fine-tuning a pre-trained TTS model to preserve its ability to generate high-quality audio samples while adapting to the target speaker’s voice. This process is commonly referred to as "voice cloning." Although related works have achieved significant success in changing the TTS model’s voice, they are still required to fine-tune from a large pre-trained model, resulting in a significant size for the voice-cloned model. In this paper, we propose applying trainable structured pruning to voice cloning. By training the structured pruning masks with voice-cloning data, we can produce a unique pruned model for each target speaker. Our experiments demonstrate that using learnable structured pruning, we can compress the model size to 7 times smaller while achieving comparable voice-cloning performance.
Sung-Feng Huang, Chia-Ping Chen, Zhi-Sheng Chen, Yu-Pao Tsai, Hung-yi Lee
ICASSP2
2022 Denoising Likelihood Score Matching for Conditional Score-based Data Generation
Chen-Hao Chao, Wei-Fang Sun, Bo-Wun Cheng, Yi-Chen Lo, Chia-Che Chang, Yu-Lun Liu 0001, Yu-Lin Chang, Chia-Ping Chen, Chun-Yi Lee
ICLR8
2022 On the Efficiency of Integrating Self-Supervised Learning and Meta-Learning for User-Defined Few-Shot Keyword Spotting
abstract
User-defined keyword spotting is a task to detect new spoken terms defined by users. This can be viewed as a few-shot learning problem since it is unreasonable for users to define their desired keywords by providing many examples. To solve this problem, previous works try to incorporate self-supervised learning models or apply meta-learning algorithms. But it is unclear whether self-supervised learning and meta-learning are complementary and which combination of the two types of approaches is most effective for few-shot keyword discovery. In this work, we systematically study these questions by utilizing various self-supervised learning models and combining them with a wide variety of meta-learning algorithms. Our result shows that HuBERT combined with Matching network achieves the best result and is robust to the changes of few-shot examples.
Wei-Tsung Kao, Yuan-Kuei Wu, Chia-Ping Chen, Zhi-Sheng Chen, Yu-Pao Tsai, Hung-yi Lee
SLT3
2021 CLCC: Contrastive Learning for Color Constancy
abstract
In this paper, we present CLCC, a novel contrastive learning framework for color constancy. Contrastive learning has been applied for learning high-quality visual representations for image classification. One key aspect to yield useful representations for image classification is to design illuminant invariant augmentations. However, the illuminant invariant assumption conflicts with the nature of the color constancy task, which aims to estimate the illuminant given a raw image. Therefore, we construct effective contrastive pairs for learning better illuminant-dependent features via a novel raw-domain color augmentation. On the NUS-8 dataset, our method provides 17.5% relative improvements over a strong baseline, reaching state-of-the-art performance without increasing model complexity. Furthermore, our method achieves competitive performance on the Gehler dataset with 3× fewer parameters compared to top-ranking deep learning methods. More importantly, we show that our model is more robust to different scenes under close proximity of illuminants, significantly reducing 28.7% worst-case error in data-sparse regions. Our code is available at https://github.com/howardyclo/clcc-cvpr21.
Yi-Chen Lo, Chia-Che Chang, Hsuan-Chao Chiu, Chia-Ping Chen, Yu-Lin Chang, Kevin Jou
CVPR5
2021 Bridging Unsupervised and Supervised Depth from Focus via All-in-Focus Supervision
abstract
Depth estimation is a long-lasting yet important task in computer vision. Most of the previous works try to estimate depth from input images and assume images are all-in-focus (AiF), which is less common in real-world applications. On the other hand, a few works take defocus blur into account and consider it as another cue for depth estimation. In this paper, we propose a method to estimate not only a depth map but an AiF image from a set of images with different focus positions (known as a focal stack). We design a shared architecture to exploit the relationship between depth and AiF estimation. As a result, the proposed method can be trained either supervisedly with ground truth depth, or unsupervisedly with AiF images as supervisory signals. We show in various experiments that our method outperforms the state-of-the-art methods both quantitatively and qualitatively, and also has higher efficiency in inference time.
Ning-Hsu Wang, Ren Wang 0014, Yu-Lun Liu 0001, Yu-Lin Chang, Chia-Ping Chen, Kevin Jou
ICCV6
2021 Systems for Low-Resource Speech Recognition Tasks in Open Automatic Speech Recognition and Formosa Speech Recognition Challenges
Hung-Pang Lin, Yu-Jia Zhang, Chia-Ping Chen
Interspeech3
2021 Improving Time Delay Neural Network Based Speaker Recognition with Convolutional Block and Feature Aggregation Methods
Yu-Jia Zhang, Yih-Wen Wang, Chia-Ping Chen, Chung-Li Lu, Bo-Cheng Chan
Interspeech3
2020 Learning Camera-Aware Noise Models
Ke-Chi Chang, Ren Wang 0014, Hung-Jin Lin, Yu-Lun Liu 0001, Chia-Ping Chen, Yu-Lin Chang, Hwann-Tzong Chen
ECCV (24)5
2020 Explorable Tone Mapping Operators
abstract
Tone-mapping plays an essential role in high dynamic range (HDR) imaging. It aims to preserve visual information of HDR images in a medium with a limited dynamic range. Although many works have been proposed to provide tone-mapped results from HDR images, most of them can only perform tone-mapping in a single pre-designed way. However, the subjectivity of tone-mapping quality varies from person to person, and the preference of tone-mapping style also differs from application to application. In this paper, a learning-based multimodal tone-mapping method is proposed, which not only achieves excellent visual quality but also explores the style diversity. Based on the framework of BicycleGAN [1], the proposed method can provide a variety of expert-level tone-mapped results by manipulating different latent codes. Finally, we show that the proposed method performs favorably against state-of-the-art tone-mapping algorithms both quantitatively and qualitatively.
Chien-Chuan Su, Ren Wang 0014, Hung-Jin Lin, Yu-Lun Liu 0001, Chia-Ping Chen, Yu-Lin Chang, Soo-Chang Pei
ICPR5
2019 Speaker Characterization Using TDNN-LSTM Based Speaker Embedding
abstract
In this paper we propose speaker characterization using time delay neural networks and long short-term memory neural networks (TDNN-LSTM) speaker embedding. Three types of front-end feature extraction are investigated to find good features for speaker embedding. Three kinds of data augmentation are used to increase the amount and diversity of the training data. The proposed methods are evaluated with the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) tasks. Experimental results show that the proposed methods achieve a decision cost of 0.400 with the pooled SRE 2018 development set with a single system. In addition, by applying simple average score combination on the outputs of 12 systems, the proposed methods achieve an equal error rate (EER) of 5.56% and a minimum decision cost function of 0.423 with the SRE 2016 evaluation set.
Chia-Ping Chen, Su-Yu Zhang, Chih-Ting Yeh, Jia-Ching Wang, Tenghui Wang, Chien-Lin Huang
ICASSP1
2019 Transfer-Representation Learning for Detecting Spoofing Attacks with Converted and Synthesized Speech in Automatic Speaker Verification System
Su-Yu Chang, Kai-Cheng Wu, Chia-Ping Chen
INTERSPEECH3
2019 AI Deep Learning with Multiple Labels for Sentiment Classification of Tweets
abstract
We introduce an incremental transfer learning pipeline for AI systems for ordinal classification based on multiple labels of Tweets. In this pipeline, 5 sub-models are trained and the incremental knowledge is transferred to achieve higher model complexity and better performance. Each training example has multiple labels, with a target label for each sub-model. The first sub-model is a polarity classification model for negative, neutral, and positive sentiments. The second and third sub-models are ordinal classification models for positive and negative sentiments. The fourth sub-model is a binary classification model for neutral sentiment. The last sub-model is a seven-class model for polarity and intensity classes. The proposed method is applied on the Semantic Evaluation 2018 Task 1 Affects in Tweets Subtask V-oc (ordinal classification task). We experiment with 5 word embeddings, and apply ensemble methods to combine their outputs to boost overall performance. We use weighted average and stacking technique on the proposed systems and the DeepMoji model which is retrained for transfer learning. We achieve a Pearson correlation coefficient of 0.806 on the test data of SemEval-2018, which would have ranked the 4th in the SemEval-2018 Task 1 Subtask V-oc.
Zi Yuan Gao, Chia-Ping Chen
ISCAS2
2018 Effective Attention Mechanism in Dynamic Models for Speech Emotion Recognition
abstract
We propose to integrate the attention mechanism into deep recurrent neural network models for speech emotion recognition. This is based on the intuition that it is beneficial to emphasize the expressive part of the speech signal for emotion recognition. By introducing attention mechanism, we make the system learn how to focus on the more robust or informative segments in the input signal. The proposed recognition model is evaluated on the FAU-Aibo tasks as defined in Interspeech 2009 Emotion Challenge. Our baseline deep recurrent neural network model achieves 37.0% unweighted averaged (UA) recall rate, which is on par with the official HMM baseline system for dynamic modeling framework. The proposed integration of attention mechanism on top of the baseline deep RNN model achieves 46.3% UA recall rate. As far as we know, this is the best UA recall rate ever achieved on FAU-Aibo tasks within the dynamic modeling framework.
Po-Wei Hsiao, Chia-Ping Chen
ICASSP2
2018 Combining De-noising Auto-encoder and Recurrent Neural Networks in End-to-End Automatic Speech Recognition for Noise Robustness
abstract
In this paper, we propose an end-to-end noise-robust automatic speech recognition system through deep-learning implementation of de-noising auto-encoders and recurrent neural networks. We use batch normalization and a novel design for the front-end de-noising auto-encoder, which mimics a two-stage prediction of a single-frame clean feature vector from multi-frame noisy feature vectors. For the backend word recognition, we use an end-to-end system based on bidirectional recurrent neural network with long short-term memory cells. The LSTM-BiRNN is trained via connectionist temporal classification criterion. Its performance is compared to a baseline backend based on hidden Markov models and Gaussian mixture models (HMM-GMM). Our experimental results show that the proposed novel front-end de-noising auto-encoder outperforms the best record we can find for the Aurora 2.0 clean-condition training tasks by an absolute improvement of 1.2% (6.0% vs. 7.2%). In addition, the proposed end-to-end back-end architecture is as good as the traditional HMM-GMM back-end recognizer.
Tzu-Hsuan Ting, Chia-Ping Chen
SLT2
2017 Speech emotion recognition with skew-robust neural networks
abstract
We propose a neural-network training algorithm that is robust to data imbalance in classification. In our proposed algorithm, weights are introduced to training examples, effectively modifying the trajectory traversed in the parameter space during the learning process. Furthermore, the proposed algorithm would reduce to the normal stochastic gradient decent learning if the data is balanced. On the FAU-Aibo database, which is known to be used in Interspeech Emotion Challenge, the proposed method achieves an unweighted average (UA) recall rate of 45.3% on the 5-class speech emotion recognition task. Within the static modeling framework, where each example is represented as a fixed-length vector, this performance is one of the best performance ever achieved on the 5-class task.
Po-Yuan Shih, Chia-Ping Chen, Hsin-Min Wang
ICASSP2
2017 Speech emotion recognition with ensemble learning methods
abstract
In this paper, we propose to apply ensemble learning methods on neural networks to improve the performance of speech emotion recognition tasks. The basic idea is to first divide unbalanced data set into balanced subsets and then combine the predictions of the models trained on these subsets. Several methods regarding the decomposition of data and the exploitation of model predictions are investigated in this study. On the public-domain FAU-Aibo database, which is used in Interspeech Emotion Challenge evaluation, the best performance we achieve is an unweighted average (UA) recall rate of 45.5% for the 5-class classification task. Furthermore, such performance is achieved with a feature space of 40-dimension. Compared to the baseline system with 384-dimension feature vector per example and an UA of 38.9%, such a performance is very impressive. Indeed, this is one of the best performances on FAU-Aibo within the static modeling framework.
Po-Yuan Shih, Chia-Ping Chen, Chung-Hsien Wu 0001
ICASSP2
2016 Integration of orthogonal feature detectors in parameter learning of artificial neural networks to improve robustness and the evaluation on hand-written digit recognition tasks
abstract
We propose to use orthogonal feature detectors in artificial neural networks for the robustness of performance under noisy conditions. The motivation is grounded on the principle that orthogonal decomposition is the most efficient among all representation of a signal. In this paper, we incorporate orthogonalization in the process of learning the network weights. In our implementation, the constraint of orthogonality is enforced by applying Gram-Schmidt processes to the feature detectors during network training. The proposed method is evaluated on MNIST database for hand-written digit recognition. The images in the training set are not corrupted, while the images in the test set are artificially corrupted with white noises. Experimental results show that the proposed orthogonalization method achieves 56.4% relative improvement in recognition error rate over a conventional learning method without orthogonalization. Given that the clean training data and the noisy test data are clearly mismatched, such an improvement with artificial neural networks is indeed very remarkable. For engineering insight, we devise a visualization tool which illuminates interesting features of the neurons learned by the proposed method.
Chia-Ping Chen, Po-Yuan Shih, Wei-Bin Liang
ICASSP1
2014 Natural speech synthesis based on hybrid approach with candidate expansion and verification
abstract
A hybrid Mandarin speech synthesis system combining concatenation-based and model-based methodology is investigated in this research. To effectively exploit a small-size corpus, the candidate sets for unit selection are expanded via clusters based on articulatory features (AF), which are estimated as the outputs of an artificial neural network. This is followed by a filtering operation incorporating residual compensation, to remove unsuitable units. Given an input text, an optimal unit sequence is decided by the minimization of a total cost, which depends on the spectral features, contextual articulatory features, formants, and pitch values. Furthermore, prosodic word verification is integrated to check the smoothness of the output speech. The units failing to pass the prosodic word verification are replaced by model-based synthesized units for better speech quality. Objective and subjective evaluations have been conducted. Comparisons among the proposed method, the HMM-based method, and the conventional hybrid method clearly show that candidate set expansion based on articulatory features lead to more units suitable for selection, and the verification process is effective in improving the naturalness of the output speech.
Chung-Hsien Wu 0001, Yi-Chin Huang, Shih-Lun Lin, Chia-Ping Chen
ICASSP4
2014 Speech emotion recognition with cross-lingual databases
abstract
In this paper, we investigate cross-lingual automatic speech emotion recognition. The basic idea is that since the emotion recognition system is based on the acoustic features only, it is possible to combine data in different languages to improve the recognition accuracy. We begin with the construction of a Mandarin database of emotional speech, which is similar to the well-known Berlin Database of Emotional Speech (EMO-DB) in the composition and size. In order to reduce the variability due to different languages and different speakers, we propose to apply histogram equalization as a data normalization method. Recognition systems based on support vector machines have been evaluated on EMO-DB. Compared to the baseline system without multi-lingual databases and data normalization, the proposed system has achieved a relative improvement of 39.9% in the emotion recognition accuracy, from 86.2% to 91.7%. The accuracy is among the best known results reported on EMO-DB, if not the best.
Bo-Chang Chiou, Chia-Ping Chen
INTERSPEECH2
2014 Polyglot Speech Synthesis Based on Cross-Lingual Frame Selection Using Auditory and Articulatory Features
abstract
In this paper, an approach for polyglot speech synthesis based on cross-lingual frame selection is proposed. This method requires only mono-lingual speech data of different speakers in different languages for building a polyglot synthesis system, thus reducing the burden of data collection. Essentially, a set of artificial utterances in the second language for a target speaker is constructed based on the proposed cross-lingual frame-selection process, and this data set is used to adapt a synthesis model in the second language to the speaker. In the cross-lingual frame-selection process, we propose to use auditory and articulatory features to improve the quality of the synthesized polyglot speech. For evaluation, a Mandarin-English polyglot system is implemented where the target speaker only speaks Mandarin. The results show that decent performance regarding voice identity and speech quality can be achieved with the proposed method.
Chia-Ping Chen, Yi-Chin Huang, Chung-Hsien Wu 0001, Kuan-De Lee
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Yet another Gaussian mixture model-based feature compensation method for robust noisy-digit recognition
abstract
We propose yet another Gaussian mixture model (YGMM) for robust speech recognition in noisy environments. The main difference between the proposed method and previously proposed GMM-based methods is that we estimate the noise features instead of the clean-speech features. In the implemented system, a condition classifier, incidentally based on GMM, is used to decide the noise type and level, and the corresponding GMM is employed to compensate for the noise-corrupted features. The proposed method and the implemented system are evaluated with the well-documented Aurora 2.0 noisy digit corpus. The results are promising. Specifically, it achieves a relative improvement in word error rate of 52.4% over the standard baseline, and 24.9% over a better baseline based on a traditional GMM-based feature compensation method.
Chia-Ping Chen, Bing-Feng Yeh
ICASSP1
2013 Query-Document Relevance Topic Models
Meng-Sung Wu, Chia-Ping Chen, Hsin-Min Wang
PAKDD (2)2
2012 Cross-lingual frame selection method for polyglot speech synthesis
abstract
A novel approach is proposed to creating a polyglot speech synthesis system without the need of collecting speech data from a bilingual (or multilingual) speaker, which is often expensive or even infeasible. Given a target speaker with data in the first language (Mandarin in this study), the basic idea is to construct artificial utterances in the second language (English) via selection of speech sample frames of the given speaker in the first language. As the speaker needs not be polyglot, this method is generally applicable to any speaker and any languages. In the search for optimal frame sequence selection, the candidate set is constrained by a decision tree for phone segments in the speech data of both languages, and the cost function depends on the context-dependent articulatory and auditory features. Evaluation results show that good performance regarding similarity (speaker identity) and naturalness (speech quality) can be achieved with the proposed method.
Chia-Ping Chen, Yi-Chin Huang, Chung-Hsien Wu 0001, Kuan-De Lee
ICASSP1
2012 Integrating Recognition and Retrieval With Relevance Feedback for Spoken Term Detection
abstract
Recognition and retrieval are typically viewed as two cascaded independent modules for spoken term detection (STD). Retrieval techniques are assumed to be applied on top of automatic speech recognition (ASR) output, with performance depending on ASR accuracy. We propose a framework that integrates recognition and retrieval and consider them jointly in order to yield better STD performance. This can be achieved either by adjusting the acoustic model parameters (model-based) or by considering detected examples (example-based) using relevance information provided by the user (user relevance feedback) or inferred by the system (pseudo-relevance feedback), either for a given query (short-term context) or by taking into account many previous queries (long-term context). Such relevance feedback approaches have long been used in text information retrieval, but are rarely considered and cannot be directly applied to the retrieval of spoken content. The proposed relevance feedback approaches are specific to spoken content retrieval and are hence very different from those developed for text retrieval, which are applied only to text symbols. We present not only these relevance feedback scenarios and approaches for STD, but also propose a framework to integrate them all together. Preliminary experiments showed significant improvements in each case.
Hung-yi Lee, Chia-Ping Chen, Lin-Shan Lee
IEEE Trans. Speech Audio Process.2
2012 Intrinsic Illumination Subspace for Lighting Insensitive Face Recognition
abstract
We introduce the intrinsic illumination subspace and its application for lighting insensitive face recognition in this paper. The intrinsic illumination subspace is constructed from illumination images of intrinsic images, which is a midlevel description of appearance images and can be useful for many visual inferences. This subspace forms a convex polyhedral cone and can be efficiently represented by a low-dimensional linear subspace, which enables an analytic generation of illumination images under varying lighting conditions. When only objects of the same class, such as faces, are concerned, a class-based generic intrinsic illumination subspace can be constructed in advance and used for novel objects of the same class. Based on this class-based generic subspace, we propose a lighting normalization method for lighting insensitive face recognition, where only a single input image is required. The generic subspace is used as a bootstrap subspace for illumination images of novel objects. Face recognition experiments are performed to demonstrate the effectiveness of the proposed lighting normalization method and verify empirically that the class-based generic subspace is applicable to novel objects. Our method is simple and fast, which makes it useful for real-time applications, embedded systems, or mobile devices with limited resources.
Chia-Ping Chen, Chu-Song Chen
IEEE Trans. Syst. Man Cybern. Part B1
2011 Improved spoken term detection with graph-based re-ranking in feature space
abstract
This paper presents a graph-based approach for spoken term detection. Each first-pass retrieved utterance is a node on a graph and the edge between two nodes is weighted by the similarity between the two utterances evaluated in feature space. The score of each node is then modified by the contributions from its neighbors by random walk or its modified version, because utterances similar to more utterances with higher scores should be given higher relevance scores. In this way the global similarity structure of all first-pass retrieved utterances can be jointly considered. Experimental results show that this new approach offers significantly better performance than the previously proposed pseudo-relevance feedback approach, which considers primarily the local similarity relationship between first-pass retrieved utterances, and these two different approaches can be cascaded to provide even better results.
Yun-Nung Chen, Chia-Ping Chen, Hung-yi Lee, Chun-an Chan, Lin-Shan Lee
ICASSP2
2011 Improved spoken term detection using support vector machines based on lattice context consistency
abstract
We propose an improved spoken term detection approach that uses support vector machines trained with lattice context consistency. The basic idea is that the same term usually have similar context, while quite different context usually implies the terms are different. Support vector machine can be trained using query context feature vectors obtained from the lattice to estimate better scores for ranking, and significant improvements can be obtained. This process can be performed iteratively and integrated with the pseudo relevance feedback in acoustic feature space proposed previously, both offering further improvements.
Hung-yi Lee, Tsung-wei Tu, Chia-Ping Chen, Chao-Yu Huang, Lin-Shan Lee
ICASSP3
2011 Real-time hand tracking on depth images
abstract
Hand tracking is a fundamental task in a gesture recognition system. Most previous works tracked the hand position on color images and relied heavily on skin color information. However, color information is very vulnerable to lighting variations and skin color varies across difference human races. Furthermore, one can not effectively discriminate faces or other skin-color-like objects from hands when using skin color detection. In this paper, we propose a hand tracking algorithm that uses depth images only, and also a hand click detection method to initialize the hand tracking automatically. We show that depth images suffice and are advantageous to real-time hand tracking. A region growing technique is applied to segment the hand region on depth images. Then a mean-shift based algorithm accurately locates the hand center in the segmented hand region. The experimental results show that the proposed tracking algorithm runs at 300+ FPS, and the average error of the tracked 3D hand positions is less than 1 centimeter. The proposed method enables a plethora of potential applications to natural Human-Computer Interaction (HCI), and is adequate for embedded systems of consumer electronics because of its low complexity and low bandwidth requirement.
Chia-Ping Chen, Ping-Han Lee, Yu-Pao Tsai, Shawmin Lei
VCIP1
2010 MOMI-Cosegmentation: Simultaneous Segmentation of Multiple Objects among Multiple Images
Wen-Sheng Chu, Chia-Ping Chen, Chu-Song Chen
ACCV (1)2
2010 Turning Rust into Gold: An ancient artifact as an interactive artwork
abstract
Turning Rust into Gold is inspired by a Chinese antique Mao-Kung Ting (cauldron) treasured by the National Palace Museum in Taiwan. Having a five-hundred-character inscription cast inside, and its weathered appearance made the Mao-Kung very unique. Motivated by revealing the great nature of the artifact and interpreting it into a meaningful narrative, we have proposed an interactive multimedia system that facilitates effective communication between museum audiences and the Mao-Kung Ting. Three technologies have been implemented to emphasize the weathered appearance of the bronze. De-/weathering simulation techniques have been deployed to revive the bronze to its original shiny gold color; while breath-based biofeedback and haptic technology have been utilized as user interfaces to trigger the de-weathering process of the Mao-Kung Ting. Also, the interactive scenarios have been designed with the Chinese cultural context and philosophy Qi, enabling users more easily fall into the Chinese civilization. The paper aims to present the development of the artwork Turing Rust into Gold, in order to further contribute to the feasibility of incorporating new media art in a historical museum context, and bring a new horizon in the museum sector.
Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Liang-Chun Lin, I-Ling Liu, Meng-Chieh Yu, Chu-Song Chen, Jiaping Wang
ICME4
2010 Improved spoken term detection by feature space pseudo-relevance feedback
abstract
Abstract In this paper, we propose an improved approach for spokenterm detection using pseudo-relevance feedback. To remedy theproblem of unmatched acoustic models with respect to spokenutterances produced under different acoustic conditions, whichmay give relatively poor recognition output, we integrate therelevance scores derived from the lattices with the DTW dis-tances derived from the feature space of MFCC parametersor phonetic posteriorgrams. These DTW distances are evalu-ated for a carefully selected set of pseudo-relevant utterances,which obtained from the first-pass returned list given by thesearch engine. The utterances on the first-pass returned list arethen reranked accordingly and finally shown to the user. Veryencouraging, performance improvements were obtained in thepreliminary experiments, especially when the acoustic modelsare poorly matched to the spoken utterances.Index Terms: spoken term detection, pseudo-relevance feed-back 1. Introduction Spoken term detection is to return a list of spoken utterancescontaining the term requested by the user. In many approachesof spoken term detection, the spoken utterances are first recog-nized and transformed into transcriptions or lattices by speechrecognition technologies, and then the search engine looksthrough all the transcriptions or lattices very similar to the text-based information retrieval. In this process much of the in-formation in the acoustic signals may be lost in the stage ofspeech recognition, especially when the acoustic models usedare not well matched to the characteristics of the acoustic sig-nals, which naturally results in degraded recognition accuracyand poor detection performance. This is very common in thescenario of spoken term detection, because the huge quantitiesof spoken utterances available over the Internet are naturallyproduced by many different people under many different acous-tic conditions, it is thus very difficult to train a set of acousticmodels well matched to so many different acoustic conditions.As a result, when the relevance scores such as the posteriorprobabilities of the query term derived from transcriptions orlattices are used to rank the retrieved utterances, it is hard tojudge whether a word hypothesis of the query in the transcrip-tions or lattices is a positive target or a false alarm when therecognition output is unreliable. Although many efficient ap-proaches [1, 2, 3] have been proposed to enhance the detectionperformance due to the relatively poor recognition output, thecompensative information straightly from the feature space isnecessary.In text-based information retrieval, even if the texts to beretrieved include all precise words, it is still difficult to retrieveall documents relevant to the query term because many of themdo not include the very short query term entered by the user.However, because many related terms may co-occur in manyrelated documents, a document containing some words appear-ing in some documents identified to be relevant by the searchengine may have high probability to be relevant, even if it doesnot include the query term. For example, a document includingthe words ”George Bush”, ”US”, ”Middle East” may be relevantto a query term of ”White House”, even if it does not includethe query term of ”White House”. In other words, it is possi-ble to retrieve the relevant documents without the query termsince they are ”similar” to some retrieved relevant documentsin some way. Pseudo-relevance feedback, also known as blindrelevance feedback, is one way to realize the above idea. In thisapproach, it is assumed that the set of documents appearing onthe top of the retrieved document list are relevant (or ”pseudo-relevant”), so documents somehow similar to those ”pseudo-relevant” documents can be retrieved, for example, by expand-ing the query with keywords from those ”pseudo-relevant” doc-uments [4]. Similar idea of pseudo-relevance feedback has beenapplied on spoken term detection [5].In this paper, we try to perform similar pseudo-relevancefeedback for spoken term detection as shown in Figure 1. Theupper half of Figure 1 is the conventional spoken term detec-tion. MFCC features were obtained from all spoken utterancesin the archive, speech recognition produces lattices for the ut-terances, and the retrieved engine selects the utterances basedon the relevance scores evaluated from the lattices with respectto the query Qentered by the user. The approach proposedhere in this paper is shown in the lower half of Figure 1. Thefirst-pass returned list is not shown to the user, but instead a”pseudo-relevant utterance set X
Chia-Ping Chen, Hung-yi Lee, Ching-Feng Yeh, Lin-Shan Lee
INTERSPEECH1
2010 Improved spoken term detection by discriminative training of acoustic models based on user relevance feedback
Hung-yi Lee, Chia-Ping Chen, Ching-Feng Yeh, Lin-Shan Lee
INTERSPEECH2
2010 Empirical mode decomposition for noise-robust automatic speech recognition
Kuo-Hao Wu, Chia-Ping Chen
INTERSPEECH2
2010 Transformational Breathing between Present and Past: Virtual Exhibition System of the Mao-Kung Ting
Chun-Ko Hsieh, Xin Tong 0001, Yi-Ping Hung, Chia-Ping Chen, Ju-Chun Ko, Meng-Chieh Yu, Han-Hung Lin, Szu-Wei Wu, Yi-Yu Chung, Liang-Chun Lin, Ming-Sui Lee, Chu-Song Chen, Jiaping Wang, Quo-Ping Lin, I-Ling Liu
MMM4
2010 A framework integrating different relevance feedback scenarios and approaches for spoken term detection
abstract
This paper presents a new framework integrating different relevance feedback scenarios (pseudo relevance feedback and user relevance feedback in short- and long-term context) and different approaches (model- and example-based) in a spoken term detection system, and shows the retrieval performance can be improved step by step. It is found that short-term context user relevance feedback can further improve the retrieval performance after pseudo relevance feedback, regardless of whether the acoustic models have been adapted by matched data or long-term context user relevance feedback or not. Moreover, model-based and example-based methods are shown to be additive when integrated in short-term context user relevance feedback scenario.
Hung-yi Lee, Chia-Ping Chen, Ching-Feng Yeh, Lin-Shan Lee
SLT2
2009 Speaker diarization using divide-and-conquer
abstract
Speaker diarization systems usually consist of two core components: speaker segmentation and speaker clustering. The current state-of-the-art speaker diarization systems usually apply hierarchical agglomerative clustering (HAC) for speaker clustering after segmentation. However, HAC’s quadratic computational complexity with respect to the number of data samples inevitably limits its application in large-scale data sets. In this paper, we propose a divide-and-conquer (DAC) framework for speaker diarization. It recursively partitions the input speech stream into two sub-streams, performs diarization on them separately, and then combines the diarization results obtained from them using HAC. The results of experiments conducted on RT-02 and RT-03 broadcast news data show that the proposed framework is faster than the conventional segmentation and clustering-based approach while achieving comparable diarization accuracy. Moreover, the proposed framework obtains a higher speedup over the conventional approach on a larger test data set. Index Terms: speaker diarization, speaker segmentation, speaker clustering, divide-and-conquer
Shih-Sian Cheng, Chun-Han Tseng, Chia-Ping Chen, Hsin-Min Wang
INTERSPEECH3
2009 Noise-robust feature extraction based on forward masking
Sheng-Chiuan Chiou, Chia-Ping Chen
INTERSPEECH2
2007 MVA Processing of Speech Features
abstract
In this paper, we investigate a technique consisting of mean subtraction, variance normalization and time sequence filtering. Unlike other techniques, it applies auto-regression moving-average (ARMA) filtering directly in the cepstral domain. We call this technique mean subtraction, variance normalization, and ARMA filtering (MVA) post-processing, and speech features with MVA post-processing are called MVA features. Overall, compared to raw features without post-processing, MVA features achieve an error rate reduction of 45% on matched tasks and 65% on mismatched tasks on the Aurora 2.0 noisy speech database, and an average 57% error reduction on the Aurora 3.0 database. These improvements are comparable to the results of much more complicated techniques even though MVA is relatively simple and requires practically no additional computational cost. In this paper, in addition to describing MVA processing, we also present a novel analysis of the distortion of mel-frequency cepstral coefficients and the log energy in the presence of different types of noise. The effectiveness of MVA is extensively investigated with respect to several variations: the configurations used to extract and the type of raw features, the domains where MVA is applied, the filters that are used, the ARMA filter orders, and the causality of the normalization process. Specifically, it is argued and demonstrated that MVA works better when applied to the zeroth-order cepstral coefficient than to log energy, that MVA works better in the cepstral domain, that an ARMA filter is better than either a designed finite impulse response filter or a data-driven filter, and that a five-tap ARMA filter is sufficient to achieve good performance in a variety of settings. We also investigate and evaluate a multi-domain MVA generalization
Chia-Ping Chen, Jeff A. Bilmes
IEEE Trans. Speech Audio Process.1
2006 The 4-Source Photometric Stereo Under General Unknown Lighting
Chia-Ping Chen, Chu-Song Chen
ECCV (3)1
2006 Chinese input method based on reduced Mandarin phonetic alphabet
Chun-Han Tseng, Chia-Ping Chen
INTERSPEECH2
2005 Speech Feature Smoothing for Robust ASR
abstract
We evaluate smoothing within the context of the MVA (mean subtraction, variance normalization, and ARMA filtering) post-processing scheme for noise-robust automatic speech recognition. MVA has shown great success in the past on the Aurora 2.0 and 3.0 corpora, even though it is computationally inexpensive. MVA is applied to many acoustic feature extraction methods, and is evaluated using Aurora 2.0. We evaluate MVA post-processing on MFCCs, LPCs, PLPs, RASTA, Tandem, modulation-filtered spectrogram, and modulation cross-correlogram features. We conclude that, while effectiveness does depend on the extraction method, the majority of features benefit significantly from MVA, and the smoothing ARMA filter is an important component. It appears that the effectiveness of normalization and smoothing depends on the domain in which it is applied, being most fruitfully applied just before being scored by a probabilistic model. Moreover, since it is both effective and simple, our ARMA filter should be considered a candidate method in most noise-robust speech recognition tasks.
Chia-Ping Chen, Jeff A. Bilmes, Daniel P. W. Ellis
ICASSP (1)1
2005 Lighting Normalization with Generic Intrinsic Illumination Subspace for Face Recognition
abstract
In this paper, we introduce the concept of intrinsic illumination subspace which is based on the intrinsic images. This intrinsic illumination subspace enables an analytic generation of the illumination images under varying lighting conditions. When objects of the same class are concerned, our method allows a class-based generic intrinsic illumination subspace to be constructed in advance. We propose a lighting normalization method based on the generic intrinsic illumination subspace, which is used as a bootstrap subspace for novel images. Face recognition experiments are performed to demonstrate the effectiveness of our method.
Chia-Ping Chen, Chu-Song Chen
ICCV1
2005 Focused word segmentation for ASR
abstract
We propose a new set of features based on the temporal statistics of the spectral entropy of speech. We show why these features make good inputs for a speech detector. Moreover, we propose a back-end that uses the evidence from the above features in a ‘focused’ manner. Subsequently, by means of recognition experiments we show that using the above back-end leads to significant performance improvements, but merely appending the features to the standard feature vector does not improve performance. We also report a 10% average improvement in word error rate over our baseline for the highly mis-matched case in the Aurora3.0 corpus.
Amarnag Subramanya, Jeff A. Bilmes, Chia-Ping Chen
INTERSPEECH3
2004 Image set compression through minimal-cost prediction structures
abstract
We propose a new scheme for compressing on image set by building its minimal-cost prediction structure. Existing prediction-based video coding methods can be easily extended and incorporated into this scheme to achieve higher compression efficiency. According to this prediction structure, we also develop a progressive transmission approach for interactive object movie (OM) browsing.
Chia-Ping Chen, Chu-Song Chen, Kuo-Liang Chung, Hsueh-I Lu, Gregory Y. Tang
ICIP1
2002 Low-resource noise-robust feature post-processing on Aurora 2.0
Chia-Ping Chen, Jeff A. Bilmes, Katrin Kirchhoff
INTERSPEECH1
2002 Frontend post-processing and backend model enhancement on the Aurora 2.0/3.0 databases
Chia-Ping Chen, Karim Filali, Jeff A. Bilmes
INTERSPEECH1