EDBT 2026 Demo / reviewers in the wild / expert
Jen-Tzung Chien
dblp:03/3569
· DBLP profile ↗
205ranked-venue papers
123as first author
52since 2021 · last 2026
0000-0003-3466-8941ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 138 · 71 first-author · 33 since 2021Artificial intelligence and machine learning · 129 · 88 first-author · 36 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Risk-Aware Bilingual Spoken Dialogue for Campus Mental Health SupportabstractThis work presented a web-based system which introduces an active-listening strategy in a spoken dialogue for self-disclosure to support mental health of a campus user. To enhance the system usability and safety, this demo is developed to conduct the bilingual (Mandarin/English) spoken dialogue where a high-risk dialogue detection during speech interaction is reliably augmented. In particular, a prompt-driven GPT classifier identifies the utterances indicating self-harm or suicide intent and triggers safety alerts with help center and counselor notification. We also integrate a TTS module for Taiwanese Mandarin and standard English, and redesign the user interface to automatically pop up alert messages when high-risk dialogue is detected. In addition, we collect speech data under diverse mental dialogue scenarios with bilingual speech to enable system analysis, evaluation and refinement. Overall, these extensions build a framework that promotes empathetic interactions, enables timely alert in critical cases, and improves the accessibility for diverse users. You-Teng Lin, Li-Yang Zhang, Yi-Tang Chen, Jen-Tzung Chien |
AAAI | 4 |
| 2026 | Outlier-Aware Contrastive LearningabstractContrastive learning aims to learn an embedding space with sample discrimination where similar samples attract together while dissimilar samples repulse apart. However, the issue of sampling bias likely happens and degrades the classification performance when a contrast model is trained with the leakage caused by similar samples but from different classes or dissimilar samples from the same class. Out-of-distribution (OOD) detection provides a meaningful scheme to detect and mask those false negative samples for debiasing in an outlier-aware contrastive loss for high-fidelity contrastive learning. Sample debiasing is feasible to reduce the upper bound of contrastive loss. Also, the previous OOD detector was trained from auxiliary collection of OOD samples. In real world, the prior knowledge of OOD samples is commonly unavailable. This study presents new outlier-aware detection and contrast models through generation and augmentation of those samples near the boundary between in-distribution (ID) and OOD. These synthesized samples are located right outside ID, and their Gaussian embeddings sufficiently reflect OOD behaviors. An OOD detector is learned by using ID samples and synthesized OOD samples with the learning objective towards contrastive OOD detection and debiased contrast model. The experiments are conducted to illustrate the merit of the proposed outlier-aware contrastive learning. Jen-Tzung Chien, Kuan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Contrastive Mixture Diffusion ModelsabstractContinuous diffusion model is feasible to perform smoothed transitions for flexible text generation in a continuous latent space. A challenging issue in continuous diffusion is to implement the denoising process and manipulate the generated tokens to reflect physical phenomenon in diffusion steps towards a contextually and semantically meaningful sentence. Basically, a high-quality text generation can be fulfilled by an easy-first strategy, which prioritizes the generation of high-frequency simple or general tokens in the beginning and then low-frequency complex or specific tokens in the end via mask language model. This paper introduces the mask noise as an absorbing state and develops a new continuous-discrete denoising process in a diffusion model. Each diffusion step is performed by generating either a continuous word embedding or a discrete mask token where the latent transition is modeled by a Gaussian-Dirac mixture distribution. The easy-first text generation is then implemented and strengthened via a contrastive loss to disentangle the generation between simple and complex tokens. A contrastive mixture diffusion model is accordingly exploited by minimizing a variational bound of negative log likelihood, which is regularized by aligning the denoising network posterior with the Gaussian-Dirac posterior in each diffusion transition based on an approximate Kullback-Leibler divergence. The experiments on text paraphrasing and other tasks demonstrate the effectiveness and efficiency of sentence generation by using the proposed method where the easy-first strategy in generation behavior is illustrated. Jen-Tzung Chien |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | CGDD: Contrastive Gaussian-Dirac Diffusion ModelabstractThis paper proposes the contrastive Gaussian-Dirac diffusion (CGDD) model, which introduces a novel approach to hybrid noise diffusion by combining continuous and discrete noise processes for text generation tasks. Leveraging the contrastive learning, the proposed CGDD structures the embedding space based on the word frequency, promoting an easy-first generation approach that aligns with the observed frequency of word usage. Furthermore, a new approximation is presented for the loss function in the continuous-discrete diffusion setting, addressing the limitations of the previous models that rely solely on the Gaussian distributions. Experimental results demonstrate that the proposed approaches are effective based on the evaluations over multiple measurements. Hsin-Yi Lin, Jen-Tzung Chien |
ICASSP | 3 |
| 2025 | Attention Disentanglement for Semantic Diffusion Modeling in Text-to-Image GenerationabstractText-to-image model has been recently improved to generate the semantically rich high-quality images by strengthening natural language processing via transformer in a stable diffusion process. However, the challenges are still remained in accurately rendering the objects, colors and compositions, and in precisely resolving the ambiguities in textual descriptions. This paper presents an attention disentanglment for semantic diffusion where the semantic consistency is enhanced for text-to-image generation. By utilizing the cross-attention maps in stable diffusion, this method is feasible to control the semantic features during generation process without the need for retraining. Additionally, a vision-language model is merged to implement this process to ensure that the generated images are closely aligned with the input prompts. This study further presents the attention disentanglement for semantic diffusion where the schemes of syntax parsing and information-theoretic learning are implemented to separate the semantic attributes and relations, thereby improving the accuracy and coherence of the generated images in the experiments. Hsiang-Chun Yu, Jen-Tzung Chien |
ICASSP | 2 |
| 2025 | Reinforced Retrieval-Augmented Generation in Large Language ModelsabstractLarge language model (LLM) has been recognized as a powerful learning machine for natural language generation. However, LLM likely generates the factually incorrect or irrelevant responses, which result in an issue of hallucination. This issue is sometimes risky especially when the application domains require high accuracy with safe control in the generated responses. This paper presents a reinforced retrieval method to build a trustworthy generative model where the reinforcement learning based on proximal policy optimization is developed to implement the reinforced retrieval-augmented generation. Such a solution is able to dynamically optimize the quality, the consistency and the completion in content generation. By balancing the fluency and accuracy through a tailored reward function, this model is optimized to pave an avenue to mitigate hallucination in using LLMs for text generation. Experimental results on NQ and TriviaQA datasets show that the proposed reinforced retrieval-augmented generation outperforms the existing models particularly in both long-form and short-form question-answering tasks in terms of various evaluation metrics. Jen-Tzung Chien, Zhi-Xuan Tai |
IJCNN | 1 |
| 2025 | Disentangling Speaker and Content in Pre-trained Speech Models with Latent Diffusion for Robust Speaker Verification
Zhe Li 0030, Man-Wai Mak, Jen-Tzung Chien, Mert Pilanci, Zezhong Jin, Helen M. Meng |
INTERSPEECH | 3 |
| 2025 | CAPR: Confidence-Aware Prompt Refinement in Large Language Models
Jen-Tzung Chien, Po-Chun Huang |
INTERSPEECH | 1 |
| 2024 | Attention-Guided Adaptation for Code-Switching Speech RecognitionabstractThe prevalence of the powerful multilingual models, such as Whisper, has significantly advanced the researches on speech recognition. However, these models often struggle with handling the code-switching setting, which is essential in multilingual speech recognition. Recent studies have attempted to address this setting by separating the modules for different languages to ensure distinct latent representations for languages. Some other methods considered the switching mechanism based on language identification. In this study, a new attention-guided adaptation is proposed to conduct parameter-efficient learning for bilingual ASR. This method selects those attention heads in a model which closely express language identities and then guided those heads to be correctly attended with their corresponding languages. The experiments on the Mandarin-English code-switching speech corpus show that the proposed approach achieves a 14.2% mixed error rate, surpassing state-of-the-art method, where only 5.6% additional parameters over Whisper are trained. Bobbi Aditya, Mahdin Rohmatillah, Liang-Hsuan Tai, Jen-Tzung Chien |
ICASSP | 4 |
| 2024 | Towards a Unified View of Adversarial Training: A Contrastive PerspectiveabstractAdversarial training (AT) has been an effective approach to build a defensive model against adversarial attacks. However, most researches on AT are conducted in a supervised learning manner where true labels of training data are required. To relax this issue, the unsupervised AT through self-supervised learning is developed. In particular, this study presents an unsupervised AT by exploiting the concept of instance discrimination in contrastive learning where the unsupervised learning is implemented but closely connected to a supervised scheme for discrimination or classification. By utilizing such an implicit and inherent connection, a unified view is addressed for a new unsupervised AT where the contrastive learning objective is consolidated and strengthened from a classification perspective. A unified framework for supervised and unsupervised AT is extended. In the experiments, this method achieves state-of-the-art results in adversarial training of image classifier under different settings. Jen-Tzung Chien, Yuan-An Chen |
ICASSP | 1 |
| 2024 | Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker VerificationabstractContrastive self-supervised learning (CSL) for speaker verification (SV) has drawn increasing interest recently due to its ability to exploit unlabeled data. Performing data augmentation on raw waveforms, such as adding noise or reverberation, plays a pivotal role in achieving promising results in SV. Data augmentation, however, demands meticulous calibration to ensure intact speaker-specific information, which is difficult to achieve without speaker labels. To address this issue, we introduce a novel framework by incorporating clean and augmented segments into the contrastive training pipeline. The clean segments are repurposed to pair with noisy segments to form additional positive and negative pairs. Moreover, the contrastive loss is weighted to increase the difference between the clean and augmented embeddings of different speakers. Experimental results on Voxceleb1 suggest that the proposed framework can achieve a remarkable 19% improvement over the conventional methods, and it surpasses many existing state-of-the-art techniques. Chong-Xin Gan, Man-Wai Mak, Weiwei Lin 0002, Jen-Tzung Chien |
ICASSP | 4 |
| 2024 | Revise the NLU: A Prompting Strategy for Robust Dialogue SystemabstractThe advent of large language models (LLMs), such as GPT 3.5, has demonstrated significant potential, especially when paired with the prompt engineering techniques. However, while this setting excels in zero-shot or few-shot scenario, the direct utilization of LLMs in multi-domain task-oriented dialogue (TOD) systems often falls short compared to smaller task-specific models in standard evaluations. This indicates the need for further exploration on how to harness the power of LLMs effectively to multi-domain TOD systems. This paper addresses the aforementioned challenge by introducing a novel prompting strategy to enhance the robustness of the existing text-based multi-domain TOD systems. This strategy aims to revise the outputs of natural language understanding (NLU) component through a series of prompting steps. By capitalizing on NLU outputs, a simple and straightforward prompt design can be carried out. Experimental results illustrate the benefit of the proposed strategy in improving robustness of the multi-domain TOD system. Mahdin Rohmatillah, Jen-Tzung Chien |
ICASSP | 2 |
| 2024 | Contrastive Speaker Embedding With Sequential DisentanglementabstractContrastive speaker embedding assumes that the contrast between the positive and negative pairs of speech segments is attributed to speaker identity only. However, this assumption is incorrect because speech signals contain not only speaker identity but also linguistic content. In this paper, we propose a contrastive learning framework with sequential disentanglement to remove linguistic content by incorporating a disentangled sequential variational autoencoder (DSVAE) into the conventional SimCLR framework. The DSVAE aims to disentangle speaker factors from content factors in an embedding space so that only the speaker factors are used for constructing a contrastive loss objective. Because content factors have been removed from the contrastive learning, the resulting speaker embeddings will be content-invariant. Experimental results on VoxCeleb1-test show that the proposed method consistently outperforms SimCLR. This suggests that applying sequential disentanglement is beneficial to learning speaker-discriminative embeddings. Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien |
ICASSP | 3 |
| 2024 | Mask Consistency and Contrast Regularization for Prompt-Based LearningabstractThis paper presents a new contrastive semi-supervised learning approach to data efficient language model where the fine-tuned mask language model is constructed with hard prompts for sentence representation. Importantly, the prompt-based pseudo labeling is merged in a mask language model (MLM) with data augmentation through a contrast loss which is utilized to pull together those similar samples and push apart the samples which do not belong to their augmentations. Furthermore, the mask consistency training is implemented for fine-tuned MLM to pursue the closeness between word prediction from weak and strong augmentations. Accordingly, this study develops a data efficient solution which improves the model generalization and leverages the rich information from unlabeled data for few-shot text classification. The experiments on natural language understanding in few-shot and semi-supervised settings show that the proposed method considerably improves the performance by using the contrastive prompt-based learning and the mask consistency training. Jen-Tzung Chien, Chien-Ching Chen |
IJCNN | 1 |
| 2024 | False Negative Masking for Debiasing in Contrastive LearningabstractContrastive learning has been popular to carry out self-supervised learning where a meaningful representation with instance discrimination is learned without any label information. However, recent studies have found that there might exist some false negative samples in training data, which have the same label as that of anchor. This phenomenon, also known as the sampling bias, considerably degrades the system performance in a downstream task. Accordingly, it is crucial to identify and reject those false negative samples without accessing their labels. This study deals with such a challenging issue and presents an approach based on the out-of-distribution (OOD) detection which identifies and masks the false negative samples. In general, the samples from the same class are seen as in-distribution (ID) data whereas the samples from the other classes are viewed as OOD data. Therefore, this study presents a new contrastive learning by detecting those false negative samples and masking them in calculation of contrastive loss during optimization. Compared with the original contrastive learning, the proposed method is illustrated with a tighter upper bound in false negative masking. The experiments demonstrate that the proposed method can achieve competitive results. Jen-Tzung Chien, Kuan Chen |
IJCNN | 1 |
| 2024 | Contrastive Meta Learning for Soft Prompts Using Dynamic MixupabstractSoft prompting is crucial in few-shot settings which are essential to carry out parameter efficient learning for natural language understanding (NLU). A new trend of solutions has been recently developed by merging meta learning schemes over different domains. This study presents a new metric-based meta learning where the contrastive perspective is implemented to enhance the discrimination among confusing classes, and accordingly a rapid domain adaptation is feasible to work for unseen tasks in text classification. To further address the generalization issue, this paper proposes a mixup augmentation where the mixup ratio is automatically estimated according to the maximum entropy principle. The experimental results on a series of NLU tasks show the merit of the proposed method in terms of classification accuracy, parameter efficiency and latent visualization in most of few-shot settings. Jen-Tzung Chien, Hsin-Ti Wang, Ching-Hsien Lee |
IJCNN | 1 |
| 2024 | Collaborative Contrastive Learning for Hypothesis Domain AdaptationabstractInterspeech 2024, 1-5 September 2024, Kos, Greece Jen-Tzung Chien, I-Ping Yeh, Man-Wai Mak |
INTERSPEECH | 1 |
| 2024 | Modality Translation Learning for Joint Speech-Text Model
Pin-Yen Liu, Jen-Tzung Chien |
INTERSPEECH | 2 |
| 2024 | Reliable dialogue system for facilitating student-counselor communication
Mahdin Rohmatillah, Bryan Gautama Ngo, Willianto Sulaiman, Po-Chuan Chen, Jen-Tzung Chien |
INTERSPEECH | 5 |
| 2024 | Cross-Modality Diffusion Modeling and Sampling for Speech RecognitionabstractThe diffusion model excels as a generative model for continuous data within a single modality.To extend its effectiveness to speech recognition, where the continuous speech frames are used as the condition to generate the discrete word tokens, building a conditional diffusion across discrete state space becomes crucial.This paper introduces a non-autoregressive discrete diffusion model, enabling parallel generation of a word string corresponding to a speech signal through iterative diffusion steps.An acoustic transformer encoder identifies the speech representation, serving as the condition for a denoising transformer decoder to predict the whole discrete sequence.To address the redundancy reduction in cross-modality diffusion, an additional feature decorrelation objective is integrated during optimization.This paper further reduces the inference time by using a fast sampling approach.The experiments on speech recognition illustrate the merit of the proposed method. Chia-Kai Yeh, Ching-Hsien Hsu, Jen-Tzung Chien |
INTERSPEECH | 4 |
| 2024 | Taming NLU Noise: Student-Teacher Learning for Robust Dialogue PolicyabstractDialogue policy is a crucial component of dialogue systems, responsible for determining system responses based on user inputs. While reinforcement learning (RL) can effectively optimize the dialogue policy, the system performance in real-world settings is heavily influenced by an earlier component for natural language understanding (NLU). Once the NLU produces a wrong information, the dialogue policy will be affected to degrade the performance. To enhance the robustness of dialogue policy, this paper proposes integrating RL optimization with a noisy student-teacher learning, taming the noise generated by NLU. To prevent overconfidence during knowledge transfer from the teacher, we introduce a dual-teacher mechanism where knowledge distillation is carried out by using dynamic changes in the samples stored in the replay buffer which leverages the exploration-exploitation paradigm from RL. Evaluations on multi-domain multi-turn dialogue tasks demonstrate the effectiveness of this approach which shows the increased robustness to noisy NLU outputs and accordingly the improved overall system performance. Mahdin Rohmatillah, Jen-Tzung Chien |
SLT | 2 |
| 2024 | Latent Semantic and Disentangled AttentionabstractSequential learning using transformer has achieved state-of-the-art performance in natural language tasks and many others. The key to this success is the multi-head self attention which encodes and gathers the features from individual tokens of an input sequence. The mapping or decoding is performed to produce an output sequence via cross attention. There are threefold weaknesses by using such an attention framework. First, since the attention would mix up the features of different tokens in input and output sequences, it is likely that redundant information exists in sequence data representation. Second, the patterns of attention weights among different heads tend to be similar. The model capacity is bounded. Third, the robustness in an encoder-decoder network against the model uncertainty is disregarded. To handle these weaknesses, this paper presents a Bayesian semantic and disentangled mask attention to learn latent disentanglement in multi-head attention where the redundant features in transformer are compensated with the latent topic information. The attention weights are filtered by a mask which is optimized through semantic clustering. This attention mechanism is implemented according to Bayesian learning for clustered disentanglement. The experiments on machine translation and speech recognition show the merit of Bayesian clustered disentanglement for mask attention. Jen-Tzung Chien, Yu-Han Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Contrastive Self-Supervised Speaker Embedding With Sequential DisentanglementabstractContrastive self-supervised learning has been widely used in speaker embedding to address the labeling challenge. Contrastive speaker embedding assumes that the contrast between the positive and negative pairs of speech segments is attributed to speaker identity only. However, this assumption is incorrect because speech signals contain not only speaker identity but also linguistic content. In this paper, we propose a contrastive learning framework with sequential disentanglement to remove linguistic content by incorporating a disentangled sequential variational autoencoder (DSVAE) into the conventional contrastive learning framework. The DSVAE aims to disentangle speaker factors from content factors in an embedding space so that the speaker factors become the main contributor to the contrastive loss. Because content factors have been removed from contrastive learning, the resulting speaker embeddings will be content-invariant. The learned embeddings are also robust to language mismatch. It is shown that the proposed method consistently outperforms the conventional contrastive speaker embedding on the VoxCeleb1 and CN-Celeb datasets. This finding suggests that applying sequential disentanglement is beneficial to learning speaker-discriminative embeddings. Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Learning Flow-Based DisentanglementabstractFace reenactment aims to generate the talking face images of a target person given by a face image of source person. It is crucial to learn latent disentanglement to tackle such a challenging task through domain mapping between source and target images. The attributes or talking features due to domains or conditions become adjustable to generate target images from source images. This article presents an information-theoretic attribute factorization (AF) where the mixed features are disentangled for flow-based face reenactment. The latent variables with flow model are factorized into the attribute-relevant and attribute-irrelevant components without the need of the paired face images. In particular, the domain knowledge is learned to provide the condition to identify the talking attributes from real face images. The AF is guided in accordance with multiple losses for source structure, target structure, random-pair reconstruction, and sequential classification. The random-pair reconstruction loss is calculated by means of exchanging the attribute-relevant components within a sequence of face images. In addition, a new mutual information flow is constructed for disentanglement toward domain mapping, condition irrelevance, and condition relevance. The disentangled features are learned and controlled to generate image sequence with meaningful interpretation. Experiments on mouth reenactment illustrate the merit of individual and hybrid models for conditional generation and mapping based on the informative AF. Jen-Tzung Chien, Sheng-Jhe Huang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Adversarial Augmentation For Adapter LearningabstractThe recent pre-trained models have achieved state-of-the-art results for natural language understanding (NLU) and automatic speech recognition (ASR). However, the pre-trained models likely suffer from the overfitting problem when adapting the model to a low-resource target domain. This study handles this low-resource setting by training an adversarial adapter based on a pre-trained backbone model. The adversarial training is performed by implementing the data augmentation rather than enhancing the adversarial robustness. The proposed method leverages adversarial training to collect augmented data to reinforce adapter learning with a smoothed decision boundary. The size of trainable parameters is tightly controlled to alleviate the overfitting to enhance the model capability. In the experiments, this work considerably improves the performance in NLU tasks. The adversarial adapter learning is further extended for ASR to show the merit of this method in terms of efficiency and accuracy. Jen-Tzung Chien, Wei-Yu Sun |
ASRU | 1 |
| 2023 | Meta Learning for Domain Agnostic Soft PromptabstractThe prompt-based learning, as used in GPT-3, has become a popular approach to extract knowledge from a powerful pre-trained language model (PLM) for natural language understanding tasks. However, either applying the hard prompt for sentences by defining a collection of human-engineering prompt templates or directly optimizing the soft or continuous prompt with labeled data may not really generalize well for unseen domain data. To cope with this issue, this paper presents a new prompt-based unsupervised domain adaptation where the learned soft prompt is able to boost the frozen pre-trained language model to deal with the input tokens from unseen domains. Importantly, the meta learning and optimization is developed to carry out the domain agnostic soft prompt where the loss for masked language model is minimized. The experiments on multi-domain natural language understanding tasks show the merits of the proposed method. Ming-Yen Chen, Mahdin Rohmatillah, Ching-Hsien Lee, Jen-Tzung Chien |
ICASSP | 4 |
| 2023 | Self-Supervised Adversarial Training for Contrastive Sentence EmbeddingabstractThe defense against adversarial attacks was originally proposed for computer vision, and recently such an adversarial training (AT) has been emerging for natural language understanding. In an AT process, the adversarial perturbations are added on the input word embeddings as the noisy data which are included to allow the trained model to be noise invariant and accordingly improve the model generalization. However, the performance of existing works was bounded under the supervised or semi-supervised setting. In addition, the contrastive learning (CL) has obtained a significant performance in a self-supervised pre-training for language models. This paper presents a novel method to re-formulate CL to meet a self-supervised classification objective. Using this new formula, a self-supervised AT method is proposed for training an efficient sentence encoder. Experiments show that the pro-posed CL can improve the previous methods to find unsupervised sentence embeddings. With the help of AT, this method further surpasses the previous supervised methods. Jen-Tzung Chien, Yuan-An Chen |
ICASSP | 1 |
| 2023 | Variational Skill Embeddings for Meta Reinforcement LearningabstractMeta reinforcement learning (meta-RL) aims to learn useful prior knowledge across tasks which can be generalized to unseen but similar tasks with only a small number of adaptation steps. Traditionally, the gradient-based metal RL was proposed to use the gradients to learn the parameters of an adaptive policy from different tasks which likely lacked sample efficiency. Recently, the context-based meta-RL improved the efficiency by learning the embeddings of the trajectories based on context representation. The learned policy can be adapted to new tasks, but the performance is bounded due to a simple context encoder. To deal with this insufficiency, this paper presents a novel regularized meta-RL where the generalization of policy is enhanced through a context-based meta-RL where the conditional variational autoencoder consisting of a context-skill encoder and a soft-actor-critic decoder is implemented. The proposed method pursues the model regularization by discovering the shared skill patterns across tasks in implementation of context-based meta-RL. The experiments on a number of benchmark tasks show the merit of variational skill embeddings for regularized meta-RL. Jen-Tzung Chien, Weiwei Lai |
IJCNN | 1 |
| 2023 | Variational Disentangled Attention and Regularization for Visual DialogabstractOne of the most important challenges in a visual dialog is to effectively extract the information from a given image and its historical conversation which are related to the current question. Many studies adopt the soft attention mechanism in different information sources due to its simplicity and ease of optimization. However, some of visual dialogs are observed in a single round. This implies that there is no substantial correlation between individual rounds of questions and answers. This paper presents a unified approach to disentangled attention to deal with context-free visual dialogs. The question is disentangled in latent representation. In particular, an informative regularization is imposed to strengthen the dependence between vision and language by pretraining on the visual question answering before transferring to visual dialog. Importantly, a novel variational attention mechanism is developed and implemented by a local reparameterization trick which carries out a discrete attention to identify the relevant conversations in a visual dialog. A set of experiments are evaluated to illustrate the merits of the proposed attention and regularization schemes for visual dialogs. Jen-Tzung Chien, Hsiu-Wei Tien |
IJCNN | 1 |
| 2023 | Contrastive Disentangled Learning for Memory-Augmented Transformer
Jen-Tzung Chien, Shang-En Li |
INTERSPEECH | 1 |
| 2023 | Promoting Mental Self-Disclosure in a Spoken Dialogue System
Mahdin Rohmatillah, Bobbi Aditya, Li-Jen Yang, Bryan Gautama Ngo, Willianto Sulaiman, Jen-Tzung Chien |
INTERSPEECH | 6 |
| 2023 | Parameter-Efficient Learning for Text-to-Speech Accent AdaptationabstractThis paper presents a parameter-efficient learning (PEL) to develop a low-resource accent adaptation for text-to-speech (TTS).A resource-efficient adaptation from a frozen pre-trained TTS model is developed by using only 1.2% to 0.8% of original trainable parameters to achieve competitive performance in voice synthesis.Motivated by a theoretical foundation of optimal transport (OT), this study carries out PEL for TTS where an auxiliary unsupervised loss based on OT is introduced to maximize a difference between the pre-trained source domain and the (unseen) target domain, in addition to its supervised training loss.Further, we leverage upon this unsupervised loss refinement to boost system performance via either sliced Wasserstein distance or maximum mean discrepancy.The merit of this work is demonstrated by fulfilling PEL solutions based on residual adapter learning, and model reprogramming when evaluating the Mandarin accent adaptation.Experiment results show that the proposed methods can achieve competitive naturalness with parameter-efficient decoder fine-tuning, and the auxiliary unsupervised loss improves model performance empirically. Li-Jen Yang, Chao-Han Huck Yang, Jen-Tzung Chien |
INTERSPEECH | 3 |
| 2023 | Learning Continuous-Time Dynamics With AttentionabstractLearning the hidden dynamics from sequence data is crucial. Attention mechanism can be introduced to spotlight on the region of interest for sequential learning. Traditional attention was measured between a query and a sequence based on a discrete-time state trajectory. Such a mechanism could not characterize the irregularly-sampled sequence data. This paper presents an attentive differential network (ADN) where the attention over continuous-time dynamics is developed. The continuous-time attention is performed over the dynamics at all time. The missing information in irregular or sparse samples can be seamlessly compensated and attended. Self attention is computed to find the attended state trajectory. However, the memory cost for attention score between a query and a sequence is demanding since self attention treats all time instants as query points in an ordinary differential equation solver. This issue is tackled by imposing the causality constraint in causal ADN (CADN) where the query is merged up to current time. To enhance the model robustness, this study further explores a latent CADN where the attended dynamics are calculated in an encoder-decoder structure via Bayesian learning. Experiments on the irregularly-sampled actions, dialogues and bio-signals illustrate the merits of the proposed methods in action recognition, emotion recognition and mortality prediction, respectively. Jen-Tzung Chien, Yi-Hsiang Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Bayesian asymmetric quantized neural networksabstractThis paper develops a robust model compression for neural networks via parameter quantization. Traditionally, quantized neural networks (QNN) were constructed by binary or ternary weights where the weights were deterministic. This paper generalizes QNN in two directions. First, M-ary QNN is developed to adjust the balance between memory storage and model capacity. The representation values and the quantization partitions in M-ary quantization are mutually estimated to enhance the resolution of gradients in neural network training. A flexible quantization with asymmetric partitions is formulated. Second, the variational inference is incorporated to implement the Bayesian asymmetric QNN. The uncertainty of weights is faithfully represented to enhance the robustness of the trained model in presence of heterogeneous data. Importantly, the multiple spike-and-slab prior is proposed to represent the quantization levels in Bayesian asymmetric learning. M-ary quantization is then optimized by maximizing the evidence lower bound of classification network. An adaptive parameter space is built to implement Bayesian quantization and neural representation. The experiments on various image recognition tasks show that M-ary QNN achieves similar performance as the full-precision neural network (FPNN), but the memory cost and the test time are significantly reduced relative to FPNN. The merit of Bayesian M-ary QNN using multiple spike-and-slab prior is investigated. Jen-Tzung Chien, Su-Ting Chang |
Pattern Recognit. | 1 |
| 2023 | Hierarchical Reinforcement Learning With Guidance for Multi-Domain Dialogue PolicyabstractAchieving high performance in a multi-domain dialogue system with low computation is undoubtedly challenging. Previous works applying an end-to-end approach have been very successful. However, the computational cost remains a major issue since the large-sized language model using GPT-2 is required. Meanwhile, the optimization for individual components in the dialogue system has not shown promising result, especially for the component of dialogue management due to the complexity of multi-domain state and action representation. To cope with these issues, this article presents an efficient guidance learning where the imitation learning and the hierarchical reinforcement learning (HRL) with human-in-the-loop are performed to achieve high performance via an inexpensive dialogue agent. The behavior cloning with auxiliary tasks is exploited to identify the important features in latent representation. In particular, the proposed HRL is designed to treat each goal of a dialogue with the corresponding sub-policy so as to provide efficient dialogue policy learning by utilizing the guidance from human through action pruning and action evaluation, as well as the reward obtained from the interaction with the simulated user in the environment. Experimental results on ConvLab-2 framework show that the proposed method achieves state-of-the-art performance in dialogue policy optimization and outperforms the GPT-2 based solutions in end-to-end system evaluation. Mahdin Rohmatillah, Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Augmentation Strategy Optimization for Language UnderstandingabstractThis paper presents a new language processing and understanding where an adaptive data augmentation strategy for individual documents is proposed instead of using one universal policy for the whole dataset. Importantly, a reinforcement learning and understanding method is exploited for document classification where the document encoder, augmenter and classifier are jointly optimized. In particular, a new reward function based on the consistency loss maximization is presented to assure the diversity of the generated documents. Using this method, the reward for adaptive augmentation policy is immediately calculated for every augmented instance without the need of waiting the child model performance metrics as the reward. The experiments on various classification tasks with a strong baseline model show that the augmentation strategy optimization can improve the model training process by providing meaningful augmentation data which eventually result in desirable evaluation performance. Furthermore, the extensive studies on the behavior of policy in different settings are provided in order to assure the diversity of the augmented data that was obtained by the proposed method. Chang-Ting Chu, Mahdin Rohmatillah, Ching-Hsien Lee, Jen-Tzung Chien |
ICASSP | 4 |
| 2022 | Adversarial Mask Transformer for Sequential LearningabstractMask language model has been successfully developed to build a transformer for robust language understanding. The transformer-based language model has achieved excellent results in various downstream applications. However, typical mask language model is trained by predicting the randomly masked words and is used to transfer the knowledge from rich-resource pre-training task to low-resource downstream tasks. This study incorporates a rich contextual embedding from pre-trained model and strengthens the attention layers for sequence-to-sequence learning. In particular, an adversarial mask mechanism is presented to deal with the shortcoming of random mask and accordingly enhance the robustness in word prediction for language understanding. The adversarial mask language model is trained in accordance with a minimax optimization over the word prediction loss. The worst-case mask is estimated to build an optimal and robust language model. The experiments on two machine translation tasks show the merits of the adversarial mask transformer. Hou Lio, Shang-En Li, Jen-Tzung Chien |
ICASSP | 3 |
| 2022 | Bayesian Transformer Using Disentangled Mask Attention
Jen-Tzung Chien, Yu-Han Huang |
INTERSPEECH | 1 |
| 2022 | Hierarchical and Self-Attended Sequence AutoencoderabstractIt is important and challenging to infer stochastic latent semantics for natural language applications. The difficulty in stochastic sequential learning is caused by the posterior collapse in variational inference. The input sequence is disregarded in the estimated latent variables. This paper proposes three components to tackle this difficulty and build the variational sequence autoencoder (VSAE) where sufficient latent information is learned for sophisticated sequence representation. First, the complementary encoders based on a long short-term memory (LSTM) and a pyramid bidirectional LSTM are merged to characterize global and structural dependencies of an input sequence, respectively. Second, a stochastic self attention mechanism is incorporated in a recurrent decoder. The latent information is attended to encourage the interaction between inference and generation in an encoder-decoder training procedure. Third, an autoregressive Gaussian prior of latent variable is used to preserve the information bound. Different variants of VSAE are proposed to mitigate the posterior collapse in sequence modeling. A series of experiments are conducted to demonstrate that the proposed individual and hybrid sequence autoencoders substantially improve the performance for variational sequential learning in language modeling and semantic understanding for document classification and summarization. Jen-Tzung Chien, Chun-Wei Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Contrastive Adversarial Domain Adaptation Networks for Speaker RecognitionabstractDomain adaptation aims to reduce the mismatch between the source and target domains. A domain adversarial network (DAN) has been recently proposed to incorporate adversarial learning into deep neural networks to create a domain-invariant space. However, DAN's major drawback is that it is difficult to find the domain-invariant space by using a single feature extractor. In this article, we propose to split the feature extractor into two contrastive branches, with one branch delegating for the class-dependence in the latent space and another branch focusing on domain-invariance. The feature extractor achieves these contrastive goals by sharing the first and last hidden layers but possessing decoupled branches in the middle hidden layers. For encouraging the feature extractor to produce class-discriminative embedded features, the label predictor is adversarially trained to produce equal posterior probabilities across all of the outputs instead of producing one-hot outputs. We refer to the resulting domain adaptation network as "contrastive adversarial domain adaptation network (CADAN)." We evaluated the embedded features' domain-invariance via a series of speaker identification experiments under both clean and noisy conditions. Results demonstrate that the embedded features produced by CADAN lead to a 33% improvement in speaker identification accuracy compared with the conventional DAN. Longxin Li, Man-Wai Mak, Jen-Tzung Chien |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Continuous-Time Attention for Sequential LearningabstractAttention mechanism is crucial for sequential learning where a wide range of applications have been successfully developed. This mechanism is basically trained to spotlight on the region of interest in hidden states of sequence data. Most of the attention methods compute the attention score through relating between a query and a sequence where the discrete-time state trajectory is represented. Such a discrete-time attention could not directly attend the continuous-time trajectory which is represented via neural differential equation (NDE) combined with recurrent neural network. This paper presents a new continuous-time attention method for sequential learning which is tightly integrated with NDE to construct an attentive continuous-time state machine. The continuous-time attention is performed at all times over the hidden states for different kinds of irregular time signals. The missing information in sequence data due to sampling loss, especially in presence of long sequence, can be seamlessly compensated and attended in learning representation. The experiments on irregular sequence samples from human activities, dialogue sentences and medical features show the merits of the proposed continuous-time attention for activity recognition, sentiment classification and mortality prediction, respectively. Jen-Tzung Chien, Yi-Hsiang Chen |
AAAI | 1 |
| 2021 | Variational Sequential Modeling, Learning and UnderstandingabstractNormalizing flow comprises of a series of invertible transformations. With careful design in transformations, it can generate images or speeches with fast sampling speed. Inference can also be efficient in maximum likelihood manner. In addition to generating scenes or human faces, it can be used to transform probability distributions. On the other hand, learning latent structures of sentences in a global manner is always challenging. Variational autoencoder (VAE) is haunted by the issue of posterior collapse, where the latent space is poorly learned. To improve inference and generation of VAE in learning sequence data, we propose the amortized flow posterior variational recurrent autoencoder (AFP-VRAE). Variational recurrent autoencoder (VRAE) has RNN based encoder and decoder and learns global representations of sentences. To learn latent space that well preserves the semantic information of data, we use the normalizing flow to generate flexible variational distributions. Furthermore, we adopt the amortized regularization to encode similar embeddings to neighboring latent representations, and we use the skip connections to reinforce the representations to predict every output directly. The benefits can be shown in the experiments as we evaluate the models for language modeling, sentiment analysis and document summarization. AFP-VRAE reports good results on variational modeling for sequence data. Jen-Tzung Chien, Chih-Jung Tsai |
ASRU | 1 |
| 2021 | Multitask Generative Adversarial Imitation Learning for Multi-Domain Dialogue SystemabstractIn the task-oriented dialogue system, dialog policy plays an important role since it determines the suitable actions based on the user's goals. However, in real situations, user's goals are varying so that the system needs to deal with the complex optimization problem for dialog policy. This paper presents a novel approach to build the multi-domain dialog system based on the multitask generative adversarial imitation learning (MGAIL). MGAIL combines hierarchical reinforcement learning and generative adversarial imitation learning where a mixture of generators are represented for multitask learning. Unlike the traditional imitation learning, this method decomposes each of complex tasks into several subtasks and builds the policy in a hierarchical way to relax the agent in handling multiple complex tasks. Experiments on a multi-domain dialogue system using MultiWOZ 2.1 under ConvLab-2 frame-work show that the proposed method outperforms the other reinforcement learning methods in system-wise evaluation in terms of complete rate, success rate and book rate. Chuan-En Hsu, Mahdin Rohmatillah, Jen-Tzung Chien |
ASRU | 3 |
| 2021 | Corrective Guidance and Learning for Dialogue ManagementabstractEstablishing robust dialogue policy with low computation cost is challenging, especially for multi-domain task-oriented dialogue management due to the high complexity in state and action spaces. The previous works mostly using the deterministic policy optimization only attain moderate performance. Meanwhile, state-of-the-art result that uses end-to-end approach is computationally demanding since it utilizes a large-scaled language model based on the generative pre-trained transformer-2 (GPT-2). In this study, a new learning procedure consisting of three learning stages is presented to improve multi-domain dialogue management with corrective guidance. Firstly, the behavior cloning with an auxiliary task is developed to build a robust pre-trained model by mitigating the causal confusion problem in imitation learning. Next, the pre-trained model is rectified by using reinforcement learning via the proximal policy optimization. Lastly, human-in-the-loop learning strategy is fulfilled to enhance the agent performance by directly providing corrective feedback from rule-based agent so that the agent is prevented to trap in confounded states. The experiments on end-to-end evaluation show that the proposed learning method achieves state-of-the-art result by performing nearly identical to the rule-based agent. This method outperforms the second place of 9th dialog system technology challenge (DSTC9) track 2 that uses GPT-2 as the core model in dialogue management. Mahdin Rohmatillah, Jen-Tzung Chien |
CIKM | 2 |
| 2021 | Continuous-Time Self-Attention in Neural Differential EquationabstractNeural differential equation (NDE) is recently developed as a continuous-time state machine which can faithfully represent the irregularly-sampled sequence data. NDE is seen as a substantial extension of recurrent neural network (RNN) which conducts discrete-time state representation for regularly-sampled data. This study presents a new continuous-time attention to improve sequential learning where the region of interest in continuous-time state trajectory over observed as well as missing samples is sufficiently attended. However, the attention score, calculated by relating between a query and a sequence, is memory demanding because self-attention should treat all time observations as query vectors to feed them into ordinary differential equation (ODE) solver. To deal with this issue, we develop a new form of dynamics for continuous-time attention where the causality property is adopted such that query vector is fed into ODE solver up to current time. The experiments on irregularly-sampled human activities and medical features show that this method obtains desirable performance with efficient memory consumption. Jen-Tzung Chien, Yi-Hsiang Chen |
ICASSP | 1 |
| 2021 | Dualformer: A Unified Bidirectional Sequence-to-Sequence LearningabstractThis paper presents a new dual domain mapping based on a unified bidirectional sequence-to-sequence (seq2seq) learning. Traditionally, dual learning in domain mapping was constructed with intrinsic connection where the conditional generative models in two directions were mutually leveraged and combined. The additional feedback from the other generation direction was used to regularize sequential learning in original direction of domain mapping. Domain matching between source sequence and target sequence was accordingly improved. However, the reconstructions for knowledge in two domains were ignored. The dual information based on separate models in two training directions was not sufficiently discovered. To cope with this weakness, this study proposes a closed-loop seq2seq learning where domain mapping and domain knowledge are jointly learned. In particular, a new feature-level dual learning is incorporated to build a dualformer where feature integration and feature reconstruction are further performed to bridge dual tasks. Experiments demonstrate the merit of the proposed dualformer for machine translation based on the multi-objective seq2seq learning. Jen-Tzung Chien, Wei-Hsiang Chang |
ICASSP | 1 |
| 2021 | Attribute Decomposition for Flow-Based Domain MappingabstractDomain mapping aims to estimate a sophisticated mapping between source and target domains. Finding the specialized attribute in latent representation plays a key role to attain a desirable performance. However, the entangled features usually contain the mixed attribute which can not be easily decomposed in an unsupervised manner. To handle the mixed features for better generation, this paper presents an attribute decomposition based on the sequence data and carries out the flow-based image domain mapping. The latent variables, characterized by flow model, are decomposed into the attribute-relevant and attribute-irrelevant components. The decomposition is guided by multiple objectives including structural-perceptual loss, cycle consistency loss, sequential random-pair reconstruction loss and sequential classification loss where the paired training data for domain mapping are not required. Importantly, the sequential random-pair reconstruction loss is formulated by means of exchanging the attribute-relevant components within a sequence of images. As a result, the source images with the attributes of reference images can be smoothly transferred to the corresponding target images. Experiments on talking face synthesis show the merit of attribute decomposition in domain mapping. Sheng-Jhe Huang, Jen-Tzung Chien |
ICASSP | 2 |
| 2021 | Variational Dialogue Generation with Normalizing FlowsabstractConditional variational autoencoder (cVAE) has shown promising performance in dialogue generation. However, there still exists two issues in dialog cVAE model. The first issue is the Kullback-Leiblier (KL) vanishing problem which results in degenerating cVAE into a simple recurrent neural network. The second issue is the assumption of isotropic Gaussian prior for latent variable which is too simple to assure diversity of the generated responses. To handle these issues, a simple distribution should be transformed into a complex distribution and simultaneously the value of KL divergence should be preserved. This paper presents the dialogue flow VAE (DF-VAE) for variational dialogue generation. In particular, KL vanishing is tackled by a new normalizing flow. An inverse autoregressive flow is proposed to transform isotropic Gaussian prior to a rich distribution. In the experiments, the proposed DF-VAE is significantly better than the other methods in terms of different evaluation metrics. The diversity of generated dialogue responses is enhanced. Ablation study is conducted to illustrate the merit of the proposed flow models. Tien-Ching Luo, Jen-Tzung Chien |
ICASSP | 2 |
| 2021 | Collaborative Regularization for Bidirectional Domain MappingabstractLearning both domain mapping and domain knowledge is crucial for different sequence-to-sequence (seq2seq) tasks. Traditionally, seq2seq model only characterized domain mapping while the knowledge in source and target domains was ignored. To strengthen seq2seq representation, this study presents a unified transformer for bidirectional domain mapping where collaborative regularization is imposed. This regularization enforces the bidirectional mapping constraint and avoids the model from overfitting for better generalization. Importantly, the unified learning objective is optimized for collaborative learning among different modules in two domains with two learning directions. Experiments on machine translation demonstrate the merit of unified transformer by comparing with the existing methods under different tasks and settings. Jen-Tzung Chien, Wei-Hsiang Chang |
IJCNN | 1 |
| 2021 | Stochastic Temporal Difference Learning for Sequence DataabstractPlanning is crucial to train an agent via model-based reinforcement learning who can predict distant observations to reflect his/her past experience. Such a planning method is theoretically and computationally attractive in comparison with traditional learning which relies on step-by-step prediction. However, it is more challenging to build a learning machine which can predict and plan randomly across multiple time steps rather than act step by step. To reflect this flexibility in learning process, we need to predict future states directly without going through all intermediate states. Accordingly, this paper develops the stochastic temporal difference learning where the sequence data are represented with multiple jumpy states while the stochastic state space model is learned by maximizing the evidence lower bound of log likelihood of training data. A general solution with various number of jumpy states is developed and formulated. Experiments demonstrate the merit of the proposed sequential machine to find predictive states to roll forward with jumps as well as predict words. Jen-Tzung Chien, Yi-Chung Chiu |
IJCNN | 1 |
| 2021 | Online Compressive Transformer for End-to-End Speech Recognition
Chi-Hang Leong, Yu-Han Huang, Jen-Tzung Chien |
Interspeech | 3 |
| 2021 | Causal Confusion Reduction for Robust Multi-Domain Dialogue Policy
Mahdin Rohmatillah, Jen-Tzung Chien |
Interspeech | 2 |
| 2020 | Neural Bayesian Information ProcessingabstractDeep learning is developed as a learning process from source inputs to target outputs where the inference or optimization is performed over an assumed deterministic model with deep structure. A wide range of temporal and spatial data in language and vision are treated as the inputs or outputs to build such a complicated mapping in different information systems. A systematic and elaborate transfer is required to meet the mapping between source and target domains. Also, the semantic structure in natural language and computer vision may not be well represented or trained in mathematical logic or computer programs. The distribution function in discrete or continuous latent variable model for words, sentences, images or videos may not be properly decomposed or estimated. The system robustness to heterogeneous environments may not be assured. This tutorial addresses the fundamentals and advances in statistical models and neural networks, and presents a series of deep Bayesian solutions including variational Bayes, sampling method, Bayesian neural network, variational auto-encoder (VAE), stochastic recurrent neural network, sequence-to-sequence model, attention mechanism, end-to-end network, stochastic temporal convolutional network, temporal difference VAE, normalizing flow and neural ordinary differential equation. Enhancing the prior/posterior representation is addressed in different latent variable models. We illustrate how these models are connected and why they work for a variety of applications on complex patterns in language and vision. The word, sentence and image embeddings are merged with semantic constraint or structural information. Bayesian learning is formulated in the optimization procedure where the posterior collapse is tackled. An informative latent space is trained to incorporate deep Bayesian learning in various information systems. Jen-Tzung Chien |
CIKM | 1 |
| 2020 | Information Maximized Variational Domain Adversarial Learning for Speaker VerificationabstractDomain mismatch is a common problem in speaker verification. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) to reduce domain mismatch by incorporating an InfoVAE into domain adversarial training (DAT). DAT aims to produce speaker discriminative and domain-invariant features. The InfoVAE has two roles. First, it performs variational regularization on the learned features so that they follow a Gaussian distribution, which is essential for the standard PLDA backend. Second, it preserves mutual information between the features and the training set to extract extra speaker discriminative information. Experiments on both SRE16 and SRE18-CMN2 show that the InfoVDANN outperforms the recent VDANN, which suggests that increasing the mutual information between the latent features and input features enables the InfoVDANN to extract extra speaker information that is otherwise not possible. Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien |
ICASSP | 3 |
| 2020 | M-ARY Quantized Neural NetworksabstractParameter quantization is crucial for model compression. This paper generalizes the binary and ternary quantizations to M-ary quantization for adaptive learning of the quantized neural networks. To compensate the performance loss, the representation values and the quantization partitions of model parameters are jointly trained to optimize the resolution of gradients for parameter updating where the non-differentiable function in back-propagation algorithm is tackled. An asymmetric quantization is implemented. The restriction in parameter quantization is sufficiently relaxed. The resulting M-ary quantization scheme is general and adaptive with different M. Training of the M-ary quantized neural network (MQNN) can be tuned to balance the tradeoff between system performance and memory storage. Experimental results show that MQNN is able to achieve comparable image classification performance with full-precision neural network (FPNN), but the memory storage can be far less than that in FPNN. Jen-Tzung Chien, Su-Ting Chang |
ICME | 1 |
| 2020 | Stochastic Convolutional Recurrent NetworksabstractRecurrent neural network (RNN) has been widely used for sequential learning which has achieved a great success in different tasks. The temporal convolutional network (TCN), a variant of one-dimensional convolutional neural network (CNN), was also developed for sequential learning in presence of sequence data. RNN and TCN typically captures long-term and short-term features in temporal or spatial domain, respectively. This paper presents a new sequential learning, called the convolutional recurrent network (CRN), which fulfills TCN as an encoder and RNN as a decoder so that the global semantics as well as the local dependencies are simultaneously characterized from sequence data. To facilitate the interpretation and robustness in neural models, we further develop the stochastic modeling for CRN based on variational inference. The merits of CNN and RNN are then incorporated in inference of latent space which sufficiently produces a generative model for sequential prediction. Experiments on language model shows the effectiveness of stochastic CRN when compared with the other sequential machines. Jen-Tzung Chien, Yu-Min Huang |
IJCNN | 1 |
| 2020 | Stochastic Curiosity Maximizing ExplorationabstractDeep reinforcement learning (RL) is known as an emerging research trend in machine learning for autonomous systems. In real-world scenarios, the extrinsic rewards, acquired from the environment for learning an agent, are usually missing or extremely sparse. Such an issue of sparse reward constrains the learning capability of agent because the agent only updates the policy when the goal state is successfully attained. It is always challenging to implement an efficient exploration in RL algorithms. To tackle the sparse reward and inefficient exploration, the agent needs other helpful information to update its policy even when there is no interaction with the environment. This paper proposes the stochastic curiosity maximizing exploration (SCME), a learning strategy explored to allow the agent to act as human. We cope with the sparse reward problem by encouraging the agent to explore future diversity. To do so, a latent dynamic system is developed to acquire the latent states and latent actions to predict the variations in future conditions. The mutual information and the prediction error in the predicted states and actions are calculated as the intrinsic rewards. The agent based on SCME is therefore learned by maximizing these rewards to improve sample efficiency for exploration. The experiments on PyDial and Super Mario Bros show the benefits of the proposed SCME in dialogue system and computer game, respectively. Jen-Tzung Chien, Po-Chien Hsu |
IJCNN | 1 |
| 2020 | Stochastic Adversarial Learning for Domain AdaptationabstractLearning across domains is challenging especially when test data in target domain are sparse, heterogeneous and unlabeled. This challenge is even severe when building a deep stochastic neural model. This paper presents a stochastic semi-supervised learning for domain adaptation by using labeled data from source domain and unlabeled data from target domain. There are twofold novelties in the proposed method. First, a graphical model is constructed to identify the random latent features for classes as well as domains which are learned by variational inference. Second, we learn the class features which are discriminative among classes and simultaneously invariant to both domains. An adversarial neural model is introduced to pursue domain invariance. The domain features are explicitly learned to purify the extraction of class features for an improved classification. The experiments on sentiment classification illustrate the merits of the proposed stochastic adversarial domain adaptation. Jen-Tzung Chien, Ching-Wei Huang |
IJCNN | 1 |
| 2020 | Amortized Mixture Prior for Variational Sequence GenerationabstractVariational autoencoder (VAE) is a popular latent variable model for data generation. However, in natural language applications, VAE suffers from the posterior collapse in optimization procedure where the model posterior likely collapses to a standard Gaussian prior which disregards latent semantics from sequence data. The recurrent decoder accordingly generates du-plicate or noninformative sequence data. To tackle this issue, this paper adopts the Gaussian mixture prior for latent variable, and simultaneously fulfills the amortized regularization in encoder and skip connection in decoder. The noise robust prior, learned from the amortized encoder, becomes semantically meaningful. The prediction of sequence samples, due to skip connection, becomes contextually precise at each time. The amortized mixture prior (AMP) is then formulated in construction of variational recurrent autoencoder (VRAE) for sequence generation. Experiments on different tasks show that AMP-VRAE can avoid the posterior collapse, learn the meaningful latent features and improve the inference and generation for semantic representation. Jen-Tzung Chien, Chih-Jung Tsai |
IJCNN | 1 |
| 2020 | Stochastic Convolutional Recurrent Networks for Language Modeling
Jen-Tzung Chien, Yu-Min Huang |
INTERSPEECH | 1 |
| 2020 | Stochastic Curiosity Exploration for Dialogue Systems
Jen-Tzung Chien, Po-Chien Hsu |
INTERSPEECH | 1 |
| 2020 | Strategies for End-to-End Text-Independent Speaker Verificationabstract21st Annual Conference of the International Speech Communication Association, INTERSPEECH 2020, 25-29 October 2020, Shanghai, China Weiwei Lin 0002, Man-Wai Mak, Jen-Tzung Chien |
INTERSPEECH | 3 |
| 2020 | Deep Bayesian Multimedia LearningabstractDeep learning has been successfully developed as a complicated learning process from source inputs to target outputs in presence of multimedia environments. The inference or optimization is performed over an assumed deterministic model with deep structure. A wide range of temporal and spatial data in language and vision are treated as the inputs or outputs to build such a domain mapping for multimedia applications. A systematic and elaborate transfer is required to meet the mapping between source and target domains. Also, the semantic structure in natural language and computer vision may not be well represented or trained in mathematical logic or computer programs. The distribution function in discrete or continuous latent variable model for words, sentences, images or videos may not be properly decomposed or estimated. The system robustness to heterogeneous environments may not be assured. This tutorial addresses the fundamentals and advances in statistical models and neural networks for domain mapping, and presents a series of deep Bayesian solutions including variational Bayes, sampling method, Bayesian neural network, variational auto-encoder (VAE), stochastic recurrent neural network, sequence-to-sequence model, attention mechanism, end-to-end network, stochastic temporal convolutional network, temporal difference VAE, normalizing flow and neural ordinary differential equation. Enhancing the prior/posterior representation is addressed in different latent variable models. We illustrate how these models are connected and why they work for a variety of applications on complex patterns in language and vision. The word, sentence and image embeddings are merged with semantic constraint or structural information. Bayesian learning is formulated in the optimization procedure where the posterior collapse is tackled. An informative latent space is trained to incorporate deep Bayesian learning in various information systems. Jen-Tzung Chien |
ACM Multimedia | 1 |
| 2020 | Deep Bayesian Data MiningabstractThis tutorial addresses the fundamentals and advances in deep Bayesian mining and learning for natural language with ubiquitous applications ranging from speech recognition to document summarization, text classification, text segmentation, information extraction, image caption generation, sentence generation, dialogue control, sentiment classification, recommendation system, question answering and machine translation, to name a few. Traditionally, "deep learning" is taken to be a learning process where the inference or optimization is based on the real-valued deterministic model. The "semantic structure" in words, sentences, entities, actions and documents drawn from a large vocabulary may not be well expressed or correctly optimized in mathematical logic or computer programs. The "distribution function" in discrete or continuous latent variable model for natural language may not be properly decomposed or estimated. This tutorial addresses the fundamentals of statistical models and neural networks, and focus on a series of advanced Bayesian models and deep models including hierarchical Dirichlet process, Chinese restaurant process, hierarchical Pitman-Yor process, Indian buffet process, recurrent neural network (RNN), long short-term memory, sequence-to-sequence model, variational auto-encoder (VAE), generative adversarial network (GAN), attention mechanism, memory-augmented neural network, skip neural network, temporal difference VAE, stochastic neural network, stochastic temporal convolutional network, predictive state neural network, and policy neural network. Enhancing the prior/posterior representation is addressed. We present how these models are connected and why they work for a variety of applications on symbolic and complex patterns in natural language. The variational inference and sampling method are formulated to tackle the optimization for complicated models. The word and sentence embeddings, clustering and co-clustering are merged with linguistic and semantic constraints. A series of case studies, tasks and applications are presented to tackle different issues in deep Bayesian mining, searching, learning and understanding. At last, we will point out a number of directions and outlooks for future studies. This tutorial serves the objectives to introduce novices to major topics within deep Bayesian learning, motivate and explain a topic of emerging importance for data mining and natural language understanding, and present a novel synthesis combining distinct lines of machine learning work. Jen-Tzung Chien |
WSDM | 1 |
| 2020 | Variational Domain Adversarial Learning With Mutual Information Maximization for Speaker VerificationabstractDomain mismatch is a common problem in speaker verification (SV) and often causes performance degradation. For the system relying on the Gaussian PLDA backend to suppress the channel variability, the performance would be further limited if there is no Gaussianity constraint on the learned embeddings. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) that incorporates an InfoVAE into domain adversarial training (DAT) to reduce domain mismatch and simultaneously meet the Gaussianity requirement of the PLDA backend. Specifically, DAT is applied to produce speaker discriminative and domain-invariant features, while the InfoVAE performs variational regularization on the embedded features so that they follow a Gaussian distribution. Another benefit of the InfoVAE is that it avoids posterior collapse in VAEs by preserving the mutual information between the embedded features and the training set so that extra speaker information can be retained in the features. Experiments on both SRE16 and SRE18-CMN2 show that the InfoVDANN outperforms the recent VDANN, which suggests that increasing the mutual information between the embedded features and input features enables the InfoVDANN to extract extra speaker information that is otherwise not possible. Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Bayesian Adversarial Learning for Speaker RecognitionabstractThis paper presents a new generative adversarial network (GAN) which artificially generates the i-vectors to compensate the imbalanced or insufficient data in speaker recognition based on the probabilistic linear discriminant analysis. Theoretically, GAN is powerful to generate the artificial data which are misclassified as the real data. However, GAN suffers from the mode collapse problem in two-player optimization over generator and discriminator. This study deals with this challenge by improving the model regularization through characterizing the weight uncertainty in GAN. A new Bayesian GAN is implemented to learn a regularized model from diverse data where the strong modes are flattened via the marginalization. In particular, we present a variational GAN (VGAN) where the encoder, generator and discriminator are jointly estimated according to the variational inference. The computation cost is significantly reduced. To assure the preservation of gradient values, the learning objective based on Wasserstein distance is further introduced. The issues of model collapse and gradient vanishing are alleviated. Experiments on NIST i-vector Speaker Recognition Challenge demonstrate the superiority of the proposed VGAN to the variational autoencoder, the standard GAN and the Bayesian GAN based on the sampling method. The learning efficiency and generation performance are evaluated. Jen-Tzung Chien, Chun Lin Kuo |
ASRU | 1 |
| 2019 | Markov Recurrent Neural Network Language ModelabstractRecurrent neural network (RNN) has achieved a great success in language modeling where the temporal information based on deterministic state is continuously extracted and evolved through time. Such a simple deterministic transition function using input-to-hidden and hidden-to-hidden weights is usually insufficient to reflect the diversities and variations of latent variable structure behind the heterogeneous natural language. This paper presents a new stochastic Markov RNN (MRNN) to strengthen the learning capability in language model where the trajectory of word sequences is driven by a neural Markov process with Markov state transitions based on a K-state long short-term memory model. A latent state machine is constructed to characterize the complicated semantics in the structured lexical patterns. Gumbel-softmax is introduced to implement the stochastic backpropatation algorithm with discrete states. The parallel computation for rapid realization of MRNN is presented. The variational Bayesian learning procedure is implemented. Experiments demonstrate the merits of stochastic and diverse representation using MRNN language model where the overhead of parameters and computations is limited. Jen-Tzung Chien, Che-Yu Kuo |
ASRU | 1 |
| 2019 | Stochastic Markov Recurrent Neural Network for Source SeparationabstractMonaural source separation based on recurrent neural network is learned to characterize the sequential patterns in source signals based on dynamic states which are propagated through time. The hidden states are assumed to be deterministic along a single path where a shared long short-term memory (LSTM) is used. Such assumptions may not faithfully reflect the randomness and the variety of temporal features in mixed signals. To strengthen the capability of LSTM in source separation, we propose a stochastic Markov LSTM where the regression from the mixed signal to its source signals is learned with a stochastic indicator of Markov state which selects the state-dependent LSTM for signal separation at each time. A set of LSTMs is discovered to capture the structural diversity of temporal signals or the stochastic trajectory of state transitions for sequential prediction. A new state machine is constructed to learn the complicated latent semantics in heterogeneous and structural mappings between mixed signals and source signals. The Gumbel-softmax sampling is implemented in the backpropagation algorithm with discrete Markov states. Experiments on speech enhancement illustrate the merit of the proposed stochastic Markov LSTM in terms of short-term objective intelligibility measure of the separated speech. Jen-Tzung Chien, Che-Yu Kuo |
ICASSP | 1 |
| 2019 | Variational and Hierarchical Recurrent AutoencoderabstractDespite a great success in learning representation for image data, it is challenging to learn the stochastic latent features from natural language based on variational inference. The difficulty in stochastic sequential learning is due to the posterior collapse caused by an autoregressive decoder which is prone to be too strong to learn sufficient latent information during optimization. To compensate this weakness in learning procedure, a sophisticated latent structure is required to assure good convergence so that random features are sufficiently captured for sequential decoding. This study presents a new variational recurrent autoencoder (VRAE) for sequence reconstruction. There are two complementary encoders consisting of a long short-term memory (LSTM) and a pyramid bidirectional LSTM which are merged to discover the global and local dependencies in a hierarchical latent variable model, respectively. Experiments on Penn Treebank and Yelp 2013 demonstrate that the proposed hierarchical VRAE is able to learn the complementary representation as well as tackle the posterior collapse in stochastic sequential learning. The performance of recurrent autoencoder is substantially improved in terms of perplexity. Jen-Tzung Chien, Chun-Wei Wang |
ICASSP | 1 |
| 2019 | Semi-supervised Nuisance-attribute Networks for Domain AdaptationabstractHow to overcome the training and test data mismatch in speaker verification systems has been a focus of research recently. In this paper, we propose a semi-supervised nuisance attribute network (SNAN) to reduce the domain mismatch in i-vectors and x-vectors. SNANs are based on the idea of nuisance attribute removal in inter-dataset variability compensation (IDVC). But instead of measuring the domain variability through the dataset means, SNANs use the maximum mean discrepancy (MMD) as part of their loss function, which enables the network to find nuisance directions in which domain variability is measured up to infinite moment. The architecture of SNANs also allows us to incorporate the out-of-domain speaker labels into the semi-supervised training process through the center loss and triplet loss. Using SNANs as a preprocessing step for PLDA training, we achieve a relative improvement of 11.8% in EER on NIST 2016 SRE compared to PLDA without adaptation. We also found that the semi-supervised approach can further improve SNANs' performance. Weiwei Lin 0002, Man-Wai Mak, Youzhi Tu, Jen-Tzung Chien |
ICASSP | 4 |
| 2019 | Meta Learning for Hyperparameter Optimization in Dialogue System
Jen-Tzung Chien, Wei Xiang Lieow |
INTERSPEECH | 1 |
| 2019 | Self Attention in Variational Sequential Learning for Summarization
Jen-Tzung Chien, Chun-Wei Wang |
INTERSPEECH | 1 |
| 2019 | Variational Domain Adversarial Learning for Speaker Verificationabstract20th Annual Conference of the International Speech Communication Association: Crossroads of Speech and Language, INTERSPEECH 2019, Graz, Austria, 15-19 September 2019 Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien |
INTERSPEECH | 3 |
| 2019 | Deep Bayesian Mining, Learning and UnderstandingabstractThis tutorial addresses the advances in deep Bayesian mining and learning for natural language with ubiquitous applications ranging from speech recognition to document summarization, text classification, text segmentation, information extraction, image caption generation, sentence generation, dialogue control, sentiment classification, recommendation system, question answering and machine translation, to name a few. Traditionally, "deep learning" is taken to be a learning process where the inference or optimization is based on the real-valued deterministic model. The "semantic structure" in words, sentences, entities, actions and documents drawn from a large vocabulary may not be well expressed or correctly optimized in mathematical logic or computer programs. The "distribution function" in discrete or continuous latent variable model for natural language may not be properly decomposed or estimated. This tutorial addresses the fundamentals of statistical models and neural networks, and focus on a series of advanced Bayesian models and deep models including hierarchical Dirichlet process, Chinese restaurant process, hierarchical Pitman-Yor process, Indian buffet process, recurrent neural network, long short-term memory, sequence-to-sequence model, variational auto-encoder, generative adversarial network, attention mechanism, memory-augmented neural network, skip neural network, stochastic neural network, predictive state neural network, policy neural network. We present how these models are connected and why they work for a variety of applications on symbolic and complex patterns in natural language. The variational inference and sampling method are formulated to tackle the optimization for complicated models. The word and sentence embeddings, clustering and co-clustering are merged with linguistic and semantic constraints. A series of case studies are presented to tackle different issues in deep Bayesian mining, learning and understanding. At last, we will point out a number of directions and outlooks for future studies. Jen-Tzung Chien |
KDD | 1 |
| 2019 | Neural adversarial learning for speaker recognition
Jen-Tzung Chien, Kang-Ting Peng |
Comput. Speech Lang. | 1 |
| 2019 | Image-text dual neural network with decision strategy for small-sample image classification
Fangyi Zhu, Zhanyu Ma, Guang Chen 0003, Jen-Tzung Chien, Jing-Hao Xue, Jun Guo 0002 |
Neurocomputing | 5 |
| 2018 | Spectro-Temporal Neural Factorization for Speech DereverberationabstractThis study presents a spectro-temporal neural factorization (STNF) for speech dereverberation. Traditionally, a contextual window of spectro-temporal reverberant speech was unfolded into a one-way vector which was fed into a neural network to estimate the spectra of source speech at each time frame. Model parameters were trained by using the vectorized error backpropagation algorithm. System performance is constrained because contextual correlations and common factors in frequency and time horizons are disregarded. To compensate this weakness, a spectro-temporal factorization is incorporated to preserve the structural information in neural network training based on bi-factorized error backpropagation where the spectral and temporal factor matrices are estimated. Affine transformation in one-way neural network is generalized to the bilinear decomposition in bi-factorized neural network. The spectro-temporal features are extracted and forwarded to fully-connected layers for regression outputs. Such a STNF is further improved by merging with long short-term memory layer to capture the temporal features. Experiments results on 2014 REVERB Challenge demonstrate the meaningfulness of the factorized features and the merit of integrating these features for speech dereverberation. Jen-Tzung Chien, Kuan-Ting Kuo |
ICASSP | 1 |
| 2018 | Recall Neural Network for Source SeparationabstractThis paper presents a novel memory-augmented neural network for single-channel source separation. We propose a recall neural network (RCNN) where a couple of external memories are realized for sequence-to-sequence learning based on an encoder and a decoder. These memories are learned in a two-pass sensing procedure where the mixed signal is encoded and then decoded (or recalled) as context vectors by using a bidirectional long short-term memory (LSTM) and a LSTM, respectively. These context vectors are integrated in a gating layer. A set of attention weights are calculated to attend the hidden state of decoder to implement a recurrent neural network for source separation. A gated attention mechanism is carried out to fulfill a specialized memory network. The regression errors due to two passes of sensing procedure and one pass of gated attention are jointly minimized to estimate the weight parameters of different components in different layers. The experiments on multi-speaker speech enhancement show that the proposed RCNN consistently outperforms LSTM and neural Turing machine in different settings. Jen-Tzung Chien, Kai-Wei Tsou |
ICASSP | 1 |
| 2018 | Latent Dirichlet mixture model
Jen-Tzung Chien, Chao-Hsi Lee, Zheng-Hua Tan |
Neurocomputing | 1 |
| 2018 | Recent advances in machine learning for non-Gaussian data processing
Zhanyu Ma, Jen-Tzung Chien, Zheng-Hua Tan, Yi-Zhe Song, Jalil Taghia, Ming Xiao 0001 |
Neurocomputing | 2 |
| 2018 | Deep Unfolding for Topic ModelsabstractDeep unfolding provides an approach to integrate the probabilistic generative models and the deterministic neural networks. Such an approach is benefited by deep representation, easy interpretation, flexible learning and stochastic modeling. This study develops the unsupervised and supervised learning of deep unfolded topic models for document representation and classification. Conventionally, the unsupervised and supervised topic models are inferred via the variational inference algorithm where the model parameters are estimated by maximizing the lower bound of logarithm of marginal likelihood using input documents without and with class labels, respectively. The representation capability or classification accuracy is constrained by the variational lower bound and the tied model parameters across inference procedure. This paper aims to relax these constraints by directly maximizing the end performance criterion and continuously untying the parameters in learning process via deep unfolding inference (DUI). The inference procedure is treated as the layer-wise learning in a deep neural network. The end performance is iteratively improved by using the estimated topic parameters according to the exponentiated updates. Deep learning of topic models is therefore implemented through a back-propagation procedure. Experimental results show the merits of DUI with increasing number of layers compared with variational inference in unsupervised as well as supervised topic models. Jen-Tzung Chien, Chao-Hsi Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Bayesian Nonparametric Learning for Hierarchical and Sparse TopicsabstractThis paper presents the Bayesian nonparametric (BNP) learning for hierarchical and sparse topics from natural language. Traditionally, the Indian buffet process provides the BNP prior on a binary matrix for an infinite latent feature model consisting of a flat layer of topics. The nested model paves an avenue to construct a tree model instead of a flat-layer model. This paper presents the nested Indian buffet process (nIBP) to achieve the sparsity and flexibility in topic model where the model complexity and topic hierarchy are learned from the groups of words. The mixed membership modeling is conducted by representing a document using the tree nodes or dishes that a document or a customer chooses according to the nIBP scenario. A tree stick-breaking process is implemented to select topic weights from a subtree for flexible topic modeling. Such an nIBP relaxes the constraint of adopting a single tree path in the nested Chinese restaurant process (nCRP) and, therefore, improves the variety of topic representation for heterogeneous documents. A Gibbs sampling procedure is developed to infer the nIBP topic model. Compared to the nested hierarchical Dirichlet process (nhDP), the compactness of the estimated topics in a tree using nIBP is improved. Experimental results show that the proposed nIBP reduces the error rate of nCRP and nhDP by 18% and 8% on Reuters task for document classification, respectively. Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Multisource I-Vectors Domain Adaptation Using Maximum Mean Discrepancy Based AutoencodersabstractLike many machine learning tasks, the performance of speaker verification (SV) systems degrades when training and test data come from very different distributions. What's more, both training and test data themselves could be composed of heterogeneous subsets. These multisource mismatches are detrimental to SV performance. This paper proposes incorporating maximum mean discrepancy (MMD) into the loss function of autoencoders to reduce these mismatches. MMD is a nonparametric method for measuring the distance between two probability distributions. With a properly chosen kernel, MMD can match up to infinite moments of data distributions. We generalize MMD to measure the discrepancies of multiple distributions. We call the generalized MMD domainwise MMD. Using domainwise MMD as an objective function, we propose two autoencoders, namely nuisance-attribute autoencoder (NAE) and domain-invariant autoencoder (DAE), for multisource i-vector adaptation. NAE encodes the features that cause most of the multisource mismatch measured by domainwise MMD. DAE directly encodes the features that minimize the multisource mismatch. Using these MMD-based autoencoders as a preprocessing step for PLDA training, we achieve a relative improvement of 19.2% EER on the NIST 2016 SRE compared to PLDA without adaptation. We also found that MMD-based autoencoders are more robust to unseen domains. In the domain robustness experiments, MMD-based autoencoders show 6.8% and 5.2% improvements over IDVC on female and male Cantonese speakers, respectively. Weiwei Lin 0002, Man-Wai Mak, Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Tensor-Factorized Neural NetworksabstractThe growing interests in multiway data analysis and deep learning have drawn tensor factorization (TF) and neural network (NN) as the crucial topics. Conventionally, the NN model is estimated from a set of one-way observations. Such a vectorized NN is not generalized for learning the representation from multiway observations. The classification performance using vectorized NN is constrained, because the temporal or spatial information in neighboring ways is disregarded. More parameters are required to learn the complicated data structure. This paper presents a new tensor-factorized NN (TFNN), which tightly integrates TF and NN for multiway feature extraction and classification under a unified discriminative objective. This TFNN is seen as a generalized NN, where the affine transformation in an NN is replaced by the multilinear and multiway factorization for tensor-based NN. The multiway information is preserved through layerwise factorization. Tucker decomposition and nonlinear activation are performed in each hidden layer. The tensor-factorized error backpropagation is developed to train TFNN with the limited parameter size and computation time. This TFNN can be further extended to realize the convolutional TFNN (CTFNN) by looking at small subtensors through the factorized convolution. Experiments on real-world classification tasks demonstrate that TFNN and CTFNN attain substantial improvement when compared with an NN and a convolutional NN, respectively. Jen-Tzung Chien, Yi-Ting Bao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Adversarial manifold learning for speaker recognitionabstractThis paper presents an adversarial manifold learning (AML) for speaker recognition based on the probabilistic linear discriminant analysis (PLDA) using i-vectors. PLDA basically consists of an encoder for finding the latent variables and a decoder for reconstructing the i-vectors. AML is developed and incorporated in deep learning for a latent variable model. Low-dimensional latent space is therefore constructed according to an adversarial learning with neighbor embedding. This AML-PLDA is formulated to jointly optimize three learning objectives including a reconstruction error based on PLDA, a subspace learning for neighbor embedding and an adversarial loss caused by a discriminator and a generator. Using the deep neural networks, the generator is trained to fool the discriminator with its generated samples in latent space. The parameters in encoder, decoder and discriminator are jointly estimated by using the stochastic gradient descent algorithm. The experiments on speaker recognition show the merit of AML-PLDA in manifold learning and pattern classification. Jen-Tzung Chien, Kang-Ting Peng |
ASRU | 1 |
| 2017 | Variational manifold learning for speaker recognitionabstractThis paper presents a variational manifold learning for speaker recognition based on the probabilistic linear discriminant analysis (PLDA) using i-vectors. A latent variable model is introduced to compensate the constraints of the linearity in PLDA scoring and the high dimensionality in using i-vectors. A deep variational learning is formulated to jointly optimize three objectives including a regularization for variational distributions, a reconstruction based on PLDA and a manifold learning for neighbor embedding. A stochastic gradient variational Bayesian algorithm is developed to optimize the variational lower bound of log likelihood where the expectation in the objectives is estimated via a sampling method. Interestingly, the latent variables in the proposed variational manifold PLDA (vm-PLDA) are capable of decoding or reconstructing the i-vectors. The experiments on visualization and speaker recognition show the merits of vm-PLDA in manifold learning and classification. Jen-Tzung Chien, Cheng-Wei Hsu |
ICASSP | 1 |
| 2017 | Power-law stochastic neighbor embeddingabstractStochastic neighbor embedding (SNE) aims to transform the observations in high-dimensional space into a low-dimensional space which preserves neighbor identities by minimizing the Kullback-Leibler divergence of the pairwise distributions between two spaces where Gaussian distributions are assumed. Data visualization could be improved by adopting the t-SNE where Student t distribution is used in the low-dimensional space. However, data pairs in the latent space are forced to be squeezed due to the loss of dimensions. This study incorporates the power-law distribution into construction of the p-SNE. Such an unsupervised p-SNE increases the physical forces in neighbor embedding so that the neighbors in the low-dimensional space can be adjusted flexibly to reflect the neighboring in the high-dimensional space. The experiments on three learning tasks illustrate that the manifold or data structure using the proposed p-SNE is preserved in better shape than that using SNE and t-SNE. Huan-Hsin Tseng, Issam El-Naqa, Jen-Tzung Chien |
ICASSP | 3 |
| 2017 | Variational Recurrent Neural Networks for Speech Separation
Jen-Tzung Chien, Kuan-Ting Kuo |
INTERSPEECH | 1 |
| 2017 | Stochastic Recurrent Neural Network for Speech Recognition
Jen-Tzung Chien |
INTERSPEECH | 1 |
| 2017 | Deep Neural Factorization for Speech Recognition
Jen-Tzung Chien |
INTERSPEECH | 1 |
| 2017 | Discriminative subspace modeling of SNR and duration variabilities for robust speaker verification
Na Li 0012, Man-Wai Mak, Weiwei Lin 0002, Jen-Tzung Chien |
Comput. Speech Lang. | 4 |
| 2017 | Fast scoring for PLDA with uncertainty propagation via i-vector grouping
Weiwei Lin 0002, Man-Wai Mak, Jen-Tzung Chien |
Comput. Speech Lang. | 3 |
| 2017 | DNN-Driven Mixture of PLDA for Robust Speaker VerificationabstractThe mismatch between enrollment and test utterances due to different types of variabilities is a great challenge in speaker verification. Based on the observation that the SNR-level variability or channel-type variability causes heterogeneous clusters in i-vector space, this paper proposes to apply supervised learning to drive or guide the learning of probabilistic linear discriminant analysis (PLDA) mixture models. Specifically, a deep neural network (DNN) is trained to produce the posterior probabilities of different SNR levels or channel types given i-vectors as input. These posteriors then replace the posterior probabilities of indicator variables in the mixture of PLDA. The discriminative training causes the mixture model to perform more reasonable soft divisions of the i-vector space as compared to the conventional mixture of PLDA. During verification, given a test i-vector and a target-speaker's i-vector, the marginal likelihood for the same-speaker hypothesis is obtained by summing the component likelihoods weighted by the component posteriors produced by the DNN, and likewise for the different-speaker hypothesis. Results based on NIST 2012 SRE demonstrate that the proposed scheme leads to better performance under more realistic situations where both training and test utterances cover a wide range of SNRs and different channel types. Unlike the previous SNR-dependent mixture of PLDA which only focuses on SNR mismatch, the proposed model is more general and is potentially applicable to addressing different types of variability in speech. Na Li 0012, Man-Wai Mak, Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Deep discriminative manifold learningabstractThis paper presents a new non-linear dimensionality reduction with stochastic neighbor embedding. A deep neural network is developed for discriminative manifold learning where the class information in transformed low-dimensional space is preserved. Importantly, the objective function for deep manifold learning is formed as the Kullback-Leibler divergence between the probability measures of the labeled samples in high-dimensional and low-dimensional spaces. Different from conventional methods, the derived objective does not require the empirically-tuned parameter. This objective is optimized to attractive those samples from the same class to be close together and simultaneously impose those samples from different classes to be far apart. In the experiments on image and audio tasks, we illustrate the effectiveness of the proposed discriminative manifold learning in terms of visualization and classification performance. Jen-Tzung Chien, Ching-Huai Chen |
ICASSP | 1 |
| 2016 | Deep unfolding inference for supervised topic modelabstractConventional supervised topic model for multi-class classification is inferred via the variational inference algorithm where the model parameters are estimated by maximizing the lower bound of the logarithm of marginal likelihood function over input documents and labels. The classification accuracy is constrained by the variational lower bound. In this study, we aim to improve the classification accuracy by relaxing this constraint through directly maximizing the negative cross entropy error function via a deep unfolding inference (DUI). The inference procedure for class posterior is treated as the layer-wise learning in a deep neural network. The classification accuracy in DUI is accordingly increased by using the estimated topic parameters according to the exponentiated updates. Deep learning of supervised topic model is achieved through an error back-propagation algorithm. Experimental results show the superiority of DUI to variational Bayes inference in supervised topic model. Chao-Hsi Lee, Jen-Tzung Chien |
ICASSP | 2 |
| 2016 | Automatic speech recognition for acoustical analysis and assessment of cantonese pathological voice and speechabstractThis paper describes the application of state-of-the-art automatic speech recognition (ASR) systems to objective assessment of voice and speech disorders. Acoustical analysis of speech has long been considered a promising approach to non-invasive and objective assessment of people. In the past the types and amount of speech materials used for acoustical assessment were very limited. With the ASR technology, we are able to perform acoustical and linguistic analyses with a large amount of natural speech from impaired speakers. The present study is focused on Cantonese, which is a major Chinese dialect. Two representative disorders of speech production are investigated: dysphonia and aphasia. ASR experiments are carried out with continuous and spontaneous speech utterances from Cantonese-speaking patients. The results confirm the feasibility and potential of using natural speech for acoustical assessment of voice and speech disorders, and reveal the challenging issues in acoustic modeling and language modeling of pathological speech. Tan Lee, Yuanyuan Liu 0002, Pei-Wen Huang, Jen-Tzung Chien, Wang-Kong Lam, Yu Ting Yeung, Thomas K. T. Law, Kathy Yuet-Sheung Lee, Anthony Pak-Hin Kong, Sam-Po Law |
ICASSP | 4 |
| 2016 | Discriminative deep recurrent neural networks for monaural speech separationabstractDeep neural network is now a new trend towards solving different problems in speech processing. In this paper, we propose a discriminative deep recurrent neural network (DRNN) model for monaural speech separation. Our idea is to construct DRNN as a regression model to discover the deep structure and regularity for signal reconstruction from a mixture of two source spectra. To reinforce the discrimination capability between two separated spectra, we estimate DRNN separation parameters by minimizing an integrated objective function which consists of two measurements. One is the within-source reconstruction errors due to the individual source spectra while the other conveys the discrimination information which preserves the mutual difference between two source spectra during the supervised training procedure. This discrimination information acts as a kind of regularization so as to maintain between-source separation in monaural source separation. In the experiments, we demonstrate the effectiveness of the proposed method for speech separation compared with the other methods. Guan-Xiang Wang, Chung-Chien Hsu, Jen-Tzung Chien |
ICASSP | 3 |
| 2016 | Hybrid Accelerated Optimization for Speech Recognition
Jen-Tzung Chien, Pei-Wen Huang, Tan Lee |
INTERSPEECH | 1 |
| 2016 | Discriminative Layered Nonnegative Matrix Factorization for Speech Separation
Chung-Chien Hsu, Tai-Shih Chi, Jen-Tzung Chien |
INTERSPEECH | 3 |
| 2016 | Deep neural network driven mixture of PLDA for robust i-vector speaker verificationabstractIn speaker recognition, the mismatch between the enrollment and test utterances due to noise with different signal-to-noise ratios (SNRs) is a great challenge. Based on the observation that noise-level variability causes the i-vectors to form heterogeneous clusters, this paper proposes using an SNR-aware deep neural network (DNN) to guide the training of PLDA mixture models. Specifically, given an i-vector, the SNR posterior probabilities produced by the DNN are used as the posteriors of indicator variables of the mixture model. As a result, the proposed model provides a more reasonable soft division of the i-vector space compared to the conventional mixture of PLDA. During verification, given a test trial, the marginal likelihoods from individual PLDA models are linearly combined by the posterior probabilities of SNR levels computed by the DNN. Experimental results for SNR mismatch tasks based on NIST 2012 SRE suggest that the proposed model is more effective than PLDA and conventional mixture of PLDA for handling heterogeneous corpora. Na Li 0012, Man-Wai Mak, Jen-Tzung Chien |
SLT | 3 |
| 2016 | Bayesian Factorization and Learning for Monaural Source SeparationabstractThis paper presents a new Bayesian nonnegative matrix factorization (NMF) for monaural source separation. Using this approach, the reconstruction error based on NMF is represented by a Poisson distribution, and the NMF parameters, consisting of the basis and weight matrices, are characterized by the exponential priors. A variational Bayesian inference procedure is developed to learn variational parameters and model parameters. The randomness in separation process is faithfully represented so that the system robustness to model variations in heterogeneous environments could be achieved. Importantly, the exponential prior parameters are used to impose sparseness in basis representation. The variational lower bound of log marginal likelihood is adopted as the objective to control model complexity. The dependencies of variational objective on model parameters are fully characterized in the derived closed-form solution. A clustering algorithm is performed to find the groups of bases for unsupervised source separation. The experiments on speech/music separation and singing voice separation show that the proposed Bayesian NMF (BNMF) with adaptive basis representation outperforms the NMF with fixed number of bases and the other BNMFs in terms of signal-to-distortion ratio and the global normalized source to distortion ratio. Jen-Tzung Chien, Po-Kai Yang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Mixture of PLDA for Noise Robust I-Vector Speaker VerificationabstractIn real-world environments, noisy utterances with variable noise levels are recorded and then converted to i-vectors for cosine distance or PLDA scoring. This paper investigates the effect of noise-level variability on i-vectors. It demonstrates that noise-level variability causes the i-vectors to shift, causing the noise contaminated i-vectors to form clusters in the i-vector space. It also demonstrates that optimal subspaces for discriminating speakers are noise-level dependent. Based on these observations, this paper proposes using signal-to-noise ratio (SNR) of utterances as guidance for training mixture of PLDA models. To maximize the coordination among the PLDA models, mixtures of PLDA models are trained simultaneously via an EM algorithm using the utterances contaminated with noise at various levels. For scoring, given a test i-vector, the marginal likelihoods from individual PLDA models are linearly combined by the posterior probabilities of the test utterance's SNR. Verification scores are the ratio of the marginal likelihoods. Results based on NIST 2012 SRE suggest that the SNR-dependent mixture of PLDA is not only suitable for the situations where the test utterances exhibit a wide range of SNR, but also beneficial for the test utterances with unknown SNR distribution. Supplementary materials containing full derivations of the EM algorithms and scoring functions can be found in http://bioinfo.eie.polyu.edu.hk/mPLDA/SuppMaterials.pdf. Man-Wai Mak, Xiaomin Pang, Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Hierarchical Theme and Topic ModelingabstractConsidering the hierarchical data groupings in text corpus, e.g., words, sentences, and documents, we conduct the structural learning and infer the latent themes and topics for sentences and words from a collection of documents, respectively. The relation between themes and topics under different data groupings is explored through an unsupervised procedure without limiting the number of clusters. A tree stick-breaking process is presented to draw theme proportions for different sentences. We build a hierarchical theme and topic model, which flexibly represents the heterogeneous documents using Bayesian nonparametrics. Thematic sentences and topical words are extracted. In the experiments, the proposed method is evaluated to be effective to build semantic tree structure for sentences and the corresponding words. The superiority of using tree model for selection of expressive sentences for document summarization is illustrated. Jen-Tzung Chien |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Bayesian Recurrent Neural Network for Language ModelingabstractA language model (LM) is calculated as the probability of a word sequence that provides the solution to word prediction for a variety of information systems. A recurrent neural network (RNN) is powerful to learn the large-span dynamics of a word sequence in the continuous space. However, the training of the RNN-LM is an ill-posed problem because of too many parameters from a large dictionary size and a high-dimensional hidden layer. This paper presents a Bayesian approach to regularize the RNN-LM and apply it for continuous speech recognition. We aim to penalize the too complicated RNN-LM by compensating for the uncertainty of the estimated model parameters, which is represented by a Gaussian prior. The objective function in a Bayesian classification network is formed as the regularized cross-entropy error function. The regularized model is constructed not only by calculating the regularized parameters according to the maximum a posteriori criterion but also by estimating the Gaussian hyperparameter by maximizing the marginal likelihood. A rapid approximation to a Hessian matrix is developed to implement the Bayesian RNN-LM (BRNN-LM) by selecting a small set of salient outer-products. The proposed BRNN-LM achieves a sparser model than the RNN-LM. Experiments on different corpora show the robustness of system performance by applying the rapid BRNN-LM under different conditions. Jen-Tzung Chien, Yuan-Chu Ku |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | The shared dirichlet priors for bayesian language modelingabstractWe present a new full Bayesian approach for language modeling based on the shared Dirichlet priors. This model is constructed by introducing the Dirichlet distribution to represent the uncertainty of n-gram parameters in training phase as well as in test time. Given a set of training data, the marginal likelihood over n-gram probabilities is illustrated in a form of linearly-interpolated n-grams. The hyperparameters in Dirichlet distributions are interpreted as the prior backoff information which is shared for the group of n-gram histories. This study estimates the shared hyperparameters by maximizing the marginal distribution of n-gram given the training data. Such Bayesian language model is connected to the smoothed language model. Experimental results show the superiority of the proposed method to the other methods in terms of perplexity and word error rate. Jen-Tzung Chien |
ICASSP | 1 |
| 2015 | Deep recurrent regularization neural network for speech recognitionabstractThis paper presents a deep recurrent regularization neural network (DRRNN) for speech recognition. Our idea is to build a regularization neural network acoustic model by conducting the hybrid Tikhonov and weight-decay regularization which compensates the variations due to the input speech as well as the model parameters in the restricted Boltzmann machine as a pre-training stage for feature learning and structural modeling. In addition, a new backpropagation through time (BPTT) algorithm is developed by extending the truncated minibatch training for recurrent neural network where the minibatch BPTT is not only performed in recurrent layer but also in feedforward layer. The DRRNN acoustic model is accordingly established to capture the temporal correlation in a regularization neural network. Experimental results on the tasks of RM and Aurora4 show the effectiveness and robustness of using DRRNN for speech recognition. Jen-Tzung Chien, Tsai-Wei Lu |
ICASSP | 1 |
| 2015 | Modulation Wiener filter for improving speech intelligibilityabstractThis paper presents a single-channel high-dimensional Wiener filter in the spectro-temporal modulation domain. Unlike other conventional noise reduction techniques, the proposed algorithm not only reduces noise but also enhances the “textures” of the speech signal. A non-iterative decision-directed noise estimation method is adopted to estimate the modulation SNR for the modulation-domain Wiener filter. The efficacy of the proposed algorithm in enhancing speech intelligibility is assessed using the short-time objective intelligibility (STOI) measure. Statistical analysis results demonstrate that our proposed algorithm can improve STOI scores in speech-shape noise (SSN) and white noise conditions, but not in babble noise condition, while the conventional Wiener filter fails to improve STOI scores in all three noise conditions. Chung-Chien Hsu, Kah-Meng Cheong, Jen-Tzung Chien, Tai-Shih Chi |
ICASSP | 3 |
| 2015 | Layered nonnegative matrix factorization for speech separation
Chung-Chien Hsu, Jen-Tzung Chien, Tai-Shih Chi |
INTERSPEECH | 2 |
| 2015 | Laplace Group Sensing for Acoustic ModelsabstractThis paper presents the group sparse learning for acoustic models where a sequence of acoustic features is driven by Markov chain and each feature vector is represented by groups of basis vectors. The group of common bases represents the features across Markov states within a regression class. The group of individual basis compensates the intra-state residual information. Laplace distribution is used as the sparse prior of sensing weights for group basis representation. Laplace parameter serves as regularization parameter or automatic relevance determination which controls the selection of relevant bases for acoustic modeling. The groups of regularization parameters and basis vectors are estimated from training data by maximizing the marginal likelihood over sensing weights which is implemented by Laplace approximation using the Hessian matrix and the maximum a posteriori parameters. Model uncertainty is compensated through full Bayesian treatment. The connection of Laplace group sensing to lasso regularization is illustrated. Experiments on noisy speech recognition show the robustness of group sparse acoustic models in presence of different noise types and SNRs. Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Hierarchical Pitman-Yor-Dirichlet Language ModelabstractProbabilistic models are often viewed as insufficiently expressive because of strong limitation and assumption on the probabilistic distribution and the fixed model complexity. Bayesian nonparametric learning pursues an expressive probabilistic representation based on the nonparametric prior and posterior distributions with less assumption-laden approach to inference. This paper presents a hierarchical Pitman-Yor-Dirichlet (HPYD) process as the nonparametric priors to infer the predictive probabilities of the smoothed n-grams with the integrated topic information. A metaphor of hierarchical Chinese restaurant process is proposed to infer the HPYD language model (HPYD-LM) via Gibbs sampling. This process is equivalent to implement the hierarchical Dirichlet process-latent Dirichlet allocation (HDP-LDA) with the twisted hierarchical Pitman-Yor LM (HPY-LM) as base measures. Accordingly, we produce the power-law distributions and extract the semantic topics to reflect the properties of natural language in the estimated HPYD-LM. The superiority of HPYD-LM to HPY-LM and other language models is demonstrated by the experiments on model perplexity and speech recognition. Jen-Tzung Chien |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | The nested indian buffet process for flexible topic modeling
Jen-Tzung Chien, Ying-Lan Chang |
INTERSPEECH | 1 |
| 2014 | Binary mask estimation based on frequency modulations
Chung-Chien Hsu, Jen-Tzung Chien, Tai-Shih Chi |
INTERSPEECH | 2 |
| 2014 | Bayesian factorization and selection for speech and music separation
Po-Kai Yang, Chung-Chien Hsu, Jen-Tzung Chien |
INTERSPEECH | 3 |
| 2014 | Bayesian recurrent neural network language modelabstractThis paper presents a Bayesian approach to construct the recurrent neural network language model (RNN-LM) for speech recognition. Our idea is to regularize the RNN-LM by compensating the uncertainty of the estimated model parameters which is represented by a Gaussian prior. The objective function in Bayesian RNN (BRNN) is formed as the regularized cross entropy error function. The regularized model is not only constructed by training the regularized parameters according to the maximum a posteriori criterion but also estimating the Gaussian hyperparameter by maximizing the marginal likelihood. A rapid approximation to Hessian matrix is developed by selecting a small set of salient outer-products and illustrated to be effective for BRNN-LM. BRNN-LM achieves sparser model than RNN-LM. Experiments on different corpora show promising improvement by applying BRNN-LM using different amount of training data. Jen-Tzung Chien, Yuan-Chu Ku |
SLT | 1 |
| 2014 | Tikhonov regularization for deep neural network acoustic modelingabstractDeep neural network (DNN) has been widely demonstrated to achieve high performance in different speech recognition tasks. This paper focuses on the issue of model regularization in DNN acoustic model. Our idea is to compensate for the perturbations over training samples in the restricted Boltzmann machine (RBM) which is applied as a pre-training stage for unsupervised feature learning and structural modeling. We introduce the Tikhonov regularization in pre-training procedure and pursue the invariance property of objective function over the variations in input samples. This Tikhonov regularization is further combined with the regularization based on weight decay. The error function in supervised cross-entropy training is accordingly reduced. Experimental results on using RM and Aurora4 tasks show that hybrid regularization in RBM pre-training improves the training condition in DNN acoustic model and the robustness in speech recognition performance. Jen-Tzung Chien, Tsai-Wei Lu |
SLT | 1 |
| 2013 | Bayesian latent variable models for speech recognitionabstractWe present a Bayesian framework to learn prior and posterior distributions for latent variable models. Our goal is to deal with model regularization and achieve desirable prediction using heterogeneous speech data. A variational Bayesian expectation-maximization algorithm is developed to establish a latent variable model based on the exponential family distributions. This algorithm does not only estimate model parameters but also their hyperparameters which reflect the model uncertainties. The uncertainty is compensated to construct a variety of regularized models. We realize this full Bayesian framework for uncertainty decoding of speech signals. Compared to maximum likelihood method and Bayesian approach with heuristically-selected hyperparameters, the proposed method achieves higher speech recognition accuracy especially in case of sparse and noisy training data. Jen-Tzung Chien, Peng Liu 0001 |
ICASSP | 1 |
| 2013 | Hierarchical pitman-yor and dirichlet process for language model
Jen-Tzung Chien, Ying-Lan Chang |
INTERSPEECH | 1 |
| 2013 | Nonstationary Source Separation Using Sequential and Variational Bayesian LearningabstractIndependent component analysis (ICA) is a popular approach for blind source separation where the mixing process is assumed to be unchanged with a fixed set of stationary source signals. However, the mixing system and source signals are nonstationary in real-world applications, e.g., the source signals may abruptly appear or disappear, the sources may be replaced by new ones or even moving by time. This paper presents an online learning algorithm for the Gaussian process (GP) and establishes a separation procedure in the presence of nonstationary and temporally correlated mixing coefficients and source signals. In this procedure, we capture the evolved statistics from sequential signals according to online Bayesian learning. The activity of nonstationary sources is reflected by an automatic relevance determination, which is incrementally estimated at each frame and continuously propagated to the next frame. We employ the GP to characterize the temporal structures of time-varying mixing coefficients and source signals. A variational Bayesian inference is developed to approximate the true posterior for estimating the nonstationary ICA parameters and for characterizing the activity of latent sources. The differences between this ICA method and the sequential Monte Carlo ICA are illustrated. In the experiments, the proposed algorithm outperforms the other ICA methods for the separation of audio signals in the presence of different nonstationary scenarios. Jen-Tzung Chien, Hsin-Lung Hsieh |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Group Sparse Hidden Markov Models for Speech Recognition
Jen-Tzung Chien, Cheng-Chun Chiang |
INTERSPEECH | 1 |
| 2012 | Bayesian Group Sparse Learning for Nonnegative Matrix Factorization
Jen-Tzung Chien, Hsin-Lung Hsieh |
INTERSPEECH | 1 |
| 2012 | Topic-Based Hierarchical SegmentationabstractLatent Dirichlet allocation (LDA) is a new paradigm of topic model which is powerful to capture the latent topic information from natural language. However, the topic information in text streams, e.g. meeting recording, lecture transcription and conversational dialogue, are inherently heterogeneous and nonstationary without explicit boundaries. It is difficult to train a precise topic model from the observed text streams. Furthermore, the usage of words in different paragraphs within a document is varied with different composition styles. In this paper, we present a new hierarchical segmentation model (HSM) where the heterogeneous topic information in stream level and the word variations in document level are characterized. We incorporate the contextual topic information in stream-level segmentation. The topic similarity between sentences is used to form a beta distribution reflecting the prior knowledge of document boundaries in a text stream. The distribution of segmentation variable is adaptively updated to achieve flexible segmentation and is used to group coherent sentences into a topic-specific document. For each pseudo-document, we further use a Markov chain to detect the stylistic segments within a document. The words in a segment are accordingly generated by the same composition style, which differs from the style of the next segment. Each segment is represented by a Markov state, and so the word variations within a document are compensated. The whole model is trained by a variational Bayesian EM procedure and is evaluated on using TDT2 corpus. Experimental results show benefits by using the proposed HSM in terms of perplexity, segmentation error, detection accuracy and F measure. Jen-Tzung Chien, Chuang-Hua Chueh |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Convex Divergence ICA for Blind Source SeparationabstractIndependent component analysis (ICA) is vital for unsupervised learning and blind source separation (BSS). The ICA unsupervised learning procedure attempts to demix the observation vectors and identify the salient features or mixture sources. This work presents a novel contrast function for evaluating the dependence among sources. A convex divergence measure is developed by applying the convex functions to the Jensen's inequality. Adjustable with a convexity parameter, this inequality-based divergence measure has a wide range of the steepest descents to reach its minimum value. A convex divergence ICA (C-ICA) is constructed and a nonparametric C-ICA algorithm is derived with different convexity parameters where the non-Gaussianity of source signals is characterized by the Parzen window-based distribution. Experimental results indicate that the specialized C-ICA significantly reduces the number of learning epochs during estimation of the demixing matrix. The convergence speed is improved by using the scaled natural gradient algorithm. Experiments on the BSS of instantaneous, noisy and convolutive mixtures of speech and music signals further demonstrate the superiority of the proposed C-ICA to JADE, Fast-ICA, and the nonparametric ICA based on mutual information. Jen-Tzung Chien, Hsin-Lung Hsieh |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Multi-View and Multi-Objective Semi-Supervised Learning for HMM-Based Automatic Speech RecognitionabstractCurrent hidden Markov acoustic modeling for large-vocabulary continuous speech recognition (LVCSR) heavily relies on the availability of abundant labeled transcriptions. Given that speech labeling is both expensive and time-consuming while there is a huge amount of unlabeled data easily available nowadays, the semi-supervised learning (SSL) from both labeled and unlabeled data aiming to reduce the development cost for LVCSR becomes more important than ever. In this paper, a new SSL approach is proposed which exploits the cross-view transfer learning for LVCSR through a committee machine consisting of multiple views learned from different acoustic features and randomized decision trees. In addition, a multi-objective learning scheme is developed in each view by maximizing a hybrid information-theoretic criterion which is established by the relative entropy between labeled data and their labels and the entropy of unlabeled data. The multi-objective scheme is then generalized to a unified SSL framework which can be interpreted into a variety of learning strategies under different weighting schemes. Experiments conducted on English Broadcast News using 50 hours of transcribed speech with 50 hours and 150 hours of untranscribed speech show the benefits of proposed approaches. Jing Huang 0019, Jen-Tzung Chien |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Bayesian Sensing Hidden Markov ModelsabstractIn this paper, we introduce Bayesian sensing hidden Markov models (BS-HMMs) to represent sequential data based on a set of state-dependent basis vectors. The goal of this work is to perform Bayesian sensing and model regularization for heterogeneous training data. By incorporating a prior density on sensing weights, the relevance of different bases to a feature vector is determined by the corresponding precision parameters. The BS-HMM parameters, consisting of the basis vectors, the precision matrices of sensing weights and the precision matrices of reconstruction errors, are jointly estimated by maximizing the likelihood function, which is marginalized over the weight priors. We derive recursive solutions for the three parameters, which are expressed via maximum a posteriori estimates of the sensing weights. We specifically optimize BS-HMMs for large-vocabulary continuous speech recognition (LVCSR) by introducing a mixture model of BS-HMMs and by adapting the basis vectors to different speakers. Discriminative training of BS-HMMs in the model domain and the feature domain is also proposed. Experimental results on an LVCSR task show consistent improvements due to the three sets of BS-HMM parameters and demonstrate how the extensions of mixture models, speaker adaptation, and discriminative training achieve better recognition results compared to those of conventional HMMs based on Gaussian mixture models. George Saon, Jen-Tzung Chien |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Introduction to the Special Section on Deep Learning for Speech and Language ProcessingabstractCurrent speech recognition systems, for example, typically use Gaussian mixture models (GMMs), to estimate the observation (or emission) probabilities of hidden Markov models (HMMs), and GMMs are generative models that have only one layer of latent variables. Instead of developing more powerful models, most of the research effort has gone into finding better ways of estimating the GMM parameters so that error rates are decreased or the margin between different classes is increased. The same observation holds for natural language processing (NLP) in which maximum entropy (MaxEnt) models and conditional random fields (CRFs) have been popular for the last decade. Both of these approaches use shallow models whose success largely depends on the use of carefully handcrafted features. Dong Yu 0001, Geoffrey E. Hinton, Nelson Morgan, Jen-Tzung Chien, Shigeki Sagayama |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Some properties of Bayesian sensing hidden Markov modelsabstractIn Bayesian sensing hidden Markov models (BSHMMs) the acoustic feature vectors are represented by a set of state-dependent basis vectors and by time-dependent sensing weights. The Bayesian formulation comes from assuming state-dependent zero mean Gaussian priors for the weights and from using marginal likelihood functions obtained by integrating out the weights. Here, we discuss two properties of BSHMMs. The first property is that the marginal likelihood is Gaussian with a factor analyzed covariance matrix with the basis providing a low-rank correction to the diagonal covariance of the reconstruction errors. The second property, termed automatic relevance determination, provides a method for discarding basis vectors that are not relevant for encoding feature vectors. This allows model complexity control where one can initially train a large model and then prune it to a smaller size by removing the basis vectors which correspond to the largest precision values of the sensing weights. The last property turned out to be useful in successfully deploying models trained on 1800 hours of data during the 2011 DARPA GALE Arabic broadcast news transcription evaluation. George Saon, Jen-Tzung Chien |
ASRU | 2 |
| 2011 | Multi-view and multi-objective semi-supervised learning for large vocabulary continuous speech recognitionabstractCurrent hidden Markov acoustic modeling for large vocabulary continuous speech recognition (LVCSR) relies on the availability of abundant labeled transcriptions. Given that speech labeling is both expensive and time-consuming while there is a huge amount of unlabeled data easily available nowadays, semi-supervised learning (SSL) from both labeled and unlabeled data which aims to reduce the development cost for LVCSR becomes more important than ever. In this paper, we propose SSL for LVCSR by using the multiple views learned from different acoustic features and randomized decision trees. In addition, we develop the multi-objective learning of HMM-based acoustic models by optimizing a hybrid criterion which is established by the combination of the discriminative mutual information from labeled data and the entropy from unlabeled data. Experiments conducted on Broadcast News show the benefits of proposed methods. Jing Huang 0019, Jen-Tzung Chien |
ICASSP | 3 |
| 2011 | Nonstationary and temporally correlated source separation using Gaussian processabstractBlind source separation (BSS) is a process to reconstruct source signals from the mixed signals. The standard BSS methods assume a fixed set of stationary source signals with the fixed distribution functions. However, in practical mixing systems, the source signals are nonstationary and temporally correlated; e.g. source signal may be abruptly active or inactive or even replaced by a new one. The mixing system is also time-varying. In this paper, we present a novel Gaussian process (GP) to characterize the time-varying mixing coefficients and the temporally correlated source signals. An online variational Bayesian algorithm is established to learn the noisy mixing process where GP priors are adopted to express the correlated sources as well as the mixing matrix. Experimental results demonstrate the effectiveness of proposed method in speech separation under different scenarios. Hsin-Lung Hsieh, Jen-Tzung Chien |
ICASSP | 2 |
| 2011 | Bayesian sensing hidden Markov models for speech recognitionabstractWe introduce Bayesian sensing hidden Markov models (BS-HMMs) to represent speech data based on a set of state-dependent basis vectors. By incorporating the prior density of sensing weights, the relevance of a feature vector to different bases is determined by the corresponding precision parameters. The BS-HMM parameters, consisting of the basis vectors, the precision matrices of sensing weights and the precision matrices of reconstruction errors, are jointly estimated by maximizing the likelihood function, which is marginalized over the weight priors. We derive recursive solutions for the three parameters, which are expressed via maximum a posteriori estimates of the sensing weights. Experimental results on an LVCSR task show consistent gains over conventional HMMs with Gaussian mixture models for both ML and discriminative training scenarios. George Saon, Jen-Tzung Chien |
ICASSP | 2 |
| 2011 | Discriminative training for Bayesian sensing hidden Markov modelsabstractWe describe feature space and model space discriminative training for a new class of acoustic models called Bayesian sensing hidden Markov models (BS-HMMs). In BS-HMMs, speech data is represented by a set of state-dependent basis vectors. The relevance of a feature vector to different bases is determined by the precision matrices of the sensing weights. The basis vectors and the precision matrices of the reconstruction errors are jointly estimated by optimizing a maximum mutual information (MMI) criterion. Additionally, we discuss the training of an fMPE-style discriminative feature transformation under the same criterion given these models. Experimental results on an LVCSR task show that the proposed models outperform discriminatively trained conventional HMMs with Gaussian mixture models (GMMs). Cross-adapting the baseline GMM-HMMs to the BS-HMM output yields a 6% relative gain which indicates that the two systems make different errors. George Saon, Jen-Tzung Chien |
ICASSP | 2 |
| 2011 | Dirichlet Class Language Models for Speech RecognitionabstractLatent Dirichlet allocation (LDA) was successfully developed for document modeling due to its generalization to unseen documents through the latent topic modeling. LDA calculates the probability of a document based on the bag-of-words scheme without considering the order of words. Accordingly, LDA cannot be directly adopted to predict words in speech recognition systems. This work presents a new Dirichlet class language model (DCLM), which projects the sequence of history words onto a latent class space and calculates a marginal likelihood over the uncertainties of classes, which are expressed by Dirichlet priors. A Bayesian class-based language model is established and a variational Bayesian procedure is presented for estimating DCLM parameters. Furthermore, the long-distance class information is continuously updated using the large-span history words and is dynamically incorporated into class mixtures for a cache DCLM. Different language models are experimentally evaluated using the Wall Street Journal (WSJ) corpus. The amount of training data and the size of vocabulary are evaluated. We find that the cache DCLM effectively characterizes the unseen -gram events and stores the class information for long-distance language modeling. This approach outperforms the other class-based and topic-based language models in terms of perplexity and recognition accuracy. The DCLM and cache DCLM achieved relative gain of word error rate by 3% to 5% over the LDA topic-based language model with different sizes of training data . Jen-Tzung Chien, Chuang-Hua Chueh |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Topic cache language model for speech recognitionabstractTraditional n-gram language models suffer from insufficient long-distance information. The cache language model, which captures the dynamics of word occurrences in a cache, is feasible to compensate this weakness. This paper presents a new topic cache model for speech recognition based on the latent Dirichlet language model where the latent topic structure is explored from n-gram events and employed for word prediction. In particular, the long-distance topic information is continuously updated from the large-span historical words and dynamically incorporated in generating the topic mixtures through Bayesian learning. The topic cache language model does effectively characterize the unseen n-gram events and catch the topic cache for long-distance language modeling. In the experiments on Wall Street Journal corpus, the proposed method achieves better performance than baseline n-gram and the other related language models in terms of perplexity and recognition accuracy. Chuang-Hua Chueh, Jen-Tzung Chien |
ICASSP | 2 |
| 2010 | Online Bayesian learning for dynamic source separationabstractIndependent component analysis (ICA) is a popular approach for blind source separation when the source signals are stationary with fixed distribution functions. However, the source signals are nonstationary in real-world applications, e.g. the source signals may abruptly appear or disappear or even the number of sources may be changed by time. This study presents a nonstationary ICA for dynamic source separation through an online Bayesian learning procedure. In this procedure, we capture the evolved statistics of independent sources from the online observed signals. The mixing matrix is incrementally compensated at each frame and continuously propagated to the next frame. A variational Bayesian algorithm is established to estimate the nonstationary ICA parameters. The number of latent sources is automatically determined at each frame. In the experiments, the proposed method effectively recovers the source speech signals from different speakers in presence of different mixing scenarios. Hsin-Lung Hsieh, Jen-Tzung Chien |
ICASSP | 2 |
| 2010 | A new nonnegative matrix factorization for independent component analysisabstractNonnegative matrix factorization (NMF) is known as a parts-based linear representation for nonnegative data. This method has been applied for blind source separation (BSS) when the sources are nonnegative. This paper presents a new NMF method for independent component analysis (ICA), which is useful for BSS without the nonnegativity constraint. Using this method, we transform the sources by their cumulative distribution functions (CDFs) and perform the nonparametric quantization to construct a nonnegative matrix where each entry represents the joint probability density of two transformed signals. The NMF procedure is accordingly realized to find the ICA demixing matrix. The independence between sources is maximized towards attaining the uniformity in the joint probability density. In the experiments on the separation of signal and music signals, we show the effectiveness of the proposed NMF-ICA compared to the infomax ICA and FastICA algorithms. Hsin-Lung Hsieh, Jen-Tzung Chien |
ICASSP | 2 |
| 2010 | Variational inference for conditional random fieldsabstractConditional random fields (CRFs) have been popular for contextual pattern classification. This paper presents two variational inference methods for direct approximation of a conditional probability instead of indirect calculation through Viterbi approximation of a marginal probability. The CRFs with the factorized variational inference (FVI) and the structured variational inference (SVI) are proposed and investigated for human motion recognition. In general, FVI assumes a factorization of variational distributions of individual states for representation of conditional probability while SVI preserves the state structure in the variational distribution. In the experiments on using IDIAP human motion database, we found that CRFs using variation inference methods performed better than baseline CRFs using Viterbi approximation. CRFs with SVI obtained higher classification accuracy than those with FVI. Chih-Pin Liao, Jen-Tzung Chien |
ICASSP | 2 |
| 2010 | A new topic-bridged model for transfer learningabstractIn real-world information systems, there are abundant unlabeled data but sparse labeled data. It is challenging to construct an adaptive model to classify a large amount of documents containing different domains. The classifiers trained from a source domain shall perform poorly for the test data in a target domain due to the domain mismatch. In this study, we build a topic-bridged latent Dirichlet allocation (TLDA) model from a variety of labeled and unlabeled documents and perform the transfer learning for document classification. The severe change of word distributions is compensated by bridging the latent topics of source and target data which are drawn by the Dirichlet priors. A variational inference procedure is performed for semi-supervised learning. In the experiments on text categorization using 20 Newsgroups dataset, the proposed TLDA model achieved higher classification performance compared to the other methods. Meng-Sung Wu, Jen-Tzung Chien |
ICASSP | 2 |
| 2010 | Online Gaussian process for nonstationary speech separation
Hsin-Lung Hsieh, Jen-Tzung Chien |
INTERSPEECH | 2 |
| 2010 | Joint acoustic and language modeling for speech recognition
Jen-Tzung Chien, Chuang-Hua Chueh |
Speech Commun. | 1 |
| 2009 | Latent Dirichlet learning for document summarizationabstractAutomatic summarization is developed to extract the representative contents or sentences from a large corpus of documents. This paper presents a new hierarchical representation of words, sentences and documents in a corpus, and infers the Dirichlet distributions for latent topics and latent themes in word level and sentence level, respectively. The sentence-based latent Dirichlet allocation (SLDA) is accordingly established for document summarization. Different from the vector space summarization, SLDA is built to fit the fine structure of text documents, and is specifically designed for sentence selection. SLDA acts as a sentence mixture model with a mixture of Dirichlet themes, which are used to generate the latent topics in observed words. The theme model is inherent to distinguish sentences in a summarization system. In the experiments, the proposed SLDA outperforms other methods for document summarization in terms of precision, recall and F-measure. Ying-Lang Chang, Jen-Tzung Chien |
ICASSP | 2 |
| 2009 | Bayesian large margin hidden Markov models for speech recognitionabstractThis paper presents a Bayesian learning approach to large margin classifier for hidden Markov model (HMM) based speech recognition. We build the Bayesian large margin HMMs (BLM-HMMs) and improve the model generalization for handling unknown test environments. Using BLM-HMMs, the variational Bayesian HMM parameters are estimated by maximizing lower bound of a marginal likelihood over the uncertainties of HMM parameters. The Bayesian large margin estimation is performed with frame selection mechanism, and is illustrated to meet the objective of support vector machines, i.e. maximal class margin and minimal training errors. The new objective function is not only interpreted as a discriminative criterion, but also feasible to deal with model selection and adaptive training. Experiments on phone recognition show that BLM-HMMs perform better than other generative and discriminative models. Jung-Chun Chen, Jen-Tzung Chien |
ICASSP | 2 |
| 2009 | Independent component analysis for noisy speech recognitionabstractIndependent component analysis (ICA) is not only popular for blind source separation but also for unsupervised learning when the observations can be decomposed into some independent components. These components represent the specific speaker, gender, accent, noise or environment, and act as the basis functions to span the vector space of the human voices in different conditions. Different from eigenvoices built by principal component analysis, the proposed independent voices are estimated by ICA algorithm, and are applied for efficient coding of an adapted acoustic model. Since the information redundancy is significantly reduced in independent voices, we effectively calculate a coordinate vector in independent voice space, and estimate the hidden Markov models (HMMs) for speech recognition. In the experiments, we build independent voices from HMMs under different noise conditions, and find that these voices attain larger redundancy reduction than eigenvoices. The noise adaptive HMMs generated by independent voices achieve better recognition performance than those by eigenvoices. Hsin-Lung Hsieh, Jen-Tzung Chien, Koichi Shinoda, Sadaoki Furui |
ICASSP | 2 |
| 2009 | An evidence framework for Bayesian learning of continuous-density hidden Markov modelsabstractWe present an evidence Bayesian framework, which can learn both the prior distributions and posterior distributions from data, for continuous-density hidden Markov models (CDHMM). The goal of this study is to build the regularized CDHMMs to improve model generalization, and achieve desirable recognition performance for unknown test speech. Under this framework, we develop an EM iterative procedure to estimate the marginal distribution or the evidence function for exponential family distributions. By adopting the variational Bayesian inference, we derive an empirical Bayesian solution to CDHMM parameters and their hyperparameters. Such a regularized CDHMM compensates the model uncertainty and the ill-posed conditions. Compared with maximum likelihood (ML) or other Bayesian approaches with heuristic hyperparameters, the proposed approach can utilize available data more effectively. The experiments on noisy speech recognition using Aurora2 show that the proposed Bayesian approach performs better than the baseline ML CDHMMs especially with mismatched test data or limited training data. Yu Zhang 0007, Peng Liu 0001, Jen-Tzung Chien, Frank K. Soong |
ICASSP | 3 |
| 2009 | Nonstationary latent Dirichlet allocation for speech recognition
Chuang-Hua Chueh, Jen-Tzung Chien |
INTERSPEECH | 2 |
| 2009 | Factor analyzed HMM topology for speech recognition
Chuan-Wei Ting, Jen-Tzung Chien |
INTERSPEECH | 2 |
| 2009 | Acoustic Factor Analysis for Streamed Hidden Markov ModelingabstractThis paper presents a novel streamed hidden Markov model (HMM) framework for speech recognition. The factor analysis (FA) principle is adopted to explore the common factors from acoustic features. The streaming regularities in building HMMs are governed by the correlation between cepstral features, which is inherent in common factors. Those features corresponding to the same factor are generated by the identical HMM state. Accordingly, the multiple Markov chains are adopted to characterize the variation trends in different dimensions of cepstral vectors. An FA streamed HMM (FASHMM) method is developed to relax the assumption of standard HMM topology, namely, that all features of a speech frame perform the same state emission. The proposed FASHMM is more flexible than the streamed factorial HMM (SFHMM) where the streaming was empirically determined. To reduce the number of factor loading matrices in FA, we evaluated the similarity between individual matrices to find the optimal solution to parameter clustering of FA models. A new decoding algorithm was presented to perform FASHMM speech recognition. FASHMM carries out the streamed Markov chains for a sequence of multivariate Gaussian mixture observations through the state transitions of the partitioned vectors. In the experiments, the proposed method reduced the recognition error rates significantly when compared with the standard HMM and SFHMM methods. Jen-Tzung Chien, Chuan-Wei Ting |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Minimum Rank Error Language ModelingabstractStatistical language modeling has been successfully developed for speech recognition and information retrieval. The minimum classification error (MCE) training was undertaken to enhance speech recognition performance by minimizing the word error rate. This paper presents a new minimum rank error (MRE) algorithm forn-gram language model training. Rather than speech recognition, the proposed language models are estimated forinformationretrievalby considering the metric ofaverageprecision. However, the maximization of average precision is closely linked to minimizing the rank error or optimizing the order of the ranked documents. Accordingly, this paper calculates therankerrorlossfunctionfrom the misordering pairs of relevant and irrelevant documents in the rank list. The Bayes risk due to the expected rank loss is minimized to develop the Bayesian retrieval rule forad-hocinformation retrieval. Consequently, thediscriminativetrainingof language model is performed by integrating discrimination information from individual relevant documents relative to their corresponding irrelevant documents. Experimental results on TREC collections indicate that the proposed MRE language model improves the order of relevant documents, and degrades that of irrelevant documents. The MRE method achieves significantly higher average precision for test queries than the maximum likelihood and the MCE retrieval models. Jen-Tzung Chien, Meng-Sung Wu |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | A new mutual information measure for independent component alalysisabstractIndependent component analysis (ICA) is a popular approach for blind source separation (BSS). In this study, we develop a new mutual information measure for BSS and unsupervised learning of acoustic models. The underlying concept of ICA unsupervised learning algorithm is to demix the observations vectors and identify the corresponding mixture sources. These independent sources represent the specific speaker, gender, accent, noise or environment, etc, embedded in acoustic models. The novelty of the proposed ICA is to derive a new metric of mutual information for measuring the dependence among mixture sources. We focus on building this metric based on the Jensen’s inequality, which is illustrated to use smaller number of iterations in finding the demixing matrix compared to other types of mutual information. We present a parametric ICA using the generalized Gaussian distribution to characterize the non-Gaussianity of model parameters. Also, a nonparametric ICA is established by using the Parzen window based distribution. In the experiments on BSS and noisy speech recognition, we demonstrate the effectiveness of the proposed Jensen ICA compared to FastICA and other nonparametric ICA. Jen-Tzung Chien, Hsin-Lung Hsieh, Sadaoki Furui |
ICASSP | 1 |
| 2008 | Reliable feature selection for language model adaptationabstractLanguage model adaptation aims to adapt a general model to a domain-specific model so that the adapted model can match the lexical information in test data. The minimum discrimination information (MDI) is a popular mechanism for language model adaptation through minimizing the Kullback-Leibler distance to the background model where the constraints found in adaptation data are satisfied. MDI adaptation with unigram constraints has been successfully applied for speech recognition owing to its computational efficiency. However, the unigram features only contain low-level information of adaptation articles which are too rough to attain precise adaptation performance. Accordingly, it is desirable to induce high-order features and explore delicate information for language model adaptation if the adaptation data is abundant. In this study, we focus on adaptively select the reliable features based on re-sampling and calculating the statistical confidence interval. We identify the reliable regions and build the inequality constraints for MDI adaptation. In this way, the reliable intervals can be used for adaptation so that interval estimation is achieved rather than point estimation. Also, the features can be selected automatically in the whole procedure. In the experiments, we carry out the proposed method for broadcast news transcription. We obtain significant improvement compared to conventional MDI adaptation with unigram features for different amount of adaptation data. Chuang-Hua Chueh, Jen-Tzung Chien |
ICASSP | 2 |
| 2008 | Graphical modeling of conditional random fields for human motion recognitionabstractModeling and understanding human motions are challenging in computer vision areas because the similar motions often occur at various time moments. The long-term dependences in observation data should be modeled to improve motion recognition performance. The conditional random field (CRF) is a powerful mechanism for large-span data modeling. In this paper, we present a new graphical model approach to effectively and efficiently implement CRF. Specifically, we integrate the dependent variables of a graph into a clique and build the junction tree for complex CRF structure with cycles. Using this approach, a tree inference algorithm is developed for finding the joint probability of all variables in the clique tree. In the implementation, we specify the continuous-valued hidden Markov model (HMM) parameters as the feature functions and evaluate the proposed junction tree CRF (JT-CRF) by using CMU Graphics Lab Motion Capture Database. The experimental results show that JT-CRF achieves the highest classification accuracies compared to the HMM, the maximum entropy Markov model and the linear-chain CRF. Chih-Pin Liao, Jen-Tzung Chien |
ICASSP | 2 |
| 2008 | Adaptive HMM topology for speech recognition
Chuan-Wei Ting, Kuo-Yuan Lee, Jen-Tzung Chien |
INTERSPEECH | 3 |
| 2008 | Bayesian latent topic clustering model
Meng-Sung Wu, Jen-Tzung Chien |
INTERSPEECH | 2 |
| 2008 | Latent dirichlet language model for speech recognitionabstractLatent Dirichlet allocation (LDA) has been successfully presented for document modeling and classification. LDA calculates the document probability based on bag-of-words scheme without considering the sequence of words. This model discovers the topic structure at document level, which is different from the concern of word prediction in speech recognition. In this paper, we present a new latent Dirichlet language model (LDLM) for modeling of word sequence. A new Bayesian framework is introduced by merging the Dirichlet priors to characterize the uncertainty of latent topics of n-gram events. The robust topic-based language model is established accordingly. In the experiments, we implement LDLM for continuous speech recognition and obtain better performance than probabilistic latent semantic analysis (PLSA) based language method. Jen-Tzung Chien, Chuang-Hua Chueh |
SLT | 1 |
| 2008 | Continuous topic language modeling for speech recognitionabstractContinuous representation of word sequence can effectively solve data sparseness problem in n-gram language model, where the discrete variables of words are represented and the unseen events are prone to happen. This problem is increasingly severe when extracting long-distance regularities for high-order n-gram model. Rather than considering discrete word space, we construct the continuous space of word sequence where the latent topic information is extracted. The continuous vector is formed by the topic posterior probabilities and the least-squares projection matrix from discrete word space to continuous topic space is estimated accordingly. The unseen words can be predicted through the new continuous latent topic language model. In the experiments on continuous speech recognition, we obtain significant performance improvement over the conventional topic-based language model. Chuang-Hua Chueh, Jen-Tzung Chien |
SLT | 2 |
| 2008 | Maximum Confidence Hidden Markov Modeling for Face RecognitionabstractThis paper presents a hybrid framework of feature extraction and hidden Markov modeling(HMM) for two-dimensional pattern recognition. Importantly, we explore a new discriminative training criterion to assure model compactness and discriminability. This criterion is derived from hypothesis test theory via maximizing the confidence of accepting the hypothesis that observations are from target HMM states rather than competing HMM states. Accordingly, we develop the maximum confidence hidden Markov modeling (MC-HMM) for face recognition. Under this framework, we merge a transformation matrix to extract discriminative facial features. The closed-form solutions to continuous-density HMM parameters are formulated. Attractively, the hybrid MC-HMM parameters are estimated under the same criterion and converged through the expectation-maximization procedure. From the experiments on FERET and GTFD facial databases, we find that the proposed method obtains robust segmentation in presence of different facial expressions, orientations, etc. In comparison with maximum likelihood and minimum classification error HMMs, the proposed MC-HMM achieves higher recognition accuracies with lower feature dimensions. Jen-Tzung Chien, Chih-Pin Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Factor Analyzed Subspace Modeling and SelectionabstractWe present a novel subspace modeling and selection approach for noisy speech recognition. In subspace modeling, we develop a factor analysis (FA) representation of noisy speech, which is a generalization of a signal subspace (SS) representation. Using FA, noisy speech is represented by the extracted common factors, factor loading matrix, and specific factors. The observation space of noisy speech is accordingly partitioned into a principal subspace, containing speech and noise, and a minor subspace, containing residual speech and residual noise. We minimize the energies of speech distortion in the principal subspace as well as in the minor subspace so as to estimate clean speech with residual information. Importantly, we explore the optimal subspace selection via solving the hypothesis test problems. We test the equivalence of eigenvalues in the minor subspace to select the subspace dimension. To fulfill the FA spirit, we also examine the hypothesis of uncorrelated specific factors/residual speech. The subspace can be partitioned according to a consistent confidence towards rejecting the null hypothesis. Optimal solutions are realized through the likelihood ratio tests, which arrive at the approximated chi-square distributions as test statistics. In the experiments on the Aurora2 database, the FA model significantly outperforms the SS model for speech enhancement and recognition. Subspace selection via testing the correlation of residual speech achieves higher recognition accuracies than that of testing the equivalent eigenvalues in the minor subspace. Jen-Tzung Chien, Chuan-Wei Ting |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Adaptive Bayesian Latent Semantic AnalysisabstractDue to the vast growth of data collections, the statistical document modeling has become increasingly important in language processing areas. Probabilistic latent semantic analysis (PLSA) is a popular approach whereby the semantics and statistics can be effectively captured for modeling. However, PLSA is highly sensitive to task domain, which is continuously changing in real-world documents. In this paper, a novel Bayesian PLSA framework is presented. We focus on exploiting the incremental learning algorithm for solving the updating problem of new domain articles. This algorithm is developed to improve document modeling by incrementally extracting up-to-date latent semantic information to match the changing domains at run time. By adequately representing the priors of PLSA parameters using Dirichlet densities, the posterior densities belong to the same distribution so that a reproducible prior/posterior mechanism is activated for incremental learning from constantly accumulated documents. An incremental PLSA algorithm is constructed to accomplish the parameter estimation as well as the hyperparameter updating. Compared to standard PLSA using maximum likelihood estimate, the proposed approach is capable of performing dynamic document indexing and modeling. We also present the maximum a posteriori PLSA for corrective training. Experiments on information retrieval and document categorization demonstrate the superiority of using Bayesian PLSA methods. Jen-Tzung Chien, Meng-Sung Wu |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Factor analysis of acoustic features for streamed hidden Markov modelingabstractThis paper presents a new streamed hidden Markov model (HMM) framework for speech recognition. The factor analysis (FA) is performed to discover the common factors of acoustic features. The streaming regularities are governed by the correlation between features, which is inherent in common factors. Those features corresponding to the same factor are generated by identical HMM state. Accordingly, we use multiple Markov chains to represent the variation trends in cepstral features. We develop a FA streamed HMM (FASHMM) and go beyond the conventional HMM assuming that all features at a speech frame conduct the same state emission. This streamed HMM is more delicate than the factorial HMM where the streaming was empirically determined. We also exploit a new decoding algorithm for FASHMM speech recognition. In this manner, we fulfill the flexible Markov chains for an input sequence of multivariate Gaussian mixture observations. In the experiments, the proposed method can reduce word error rate by 36% at most. Chuan-Wei Ting, Jen-Tzung Chien |
ASRU | 2 |
| 2007 | Recursive Bayesian Regression Modeling and LearningabstractThis paper presents a new Bayesian regression and learning algorithm for adaptive pattern classification. Our aim is to continuously update regression parameters to meet nonstationary environments for real-world applications. Here, a kernel regression model is used to represent two-class data. The initial regression parameters are estimated by maximizing the likelihood of training data. To activate online learning, we properly express the randomness of regression parameters as a conjugate prior, which is a normal-gamma distribution. When new adaptation data are enrolled, we can accumulate sufficient statistics and come up with a new normal-gamma distribution as the posterior distribution. We therefore exploit a recursive Bayesian algorithm for online regression and learning. Regression parameters are incrementally adapted to the newest environments. Robustness of classification rule is assured using online regression parameters. In the experiments on face recognition, the proposed regression algorithm outperforms support vector machine and relevance vector machine for different numbers of adaptation data. Jen-Tzung Chien, Jung-Chun Chen |
ICASSP (2) | 1 |
| 2007 | Predictive minimum Bayes risk classification for robust speech recognitionabstractThis paper presents a new Bayes classification rule towards minimizing the predictive Bayes risk for robust speech recognition. Conventionally, the plug-in maximum a posteriori (MAP) classification is constructed by adopting nonparametric loss function and deterministic model parameters. Speech recognition performance is limited due to the environmental mismatch and the ill-posed model. Concerning these issues, we develop the predictive minimum Bayes risk (PMBR) classification where the predictive distributions are inherent in Bayes risk. More specifically, we exploit the Bayes loss function and the predictive word posterior probability for Bayes classification. Model mismatch and randomness are compensated to improve generalization capability in speech recognition. In the experiments on car speech recognition, we estimate the prior densities of hidden Markov model parameters from adaptation data. With the prior knowledge of new environment and model uncertainty, PMBR classification is realized and evaluated to be better than MAP, MBR and Bayesian predictive classification. Jen-Tzung Chien, Koichi Shinoda, Sadaoki Furui |
INTERSPEECH | 1 |
| 2007 | Minimum rank error training for language modeling
Meng-Sung Wu, Jen-Tzung Chien |
INTERSPEECH | 2 |
| 2007 | Structural Bayesian language modeling and adaptationabstractWe propose a language modeling and adaptation framework using Bayesian structural maximum a posteriori (SMAP) principle, in which each n-gram event is embedded in a branch of a tree structure. The nodes in the first layer of this tree structure represent the unigrams, and those in the second layer represent the bigrams, and so on. Each node in the tree structure has an associated hyper-parameter representing the information about the prior distribution, and a count representing the number of times the word sequence occurs in the domain-specific data. In general, the hyper-parameters depend on the observation frequency of not only the node event but also its parent node of lower order n-gram event. Our automatic speech recognition experiments using the Wall Street Journal corpus verify that the proposed SMAP language model adaptation achieves a 5.6% relative improvement over maximum likelihood language models obtained with the same training and adaptation data sets. Sibel Yaman, Jen-Tzung Chien |
INTERSPEECH | 2 |
| 2006 | Towards Optimal Bayes Decision for Speech RecognitionabstractThis paper presents a new speech recognition framework towards fulfilling optimal Bayes decision theory, which is essential for general pattern recognition. The recognition procedure is developed through minimizing the Bayes risk, or equivalently the expected loss due to classification action. Typically, loss function measures the penalty/evidence of choosing a candidate hypothesis. This function was manually specified or empirically calculated. Here, we exploit a novel Bayes loss function via testing the hypotheses whether the classification action produces loss or not. A Bayes factor is derived to measure loss in a statistical and meaningful way. Attractively, Bayes loss function using predictive distributions is robust to the uncertainty of environments. Also, optimizing this Bayes criterion equals to minimizing classification errors of test data. The relation between the minimum classification error (MCE) classifier and the proposed optimal Bayes classifier (OBC) is bridged. Specifically, the logarithm of Bayes factor in OBC is analogous to the misclassification measure in MCE when using predictive distribution as the discriminant function. We accordingly build a robust and discriminative classification for large vocabulary continuous speech recognition. In the experiments on broadcast news transcription, the new OBC rule significantly outperforms traditional maximum a posteriori classification. Jen-Tzung Chien, Chih-Hsien Huang, Koichi Shinoda, Sadaoki Furui |
ICASSP (1) | 1 |
| 2006 | Maximum Entropy Modeling of Acoustic and Linguistic FeaturesabstractTraditionally, speech recognition system is established assuming that acoustic and linguistic information sources are independent. Parameters of hidden Markov model and n-gram are estimated individually and then plugged in a maximum a posteriori classification rule. However, acoustic and linguistic features are correlated in essence. Modeling performance is limited accordingly. This study aims to relax the independence assumption and achieve sophisticated acoustic and linguistic modeling for speech recognition. We propose an integrated approach based on maximum entropy (ME) principle where acoustic and linguistic features are optimally merged in a unified framework. The correlations between acoustic and linguistic features are explored and properly represented in the integrated models. Due to the flexibility of ME model, we can further combine other high-level linguistic features. In the experiments, we carry out the proposed methods for broadcast news transcription using MATBN database. We obtain significant improvement compared to conventional speech recognition system using individual maximum likelihood training. Chuang-Hua Chueh, Jen-Tzung Chien |
ICASSP (1) | 2 |
| 2006 | Maximum Confidence Hidden Markov ModelingabstractThis paper presents a compact and discriminative hidden Markov model (HMM) approach for general pattern classification. To achieve model compactness and discriminability, we simultaneously perform feature dimension reduction and HMM parameter estimation via maximizing the confidence of accepting the hypothesis that observations are from target HMM states rather than competing HMM states. A new discriminative training criterion is derived using hypothesis test theory. Particularly, we develop the maximum confidence hidden Markov modeling (MCHMM) framework for face recognition. Using this framework, we incorporate a transformation matrix to extract discriminative facial features. The continuous-density HMM parameters are estimated using the extracted features. Importantly, we adopt a consistent criterion to build whole framework including feature extraction and model estimation. From the experiments on ORL facial databases, we find that the proposed method obtains robust image segmentation performance in presence of different variations of facial expressions, orientations, etc. In comparison of previous HMM approaches, the proposed MCHMM achieves better recognition accuracies and image segmentation. Chih-Pin Liao, Jen-Tzung Chien |
ICASSP (5) | 2 |
| 2006 | Subspace modeling and selection for noisy speech recognition
Jen-Tzung Chien, Chuan-Wei Ting |
INTERSPEECH | 1 |
| 2006 | Association pattern language modelingabstractStatistical n-gram language modeling is popular for speech recognition and many other applications. The conventional n-gram suffers from the insufficiency of modeling long-distance language dependencies. This paper presents a novel approach focusing on mining long distance word associations and incorporating these features into language models based on linear interpolation and maximum entropy (ME) principles. We highlight the discovery of the associations of multiple distant words from training corpus. A mining algorithm is exploited to recursively merge the frequent word subsets and efficiently construct the set of association patterns. By combining the features of association patterns into n-gram models, the association pattern n-grams are estimated with a special realization to trigger pair n-gram where only the associations of two distant words are considered. In the experiments on Chinese language modeling, we find that the incorporation of association patterns significantly reduces the perplexities of n-gram models. The incorporation using ME outperforms that using linear interpolation. Association pattern n-gram is superior to trigger pair n-gram. The perplexities are further reduced using more association steps. Further, the proposed association pattern n-grams are not only able to elevate document classification accuracies but also improve speech recognition rates. Jen-Tzung Chien |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | A new independent component analysis for speech recognition and separationabstractThis paper presents a novel nonparametric likelihood ratio (NLR) objective function for independent component analysis (ICA). This function is derived through the statistical hypothesis test of independence of random observations. A likelihood ratio function is developed to measure the confidence toward independence. We accordingly estimate the demixing matrix by maximizing the likelihood ratio function and apply it to transform data into independent component space. Conventionally, the test of independence was established assuming data distributions being Gaussian, which is improper to realize ICA. To avoid assuming Gaussianity in hypothesis testing, we propose a nonparametric approach where the distributions of random variables are calculated using kernel density functions. A new ICA is then fulfilled through the NLR objective function. Interestingly, we apply the proposed NLR-ICA algorithm for unsupervised learning of unknown pronunciation variations. The clusters of speech hidden Markov models are estimated to characterize multiple pronunciations of subword units for robust speech recognition. Also, the NLR-ICA is applied to separate the linear mixture of speech and audio signals. In the experiments, NLR-ICA achieves better speech recognition performance compared to parametric and nonparametric minimum mutual information ICA. Jen-Tzung Chien, Bo-Cheng Chen |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Aggregate a posteriori linear regression adaptationabstractWe present a new discriminative linear regression adaptation algorithm for hidden Markov model (HMM) based speech recognition. The cluster-dependent regression matrices are estimated from speaker-specific adaptation data through maximizing the aggregate a posteriori probability, which can be expressed in a form of classification error function adopting the logarithm of posterior distribution as the discriminant function. Accordingly, the aggregate a posteriori linear regression (AAPLR) is developed for discriminative adaptation where the classification errors of adaptation data are minimized. Because the prior distribution of regression matrix is involved, AAPLR is geared with the Bayesian learning capability. We demonstrate that the difference between AAPLR discriminative adaptation and maximum a posteriori linear regression (MAPLR) adaptation is due to the treatment of the evidence. Different from minimum classification error linear regression (MCELR), AAPLR has closed-form solution to fulfil rapid adaptation. Experimental results reveal that AAPLR speaker adaptation does improve speech recognition performance with moderate computational cost compared to maximum likelihood linear regression (MLLR), MAPLR, MCELR and conditional maximum likelihood linear regression (CMLLR). These results are verified for supervised adaptation as well as unsupervised adaptation for different numbers of adaptation data. Jen-Tzung Chien, Chih-Hsien Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Independent component analysis using nonparametric likelihood ratio criterionabstractThe paper presents a novel nonparametric likelihood ratio criterion for independent component analysis (ICA). This criterion is derived through a statistical hypothesis test of the independence of random variables. A likelihood ratio (LR) criterion is developed to measure the strength of independence. We accordingly estimate the unmixing matrix by maximizing the LR function and apply it to transform data into independent component space. Conventionally, the test of independence was established assuming data distributions being Gaussian, which is improper to realize ICA. To prevent assuming Gaussianity in hypothesis testing, we propose a nonparametric approach where the distributions of random variables are calculated using kernel density functions and adopted for the estimation of the LR function. Finally, a new ICA is fulfilled using the nonparametric likelihood ratio (NLR) criterion. In the experiments, we apply the proposed ICA for blind source separation and speech recognition. The evaluation of using the NLR criterion shows good performance for the separation and recognition of speech signals. Jen-Tzung Chien, Bo-Cheng Chen |
ICASSP (5) | 1 |
| 2005 | Aggregate a Posteriori Linear Regression for Speaker AdaptationabstractIn this paper, we present a rapid and discriminative speaker adaptation algorithm for speech recognition. The adaptation paradigm is constructed under the popular linear regression transformation framework. Attractively, we estimate the regression matrices from the speaker-specific adaptation data according to the aggregate a posteriori criterion, which can be expressed in a form of classification error function. The goal of the proposed aggregate a posteriori linear regression (AAPLR) is to estimate the discriminative linear regression matrices for transformation-based adaptation so that the classification errors can be minimized. Different from minimum classification error linear regression (MCELR), the AAPLR algorithm has a closed-form solution to achieve rapid speaker adaptation. The experimental results reveal that AAPLR speaker adaptation does improve speech recognition performance with moderate computational cost compared to the maximum likelihood linear regression (MLLR), maximum a posteriori linear regression (MAPLR) and MCELR. Chih-Hsien Huang, Jen-Tzung Chien |
ICASSP (1) | 2 |
| 2005 | Nonsingular discriminant feature extraction for face recognitionabstractIn this paper, we present a nonsingular transformation prior to performing Fisher linear discriminant analysis (LDA). This method is used to transform general features using all eigenvectors of the scatter matrix with nonzero eigenvalues. As a result, the scatter matrix of transformed features is nonsingular. Subsequently, the discriminant transformation is applied according to LDA using the new scatter matrices. The superiority of nonsingular discriminant analysis of the between-class matrix comes from the shrinkage of within-class scatters and accordingly the enhancement of Fisher class separability. From experiments on facial databases, we find that the nonsingular discriminant feature extraction achieves significant face recognition performance compared to other LDA-related methods for a wide range of sample sizes and class numbers. Chih-Pin Liao, Jen-Tzung Chien |
ICASSP (2) | 2 |
| 2005 | Bayesian learning for latent semantic analysisabstractProbabilistic latent semantic analysis (PLSA) is a popular approach to text modeling where the semantics and statistics in documents can be effectively captured. In this paper, a novel Bayesian PLSA framework is presented. We focus on exploiting the incremental learning algorithm for solving the updating problem of new domain articles. This algorithm is developed to improve text modeling by incrementally extracting the up-to-date latent semantic information to match the changing domains at run time. The expectationmaximization (EM) algorithm is applied to resolve the quasiBayes (QB) estimate of PLSA parameters. The online PLSA is constructed to accomplish parameter estimation as well as hyperparameter updating. Compared to standard PLSA using maximum likelihood estimate, the proposed QB approach is capable of performing dynamic document indexing and classification. Also, we present the maximum a posteriori PLSA for corrective training. Experiments on evaluating model perplexities and classification accuracies demonstrate the superiority of using Bayesian PLSA. Jen-Tzung Chien, Meng-Sung Wu, Chia-Sheng Wu |
INTERSPEECH | 1 |
| 2005 | Discriminative maximum entropy language model for speech recognition
Chuang-Hua Chueh, To-Chang Chien, Jen-Tzung Chien |
INTERSPEECH | 3 |
| 2005 | Decision tree State tying using cluster validity criteriaabstractDecision tree state tying aims to perform divisive clustering, which can combine the phonetics and acoustics of speech signal for large vocabulary continuous speech recognition. A tree is built by successively splitting the observation frames of a phonetic unit according to the best phonetic questions. To prevent building over-large tree models, the stopping criterion is required to suppress tree growing. Accordingly, it is crucial to exploit the goodness-of-split criteria to choose the best questions for node splitting and test whether the splitting should be terminated or not. In this paper, we apply the Hubert's /spl Gamma/ statistic as the node splitting criterion and the T/sup 2/-statistic as the stopping criterion. The Hubert's /spl Gamma/ statistic sufficiently characterizes the clustering structure in the given data. This cluster validity criterion is adopted to select the best questions to unravel tree nodes. Further, we examine the population closeness of two split nodes with a significance level. The T/sup 2/-statistic expressed by an F distribution is determined to verify whether the mean vectors of two nodes are close together. The splitting is stopped when verified. In the experiments of Mandarin speech recognition, the proposed methods achieve better syllable recognition rates with smaller tree models compared to the conventional maximum likelihood and minimum description length criteria. Jen-Tzung Chien |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Predictive hidden Markov model selection for speech recognitionabstractThis paper surveys a series of model selection approaches and presents a novel predictive information criterion (PIC) for hidden Markov model (HMM) selection. The approximate Bayesian using Viterbi approach is applied for PIC selection of the best HMMs providing the largest prediction information for generalization of future data. When the perturbation of HMM parameters is expressed by a product of conjugate prior densities, the segmental prediction information is derived at the frame level without Laplacian integral approximation. In particular, a multivariate t distribution is attained to characterize the prediction information corresponding to HMM mean vector and precision matrix. When performing model selection in tree structure HMMs, we develop a top-down prior/posterior propagation algorithm for estimation of structural hyperparameters. The prediction information is determined so as to choose the best HMM tree model. Different from maximum likelihood (ML) and minimum description length (MDL) selection criteria, the parameters of PIC chosen HMMs are computed via maximum a posteriori estimation. In the evaluation of continuous speech recognition using decision tree HMMs, the PIC criterion outperforms ML and MDL criteria in building a compact tree structure with moderate tree size and higher recognition rate. Jen-Tzung Chien, Sadaoki Furui |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Bayesian duration modeling and learning for speech recognitionabstractWe present Bayesian duration modeling and learning for speech recognition under nonstationary speaking rates and noise conditions. In this study, the Gaussian, Poisson and gamma distributions are investigated, to characterize duration models. The maximum a posteriori (MAP) estimate of the gamma duration model is developed. To exploit the sequential learning, we adopt the Poisson duration model, incorporated with gamma prior density, which belongs to the conjugate prior family. When the adaptation data are sequentially observed, the gamma posterior density is produced for twofold advantages. One is to determine the optimal quasi-Bayes (QB) duration parameter, which can be merged in HMM's for speech recognition. The other one is to build the updating mechanism of gamma prior statistics for sequential learning. An expectation-maximization algorithm is applied to fulfill parameter estimation. In the experiments, the proposed Bayesian approaches significantly improve the speech recognition performance of Mandarin broadcast news. Batch and sequential learning are investigated for MAP and QB duration models, respectively. Jen-Tzung Chien, Chih-Hsien Huang |
ICASSP (1) | 1 |
| 2004 | Mining of association patterns for language modelingabstractLanguage modeling using n-gram is popular for speech recognition and many other applications. The conventional ngram suffers from the insufficiencies of training data, domain knowledge and long distance language dependencies. This paper presents a new approach to mining long distance word associations and incorporating their mutual information into language models. We aim to discover the associations of multiple distant words from training corpus. An efficient algorithm is exploited to merge the frequent word subsets and construct the association patterns. The resulting association pattern n-gram is general with a special realization to trigger pair n-gram where only associations of two distant words are considered. To improve the modeling, we further compensate the weaknesses of sparse training data via parameter smoothing and domain mismatch via online adaptive learning. The proposed association pattern n-gram and several hybrid models are successfully applied for speech recognition. We also find that the incorporation of mutual information of association patterns can significantly reduce the perplexities of language models. Jen-Tzung Chien, Hung-Ying Chen |
INTERSPEECH | 1 |
| 2004 | Speaker identification using probabilistic PCA model selectionabstractGaussian mixture model (GMM) techniques are popular for speaker identification. Theoretically, each Gaussian function should have a full covariance matrix. However, the diagonal covariance matrix is usually used because the inverse of diagonal covariance matrix can be easily calculated via expectation maximization (EM) algorithm. This paper proposes a new probabilistic principal component analysis (PPCA) model for speaker identification. The full covariance of speaker’s data is considered. This model is originated from factor analysis theory. The probability distributions using PPCA are well defined. In particular, GMM and PPCA are found to be equivalent when using diagonal covariance matrix. In this study, we derive a novel PPCA model selection and establish models for different speakers. Applying PPCA model selection, we can dynamically determine the numbers of speech features and mixture components. Experiments show that PPCA achieves desirable speaker recognition performance with proper model regularization. 1. Jen-Tzung Chien, Chuan-Wei Ting |
INTERSPEECH | 1 |
| 2004 | On latent semantic language modeling and smoothingabstractLanguage modeling plays a critical role for automatic speech recognition. Typically, the n-gram language models suffer from the lack of a good representation of historical words and an inability to estimate unseen parameters due to insufficient training data. In this study, we explore the application of latent semantic information (LSI) to language modeling and parameter smoothing. Our approach adopts latent semantic analysis to transform all words and documents into a common semantic space. The word-to-word, word-to-document and document-to-document relations are, accordingly, exploited for language modeling and smoothing. For language modeling, we present a new representation of historical words based on retrieval of the most relevant document. We also develop a novel parameter smoothing method, where the language models of seen and unseen words are estimated by interpolating the k nearest seen words in the training corpus. The interpolation coefficients are determined according to the closeness of words in the semantic space. As shown by experiments, the proposed modeling and smoothing methods can significantly reduce the perplexity of language models with moderate computational cost. Jen-Tzung Chien, Meng-Sung Wu, Hua-Jui Peng |
INTERSPEECH | 1 |
| 2003 | Predictive hidden Markov model selection for decision tree state tyingabstractThis paper presents a novel predictive information criterion (PIC) for hidden Markov model (HMM) selection.The PIC criterion is exploited to select the best HMMs, which provide the largest prediction information for generalization of future data.When the randomness of HMM parameters is expressed by a product of conjugate prior densities, the prediction information is derived without integral approximation.In particular, a multivariate t distribution is attained to characterize the prediction information corresponding to HMM mean vector and precision matrix.When performing HMM selection in tree structure HMMs, we develop a top-down prior/posterior propagation algorithm for estimation of structural hyperparameters.The prediction information is accordingly determined so as to choose the best HMM tree model.The parameters of chosen HMMs can be rapidly computed via maximum a posteriori (MAP) estimation.In the evaluation of continuous speech recognition using decision tree HMMs, the PIC model selection criterion performs better than conventional maximum likelihood and minimum description length criteria in building a compact tree structure with moderate tree size and higher recognition rate. Jen-Tzung Chien, Sadaoki Furui |
INTERSPEECH | 1 |
| 2003 | Linear regression based Bayesian predictive classification for speech recognitionabstractThe uncertainty in parameter estimation due to the adverse environments deteriorates the classification performance for speech recognition. It becomes crucial to incorporate the parameter uncertainty into decision so that the classification robustness can be assured. We propose a novel linear regression based Bayesian predictive classification (LRBPC) for robust speech recognition. This framework is constructed under the paradigm of linear regression adaptation of speech hidden Markov models (HMMs). Because the regression mapping between HMMs and adaptation data is ill posed, we properly characterize the uncertainty of regression parameters using a joint Gaussian distribution . A closed-form predictive distribution can be derived to set up the LRBPC decision for speech recognition. Such decision is robust compared to the plug-in maximum a posteriori (MAP) decision adopted in the maximum likelihood linear regression (MLLR) and MAP linear regression (MAPLR). Since the specified distribution belongs to the conjugate prior family, the evolutionary hyperparameters are established. With the statistically rich hyperparameters, the LRBPC achieves decision robustness. In the experiments, we find that LRBPC decision in cases of general linear regression as well as single variable linear regression attains significantly better recognition performance than MLLR and MAPLR adaptation. Jen-Tzung Chien |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Bayesian learning of speech duration modelsabstractThis paper presents the Bayesian speech duration modeling and learning for hidden Markov model (HMM) based speech recognition. We focus on the sequential learning of HMM state duration using quasi-Bayes (QB) estimate. The adapted duration models are robust to nonstationary speaking rates and noise conditions. In this study, the Gaussian, Poisson, and gamma distributions are investigated to characterize the duration models. The maximum a posteriori (MAP) estimate of gamma duration model is developed. To exploit the sequential learning, we adopt the Poisson duration model incorporated with gamma prior density, which belongs to the conjugate prior family. When the adaptation data are sequentially observed, the gamma posterior density is produced with twofold advantages. One is to determine the optimal QB duration parameter, which can be merged in HMMs for speech recognition. The other one is to build the updating mechanism of gamma prior statistics for sequential learning. EM algorithm is applied to fulfill QB parameter estimation. The adaptation of overall HMM parameters can be performed simultaneously. In the experiments, the proposed adaptive duration model improves the speech recognition performance of Mandarin broadcast news and noisy connected digits. The batch and sequential learning are respectively investigated for MAP and QB duration models. Jen-Tzung Chien, Chih-Hsien Huang |
IEEE Trans. Speech Audio Process. | 1 |
| 2002 | Compact decision trees with cluster validity for speech recognitionabstractA decision tree is built by successively splitting the observation frames of a phonetic unit according to the best phonetic questions. To prevent over-large tree models, the stopping criterion is required to suppress tree growing. It is crucial to exploit the goodness-of-split criteria to choose the best questions for node splitting and test if the hypothesis of splitting should be terminated. The robust tree models could be established. In this study, we apply the Hubert's Γ statistic as the node splitting criterion and the T2-statistic as the stopping criterion. Hubert's Γ statistic is a cluster validity measure, which characterizes the degree of clustering in the available data. This measure is useful to select the best questions to unravel tree nodes. Further, we examine the population closeness of two child nodes with a significant level, T2-statistic is determined to validate whether the corresponding mean vectors are close together. The splitting is stopped when validated. In continuous speech recognition experiments, the proposed methods achieve better recognition rates with smaller tree models compared to the maximum likelihood and minimum description length criteria. Jen-Tzung Chien, Chih-Hsien Huang, Shun-Ju Chen |
ICASSP | 1 |
| 2002 | Discriminant Waveletfaces and Nearest Feature Classifiers for Face RecognitionabstractFeature extraction, discriminant analysis, and classification rules are three crucial issues for face recognition. We present hybrid approaches to handle three issues together. For feature extraction, we apply the multiresolution wavelet transform to extract the waveletface. We also perform the linear discriminant analysis on waveletfaces to reinforce discriminant power. During classification, the nearest feature plane (NFP) and nearest feature space (NFS) classifiers are explored for robust decisions in presence of wide facial variations. Their relationships to conventional nearest neighbor and nearest feature line classifiers are demonstrated. In the experiments, the discriminant waveletface incorporated with the NFS classifier achieves the best face recognition performance. Jen-Tzung Chien, Chia-Chen Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Adaptive hierarchy of hidden Markov models for transformation-based adaptation
Jen-Tzung Chien |
Speech Commun. | 1 |
| 2002 | A Bayesian prediction approach to robust speech recognition and online environmental learning
Jen-Tzung Chien |
Speech Commun. | 1 |
| 2002 | Quasi-Bayes linear regression for sequential learning of hidden Markov modelsabstractThis paper presents an online/sequential linear regression adaptation framework for hidden Markov model (HMM) based speech recognition. Our attempt is to sequentially improve speaker-independent speech recognition system to handle the nonstationary environments via the linear regression adaptation of HMMs. A quasi-Bayes linear regression (QBLR) algorithm is developed to execute the sequential adaptation where the regression matrix is estimated using QB theory. In the estimation, we specify the prior density of regression matrix as a matrix variate normal distribution and derive the pooled posterior density belonging to the same distribution family. Accordingly, the optimal regression matrix can be easily calculated. Also, the reproducible prior/posterior pair provides a meaningful mechanism for sequential learning of prior statistics. At each sequential epoch, only the updated prior statistics and the current observed data are required for adaptation. The proposed QBLR is a general framework with maximum likelihood linear regression (MLLR) and maximum a posteriori linear regression (MAPLR) as special cases. Experiments on supervised and unsupervised speaker adaptation demonstrate that the sequential adaptation using QBLR is efficient and asymptotical to batch learning using MLLR and MAPLR in recognition performance. Jen-Tzung Chien |
IEEE Trans. Speech Audio Process. | 1 |
| 2001 | Online speaker adaptation based on quasi-Bayes linear regressionabstractThis paper presents an online/sequential linear regression adaptation framework for hidden Markov model (HMM) based speech recognition. Our attempt is to sequentially improve the speaker-independent (SI) speech recognizer to meet nonstationary environments via linear regression adaptation of SI HMMs. A quasi-Bayes linear regression (QBLR) algorithm is developed to execute online adaptation where the regression matrix is estimated using QB theory. In the estimation, we moderately specify the prior density of the regression matrix as a matrix variate normal distribution and exactly derive the pooled posterior density belonging to the same distribution family. Accordingly, the optimal regression matrix can be easily calculated. Also, the reproducible prior/posterior density pair provides a meaningful mechanism for sequential learning of prior statistics. At each sequential epoch, only the updated prior statistics and the current observed data are required for adaptation. In general, the proposed QBLR is universal and can be reduced to the well-known maximum likelihood linear regression (MLLR) and maximum a posteriori linear regression (MAPLR). Experiments show that the QBLR is effective for speaker adaptation in car environments. Jen-Tzung Chien, Chih-Hsien Huang |
ICASSP | 1 |
| 2001 | Combined linear regression adaptation and Bayesian predictive classification for robust speech recognitionabstractThe uncertainty in parameter estimation due to the adverse environments deteriorates the speech recognition performance. It becomes crucial to incorporate the parameter uncertainty into decision so that the classification robustness can be assured. In this paper, we propose a linear regression based Bayesian predictive classification (LRBPC) for robust speech recognition. This framework is constructed under the paradigm of linear regression adaptation of HMM’s. Because the regression mapping between HMM’s and adaptation data is ill posed, we properly characterize the uncertainty of regression parameters using a joint Gaussian distribution. A predictive distribution is derived to set up the LRBPC decision. Such decision is robust compared to the plug-in maximum a posteriori decision adopted in the maximum likelihood linear regression (MLLR). Since the specified distribution belongs to the conjugate prior family, the evolutionary hyperparameter is established. With the hyperparameter, the LRBPC achieves significantly better performance than MLLR adaptation in car speech recognition. Jen-Tzung Chien |
INTERSPEECH | 1 |
| 2001 | Transformation-based Bayesian predictive classification using online prior evolutionabstractThe mismatch between training and testing environments makes the necessity of speech recognizers to be adaptive both in acoustic modeling and decision making. Accordingly, the speech hidden Markov models (HMMs) should be able to incrementally capture the evolving statistics of environments using online available data. Also, it is necessary for speech recognizers to exploit the robust decision strategy, which takes the uncertainty of parameters into account. This paper presents a transformation-based Bayesian predictive classification (TBPC) where the uncertainty of the transformation parameters of the HMM mean vector and precision matrix is adequately represented by a joint multivariate prior density of normal-Wishart belonging to the conjugate family. The formulation of TBPC decision is correspondingly constructed. Due to the benefit of conjugate density, we generate the reproducible prior/posterior pair such that the hyperparameters of prior density could evolve successively to new environments using online test/adaptation data. The evolved hyperparameters could suitably describe the parameter uncertainty for TBPC decision. Therefore, a novel framework of TBPC geared with online prior evolution (OPE) capability is developed for robust speech recognition. This framework is examined to be effective as well as efficient on the recognition task of connected Chinese digits in hands-free car environments. Jen-Tzung Chien, Guo-Hong Liao |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Transformation-based Bayesian predictive classification for online environmental learning and robust speech recognition
Jen-Tzung Chien, Guo-Hong Liao |
INTERSPEECH | 1 |
| 2000 | Unsupervised hierarchical adaptation using reliable selection of cluster-dependent parameters
Jen-Tzung Chien, Jean-Claude Junqua |
Speech Commun. | 1 |
| 1999 | N-best based supervised and unsupervised adaptation for native and non-native speakers in carsabstractA new set of techniques exploiting N-best hypotheses in supervised and unsupervised adaptation are presented. These techniques combine statistics extracted from the N-best hypotheses with a weight derived from a likelihood ratio confidence measure. In the case of supervised adaptation the knowledge of the correct string is used to perform N-best based corrective adaptation. Experiments run for continuous letter recognition recorded in a car environment show that weighting N-best sequences by a likelihood ratio confidence measure provides only marginal improvement as compared to 1-best unsupervised adaptation and N-best unsupervised adaptation with equal weighting. However, an N-best based supervised corrective adaptation method weighting correct letters positively and incorrect letters negatively, resulted in a 13% decrease of the error rate as compared with supervised adaptation. The largest improvement was obtained for non-native speakers. Patrick Nguyen, Philippe Gelin, Jean-Claude Junqua, Jen-Tzung Chien |
ICASSP | 4 |
| 1999 | Extraction of reliable transformation parameters for unsupervised speaker adaptationabstractAdaptation of speaker-independent hidden Markov models (HMM’s) to a new speaker using speaker-specific data is an effective approach to reinforce speech recognition performance for the enrolled speaker. Practically, it is desirable to flexibly perform the adaptation without any knowledge or limitation on the enrolled adaptation data (e.g. data transcription, length and content). However, the inevitable transcription errors on adaptation data may cause unreliability in model adaptation. The variable amount and content of adaptation data require the algorithm to dynamically control the degrees of sharing in transformation-based adaptation. This paper presents an unsupervised hierarchical adaptation algorithm where a tree structure of HMM’s is incorporated to control the transformation sharing. To extract reliable transformation parameters, we exploit the reliability assessment criteria using the confidence measure and description length. Experiments show that the unsupervised speaker adaptation with reliability assessment can significantly improve the recognition performance for any lengths of adaptation data. Jen-Tzung Chien, Jean-Claude Junqua, Philippe Gelin |
EUROSPEECH | 1 |
| 1999 | Online hierarchical transformation of hidden Markov models for speech recognitionabstractThis paper proposes a novel framework of online hierarchical transformation of hidden Markov model (HMM) parameters for adaptive speech recognition. Our goal is to incrementally transform (or adapt) all the HMM parameters to a new acoustical environment even though most of HMM units are unseen in observed adaptation data. We establish a hierarchical tree of HMM units and apply the tree to dynamically search the transformation parameters for individual HMM mixture components. In this paper, the transformation framework formulated according to the approximate Bayesian estimate, where the prior statistics and the transformation parameters can be jointly and incrementally refreshed after each consecutive adaptation data, is presented. Using this formulation, only the refreshed prior statistics and the current block of data are needed for online transformation. In a series of speaker adaptation experiments on the recognition of 408 Mandarin syllables, we examine the effects on constructing various types of hierarchical trees. The efficiency and effectiveness of proposed method on incremental adaptation of overall HMM units are also confirmed. Besides, we demonstrate the superiority of proposed online transformation to Huo's (see ibid., vol.5, p.161-72, 1997) on-line adaptation for a wide range of adaptation data. Jen-Tzung Chien |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | On-line hierarchical transformation of hidden Markov models for speaker adaptationabstractThis paper presents a novel framework of on-line hierarchical transformation of hidden Markov models (HMM’s) for speaker adaptation. Our aim is to incrementally transform (or adapt) all the HMM parameters to a new speaker even though part of HMM units are unseen in adaptation data. The transformation paradigm is formulated according to the approximate Bayesian estimate, which the prior statistics and the transformation parameters are incrementally updated for each consecutive adaptation data. Using this formulation, the updated prior statistics and the current block of data are sufficient for on-line transformation. Further, we establish a hierarchical tree of HMM’s and use it to dynamically control the transformation sharing for each HMM unit. In the speaker adaptation experiments, we demonstrate the superiority of proposed on-line transformation to other method. Jen-Tzung Chien |
ICSLP | 1 |
| 1998 | A novel projection-based likelihood measure for noisy speech recognition
Jen-Tzung Chien, Hsiao-Chuan Wang, Lee-Min Lee |
Speech Commun. | 1 |
| 1998 | Phone-dependent channel compensated hidden Markov model for telephone speech recognitionabstractWe propose the phone-dependent channel compensated hidden Markov model (PDCC-HMM) for telephone speech recognition. The PDCC-HMM is derived by modifying the conventional hidden Markov model (HMM) with the phone-dependent channel compensation vectors. The telephone speech is recognized efficiently by using the derived PDCC-HMM. Experiments demonstrate the robustness of PDCC-HMM in speech recognition and show the significant reduction of recognition error rate by 50% compared to the conventional HMM method. Jen-Tzung Chien, Hsiao-Chuan Wang |
IEEE Signal Process. Lett. | 1 |
| 1997 | Improved Bayesian learning of hidden Markov models for speaker adaptationabstractWe propose an improved maximum a posteriori (MAP) learning algorithm of continuous-density hidden Markov model (CDHMM) parameters for speaker adaptation. The algorithm is developed by sequentially combining three adaptation approaches. First, the clusters of speaker-independent HMM parameters are locally transformed through a group of transformation functions. Then, the transformed HMM parameters are globally smoothed via the MAP adaptation. Within the MAP adaptation, the parameters of unseen units in adaptation data are further adapted by employing the transfer vector interpolation scheme. Experiments show that the combined algorithm converges rapidly and outperforms those other adaptation methods. Jen-Tzung Chien, Hsiao-Chuan Wang |
ICASSP | 1 |
| 1997 | Bayesian affine transformation of HMM parameters for instantaneous and supervised adaptation in telephone speech recognitionabstractThis paper proposes a Bayesian affine transformation of hidden Markov model (HMM) parameters for reducing the acoustic mismatch problem in telephone speech recognition. Our purpose is to transform the existing HMM parameters into its new version of specific telephone environment using affine function so as to improve the recognition rate. The maximum a posteriori (MAP) estimation which merges the prior statistics into transformation is applied for estimating the transformation parameters. Experiments demonstrate that the proposed Bayesian affine transformation is effective for instantaneous adaptation and supervised adaptation in telephone speech recognition. Model transformation using MAP estimation performs better than that using maximum-likelihood (ML) estimation. Jen-Tzung Chien, Hsiao-Chuan Wang |
EUROSPEECH | 1 |
| 1997 | Telephone speech recognition based on Bayesian adaptation of hidden Markov models
Jen-Tzung Chien, Hsiao-Chuan Wang |
Speech Commun. | 1 |
| 1997 | A hybrid algorithm for speaker adaptation using MAP transformation and adaptationabstractWe present a hybrid algorithm for adapting a set of speaker-independent hidden Markov models (HMMs) to a new speaker based on a combination of maximum a posteriori (MAP) parameter transformation and adaptation. The algorithm is developed by first transforming clusters of HMM parameters through a class of transformation functions. Then, the transformed HMM parameters are further smoothed via Bayesian adaptation. The proposed transformation/adaptation process can be iterated for any given amount of adaptation data, and it converges rapidly in terms of likelihood improvement. The algorithm also gives a better speech recognition performance than that obtained using transformation or adaptation alone for almost any practical amount of adaptation data. Jen-Tzung Chien, Hsiao-Chuan Wang |
IEEE Signal Process. Lett. | 1 |
| 1996 | Noisy speech recognition using variance adapted likelihood measureabstractBecause the norm of testing cepstral vector was shrinked in a noisy environment, the model parameters, i.e., mean vector and covariance matrix, should be adapted simultaneously. We propose a method called variance adapted likelihood measure (VALM) which adapts the mean vector using a projection-based scale factor and adapts the covariance matrix using a variance reduction function estimated from the training database. The variance reduction function can be obtained according to various phonetic units. In the hidden Markov model based experiments, the speech recognition performance is greatly improved by applying VALM. The most significant improvement is achieved when the variance reduction function is separately estimated for different state parameters. Jen-Tzung Chien, Lee-Min Lee, Hsiao-Chuan Wang |
ICASSP | 1 |
| 1996 | Estimation of channel bias for telephone speech recognitionabstractIn this study, we propose a maximum a posterior (MAP) estimation of channel bias to compensate the channel mismatch in telephone speech recognition.For a telephone speech, the channel bias is estimated by maximizing a posterior probability.Because a posterior probability is composed of a likelihood function and a prior density, we introduce a scale factor to evaluate their weights in MAP estimation.To further improve the performance, a prior channel statistics is extended to multiple components and the channel mismatch is separately compensated for different segments.Besides, a rapid MAP estimation applied in feature domain is also proposed for reducing the computational complexity.Experiments show that proposed method can significantly improve recognition rates and computational complexity. Jen-Tzung Chien, Hsiao-Chuan Wang, Lee-Min Lee |
ICSLP | 1 |
| 1995 | Channel estimation for reference model adaptation in telephone speech recognition
Jen-Tzung Chien, Lee-Min Lee, Hsiao-Chuan Wang |
EUROSPEECH | 1 |