EDBT 2026 Demo / reviewers in the wild / expert
Zhijian Ou
dblp:44/5256
· DBLP profile ↗
56ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-9018-5074ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 14 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Dual Consistency Training (DCT) strategy for polyphonic sound event detection
Ying Hu 0005, Xinchun Ma, Zhijian Ou |
Neurocomputing | 5 |
| 2025 | A Joint Network for Singing Melody Extraction from Polyphonic Music with Attention Aggregation and Self-Consistency Training
Jiabo Jing, Ying Hu 0005, Hao Huang 0009, Liang He 0003, Zhijian Ou |
INTERSPEECH | 5 |
| 2025 | Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
Xiangzhu Kong, Hao Huang 0009, Zhijian Ou |
INTERSPEECH | 3 |
| 2025 | LLM-based phoneme-to-grapheme for phoneme-based speech recognition
Te Ma, Min Bi, Saierdaer Yusuyin, Hao Huang 0009, Zhijian Ou |
INTERSPEECH | 5 |
| 2024 | UniPCM: Universal Pre-trained Conversation Model with Task-aware Automatic PromptabstractRecent researches have shown that multi-task instruction tuning after pre-training greatly improves the model’s robustness and transfer ability, which is crucial for building a high-quality dialog system. However, most previous works on multi-task instruction tuning rely heavily on human-defined input format or prompt, which is not optimal in quality and quantity.In this work, we propose to use Task-aware Automatic Prompt generation (TAP) to automatically generate high-quality prompts. Using the high-quality prompts generated, we scale the corpus of the pre-trained conversation model to 122 datasets from 15 dialog-related tasks, resulting in Universal Pre-trained Conversation Model (UniPCM), a powerful foundation model for various conversational tasks and different dialog systems. Extensive experiments have shown that UniPCM is robust to input prompts and capable of various dialog-related tasks. Moreover, UniPCM has strong transfer ability and excels at low resource scenarios, achieving SOTA results on 9 different datasets ranging from task-oriented dialog to open-domain conversation. Furthermore, we are amazed to find that TAP can generate prompts on par with those collected with crowdsourcing. Yucheng Cai, Yuchuan Wu, Shuzheng Si, Yuan Shao, Zhijian Ou |
LREC/COLING | 6 |
| 2024 | The 2nd Futuredial Challenge: Dialog Systems With Retrieval Augmented Generation (Futuredial-RAG)abstractRecently, increasing research interests have focused on retrieval augmented generation (RAG) to mitigate hallucination for large language models (LLMs). Following this trend, we launch the FutureDial-RAG challenge at SLT 2024, which aims at promoting the study of RAG for dialog systems. The challenge builds upon the MobileCS2 dataset, a real-life customer service datasets with nearly 3000 high-quality dialogs containing annotations for knowledge base query and corresponding results. Over the dataset, we define two tasks, track 1 for knowledge retrieval and track 2 for response generation, which are core research questions in dialog systems with RAG. We build baseline systems for the two tracks and design metrics to measure whether the systems can perform accurate retrieval and generate informative and coherent response. The baseline results show that it is very challenging to perform well on the two tasks, which encourages the participating teams and the community to study how to make better use of RAG for real-life dialog systems. Yucheng Cai, Yi Huang 0017, Junlan Feng, Zhijian Ou |
SLT | 6 |
| 2023 | Prompt Pool Based Class-Incremental Continual Learning for Dialog State TrackingabstractContinual learning is crucial for dialog state tracking (DST) in dialog systems, since requirements from users for new functionalities are often encountered. However, most of existing continual learning methods for DST require task identities during testing, which is a severe limit in real-world applications. In this paper, we aim to address continual learning of DST in the class-incremental scenario (namely the task identity is unknown in testing). Inspired by the recently emerging prompt tuning method that performs well on dialog systems, we propose to use the prompt pool method, where we maintain a pool of key-value paired prompts and select prompts from the pool according to the distance between the dialog history and the prompt keys. The proposed method can automatically identify tasks and select appropriate prompts during testing. We conduct experiments on Schema-Guided Dialog dataset (SGD) and another dataset collected from a real-world dialog application. Experiment results show that the prompt pool method achieves much higher joint goal accuracy than the baseline. After combining with a rehearsal buffer, the model performance can be further improved. Hong Liu 0024, Yucheng Cai, Zhijian Ou, Yi Huang 0017, Junlan Feng |
ASRU | 4 |
| 2023 | Knowledge-Retrieval Task-Oriented Dialog Systems with Semi-Supervision
Yucheng Cai, Hong Liu 0024, Zhijian Ou, Yi Huang 0017, Junlan Feng |
INTERSPEECH | 3 |
| 2023 | Exploring Energy-based Language Models with Different Architectures and Training Methods for Speech Recognition
Hong Liu 0024, Zhaobiao Lv, Zhijian Ou |
INTERSPEECH | 3 |
| 2023 | Callee: Recovering Call Graphs for Binaries with Transfer and Contrastive LearningabstractRecovering binary programs’ call graphs is crucial for inter-procedural analysis tasks and applications based on them. One of the core challenges is recognizing targets of indirect calls (i.e., indirect callees). Existing solutions all have high false positives and negatives, making call graphs inaccurate. In this paper, we propose a new solution Callee combining transfer learning and contrastive learning. The key insight is that, deep neural networks (DNNs) can automatically identify patterns concerning indirect calls. Inspired by the advances in question-answering applications, we utilize contrastive learning to answer the callsite-callee question. However, one of the toughest challenges is that DNNs need large datasets to achieve high performance, while collecting large-scale indirect-call ground truths can be computational-expensive. Therefore, we leverage transfer learning to pre-train DNNs with easy-to-collect direct calls and further fine-tune DNNs for indirect-calls. We evaluate Callee on several groups of targets, and results show that our solution could match callsites to callees with an F1-Measure of 94.6%, much better than state-of-the-art solutions. Further, we apply Callee to two applications – binary code similarity detection and hybrid fuzzing, and found it could greatly improve their performance. Wenyu Zhu, Zhiyao Feng, Jianjun Chen 0005, Zhijian Ou, Min Yang 0002, Chao Zhang 0008 |
SP | 5 |
| 2023 | Variational Latent-State GPT for Semi-Supervised Task-Oriented Dialog SystemsabstractRecently, two approaches, fine-tuning large pre-trained language models and variational training, have attracted significant interests, separately, for semi-supervised end-to-end task-oriented dialog (TOD) systems. In this paper, we propose Variational Latent-State GPT model (VLS-GPT), which is the first to combine the strengths of the two approaches. Among many options of models, we propose the generative model and the inference model for variational learning of the end-to-end TOD system, both as auto-regressive language models based on GPT-2, which can be further trained over a mix of labeled and unlabeled dialog data in a semi-supervised manner. Variational training of VLS-GPT is both statistically and computationally more challenging than previous variational learning works for sequential latent variable models, which use turn-level first-order Markovian. The inference model in VLS-GPT is non-Markovian due to the use of the Transformer architecture. In this work, we establish Recursive Monte Carlo Approximation (RMCA) to the variational objective with non-Markovian inference model and prove its unbiasedness. Further, we develop the computational strategy of sampling-then-forward-computation to realize RMCA, which successfully overcomes the memory explosion issue of using GPT in variational learning and speeds up training. Semi-supervised TOD experiments are conducted on two benchmark multi-domain datasets of different languages - MultiWOZ2.1 and CrossWOZ. VLS-GPT is shown to significantly outperform both supervised-only and semi-supervised self-training baselines. Hong Liu 0024, Yucheng Cai, Zhenru Lin, Zhijian Ou, Yi Huang 0017, Junlan Feng |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | CUSIDE: Chunking, Simulating Future Context and Decoding for Streaming ASR
Keyu An, Huahuan Zheng, Zhijian Ou, Hongyu Xiang, Guanglu Wan |
INTERSPEECH | 3 |
| 2022 | An Empirical Study of Language Model Integration for Transducer based Speech RecognitionabstractUtilizing text-only data with an external language model (ELM) in end-to-end RNN-Transducer (RNN-T) for speech recognition is challenging.Recently, a class of methods such as density ratio (DR) and internal language model estimation (ILME) have been developed, outperforming the classic shallow fusion (SF) method.The basic idea behind these methods is that RNN-T posterior should first subtract the implicitly learned internal language model (ILM) prior, in order to integrate the ELM.While recent studies suggest that RNN-T only learns some low-order language model information, the DR method uses a well-trained neural language model with full context, which may be inappropriate for the estimation of ILM and deteriorate the integration performance.Based on the DR method, we propose a loworder density ratio method (LODR) by replacing the estimation with a low-order weak language model.Extensive empirical experiments are conducted on both in-domain and cross-domain scenarios on English LibriSpeech & Tedlium-2 and Chinese WenetSpeech & AISHELL-1 datasets.It is shown that LODR consistently outperforms SF in all tasks, while performing generally close to ILME and better than DR in most tests. Huahuan Zheng, Keyu An, Zhijian Ou, Guanglu Wan |
INTERSPEECH | 3 |
| 2022 | Advancing Semi-Supervised Task Oriented Dialog Systems by JSA Learning of Discrete Latent Variable ModelsabstractDeveloping semi-supervised task-oriented dialog (TOD) systems by leveraging unlabeled dialog data has attracted increasing interests.For semi-supervised learning of latent state TOD models, variational learning is often used, but suffers from the annoying high-variance of the gradients propagated through discrete latent variables and the drawback of indirectly optimizing the target log-likelihood.Recently, an alternative algorithm, called joint stochastic approximation (JSA), has emerged for learning discrete latent variable models with impressive performances.In this paper, we propose to apply JSA to semi-supervised learning of the latent state TOD models, which is referred to as JSA-TOD.To our knowledge, JSA-TOD represents the first work in developing JSA based semi-supervised learning of discrete latent variable conditional models for such long sequential generation problems like in TOD systems.Extensive experiments show that JSA-TOD significantly outperforms its variational learning counterpart.Remarkably, semi-supervised JSA-TOD using 20% labels performs close to the full-supervised baseline on MultiWOZ2.1. Yucheng Cai, Hong Liu 0024, Zhijian Ou, Yi Huang 0017, Junlan Feng |
SIGDIAL | 3 |
| 2022 | Building Markovian Generative Architectures Over Pretrained LM Backbones for Efficient Task-Oriented Dialog SystemsabstractRecently, Transformer based pretrained language models (PLMs), such as GPT2 and T5, have been leveraged to build generative task-oriented dialog (TOD) systems. A drawback of existing PLM-based models is their non-Markov architectures across turns, i.e., the whole history is used as the conditioning input at each turn. First, this brings inefficiencies in memory and computation. Furthermore, using the whole history increases model complexity and may hurt the training efficiency, especially when facing small amounts of labeled training data (the low-resource setting). In this paper, motivated by the observation that dialog states could be viewed as Markov states, we propose to build Markovian Generative Architectures (MGA) over PLM backbones for efficient TOD systems. Experiments on MultiWOZ2.1 show that in the rich-resource setting, the proposed Markov models reduce memory and time costs without performance degradation; in the low-resource setting, the training efficiency of the Markov models is more significant. Hong Liu 0024, Yucheng Cai, Zhijian Ou, Yi Huang 0017, Junlan Feng |
SLT | 3 |
| 2021 | Multilingual and Crosslingual Speech Recognition Using Phonological-Vector Based Phone EmbeddingsabstractThe use of phonological features (PFs) potentially allows language-specific phones to remain linked in training, which is highly desirable for information sharing for multilingual and crosslingual speech recognition methods for low-resourced languages. A drawback suffered by previous methods in using phonological features is that the acoustic-to-PF extraction in a bottom-up way is itself difficult. In this paper, we propose to join phonology driven phone embedding (top-down) and deep neural network (DNN) based acoustic feature extraction (bottom-up) to calculate phone probabilities. The new method is called JoinAP (Joining of Acoustics and Phonology). Remarkably, no inversion from acoustics to phonological features is required for speech recognition. For each phone in the IPA (International Phonetic Alphabet) table, we encode its phonological features to a phonological-vector, and then apply linear or nonlinear transformation of the phonological-vector to obtain the phone embedding. A series of multilingual and crosslingual (both zero-shot and few-shot) speech recognition experiments are conducted on the CommonVoice dataset (German, French, Spanish and Italian) and the AISHLL-1 dataset (Mandarin), and demonstrate the superiority of JoinAP with nonlinear phone embeddings over both JoinAP with linear phone embeddings and the traditional method with flat phone embeddings. Chengrui Zhu, Keyu An, Huahuan Zheng, Zhijian Ou |
ASRU | 4 |
| 2021 | Deformable TDNN with Adaptive Receptive Fields for Speech RecognitionabstractTime Delay Neural Networks (TDNNs) are widely used in both DNN-HMM based hybrid speech recognition systems and recent end-to-end systems.Nevertheless, the receptive fields of TDNNs are limited and fixed, which is not desirable for tasks like speech recognition, where the temporal dynamics of speech are varied and affected by many factors.In this paper, we propose to use deformable TDNNs for adaptive temporal dynamics modeling in end-to-end speech recognition.Inspired by deformable ConvNets, deformable TDNNs augment the temporal sampling locations with additional offsets and learn the offsets automatically based on the ASR criterion, without additional supervision.Experiments show that deformable TDNNs obtain state-of-the-art results on WSJ benchmarks (1.42%/3.45%WER on WSJ eval92/dev93 respectively), outperforming standard TDNNs significantly.Furthermore, we propose the latency control mechanism for deformable TDNNs, which enables deformable TDNNs to do streaming ASR without accuracy degradation. Keyu An, Zhijian Ou |
Interspeech | 3 |
| 2021 | The SLT 2021 Children Speech Recognition Challenge: Open Datasets, Rules and BaselinesabstractAutomatic speech recognition (ASR) has been significantly advanced with the use of deep learning and big data. How-ever improving robustness, including achieving equally good performance on diverse speakers and accents, is still a challenging problem. In particular, the performance of children speech recognition (CSR) still lags behind due to 1) the speech and language characteristics of children's voice are substantially different from those of adults and 2) sizable open dataset for children speech is still not available in the research community. To address these problems, we launch the Children Speech Recognition Challenge (CSRC), as a flagship satellite event of IEEE SLT 2021 workshop. The challenge will release about 400 hours of Mandarin speech data for registered teams and set up two challenge tracks and provide a common testbed to benchmark the CSR performance. In this paper, we introduce the datasets, rules, evaluation method as well as baselines. Fan Yu 0002, Zhuoyuan Yao, Keyu An, Lei Xie 0001, Zhijian Ou, Xiulin Li, Guanqiong Miao |
SLT | 6 |
| 2021 | Efficient Neural Architecture Search for End-to-End Speech Recognition Via Straight-Through GradientsabstractNeural Architecture Search (NAS), the process of automating architecture engineering, is an appealing next step to advancing end-to-end Automatic Speech Recognition (ASR), replacing expert-designed networks with learned, task-specific architectures. In contrast to early computational-demanding NAS methods, recent gradient-based NAS methods, e.g., DARTS (Differentiable ARchiTecture Search), SNAS (Stochastic NAS) and ProxylessNAS, significantly improve the NAS efficiency. In this paper, we make two contributions. First, we rigorously develop an efficient NAS method via Straight-Through (ST) gradients, called ST-NAS. Basically, ST-NAS uses the loss from SNAS but uses ST to back-propagate gradients through discrete variables to optimize the loss, which is not revealed in ProxylessNAS. Using ST gradients to support sub-graph sampling is a core element to achieve efficient NAS beyond DARTS and SNAS. Second, we successfully apply ST-NAS to end-to-end ASR. Experiments over the widely benchmarked 80-hour WSJ and 300-hour Switchboard datasets show that the ST-NAS induced architectures significantly outperform the human-designed architecture across the two datasets. Strengths of ST-NAS such as architecture transferability and low computation cost in memory and time are also reported. Huahuan Zheng, Keyu An, Zhijian Ou |
SLT | 3 |
| 2020 | Task-Oriented Dialog Systems That Consider Multiple Appropriate Responses under the Same ContextabstractConversations have an intrinsic one-to-many property, which means that multiple responses can be appropriate for the same dialog context. In task-oriented dialogs, this property leads to different valid dialog policies towards task completion. However, none of the existing task-oriented dialog generation approaches takes this property into account. We propose a Multi-Action Data Augmentation (MADA) framework to utilize the one-to-many property to generate diverse appropriate dialog responses. Specifically, we first use dialog states to summarize the dialog history, and then discover all possible mappings from every dialog state to its different valid system actions. During dialog system training, we enable the current dialog state to map to all valid system actions discovered in the previous process to create additional state-action pairs. By incorporating these additional pairs, the dialog policy learns a balanced action distribution, which further guides the dialog model to generate diverse responses. Experimental results show that the proposed framework consistently improves dialog policy diversity, and results in improved response diversity and appropriateness. Our model obtains state-of-the-art results on MultiWOZ. Yichi Zhang 0001, Zhijian Ou, Zhou Yu 0005 |
AAAI | 2 |
| 2020 | Paraphrase Augmented Task-Oriented Dialog GenerationabstractNeural generative models have achieved promising performance on dialog generation tasks if given a huge data set.However, the lack of high-quality dialog data and the expensive data annotation process greatly limit their application in real-world settings.We propose a paraphrase augmented response generation (PARG) framework that jointly trains a paraphrase model and a response generation model to improve the dialog generation performance.We also design a method to automatically construct paraphrase training data set based on dialog state and dialog act labels.PARG is applicable to various dialog generation models, such as TSCP (Lei et al., 2018) and DAMD (Zhang et al., 2019).Experimental results show that the proposed framework improves these state-of-the-art dialog models further on CamRest676 and MultiWOZ.PARG also significantly outperforms other data augmentation methods in dialog generation tasks, especially under low resource settings.1 2 Silin Gao, Yichi Zhang 0001, Zhijian Ou, Zhou Yu 0005 |
ACL | 3 |
| 2020 | A Probabilistic End-To-End Task-Oriented Dialog Model with Latent Belief States towards Semi-Supervised LearningabstractStructured belief states are crucial for user goal tracking and database query in task-oriented dialog systems.However, training belief trackers often requires expensive turn-level annotations of every user utterance.In this paper we aim at alleviating the reliance on belief state labels in building end-to-end dialog systems, by leveraging unlabeled dialog data towards semi-supervised learning.We propose a probabilistic dialog model, called the LAtent BElief State (LABES) model, where belief states are represented as discrete latent variables and jointly modeled with system responses given user inputs.Such latent variable modeling enables us to develop semi-supervised learning under the principled variational learning framework.Furthermore, we introduce LABES-S2S, which is a copyaugmented Seq2Seq model instantiation of LABES 1 .In supervised experiments, LABES-S2S obtains strong results on three benchmark datasets of different scales.In utilizing unlabeled dialog data, semi-supervised LABES-S2S significantly outperforms both supervisedonly and semi-supervised baselines.Remarkably, we can reduce the annotation demands to 50% without performance loss on MultiWOZ. Yichi Zhang 0001, Zhijian Ou, Junlan Feng |
EMNLP (1) | 2 |
| 2020 | Integrating Discrete and Neural Features Via Mixed-Feature Trans-Dimensional Random Field Language ModelsabstractThere has been a long recognition that discrete features (n-gram features) and neural network based features have complementary strengths for language models (LMs). Improved performance can be obtained by model interpolation, which is, however, a sub-optimal two-step integration of discrete and neural features. The trans-dimensional random field (TRF) framework has the potential advantage of being able to flexibly integrate a richer set of features. However, either discrete or neural features are used alone in previous TRF LMs. This paper develops a mixed-feature TRF LM and demonstrates its advantage in integrating discrete and neural features. Various LMs are trained over PTB and Google one-billion-word datasets, and evaluated in N-best list rescoring experiments for speech recognition. Among all single LMs (i.e. without model interpolation), the mixed-feature TRF LMs perform the best, improving over both discrete TRF LMs and neural TRF LMs alone, and also being significantly better than LSTM LMs. Compared to interpolating two separately trained models with discrete and neural features respectively, the performance of mixed-feature TRF LMs matches the best interpolated model, and with simplified one-step training process and reduced training time. Silin Gao, Zhijian Ou, Huifang Xu |
ICASSP | 2 |
| 2020 | Upgrading CRFS to JRFS and its Benefits to Sequence Modeling and LabelingabstractTwo important sequence tasks are sequence modeling and labeling. Sequence modeling involves determining the probabilities of sequences, e.g. language modeling. It is still difficult to improve language modeling with additional relevant tags, e.g. part-of-speech (POS) tags. For sequence labeling, it is worthwhile to explore task-dependent semi-supervised learning to leverage a mix of labeled and unlabeled data, besides pre-training. In this paper, we propose to upgrade conditional random fields (CRFs) and obtain a joint generative model of observation and label sequences, called joint random fields (JRFs). Specifically, we propose to use the potential function in the original CRF as the potential function that defines the joint distribution. This development from CRFs to JRFs benefits both modeling and labeling of sequence data, as shown in our experiments. For example, the JRF model (using POS tags) outperforms traditional language models and avoids the need to produce hypothesized labels by a standalone POS tagger. For sequence labeling, task-dependent semi-supervised learning by JRFs consistently outperform the CRF baseline and self-training, on POS tagging, chunking and NER. Yunfu Song, Zhijian Ou, Songfan Yang |
ICASSP | 2 |
| 2020 | CAT: A CTC-CRF Based ASR Toolkit Bridging the Hybrid and the End-to-End Approaches Towards Data Efficiency and Low LatencyabstractIn this paper, we present a new open source toolkit for speech recognition, named CAT (CTC-CRF based ASR Toolkit).CAT inherits the data-efficiency of the hybrid approach and the simplicity of the E2E approach, providing a full-fledged implementation of CTC-CRFs and complete training and testing scripts for a number of English and Chinese benchmarks.Experiments show CAT obtains state-of-the-art results, which are comparable to the fine-tuned hybrid models in Kaldi but with a much simpler training pipeline.Compared to existing nonmodularized E2E models, CAT performs better on limited-scale datasets, demonstrating its data efficiency.Furthermore, we propose a new method called contextualized soft forgetting, which enables CAT to do streaming ASR without accuracy degradation.We hope CAT, especially the CTC-CRF based framework and software, will be of broad interest to the community, and can be further explored and improved. Keyu An, Hongyu Xiang, Zhijian Ou |
INTERSPEECH | 3 |
| 2020 | Improved Learning of Word Embeddings with Word Definitions and Semantic Injection
Yichi Zhang 0001, Yinpei Dai, Zhijian Ou, Huixin Wang, Junlan Feng |
INTERSPEECH | 3 |
| 2020 | Joint Stochastic Approximation and Its Application to Learning Discrete Latent Variable ModelsabstractAlthough with progress in introducing auxiliary amortized inference models, learning discrete latent variable models is still challenging. In this paper, we show that the annoying difficulty of obtaining reliable stochastic gradients for the inference model and the drawback of indirectly optimizing the target log-likelihood can be gracefully addressed in a new method based on stochastic approximation (SA) theory of the Robbins-Monro type. Specifically, we propose to directly maximize the target log-likelihood and simultaneously minimize the inclusive divergence between the posterior and the inference model. The resulting learning algorithm is called joint SA (JSA). To the best of our knowledge, JSA represents the first method that couples an SA version of the EM (expectation-maximization) algorithm (SAEM) with an adaptive MCMC procedure. Experiments on several benchmark generative modeling and structured prediction tasks show that JSA consistently outperforms recent competitive algorithms, with faster convergence, better final likelihoods, and lower variance of gradient estimates. Zhijian Ou, Yunfu Song |
UAI | 1 |
| 2020 | Semi-Supervised Seq2seq Joint-Stochastic-Approximation Autoencoders With Applications to Semantic ParsingabstractDeveloping Semi-Supervised Seq2Seq (S4) learning for sequence transduction tasks in natural language processing (NLP), e.g. semantic parsing, is challenging, since both the input and the output sequences are discrete. This discrete nature makes trouble for methods which need gradients either from the input space or from the output space. Recently, a new learning method called joint stochastic approximation is developed for unsupervised learning of fixed-dimensional autoencoders and theoretically avoids gradient propagation through discrete latent variables, which is suffered by Variational Auto-Encoders (VAEs). In this letter, we propose seq2seq Joint-stochastic-approximation AutoEncoders (JAEs) and apply them to S4learning for NLP sequence transduction tasks. Further, we propose bi-directional JAEs (called bi-JAEs) to leverage not only unpaired input sequences (which is most commonly studied) but also unpaired output sequences. Experiments on two benchmarking datasets for semantic parsing show that JAEs consistently outperform VAEs in S4 learning and bi-JAEs yield further improvements. Yunfu Song, Zhijian Ou |
IEEE Signal Process. Lett. | 2 |
| 2019 | Neural CRF Transducers for Sequence LabelingabstractConditional random fields (CRFs) have been shown to be one of the most successful approaches to sequence labeling. Various linear-chain neural CRFs (NCRFs) are developed to implement the non-linear node potentials in CRFs, but still keeping the linear-chain hidden structure. In this paper, we propose NCRF transducers, which consists of two RNNs, one extracting features from observations and the other capturing (theoretically infinite) long-range dependencies between labels. Different sequence labeling methods are evaluated over POS tagging, chunking and NER (English, Dutch). Experiment results show that NCRF transducers achieve consistent improvements over linear-chain NCRFs and RNN transducers across all the four tasks, and can improve state-of-the-art results. Zhijian Ou, Junlan Feng |
ICASSP | 2 |
| 2019 | CRF-based Single-stage Acoustic Modeling with CTC TopologyabstractIn this paper, we develop conditional random field (CRF) based single-stage (SS) acoustic modeling with connectionist temporal classification (CTC) inspired state topology, which is called CTC-CRF for short. CTC-CRF is conceptually simple, which basically implements a CRF layer on top of features generated by the bottom neural network with the special state topology. Like SS-LF-MMI (lattice-free maximum-mutual-information), CTC-CRFs can be trained from scratch (flat-start), eliminating GMM-HMM pre-training and tree-building. Evaluation experiments are conducted on the WSJ, Switchboard and Librispeech datasets. In a head-to-head comparison, the CTC-CRF model using simple Bidirectional LSTMs consistently outperforms the strong SS-LF-MMI, across all the three benchmarking datasets and in both cases of mono-phones and mono-chars. Additionally, CTC-CRFs avoid some ad-hoc operation in SS-LF-MMI. Hongyu Xiang, Zhijian Ou |
ICASSP | 2 |
| 2018 | Tracking of Enriched Dialog States for Flexible Conversational Information AccessabstractDialog state tracking (DST) is a crucial component in a task-oriented dialog system for conversational information access. A common practice in current dialog systems is to define the dialog state by a set of slot-value pairs. Such representation of dialog states and the slot-filling based DST have been widely employed, but suffer from three drawbacks. (1) The dialog state can contain only a single value for a slot, and (2) can contain only users' affirmative preference over the values for a slot. (3) Current task-based dialog systems mainly focus on the searching task, while the enquiring task is also very common in practice. The above observations motivate us to enrich current representation of dialog states and collect a brand new dialog dataset about movies, based upon which we build a new DST, called enriched DST (EDST), for flexible movie information access. The EDST supports the searching task, the enquiring task and their mixed task. We show that our new EDST method not only achieves good results on Iqiyi dataset, but also outperforms other state-of-the-art DST methods on the traditional dialog datasets, WOZ2.0 and DSTC2. Yinpei Dai, Zhijian Ou, Dawei Ren, Pengfei Yu 0001 |
ICASSP | 2 |
| 2018 | Learning Neural Trans-Dimensional Random Field Language Models with Noise-Contrastive EstimationabstractTrans-dimensional random field language models (TRF LMs) where sentences are modeled as a collection of random fields, have shown close performance with LSTM LMs in speech recognition and are computationally more efficient in inference. However, the training efficiency of neural TRF LMs is not satisfactory, which limits the scalability of TRF LMs on large training corpus. In this paper, several techniques on both model formulation and parameter estimation are proposed to improve the training efficiency and the performance of neural TRF LMs. First, TRFs are reformulated in the form of exponential tilting of a reference distribution. Second, noise-contrastive estimation (NCE) is introduced to jointly estimate the model parameters and normalization constants. Third, we extend the neural TRF LMs by marrying the deep convolutional neural network (CNN) and the bidirectional LSTM into the potential function to extract the deep hierarchical features and bidirectionally sequential features. Utilizing all the above techniques enables the successful and efficient training of neural TRF LMs on a 40x larger training set with only 1/3 training time and further reduces the WER with relative reduction of 4.7% on top of a strong LSTM LM baseline. Bin Wang 0042, Zhijian Ou |
ICASSP | 2 |
| 2018 | Improved Training Of Neural Trans-Dimensional Random field Language Models with Dynamic Noise-Contrastive EstimationabstractA new whole-sentence language model - neural trans-dimensional random field language model (neural TRF LM), where sentences are modeled as a collection of random fields, and the potential function is defined by a neural network, has been introduced and successfully trained by noise-contrastive estimation (NCE). In this paper, we extend NCE and propose dynamic noise-contrastive estimation (DNCE) to solve the two problems observed in NCE training. First, a dynamic noise distribution is introduced and trained simultaneously to converge to the data distribution. This helps to significantly cut down the noise sample number and reduce the training cost over NCE. Second, DNCE discriminates between sentences generated from the noise distribution and sentences generated from the interpolation of the data distribution and the noise distribution. This alleviates the overfitting problem caused by the sparseness of the training set. With DNCE, we can successfully and efficiently train neural TRF LMs on large corpus (about 0.8 billion words) with large vocabulary (about 568 K words). Neural TRF LMs perform as good as LSTM LMs with less parameters and being 5×~114× faster in rescoring sentences. Interpolating neural TRF LMs with LSTM LMs and n-gram LMs can further reduce the error rates. Bin Wang 0042, Zhijian Ou |
SLT | 2 |
| 2018 | Learning Trans-Dimensional Random Fields with Applications to Language ModelingabstractTo describe trans-dimensional observations in sample spaces of different dimensions, we propose a probabilistic model, called the trans-dimensional random field (TRF) by explicitly mixing a collection of random fields. In the framework of stochastic approximation (SA), we develop an effective training algorithm, called augmented SA, which jointly estimates the model parameters and normalizing constants while using trans-dimensional mixture sampling to generate observations of different dimensions. Furthermore, we introduce several statistical and computational techniques to improve the convergence of the training algorithm and reduce computational cost, which together enable us to successfully train TRF models on large datasets. The new model and training algorithm are thoroughly evaluated in a number of experiments. The word morphology experiment provides a benchmark test to study the convergence of the training algorithm and to compare with other algorithms, because log-likelihoods and gradients can be exactly calculated in this experiment. For language modeling, our experiments demonstrate the superiority of the TRF approach in being computationally more efficient in computing data probabilities by avoiding local normalization and being able to flexibly integrate a richer set of features, when compared with n-gram models and neural network models. Bin Wang 0042, Zhijian Ou, Zhiqiang Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Language modeling with neural trans-dimensional random fieldsabstractTrans-dimensional random field language models (TRF LMs) have recently been introduced, where sentences are modeled as a collection of random fields. The TRF approach has been shown to have the advantages of being computationally more efficient in inference than LSTM LMs with close performance and being able to flexibly integrate rich features. In this paper we propose neural TRFs, beyond of the previous discrete TRFs that only use linear potentials with discrete features. The idea is to use nonlinear potentials with continuous features, implemented by neural networks (NNs), in the TRF framework. Neural TRFs combine the advantages of both NNs and TRFs. The benefits of word embedding, nonlinear feature learning and larger context modeling are inherited from the use of NNs. At the same time, the strength of efficient inference by avoiding expensive softmax is preserved. A number of technical contributions, including employing deep convolutional neural networks (CNNs) to define the potentials and incorporating the joint stochastic approximation (JSA) strategy in the training algorithm, are developed in this work, which enable us to successfully train neural TRF LMs. Various LMs are evaluated in terms of speech recognition WERs by rescoring the 1000-best lists of WSJ'92 test data. The results show that neural TRF LMs not only improve over discrete TRF LMs, but also perform slightly better than LSTM LMs with only one fifth of parameters and 16x faster inference efficiency. Bin Wang 0042, Zhijian Ou |
ASRU | 2 |
| 2017 | Joint Bayesian Gaussian Discriminant Analysis for speaker verificationabstractState-of-the-art i-vector based speaker verification relies on variants of Probabilistic Linear Discriminant Analysis (PLDA) for discriminant analysis. We are mainly motivated by the recent work of the joint Bayesian (JB) method, which is originally proposed for discriminant analysis in face verification. We apply JB to speaker verification and make three contributions beyond the original JB. 1) In contrast to the EM iterations with approximated statistics in the original JB, the EM iterations with exact statistics are employed and give better performance. 2) We propose to do simultaneous diagonalization (SD) of the within-class and between-class covariance matrices to achieve efficient testing, which has broader application scope than the SVD-based efficient testing method in the original JB. 3) We scrutinize similarities and differences between various Gaussian PLDAs and JB, complementing the previous analysis of comparing JB only with Prince-Elder PLDA. Extensive experiments are conducted on NIST SRE10 core condition 5, empirically validating the superiority of JB with faster convergence rate and 9 - 13% EER reduction compared with state-of-the-art PLDA. Yiyan Wang, Zhijian Ou |
ICASSP | 3 |
| 2016 | Scalable Discovery of Audio Fingerprint Motifs in Broadcast Streams With Determinantal Point Process Based Motif ClusteringabstractIn this paper, we study the scalable discovery of audio repetitive patterns/motifs in long broadcast streams, where two segments are said to be repetitive if their audio fingerprints are close to each other. In this task, as we are confined to handle limited variability, we can adapt an audio hashing technique, originally proposed for searching a given music clip in music tracks, to successfully devise a linear complexity similarity matching method with a new step of repeated interval formation. This is the first contribution of this paper. As the similarity matching is super fast and thus coarse, there are false alarms in the large number of pairwise matches generated, which constitute a major source of noise. We propose applying subset selection to the original set of pairwise matches based on determinantal point processes (DPPs), as a filtering step, to reduce the noise. The selected subset of pairwise matches is then subjected to motif clustering. We successfully apply DPP-based subset selection to improve motif clustering, which has a nice property that favors both quality and diversity. This is the second contribution of this paper. The proposed method is thoroughly evaluated on a 9-hour real-world audio stream and is compared with several reference methods. The bootstrap technique is used for the significance test. It is shown that the similarity matching is computationally very efficient (above 100 times faster than real time), and the filtering step with DPPs can significantly improve the precision of motif discovery, without sacrificing the recall performance. Zhijian Ou |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Trans-dimensional Random Fields for Language ModelingabstractBin Wang, Zhijian Ou, Zhiqiang Tan. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Bin Wang 0042, Zhijian Ou, Zhiqiang Tan |
ACL (1) | 2 |
| 2014 | Improvement of Probabilistic Acoustic Tube model for speech decompositionabstractCurrent model-based speech analysis tends to be incomplete - only a part of parameters of interest (e.g. only the pitch or vocal tract) are modeled, while the rest that might as well be important are disregarded. The drawback is that without joint modeling of parameters that are correlated, the analysis on speech parameters may be inaccurate or even incorrect. Under this motivation, we have proposed such a model called PAT (Probabilistic Acoustic Tube), where pitch, vocal tract and energy are jointly modeled. This paper proposes an improved version of PAT model, named PAT2, where both signal and probabilistic modeling are tremendously renovated. Compared to related works, PAT2 is much more comprehensive, which incorporates mixed excitation, glottal wave and phase modeling. Experimental results show its ability in decomposing speech into desirable parameters and its potential for speech synthesis. Yang Zhang 0001, Zhijian Ou, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2012 | Combining eigenvoice speaker modeling and VTS-based environment compensation for robust speech recognitionabstractEigenvoice and vector Taylor series (VTS) are good models for speaker differences and environmental variations separately. However, speaker and environmental variation always coexist in real-world speech. In this paper, we propose to combine eigenvoice and VTS. Specifically, we introduce eigenvoice speaker modeling for the clean speech into VTS's nonlinear mismatch function. In contrast, the standard VTS uses speaker-independent modeling to represent the clean speech, regardless of speaker differences. The eigenvoice coefficients and the noise model parameters are jointly estimated in the new approach. Experimental results on the Aurora2 task show the improved performances of combining eigenvoice and VTS and demonstrate its ability for speaker and noise factorization. Zhijian Ou, Kan Deng |
ICASSP | 1 |
| 2012 | CRF-based confidence measures of recognized candidates for lattice-based audio indexingabstractThe use of forward-backward (FB) computation based posterior probabilities as confidence measures (CMs) for all recognized candidates in a lattice seems to be common across various lattice-based audio indexing systems. However, a major limitation with this approach is that its performance for CMs cannot be improved easily, since it relies almost entirely on a single information source - the acoustic and language-model probabilities. In this paper, we propose to formulate computing CMs in the lattice case as a multi-class sequential labeling problem, using conditional random fields (CRFs) as the underlying model. In this approach, various relevant features including the FB posterior probabilities could be combined together. Note that CRFs are well suited to label sequence data and some features are defined over a word sequence. This paper presents how we resolve these two issues in the lattice case, beyond others' previous work in CRF-based CMs for the 1-best case. Once properly implemented, the proposed approach achieves significant performance improvements for both CMs in the lattice case and lattice-based audio indexing. Zhijian Ou, Huaqing Luo |
ICASSP | 1 |
| 2011 | Combining HMM-based melody extraction and NMF-based soft masking for separating voice and accompaniment from monaural audioabstractModern monaural voice and accompaniment separation systems usually consist of two main modules: melody extraction and time frequency masking. A main distinction between different separation systems lies in what approaches are used for the two modules. Popular techniques for melody extraction include hidden Markov models (HMMs) and non-negative matrix factorization (NMF), and masking includes hard and soft masking. This paper investigates the flaw of NMF-based melody extraction, and proposes the combination of HMM-based melody extraction (equipped with a newly-defined feature) and NMF-based soft masking. Evaluations on two publicly available databases show that the proposed system reaches state-of the-art performance and outperforms several other combinations. Zhijian Ou |
ICASSP | 2 |
| 2010 | Variational nonparametric Bayesian Hidden Markov ModelabstractThe Hidden Markov Model (HMM) has been widely used in many applications such as speech recognition. A common challenge for applying the classical HMM is to determine the structure of the hidden state space. Based on the Dirichlet Process, a nonparametric Bayesian Hidden Markov Model is proposed, which allows an infinite number of hidden states and uses an infinite number of Gaussian components to support continuous observations. An efficient variational inference method is also proposed and applied on the model. Our experiments demonstrate that the variational Bayesian inference on the new model can discover the HMM hidden structure for both synthetic data and real-world applications. Nan Ding 0002, Zhijian Ou |
ICASSP | 2 |
| 2010 | Spoken English assessment system for non-native speakers using acoustic and prosodic featuresabstractThe absence of real-time and targeted feedback is often critical in spoken foreign language learning. Computer-assisted language assessment systems are playing an ever more important role in this domain. This work considers the idiosyncratic pronunciation patterns of Chinese English speakers and uses both acoustic and prosody features to capture pronunciation, word stress, and rhythm information. The proposed system uses a. automatic speech recognition and alignment for pronunciation assessment, b. a set of special features with appropriate normalization for word stress detection, and c. a prosody phrase prediction model for rhythm assessment; and is shown to give immediate and accurate analyses to speakers to improve learning efficiency. Qin Shi 0001, Shilei Zhang, Stephen M. Chu, Ji Xiao, Zhijian Ou |
INTERSPEECH | 6 |
| 2008 | Caption-aided speech detection in videosabstractThis paper presents a novel audio-visual fusion method for speech detection, which is an important front-end for content-based video processing. This approach aims to extract homogeneous speech segments from the accompanying audio stream in real-world movie/TV videos with the help of video captions. Note that captions are mainly created to help viewers to follow the dialog, rather than to accurately locate the speech regions. We propose a caption-aided speech detection approach, which makes use of both caption information and audio information. The inaccurate positions of the captions are refined through using audio features (pitch and MFCCs) and BIC-based acoustic change detection. Comparison experiments against several other traditional speech detection approaches are conducted, showing that the proposed approach improves the speech detection performance greatly. Zhijian Ou, Wei Hu 0002, Tao Wang 0003, Yimin Zhang 0002 |
ICASSP | 2 |
| 2007 | Latent Correlation Analysis of HMM Parameters for Speech RecognitionabstractCorrelation between HMM parameters has been utilized for various rapid speaker adaptation, e.g. eigenvoice adaptation. The covariance matrix of the supervector which is a concatenation of all the Gaussian means in HMM, is clearly a good measure of such parameter correlation. In this paper, we propose to treat the supervector as a latent variable under HMM, and perform estimation of the hidden supervector's covariance matrix directly from the acoustic frames using EM algorithm. In contrast to traditional methods which depend on using well-trained/adapted supervector samples, the proposed method is more theoretically sound and capable of dealing well with speaker-specific data sparseness. Moreover, the idea of conducting utterance-level correlation analysis, estimating utterance eigenvoices, and performing (unsupervised) utterance adaptation is explored. Experiments on the OGI Numbers database show that the proposed approach achieves better adaptation performance than the traditional methods, and the utterance-level correlation analysis is found to be useful. Zhijian Ou |
ICASSP (4) | 1 |
| 2007 | Switching Auxiliary Chains for Speech RecognitionabstractThis letter investigates the problem of incorporating auxiliary information, e.g., pitch, zero crossing rate (ZCR), and rate-of-speech (ROS), for speech recognition using dynamic Bayesian networks. In this letter, we propose switching auxiliary chains for exploiting different auxiliary information tailored to different phonetic states. The switching function can be specified by a priori knowledge or, more flexibly, be learned from data with information-theoretic dependency selection. Experiments on the OGI Numbers database show that the new model achieves 7% word-error-rate relative reduction by jointly exploiting pitch, ZCR, and ROS, while keeping almost the same parameter size as the standard HMM. Hui Lin 0001, Zhijian Ou |
IEEE Signal Process. Lett. | 2 |
| 2007 | Closely Coupled Array Processing and Model-Based Compensation for Microphone Array Speech RecognitionabstractIn conventional microphone array speech recognition, the array processor and the speech recognizer are loosely coupled. The only connection between the two modules is the enhanced target signal output from the array processor, which then gets treated as a single input to the recognizer. In this approach, useful environmental information, which can be provided by the array processor and also needs to be exploited by the recognizer, is ignored. Inherently, the array processor can generate multiple outputs of spatially filtered signals, as a multi-input-multi-output (MIMO) module. In this paper, a closely coupled approach is proposed, in which a recognizer with model-based noise compensation exploits the reference noise outputs from a MIMO array processor. Specifically, a multichannel model-based noise compensation is presented, including the compensation procedure using the vector Taylor series (VTS) expansion and parameter estimation using the expectation-maximization (EM) algorithm. It is also shown how to construct MIMO array processors from conventional beamformers. A number of practical implementations of the conventional loosely coupled approach and the proposed closely coupled approach were tested on a publicly available database, the Multichannel Overlapping Number Corpus (MONC). Experimental results showed that the proposed closely coupled approach significantly improved the speech recognition performance in the overlapping speech situations Xianyu Zhao, Zhijian Ou |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Partial-tied-mixture Auxiliary Chain Models for Speech Recognition Based on Dynamic Bayesian NetworksabstractIt is observed that the cepstral-based features used for speech recognition are sensitive to some auxiliary information (e.g. pitch). Encoding the auxiliary information in discrete auxiliary variables based on dynamic Bayesian networks (DBNs) typically results in an increased number of parameters. There are tradeoffs to be studied between parameter reduction and dependency modeling. In this paper, we propose a method using state-specific partial tying with information- theoretic dependency selection. This method is essentially to relax the conditional independence assumptions imposed by the full-tied-mixture model, by adding strong dependencies (i.e. those with large mutual information computed from training data). Experiments were carried out on the OGI Numbers database, considering pitch as the auxiliary information. The results show that the partial-tied-mixture auxiliary chain models can efficiently improve recognition performances with an economical way of increasing parameters. Hui Lin 0001, Zhijian Ou |
SMC | 2 |
| 2006 | Generalized Time-Series Active Search With Kullback-Leibler Distance for Audio FingerprintingabstractIn this letter, a new audio fingerprinting approach is presented. We investigate to improve robustness by more precise statistical fingerprint modeling with common component Gaussian mixture models (CCGMMs) and Kullback–Leibler (KL) distance, which is more suitable to measure the dissimilarity between two probabilistic models. To address the resulting complexity, generalized time-series active search is proposed, which supports a wide variety of distance measures between two CCGMMs, including$L_1$,$L_2$, KL, etc. Experiments show that the new approach with KL distance increases robustness to distortions (including low-quality MP3 compression, small room echo, and play-and-record) while achieving efficient search. Hui Lin 0001, Zhijian Ou |
IEEE Signal Process. Lett. | 2 |
| 2005 | Closely Coupled Array Processing and Model-Based Compensation for Microphone Array Speech RecognitionabstractIn this paper, a new microphone array speech recognition system in which the array processor and the speech recognizer are closely coupled is studied. The system includes a generalized sidelobe canceller (GSC) beamformer followed by a recognizer with vector Taylor series (VTS) compensation. The GSC beamformer provides two outputs, allowing more information to be used in the recognizer. One is the enhanced target speech output, the other is the reference noise output. VTS is used to compensate the effect of the residual noise in the GSC speech output, utilizing the GSC reference noise output. The compensation is done in a minimum mean square error (MMSE) sense. Moreover, an iteration procedure using an expectation-maximization (EM) algorithm is developed to refine the compensation parameters. Experimental results on the MONC database showed that the new system significantly improved the speech recognition performance in overlapping speech situations. Xianyu Zhao, Zhijian Ou, Minhua Chen, Zuoying Wang |
ICASSP (1) | 2 |
| 2005 | Discriminative speaker adaptation with eigenvoicesabstractEigenvoice is an effective speaker adaptation approach and capable of balancing the performance and the requirement for a large amount of adaptation data. However, the conventional Maximum Likelihood Eigen-Decomposition (MLED) method in eigenvoice adaptation is based on Maximum Likelihood (ML) criterion and suffers from the unrealistic assumption made by HMM on speech process, so alternative schemes may be more effective to improve the performance. In this paper, we propose a new discriminative adaptation algorithm called Maximum Mutual Information Eigen-Decomposition (MMIED) in which the mutual information between the training word sequences and the observation sequences is maximized. By the use of word lattice, the competing word hypotheses are taken into account to make the estimation more discriminative. MLED, MMIED and Maximum a Posteriori EigenDecomposition (MAPED) which is based on Maximum a Posteriori (MAP) criterion were all experimented to give a comprehensive comparison. Results showed that MMIED outperformed both MLED and MAPED. Zhijian Ou, Zuoying Wang |
INTERSPEECH | 2 |
| 2004 | Discriminative combination of multiple linear predictions for speech recognitionabstractIn this paper, new analyses are provided for the problems of applying linear prediction (LP) HMMs in speech recognition. It is shown that, apart from simply aggregating all predictors in one LP, which produces inconsistent results, ‘combination’ provides another useful way to implement complex dependencies. A method by discriminative combination of multiple LPs (DCoLP) is proposed, with a component-LP selection heuristic. The resulting DCoLP model was tested on a speaker-independent, large vocabulary continuous speech recognition task, and showed improved performance over the standard HMM with comparable computation cost. Zhijian Ou, Zuoying Wang |
INTERSPEECH | 1 |
| 2002 | A new combined model of statics-dynamics of speechabstractLinear prediction (LP) HMM does not make the independent and identical distribution (IID) assumption in traditional HMM; however it often produces unsatisfactory results. In this paper, a new combined model of statics-dynamics of speech is proposed, based on a new analysis of both HMM's modeling strengths and weaknesses. The new model works with LPHMM as the dynamic part and traditional IID-based HMM as the static part; in addition, easy implementation and low cost are preserved. A new effective re-estimation solution is suggested for parameter tying to achieve better discrimination. Our experiments on speaker-independent continuous speech recognition demonstrated that the combined model achieved 7.5% error rate reduction from traditional HMM. Zhijian Ou, Zuoying Wang |
ICASSP | 1 |
| 2002 | A combined model of statics-dynamics of speech optimized using maximum mutual information
Zhijian Ou, Zuoying Wang |
INTERSPEECH | 1 |
| 2001 | A new DP-like speaker clustering algorithm
Zhijian Ou, Zuoying Wang |
INTERSPEECH | 1 |