EDBT 2026 Demo / reviewers in the wild / expert
Kyu Jeong Han
dblp:44/2176
· DBLP profile ↗
42ranked-venue papers
21as first author
10since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 19 first-author · 8 since 2021Artificial intelligence and machine learning · 27 · 14 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Wav2Seq: Pre-Training Speech-to-Text Encoder-Decoder Models Using Pseudo LanguagesabstractWe introduce Wav2Seq, the first self-supervised approach to pre-train both parts of encoder-decoder models for speech data. We induce a pseudo language as a compact discrete representation, and formulate a self-supervised pseudo speech recognition task — transcribing audio inputs into pseudo subword sequences. This process stands on its own, or can be applied as low-cost second-stage pre-training. We experiment with automatic speech recognition (ASR), spoken named entity recognition, and speech-to-text translation. We set new state-of-the-art results for end-to-end spoken named entity recognition, and show consistent improvements on 8 language pairs for speech-to-text translation, even when competing methods use additional text data for training. On ASR, our approach enables encoder-decoder methods to benefit from pre-training for all parts of the network, and shows comparable performance to highly optimized recent methods. Felix Wu, Kwangyoun Kim, Shinji Watanabe 0001, Kyu Jeong Han, Ryan McDonald, Kilian Q. Weinberger, Yoav Artzi |
ICASSP | 4 |
| 2022 | SRU++: Pioneering Fast Recurrence with Attention for Speech RecognitionabstractThe Transformer architecture has been well adopted as a dominant architecture in most sequence transduction tasks including automatic speech recognition (ASR), since its attention mechanism excels in capturing long-range dependencies. While models built solely upon attention can be better parallelized than regular RNN, a novel network architecture, SRU++, was recently proposed. By combining the fast recurrence and attention mechanism, SRU++ exhibits strong capability in sequence modeling and achieves near-state-of-the-art results in various language modeling and machine translation tasks with improved compute efficiency. In this work, we present the advantages of applying SRU++ in ASR tasks by comparing with Conformer across multiple ASR benchmarks and study how the benefits can be generalized to long-form speech inputs. On the popular LibriSpeech benchmark, our SRU++ model achieves 2.0% / 4.7% WER on test-clean / test-other, showing competitive performances compared with the state-of-the-art Conformer encoder under the same set-up. Specifically, SRU++ can surpass Conformer on long-form speech input with a large margin, based on our analysis. Tao Lei 0001, Kwangyoun Kim, Kyu Jeong Han, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2022 | SLUE: New Benchmark Tasks For Spoken Language Understanding Evaluation on Natural SpeechabstractProgress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest has been growing in higher-level spoken language understanding tasks, including using end-to-end models, but there are fewer annotated datasets for such tasks. At the same time, recent work shows the possibility of pre-training generic representations and then fine-tuning for several tasks using relatively little labeled data. We propose to create a suite of benchmark tasks for Spoken Language Understanding Evaluation (SLUE) consisting of limited-size labeled training sets and corresponding evaluation sets. This resource would allow the research community to track progress, evaluate pre-trained representations for higher-level tasks, and study open questions such as the utility of pipeline versus end-to-end approaches. We present the first phase of the SLUE benchmark suite, consisting of named entity recognition, sentiment analysis, and ASR on the corresponding datasets. We focus on naturally produced (not read or synthesized) speech, and freely available datasets. We pro-vide new transcriptions and annotations on subsets of the VoxCeleb and VoxPopuli datasets, evaluation metrics and results for baseline models, and an open-source toolkit to reproduce the baselines and evaluate new models. Suwon Shon, Ankita Pasad, Felix Wu, Pablo Brusco, Yoav Artzi, Karen Livescu, Kyu Jeong Han |
ICASSP | 7 |
| 2022 | Performance-Efficiency Trade-Offs in Unsupervised Pre-Training for Speech RecognitionabstractThis paper is a study of performance-efficiency trade-offs in pre-trained models for automatic speech recognition (ASR). We focus on wav2vec 2.0, and formalize several architecture designs that influence both the model performance and its efficiency. Putting together all our observations, we introduce SEW-D (Squeezed and Efficient Wav2vec with Disentangled Attention), a pre-trained model architecture with significant improvements along both performance and efficiency dimensions across a variety of training setups. For example, under the 100h-960h semi-supervised setup on LibriSpeech, SEW-D achieves a 1.9x inference speedup compared to wav2vec 2.0, with a 13.5% relative reduction in word error rate. With a similar inference time, SEW reduces word error rate by 25–50% across different model sizes. Felix Wu, Kwangyoun Kim, Kyu Jeong Han, Kilian Q. Weinberger, Yoav Artzi |
ICASSP | 4 |
| 2022 | On the Use of External Data for Spoken Named Entity RecognitionabstractAnkita Pasad, Felix Wu, Suwon Shon, Karen Livescu, Kyu Han. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ankita Pasad, Felix Wu, Suwon Shon, Karen Livescu, Kyu Jeong Han |
NAACL-HLT | 5 |
| 2022 | E-Branchformer: Branchformer with Enhanced Merging for Speech RecognitionabstractConformer, combining convolution and self-attention sequentially to capture both local and global information, has shown remarkable performance and is currently regarded as the state-of-the-art for automatic speech recognition (ASR). Several other studies have explored integrating convolution and self-attention but they have not managed to match Conformer's performance. The recently introduced Branchformer achieves comparable performance to Conformer by using dedicated branches of convolution and self-attention and merging local and global context from each branch. In this paper, we propose E-Branchformer, which enhances Branchformer by applying an effective merging method and stacking additional point-wise modules. E-Branchformer sets new state-of-the-art word error rates (WERs) 1.81% and 3.65% on LibriSpeech test-clean and test-other sets without using any external training data. Kwangyoun Kim, Felix Wu, Yifan Peng 0003, Prashant Sridhar, Kyu Jeong Han, Shinji Watanabe 0001 |
SLT | 6 |
| 2022 | A review of speaker diarization: Recent advances with deep learning
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu Jeong Han, Shinji Watanabe 0001, Shri Narayanan |
Comput. Speech Lang. | 4 |
| 2021 | Multistream CNN for Robust Acoustic ModelingabstractThis paper proposes multistream CNN, a novel neural network architecture for robust acoustic modeling in speech recognition tasks. The proposed architecture processes input speech with diverse temporal resolutions by applying different dilation rates to convolutional neural networks across multiple streams to achieve the robustness. The dilation rates are selected from the multiples of a sub-sampling rate of 3 frames. Each stream stacks TDNN-F layers (a variant of 1D CNN), and output embedding vectors from the streams are concatenated then projected to the final layer. We validate the effectiveness of the proposed multistream CNN architecture by showing consistent improvements against Kaldi’s best TDNN-F model across various data sets. Multistream CNN improves the WER of the test-other set in the LibriSpeech corpus by 12% (relative). On custom data from ASAPP’s production ASR system for a contact center, it records a relative WER improvement of 11% for customer channel audio to prove its robustness to data in the wild. In terms of real-time factor, multistream CNN outperforms the baseline TDNN-F by 15%, which also suggests its practicality on production systems. When combined with self-attentive SRU LM rescoring, multistream CNN contributes for ASAPP to achieve the best WER of 1.75% on test-clean in LibriSpeech. Kyu Jeong Han, Venkata Krishna Naveen Tadala, Daniel Povey |
ICASSP | 1 |
| 2021 | Multi-Mode Transformer Transducer with Stochastic Future ContextabstractAutomatic speech recognition (ASR) models make fewer errors when more surrounding speech information is presented as context.Unfortunately, acquiring a larger future context leads to higher latency.There exists an inevitable trade-off between speed and accuracy.Naïvely, to fit different latency requirements, people have to store multiple models and pick the best one under the constraints.Instead, a more desirable approach is to have a single model that can dynamically adjust its latency based on different constraints, which we refer to as Multimode ASR.A Multi-mode ASR model can fulfill various latency requirements during inference -when a larger latency becomes acceptable, the model can process longer future context to achieve higher accuracy and when a latency budget is not flexible, the model can be less dependent on future context but still achieve reliable accuracy.In pursuit of Multi-mode ASR, we propose Stochastic Future Context, a simple training procedure that samples one streaming configuration in each iteration.Through extensive experiments on AISHELL-1 and Lib-riSpeech datasets, we show that a Multi-mode ASR model rivals, if not surpasses, a set of competitive streaming baselines trained with different latency budgets. Kwangyoun Kim, Felix Wu, Prashant Sridhar, Kyu Jeong Han, Shinji Watanabe 0001 |
Interspeech | 4 |
| 2021 | Leveraging Pre-Trained Language Model for Speech Sentiment AnalysisabstractIn this paper, we explore the use of pre-trained language models to learn sentiment information of written texts for speech sentiment analysis.First, we investigate how useful a pre-trained language model would be in a 2-step pipeline approach employing Automatic Speech Recognition (ASR) and transcripts-based sentiment analysis separately.Second, we propose a pseudo label-based semi-supervised training strategy using a language model on an end-to-end speech sentiment approach to take advantage of a large, but unlabeled speech dataset for training.Although spoken and written texts have different linguistic characteristics, they can complement each other in understanding sentiment.Therefore, the proposed system can not only model acoustic characteristics to bear sentimentspecific information in speech signals, but learn latent information to carry sentiments in the text representation.In these experiments, we demonstrate the proposed approaches improve F1 scores consistently compared to systems without a language model.Moreover, we also show that the proposed framework can reduce 65% of human supervision by leveraging a large amount of data without human sentiment annotation and boost performance in a low-resource condition where the human sentiment annotation is not available enough. Suwon Shon, Pablo Brusco, Kyu Jeong Han, Shinji Watanabe 0001 |
Interspeech | 4 |
| 2020 | ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech RecognitionabstractIn this paper we present state-of-the-art (SOTA) performance on the LibriSpeech corpus with two novel neural network architectures, a multistream CNN for acoustic modeling and a selfattentive simple recurrent unit (SRU) for language modeling.In the hybrid ASR framework, the multistream CNN acoustic model processes an input of speech frames in multiple parallel pipelines where each stream has a unique dilation rate for diversity.Trained with the SpecAugment data augmentation method, it achieves relative word error rate (WER) improvements of 4% on test-clean and 14% on test-other.We further improve the performance via N -best rescoring using a 24-layer self-attentive SRU language model, achieving WERs of 1.75% on test-clean and 4.46% on test-other. Joshua Shapiro, Jeremy Wohlwend, Kyu Jeong Han, Tao Lei 0001 |
INTERSPEECH | 4 |
| 2020 | Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum EigengapabstractIn this study, we propose a new spectral clustering framework that can auto-tune the parameters of the clustering algorithm in the context of speaker diarization. The proposed framework uses normalized maximum eigengap (NME) values to estimate the number of clusters and the parameters for the threshold of the elements of each row in an affinity matrix during spectral clustering, without the use of parameter tuning on the development set. Even through this hands-off approach, we achieve a comparable or better performance across various evaluation sets than the results found using traditional clustering methods that apply careful parameter tuning and development data. A relative improvement of 17% in the speaker error rate on the well-known CALLHOME evaluation set shows the effectiveness of our proposed spectral clustering with auto-tuning. Tae Jin Park, Kyu Jeong Han, Manoj Kumar 0007, Shri Narayanan |
IEEE Signal Process. Lett. | 2 |
| 2019 | State-of-the-Art Speech Recognition Using Multi-Stream Self-Attention with Dilated 1D ConvolutionsabstractSelf-attention has been a huge success for many downstream tasks in NLP, which led to exploration of applying self-attention to speech problems as well. The efficacy of self-attention in speech applications, however, seems not fully blown yet since it is challenging to handle highly correlated speech frames in the context of self-attention. In this paper we propose a new neural network model architecture, namely multi-stream self-attention, to address the issue thus make the self-attention mechanism more effective for speech recognition. The proposed model architecture consists of parallel streams of self-attention encoders, and each stream has layers of 1D convolutions with dilated kernels whose dilation rates are unique given stream, followed by a self-attention layer. The self-attention mechanism in each stream pays attention to only one resolution of input speech frames and the attentive computation can be more efficient. In a later stage, outputs from all the streams are concatenated then linearly projected to the final embedding. By stacking the proposed multi-stream self-attention encoder blocks and rescoring the resultant lattices with neural network language models, we achieve the word error rate of 2.2% on the test-clean dataset of the LibriSpeech corpus, the best number reported thus far on the dataset. Kyu Jeong Han, Ramon Prieto |
ASRU | 1 |
| 2019 | Multi-Stride Self-Attention for Speech Recognition
Kyu Jeong Han, Jing Huang 0019, Yun Tang 0002, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 1 |
| 2019 | Survey Talk: When Attention Meets Speech Applications: Speech & Speaker Recognition Perspective
Kyu Jeong Han, Ramon Prieto |
INTERSPEECH | 1 |
| 2019 | Speaker Diarization with Lexical InformationabstractThis work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with speaker embeddings into a speaker clustering process to improve the overall diarization accuracy. To integrate lexical and acoustic information in a comprehensive way during clustering, we introduce an adjacency matrix integration for spectral clustering. Since words and word boundary information for word-level speaker turn probability estimation are provided by a speech recognition system, our proposed method works without any human intervention for manual transcriptions. We show that the proposed method improves diarization performance on various evaluation datasets compared to the baseline diarization system using acoustic information only in speaker embeddings. Tae Jin Park, Kyu Jeong Han, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2018 | Densely Connected Networks for Conversational Speech Recognition
Kyu Jeong Han, Akshay Chandrashekaran, Jungsuk Kim, Ian Lane |
INTERSPEECH | 1 |
| 2017 | Deep Learning-Based Telephony Speech Recognition in the Wild
Kyu Jeong Han, Seongjun Hahm, Byung-Hak Kim, Jungsuk Kim, Ian Lane |
INTERSPEECH | 1 |
| 2016 | Semi-Supervised Speaker Adaptation for In-Vehicle Speech Recognition with Deep Neural Networks
Wonkyum Lee, Kyu Jeong Han, Ian Lane |
INTERSPEECH | 2 |
| 2014 | Robust language identification using convolutional neural network featuresabstractThe language identification (LID) task in the Robust Automatic Transcription of Speech (RATS) program is challenging due to the noisy nature of the audio data collected over highly degraded radio communication channels as well as the use of short duration speech segments for testing. In this paper, we report the recent advances made in the RATS LID task by using bottleneck features from a convolutional neural network (CNN). The CNN, which is trained with labelled data from one of target languages, generates bottleneck features which are used in a Gaussian mixture model (GMM)-ivector LID system. The CNN bottleneck features provide substantial complimentary information to the conventional acoustic features even on languages not seen in its training. Using these bottleneck features in conjunction with acoustic features, we obtain significant improvements (average relative improvements of 25% in terms of equal error rate (EER) compared to the corresponding acoustic system) for the LID task. Furthermore, these improvements are consistent for various choices of acoustic features as well as speech segment durations. Sriram Ganapathy, Kyu Jeong Han, Samuel Thomas 0001, Mohamed Kamal Omar, Maarten Van Segbroeck, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | TRAP language identification system for RATS phase II evaluationabstractAutomatic language identification or detection of audio data has become an important preprocessing step for speech/speaker recognition and audio data mining. In many surveillance applications, language detection has to be performed on highly degraded audio inputs. In this paper, we present our work on language detection in highly degraded radio channel scenarios. We provide a brief description of the Targeted Robust Audio Processing (TRAP) language detection system builtfor the Phase II Evaluationof the RobustAutomatic Transcription of Speech (RATS) program. This system is a combination of 15 systems with different frontends and speech activity decisions. We also analyze the usefulness of multi-layer perceptron (MLP) based non-linear projection of i-vectors before SVM classification. The proposed backend reduces the Equal Error Rate (EER) by 11%–25% relative compared to the baseline PCA-based feature representation for SVM classification, on the RATS test data consisting of data from eight highfrequency radio communication channels. Index Terms: Language identification (detection), highly degraded radio channel, RATS, i-vector, multi-layer perceptron. Kyu Jeong Han, Sriram Ganapathy, Ming Li 0026, Mohamed Kamal Omar, Shri Narayanan |
INTERSPEECH | 1 |
| 2013 | Automatic speaker age and gender recognition using acoustic and prosodic level information fusion
Ming Li 0026, Kyu Jeong Han, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2012 | Keyword-conditioned phone N-gram modeling with contextual information for speaker verificationabstractIn this paper we present our current work on automatic speaker recognition using keyword-conditioned phone N-gram modeling. We propose the use of contextual information around keywords in modeling a speaker's pronunciation characteristics at a phonetic level. Our approach is to add time margins around keywords when aligning keyword regions with keyword-specific phone events for feature vector generation. Including such additional information by incorporating time margins can capture idiosyncratic pronunciation information and is shown to help our keyword-conditioned phonetic speaker verification system achieve more than 50% (relative) performance improvement. This leads our high-level speaker verification system (i.e., fusion of non-conditioned and keyword-conditioned phonetic speaker verification systems) to currently achieve the best published result for the English 8-conversation enrollment telephony task of the 2008 NIST Speaker Recognition Evaluation for systems utilizing features not based directly on low-level acoustic information. Kyu Jeong Han, Jason W. Pelecanos, Mohamed Kamal Omar |
ICASSP | 1 |
| 2012 | Frame-based phonotactic Language IdentificationabstractThis paper describes a frame-based phonotactic Language Identification (LID) system, which was used for the LID evaluation of the Robust Automatic Transcription of Speech (RATS) program by the Defense Advanced Research Projects Agency (DARPA). The proposed approach utilizes features derived from frame-level phone log-likelihoods from a phone recognizer. It is an attempt to capture not only phone sequence information but also short-term timing information for phone N-gram events, which is lacking in conventional phonotactic LID systems that simply count phone N-gram events. Based on this new method, we achieved 26% relative improvement in terms of Cavgfor the RATS LID evaluation data compared to phone N-gram counts modeling. We also observed that it had a significant impact on score combination with our best acoustic system based on Mel-Frequency Cepstral Coefficients (MFCCs). Kyu Jeong Han, Jason W. Pelecanos |
SLT | 1 |
| 2011 | Forensically inspired approaches to automatic speaker recognitionabstractThis paper presents ongoing research leveraging forensic methods for automatic speaker recognition. Some of the methods forensic scientists employ include identifying speaker distinctive audio segments and comparing these segments using features such as pitch, formant, and other information. Other approaches have also involved performing a phonetic analysis to recognize idiolectal attributes, and an implicit analysis of the demographics of speakers. Inspired by these forensic phonetic approaches, we target three threads of work; hot-spot analysis, speaker style and pronunciation modelling, and demographics analysis. As a result of this work we show that a phonetic analysis conditioned on select speech events (or hot-spots) can outperform a phonetic analysis performed over all speech without conditioning. In the area of pronunciation modelling, one set of results demonstrate significantly improved robustness by exploiting phonetic structure in an automatic speech recognition system. For demographics analysis, we present state-of-the-art results of systems capable of detecting dialect, non-nativeness and native language. Kyu Jeong Han, Mohamed Kamal Omar, Jason W. Pelecanos, Cezar Pendus, Sibel Yaman, Weizhong Zhu |
ICASSP | 1 |
| 2010 | An improved cluster model selection method for agglomerative hierarchical speaker clustering using incremental Gaussian mixture models
Kyu Jeong Han, Shri Narayanan |
INTERSPEECH | 1 |
| 2010 | A variable frame length and rate algorithm based on the spectral kurtosis measure for speaker verification
Chi-Sang Jung, Kyu Jeong Han, Hyunson Seo, Shri Narayanan, Hong-Goo Kang |
INTERSPEECH | 2 |
| 2010 | Combining five acoustic level modeling methods for automatic speaker age and gender recognitionabstractAutomatic recognition of paralinguistic information from speech is important. Speaker identity, gender, age range, emotional state, etc. Guide human computer interaction systems to automatically adapt to different user needs. Ming Li 0026, Chi-Sang Jung, Kyu Jeong Han |
INTERSPEECH | 3 |
| 2010 | A cluster-profile representation of emotion using agglomerative hierarchical clusteringabstractThe proper representation of emotion is critical to automatic classification systems. In previous research, we demonstrated that emotion profile (EP) based representations are effective for this task. In EP-based representations, emotions are expressed in terms of underlying affective components from the subset of anger, happiness, neutrality, and sadness. The current study explores cluster profiles (CP), an alternate profile representa-tion in which the components are no longer semantic labels, but clusters inherent in the feature space. This unsupervised clus-tering of the feature space permits the application of a system-level semi-supervised learning paradigm. The results demon-strate that CPs are similarly discriminative to EPs (EP classifica-tion accuracy: 68.37 % vs. 69.25 % for the CP-based classifica-tion). This suggests that exhaustive labeling of a representative training corpus may not be necessary for emotion classification tasks. Emily Mower Provost, Kyu Jeong Han, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 2 |
| 2009 | Improved speaker diarization of meeting speech with recurrent selection of representative speech segments and participant interaction pattern modelingabstractIn this work we describe two distinct novel improvements to our speaker diarization system, previously proposed for analysis of meeting speech. The first approach focuses on recurrent selection of representative speech segments for speaker clustering while the other is based on participant interaction pattern modeling. The former selects speech segments with high relevance to speaker clustering, especially from a robust cluster modeling perspective, and keeps updating them throughout clustering procedures. The latter statistically models conversation patterns between meeting participants and applies it as a priori information when refining diarization results. Experimental results reveal that the two proposed approaches provide performance enhancement by 29.82% (relative) in terms of diarization error rate in tests on 13 meeting excerpts from various meeting speech corpora. Index Terms: speaker diarization, representative speech segments, participant interaction pattern modeling Kyu Jeong Han, Shri Narayanan |
INTERSPEECH | 1 |
| 2009 | Signature cluster model selection for incremental Gaussian mixture cluster modeling in agglomerative hierarchical speaker clusteringabstractAgglomerative hierarchical speaker clustering (AHSC) has been widely used for classifying speech data by speaker charac-teristics. Its bottom-up, one-way structure of merging the clos-est cluster pair at every recursion step, however, makes it diffi-cult to recover from incorrect merging. Hence, making AHSC robust to incorrect merging is an important issue. In this pa-per we address this problem in the framework of AHSC based on incremental Gaussian mixture models, which we previously introduced for better representing variable cluster size. Specif-ically, to minimize contamination in cluster models by hetero-geneous data, we select and keep updating a representative (or signature) model for each cluster during AHSC. Experiments on meeting speech excerpts (4 hours total) verify that the proposed approach improves average speaker clustering performance by approximately 20 % (relative). Index Terms: agglomerative hierarchical speaker clustering, incremental Gaussian mixture model, signature cluster model selection 1. Kyu Jeong Han, Shri Narayanan |
INTERSPEECH | 1 |
| 2009 | A Low-Complexity Dynamic Face-Voice Feature Fusion Approach to Multimodal Person RecognitionabstractIn this paper, we show the importance of face-voice correlation for audio-visual person recognition. We evaluate the performance of a system which uses the correlation between audio-visual features during speech against audio-only, video-only and audio-visual systems which use audio and visual features independently neglecting the interdependency of a person's spoken utterance and the associated facial movements. Experiments performed on the Vid-TIMIT dataset show that the proposed multimodal scheme has lower error rate than all other comparison conditions and is more robust against replay attacks. The simplicity of the fusion technique also allows the use of only one classifier which greatly simplifies system design and allows for a simple real-time DSP implementation. Dhaval Shah, Kyu Jeong Han, Shri Narayanan |
ISM | 2 |
| 2008 | Novel inter-cluster distance measure combining GLR and ICR for improved agglomerative hierarchical speaker clusteringabstractAgglomerative hierarchical clustering (AHC) has been a popular strategy for speaker clustering, due to its simple structure but acceptable level of performance. One of the main challenges in AHC that affects clustering performance is how to select the closest cluster pair for merging at every recursion. For this, generalized likelihood ratio (GLR) has been widely adopted as an inter-cluster distance measure. However, it tends to be affected by the size of the clusters considered, which could result in erroneous selection of the cluster pair to be merged during AHC. To tackle this problem, we propose a novel alternative to GLR in this paper, which is a combination of GLR and information change rate (ICR) that we recently introduced for addressing the aforementioned tendency of GLR. Experiments on various meeting speech data show that this combined measure improves clustering performance on average by around 30% (relative). Kyu Jeong Han, Shri Narayanan |
ICASSP | 1 |
| 2008 | Agglomerative hierarchical speaker clustering using incremental Gaussian mixture cluster modeling
Kyu Jeong Han, Shri Narayanan |
INTERSPEECH | 1 |
| 2008 | Multimodal Speaker Segmentation in Presence of Overlapped Speech SegmentsabstractWe propose a multimodal speaker segmentation algorithm with two main contributions: First, we suggest a hidden Markov model architecture that performs fusion of the three modalities: a multi-camera system for participant localization, a microphone array for speaker localization, and a speaker identification system; Second, we present a novel method for dealing with overlapped speech segments through a likelihood model of the microphone array observations that uses multiple local maxima of the Steered Power Response Generalized Cross Correlation Phase Transform (SPR-GCC-PHAT) function in the Joint Probabilistic Data Association (JPDA) framework. Results show that the proposed method outperforms standard speaker segmentation systems based on: (a) speaker identification and; (b) microphone array processing, for datasets with the significant portion (27.4%) of overlapped speech, and scores as high as 94.4% on the F-measure scale. Viktor Rozgic, Kyu Jeong Han, Panayiotis G. Georgiou, Shri Narayanan |
ISM | 2 |
| 2008 | The SAIL speaker diarization system for analysis of spontaneous meetingsabstractIn this paper, we propose a novel approach to speaker diarization of spontaneous meetings in our own multimodal SmartRoom environment. The proposed speaker diarization system first applies a sequential clustering concept to segmentation of a given audio data source, and then performs agglomerative hierarchical clustering for speaker-specific classification (or speaker clustering) of speech segments. The speaker clustering algorithm utilizes an incremental Gaussian mixture cluster modeling strategy, and a stopping point estimation method based on information change rate. Through experiments on various meeting conversation data of approximately 200 minutes total length, this system is demonstrated to provide diarization error rate of 18.90% on average. Kyu Jeong Han, Panayiotis G. Georgiou, Shri Narayanan |
MMSP | 1 |
| 2008 | Strategies to Improve the Robustness of Agglomerative Hierarchical Clustering Under Data Source Variation for Speaker DiarizationabstractMany current state-of-the-art speaker diarization systems exploit agglomerative hierarchical clustering (AHC) as their speaker clustering strategy, due to its simple processing structure and acceptable level of performance. However, AHC is known to suffer from performance robustness under data source variation. In this paper, we address this problem. We specifically focus on the issues associated with the widely used clustering stopping method based on Bayesian information criterion (BIC) and the merging-cluster selection scheme based on generalized likelihood ratio (GLR). First, we propose a novel alternative stopping method for AHC based on information change rate (ICR). Through experiments on several meeting corpora, the proposed method is demonstrated to be more robust to data source variation than the BIC-based one. The average improvement obtained in diarization error rate (DER) by this method is 8.76% (absolute) or 35.77% (relative). We also introduce a selective AHC (SAHC) in the paper, which first runs AHC with the ICR-based stopping method only on speech segments longer than 3 s and then classifies shorter speech segments into one of the clusters given by the initial AHC. This modified version of AHC is motivated by our previous analysis that the proportion of short speech turns (or segments) in a data source is a significant factor contributing to the robustness problem arising in the GLR-based merging-cluster selection scheme. The additional performance improvement obtained by SAHC is 3.45% (absolute) or 14.08% (relative) in terms of averaged DER. Kyu Jeong Han, Samuel Kim, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Robust speaker clustering strategies to data source variation for improved speaker diarizationabstractAgglomerative hierarchical clustering (AHC) has been widely used in speaker diarization systems to classify speech segments in a given data source by speaker identity, but is known to be not robust to data source variation. In this paper, we identify one of the key potential sources of this variability that negatively affects clustering error rate (CER), namely short speech segments, and propose three solutions to tackle this issue. Through experiments on various meeting conversation excerpts, the proposed methods are shown to outperform simple AHC in terms of relative CER improvements in the range of 17-32%. Kyu Jeong Han, Samuel Kim, Shri Narayanan |
ASRU | 1 |
| 2007 | A robust stopping criterion for agglomerative hierarchical clustering in a speaker diarization systemabstractAgglomerative hierarchical clustering (AHC) is an unsupervised classification strategy of merging the closest pair of clusters recursively, and has been widely used in speaker diarization systems to classify speech segments by speaker identity. The most critical part in AHC is how to automatically stop the recursive process at the point when clustering error rate reaches its lowest possible value, for which a BIC-based stopping criterion has been widely used. However, this criterion is not robust to data source variation. In this paper, we examine the criterion to establish the cause for the robustness issue and, based on this, propose an improved stopping criterion. Experimental results based on meeting conversation excerpts randomly chosen from various meeting speech corpora indicate that the proposed criterion is superior to the BIC-based one, showing that clustering error rate is improved on average by 7.28 % (absolute) and Kyu Jeong Han, Shri Narayanan |
INTERSPEECH | 1 |
| 2004 | A distributed speech recognition system in multi-user environments
Kyu Jeong Han, Shri Narayanan, Naveen Srinivasamurthy |
INTERSPEECH | 1 |
| 2004 | Robust speech recognition over packet networks: an overview
Naveen Srinivasamurthy, Kyu Jeong Han, Shri Narayanan |
INTERSPEECH | 2 |
| 2002 | Iterative decoding of a differential space-time block code with low complexityabstractLike differential phase shift keying (DPSK), a differential space-time block code has recursive encoding structure and then, when it is concatenated serially with a channel code, additional coding gain from the concatenation is obtained. To obtain the additional gain, recent research has introduced serial concatenation of the differential space-time block code and channel code, and its sub-optimal iterative receiver. The receiver consists of two APP decoders, which not only has considerable computational complexity, but also is not suitable, especially for the differential space-time block code. This paper proposes a simpler, more suitable decoder for the differential space-time block code in the serial concatenation than the APP decoder. The core idea of the proposed decoder is to employ a maximum likelihood decoding method. With low complexity, a receiver with the proposed decoder shows performance almost identical to the conventional one in a time-correlated fading channel. Kyu Jeong Han, Jae Hong Lee |
VTC Spring | 1 |