EDBT 2026 Demo / reviewers in the wild / expert
Suwon Shon
dblp:81/11275
· DBLP profile ↗
34ranked-venue papers
20as first author
11since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 19 first-author · 8 since 2021Artificial intelligence and machine learning · 19 · 10 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?abstractReference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of the recording.In this paper, we examine whether summaries based on annotators listening to the recordings differ from those based on annotators reading transcripts.Using existing intrinsic evaluation based on human evaluation, automatic metrics, LLM-based evaluation, and a retrieval-based reference-free method.We find that summaries are indeed different based on the source modality, and that speechbased summaries are more factually consistent and information-selective than transcript-based summaries.Meanwhile, transcript-based summaries are impacted by recognition errors in the source, and expert-written summaries are more informative and reliable.We make all the collected data and analysis code public 1 to facilitate the reproduction of our work and advance research in this area. Roshan S. Sharma, Suwon Shon, Mark Lindsey, Hira Dhamyal, Bhiksha Raj |
ACL (1) | 2 |
| 2024 | Generative Context-Aware Fine-Tuning of Self-Supervised Speech ModelsabstractWhen performing tasks like automatic speech recognition or spoken language understanding for a given utterance, access to preceding text or audio provides contextual information that can improve performance. Considering the recent advances in generative large language models (LLM), we hypothesize that an LLM could generate useful context information using the preceding text. With appropriate prompts, LLM could generate a prediction of the next sentence or abstractive text like titles or topics. In this paper, we study the use of LLM-generated context information and propose an approach to distill the generated information during fine-tuning of self-supervised speech models, which we refer to as generative context-aware fine-tuning. This approach allows the fine-tuned model to make improved predictions without access to the true surrounding segments or to the LLM at inference time, while requiring only a very small additional context module. We evaluate the proposed approach using the SLUE and Libri-light benchmarks for several downstream tasks: automatic speech recognition, named entity recognition, and sentiment analysis. The results show that generative context-aware fine-tuning outperforms a context injection fine-tuning approach that accesses the ground-truth previous text, and is competitive with a generative context injection fine-tuning approach that requires the LLM at inference time. Suwon Shon, Kwangyoun Kim, Prashant Sridhar, Yi-Te Hsu, Shinji Watanabe 0001, Karen Livescu |
ICASSP | 1 |
| 2024 | Improving ASR Contextual Biasing with Guided AttentionabstractIn this paper, we propose a Guided Attention (GA) auxiliary training loss, which improves the effectiveness and robustness of automatic speech recognition (ASR) contextual biasing without introducing additional parameters. A common challenge in previous literature is that the word error rate (WER) reduction brought by contextual biasing diminishes as the number of bias phrases increases. To address this challenge, we employ a GA loss as an additional training objective besides the Transducer loss. The proposed GA loss aims to teach the cross attention how to align bias phrases with text tokens or audio frames. Compared to studies with similar motivations, the proposed loss operates directly on the cross attention weights and is easier to implement. Through extensive experiments based on Conformer Transducer with Contextual Adapter, we demonstrate that the proposed method not only leads to a lower WER but also retains its effectiveness as the number of bias phrases increases. Specifically, the GA loss decreases the WER of rare vocabularies by up to 19.2% on LibriSpeech compared to the contextual biasing baseline, and up to 49.3% compared to a vanilla Transducer. Jiyang Tang, Kwangyoun Kim, Suwon Shon, Felix Wu, Prashant Sridhar |
ICASSP | 3 |
| 2024 | Convolution-Augmented Parameter-Efficient Fine-Tuning for Speech Recognition
Kwangyoun Kim, Suwon Shon, Yi-Te Hsu, Prashant Sridhar, Karen Livescu, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2024 | DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language UnderstandingabstractThe integration of pre-trained text-based large language models (LLM) with speech input has enabled instruction-following capabilities for diverse speech tasks.This integration requires the use of a speech encoder, a speech adapter, and an LLM, trained on diverse tasks.We propose the use of discrete speech units (DSU), rather than continuous-valued speech encoder outputs, that are converted to the LLM token embedding space using the speech adapter.We generate DSU using a selfsupervised speech encoder followed by k-means clustering.The proposed model shows robust performance on speech inputs from seen/unseen domains and instruction-following capability in spoken question answering.We also explore various types of DSU extracted from different layers of the self-supervised speech encoder, as well as Mel frequency Cepstral Coefficients (MFCC).Our findings suggest that the ASR task and datasets are not crucial in instruction-tuning for spoken question answering tasks. Suwon Shon, Kwangyoun Kim, Yi-Te Hsu, Prashant Sridhar, Shinji Watanabe 0001, Karen Livescu |
INTERSPEECH | 1 |
| 2023 | SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding TasksabstractSuwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad, Felix Wu, Roshan S Sharma, Wei-Lun Wu, Hung-yi Lee, Karen Livescu, Shinji Watanabe. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Suwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad, Felix Wu, Roshan S. Sharma, Wei-Lun Wu, Hung-yi Lee, Karen Livescu, Shinji Watanabe 0001 |
ACL (1) | 1 |
| 2023 | Context-Aware Fine-Tuning of Self-Supervised Speech ModelsabstractSelf-supervised pre-trained transformers have improved the state of the art on a variety of speech tasks. Due to the quadratic time and space complexity of self-attention, they usually operate at the level of relatively short (e.g., utterance) segments. In this paper, we study the use of context, i.e., surrounding segments, during fine-tuning and propose a new approach called context-aware fine-tuning. We attach a context module on top of the last layer of a pre-trained model to encode the whole segment into a context embedding vector which is then used as an additional feature for the final prediction. During the fine-tuning stage, we introduce an auxiliary loss that encourages this context embedding vector to be similar to context vectors of surrounding segments. This allows the model to make predictions without access to these surrounding segments at inference time and requires only a tiny overhead compared to standard fine-tuned models. We evaluate the proposed approach using the SLUE and Librilight benchmarks for several downstream tasks: Automatic speech recognition (ASR), named entity recognition (NER), and sentiment analysis (SA). The results show that context-aware fine-tuning not only outperforms a standard fine-tuning baseline but also rivals a strong context injection baseline that uses neighboring speech segments during inference. Suwon Shon, Felix Wu, Kwangyoun Kim, Prashant Sridhar, Karen Livescu, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2023 | A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks
Yifan Peng 0003, Kwangyoun Kim, Felix Wu, Brian Yan, Siddhant Arora, Jiyang Tang, Suwon Shon, Prashant Sridhar, Shinji Watanabe 0001 |
INTERSPEECH | 8 |
| 2022 | SLUE: New Benchmark Tasks For Spoken Language Understanding Evaluation on Natural SpeechabstractProgress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest has been growing in higher-level spoken language understanding tasks, including using end-to-end models, but there are fewer annotated datasets for such tasks. At the same time, recent work shows the possibility of pre-training generic representations and then fine-tuning for several tasks using relatively little labeled data. We propose to create a suite of benchmark tasks for Spoken Language Understanding Evaluation (SLUE) consisting of limited-size labeled training sets and corresponding evaluation sets. This resource would allow the research community to track progress, evaluate pre-trained representations for higher-level tasks, and study open questions such as the utility of pipeline versus end-to-end approaches. We present the first phase of the SLUE benchmark suite, consisting of named entity recognition, sentiment analysis, and ASR on the corresponding datasets. We focus on naturally produced (not read or synthesized) speech, and freely available datasets. We pro-vide new transcriptions and annotations on subsets of the VoxCeleb and VoxPopuli datasets, evaluation metrics and results for baseline models, and an open-source toolkit to reproduce the baselines and evaluate new models. Suwon Shon, Ankita Pasad, Felix Wu, Pablo Brusco, Yoav Artzi, Karen Livescu, Kyu Jeong Han |
ICASSP | 1 |
| 2022 | On the Use of External Data for Spoken Named Entity RecognitionabstractAnkita Pasad, Felix Wu, Suwon Shon, Karen Livescu, Kyu Han. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ankita Pasad, Felix Wu, Suwon Shon, Karen Livescu, Kyu Jeong Han |
NAACL-HLT | 3 |
| 2021 | Leveraging Pre-Trained Language Model for Speech Sentiment AnalysisabstractIn this paper, we explore the use of pre-trained language models to learn sentiment information of written texts for speech sentiment analysis.First, we investigate how useful a pre-trained language model would be in a 2-step pipeline approach employing Automatic Speech Recognition (ASR) and transcripts-based sentiment analysis separately.Second, we propose a pseudo label-based semi-supervised training strategy using a language model on an end-to-end speech sentiment approach to take advantage of a large, but unlabeled speech dataset for training.Although spoken and written texts have different linguistic characteristics, they can complement each other in understanding sentiment.Therefore, the proposed system can not only model acoustic characteristics to bear sentimentspecific information in speech signals, but learn latent information to carry sentiments in the text representation.In these experiments, we demonstrate the proposed approaches improve F1 scores consistently compared to systems without a language model.Moreover, we also show that the proposed framework can reduce 65% of human supervision by leveraging a large amount of data without human sentiment annotation and boost performance in a low-resource condition where the human sentiment annotation is not available enough. Suwon Shon, Pablo Brusco, Kyu Jeong Han, Shinji Watanabe 0001 |
Interspeech | 1 |
| 2020 | ADI17: A Fine-Grained Arabic Dialect Identification DatasetabstractIn this paper, we describe a method to collect dialectal speech from YouTube videos to create a large-scale Dialect Identification (DID) dataset. Using this method, we collected dialectal Arabic from known YouTube channels from 17 Arabic speaking countries in the Middle East and Northern Africa. After a refinement process, a total of 3,000 hours of speech was available for training DID systems, with an additional 57 hours of speech for development and testing. For detailed evaluations, the DID data was divided into three sub-categories based on the segment duration: short (less than 5s), medium (5-20s), and long (over 20s). We compare state-of-the-art DID techniques on these data, and also analyze a DID system trained on these data. Since the training and test data share the same channel domain, we also used the Multi-Genre Broadcast 3 (MGB-3) test set to evaluate on domain mismatched condition. Suwon Shon, Ahmed Ali 0002, Younes Samih, Hamdy Mubarak, James R. Glass |
ICASSP | 1 |
| 2020 | What Does an End-to-End Dialect Identification Model Learn About Non-Dialectal Information?
Shammur Absar Chowdhury, Ahmed Ali 0002, Suwon Shon, James R. Glass |
INTERSPEECH | 3 |
| 2020 | Multimodal Association for Speaker Verification
Suwon Shon, James R. Glass |
INTERSPEECH | 1 |
| 2019 | The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic SpeechabstractThis paper describes the fifth edition of the Multi-Genre Broadcast Challenge (MGB-5), an evaluation focused on Arabic speech recognition and dialect identification. MGB-5 extends the previous MGB-3 challenge in two ways: first it focuses on Moroccan Arabic speech recognition; second the granularity of the Arabic dialect identification task is increased from 5 dialect classes to 17, by collecting data from 17 Arabic speaking countries. Both tasks use YouTube recordings to provide a multi-genre multi-dialectal challenge in the wild. Moroccan speech transcription used about 13 hours of transcribed speech data, split across training, development, and test sets, covering 7-genres: comedy, cooking, family/kids, fashion, drama, sports, and science (TEDx). The fine-grained Arabic dialect identification data was collected from known YouTube channels from 17 Arabic countries. 3,000 hours of this data was released for training, and 57 hours for development and testing. The dialect identification data was divided into three sub-categories based on the segment duration: short (under 5 s), medium (5-20 s), and long (>20 s). Overall, 25 teams registered for the challenge, and 9 teams submitted systems for the two tasks. We outline the approaches adopted in each system and summarize the evaluation results. Ahmed Ali 0002, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James R. Glass, Steve Renals, Khalid Choukri |
ASRU | 2 |
| 2019 | Domain Mismatch Robust Acoustic Scene Classification Using Channel Information ConversionabstractIn recent acoustic scene classification (ASC) research field, training and test device channel mismatch have become an issue for the real world implementation. To address the issue, this paper proposes a channel domain conversion using factorized hierarchical variational autoencoder. Proposed method adapts both the source and target domain to a pre-defined specific domain. Unlike the conventional approach, the relationship between the target and source domain and information of each domain are not required in the adaptation process. Based on the experimental results using the IEEE Detection and Classification of Acoustic Scenes and Event 2018 task 1-B dataset and the baseline system, it is shown that the proposed approach can mitigate the channel mismatching issue of different recording devices. Seongkyu Mun, Suwon Shon |
ICASSP | 2 |
| 2019 | Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target DomainabstractEnd-to-end deep learning language or dialect identification systems operate on the spectrogram or other acoustic feature and directly generate identification scores for each class. An important issue for end-to-end systems is to have some knowledge of the application domain, because the system can be vulnerable to use cases that were not seen in the training phase; such a scenario is often referred to as a domain mismatched condition. In general, we assume that there is enough variation in the training dataset to expose the system to multiple domains. In this work, we study how to best make use a training dataset in order to have maximum effectiveness on unknown target domains. Our goal is to process the input without any knowledge of the target domain while preserving robust performance on other domains as well. To accomplish this objective, we propose a domain attentive fusion approach for end-to-end dialect/language identification systems. To help with experimentation, we collect a dataset from three different domains, and create experimental protocols for a domain mismatched condition. The results of our proposed approach, which were tested on a variety of broadcast and YouTube data, shows significant performance gain compared to traditional approaches, even without any prior target domain information. Suwon Shon, Ahmed Ali 0002, James R. Glass |
ICASSP | 1 |
| 2019 | Noise-tolerant Audio-visual Online Person Verification Using an Attention-based Neural Network FusionabstractIn this paper, we present a multi-modal online person verification system using both speech and visual signals. Inspired by neuroscientific findings on the association of voice and face, we propose an attention-based end-to-end neural network that learns multi-sensory association for the task of person verification. The attention mechanism in our proposed network learns to conditionally select a salient modality between speech and facial representations that provides a balance between complementary inputs. By virtue of this capability, the network is robust to missing or corrupted data from either modality. In the VoxCeleb2 dataset, we show that our method performs favorably against competing multi-modal methods. Even for extreme cases of large corruption or missing data on either modality, our method demonstrates robustness over other unimodal methods. Suwon Shon, Tae-Hyun Oh, James R. Glass |
ICASSP | 1 |
| 2019 | MCE 2018: The 1st Multi-Target Speaker Detection and Identification Challenge EvaluationabstractThe Multi-target Challenge aims to assess how well current speech technology is able to determine whether or not a recorded utterance was spoken by one of a large number of blacklisted speakers. It is a form of multi-target speaker detection based on real-world telephone conversations. Data recordings are generated from call center customer-agent conversations. The task is to measure how accurately one can detect 1) whether a test recording is spoken by a blacklisted speaker, and 2) which specific blacklisted speaker was talking. This paper outlines the challenge and provides its baselines, results, and discussions. Suwon Shon, Najim Dehak, Douglas A. Reynolds, James R. Glass |
INTERSPEECH | 1 |
| 2019 | Large-Scale Speaker Retrieval on Random Speaker Variability SubspaceabstractThis paper describes a fast speaker search system to retrieve segments of the same voice identity in the large-scale data.A recent study shows that Locality Sensitive Hashing (LSH) enables quick retrieval of a relevant voice in the large-scale data in conjunction with i-vector while maintaining accuracy.In this paper, we proposed Random Speaker-variability Subspace (RSS) projection to map a data into LSH based hash tables.We hypothesized that rather than projecting on completely random subspace without considering data, projecting on randomly generated speaker variability space would give more chance to put the same speaker representation into the same hash bins, so we can use less number of hash tables.Multiple RSS can be generated by randomly selecting a subset of speakers from a large speaker cohort.From the experimental result, the proposed approach shows 100 times and 7 times faster than the linear search and LSH, respectively. Suwon Shon, Younggun Lee, Taesu Kim |
INTERSPEECH | 1 |
| 2019 | VoiceID Loss: Speech Enhancement for Speaker VerificationabstractIn this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly used loss functions for speech enhancement such as the L2 loss, the VoiceID loss is based on the feedback from a speaker verification model to generate a ratio mask. The generated ratio mask is multiplied pointwise with the original spectrogram to filter out unnecessary components for speaker verification. In the experiments, we observed that the enhancement network, after training with the VoiceID loss, is able to ignore a substantial amount of time-frequency bins, such as those dominated by noise, for verification. The resulting model consistently improves the speaker verification system on both clean and noisy conditions. Suwon Shon, Hao Tang 0002, James R. Glass |
INTERSPEECH | 1 |
| 2019 | State-of-the-Art Speaker Recognition for Telephone and Video Speech: The JHU-MIT Submission for NIST SRE18
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, Réda Dehak, L. Paola García-Perera, Daniel Povey, Pedro A. Torres-Carrasquillo, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 9 |
| 2019 | Time-Contrastive Learning Based Deep Bottleneck Features for Text-Dependent Speaker VerificationabstractThere are a number of studies about extraction of bottleneck (BN) features from deep neural networks (DNNs) trained to discriminate speakers, pass-phrases, and triphone states for improving the performance of text-dependent speaker verification (TD-SV). However, a moderate success has been achieved. A recent study presented a time contrastive learning (TCL) concept to explore the non-stationarity of brain signals for classification of brain states. Speech signals have similar non-stationarity property, and TCL further has the advantage of having no need for labeled data. We therefore present a TCL based BN feature extraction method. The method uniformly partitions each speech utterance in a training dataset into a predefined number of multi-frame segments. Each segment in an utterance corresponds to one class, and class labels are shared across utterances. DNNs are then trained to discriminate all speech frames among the classes to exploit the temporal structure of speech. In addition, we propose a segment-based unsupervised clustering algorithm to re-assign class labels to the segments. TD-SV experiments were conducted on the RedDots challenge database. The TCL-DNNs were trained using speech data of fixed pass-phrases that were excluded from the TD-SV evaluation set, so the learned features can be considered phrase-independent. We compare the performance of the proposed TCL BN feature with those of short-time cepstral features and BN features extracted from DNNs discriminating speakers, pass-phrases, speaker+pass-phrase, as well as monophones whose labels and boundaries are generated by three different automatic speech recognition (ASR) systems. Experimental results show that the proposed TCL-BN outperforms cepstral features and speaker+pass-phrase discriminant BN features, and its performance is on par with those of ASR derived BN features. Moreover, the clustering method improves the TD-SV performance of TCL-BN and ASR derived BN features with respect to their standalone counterparts. We further study the TD-SV performance of fusing cepstral and BN features. Achintya Kumar Sarkar, Zheng-Hua Tan, Hao Tang 0002, Suwon Shon, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Exploiting Convolutional Neural Networks for Phonotactic Based Dialect IdentificationabstractIn this paper, we investigate different approaches for Dialect Identification (DID) in Arabic broadcast speech. Dialects differ in their inventory of phonological segments. This paper proposes a new phonotactic based feature representation approach which enables discrimination among different occurrences of the same phone n-grams with different phone duration and probability statistics. To achieve further gain in accuracy we used multi-lingual phone recognizers, trained separately on Arabic, English, Czech, Hungarian and Russian languages. We use Support Vector Machines (SVMs), and Convolutional Neural Networks (CNN s) as backend classifiers throughout the study. The final system fusion results in 24.7% and 19.0% relative error rate reduction compared to that of a conventional phonotactic DID, and i-vectors with bottleneck features. Maryam Najafian, Sameer Khurana, Suwon Shon, Ahmed Ali 0002, James R. Glass |
ICASSP | 3 |
| 2018 | Unsupervised Representation Learning of Speech for Dialect IdentificationabstractIn this paper, we explore the use of a factorized hierarchical variational autoencoder (FHVAE) model to learn an unsupervised latent representation for dialect identification (DID). An FHVAE can learn a latent space that separates the more static attributes within an utterance from the more dynamic attributes by encoding them into two different sets of latent variables. Useful factors for dialect identification, such as phonetic or linguistic content, are encoded by a segmental latent variable, while irrelevant factors that are relatively constant within a sequence, such as a channel or a speaker information, are encoded by a sequential latent variable. The disentanglement property makes the segmental latent variable less susceptible to channel and speaker variation, and thus reduces degradation from channel domain mismatch. We demonstrate that on fully-supervised DID tasks, an end-to-end model trained on the features extracted from the FHVAE model achieves the best performance, compared to the same model trained on conventional acoustic features and an i-vector based system. Moreover, we also show that the proposed approach can leverage a large amount of unlabeled data for FHVAE training to learn domain-invariant features for DID, and significantly improve the performance in a low-resource condition, where the labels for the in-domain data are not available. Suwon Shon, Wei-Ning Hsu, James R. Glass |
SLT | 1 |
| 2018 | Frame-Level Speaker Embeddings for Text-Independent Speaker Recognition and Analysis of End-to-End ModelabstractIn this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently with linear activation in the embedding layer. To understand how the speaker recognition model operates with text-independent input, we modify the structure to extract frame-level speaker embeddings from each hidden layer. We feed utterances from the TIMIT dataset to the trained network and use several proxy tasks to study the networks ability to represent speech input and differentiate voice identity. We found that the networks are better at discriminating broad phonetic classes than individual phonemes. In particular, frame-level embeddings that belong to the same phonetic classes are similar (based on cosine distance) for the same speaker. The frame level representation also allows us to analyze the networks at the frame level, and has the potential for other analyses to improve speaker recognition. Suwon Shon, Hao Tang 0002, James R. Glass |
SLT | 1 |
| 2017 | MIT-QCRI Arabic dialect identification system for the 2017 multi-genre broadcast challengeabstractIn order to successfully annotate the Arabic speech content found in open-domain media broadcasts, it is essential to be able to process a diverse set of Arabic dialects. For the 2017 Multi-Genre Broadcast challenge (MGB-3) there were two possible tasks: Arabic speech recognition, and Arabic Dialect Identification (ADI). In this paper, we describe our efforts to create an ADI system for the MGB-3 challenge, with the goal of distinguishing amongst four major Arabic dialects, as well as Modern Standard Arabic. Our research focused on dialect variability and domain mismatches between the training and test domain. In order to achieve a robust ADI system, we explored both Siamese neural network models to learn similarity and dissimilarities among Arabic dialects, as well as i-vector post-processing to adapt domain mismatches. Both Acoustic and linguistic features were used for the final MGB-3 submissions, with the best primary system achieving 75% accuracy on the official 10hr test set. Suwon Shon, Ahmed Ali 0002, James R. Glass |
ASRU | 1 |
| 2017 | Deep Neural Network based learning and transferring mid-level audio features for acoustic scene classificationabstractDeep Neural Network (DNN) based transfer learning has been shown to be effective in Visual Object Classification (VOC) for complementing the deficit of target domain training samples by adapting classifiers that have been pre-trained for other large-scaled DataBase (DB). Although there exists an abundance of acoustic data, it can also be said that datasets of specific acoustic scenes are sparse for training Acoustic Scene Classification (ASC) models. By exploiting VOC DNN's ability of learning beyond its pre-trained environments, this paper proposes DNN based transfer learning for ASC. Effectiveness of the proposed method is demonstrated on the database of IEEE DCASE Challenge 2016 Task 1 and home surveillance environment via representative experiments. Its improved performance is verified by comparing it to prominent conventional methods. Seongkyu Mun, Suwon Shon, Wooil Kim, David K. Han, Hanseok Ko |
ICASSP | 2 |
| 2017 | Recursive Whitening Transformation for Speaker Recognition on Language Mismatched ConditionabstractRecently in speaker recognition, performance degradation due to the channel domain mismatched condition has been actively addressed. However, the mismatches arising from language is yet to be sufficiently addressed. This paper proposes an approach which employs recursive whitening transformation to mitigate the language mismatched condition. The proposed method is based on the multiple whitening transformation, which is intended to remove un-whitened residual components in the dataset associated with i-vector length normalization. The experiments were conducted on the Speaker Recognition Evaluation 2016 trials of which the task is non-English speaker recognition using development dataset consist of both a large scale out-of-domain (English) dataset and an extremely low-quantity in-domain (non-English) dataset. For performance comparison, we develop a state-of- the-art system using deep neural network and bottleneck feature, which is based on a phonetically aware model. From the experimental results, along with other prior studies, effectiveness of the proposed method on language mismatched condition is validated. Suwon Shon, Seongkyu Mun, Hanseok Ko |
INTERSPEECH | 1 |
| 2017 | Autoencoder Based Domain Adaptation for Speaker Recognition Under Insufficient Channel InformationabstractIn real-life conditions, mismatch between development and test domain degrades speaker recognition performance. To solve the issue, many researchers explored domain adaptation approaches using matched in-domain dataset. However, adaptation would be not effective if the dataset is insufficient to estimate channel variability of the domain. In this paper, we explore the problem of performance degradation under such a situation of insufficient channel information. In order to exploit limited in-domain dataset effectively, we propose an unsupervised domain adaptation approach using Autoencoder based Domain Adaptation (AEDA). The proposed approach combines an autoencoder with a denoising autoencoder to adapt resource-rich development dataset to test domain. The proposed technique is evaluated on the Domain Adaptation Challenge 13 experimental protocols that is widely used in speaker recognition for domain mismatched condition. The results show significant improvements over baselines and results from other prior studies. Suwon Shon, Seongkyu Mun, Wooil Kim, Hanseok Ko |
INTERSPEECH | 1 |
| 2016 | Deep Neural Network Bottleneck Features for Acoustic Event Recognition
Seongkyu Mun, Suwon Shon, Wooil Kim, Hanseok Ko |
INTERSPEECH | 2 |
| 2015 | Maximum likelihood Linear Dimension Reduction of heteroscedastic feature for robust Speaker RecognitionabstractThis paper analyzes heteroscedasticity in i-vector for robust forensics and surveillance speaker recognition system. Linear Discriminant Analysis (LDA), a widely-used linear dimension reduction technique, assumes that classes are homoscedastic within a same covariance. In this paper it is assumed that general speech utterances contain both homoscedastic and heteroscedastic elements. We show the validity of this assumption by employing several analyses and also demonstrate that dimension reduction using principal components is feasible. To effectively handle the presence of heteroscedastic and homoscedastic elements, we propose a fusion approach of applying both LDA and Heteroscedastic-LDA (HLDA). The experiments are conducted to show its effectiveness and compare to other methods using the telephone database of National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) 2010 extended. Suwon Shon, Seongkyu Mun, David K. Han, Hanseok Ko |
AVSS | 1 |
| 2014 | Generalized cross-correlation based noise robust abnormal acoustic event localization utilizing non-negative matrix factorizationabstractIn this paper, robust sound source localization for surveillance system is presented. In particular, we propose an algorithm for abnormal acoustic event localization using non-negative matrix factorization based frequency bin weighting. Based on the abnormal acoustic event localization experiments in real acoustic environment, the proposed algorithm's excellent strength is validated in terms of representative performance measures compared to the conventional method. Sungkyu Moon, Suwon Shon, Wooil Kim, David K. Han |
AVSS | 2 |
| 2013 | Abnormal acoustic event localization based on selective frequency bin in high noise environment for audio surveillanceabstractIn this paper, a method for source localization for surveillance system is presented. In particular, we propose an algorithm for abnormal acoustic event localization based on a novel approach of relevant frequency bin selections by statistical analyses. By means of selective frequency bin, it becomes possible to localize the event more accurately in high noise environment with low computational complexity. The effectiveness is verified through the experimental results in varied noise environments with different levels of Signal to Noise Ratio (SNR). Suwon Shon, David K. Han, Hanseok Ko |
AVSS | 1 |