VLDB 2026 Research / reviewers in the wild / expert
Alex Park 0001
dblp:69/2863 · also Alex Seungryong Park
· DBLP profile ↗
14ranked-venue papers
5as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Locale Encoding for Scalable Multilingual Keyword Spotting ModelsabstractA Multilingual Keyword Spotting (KWS) system detects spoken keywords over multiple locales. Conventional monolingual KWS approaches do not scale well to multilingual scenarios because of high development/maintenance costs and lack of resource sharing. To overcome this limit, we propose two locale-conditioned universal models with locale feature concatenation and feature-wise linear modulation (FiLM). We compare these models with two baseline methods: locale-specific monolingual KWS, and a single universal model trained over all data. Experiments over 10 localized language datasets show that locale-conditioned models substantially improve accuracy over baseline methods across all locales in different noise conditions. FiLM performed the best, improving on average FRR by 61% (relative) compared to monolingual KWS models of similar sizes. Pai Zhu, Hyun Jin Park, Alex Park 0001, Angelo Scorza Scarpati, Ignacio López-Moreno |
ICASSP | 3 |
| 2022 | Production federated keyword spotting via distillation, filtering, and joint federated-centralized training
Andrew Hard, Kurt Partridge, Neng Chen, Sean Augenstein, Aishanee Shah, Hyun Jin Park, Alex Park 0001, Sara Ng, Jessica Nguyen, Ignacio López-Moreno, Rajiv Mathews, Françoise Beaufays |
INTERSPEECH | 7 |
| 2022 | A Conformer-based Waveform-domain Neural Acoustic Echo Canceller Optimized for ASR AccuracyabstractAcoustic Echo Cancellation (AEC) is essential for accurate recognition of queries spoken to a smart speaker that is playing out audio.Previous work has shown that a neural AEC model operating on log-mel spectral features (denoted "logmel" hereafter) can greatly improve Automatic Speech Recognition (ASR) accuracy when optimized with an auxiliary loss utilizing a pre-trained ASR model encoder.In this paper, we develop a conformer-based waveform-domain neural AEC model inspired by the "TasNet" architecture.The model is trained by jointly optimizing Negative Scale-Invariant SNR (SISNR) and ASR losses on a large speech dataset.On a realistic rerecorded test set, we find that cascading a linear adaptive AEC and a waveform-domain neural AEC is very effective, giving 56-59% word error rate (WER) reduction over the linear AEC alone.On this test set, the 1.6M parameter waveform-domain neural AEC also improves over a larger 6.5M parameter logmeldomain neural AEC model by 20-29% in easy to moderate conditions.By operating on smaller frames, the waveform neural model is able to perform better at smaller sizes and is better suited for applications where memory is limited. Sankaran Panchapagesan, Arun Narayanan, Turaj Zakizadeh Shabestary, Nathan Howard, Alex Park 0001, James Walker, Alexander Gruenstein |
INTERSPEECH | 6 |
| 2021 | A Conformer-Based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech SeparationabstractWe present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. This is achieved by using a contextual enhancement neural network that can optionally make use of different types of side inputs: (1) a reference signal of the playback audio, which is necessary for echo cancellation; (2) a noise context, which is useful for speech enhancement; and (3) an embedding vector representing the voice characteristic of the target speaker of interest, which is not only critical in speech separation, but also helpful for echo cancellation and speech enhancement. We present detailed evaluations to show that the joint model performs almost as well as the task-specific models, and significantly reduces word error rate in noisy conditions even when using a large-scale state-of-the-art ASR model. Compared to the noisy baseline, the joint model reduces the word error rate in low signal-to-noise ratio conditions by at least 71% on our echo cancellation dataset, 10% on our noisy dataset, and 26% on our multi-speaker dataset. Compared to task-specific models, the joint model performs within 10% on our echo cancellation dataset, 2% on the noisy dataset, and 3% on the multi-speaker dataset. Tom O'Malley, Arun Narayanan, Alex Park 0001, James Walker, Nathan Howard |
ASRU | 4 |
| 2021 | A Neural Acoustic Echo Canceller Optimized Using An Automatic Speech Recognizer and Large Scale Synthetic DataabstractWe consider the problem of recognizing speech utterances spoken to a device which is generating a known sound waveform; for example, recognizing queries issued to a digital assistant which is generating responses to previous user inputs. Previous work has proposed building acoustic echo cancellation (AEC) models for this task that optimize speech enhancement metrics using both neural network as well as signal processing approaches.Since our goal is to recognize the input speech, we consider enhancements which improve word error rates (WERs) when the predicted speech signal is passed to an automatic speech recognition (ASR) model. First, we augment the loss function with a term that produces outputs useful to a pre-trained ASR model and show that this augmented loss function improves WER metrics. Second, we demonstrate that augmenting our training dataset of real world examples with a large synthetic dataset improves performance. Crucially, applying SpecAugment style masks to the reference channel during training aids the model in adapting from synthetic to real domains. In experimental evaluations, we find the proposed approaches improve performance, on average, by 57% over a signal processing baseline and 45% over the neural AEC model without the proposed changes. Nathan Howard, Alex Park 0001, Turaj Zakizadeh Shabestary, Alexander Gruenstein, Rohit Prabhavalkar |
ICASSP | 2 |
| 2007 | Making Sense of Sound: Unsupervised Topic Segmentation over Acoustic Input
Igor Malioutov, Alex Park 0001, Regina Barzilay, James R. Glass |
ACL | 2 |
| 2006 | Unsupervised Word Acquisition from Speech using Pattern DiscoveryabstractIn this paper, we present an unsupervised method for automatically discovering words from speech using a combination of acoustic pattern discovery, graph clustering, and baseform searching. The algorithm we propose represents an alternative to traditional methods of speech recognition and makes use of the acoustic similarity of multiple realizations of the same words or phrases. On a set of three academic lectures on different subjects, we show that the clustering component of the algorithm is able to successfully generate word clusters that have good coverage of subject-relevant words. Moreover, we illustrate how to use the cluster nodes to retrieve the word identity of each cluster from a large baseform dictionary. Results indicate that this algorithm may prove useful for applications such as vocabulary initialization, speech summarization, or augmentation of existing recognition systems Alex Park 0001, James R. Glass |
ICASSP (1) | 1 |
| 2006 | A Novel DTW-Based Distance Measure for speaker SegmentationabstractWe present a novel distance measure for comparing two speech segments that uses a local version of the well-known DTW algorithm. Our approach is based on the idea of finding word-level speech patterns that are repeated by the same speaker. Using this distance measure, we develop a speaker segmentation procedure and apply it to the task of segmenting multi-speaker lectures. We demonstrate that our approach is able to generate segmentations that correlate well to independently generated human segmentations. In experiments performed on over ten hours of multi-speaker lecture data, we were able to find speaker change points with precision and recall rates of 80% and 100%, respectively. Alex Park 0001, James R. Glass |
SLT | 1 |
| 2005 | Automatic Processing of Audio Lectures for Information Retrieval: Vocabulary Selection and Language ModelingabstractThis paper describes our initial progress towards developing a system for automatically transcribing and indexing audio-visual academic lectures for audio information retrieval. We investigate the problem of how to combine generic spoken data sources with subject-specific text sources for processing lecture speech. In addition to word recognition experiments, we perform audio information retrieval simulations to characterize retrieval performance when using errorful automatic transcriptions. Given an appropriately selected vocabulary, we observe that good retrieval performance can be obtained even with high recognition error rates. For language model training, we observe that the addition of spontaneous speech data to subject-specific written material results in more accurate transcriptions, but has a marginal effect on retrieval performance. 1. Alex Park 0001, Timothy J. Hazen, James R. Glass |
ICASSP (1) | 1 |
| 2004 | A comparison of normalization and training approaches for ASR-dependent speaker identificationabstractIn this paper we discuss a speaker identification approach, called ASR-dependent speaker identification, that incorporates phonetic knowledge into the models for each speaker. This approach differs from traditional methods for performing textindependent speaker identification, such as global Gaussian mixture modeling, that typically ignore the phonetic content of the speech signal. We introduce a new score normalization approach, called phone adaptive normalization, which improves upon our previous speaker adaptive normalization technique. This paper also examines the use of automatically generated transcriptions during the training of our speaker models. Experiments show that speaker models trained using automatically generated transcriptions achieve the same performance as models trained using manually generated transcriptions. 1. Alex Park 0001, Timothy J. Hazen |
INTERSPEECH | 1 |
| 2003 | Towards robust person recognition on handheld devices using face and speaker identification technologiesabstractMost face and speaker identification techniques are tested on data collected in controlled environments using high quality cameras and microphones. However, the use of these technologies in variable environments and with the help of the inexpensive sound and image capture hardware present in mobile devices presents an additional challenge. In this study, we investigate the application of existing face and speaker identification techniques to a person identification task on a handheld device. These techniques have proven to perform accurately on tightly constrained experiments where the lighting conditions, visual backgrounds, and audio environments are fixed and specifically adjusted for optimal data quality. When these techniques are applied on mobile devices where the visual and audio conditions are highly variable, degradations in performance can be expected. Under these circumstances, the combination of multiple biometric modalities can improve the robustness and accuracy of the person identification task. In this paper, we present our approach for combining face and speaker identification technologies and experimentally demonstrate a fused multi-biometric system which achieves a 50% reduction in equal error rate over the better of the two independent systems. Timothy J. Hazen, Eugene Weinstein, Alex Park 0001 |
ICMI | 3 |
| 2003 | Integration of speaker recognition into conversational spoken dialogue systemsabstractIn this paper we examine the integration of speaker identification/verification technology into two dialogue systems developed at MIT: the Mercury air travel reservation system and the Orion task delegation system. These systems both utilize information collected from registered users that is useful in personalizing the system to specific users and that must be securely protected from imposters. Two speaker recognition systems, the MIT Lincoln Laboratory textindependent GMM based system and the MIT Laboratory for Computer Science text-constrained speaker-adaptive ASRbased system, are evaluated and compared within the context of these conversational systems. 1. Timothy J. Hazen, Douglas A. Jones, Alex Park 0001, Linda C. Kukolich, Douglas A. Reynolds |
INTERSPEECH | 3 |
| 2002 | ASR dependent techniques for speaker identificationabstractTraditional text independent speaker recognition systems are based on Gaussian Mixture Models (GMMs) trained globally over all speech from a given speaker. In this paper, we describe alternative methods for performing speaker identification that utilize domain dependent automatic speech recognition (ASR) to provide a phonetic segmentation of the test utterance. When evaluated on YOHO, several of these approaches were able outperform previously published results on the speaker ID task. On a more difficult conversational speech task, we were able to use a combination of classifiers to reduce identification error rates on single test utterances. Over multiple utterances, the ASR dependent approaches performed significantly better than the ASR independent methods. Using an approach we call speaker adaptive modelling for speaker identification, we were able to reduce speaker identification error rates by 39 % over a baseline GMM approach when observing five test utterances from a speaker. 1. Alex Park 0001, Timothy J. Hazen |
INTERSPEECH | 1 |
| 2001 | FST-based recognition techniques for multi-lingual and multi-domain spontaneous speechabstractIn this paper we present techniques for building multi-domain and multi-lingual recognizers within a finite-state transducer (FST) framework. The flexibility of the FST approach is also demonstrated on the task of incorporating networks modeling different types of non-speech events into an existing word lattice network. The ability to create robust multi-domain and/or multi-lingual recognizers for spontaneous speech will enable a conversational system to switch seamlessly and automatically among different domains and/or languages. Preliminary results using a bi-domain recognizer exhibit only small recognition accuracy degradation in comparison to domain-dependent recognition. Similarly promising results were observed using a bilingual recognizer which performs simultaneous language identification and recognition. When using the FST techniques to add non-speech models to the recognizer, experiments show a 10% reduction in word error rate across all utterances and a 30% reduction on utterances containing non-speech events. 1. Timothy J. Hazen, I. Lee Hetherington, Alex Park 0001 |
INTERSPEECH | 3 |