VLDB 2026 Research / reviewers in the wild / expert
Hao Tang 0002
dblp:07/5751-2
· DBLP profile ↗
46ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0002-2445-2605ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 30 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A framework for analyzing concept representations in neural modelsabstractUnderstanding how neural models represent human-interpretable concepts is challenging.Prior work has explored linear concept subspaces from diverse perspectives, such as probing and concept erasure.We introduce a unified framework to study these subspaces along two axes: containment, which tests if a concept is fully represented in a subspace but not outside it, and disentanglement, which tests for isolation from other concepts.In experiments on both text and speech models, we first highlight that concept subspaces may not be uniquely determined, and discuss the implications for concept subspace analysis.Then, we compare properties of concept subspaces estimated using five estimators, proposed in different communities.We find that (1) the choice of estimator impacts the containment and disentanglement properties; (2) the state-of-theart concept erasure method, LEACE, performs well on both testing axes, but still struggles to generalize to unseen data; and (3) in HuBERT speech representations, phone information is both contained and disentangled from speaker information, while speaker information is hard to contain in a compact subspace, despite being disentangled from phones. 1 1 We release the source code at https://github.com/ burin-n/concept_space. Burin Naowarat, Hao Tang 0002, Sharon Goldwater |
CoNLL | 2 |
| 2025 | Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech TransformersabstractTransformer-based self-supervised models have achieved remarkable success in speech processing, but their large size and high inference cost present significant challenges for real-world deployment. While numerous compression techniques have been proposed, inconsistent evaluation metrics make it difficult to compare their practical effectiveness. In this work, we conduct a comprehensive study of four common compression methods, including weight pruning, head pruning, low-rank approximation, and knowledge distillation on self-supervised speech Transformers. We evaluate each method under three key metrics: parameter count, multiply-accumulate operations, and real-time factor. Results show that each method offers distinct advantages. In addition, we contextualize recent compression techniques, comparing DistilHuBERT, FitHuBERT, LightHuBERT, ARMHuBERT, and STaRHuBERT under the same framework, offering practical guidance on compression for deployment. Tzu-Quan Lin, Tsung-Huan Yang, Chun-Yao Chang, Kuang-Ming Chen, Tzu-hsun Feng, Hung-yi Lee, Hao Tang 0002 |
ASRU | 7 |
| 2025 | Whisper Has an Internal Word AlignerabstractThere is an increasing interest in obtaining accurate word-level timestamps from strong automatic speech recognizers, in particular Whisper. Existing approaches either require additional training or are simply not competitive. The evaluation in prior work is also relatively loose, typically using a tolerance of more than 200 ms. In this work, we discover attention heads in Whisper that capture accurate word alignments and are distinctively different from those that do not. Moreover, we find that using characters produces finer and more accurate alignments than using wordpieces. Based on these findings, we propose an unsupervised approach to extracting word alignments by filtering attention heads while teacher forcing Whisper with characters. Our approach not only does not require training but also produces word alignments that are more accurate than prior work under a stricter tolerance between 20 ms and $100 \mathrm{~ms}$.11The source code is available at https://github.com/30stomercury/whisper-char-alignment Sung-Lin Yeh, Yen Meng, Hao Tang 0002 |
ASRU | 3 |
| 2025 | Effective Context in Neural Speech ModelsabstractModern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted to measure how much context these models actually use, i.e., the effective context. Here, we propose two approaches to measuring the effective context, and use them to analyze different speech Transformers. For supervised models, we find that the effective context correlates well with the nature of the task, with fundamental frequency tracking, phone classification, and word classification requiring increasing amounts of effective context. For self-supervised models, we find that effective context increases mainly in the early layers, and remains relatively short---similar to the supervised phone model. Given that these models do not use a long context during prediction, we show that HuBERT can be run in streaming mode without modification to the architecture and without further fine-tuning. Yen Meng, Sharon Goldwater, Hao Tang 0002 |
INTERSPEECH | 3 |
| 2024 | A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
Oli Danyi Liu, Hao Tang 0002, Naomi Feldman, Sharon Goldwater |
CogSci | 2 |
| 2024 | DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
Tzu-Quan Lin, Hung-yi Lee, Hao Tang 0002 |
INTERSPEECH | 3 |
| 2024 | Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representationsabstractSelf-supervised speech representations can hugely benefit downstream speech technologies, yet the properties that make the museful are still poorly understood. Two candidate properties related to the geometry of the representation space have been hypothesized to correlate well with downstream tasks: (1) the degree of orthogonality between the subspaces spanned by the speaker centroids and phone centroids, and (2) the isotropy of the space, i.e., the degree to which all dimensions are effectively utilized. To study them, we introduce a new measure, Cumulative Residual Variance (CRV), which can be used to assess both properties. Using linear classifiers for speaker and phone ID to probe the representations of six different self-supervised models and two untrained baselines, we ask whether either orthogonality or isotropy correlate with linear probing accuracy. We find that both measures correlate with phonetic probing accuracy, though our results on isotropy are more nuanced. Mukhtar Mohamed, Oli Danyi Liu, Hao Tang 0002, Sharon Goldwater |
INTERSPEECH | 3 |
| 2024 | Property Neurons in Self-Supervised Speech TransformersabstractThere have been many studies on analyzing self-supervised speech Transformers, in particular, with layer-wise analysis. It is, however, desirable to have an approach that can pinpoint exactly a subset of neurons that is responsible for a particular property of speech, being amenable to model pruning and model editing. In this work, we identify a set of property neurons in the feedforward layers of Transformers to study how speech-related properties, such as phones, gender, and pitch, are stored. When removing neurons of a particular property (a simple form of model editing), the respective downstream performance significantly degrades, showing the importance of the property neurons. We apply this approach to pruning the feedforward layers in Transformers, where most of the model parameters are. We show that protecting property neurons during pruning is significantly more effective than normbased pruning. Tzu-Quan Lin, Guan-Ting Lin, Hung-yi Lee, Hao Tang 0002 |
SLT | 4 |
| 2024 | A Simple HMM with Self-Supervised Representations for Phone SegmentationabstractDespite the recent advance in self-supervised representations, unsupervised phonetic segmentation remains challenging. Most approaches focus on improving phonetic representations with self-supervised learning, with the hope that the improvement can transfer to phonetic segmentation. In this paper, contrary to recent approaches, we show that peak detection on Mel spectrograms is a strong baseline, better than many self-supervised approaches. Based on this finding, we propose a simple hidden Markov model that uses self-supervised representations and features at the boundaries for phone segmentation. Our results demonstrate consistent improvements over previous approaches, with a generalized formulation allowing versatile design adaptations. Gene-Ping Yang, Hao Tang 0002 |
SLT | 2 |
| 2024 | Estimating the Completeness of Discrete Speech UnitsabstractRepresenting speech with discrete units has been widely used in speech codec and speech generation. However, there are several unverified claims about self-supervised discrete units, such as disentangling phonetic and speaker information with k-means, or assuming information loss after k-means. In this work, we take an information-theoretic perspective to answer how much information is present (information completeness) and how much information is accessible (information accessibility), before and after residual vector quantization. We show a lower bound for information completeness and estimate completeness on discretized HuBERT representations after residual vector quantization. We find that speaker information is sufficiently present in HuBERT discrete units, and that phonetic information is sufficiently present in the residual, showing that vector quantization does not achieve disentanglement. Our results offer a comprehensive assessment on the choice of discrete units, and suggest that a lot more information in the residual should be mined rather than discarded. Sung-Lin Yeh, Hao Tang 0002 |
SLT | 2 |
| 2023 | MelHuBERT: A Simplified Hubert on Mel SpectrogramsabstractSelf-supervised models have had great success in learning speech representations that can generalize to various downstream tasks. However, most self-supervised models require a large amount of compute and multiple GPUs to train, significantly hampering the development of self-supervised learning. In an attempt to reduce the computation of training, we revisit the training of HuBERT, a highly successful self-supervised model. We improve and simplify several key components, including the loss function, input representation, and training in multiple stages. Our model, MelHuBERT, is able to achieve favorable performance on phone recognition, speaker identification, and automatic speech recognition against HuBERT, while saving 31.2% of the pre-training time, or equivalently 33.5% MACs per one second speech. The code and pretrained models are available in https://github.com/nervjack2/MelHuBERT. Tzu-Quan Lin, Hung-yi Lee, Hao Tang 0002 |
ASRU | 3 |
| 2023 | Towards Matching Phones and Speech RepresentationsabstractLearning phone types from phone instances has been a long-standing problem, while still being open. In this work, we revisit this problem in the context of self-supervised learning, and pose it as the problem of matching cluster centroids to phone embeddings. We study two key properties that enable matching, namely, whether cluster centroids of self-supervised representations reduce the variability of phone instances and respect the relationship among phones. We then use the matching result to produce pseudo-labels and introduce a new loss function for improving self-supervised representations. Our experiments show that the matching result captures the relationship among phones. Training the new loss function jointly with the regular self-supervised losses, such as APC and CPC, significantly improves the downstream phone classification. Gene-Ping Yang, Hao Tang 0002 |
ASRU | 2 |
| 2023 | Analyzing Acoustic Word Embeddings from Pre-Trained Self-Supervised Speech ModelsabstractGiven the strong results of self-supervised models on various tasks, there have been surprisingly few studies exploring self-supervised representations for acoustic word embeddings (AWE), fixed-dimensional vectors representing variable-length spoken word segments. In this work, we study several pre-trained models and pooling methods for constructing AWEs with self-supervised representations. Owing to the contextualized nature of self-supervised representations, we hy-pothesize that simple pooling methods, such as averaging, might already be useful for constructing AWEs. When evaluating on a standard word discrimination task, we find that HuBERT representations with mean-pooling rival the state of the art on English AWEs. More surprisingly, despite being trained only on English, HuBERT representations evaluated on Xitsonga, Mandarin, and French consistently outperform the multilingual model XLSR-53 (as well as Wav2Vec 2.0 trained on English). Ramon Sanabria, Hao Tang 0002, Sharon Goldwater |
ICASSP | 2 |
| 2023 | Learning Dependencies of Discrete Speech Representations with Neural Hidden Markov ModelsabstractWhile discrete latent variable models have had great success in self-supervised learning, most models assume that frames are independent. Due to the segmental nature of phonemes in speech perception, modeling dependencies among latent variables at the frame level can potentially improve the learned representations on phonetic-related tasks. In this work, we assume Markovian dependencies among latent variables, and propose to learn speech representations with neural hidden Markov models. Our general framework allows us to compare to self-supervised models that assume independence, while keeping the number of parameters fixed. The added dependencies improve the accessibility of phonetic information, phonetic segmentation, and the cluster purity of phones, showcasing the benefit of the assumed dependencies. Sung-Lin Yeh, Hao Tang 0002 |
ICASSP | 2 |
| 2023 | Conditioning and Sampling in Variational Diffusion Models for Speech Super-ResolutionabstractRecently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resolution audio. In this paper, we propose a novel sampling algorithm that communicates the information of the low-resolution audio via the reverse sampling process of DMs. The proposed method can be a drop-in replacement for the vanilla sampling process and can significantly improve the performance of the existing works. Moreover, by coupling the proposed sampling method with an unconditional DM, i.e., a DM with no auxiliary inputs to its noise predictor, we can generalize it to a wide range of SR setups. We also attain state-of-the-art results on the VCTK Multi-Speaker benchmark with this novel formulation. Chin-Yun Yu, Sung-Lin Yeh, György Fazekas, Hao Tang 0002 |
ICASSP | 4 |
| 2023 | Self-supervised Predictive Coding Models Encode Speaker and Phonetic Information in Orthogonal SubspacesabstractSelf-supervised speech representations are known to encode both speaker and phonetic information, but how they are distributed in the high-dimensional space remains largely unexplored. We hypothesize that they are encoded in orthogonal subspaces, a property that lends itself to simple disentanglement. Applying principal component analysis to representations of two predictive coding models, we identify two subspaces that capture speaker and phonetic variances, and confirm that they are nearly orthogonal. Based on this property, we propose a new speaker normalization method which collapses the subspace that encodes speaker information, without requiring transcriptions. Probing experiments show that our method effectively eliminates speaker information and outperforms a previous baseline in phone discrimination tasks. Moreover, the approach generalizes and can be used to remove information of unseen speakers. Oli Danyi Liu, Hao Tang 0002, Sharon Goldwater |
INTERSPEECH | 2 |
| 2023 | Acoustic Word Embeddings for Untranscribed Target Languages with Continued Pretraining and Learned PoolingabstractAcoustic word embeddings are typically created by training a pooling function using pairs of word-like units. For unsupervised systems, these are mined using k-nearest neighbor (KNN) search, which is slow. Recently, mean-pooled representations from a pre-trained self-supervised English model were suggested as a promising alternative, but their performance on target languages was not fully competitive. Here, we explore improvements to both approaches: we use continued pre-training to adapt the self-supervised model to the target language, and we use a multilingual phone recognizer (MPR) to mine phone n-gram pairs for training the pooling function. Evaluating on four languages, we show that both methods outperform a recent approach on word discrimination. Moreover, the MPR method is orders of magnitude faster than KNN, and is highly data efficient. We also show a small improvement from performing learned pooling on top of the continued pre-trained representations. Ramon Sanabria, Ondrej Klejch, Hao Tang 0002, Sharon Goldwater |
INTERSPEECH | 3 |
| 2023 | Improving Seq2Seq TTS Frontends With Transcribed Speech AudioabstractDue to the data inefficiency and low speech quality of grapheme-based end-to-end text-to-speech (TTS), having a separate high-performance TTS linguistic frontend is still commonly regarded as necessary. However, a TTS frontend is itself difficult to build and maintain, since it requires abundant linguistic knowledge for its construction. In this paper, we start by bootstrapping an integrated sequence-to-sequence (Seq2Seq) TTS frontend using a pre-existing pipeline-based frontend and large amounts of unlabelled normalized text, achieving promising memorization and generalisation abilities. To overcome the performance limitation imposed by the pipeline-based frontend, this work proposes a Forced Alignment (FA) method to decode the pronunciations from transcribed speech audio and then use them to update the Seq2Seq frontend. Our experiments demonstrate the effectiveness of our proposed FA method, which can significantly improve the word token accuracy from 52.6% to 91.2% for out-of-dictionary words. In addition, it can also correct the pronunciation of homographs from transcribed speech audio and potentially improve the homograph disambiguation performance of the Seq2Seq frontend. Korin Richmond, Hao Tang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Supervised Attention in Sequence-to-Sequence Models for Speech RecognitionabstractAttention mechanism in sequence-to-sequence models is designed to model the alignments between acoustic features and output tokens in speech recognition. However, attention weights produced by models trained end to end do not always correspond well with actual alignments, and several studies have further argued that attention weights might not even correspond well with the relevance attribution of frames. Regardless, visual similarity between attention weights and alignments is widely used during training as an indicator of the models quality. In this paper, we treat the correspondence between attention weights and alignments as a learning problem by imposing a supervised attention loss. Experiments have shown significant improved performance, suggesting that learning the alignments well during training critically determines the performance of sequence-to-sequence models. Gene-Ping Yang, Hao Tang 0002 |
ICASSP | 2 |
| 2022 | Phonetic Analysis of Self-supervised Representations of English SpeechabstractWe present an analysis of discrete units discovered via self-supervised representation learning on English speech. We focus on units produced by a pre-trained HuBERT model due to its wide adoption in ASR, speech synthesis, and many other tasks. Whereas previous work has evaluated the quality of such quantization models in aggregate over all phones for a given language, we break our analysis down into broad phonetic classes, taking into account specific aspects of their articulation when considering their alignment to discrete units. We find that these units correspond to sub-phonetic events, and that fine dynamics such as the distinct closure and release portions of plosives tend to be represented by sequences of discrete units. Our work provides a reference for the phonetic properties of discrete units discovered by HuBERT, facilitating analyses of many speech applications based on this model. Dan Wells, Hao Tang 0002, Korin Richmond |
INTERSPEECH | 2 |
| 2022 | Autoregressive Co-Training for Learning Discrete Speech Representation
Sung-Lin Yeh, Hao Tang 0002 |
INTERSPEECH | 2 |
| 2022 | On Compressing Sequences for Self-Supervised Speech ModelsabstractCompressing self-supervised models has become increasingly necessary, as self-supervised models become larger. While previous approaches have primarily focused on compressing the model size, shortening sequences is also effective in reducing the computational cost. In this work, we study fixed-length and variable-length subsampling along the time axis in self-supervised learning. We explore how individual downstream tasks are sensitive to input frame rates. Subsampling while training self-supervised models not only improves the overall performance on downstream tasks under certain frame rates, but also brings significant speed-up in inference. Variable-length subsampling performs particularly well under low frame rates. In addition, if we have access to phonetic boundaries, we find no degradation in performance for an average frame rate as low as 10 Hz. Yen Meng, Hsuan-Jui Chen, Jiatong Shi, Shinji Watanabe 0001, L. Paola García-Perera, Hung-yi Lee, Hao Tang 0002 |
SLT | 7 |
| 2020 | Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHATabstractThis paper proposes a straightforward 2-D method to spatially calibrate the visual field of a camera with the auditory field of an array microphone by generating and overlaying an acoustic image over an optical image. Using a low-cost microphone array and an off-the-shelf camera, we show that polynomial regression can deal efficiently with non-linear camera distortion, and that a recently proposed sound source localization method for real-time processing, SVD-PHAT, can be adapted for this task. François Grondin, Hao Tang 0002, James R. Glass |
ICASSP | 2 |
| 2020 | Vector-Quantized Autoregressive Predictive CodingabstractAutoregressive Predictive Coding (APC), as a self-supervised objective, has enjoyed success in learning representations from large amounts of unlabeled data, and the learned representations are rich for many downstream tasks.However, the connection between low self-supervised loss and strong performance in downstream tasks remains unclear.In this work, we propose Vector-Quantized Autoregressive Predictive Coding (VQ-APC), a novel model that produces quantized representations, allowing us to explicitly control the amount of information encoded in the representations.By studying a sequence of increasingly limited models, we reveal the constituents of the learned representations.In particular, we confirm the presence of information with probing tasks, while showing the absence of information with mutual information, uncovering the model's preference in preserving speech information as its capacity becomes constrained.We find that there exists a point where phonetic and speaker information are amplified to maximize a selfsupervised objective.As a byproduct, the learned codes for a particular model capacity correspond well to English phones. Yu-An Chung, Hao Tang 0002, James R. Glass |
INTERSPEECH | 2 |
| 2019 | An Unsupervised Autoregressive Model for Speech Representation LearningabstractThis paper proposes a novel unsupervised autoregressive neural model for learning generic speech representations.In contrast to other speech representation learning methods that aim to remove noise or speaker variabilities, ours is designed to preserve information for a wide range of downstream tasks.In addition, the proposed model does not require any phonetic or word boundary labels, allowing the model to benefit from large quantities of unlabeled data.Speech representations learned by our model significantly improve performance on both phone classification and speaker verification over the surface features and other supervised and unsupervised approaches.Further analysis shows that different levels of speech information are captured by our model at different layers.In particular, the lower layers tend to be more discriminative for speakers, while the upper layers provide more phonetic content. Yu-An Chung, Wei-Ning Hsu, Hao Tang 0002, James R. Glass |
INTERSPEECH | 3 |
| 2019 | A Deep Residual Network for Large-Scale Acoustic Scene Analysis
Logan Ford, Hao Tang 0002, François Grondin, James R. Glass |
INTERSPEECH | 2 |
| 2019 | VoiceID Loss: Speech Enhancement for Speaker VerificationabstractIn this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly used loss functions for speech enhancement such as the L2 loss, the VoiceID loss is based on the feedback from a speaker verification model to generate a ratio mask. The generated ratio mask is multiplied pointwise with the original spectrogram to filter out unnecessary components for speaker verification. In the experiments, we observed that the enhancement network, after training with the VoiceID loss, is able to ignore a substantial amount of time-frequency bins, such as those dominated by noise, for verification. The resulting model consistently improves the speaker verification system on both clean and noisy conditions. Suwon Shon, Hao Tang 0002, James R. Glass |
INTERSPEECH | 2 |
| 2019 | Time-Contrastive Learning Based Deep Bottleneck Features for Text-Dependent Speaker VerificationabstractThere are a number of studies about extraction of bottleneck (BN) features from deep neural networks (DNNs) trained to discriminate speakers, pass-phrases, and triphone states for improving the performance of text-dependent speaker verification (TD-SV). However, a moderate success has been achieved. A recent study presented a time contrastive learning (TCL) concept to explore the non-stationarity of brain signals for classification of brain states. Speech signals have similar non-stationarity property, and TCL further has the advantage of having no need for labeled data. We therefore present a TCL based BN feature extraction method. The method uniformly partitions each speech utterance in a training dataset into a predefined number of multi-frame segments. Each segment in an utterance corresponds to one class, and class labels are shared across utterances. DNNs are then trained to discriminate all speech frames among the classes to exploit the temporal structure of speech. In addition, we propose a segment-based unsupervised clustering algorithm to re-assign class labels to the segments. TD-SV experiments were conducted on the RedDots challenge database. The TCL-DNNs were trained using speech data of fixed pass-phrases that were excluded from the TD-SV evaluation set, so the learned features can be considered phrase-independent. We compare the performance of the proposed TCL BN feature with those of short-time cepstral features and BN features extracted from DNNs discriminating speakers, pass-phrases, speaker+pass-phrase, as well as monophones whose labels and boundaries are generated by three different automatic speech recognition (ASR) systems. Experimental results show that the proposed TCL-BN outperforms cepstral features and speaker+pass-phrase discriminant BN features, and its performance is on par with those of ASR derived BN features. Moreover, the clustering method improves the TD-SV performance of TCL-BN and ASR derived BN features with respect to their standalone counterparts. We further study the TD-SV performance of fusing cepstral and BN features. Achintya Kumar Sarkar, Zheng-Hua Tan, Hao Tang 0002, Suwon Shon, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Unsupervised Adaptation with Interpretable Disentangled Representations for Distant Conversational Speech RecognitionabstractThe current trend in automatic speech recognition is to leverage large amounts of labeled data to train supervised neural network models. Unfortunately, obtaining data for a wide range of domains to train robust models can be costly. However, it is relatively inexpensive to collect large amounts of unlabeled data from domains that we want the models to generalize to. In this paper, we propose a novel unsupervised adaptation method that learns to synthesize labeled data for the target domain from unlabeled in-domain data and labeled out-of-domain data. We first learn without supervision an interpretable latent representation of speech that encodes linguistic and nuisance factors (e.g., speaker and channel) using different latent variables. To transform a labeled out-of-domain utterance without altering its transcript, we transform the latent nuisance variables while maintaining the linguistic variables. To demonstrate our approach, we focus on a channel mismatch setting, where the domain of interest is distant conversational speech, and labels are only available for close-talking speech. Our proposed method is evaluated on the AMI dataset, outperforming all baselines and bridging the gap between unadapted and in-domain models by over 77% without using any parallel data. Wei-Ning Hsu, Hao Tang 0002, James R. Glass |
INTERSPEECH | 2 |
| 2018 | A Study of Enhancement, Augmentation and Autoencoder Methods for Domain Adaptation in Distant Speech RecognitionabstractSpeech recognizers trained on close-talking speech do not generalize to distant speech and the word error rate degradation can be as large as 40% absolute.Most studies focus on tackling distant speech recognition as a separate problem, leaving little effort to adapting close-talking speech recognizers to distant speech.In this work, we review several approaches from a domain adaptation perspective.These approaches, including speech enhancement, multi-condition training, data augmentation, and autoencoders, all involve a transformation of the data between domains.We conduct experiments on the AMI data set, where these approaches can be realized under the same controlled setting.These approaches lead to different amounts of improvement under their respective assumptions.The purpose of this paper is to quantify and characterize the performance gap between the two domains, setting up the basis for studying adaptation of speech recognizers from close-talking speech to distant speech.Our results also have implications for improving distant speech recognition. Hao Tang 0002, Wei-Ning Hsu, François Grondin, James R. Glass |
INTERSPEECH | 1 |
| 2018 | Frame-Level Speaker Embeddings for Text-Independent Speaker Recognition and Analysis of End-to-End ModelabstractIn this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently with linear activation in the embedding layer. To understand how the speaker recognition model operates with text-independent input, we modify the structure to extract frame-level speaker embeddings from each hidden layer. We feed utterances from the TIMIT dataset to the trained network and use several proxy tasks to study the networks ability to represent speech input and differentiate voice identity. We found that the networks are better at discriminating broad phonetic classes than individual phonemes. In particular, frame-level embeddings that belong to the same phonetic classes are similar (based on cosine distance) for the same speaker. The frame level representation also allows us to analyze the networks at the frame level, and has the potential for other analyses to improve speaker recognition. Suwon Shon, Hao Tang 0002, James R. Glass |
SLT | 2 |
| 2018 | On Training Recurrent Networks with Truncated Backpropagation Through time in Speech RecognitionabstractRecurrent neural networks have been the dominant models for many speech and language processing tasks. However, we understand little about the behavior and the class of functions recurrent networks can realize. Moreover, the heuristics used during training complicate the analyses. In this paper, we study recurrent networks' ability to learn long-term dependency in the context of speech recognition. We consider two decoding approaches, online and batch decoding, and show the classes of functions to which the decoding approaches correspond. We then draw a connection between batch decoding and a popular training approach for recurrent networks, truncated backpropagation through time. Changing the decoding approach restricts the amount of past history recurrent networks can use for prediction, allowing us to analyze their ability to remember. Empirically, we utilize long-term dependency in subphonetic states, phonemes, and words, and show how the design decisions, such as the decoding approach, lookahead, context frames, and consecutive prediction, characterize the behavior of recurrent networks. Finally, we draw a connection between Markov processes and vanishing gradients. These results have implications for studying the long-term dependency in speech data and how these properties are learned by recurrent networks. Hao Tang 0002, James R. Glass |
SLT | 1 |
| 2017 | Multitask Learning with Low-Level Auxiliary Tasks for Encoder-Decoder Based Speech RecognitionabstractEnd-to-end training of deep learning-based models allows for implicit learning of intermediate representations based on the final task loss. However, the end-to-end approach ignores the useful domain knowledge encoded in explicit intermediate-level supervision. We hypothesize that using intermediate representations as auxiliary supervision at lower levels of deep networks may be a good way of combining the advantages of end-to-end training and more traditional pipeline approaches. We present experiments on conversational speech recognition where we use lower-level tasks, such as phoneme recognition, in a multitask training approach with an encoder-decoder model for direct character transcription. We compare multiple types of lower-level tasks and analyze the effects of the auxiliary tasks. Our results on the Switchboard corpus show that this approach improves recognition accuracy over a standard encoder-decoder model on the Eval2000 test set. Shubham Toshniwal, Hao Tang 0002, Liang Lu 0001, Karen Livescu |
INTERSPEECH | 2 |
| 2017 | Lexicon-free fingerspelling recognition from video: Data, models, and signer adaptation
Taehwan Kim 0003, Jonathan Keane, Hao Tang 0002, Jason Riggle, Gregory Shakhnarovich, Diane Brentari, Karen Livescu |
Comput. Speech Lang. | 4 |
| 2017 | ASR for Under-Resourced Languages From Probabilistic TranscriptionabstractIn many under-resourced languages it is possible to find text, and it is possible to find speech, but transcribed speech suitable for training automatic speech recognition (ASR) is unavailable. In the absence of native transcripts, this paper proposes the use of a probabilistic transcript: A probability mass function over possible phonetic transcripts of the waveform. Three sources of probabilistic transcripts are demonstrated. First, self-training is a well-established semisupervised learning technique, in which a cross-lingual ASR first labels unlabeled speech, and is then adapted using the same labels. Second, mismatched crowdsourcing is a recent technique in which nonspeakers of the language are asked to write what they hear, and their nonsense transcripts are decoded using noisy channel models of second-language speech perception. Third, EEG distribution coding is a new technique in which nonspeakers of the language listen to it, and their electrocortical response signals are interpreted to indicate probabilities. ASR was trained in four languages without native transcripts. Adaptation using mismatched crowdsourcing significantly outperformed self-training, and both significantly outperformed a cross-lingual baseline. Both EEG distribution coding and text-derived phone language models were shown to improve the quality of probabilistic transcripts derived from mismatched crowdsourcing. Mark Hasegawa-Johnson, Preethi Jyothi, Daniel McCloy, Majid Mirbagheri, Giovanni M. Di Liberto, Amit Das 0007, Bradley Ekin, Chunxi Liu, Vimal Manohar, Hao Tang 0002, Edmund C. Lalor, Nancy F. Chen, Paul Hager, Tyler Kekona, Rose Sloan, Adrian K. C. Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 10 |
| 2016 | Signer-independent fingerspelling recognition with deep neural network adaptationabstractWe study the problem of recognition of fingerspelled letter sequences in American Sign Language in a signer-independent setting. Fingerspelled sequences are both challenging and important to recognize, as they are used for many content words such as proper nouns and technical terms. Previous work has shown that it is possible to achieve almost 90% accuracies on fingerspelling recognition in a signer-dependent setting. However, the more realistic signer-independent setting presents challenges due to significant variations among signers, coupled with the dearth of available training data. We investigate this problem with approaches inspired by automatic speech recognition. We start with the best-performing approaches from prior work, based on tandem models and segmental conditional random fields (SCRFs), with features based on deep neural network (DNN) classifiers of letters and phonological features. Using DNN adaptation, we find that it is possible to bridge a large part of the gap between signer-dependent and signer-independent performance. Using only about 115 transcribed words for adaptation from the target signer, we obtain letter accuracies of up to 82.7% with framelevel adaptation labels and 69.7% with only word labels. Taehwan Kim 0003, Hao Tang 0002, Karen Livescu |
ICASSP | 3 |
| 2016 | Adapting ASR for under-resourced languages using mismatched transcriptionsabstractMismatched transcriptions of speech in a target language refers to transcriptions provided by people unfamiliar with the language, using English letter sequences. In this work, we demonstrate the value of such transcriptions in building an ASR system for the target language. For different languages, we use less than an hour of mismatched transcriptions to successfully adapt baseline multilingual models built with no access to native transcriptions in the target language. The adapted models provide up to 25% relative improvement in phone error rates on an unseen evaluation set. Chunxi Liu, Preethi Jyothi, Hao Tang 0002, Vimal Manohar, Rose Sloan, Tyler Kekona, Mark Hasegawa-Johnson, Sanjeev Khudanpur |
ICASSP | 3 |
| 2016 | Efficient Segmental Cascades for Speech RecognitionabstractDiscriminative segmental models offer a way to incorporate flexible feature functions into speech recognition.However, their appeal has been limited by their computational requirements, due to the large number of possible segments to consider.Multi-pass cascades of segmental models introduce features of increasing complexity in different passes, where in each pass a segmental model rescores lattices produced by a previous (simpler) segmental model.In this paper, we explore several ways of making segmental cascades efficient and practical: reducing the feature set in the first pass, frame subsampling, and various pruning approaches.In experiments on phonetic recognition, we find that with a combination of such techniques, it is possible to maintain competitive performance while greatly reducing decoding, pruning, and training time. Hao Tang 0002, Kevin Gimpel, Karen Livescu |
INTERSPEECH | 1 |
| 2016 | Triphone State-Tying via Deep Canonical Correlation Analysis
Hao Tang 0002, Karen Livescu |
INTERSPEECH | 2 |
| 2016 | End-to-end training approaches for discriminative segmental modelsabstractRecent work on discriminative segmental models has shown that they can achieve competitive speech recognition performance, using features based on deep neural frame classifiers. However, segmental models can be more challenging to train than standard frame-based approaches. While some segmental models have been successfully trained end to end, there is a lack of understanding of their training under different settings and with different losses. Hao Tang 0002, Kevin Gimpel, Karen Livescu |
SLT | 1 |
| 2015 | Discriminative segmental cascades for feature-rich phone recognitionabstractDiscriminative segmental models, such as segmental conditional random fields (SCRFs) and segmental structured support vector machines (SSVMs), have had success in speech recognition via both lattice rescoring and first-pass decoding. However, such models suffer from slow decoding, hampering the use of computationally expensive features, such as segment neural networks or other high-order features. A typical solution is to use approximate decoding, either by beam pruning in a single pass or by beam pruning to generate a lattice followed by a second pass. In this work, we study discriminative segmental models trained with a hinge loss (i.e., segmental structured SVMs). We show that beam search is not suitable for learning rescoring models in this approach, though it gives good approximate decoding performance when the model is already well-trained. Instead, we consider an approach inspired by structured prediction cascades, which use max-marginal pruning to generate lattices. We obtain a high-accuracy phonetic recognition system with several expensive feature types: a segment neural network, a second-order language model, and second-order phone boundary features. Hao Tang 0002, Kevin Gimpel, Karen Livescu |
ASRU | 1 |
| 2014 | Log-linear dialog managerabstractWe design a log-linear probabilistic model for solving the dialog management task. In both planning and learning we optimize the same objective function: the expected reward. Rather than performing full policy optimization, we perform on-line estimation of the optimal action as a belief-propagation inference step. We employ context-free grammars to describe our variable spaces, which enables us to define rich features. To scale our approach to large variable spaces, we use particle belief propagation. Experiments show that the model is able to choose system actions that yield a high expected reward, outperforming its POMDP-like log-linear counterpart and a hand-crafted rule-based system. Hao Tang 0002, Shinji Watanabe 0001, Tim K. Marks, John R. Hershey |
ICASSP | 1 |
| 2014 | A comparison of training approaches for discriminative segmental modelsabstractSegmental models such as segmental conditional random fields have had some recent success in lattice rescoring for speech recognition. They provide a flexible framework for incorpo-rating a wide range of features across different levels of units, such as phones and words. However, such models have mainly been trained by maximizing conditional likelihood, which may not be the best proxy for the task loss of speech recognition. In addition, there has been little work on designing cost func-tions as surrogates for the word error rate. In this paper, we investigate various losses and introduce a new cost function for training segmental models. We compare lattice rescoring results for multiple tasks and also study the impact of several choices required when optimizing these losses. Index Terms: speech recognition, segmental conditional ran-dom fields, empirical Bayes risk, large-margin training Hao Tang 0002, Kevin Gimpel, Karen Livescu |
INTERSPEECH | 1 |
| 2012 | Discriminative Pronunciation Modeling: A Large-Margin, Feature-Rich Approach
Hao Tang 0002, Joseph Keshet, Karen Livescu |
ACL (1) | 1 |
| 2010 | An initial attempt for phoneme recognition using Structured Support Vector Machine (SVM)abstractStructured Support Vector Machine (SVM) is a recently developed extension of the very successful SVM approach, which can efficiently classify structured pattern with maximized margin. This paper presents an initial attempt for phoneme recognition using structured SVM. We simply learn the basic framework of HMMs in configuring the structured SVM. In the preliminary experiments with TIMIT corpus, the proposed approach was able to offer an absolute performance improvement of 1.33% over HMMs even with a highly simplified initial approach, probably because of the concept of maximized margin of SVM. We see the potential of this approach because of the high generality, high flexibility, and high power of structured SVM. Hao Tang 0002, Chao-Hong Meng, Lin-Shan Lee |
ICASSP | 1 |
| 2009 | Spoken term detection from bilingual spontaneous speech using code-switched lattice-based structures for words and subword unitsabstractThis paper presents the first work known publicly on spoken term detection from bilingual spontaneous speech using code-switched lattice-based structures for word and subword units. The corpus used is the lectures with Chinese as the host language and English as the guest language recorded for a real course offered in National Taiwan University. The techniques reported here have been successfully implemented and tested in a real lecture system now available on-line over the Internet. We also present the approaches of using word fragment as the subword unit for English, and analyse the difficult issues when code-switched lattice-based structures for subword units are used for tasks involving languages of quite different natures. Hung-yi Lee, Yueh-Lien Tang, Hao Tang 0002, Lin-Shan Lee |
ASRU | 3 |