Chao Wang 0018

dblp:188/7759-18 · DBLP profile ↗
← Back
53ranked-venue papers
12as first author
12since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 9 first-author · 12 since 2021Artificial intelligence and machine learning · 32 · 10 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 FedRPO: Federated Relaxed Pareto Optimization for Acoustic Event Classification
abstract
Performance and robustness of real-world Acoustic Event Classification (AEC) solutions depend on ability to train on diverse data from wide range of end-point devices and acoustic environments. Federated Learning (FL) provides a framework to leverage annotated and non-annotated AEC data from servers and client devices in a privacy preserving manner. In this work we propose a novel Federated Relaxed Pareto Optimization (FedRPO) method for semi-supervised FL with heterogeneous client data. In contrast to federated averaging class of FL algorithms (fedAvg) that perform unconstrained weighted aggregation across all data sources, FedRPO enables special treatment of data with high quality annotations vs. data with pseudo-labels of unknown, varying qualities. In particular, FedRPO computes the updates to the global model solving a constrained linear program, with explicit Pareto constraints to prevent performance degradation on annotated data, and controlled relaxation of the Pareto constraints on pseudo-labeled data to prevent learning of patterns in conflict with the annotated data. We show FedRPO significantly outperforms FedAvg on Amazon internal de-identified dataset on AEC tasks. On supervised learning, FedRPO improved precision by 32.5% over FedAvg when maintaining recall at 90%. Combined with FixMatch [1] for semi-supervised learning, FedRPO outperformed FedAvg on precision by 50.5% at 90% recall.
Meng Feng, Chieh-Chi Kao, Qingming Tang, Amit Solomon, Viktor Rozgic, Chao Wang 0018
ICASSP6
2023 Weight-Sharing Supernet for Searching Specialized Acoustic Event Classification Networks Across Device Constraints
abstract
Acoustic Event Classification (AEC) has been widely used in devices such as smart speakers and mobile phones for home safety or accessibility support [1]. As AEC models run on more and more devices with diverse computation resource constraints, it became increasingly expensive to develop models that are tuned to achieve optimal accuracy/computation trade-off for each given computation resource constraint. In this paper, we introduce a Once-For-All (OFA) Neural Architecture Search (NAS) framework for AEC. Specifically, we first train a weight-sharing supernet that supports different model architectures, followed by automatically searching for a model given specific computational resource constraints. Our experimental results showed that by just training once, the resulting model from NAS significantly outperforms both models trained individually from scratch and knowledge distillation (25.4% and 7.3% relative improvement). We also found that the benefit of weight-sharing supernet training of ultra-small models comes not only from searching but from optimization.
Guan-Ting Lin, Qingming Tang, Chieh-Chi Kao, Viktor Rozgic, Chao Wang 0018
ICASSP5
2023 Towards Paralinguistic-Only Speech Representations for End-to-End Speech Emotion Recognition
Georgios Ioannides, Michael Owen, Andrew Fletcher, Viktor Rozgic, Chao Wang 0018
INTERSPEECH5
2022 Federated Self-Supervised Learning for Acoustic Event Classification
abstract
Standard acoustic event classification (AEC) solutions require large-scale collection of data from client devices for model optimization. Federated learning (FL) is a compelling frame- work that decouples data collection and model training to enhance customer privacy. In this work, we investigate the feasibility of applying FL to improve AEC performance while no customer data can be directly uploaded to the server. We assume no pseudo labels can be inferred from on-device user inputs, aligning with the typical use cases of AEC. We adapt self-supervised learning to the FL framework for on-device continual learning of representations, and it results in improved performance of the downstream AEC classifiers with- out labeled/pseudo-labeled data available. Compared to the baseline w/o FL, the proposed method improves precision up to 20.3% relatively while maintaining the recall. Our work differs from prior work in FL in that our approach does not require user-generated learning targets, and the data we use is collected from our Beta program and is de-identified, to maximally simulate the production settings.
Meng Feng, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018
ICASSP7
2022 Sentiment-Aware Automatic Speech Recognition Pre-Training for Enhanced Speech Emotion Recognition
abstract
We propose a novel multi-task pre-training method for Speech Emotion Recognition (SER). We pre-train SER model simultaneously on Automatic Speech Recognition (ASR) and sentiment classification tasks to make the acoustic ASR model more "emotion aware". We generate targets for the sentiment classification using text-to-sentiment model trained on publicly available data. Finally, we fine-tune the acoustic ASR on emotion annotated speech data. We evaluated the proposed approach on MSP-Podcast dataset, where we achieved the best reported concordance correlation coefficient (CCC) of 0.41 for valence prediction.
Ayoub Ghriss, Viktor Rozgic, Elizabeth Shriberg, Chao Wang 0018
ICASSP5
2022 Confidence Estimation for Speech Emotion Recognition Based on the Relationship Between Emotion Categories and Primitives
abstract
Confidence estimation for Speech Emotion Recognition (SER) is instrumental in improving the reliability in the behavior of downstream applications. In this work we propose (1) a novel confidence metric for SER based on the relationship between emotion primitives: arousal, valence, and dominance (AVD) and emotion categories (ECs), (2) EmoConfidNet - a DNN trained alongside the EC recognizer to predict the proposed confidence metric, and (3) a data filtering technique used to enhance the training of EmoConfidNet and the EC recognizer. For each training sample, we calculate distances from corresponding AVD annotation vectors to centroids of each EC in the AVD space, and define EC confidences as functions of the evaluated distances. EmoConfidNet is trained to predict confidence from the same acoustic representations used to train the EC recognizer. EmoConfidNet outperforms state-of-the-art confidence estimation methods on the MSP-Podcast and IEMOCAP datasets. For a fixed EC recognizer, after we reject the same number of low confidence predictions using EmoConfidNet, we achieve a higher F1 and unweighted average recall (UAR) than when rejecting using other methods.
Yang Li 0149, Constantinos Papayiannis, Viktor Rozgic, Elizabeth Shriberg, Chao Wang 0018
ICASSP5
2022 Wikitag: Wikipedia-Based Knowledge Embeddings Towards Improved Acoustic Event Classification
abstract
Acoustic event classification (AEC) is the task of determining whether certain events occur in an audio clip. Inspired by previous research [1], [2], [3] that embeddings from event labels can be leveraged to facilitate the learning of new detectors with no or limited audio samples, we introduce Wikipedia-based text embeddings as auxiliary information to improve AEC. We describe how to extract label embeddings from multiple Wikipedia texts, and formulate the multi-view aligned AEC problem based on VGGish model. We show that our "wikiTAG" embeddings encode rich semantic information and are more informative than label embeddings for AEC tasks. Compared to a supervised baseline on AudioSet, the multi-view model with "wikiTAG" embeddings achieves 7.3% and 1.3% relative improvement in mean average precision (mAP) using 10% and full AudioSet for training, respectively. To the author’s knowledge, this is the first work in the AEC domain on building large-scale label representations by leveraging Wikipedia data in a systematic fashion.
Qingming Tang, Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018
ICASSP6
2022 Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology
abstract
Acoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is not available in audios and plain tags. We show that by organizing audio representations with a human-curated tree ontology, we can improve the quality of the learned audio representations for downstream AEC tasks. We use consistency training to use large amounts of unlabeled data for structured representation manifold learning. Experimental results indicate that our framework learns high quality representations which enable us to achieve comparable performance in discriminative tasks as fully supervised baselines. Moreover, our framework can better handle audios with unseen tags by confidently assigning a super-category (internal node like "animal" in Fig. 1) tag to the audio.
Arman Zharmagambetov, Qingming Tang, Chieh-Chi Kao, Ming Sun 0007, Viktor Rozgic, Jasha Droppo, Chao Wang 0018
ICASSP8
2022 Impact of Acoustic Event Tagging on Scene Classification in a Multi-Task Learning Framework
Rahil Parikh, Harshavardhan Sundar, Ming Sun 0007, Chao Wang 0018, Spyridon Matsoukas
INTERSPEECH4
2021 Unsupervised and Semi-Supervised Few-Shot Acoustic Event Classification
abstract
Few-shot Acoustic Event Classification (AEC) aims to learn a model to recognize novel acoustic events using very limited labeled data. Previous works utilize supervised pre-training as well as meta-learning approaches, which heavily rely on labeled data. Here, we study unsupervised and semi-supervised learning approaches for few-shot AEC. Our work builds upon recent advances in unsupervised representation learning introduced for speech recognition and language modeling. We learn audio representations from a large amount of unlabeled data, and use the resulting representations for few-shot AEC. We further extend our model in a semi-supervised fashion. Our unsupervised representation learning approach outperforms supervised pre-training methods, and our semi-supervised learning approach outperforms meta-learning methods for few-shot AEC. We also show that our work is more robust under domain mismatch.
Hsin-Ping Huang, Krishna C. Puvvada, Ming Sun 0007, Chao Wang 0018
ICASSP4
2021 Contrastive Unsupervised Learning for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is a key technology to enable more natural human-machine communication. However, SER has long suffered from a lack of public large-scale labeled datasets. To circumvent this problem, we investigate how unsupervised representation learning on unlabeled datasets can benefit SER. We show that the contrastive predictive coding (CPC) method can learn salient representations from unlabeled datasets, which improves emotion recognition performance. In our experiments, this method achieved state-of-the-art concordance correlation coefficient (CCC) performance for all emotion primitives (activation, valence, and dominance) on IEMOCAP. Additionally, on the MSP-Podcast dataset, our method obtained considerable performance improvements compared to baselines.
Joshua Levy, Andreas Stolcke, Viktor Rozgic, Spyridon Matsoukas, Constantinos Papayiannis, Daniel Bone, Chao Wang 0018
ICASSP9
2021 Multi-Task Self-Supervised Pre-Training for Music Classification
abstract
Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. Besides, models learned from labeled dataset often embed biases specific to that particular dataset. Therefore, unsupervised learning techniques become popular approaches in solving machine listening problems. Particularly, a self-supervised learning technique utilizing reconstructions of multiple hand-crafted audio features has shown promising results when it is applied to speech domain such as emotion recognition and automatic speech recognition (ASR). In this paper, we apply self-supervised and multi-task learning methods for pre-training music encoders, and explore various design choices including encoder architectures, weighting mechanisms to combine losses from multiple tasks, and worker selections of pretext tasks. We investigate how these design choices interact with various downstream music classification tasks. We find that using various music specific workers altogether with weighting mechanisms to balance the losses during pre-training helps improve and generalize to the downstream tasks.
Ho-Hsiang Wu, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Brian McFee, Juan Pablo Bello, Chao Wang 0018
ICASSP7
2020 A Comparison of Pooling Methods on LSTM Models for Rare Acoustic Event Classification
abstract
Acoustic event classification (AEC) and acoustic event detection (AED) refer to the task of detecting whether specific target events occur in audios. As long short-term memory (LSTM) leads to state-of-the-art results in various speech related tasks, it is employed as a popular solution for AEC as well. This paper focuses on investigating the dynamics of LSTM model on AEC tasks. It includes a detailed analysis on LSTM memory retaining, and a benchmarking of nine different pooling methods on LSTM models using 1.7M generated mixture clips of multiple events with different signal-to-noise ratios. This paper focuses on understanding: 1) utterance-level classification accuracy; 2) sensitivity to event position within an utterance. The analysis is done on the dataset for the detection of rare sound events from DCASE 2017 Challenge. We find max pooling on the prediction level to perform the best among the nine pooling approaches in terms of classification accuracy and insensitivity to event position within an utterance. To authors’ best knowledge, this is the first kind of such work focused on LSTM dynamics for AEC tasks.
Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018
ICASSP4
2020 Few-Shot Acoustic Event Detection Via Meta Learning
abstract
We study few-shot acoustic event detection (AED) in this paper. Few-shot learning enables detection of new events with very limited labeled data. Compared to other research areas like computer vision, few-shot learning for audio recognition has been under-studied. We formulate few-shot AED problem and explore different ways of utilizing traditional supervised methods for this setting as well as a variety of meta-learning approaches, which are conventionally used to solve few-shot classification problem. Compared to supervised baselines, meta-learning models achieve superior performance, thus showing its effectiveness on generalization to new audio events. Our analysis including impact of initialization and domain discrepancy further validate the advantage of meta-learning approaches in few-shot AED.
Bowen Shi 0002, Ming Sun 0007, Krishna C. Puvvada, Chieh-Chi Kao, Spyridon Matsoukas, Chao Wang 0018
ICASSP6
2020 Raw Waveform Based End-to-end Deep Convolutional Network for Spatial Localization of Multiple Acoustic Sources
abstract
In this paper, we present an end-to-end deep convolutional neural network operating on multi-channel raw audio data to localize multiple simultaneously active acoustic sources in space. Previously reported deep learning based approaches work well in localizing a single source directly from multi-channel raw-audio, but are not easily extendable to localize multiple sources due to the well known permutation problem. We propose a novel encoding scheme to represent the spatial coordinates of multiple sources, which facilitates 2D localization of multiple sources in an end-to-end fashion, avoiding the permutation problem and achieving arbitrary spatial resolution. Experiments on a simulated data set and real recordings from the AV16.3 Corpus demonstrate that the proposed method generalizes well to unseen test conditions, and outperforms a recent time difference of arrival (TDOA) based multiple source localization approach reported in the literature.
Harshavardhan Sundar, Ming Sun 0007, Chao Wang 0018
ICASSP4
2020 Intra-Utterance Similarity Preserving Knowledge Distillation for Audio Tagging
abstract
Knowledge Distillation (KD) is a popular area of research for reducing the size of large models while still maintaining good performance.The outputs of larger teacher models are used to guide the training of smaller student models.Given the repetitive nature of acoustic events, we propose to leverage this information to regulate the KD training for Audio Tagging.This novel KD method, Intra-Utterance Similarity Preserving KD (IUSP), shows promising results for the audio tagging task.It is motivated by the previously published KD method: Similarity Preserving KD (SP).However, instead of preserving the pairwise similarities between inputs within a mini-batch, our method preserves the pairwise similarities between the frames of a single input utterance.Our proposed KD method, IUSP, shows consistent improvements over SP across student models of different sizes on the DCASE 2019 Task 5 dataset for audio tagging.There is a 27.1% to 122.4% percent increase in improvement of micro AUPRC over the baseline relative to SPs improvement of over the baseline.
Chun-Chieh Chang, Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018
INTERSPEECH4
2020 Semi-Supervised ASR by End-to-End Self-Training
abstract
While deep learning based end-to-end automatic speech recognition (ASR) systems have greatly simplified modeling pipelines, they suffer from the data sparsity issue.In this work, we propose a self-training method with an end-to-end system for semi-supervised ASR.Starting from a Connectionist Temporal Classification (CTC) system trained on the supervised data, we iteratively generate pseudo-labels on a mini-batch of unsupervised utterances with the current model, and use the pseudo-labels to augment the supervised data for immediate model update.Our method retains the simplicity of end-to-end ASR systems, and can be seen as performing alternating optimization over a well-defined learning objective.We also perform empirical investigations of our method, regarding the effect of data augmentation, decoding beamsize for pseudo-label generation, and freshness of pseudo-labels.On a commonly used semi-supervised ASR setting with the Wall Street Journal (WSJ) corpus, our method gives 14.4% relative WER improvement over a carefully-trained base system with data augmentation, reducing the performance gap between the base system and the oracle system by 46%.
Chao Wang 0018
INTERSPEECH3
2020 A Joint Framework for Audio Tagging and Weakly Supervised Acoustic Event Detection Using DenseNet with Global Average Pooling
abstract
This paper proposes a network architecture mainly designed for audio tagging, which can also be used for weakly supervised acoustic event detection (AED).The proposed network consists of a modified DenseNet as the feature extractor, and a global average pooling (GAP) layer to predict frame-level labels at inference time.This architecture is inspired by the work proposed by Zhou et al., a well-known framework using GAP to localize visual objects given image-level labels.While most of the previous works on weakly supervised AED used recurrent layers with attention-based mechanism to localize acoustic events, the proposed network directly localizes events using the feature map extracted by DenseNet without any recurrent layers.In the audio tagging task of DCASE 2017, our method significantly outperforms the state-of-the-art method in F1 score by 5.3% on the dev set, and 6.0% on the eval set in terms of absolute values.For weakly supervised AED task in DCASE 2018, our model outperforms the state-of-the-art method in event-based F1 by 8.1% on the dev set, and 0.5% on the eval set in terms of absolute values, by using data augmentation and tri-training to leverage unlabeled data.
Chieh-Chi Kao, Bowen Shi 0002, Ming Sun 0007, Chao Wang 0018
INTERSPEECH4
2020 Acoustic Scene Analysis with Multi-Head Attention Networks
abstract
Acoustic Scene Classification (ASC) is a challenging task, as a single scene may involve multiple events that contain complex sound patterns.For example, a cooking scene may contain several sound sources including silverware clinking, chopping, frying, etc.What complicates ASC more is that classes of different activities could have overlapping sounds patterns (e.g. both cooking and dishwashing could have silverware clinking sound).In this paper, we propose a multihead attention network to model the complex temporal input structures for ASC.The proposed network takes the audio's time-frequency representation as input, and it leverages standard VGG plus LSTM layers to extract high-level feature representation.Further more, it applies multiple attention heads to summarize various patterns of sound events into fixed dimensional representation, for the purpose of final scene classification.The whole network is trained in an end-to-end fashion with backpropagation.Experimental results confirm that our model discovers meaningful sound patterns through the attention mechanism, without using explicit supervision in the alignment.We evaluated our proposed model using DCASE 2018 Task 5 dataset, and achieved competitive performance on par with previous winner's results.
Ming Sun 0007, Chao Wang 0018
INTERSPEECH4
2019 Multimodal and Multi-view Models for Emotion Recognition
abstract
Studies on emotion recognition (ER) show that combining lexical and acoustic information results in more robust and accurate models.The majority of the studies focus on settings where both modalities are available in training and evaluation.However, in practice, this is not always the case; getting ASR output may represent a bottleneck in a deployment pipeline due to computational complexity or privacyrelated constraints.To address this challenge, we study the problem of efficiently combining acoustic and lexical modalities during training while still providing a deployable acoustic model that does not require lexical inputs.We first experiment with multimodal models and two attention mechanisms to assess the extent of the benefits that lexical information can provide.Then, we frame the task as a multi-view learning problem to induce semantic information from a multimodal model into our acoustic-only network using a contrastive loss function.Our multimodal model outperforms the previous state of the art on the USC-IEMOCAP dataset reported on lexical and acoustic information.Additionally, our multi-view-trained acoustic network significantly surpasses models that have been exclusively trained with acoustic features.
Gustavo Aguilar, Viktor Rozgic, Chao Wang 0018
ACL (1)4
2019 Deep Embeddings for Rare Audio Event Detection with Imbalanced Data
abstract
In this paper, we present a method to handle data imbalance for classification with neural networks, and apply it to acoustic event detection (AED) problem. The common approach to tackle data imbalance is to use class-weights in the objective function while training. An existing more sophisticated approach is to map the input to clusters in an embedding space, so that learning is locally balanced by incorporating inter-cluster and inter-class margins. On these lines, we propose a method to learn the embedding using a novel objective function, called triple-header cross entropy. Our scheme integrates in a simple way with back-propagation based training, and is computationally more efficient than general hinge-loss based embedding learning schemes. The empirical evaluation results demonstrate the effectiveness of the proposed method for AED with imbalanced training data.
Vipul Arora 0001, Ming Sun 0007, Chao Wang 0018
ICASSP3
2019 Improving Emotion Classification through Variational Inference of Latent Variables
abstract
Conventional models for emotion recognition from speech signal are trained in supervised fashion using speech utterances with emotion labels. In this study we hypothesize that speech signal depends on multiple latent variables including the emotional state, age, gender, and speech content. We propose an Adversarial Autoencoder (AAE) to perform variational inference over the latent variables and reconstruct the input feature representations. Reconstruction of feature representations is used as an auxiliary task to aid the primary emotion recognition task. Experiments on the IEMOCAP dataset demonstrate that the auxiliary learning tasks improve emotion classification accuracy compared to a baseline supervised classifier. Further, we demonstrate that the proposed learning approach can be used for the end-to-end speech emotion recognition, as its applicable for models that operate on frame-level inputs.
Srinivas Parthasarathy, Viktor Rozgic, Ming Sun 0007, Chao Wang 0018
ICASSP4
2019 Semi-supervised Acoustic Event Detection Based on Tri-training
abstract
This paper presents our work of training acoustic event detection (AED) models using unlabeled dataset. Recent acoustic event detectors are based on large-scale neural networks, which are typically trained with huge amounts of labeled data. Labels for acoustic events are expensive to obtain, and relevant acoustic event audios can be limited, especially for rare events. In this paper we leverage an Internet-scale un-labeled dataset with potential domain shift to improve the detection of acoustic events. Based on the classic tri-training approach, our proposed method shows accuracy improvement over both the supervised training baseline, and semi-supervised self-training set-up, in all pre-defined acoustic event detection tasks. As our approach relies on ensemble models, we further show the improvements can be distilled to a single model via knowledge distillation, with the resulting single student model maintaining high accuracy of teacher ensemble models.
Bowen Shi 0002, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018
ICASSP6
2019 Hierarchical Residual-pyramidal Model for Large Context Based Media Presence Detection
abstract
We study media presence detection, that is, learning to recognize if a sound segment (typically lasting for a few seconds) of a long recorded stream contains media (TV) sound. This problem is difficult because non-media sound sources can be quite diverse (e.g. human voicing, non-vocal sounds and non-human sounds), and the recorded sound can be a mixture of media and non-media sound.Different from speech recognition, where the recognizer needs to detect local phonetic variation, the key features used to distinguish media and non-media sounds are non-local features. Motivated by this, we propose a hierarchical model to learn representation of each pre-chunked segment within a long recorded stream jointly, and encourage every local representation to be not sensitive to variations within each segment. We also further explore the effects of techniques including stream based normalization and iteratively imputing missing labels of training dataset. Experimental results indicate that our proposed contextual based methods are effective for media presence detection.
Qingming Tang, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Chao Wang 0018
ICASSP5
2019 Sub-Band Convolutional Neural Networks for Small-Footprint Spoken Term Classification
abstract
This paper proposes a Sub-band Convolutional Neural Network for spoken term classification.Convolutional neural networks (CNNs) have proven to be very effective in acoustic applications such as spoken term classification, keyword spotting, speaker identification, acoustic event detection, etc.Unlike applications in computer vision, the spatial invariance property of 2D convolutional kernels does not fit acoustic applications well since the meaning of a specific 2D kernel varies a lot along the feature axis in an input feature map.We propose a sub-band CNN architecture to apply different convolutional kernels on each feature sub-band, which makes the overall computation more efficient.Experimental results show that the computational efficiency brought by sub-band CNN is more beneficial for smallfootprint models.Compared to a baseline full band CNN for spoken term classification on a publicly available Speech Commands dataset, the proposed sub-band CNN architecture reduces the computation by 39.7% on commands classification, and 49.3% on digits classification with accuracy maintained.
Chieh-Chi Kao, Ming Sun 0007, Shiv Vitaladevuni, Chao Wang 0018
INTERSPEECH5
2019 Compression of Acoustic Event Detection Models with Quantized Distillation
abstract
Acoustic Event Detection (AED), aiming at detecting categories of events based on audio signals, has found application in many intelligent systems. Recently deep neural network significantly advances this field and reduces detection errors to a large scale. However how to efficiently execute deep models in AED has received much less attention. Meanwhile state-of-the-art AED models are based on large deep models, which are computational demanding and challenging to deploy on devices with constrained computational resources. In this paper, we present a simple yet effective compression approach which jointly leverages knowledge distillation and quantization to compress larger network (teacher model) into compact network (student model). Experimental results show proposed technique not only lowers error rate of original compact network by 15% through distillation but also further reduces its model size to a large extent (2% of teacher, 12% of full-precision student) through quantization.
Bowen Shi 0002, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018
INTERSPEECH6
2018 R-CRNN: Region-based Convolutional Recurrent Neural Network for Audio Event Detection
abstract
This paper proposes a Region-based Convolutional Recurrent Neural Network (R-CRNN) for audio event detection (AED).The proposed network is inspired by Faster-RCNN [1], a wellknown region-based convolutional network framework for visual object detection.Different from the original Faster-RCNN, a recurrent layer is added on top of the convolutional network to capture the long-term temporal context from the extracted highlevel features.While most of the previous works on AED generate predictions at frame level first, and then use post-processing to predict the onset/offset timestamps of events from a probability sequence; the proposed method generates predictions at event level directly and can be trained end-to-end with a multitask loss, which optimizes the classification and localization of audio events simultaneously.The proposed method is tested on DCASE 2017 Challenge dataset [2].To the best of our knowledge, R-CRNN is the best performing single-model method among all methods without using ensembles both on development and evaluation sets.Compared to the other region-based network for AED (R-FCN [3]) with an event-based error rate (ER) of 0.18 on the development set, our method reduced the ER to half.
Chieh-Chi Kao, Ming Sun 0007, Chao Wang 0018
INTERSPEECH4
2018 Detecting Media Sound Presence in Acoustic Scenes
Constantinos Papayiannis, Justice Amoh, Viktor Rozgic, Shiva Sundaram, Chao Wang 0018
INTERSPEECH5
2018 A Simple Model for Detection of Rare Sound Events
abstract
We propose a simple recurrent model for detecting rare sound events, when the time boundaries of events are available for training. Our model optimizes the combination of an utterance-level loss, which classifies whether an event occurs in an utterance, and a frame-level loss, which classifies whether each frame corresponds to the event when it does occur. The two losses make use of a shared vectorial representation the event, and are connected by an attention mechanism. We demonstrate our model on Task 2 of the DCASE 2017 challenge, and achieve competitive performance.
Chieh-Chi Kao, Chao Wang 0018
INTERSPEECH3
2011 Your Mobile Virtual Assistant Just Got Smarter!
abstract
A Mobile Virtual Assistant (MVA) is a communication agent that recognizes and understands free speech, and performs actions such as retrieving information and completing transactions. One essential characteristic of MVAs is their ability to learn and adapt without supervision. This paper describes our ongoing research in developing more intelligent MVAs that recognize and understand very large vocabulary speech input across a variety of tasks. In particular, we present our architecture for unsupervised acoustic and language model adaptation. Experimental results show that unsupervised acoustic model learning approaches the performance of supervised learning when adapting on 40-50 device-specific utterances. Unsupervised language model learning results in an 8% absolute drop in word error rate.
Mazin Gilbert, Iker Arizmendi, Enrico Bocchieri, Diamantino Caseiro, Vincent Goffin, Andrej Ljolje, Mike Phillips, Chao Wang 0018, Jay G. Wilpon
INTERSPEECH8
2010 A hybrid architecture for mobile voice user interfaces
Imre Kiss, Joseph Polifroni, Chao Wang 0018, Ghinwa F. Choueiter, Mike Phillips
INTERSPEECH3
2010 Good grief, i can speak it! preliminary experiments in audio restaurant reviews
abstract
In this paper, we introduce a new envisioned application for speech which allows users to enter restaurant reviews orally via their mobile device, and, at a later time, update a shared and growing database of consumer-provided information about restaurants. During the intervening period, a speech recognition and NLP based system has analyzed their audio recording both to extract key descriptive phrases and to compute sentiment ratings based on the evidence provided in the audio clip. We report here on our preliminary work moving towards this goal. Our experiments demonstrate that multi-aspect sentiment ranking works surprisingly well on speech output, even in the presence of recognition errors. We also present initial experiments on integrated sentence boundary detection and key phrase extraction from recognition output.
Joseph Polifroni, Stephanie Seneff, S. R. K. Branavan, Chao Wang 0018, Regina Barzilay
SLT4
2007 A Spoken Translation Game for Second Language Learning
Chao Wang 0018, Stephanie Seneff
AIED1
2007 Chinese Syntactic Reordering for Statistical Machine Translation
Chao Wang 0018, Michael Collins 0001, Philipp Koehn
EMNLP-CoNLL1
2007 Automatic Assessment of Student Translations for Foreign Language Tutoring
Chao Wang 0018, Stephanie Seneff
HLT-NAACL1
2006 Scalable and portable web-based multimodal dialogue interaction with geographical databases
abstract
We describe work towards developing a scalable and portable framework for enabling map-based multimodal dialogue interaction over the web. Working in the context of a restaurant-guide system, we show how large information databases harvested from the web can be accommodated in our speech recognizer, parser, and web-based GUI. We compare two dynamic language modeling techniques, which calculate context-dependent weights for the large sets of proper nouns associated with geographical entities such as restaurants and streets. We show that the more fine-grained approach results in a 7.8% reduction in concept error rate. Index Terms: multimodal dialogue system, language modeling, restaurants, maps, world wide web
Alexander Gruenstein, Stephanie Seneff, Chao Wang 0018
INTERSPEECH3
2006 High-quality speech translation in the flight domain
abstract
Portability is an important issue to the viability of a domainspecific translation approach. This paper describes an English to Chinese translation system for flight-domain queries, utilizing an interlingua translation framework that has been successfully applied in the weather domain. Portability of various components is tested, and new technologies to handle parse ambiguities and illformed inputs are developed to enhance the translation framework. Evaluation of translation quality is conducted manually on a set of 432 unseen flight-domain utterances, which are translated into Chinese using a formal method and a new robust back-off method in tandem. We achieved 96.7 % sentence accuracy with a rejection rate of 7.6 % on manual transcripts, and 89.1 % accuracy with an 8.6 % rejection rate on speech input. A game for language learning using the translation capability is currently under development. Index Terms: speech translation, domain portability, natural language understanding, natural language generation.
Chao Wang 0018, Stephanie Seneff
INTERSPEECH1
2005 Context-sensitive statistical language modeling
abstract
We present context-sensitive dynamic classes ‐ a novel mechanism for integrating contextual information from spoken dialogue into a class n-gram language model. We exploit the dialogue system’s information state to populate dynamic classes, thus percolating contextual constraints to the recognizer’s language model in real time. We describe a technique for training a language model incorporating context-sensitive dynamic classes which considerably reduces word error rate under several conditions. Significantly, our technique does not partition the language model based on potentially artificial dialogue state distinctions; rather, it accommodates both strong and weak expectations via dynamic manipulation of a single model.
Alexander Gruenstein, Chao Wang 0018, Stephanie Seneff
INTERSPEECH2
2005 Language model data filtering via user simulation and dialogue resynthesis
abstract
In this paper, we address the issue of generating language model training data during the initial stages of dialogue system development. The process begins with a large set of sentence templates, automatically adapted from other application domains. We propose two methods to filter the raw data set to achieve a desired probability distribution of the semantic content, both on the sentence level and on the class level. The first method utilizes user simulation technology, which obtains the probability model via an interplay between a probabilistic user model and the dialogue system. The second method synthesizes novel dialogue interactions by modeling after a small set of dialogues produced by the developers during the course of system refinement. We evaluated our methodology by speech recognition performance on a set of 520 unseen utterances from naive users interacting with a restaurant domain dialogue system. 1.
Chao Wang 0018, Stephanie Seneff, Grace Chung
INTERSPEECH1
2005 Statistical modeling of phonological rules through linguistic hierarchies
Stephanie Seneff, Chao Wang 0018
Speech Commun.2
2004 Combining linguistic knowledge and acoustic information in automatic pronunciation lexicon generation
abstract
This paper describes several experiments aimed at the long term goal of enabling a spoken conversational system to automatically improve its pronunciation lexicon over time through direct interactions with end users and from available Web sources. We selected a set of 200 rare words from the OGI corpus of spoken names, and performed several experiments combining spelling and pronunciation information to hypothesize phonemic baseforms for these words. We evaluated the quality of the resulting baseforms through a series of recognition experiments, using the 200 words in an isolated word recognition task. We also report here on a modification to our letter-to-sound system, utilizing a letter-phoneme -gram language model, either alone or in combination with our original “column-bigram” model, for additional linguistic constraint and robustness. Our experiments confirm our expectation that acoustic information drawn from spoken examples of the words can greatly improve the quality of the baseforms, as measured by the recognition error rate. Our ultimate goal is to allow a spoken dialogue system to automatically expand and improve its baseforms over time as users introduce new words or supply spoken pronunciations of existing words.
Grace Chung, Chao Wang 0018, Stephanie Seneff, Edward Filisko, Min Tang 0005
INTERSPEECH2
2004 A dynamic vocabulary spoken dialogue interface
abstract
Mixed-initiative spoken dialogue systems today generally allow users to query with a fixed vocabulary and grammar that is determined prior to run-time. This paper presents a spoken dialogue interface enhanced with a dynamic vocabulary capability. One or more word classes can be made dynamic in the speech recognizer and natural language (NL) grammar so that a context-specific vocabulary subset can be incorporated on-the-fly as the context of the dialogue changes, at each dialogue turn. Described is a restaurant information domain which continually updates the restaurant name class, given the dialogue context. We examine progress made to the speech recognizer, natural language parser and dialogue manager in order to support the dynamic vocabulary capability, and present preliminary experimental results conducted from simulated dialogues.
Stephanie Seneff, Chao Wang 0018, I. Lee Hetherington, Grace Chung
INTERSPEECH2
2004 An interactive English pronunciation dictionary for Korean learners
abstract
We present research towards developing a pronunciation dictionary that features sensitivity to learners’ native phonology, specifically designed for Korean learners of English-as-a-Foreign-Language (EFL). We envision a future system that can record and process learners’ imitationof thedictionary pronunciation and instantly provide segmental and prosodic feedback on accent. Towards this goal, we have designed and collected a speech corpus to address the phonological and prosodic issues of Korean EFL learners. We leverage the SUMMIT speech recognizer’s ability to model phonological rules to automatically identify non-native phonological phenomena. These phonological rules were carefully constructed to account for the influence of learners’ native language (Korean) on the target language (English). Feedback messages are provided to the learner to point out the non-native phonological variations detected by the speech recognizer in order to help them improve segmental pronunciation. Instructions are also given to the user on the prosodic aspects of the pronunciation, which are based on detected duration and cues. We evaluated the effectiveness of the feedback mechanism by rating 222 English utterances from six native Korean subjects, before and after receiving native-language dependent feedback messages. Human raters judged 61% of the utterances as improved after feedback.
Chao Wang 0018, Mitchell Peabody, Stephanie Seneff, Jong-mi Kim
INTERSPEECH1
2003 Empowering end users to personalize dialogue systems through spoken interaction
abstract
This paper describes recent advances we have made towards the goal of empowering end users to automatically expand the knowledge base of a dialogue system through spoken interaction, in order to personalize it to their individual needs. We describe techniques used to incrementally reconfigure a preloaded trained natural language grammar, as well as the lexicon and language models for the speech recognition system. We also report on advances in the technology to integrate a spoken pronunciation with a spoken spelling, in order to improve spelling accuracy. While the original algorithm was designed for a “speak and spell ” input mode, we have shown here that the same methods can be applied to separately uttered spoken and spelled forms of the word. By concatenating the two waveforms, we can take advantage of the mutual constraints realized in an integrated composite FST. Using an OGI corpus of separately spoken and spelled names, we have demonstrated letter error rates of under 6 % for in-vocabulary words and under 11 % for words not contained in the training lexicon, a 44 % reduction in error rate over that achieved without use of the spoken form. We anticipate applying this technique to unknown words embedded in a larger context, followed by solicited spellings. 1.
Stephanie Seneff, Grace Chung, Chao Wang 0018
INTERSPEECH3
2003 Automatic induction of n-gram language models from a natural language grammar
abstract
This paper details our work in developing a technique which can automatically generate class n-gram language models from natural language (NL) grammars in dialogue systems. The procedure eliminates the need for double maintenance of the recognizer language model and NL grammar. The resulting language model adopts the standard class n-gram framework for computational efficiency. Moreover, both the n-gram classes and training sentences are enhanced with semantic/syntactic tags defined in the NL grammar, such that the trained language model preserves the distinctive statistics associated with different word senses. We have applied this approach in several different domains and languages, and have evaluated it on our most mature dialogue systems to assess its competitiveness with preexisting n-gram language models. The speech recognition performances with the new language model are in fact the best we have achieved in both the JUPITER weather domain and the MERCURY flight reservation domain. 1.
Stephanie Seneff, Chao Wang 0018, Timothy J. Hazen
INTERSPEECH2
2003 Automatic Acquisition of Names Using Speak and Spell Mode in Spoken Dialogue Systems
Grace Chung, Stephanie Seneff, Chao Wang 0018
HLT-NAACL3
2001 Voice transformations: from speech synthesis to mammalian vocalizations
abstract
This paper describes a phase vocoder based technique for voice transformation. This method provides a flexible way to manipulate various aspects of the input signal, e.g., fundamental frequency of voicing, duration, energy, and formant positions, without explicit extraction. The modifications to the signal can be specific to any feature dimensions, and can vary dynamically over time. There are many potential applications for this technique. In concatenative speech synthesis, the method can be applied to transform the speech corpus to different voice characteristics, or to smooth any pitch or formant discontinuities between concatenation boundaries. The method can also be used as a tool for language learning. We can modify the prosody of the student 's own speech to match that from a native speaker, and use the result as guidance for improvements. The technique can also be used to convert other biological signals, such as killer whale vocalizations, to a signal that is more appropriate for human auditory perception. Our initial experiments show encouraging results for all of these applications. 1.
Min Tang 0005, Chao Wang 0018, Stephanie Seneff
INTERSPEECH2
2001 Lexical stress modeling for improved speech recognition of spontaneous telephone speech in the jupiter domain
abstract
This paper examines an approach of using lexical stress models to improve the speech recognition performance on spontaneous telephone speech. We analyzed the correlation of various pitch, energy, and duration measurements with lexical stress on a large corpus of spontaneous utterances, and identified the most informative features of stress using classification experiments. We incorporated the stress models into the recognizer first-pass Viterbi search and obtained modest but statistically significant improvements over a state-of-the-art real-time performance on the JUPITER domain. 1.
Chao Wang 0018, Stephanie Seneff
INTERSPEECH1
2000 Robust pitch tracking for prosodic modeling in telephone speech
abstract
In this paper, we introduce a pitch detection algorithm that is particularly robust for telephone speech and prosodic modeling. The algorithm uses a logarithmically sampled spectral representation of speech, similar to that in the subharmonic summation approach. Constraints for logF/sub 0/ and /spl Delta/logF/sub 0/ are combined in a dynamic programming search to find an optimum pitch track. The search algorithm is able to find a continuous pitch contour regardless of the voicing status, while a separate voicing decision module computes the probability of voicing per frame. We evaluated the algorithm using the Keele pitch extraction reference database under both studio and telephone conditions. Our algorithm is very robust to channel degradation, and compares favorably to XWAVES under telephone conditions. It also significantly outperforms XWAVES when used for tone classification on a telephone quality Mandarin digit corpus.
Chao Wang 0018, Stephanie Seneff
ICASSP1
2000 MUXING: a telephone-access Mandarin conversational system
abstract
MUXING is a telephone-based conversational system that allows users to access weather information in Mandarin Chinese over the telephone. Although MUXING utilizes the same architecture as well as most of the same human language technology components as its English predecessor, JUPITER, some modifications to the system were necessary to account for differences between English and Mandarin Chinese. In addition, the weather database needed to be modified to reflect regions of greater interest to potential Chinese users. This paper describes our system development effort, paying particular attention to Mandarinspecific changes to the original JUPITER system. 1. INTRODUCTION For the past decade, our group has been conducting research leading to the development of conversational systems that enable users to access and manage information using spoken dialogue. In this context, multilinguality has always been an important research topic. Our approach to developing multilingual conversational...
Chao Wang 0018, D. Scott Cyphers, Xiaolong Mou, Joseph Polifroni, Stephanie Seneff, Jon Yi, Victor Zue
INTERSPEECH1
2000 Improved tone recognition by normalizing for coarticulation and intonation effects
abstract
We have previously demonstrated that tone modeling improved speech recognition on a digit corpus [7]. In this work, we further improve tone recognition by normalizing for both tone coarticulation and intonation effects. The tone classification errors on continuous digit strings were reduced by 26.1% from the baseline, when the effects of # # downdrift, phrase boundary and tone coarticulation were normalized. We also applied the same approach to conversational speech from the YINHE domain [6], and obtained similar improvements. The word error rate on spontaneous YINHE data was reduced by 16.5% when a simple fourtone model was applied to resort recognizer 10-best outputs. 1. INTRODUCTION Tone is a natural target for prosodic modeling in tonal languages, because of its important role in lexical access. There are four lexical tones in Mandarin Chinese, each corresponding to a canonical # # contour pattern: "high-level", "high-rising", "low-dipping" and "high-falling". However, tones in...
Chao Wang 0018, Stephanie Seneff
INTERSPEECH1
1998 A study of tones and tempo in continuous Mandarin digit strings and their application in telephone quality speech recognition
abstract
Prosodic cues (namely, fundamental frequency, energy and duration) provide important information for speech. For a tonal language such as Chinese, fundamental frequency ( ) plays a critical role in characterizing tone as well, which is an essential phonemic feature. In this paper, we describe our work on duration and tone modeling for telephone-quality continuous Mandarin digits, and the application of these models to improve recognition. The duration modeling includes a speaking-rate normalization scheme. A novel extraction algorithm is developed, and parameters based on orthonormal decomposition of contour are extracted for tone recognition. Context dependency is expressed by "tri-tone" models clustered into broader classes. A 20.0% error rate is achieved for four-tone classification. Over a baseline recognition performance of 5.1% word error rate, we achieve 31.4% error reduction with duration models, 23.5% error reduction with tone models, and 39.2% error reduction with duration and tone models combined.
Chao Wang 0018, Stephanie Seneff
ICSLP1
1997 YINHE: a Mandarin Chinese version of the GALAXY system
abstract
The galaxy system is a human-computer conversational system providing a spoken language interface for accessing on-line information. It was initially implemented for English in travel-related domains, including air travel, local city navigation, and weather. We began an effort to develop multilingual systems within the framework of galaxy several years ago. This paper describes our recent work on porting the system to Mandarin Chinese, including speech recognition, language understanding, and language generation components. Overall, the system produced reasonable responses nearly 70% of the time for spontaneous test data collected in a wizard environment. 1. INTRODUCTION The galaxy system is a client/server architecture for computer conversational systems [1]. In designing galaxy, we drew heavily on experience gained in the development of galaxy's predecessor, voyager [2]. Voyager was not initially designed to easily support multiple languages, but through a trial-and-error process...
Chao Wang 0018, James R. Glass, Helen M. Meng, Joseph Polifroni, Stephanie Seneff, Victor Zue
EUROSPEECH1