Xiaodan Zhuang

dblp:24/721 · DBLP profile ↗
← Back
37ranked-venue papers
12as first author
5since 2021 · last 2025
0009-0004-5481-8630ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 11 first-author · 5 since 2021Artificial intelligence and machine learning · 18 · 8 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
abstract
This paper presents an efficient decoding approach for end-to-end automatic speech recognition (E2E-ASR) with large language models (LLMs). Although shallow fusion is the most common approach to incorporate language models into E2E-ASR decoding, we face two practical problems with LLMs. (1) LLM inference is computationally costly. (2) There may be a vocabulary mismatch between the ASR model and the LLM. To resolve this mismatch, we need to retrain the ASR model and/or the LLM, which is at best time-consuming and in many cases not feasible. We propose delayed fusion, which applies LLM scores to ASR hypotheses with a delay during decoding and enables easier use of pre-trained LLMs in ASR tasks. This method can reduce not only the number of hypotheses scored by the LLM but also the number of LLM inference calls. It also allows re-tokenizion of ASR hypotheses during decoding if ASR and LLM employ different tokenizations. We demonstrate that delayed fusion provides improved decoding speed and accuracy compared to shallow fusion and N-best rescoring using the LibriHeavy ASR corpus and three public LLMs, OpenLLaMA 3B & 7B and Mistral 7B.
Takaaki Hori, Martin Kocour, Adnan Haider, Erik McDermott, Xiaodan Zhuang
ICASSP5
2024 Optimizing Byte-Level Representation For End-To-End ASR
abstract
We propose a novel approach to optimizing a byte-level representation for end-to-end automatic speech recognition (ASR). Byte-level representation is often used by large scale multilingual ASR systems when the character set of the supported languages is large. The compactness and universality of bytelevel representation allow the ASR models to use smaller output vocabularies and therefore, provide more flexibility. UTF8 is a commonly used byte-level representation for multilingual ASR, but it is not designed to optimize machine learning tasks directly. By using auto-encoder and vector quantization, we show that we can optimize a byte-level representation for ASR and achieve better accuracy. Our proposed framework can incorporate information from different modalities, and provides an error correction mechanism. In an English/Mandarin dictation task, we show that a bilingual ASR model built with this approach can outperform UTF-8 representation by 5% relative in error rate.
Roger Hsiao, Liuhui Deng, Erik McDermott, Ruchir Travadi, Xiaodan Zhuang
SLT5
2023 Variable Attention Masking for Configurable Transformer Transducer Speech Recognition
abstract
This work studies the use of attention masking in transformer transducer based speech recognition for building a single configurable model for different deployment scenarios. We present a comprehensive set of experiments comparing fixed masking, where the same attention mask is applied at every frame, with chunked masking, where the attention mask for each frame is determined by chunk boundaries, in terms of recognition accuracy and latency. We then explore the use of variable masking, where the attention masks are sampled from a target distribution at training time, to build models that can work in different configurations. Finally, we investigate how a single configurable model can be used to perform both first pass streaming recognition and second pass acoustic rescoring. Experiments show that chunked masking achieves a better accuracy vs latency trade-off compared to fixed masking, both with and without FastEmit. We also show that variable masking improves the accuracy by up to 8% relative in the acoustic re-scoring scenario.
Pawel Swietojanski, Dogan Can, Thiago Fraga da Silva, Arnab Ghoshal, Takaaki Hori, Roger Hsiao, Henry Mason, Erik McDermott, Honza Silovsky, Ruchir Travadi, Xiaodan Zhuang
ICASSP12
2023 Approximate Nearest Neighbour Phrase Mining for Contextual Speech Recognition
Maurits J. R. Bleeker, Pawel Swietojanski, Xiaodan Zhuang
INTERSPEECH4
2021 Frame-Level Specaugment for Deep Convolutional Neural Networks in Hybrid ASR Systems
abstract
Inspired by SpecAugment - a data augmentation method for end-to-end ASR systems, we propose a frame-level SpecAugment method (f-SpecAugment) to improve the performance of deep convolutional neural networks (CNN) for hybrid HMM based ASR systems. Similar to the utterance level SpecAugment, f-SpecAugment performs three transformations: time warping, frequency masking, and time masking. Instead of applying the transformations at the utterance level, f-SpecAugment applies them to each convolution window independently during training. We demonstrate that f-SpecAugment is more effective than the utterance level SpecAugment for deep CNN based hybrid models. We evaluate the proposed f-SpecAugment on 50-layer Self-Normalizing Deep CNN (SNDCNN) acoustic models trained with up to 25000 hours of training data. We observe f-SpecAugment reduces WER by 0.5-4.5% relatively across different ASR tasks for four languages. As the benefits of augmentation techniques tend to diminish as training data size increases, the large scale training reported is important in understanding the effectiveness of f-SpecAugment. Our experiments demonstrate that even with 25k training data, f-SpecAugment is still effective. We also demonstrate that f-SpecAugment has benefits approximately equivalent to doubling the amount of training data for deep CNNs.
Xiaodan Zhuang, Daben Liu
SLT3
2020 SNDCNN: Self-Normalizing Deep CNNs with Scaled Exponential Linear Units for Speech Recognition
abstract
Very deep CNNs achieve state-of-the-art results in both computer vision and speech recognition, but are difficult to train. The most popular way to train very deep CNNs is to use shortcut connections (SC) together with batch normalization (BN). Inspired by Self-Normalizing Neural Networks, we propose the self-normalizing deep CNN (SNDCNN) based acoustic model topology, by removing the SC/BN and replacing the typical RELU activations with scaled exponential linear unit (SELU) in ResNet-50. SELU activations make the network self-normalizing and remove the need for both shortcut connections and batch normalization. Compared to ResNet50, we can achieve the same or lower (up to 4.5% relative) word error rate (WER) while boosting both training and inference speed by 60%-80%. We also explore other model inference optimization schemes to further reduce latency for production use.
Zhen Huang 0001, Tim Ng, Leo Liu, Henry Mason, Xiaodan Zhuang, Daben Liu
ICASSP5
2019 Exploring Retraining-free Speech Recognition for Intra-sentential Code-switching
abstract
Code Switching refers to the phenomenon of changing languages within a sentence or discourse, and it represents a challenge for conventional automatic speech recognition systems deployed to tackle a single target language. The code switching problem is complicated by the lack of multi-lingual training data needed to build new and ad hoc multi-lingual acoustic and language models. In this work, we present a prototype research code-switching speech recognition system that leverages existing monolingual acoustic and language models, i.e., no ad hoc training is needed. To generate high quality pronunciation of foreign language words in the native language phoneme set, we use a combination of existing acoustic phone decoders and an LSTM-based grapheme-to-phoneme model. In addition, a code-switching language model was developed by using translated word pairs to borrow statistics from the native language model. We demonstrate that our approach handles accented foreign pronunciations better than techniques based on human labeling. Our best system reduces the WER from 34.4%, obtained with a conventional monolingual speech recognition system, to 15.3% on an intra-sentential code-switching task, without harming the monolingual accuracy.
Zhen Huang 0001, Xiaodan Zhuang, Daben Liu, Xiaoqiang Xiao, Sabato Marco Siniscalchi
ICASSP2
2017 Improving DNN Bluetooth Narrowband Acoustic Models by Cross-Bandwidth and Cross-Lingual Initialization
Xiaodan Zhuang, Arnab Ghoshal, Antti-Veikko Rosti, Matthias Paulik, Daben Liu
INTERSPEECH1
2014 Zero-Shot Event Detection Using Multi-modal Fusion of Weakly Supervised Concepts
abstract
Current state-of-the-art systems for visual content analysis require large training sets for each class of interest, and performance degrades rapidly with fewer examples. In this paper, we present a general framework for the zeroshot learning problem of performing high-level event detection with no training exemplars, using only textual descriptions. This task goes beyond the traditional zero-shot framework of adapting a given set of classes with training data to unseen classes. We leverage video and image collections with free-form text descriptions from widely available web sources to learn a large bank of concepts, in addition to using several off-the-shelf concept detectors, speech, and video text for representing videos. We utilize natural language processing technologies to generate event description features. The extracted features are then projected to a common high-dimensional space using text expansion, and similarity is computed in this space. We present extensive experimental results on the large TRECVID MED [26] corpus to demonstrate our approach. Our results show that the proposed concept detection methods significantly outperform current attribute classifiers such as Classemes [34], ObjectBank [21], and SUN attributes[28] . Further, we find that fusion, both within as well as between modalities, is crucial for optimal performance.
Shuang Wu 0003, Sravanthi Bondugula, Florian Luisier, Xiaodan Zhuang, Pradeep Natarajan
CVPR4
2014 Text Classification via iVector Based Feature Representation
abstract
In this paper, we address the problem of text classification: classifying modern machine-printed text, handwritten text and historical typewritten text from degraded noisy documents. We propose a novel text classification approach based on iVector, a newly developed concept in speaker verification. To a given text line, the iVector is a fixed-length feature vector representation, transformed from a high-dimensional super vector based on means of Gaussian mixture model (GMM), where the text dependent component is separated from a universal background model (UBM) and can be represented by a low dimensional set of factors. We classify the text lines with a discriminative classifier - support vector machine (SVM) in iVector space. A baseline approach of text classification using GMM in feature space is also presented for evaluation purpose. Experimental results on an Arabic document database show accuracy of 92.04% for text line classification using the proposed method. Furthermore, the relative word error rate (WER) of 9.6% is decreased in optical character recognition (OCR) when coupled with the proposed iVector-SVM classifier. The proposed iVector-SVM approach is language independent, thus, can be applied to other scripts as well.
Shengxin Zha, Xujun Peng, Huaigu Cao, Xiaodan Zhuang, Pradeep Natarajan, Premkumar Natarajan
Document Analysis Systems4
2014 Text detection and recognition in natural scenes and consumer videos
abstract
We propose an end-to-end system for text detection and recognition in natural scenes and consumer videos. Maximally Stable Extremal Regions which are robust to illumination and viewpoint variations are selected as text candidates. Rich shape descriptors such as Histogram of Oriented Gradients, Gabor filter, corners and geometrical features are used to represent the candidates and classified using a support vector machine. Positively labeled candidates serve as anchor regions for word formation. We then group candidate regions based on geometric and color properties to form word boundaries. To speed up the system for practical applications, we use Partial Least Squares approach for dimensionality reduction. The detected words are binarized, filtered and passed to a hidden Markov model based Optical Character Recognition (OCR) system for recognition. We show significant improvement in text detection and recognition tasks over previous approaches on a large consumer video dataset. Furthermore, the event detection system built upon the OCR output of this approach outperformed multiple other OCR-only based submissions in the recently concluded NIST TRECVID 2013 multimedia event detection evaluations.
Xujun Peng, Xiaodan Zhuang, Pradeep Natarajan, Huaigu Cao
ICASSP3
2014 Effective representations for leveraging language content in multimedia event detection
abstract
Language content in videos from speech and overlaid or inscene video text can provide high precision signals for video event detection and retrieval. However, sporadic occurrence, content that is unrelated to the events of interest, and high error rates of current speech and text recognition systems on consumer domain video make it difficult to exploit these channels. In this paper, we study different representations of language content to address these challenges. First, we utilize likelihood weighted word lattices obtained from a Hidden Markov Model (HMM) based decoding engine to encode many alternate hypotheses, rather than relying on noisy single best hypotheses. Second, we utilize an event-independent modified term frequency-inverse document frequency (TF-IDF) weighting scheme to obtain the final feature vector. We present detailed experimental results on the TRECVID MED 2013 dataset containing ~150000 videos, and show that our representation significantly outperforms alternate representations for both speech and video text.
Shuang Wu 0003, Xiaodan Zhuang, Pradeep Natarajan
ICASSP2
2014 Improving speech-based PTSD detection via multi-view learning
abstract
We demonstrate that by applying multi-view learning algorithms one can usefully leverage highly informative, highcost, psychophysiological data collected in a laboratory setting, to improve PTSD screening in the field, where only less-informative, low-cost, speech data are available. Cost metrics reflect resource requirements as well as subject receptivity to data collection. The speech-based representation involves distress indicator extraction from automatic speech recognition output, and a compact holistic audio representation based on the i-vector method. A prototype PTSD screening system was developed that benefits from highly informative EEG data yet, in the field, only relies on subjects' spoken commentary in response to open ended questions. Such a system can deliver screening with significantly increased engagement, to a broader population, leading to earlier intervention and improved outcomes. Using a recent dataset collected for multi-modal computer-aided diagnosis of PTSD, we demonstrate that the proposed method significantly improves speech-based PTSD detection, without requiring costly and aversive procedures at deployment.
Xiaodan Zhuang, Viktor Rozgic, Michael Crystal, Brian Marx
SLT1
2013 Audio self organized units for high-level event detection
Xiaodan Zhuang, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH1
2013 Probabilistic trainable segmenter for call center audio using multiple features
Nina Zinovieva, Xiaodan Zhuang, Pat Peterson, Joe Alwan, Rohit Prasad
INTERSPEECH2
2013 Compact bag-of-words visual representation for effective linear classification
abstract
Bag-of-words approaches have been shown to achieve state-of-the-art performance in large-scale multimedia event detection. However, the commonly used histogram representation of bag-of-words requires large codebook sizes and expensive nonlinear kernel based classifiers for optimal performance. To address these two issues, we present a two-part generative model for compact visual representation, based on the i-vector approach recently proposed for speech and audio modeling. First, we use a Gaussian mixture model (GMM) to model the joint distribution of local descriptors. Second, we use a low-dimensional factor representation that constrains the GMM parameters to a subspace that preserves most of the information. We further extend this method to incorporate overlapping spatial regions, forming a highly compact visual representation that achieves superior performance with fast linear classifiers. We evaluate the method on a large video dataset used in the TRECVID 2011 MED evaluation. With linear classifiers, the proposed representation, with one-tenth of the storage footprint, outperforms soft quantization histograms used in the top performing TRECVID 2011 MED systems.
Xiaodan Zhuang, Shuang Wu 0003, Pradeep Natarajan
ACM Multimedia1
2013 Scene image categorization and video event detection using Naive Bayes Nearest Neighbor
abstract
We present a detailed study of Naive Bayes Nearest Neighbor (NBNN) proposed by Boiman et al., with application to scene categorization and video event detection. Our study indicates that using Dense-SIFT along with dimensionality reduction using PCA enables NBNN to obtain state-of-the-art results. We demonstrate this on two tasks: (1) scene image categorization on the UIUC 8 Sports Events Image Dataset (obtaining 84.67%) and the MIT 67 Indoor Scene Image Dataset (obtaining 48.84%); and (2) detecting videos depicting certain events of interest on the challenging MED'11 video dataset with only 15 positive training videos per event. We present an extension referred to as sparse-NBNN that constrains the number of training images that can used to match with a given test image for the image-to-class distance computation. Experiments indicate that this improves upon NBNN for handling of imbalanced training data.
Shiv Vitaladevuni, Pradeep Natarajan, Shuang Wu 0003, Xiaodan Zhuang, Rohit Prasad, Premkumar Natarajan
WACV4
2013 Saliency-maximized audio visualization and efficient audio-visual browsing for faster-than-real-time human acoustic event detection
abstract
Browsing large audio archives is challenging because of the limitations of human audition and attention. However, this task becomes easier with a suitable visualization of the audio signal, such as a spectrogram transformed to make unusual audio events salient. This transformation maximizes the mutual information between an isolated event's spectrogram and an estimate of how salient the event appears in its surrounding context. When such spectrograms are computed and displayed with fluid zooming over many temporal orders of magnitude, sparse events in long audio recordings can be detected more quickly and more easily. In particular, in a 1/10-real-time acoustic event detection task, subjects who were shown saliency-maximized rather than conventional spectrograms performed significantly better. Saliency maximization also improves the mutual information between the ground truth of nonbackground sounds and visual saliency, more than other common enhancements to visualization.
Kai-Hsiang Lin, Xiaodan Zhuang, Camille Goudeseune, Sarah King, Mark Hasegawa-Johnson, Thomas S. Huang
ACM Trans. Appl. Percept.2
2012 Multimodal feature fusion for robust event detection in web videos
abstract
Combining multiple low-level visual features is a proven and effective strategy for a range of computer vision tasks. However, limited attention has been paid to combining such features with information from other modalities, such as audio and videotext, for large scale analysis of web videos. In our work, we rigorously analyze and combine a large set of low-level features that capture appearance, color, motion, audio and audio-visual co-occurrence patterns in videos. We also evaluate the utility of high-level (i.e., semantic) visual information obtained from detecting scene, object, and action concepts. Further, we exploit multimodal information by analyzing available spoken and videotext content using state-of-the-art automatic speech recognition (ASR) and videotext recognition systems. We combine these diverse features using a two-step strategy employing multiple kernel learning (MKL) and late score level fusion methods. Based on the TRECVID MED 2011 evaluations for detecting 10 events in a large benchmark set of ~45000 videos, our system showed the best performance among the 19 international teams.
Pradeep Natarajan, Shuang Wu 0003, Shiv Vitaladevuni, Xiaodan Zhuang, Stavros Tsakalidis, Unsang Park, Rohit Prasad, Premkumar Natarajan
CVPR4
2012 Multi-channel Shape-Flow Kernel Descriptors for Robust Video Event Detection and Retrieval
Pradeep Natarajan, Shuang Wu 0003, Shiv Vitaladevuni, Xiaodan Zhuang, Unsang Park, Rohit Prasad, Premkumar Natarajan
ECCV (2)4
2012 Improving faster-than-real-time human acoustic event detection by saliency-maximized audio visualization
abstract
We propose a saliency-maximized audio spectrogram as a representation that lets human analysts quickly search for and detect events in audio recordings. By rendering target events as visually salient patterns, this representation minimizes the time and effort needed to examine a recording. In particular, we propose a transformation of a conventional spectrogram that maximizes the mutual information between the spectrograms of isolated target events and the estimated saliency of the overall visual representation. When subjects are shown spectrograms that are saliency-maximized, they perform significantly better in a 1/10-real-time acoustic event detection task.
Kai-Hsiang Lin, Xiaodan Zhuang, Camille Goudeseune, Sarah King, Mark Hasegawa-Johnson, Thomas S. Huang
ICASSP2
2012 Robust Event Detection From Spoken Content In Consumer Domain Videos
Stavros Tsakalidis, Xiaodan Zhuang, Roger Hsiao, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH2
2012 Compact Audio Representation for Event Detection in Consumer Media
Xiaodan Zhuang, Stavros Tsakalidis, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH1
2011 Improving acoustic event detection using generalizable visual features and multi-modality modeling
abstract
Acoustic event detection (AED) aims to identify both timestamps and types of multiple events and has been found to be very challenging. The cues for these events often times exist in both audio and vision, but not necessarily in a synchronized fashion. We study improving the detection and classification of the events using cues from both modalities. We propose optical flow based spatial pyramid histograms as a generalizable visual representation that does not require training on labeled video data. Hidden Markov models (HMMs) are used for audio-only modeling, and multi-stream HMMs or coupled HMMs (CHMM) are used for audio-visual joint modeling. To allow the flexibility of audio-visual state asynchrony, we explore effective CHMM training via HMM state-space mapping, parameter tying and different initialization schemes. The proposed methods successfully improve acoustic event classification and detection on a multimedia meeting room dataset containing eleven types of general non-speech events without using extra data resource other than the video stream accompanying the audio observations. Our systems perform favorably compared to previously reported systems leveraging ad-hoc visual cue detectors and localization information obtained from multiple microphones.
Po-Sen Huang, Xiaodan Zhuang, Mark Hasegawa-Johnson
ICASSP2
2011 Synthesizing visual speech trajectory with minimum generation error
abstract
In this paper, we propose a minimum generation error (MGE) training method to refine the audio-visual HMM to improve visual speech trajectory synthesis. Compared with the traditional maximum likelihood (ML) estimation, the proposed MGE training explicitly optimizes the quality of generated visual speech trajectory, where the audio-visual HMM modeling is jointly refined by using a heuristic method to find the optimal state alignment and a probabilistic descent algorithm to optimize the model parameters under the MGE criterion. In objective evaluation, compared with the ML-based method, the proposed MGE-based method achieves consistent improvement in the mean square error reduction, correlation increase, and recovery of global variance. It also improves the naturalness and audio-visual consistency perceptually in the subjective test.
Yi-Jian Wu, Xiaodan Zhuang, Frank K. Soong
ICASSP3
2010 FSM-based pronunciation modeling using articulatory phonological code
abstract
According to articulatory phonology, the gestural score is an invariant speech representation. Though the timing schemes, i.e., the onsets and offsets, of \nthe gestural activations may vary, the ensemble of these activations tends to remain unchanged, informing the speech content. "Gestural pattern vector" \n(GPV) has been proposed to encode the instantaneous gestural activations that exist across all tract variables at each time. Therefore, a gestural score with a particular timing scheme can be approximated using a GPV sequence. \n \nIn this work, we propose a pronunciation modeling method that uses a finite state machine (FSM) to represent the invariance of a gestural score. Given the "canonical" gestural score of a word with a known activation timing \nscheme, the plausible activation onsets and offsets are recursively generated and encoded as a weighted FSM. An empirical measure is used to prune out gestural activation timing schemes that deviate too much from the "canonical" gestural score. Speech recognition is achieved by matching the recovered \ngestural activations to the FSM-encoded gestural scores of different speech contents. In particular, the observation distribution of each GPV is modeled \nby an artificial neural network and Gaussian mixture tandem model. These models are used together with the FSM-based pronunciation models in a Bayesian framework. \n \nWe carry out pilot word classification experiments using synthesized data from one speaker. The proposed pronunciation modeling achieves over 90% accuracy for a vocabulary of 139 words with no training observations, outperforming direct use of the "canonical" gestural score.
Chi Hu, Xiaodan Zhuang, Mark Hasegawa-Johnson
INTERSPEECH2
2010 A minimum converted trajectory error (MCTE) approach to high quality speech-to-lips conversion
abstract
High quality speech-to-lips conversion, investigated in this work, ren-ders realistic lips movement (video) consistent with input speech (audio) without knowing its linguistic content. Instead of memoryless frame-based conversion, we adopt maximum likelihood estimation of the vi-sual parameter trajectories using an audio-visual joint Gaussian Mixture Model (GMM). We propose a minimum converted trajectory error ap-proach (MCTE) to further refine the converted visual parameters. First, we reduce the conversion error by training the joint audio-visual GMM with weighted audio and visual likelihood. Then MCTE uses the gen-eralized probabilistic descent algorithm to minimize a conversion error of the visual parameter trajectories defined on the optimal Gaussian ker-nel sequence according to the input speech. We demonstrate the effec-tiveness of the proposed methods using the LIPS 2009 Visual Speech Synthesis Challenge dataset, without knowing the linguistic (phonetic) content of the input speech. Index Terms: visual speech synthesis, speech-to-lips conversion, mini-mum conversion error, minimum generation error
Xiaodan Zhuang, Frank K. Soong, Mark Hasegawa-Johnson
INTERSPEECH1
2010 Novel Gaussianized vector representation for improved natural scene categorization
Xiaodan Zhuang, Hao Tang 0001, Mark Hasegawa-Johnson, Thomas S. Huang
Pattern Recognit. Lett.2
2010 Real-world acoustic event detection
Xiaodan Zhuang, Mark Hasegawa-Johnson, Thomas S. Huang
Pattern Recognit. Lett.1
2009 Long-time span acoustic activity analysis from far-field sensors in smart homes
abstract
Smart homes for the aging population have recently started attracting the attention of the research community. One of the problems of interest is this of monitoring the activities of daily living (ADLs) of the elderly, in order to help identify critical problems, aiming to improve their protection and general well-being. In this paper, we report on our initial attempts to recognize such activities, based on input from networks of far-field microphones distributed inside the home. We propose two approaches to the problem: The first models the entire activity, which typically covers long time spans, with a single statistical model, for example a hidden Markov model (HMM), a Gaussian mixture model (GMM), or GMM super-vectors in conjunction with support vector machines (SVMs). The second is a two-step approach: It first performs acoustic event detection (AED) to locate distinctive events, characteristic of the ADLs, and it is subsequently followed by a post-processing stage that employs activity-specific language models (LMs) to classify the output sequences of detected events into ADLs. Experiments are reported on a corpus containing a small number of acted ADLs, collected as part of the Netcarity Integrated Project inside a two-room smart home. Our results show that SVM GMM supervector modeling improves six-class ADL classification accuracy to 76%, compared to 56% achieved by the GMMs, while also outperforming HMMs by 8% absolute. Preliminary results from LM scoring of acoustic event sequences are comparable to those from GMMs on a three-class ADL classification task.
Jing Huang 0019, Xiaodan Zhuang, Vit Libal, Gerasimos Potamianos
ICASSP2
2009 Acoustic fall detection using Gaussian mixture models and GMM supervectors
abstract
We present a system that detects human falls in the home environment, distinguishing them from competing noise, by using only the audio signal from a single far-field microphone. The proposed system models each fall or noise segment by means of a Gaussian mixture model (GMM) supervector, whose Euclidean distance measures the pairwise difference between audio segments. A support vector machine built on a kernel between GMM supervectors is employed to classify audio segments into falls and various types of noise. Experiments on a dataset of human falls, collected as part of the Netcarity project, show that the method improves fall classification F-score to 67% from 59% of a baseline GMM classifier. The approach also effectively addresses the more difficult fall detection problem, where audio segment boundaries are unknown. Specifically, we employ it to reclassify confusable segments produced by a dynamic programming scheme based on traditional GMMs. Such post-processing improves a fall detection accuracy metric by 5% relative.
Xiaodan Zhuang, Jing Huang 0019, Gerasimos Potamianos, Mark Hasegawa-Johnson
ICASSP1
2009 Articulatory phonological code for word classification
abstract
We propose a framework that leverages articulatory phonology for speech recognition. “Gestural pattern vectors ” (GPV) encode the instantaneous gestural activations that exist across all tract variables at each time. Given a speech observation, recognizing the sequence of GPV recovers the ensemble of gestural activations, i.e., the gestural score. For each word in the vocabulary, we use a task dynamic model of inter-articulator speech coordination to generate the “canonical ” gestural score. Speech recognition is achieved by matching the ensemble of gestural activations. In particular, we estimate the likelihood of the recognized GPV sequence on word-dependent GPV sequence models trained using the “canonical” gestural scores. These likelihoods, weighted by confidence score of the recognized GPVs, are used in a Bayesian speech recognizer. Pilot gestural score recovery and word classification experiments are carried out using synthesized data from one speaker. The observation distribution of each GPV is modeled by an artificial neural network and Gaussian mixture tandem model. Bigram GPV sequence models are used to distinguish gestural scores of different words. Given the tract variable time functions, about 80 % of the instantaneous gestural activation is correctly recovered. Word recognition accuracy is over 85 % for a vocabulary of 139 words with no training observations. These results suggest that the proposed framework might be a viable alternative to the classic sequence-of-phones model. Index Terms: speech production, speech gesture, tandem model, artificial neural network, Gaussian mixture model
Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman
INTERSPEECH1
2008 Feature analysis and selection for acoustic event detection
abstract
Speech perceptual features, such as Mel-frequency Cepstral Coefficients (MFCC), have been widely used in acoustic event detection. However, the different spectral structures between speech and acoustic events degrade the performance of the speech feature sets. We propose quantifying the discriminative capability of each feature component according to the approximated Bayesian accuracy and deriving a discriminative feature set for acoustic event detection. Compared to MFCC, feature sets derived using the proposed approaches achieve about 30% relative accuracy improvement in acoustic event detection.
Xiaodan Zhuang, Thomas S. Huang, Mark Hasegawa-Johnson
ICASSP1
2008 A novel Gaussianized vector representation for natural scene categorization
abstract
This paper presents a novel Gaussianized vector representation for scene images by an unsupervised approach. First, each image is encoded as an ensemble of orderless bag of features, and then a global Gaussian Mixture Model (GMM) learned from all images is used to randomly distribute each feature into one Gaussian component by a multinomial trial. The parameters of the multinomial distribution are defined by the posteriors of the feature on all the Gaussian components. Finally, the normalized means of the features distributed in every Gaussian component are concatenated to form a supervector, which is a compact representation for each scene image. We prove that these super-vectors observe the standard normal distribution. Our experiments on scene categorization tasks using this vector representation show significantly improved performance compared with the bag-of-features representation.
Xiaodan Zhuang, Hao Tang 0001, Mark Hasegawa-Johnson, Thomas S. Huang
ICPR2
2008 Face age estimation using patch-based hidden Markov model supervectors
abstract
Recent studies in patch-based Gaussian Mixture Model (GMM) approaches for face age estimation present promising results. We propose using a hidden Markov model (HMM) supervector to represent face image patches, to improve from the previous GMM supervector approach by capturing the spatial structure of human faces and loosening the assumption of identical face patch distribution within a face image. The Euclidean distance of HMM supervectors constructed from two face images measures the similarity of the human faces, derived from the approximated Kullback-Leibler divergence between the joint distributions of patches with implicit unsupervised alignment of different regions in two human faces. The proposed HMM supervector approach compares favorably with the GMM supervector approach in face age estimation on a large face dataset.
Xiaodan Zhuang, Mark Hasegawa-Johnson, Thomas S. Huang
ICPR1
2008 The entropy of the articulatory phonological code: recognizing gestures from tract variables
abstract
We propose an instantaneous “gestural pattern vector ” to encode the instantaneous pattern of gesture activations across tract variables in the gestural score. The design of these gestural pattern vectors is the first step towards an automatic speech recognizer motivated by articulatory phonology, which is expected to be more invariant to speech coarticulation and reduction than conventional speech recognizers built with the sequenceof-phones assumption. We use a tandem model to recover the instantaneous gestural pattern vectors from tract variable time functions in local time windows, and achieve classification accuracy up to 84.5% for synthesized data from one speaker. Recognizing all gestural pattern vectors is equivalent to recognizing the ensemble of gestures. This result suggests that the proposed gestural pattern vector might be a viable unit in statistical models for speech recognition. Index Terms: speech production, speech gesture, tandem model, artificial neural network, Gaussian mixture model
Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman
INTERSPEECH1
2008 SIFT-Bag kernel for video event analysis
abstract
In this work, we present a SIFT-Bag based generative-todiscriminative framework for addressing the problem of video event recognition in unconstrained news videos. In the generative stage, each video clip is encoded as a bag of SIFT feature vectors, the distribution of which is described by a Gaussian Mixture Models (GMM). In the discriminative stage, the SIFT-Bag Kernel is designed for characterizing the property of Kullback-Leibler divergence between the specialized GMMs of any two video clips, and then this kernel is utilized for supervised learning in two ways. On one hand, this kernel is further refined in discriminating power for centroid-based video event classification by using the Within-Class Covariance Normalization approach, which depresses the kernel components with high-variability for video clips of the same event. On the other hand, the SIFT-Bag Kernel is used in a Support Vector Machine for margin-based video event classification. Finally, the outputs from these two classifiers are fused together for final decision. The experiments on the TRECVID 2005 corpus demonstrate that the mean average precision is boosted from the best reported 38.2 % in [36] to 60.4 % based on our new framework.
Xiaodan Zhuang, Shuicheng Yan, Shih-Fu Chang, Mark Hasegawa-Johnson, Thomas S. Huang
ACM Multimedia2