EDBT 2026 Demo / reviewers in the wild / expert
Bowen Shi 0002
dblp:169/3160-2
· DBLP profile ↗
30ranked-venue papers
13as first author
22since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 10 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Meta Audiobox Aesthetics: Unified Automatic Assessment for Speech, Music and SoundabstractQuantifying audio aesthetics is challenging due to its subjective nature, influenced by human perception and cultural context. Traditional methods rely on human listeners, leading to inconsistencies and high resource demands. This paper addresses the growing need for automated systems capable of predicting audio aesthetics without human intervention. Such systems are crucial for applications like data filtering, pseudo-labeling, and evaluating generative models.In this paper, we propose new annotation guidelines that break down human listening perspectives into four axes and develop no-reference, peritem prediction models for more nuanced audio quality assessment. Our models are evaluated against human mean opinion scores (MOS) and existing methods, demonstrating comparable or superior performance. This research not only advances the field of audio aesthetics but also provides open-source models and datasets to facilitate future work and benchmarking. Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoffman, Brian Ellis, Apoorv Vyas, Bowen Shi 0002, Sanyuan Chen, Matt Le 0001, Nick Zacharov, Carleigh Wood, Ann Lee 0001, Wei-Ning Hsu |
ASRU | 7 |
| 2025 | Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
Mu Yang, Bowen Shi 0002, Matt Le 0001, Wei-Ning Hsu, Andros Tjandra |
INTERSPEECH | 2 |
| 2024 | XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech PerceptionabstractHyoJung Han, Mohamed Anwar, Juan Pino, Wei-Ning Hsu, Marine Carpuat, Bowen Shi, Changhan Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. HyoJung Han 0001, Mohamed Anwar, Juan Pino 0001, Wei-Ning Hsu, Marine Carpuat, Bowen Shi 0002, Changhan Wang |
ACL (1) | 6 |
| 2024 | M2BART: Multilingual and Multimodal Encoder-Decoder Pre-Training for Any-to-Any Machine TranslationabstractSpeech and language models are advancing towards universality. A single model can now handle translations across 200 languages and transcriptions for over 100 languages. Universal models simplify development, deployment, and importantly, transfer knowledge to less-resourced languages or modes. This paper introduces M2BART, a streamlined multilingual and multimodal framework for encoderdecoder models. It employs a self-supervised speech tokenizer, bridging speech and text, and is pre-trained with a unified objective for both unimodal and multimodal, unsupervised and supervised data. When tested on Spanish-to-English and English-to-Hokkien translations, M2BART consistently surpassed competitors. We also showcase an innovative translation model enabling zero-shot transfers even without labeled data. Peng-Jen Chen, Bowen Shi 0002, Kelvin Niu, Ann Lee 0001, Wei-Ning Hsu |
ICASSP | 2 |
| 2024 | Generative Pre-training for Speech with Flow MatchingabstractGenerative models have gained more and more attention in recent years for their remarkable success in tasks that required estimating and sampling data distribution to generate high-fidelity synthetic data. In speech, text-to-speech synthesis and neural vocoder are good examples where generative models have shined. While generative models have been applied to different applications in speech, there exists no general-purpose generative model that models speech directly. In this work, we take a step toward this direction by showing a single pre-trained generative model can be adapted to different downstream tasks with strong performance. Specifically, we pre-trained a generative model, named SpeechFlow, on 60k hours of untranscribed speech with Flow Matching and masked conditions. Experiment results show the pre-trained generative model can be fine-tuned with task-specific data to match or surpass existing expert models on speech enhancement, separation, and synthesis. Our work suggested a foundational model for generation tasks in speech can be built with generative pre-training. Alexander H. Liu, Matt Le 0001, Apoorv Vyas, Bowen Shi 0002, Andros Tjandra, Wei-Ning Hsu |
ICLR | 4 |
| 2024 | MusicFlow: Cascaded Flow Matching for Text Guided Music GenerationabstractWe introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music audios, we construct two flow matching networks to model the conditional distribution of semantic and acoustic features. Additionally, we leverage masked prediction as the training objective, enabling the model to generalize to other tasks such as music infilling and continuation in a zero-shot manner. Experiments on MusicCaps reveal that the music generated by MusicFlow exhibits superior quality and text coherence despite being over $2\sim5$ times smaller and requiring $5$ times fewer iterative steps. Simultaneously, the model can perform other music generation tasks and achieves competitive performance in music infilling and continuation. K. R. Prajwal, Bowen Shi 0002, Matt Le 0001, Apoorv Vyas, Andros Tjandra, Mahi Luthra, Baishan Guo, Triantafyllos Afouras, David Kant, Wei-Ning Hsu |
ICML | 2 |
| 2024 | Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning
Chung-Ming Chien, Andros Tjandra, Apoorv Vyas, Matt Le 0001, Bowen Shi 0002, Wei-Ning Hsu |
INTERSPEECH | 5 |
| 2024 | Data Efficient Reflow for Few Step Audio GenerationabstractFlow matching has been successfully applied onto generative models, particularly in producing high-quality images and audio. However, the iterative sampling required for the ODE solver in flow matching-based approaches can be time-consuming. Reflow finetune, a technique derived from Rectified flow, offers a promising solution by transforming the ODE trajectory into a straight one, thereby reducing the number of sampling steps. In this paper, we focus on developing data-efficient flow-based approaches for text-to-audio generation. We found that directly applying reflow to the pre-trained flow matching-based audio generation models is typically computationally expensive. It requires over 50,000 training iterations and five times the amount of training data to achieve satisfactory results. To address this issue, we introduce a novel data-efficient reflow (DEreflow) method. This method modifies the reflow data pairs and trajectory to align with the flow matching distribution. As a result of this alignment, our approach requires significantly fewer steps (8,000 compared to 50,000) and data pairs $(0.5$ times the scale of training data compared to 5 times). Results show that the proposed DEreflow consistently outperforms the original reflow method on the text-to-audio generation task. Lemeng Wu, Zhaoheng Ni, Bowen Shi 0002, Gaël Le Lan, Anurag Kumar 0003, Varun Nagaraja, Xinhao Mei, Yunyang Xiong, Bilge Soran, Raghuraman Krishnamoorthi, Wei-Ning Hsu, Yangyang Shi, Vikas Chandra |
SLT | 3 |
| 2024 | Scaling Speech Technology to 1, 000+ LanguagesabstractExpanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to about one hundred languages which is a small fraction of the over 7,000 languages spoken around the world. The Massively Multilingual Speech (MMS) project increases the number of supported languages by 10-40x, depending on the task while providing improved accuracy compared to prior work. The main ingredients are a new dataset based on readings of publicly available religious texts and effectively leveraging self-supervised learning. We built pre-trained wav2vec 2.0 models covering 1,406 languages, a single multilingual automatic speech recognition model for 1,107 languages, speech synthesis models for the same number of languages, as well as a language identification model for 4,017 languages. Experiments show that our multilingual speech recognition model more than halves the word error rate of Whisper on 54 languages of the FLEURS benchmark while being trained on a small fraction of the labeled data. Vineel Pratap, Andros Tjandra, Bowen Shi 0002, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, Alexei Baevski, Yossi Adi, Xiaohui Zhang 0007, Wei-Ning Hsu, Alexis Conneau, Michael Auli |
J. Mach. Learn. Res. | 3 |
| 2023 | ReVISE: Self-Supervised Speech Resynthesis with Visual Input for Universal and Generalized Speech RegenerationabstractPrior works on improving speech quality with visual input typically study each type of auditory distortion separately (e.g., separation, inpainting, video-to-speech) and present tailored algorithms. This paper proposes to unify these subjects and study Generalized Speech Regeneration, where the goal is not to reconstruct the exact reference clean signal, but to focus on improving certain aspects of speech while not necessarily preserving the rest such as voice. In particular, this paper concerns intelligibility, quality, and video synchronization. We cast the problem as audio-visual speech resynthesis, which is composed of two steps: pseudo audio-visual speech recognition (P-AVSR) and pseudo text-to-speech synthesis (P-TTS). P-AVSR and P-TTS are connected by discrete units derived from a self-supervised speech model. Moreover, we utilize self-supervised audio-visual speech model to initialize P-AVSR. The proposed model is coined ReVISE. ReVISE is the first high-quality model for in-the-wild video-to-speech synthesis and achieves superior performance on all LRS3 audio-visual regeneration tasks with a single model. To demonstrates its applicability in the real world, ReVISE is also evaluated on EasyCom, an audio-visual benchmark collected under challenging acoustic conditions with only 1.6 hours of training data. Similarly, ReVISE greatly suppresses noise and improves quality. Project page: https://wnhsu.github.io/ReVISE/. Wei-Ning Hsu, Tal Remez, Bowen Shi 0002, Jacob Donley, Yossi Adi |
CVPR | 3 |
| 2023 | Comparative Layer-Wise Analysis of Self-Supervised Speech ModelsabstractMany self-supervised speech models, varying in their pre-training objective, input modality, and pre-training data, have been proposed in the last few years. Despite impressive successes on downstream tasks, we still have a limited understanding of the properties encoded by the models and the differences across models. In this work, we examine the intermediate representations for a variety of recent models. Specifically, we measure acoustic, phonetic, and word-level properties encoded in individual layers, using a lightweight analysis tool based on canonical correlation analysis (CCA). We find that these properties evolve across layers differently depending on the model, and the variations relate to the choice of pre-training objective. We further investigate the utility of our analyses for downstream tasks by comparing the property trends with performance on speech recognition and spoken language understanding tasks. We discover that CCA trends provide reliable guidance to choose layers of interest for downstream tasks and that single-layer performance often matches or improves upon using all layers, suggesting implications for more efficient use of pre-trained models.1 Ankita Pasad, Bowen Shi 0002, Karen Livescu |
ICASSP | 2 |
| 2023 | MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation
Mohamed Anwar, Bowen Shi 0002, Vedanuj Goswami, Wei-Ning Hsu, Juan Pino 0001, Changhan Wang |
INTERSPEECH | 2 |
| 2023 | Expresso: A Benchmark and Analysis of Discrete Expressive Speech ResynthesisabstractInternational audience Tu Anh Nguyen, Wei-Ning Hsu, Antony D'Avirro, Bowen Shi 0002, Itai Gat, Maryam Fazel-Zarandi, Tal Remez, Jade Copet, Gabriel Synnaeve, Michael Hassid, Felix Kreuk, Yossi Adi, Emmanuel Dupoux |
INTERSPEECH | 4 |
| 2023 | Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleabstractLarge-scale generative models such as GPT and DALL-E have revolutionized the research community. These models not only generate high fidelity outputs, but are also generalists which can solve tasks not explicitly taught. In contrast, speech generative models are still primitive in terms of scale and task generalization. In this paper, we present Voicebox, the most versatile text-guided generative model for speech at scale. Voicebox is a non-autoregressive flow-matching model trained to infill speech, given audio context and text, trained on over 50K hours of speech that are not filtered or enhanced. Similar to GPT, Voicebox can perform many different tasks through in-context learning, but is more flexible as it can also condition on future context. Voicebox can be used for mono or cross-lingual zero-shot text-to-speech synthesis, noise removal, content editing, style conversion, and diverse sample generation. In particular, Voicebox outperforms the state-of-the-art zero-shot TTS model VALL-E on both intelligibility (5.9\% vs 1.9\% word error rates) and audio similarity (0.580 vs 0.681) while being up to 20 times faster. Audio samples can be found in \url{https://voicebox.metademolab.com}. Matt Le 0001, Apoorv Vyas, Bowen Shi 0002, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, Wei-Ning Hsu |
NeurIPS | 3 |
| 2022 | Searching for fingerspelled content in American Sign LanguageabstractNatural language processing for sign language video-including tasks like recognition, translation, and search-is crucial for making artificial intelligence technologies accessible to deaf individuals, and is gaining research interest in recent years.In this paper, we address the problem of searching for fingerspelled keywords or key phrases in raw sign language videos.This is an important task since significant content in sign language is often conveyed via fingerspelling, and to our knowledge the task has not been studied before.We propose an end-to-end model for this task, FSS-Net, that jointly detects fingerspelling and matches it to a text sequence.Our experiments, done on a large public dataset of ASL fingerspelling in the wild, show the importance of fingerspelling detection as a component of a search and retrieval model.Our model significantly outperforms baseline methods adapted from prior work on related tasks. Bowen Shi 0002, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
ACL (1) | 1 |
| 2022 | Open-Domain Sign Language Translation Learned from Online VideoabstractExisting work on sign language translationthat is, translation from sign language videos into sentences in a written language-has focused mainly on (1) data collected in a controlled environment or (2) data in a specific domain, which limits the applicability to realworld settings.In this paper, we introduce Ope-nASL, a large-scale American Sign Language (ASL) -English dataset collected from online video sites (e.g., YouTube).OpenASL contains 288 hours of ASL videos in multiple domains from over 200 signers and is the largest publicly available ASL translation dataset to date.To tackle the challenges of sign language translation in realistic settings and without glosses, we propose a set of techniques including sign search as a pretext task for pre-training and fusion of mouthing and handshape features.The proposed techniques produce consistent and large improvements in translation quality, over baseline models based on prior work. 1 Bowen Shi 0002, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
EMNLP | 1 |
| 2022 | Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Bowen Shi 0002, Wei-Ning Hsu, Kushal Lakhotia, Abdel-rahman Mohamed |
ICLR | 1 |
| 2022 | Robust Self-Supervised Audio-Visual Speech RecognitionabstractAudio-based automatic speech recognition (ASR) degrades significantly in noisy environments and is particularly vulnerable to interfering speech, as the model cannot determine which speaker to transcribe.Audio-visual speech recognition (AVSR) systems improve robustness by complementing the audio stream with the visual information that is invariant to noise and helps the model focus on the desired speaker.However, previous AVSR work focused solely on the supervised learning setup; hence the progress was hindered by the amount of labeled data available.In this work, we present a self-supervised AVSR framework built upon Audio-Visual HuBERT (AV-HuBERT), a state-of-theart audio-visual speech representation learning model.On the largest available AVSR benchmark dataset LRS3, our approach outperforms prior state-of-the-art by ∼ 50% (28.0% vs. 14.1%) using less than 10% of labeled data (433hr vs. 30hr) in the presence of babble noise, while reducing the WER of an audio-based model by over 75% (25.8% vs. 5.8%) on average 1 . Bowen Shi 0002, Wei-Ning Hsu, Abdel-rahman Mohamed |
INTERSPEECH | 1 |
| 2022 | Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERTabstractThis paper investigates self-supervised pre-training for audiovisual speaker representation learning where a visual stream showing the speaker's mouth area is used alongside speech as inputs.Our study focuses on the Audio-Visual Hidden Unit BERT (AV-HuBERT) approach, a recently developed generalpurpose audio-visual speech pre-training framework.We conducted extensive experiments probing the effectiveness of pretraining and visual modality.Experimental results suggest that AV-HuBERT generalizes decently to speaker related downstream tasks, improving label efficiency by roughly ten fold for both audio-only and audio-visual speaker verification.We also show that incorporating visual information, even just the lip area, greatly improves the performance and noise robustness, reducing EER by 38% in the clean condition and 75% in noisy conditions 1 . Bowen Shi 0002, Abdel-rahman Mohamed, Wei-Ning Hsu |
INTERSPEECH | 1 |
| 2022 | u-HuBERT: Unified Mixed-Modal Speech Pretraining And Zero-Shot Transfer to Unlabeled ModalityabstractWhile audio-visual speech models can yield superior performance and robustness compared to audio-only models, their development and adoption are hindered by the lack of labeled and unlabeled audio-visual data and the cost to deploy one model per modality. In this paper, we present u-HuBERT, a self-supervised pre-training framework that can leverage both multimodal and unimodal speech with a unified masked cluster prediction objective. By utilizing modality dropout during pre-training, we demonstrate that a single fine-tuned model can achieve performance on par or better than the state-of-the-art modality-specific models. Moreover, our model fine-tuned only on audio can perform well with audio-visual and visual speech input, achieving zero-shot modality generalization for multiple speech processing tasks. In particular, our single model yields 1.2%/1.4%/27.2% speech recognition word error rate on LRS3 with audio-visual/audio/visual input. Wei-Ning Hsu, Bowen Shi 0002 |
NeurIPS | 2 |
| 2021 | Fingerspelling Detection in American Sign LanguageabstractFingerspelling, in which words are signed letter by letter, is an important component of American Sign Language. Most previous work on automatic fingerspelling recognition has assumed that the boundaries of fingerspelling regions in signing videos are known beforehand. In this paper, we consider the task of fingerspelling detection in raw, untrimmed sign language videos. This is an important step towards building real-world fingerspelling recognition systems. We propose a benchmark and a suite of evaluation metrics, some of which reflect the effect of detection on the downstream fingerspelling recognition task. In addition, we propose a new model that learns to detect fingerspelling via multi-task training, incorporating pose estimation and fingerspelling recognition (transcription) along with detection, and compare this model to several alternatives. The model outperforms all alternative approaches across all metrics, establishing a state of the art on the benchmark. Bowen Shi 0002, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
CVPR | 1 |
| 2021 | Whole-Word Segmental Speech Recognition with Acoustic Word EmbeddingsabstractSegmental models are sequence prediction models in which scores of hypotheses are based on entire variable-length segments of frames. We consider segmental models for whole-word ("acoustic-to-word") speech recognition, with the feature vectors defined using vector embeddings of segments. Such models are computationally challenging as the number of paths is proportional to the vocabulary size, which can be orders of magnitude larger than when using subword units like phones. We describe an efficient approach for end-to-end whole-word segmental models, with forward-backward and Viterbi decoding performed on a GPU and a simple segment scoring function that reduces space complexity. In addition, we investigate the use of pre-training via jointly trained acoustic word embeddings (AWEs) and acoustically grounded word embeddings (AGWEs) of written word labels. We find that word error rate can be reduced by a large margin by pre-training the acoustic segment representation with AWEs, and additional (smaller) gains can be obtained by pre-training the word prediction layer with AGWEs. Our final models improve over prior A2W models. Bowen Shi 0002, Shane Settle, Karen Livescu |
SLT | 1 |
| 2020 | Few-Shot Acoustic Event Detection Via Meta LearningabstractWe study few-shot acoustic event detection (AED) in this paper. Few-shot learning enables detection of new events with very limited labeled data. Compared to other research areas like computer vision, few-shot learning for audio recognition has been under-studied. We formulate few-shot AED problem and explore different ways of utilizing traditional supervised methods for this setting as well as a variety of meta-learning approaches, which are conventionally used to solve few-shot classification problem. Compared to supervised baselines, meta-learning models achieve superior performance, thus showing its effectiveness on generalization to new audio events. Our analysis including impact of initialization and domain discrepancy further validate the advantage of meta-learning approaches in few-shot AED. Bowen Shi 0002, Ming Sun 0007, Krishna C. Puvvada, Chieh-Chi Kao, Spyridon Matsoukas, Chao Wang 0018 |
ICASSP | 1 |
| 2020 | A Joint Framework for Audio Tagging and Weakly Supervised Acoustic Event Detection Using DenseNet with Global Average PoolingabstractThis paper proposes a network architecture mainly designed for audio tagging, which can also be used for weakly supervised acoustic event detection (AED).The proposed network consists of a modified DenseNet as the feature extractor, and a global average pooling (GAP) layer to predict frame-level labels at inference time.This architecture is inspired by the work proposed by Zhou et al., a well-known framework using GAP to localize visual objects given image-level labels.While most of the previous works on weakly supervised AED used recurrent layers with attention-based mechanism to localize acoustic events, the proposed network directly localizes events using the feature map extracted by DenseNet without any recurrent layers.In the audio tagging task of DCASE 2017, our method significantly outperforms the state-of-the-art method in F1 score by 5.3% on the dev set, and 6.0% on the eval set in terms of absolute values.For weakly supervised AED task in DCASE 2018, our model outperforms the state-of-the-art method in event-based F1 by 8.1% on the dev set, and 0.5% on the eval set in terms of absolute values, by using data augmentation and tri-training to leverage unlabeled data. Chieh-Chi Kao, Bowen Shi 0002, Ming Sun 0007, Chao Wang 0018 |
INTERSPEECH | 2 |
| 2019 | Semi-supervised Acoustic Event Detection Based on Tri-trainingabstractThis paper presents our work of training acoustic event detection (AED) models using unlabeled dataset. Recent acoustic event detectors are based on large-scale neural networks, which are typically trained with huge amounts of labeled data. Labels for acoustic events are expensive to obtain, and relevant acoustic event audios can be limited, especially for rare events. In this paper we leverage an Internet-scale un-labeled dataset with potential domain shift to improve the detection of acoustic events. Based on the classic tri-training approach, our proposed method shows accuracy improvement over both the supervised training baseline, and semi-supervised self-training set-up, in all pre-defined acoustic event detection tasks. As our approach relies on ensemble models, we further show the improvements can be distilled to a single model via knowledge distillation, with the resulting single student model maintaining high accuracy of teacher ensemble models. Bowen Shi 0002, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018 |
ICASSP | 1 |
| 2019 | Fingerspelling Recognition in the Wild With Iterative Visual AttentionabstractSign language recognition is a challenging gesture sequence recognition problem, characterized by quick and highly coarticulated motion. In this paper we focus on recognition of fingerspelling sequences in American Sign Language (ASL) videos collected in the wild, mainly from YouTube and Deaf social media. Most previous work on sign language recognition has focused on controlled settings where the data is recorded in a studio environment and the number of signers is limited. Our work aims to address the challenges of real-life data, reducing the need for detection or segmentation modules commonly used in this domain. We propose an end-to-end model based on an iterative attention mechanism, without explicit hand detection or segmentation. Our approach dynamically focuses on increasingly high-resolution regions of interest. It out-performs prior work by a large margin. We also introduce a newly collected data set of crowdsourced annotations of fingerspelling in the wild, and show that performance can be further improved with this additional data set. Bowen Shi 0002, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
ICCV | 1 |
| 2019 | On the Contributions of Visual and Textual Supervision in Low-Resource Semantic Speech RetrievalabstractRecent work has shown that speech paired with images can be used to learn semantically meaningful speech representations even without any textual supervision. In real-world low-resource settings, however, we often have access to some transcribed speech. We study whether and how visual grounding is useful in the presence of varying amounts of textual supervision. In particular, we consider the task of semantic speech retrieval in a low-resource setting. We use a previously studied data set and task, where models are trained on images with spoken captions and evaluated on human judgments of semantic relevance. We propose a multitask learning approach to leverage both visual and textual modalities, with visual supervision in the form of keyword probabilities from an external tagger. We find that visual grounding is helpful even in the presence of textual supervision, and we analyze this effect over a range of sizes of transcribed data sets. With ~5 hours of transcribed speech, we obtain 23% higher average precision when also using visual supervision. Ankita Pasad, Bowen Shi 0002, Herman Kamper, Karen Livescu |
INTERSPEECH | 2 |
| 2019 | Compression of Acoustic Event Detection Models with Quantized DistillationabstractAcoustic Event Detection (AED), aiming at detecting categories of events based on audio signals, has found application in many intelligent systems. Recently deep neural network significantly advances this field and reduces detection errors to a large scale. However how to efficiently execute deep models in AED has received much less attention. Meanwhile state-of-the-art AED models are based on large deep models, which are computational demanding and challenging to deploy on devices with constrained computational resources. In this paper, we present a simple yet effective compression approach which jointly leverages knowledge distillation and quantization to compress larger network (teacher model) into compact network (student model). Experimental results show proposed technique not only lowers error rate of original compact network by 15% through distillation but also further reduces its model size to a large extent (2% of teacher, 12% of full-precision student) through quantization. Bowen Shi 0002, Ming Sun 0007, Chieh-Chi Kao, Viktor Rozgic, Spyridon Matsoukas, Chao Wang 0018 |
INTERSPEECH | 1 |
| 2018 | American Sign Language Fingerspelling Recognition in the WildabstractWe address the problem of American Sign Language fingerspelling recognition “in the wild”, using videos collected from websites. We introduce the largest data set available so far for the problem of fingerspelling recognition, and the first using naturally occurring video data. Using this data set, we present the first attempt to recognize fingerspelling sequences in this challenging setting. Unlike prior work, our video data is extremely challenging due to low frame rates and visual variability. To tackle the visual challenges, we train a special-purpose signing hand detector using a small subset of our data. Given the hand detector output, a sequence model decodes the hypothesized fingerspelled letter sequence. For the sequence model, we explore attention-based recurrent encoder-decoders and CTC-based approaches. As the first attempt at fingerspelling recognition in the wild, this work is intended to serve as a baseline for future work on sign language recognition in realistic conditions. We find that, as expected, letter error rates are much higher than in previous work on more controlled data, and we analyze the sources of error and effects of model variants. Bowen Shi 0002, Aurora Martinez Del Rio, Jonathan Keane, Jonathan Michaux, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
SLT | 1 |
| 2017 | Multitask training with unlabeled data for end-to-end sign language fingerspelling recognitionabstractWe address the problem of automatic American Sign Language fingerspelling recognition from video. Prior work has largely relied on frame-level labels, hand-crafted features, or other constraints, and has been hampered by the scarcity of data for this task. We introduce a model for fingerspelling recognition that addresses these issues. The model consists of an auto-encoder-based feature extractor and an attention-based neural encoder-decoder, which are trained jointly. The model receives a sequence of image frames and outputs the fingerspelled word, without relying on any frame-level training labels or hand-crafted features. In addition, the auto-encoder subcomponent makes it possible to leverage unlabeled data to improve the feature learning. The model achieves 11.6% and 4.4% absolute letter accuracy improvement respectively in signer-independent and signer-adapted fingerspelling recognition over previous approaches that required frame-level training labels. Bowen Shi 0002, Karen Livescu |
ASRU | 1 |