Felix Wu

dblp:20/2322 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
16since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 11 since 2021Theory of computation · 4Computer networks · 2 · 1 first-author
YearPublicationVenuePosition
2024 Improving ASR Contextual Biasing with Guided Attention
abstract
In this paper, we propose a Guided Attention (GA) auxiliary training loss, which improves the effectiveness and robustness of automatic speech recognition (ASR) contextual biasing without introducing additional parameters. A common challenge in previous literature is that the word error rate (WER) reduction brought by contextual biasing diminishes as the number of bias phrases increases. To address this challenge, we employ a GA loss as an additional training objective besides the Transducer loss. The proposed GA loss aims to teach the cross attention how to align bias phrases with text tokens or audio frames. Compared to studies with similar motivations, the proposed loss operates directly on the cross attention weights and is easier to implement. Through extensive experiments based on Conformer Transducer with Contextual Adapter, we demonstrate that the proposed method not only leads to a lower WER but also retains its effectiveness as the number of bias phrases increases. Specifically, the GA loss decreases the WER of rare vocabularies by up to 19.2% on LibriSpeech compared to the contextual biasing baseline, and up to 49.3% compared to a vanilla Transducer.
Jiyang Tang, Kwangyoun Kim, Suwon Shon, Felix Wu, Prashant Sridhar
ICASSP4
2024 Sample-Efficient Diffusion for Text-To-Speech Synthesis
abstract
This work introduces Sample-Efficient Speech Diffusion (SESD), an algorithm for effective speech synthesis in modest data regimes through latent diffusion.It is based on a novel diffusion architecture, that we call U-Audio Transformer (U-AT), that efficiently scales to long sequences and operates in the latent space of a pre-trained audio autoencoder.Conditioned on character-aware language model representations, SESD achieves impressive results despite training on less than 1k hours of speech -far less than current state-of-the-art systems.In fact, it synthesizes more intelligible speech than the state-of-the-art auto-regressive model, VALL-E, while using less than 2% the training data.Our implementation is available at https://github.com/justinlovelace/SESD.
Justin Lovelace, Soham Ray, Kwangyoun Kim, Kilian Q. Weinberger, Felix Wu
INTERSPEECH5
2023 SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding Tasks
abstract
Suwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad, Felix Wu, Roshan S Sharma, Wei-Lun Wu, Hung-yi Lee, Karen Livescu, Shinji Watanabe. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Suwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad, Felix Wu, Roshan S. Sharma, Wei-Lun Wu, Hung-yi Lee, Karen Livescu, Shinji Watanabe 0001
ACL (1)5
2023 Structured Pruning of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding
abstract
Self-supervised speech representation learning (SSL) has shown to be effective in various downstream tasks, but SSL models are usually large and slow. Model compression techniques such as pruning aim to reduce the model size and computation without degradation in accuracy. Prior studies focus on the pruning of Transformers; however, speech models not only utilize a stack of Transformer blocks, but also combine a frontend network based on multiple convolutional layers for low-level feature representation learning. This frontend has a small size but a heavy computational cost. In this work, we propose three task-specific structured pruning methods to deal with such heterogeneous networks. Experiments on LibriSpeech and SLURP show that the proposed method is more accurate than the original wav2vec2-base with 10% to 30% less computation, and is able to reduce the computation by 40% to 50% without any degradation.
Yifan Peng 0003, Kwangyoun Kim, Felix Wu, Prashant Sridhar, Shinji Watanabe 0001
ICASSP3
2023 Context-Aware Fine-Tuning of Self-Supervised Speech Models
abstract
Self-supervised pre-trained transformers have improved the state of the art on a variety of speech tasks. Due to the quadratic time and space complexity of self-attention, they usually operate at the level of relatively short (e.g., utterance) segments. In this paper, we study the use of context, i.e., surrounding segments, during fine-tuning and propose a new approach called context-aware fine-tuning. We attach a context module on top of the last layer of a pre-trained model to encode the whole segment into a context embedding vector which is then used as an additional feature for the final prediction. During the fine-tuning stage, we introduce an auxiliary loss that encourages this context embedding vector to be similar to context vectors of surrounding segments. This allows the model to make predictions without access to these surrounding segments at inference time and requires only a tiny overhead compared to standard fine-tuned models. We evaluate the proposed approach using the SLUE and Librilight benchmarks for several downstream tasks: Automatic speech recognition (ASR), named entity recognition (NER), and sentiment analysis (SA). The results show that context-aware fine-tuning not only outperforms a standard fine-tuning baseline but also rivals a strong context injection baseline that uses neighboring speech segments during inference.
Suwon Shon, Felix Wu, Kwangyoun Kim, Prashant Sridhar, Karen Livescu, Shinji Watanabe 0001
ICASSP2
2023 Wav2Seq: Pre-Training Speech-to-Text Encoder-Decoder Models Using Pseudo Languages
abstract
We introduce Wav2Seq, the first self-supervised approach to pre-train both parts of encoder-decoder models for speech data. We induce a pseudo language as a compact discrete representation, and formulate a self-supervised pseudo speech recognition task — transcribing audio inputs into pseudo subword sequences. This process stands on its own, or can be applied as low-cost second-stage pre-training. We experiment with automatic speech recognition (ASR), spoken named entity recognition, and speech-to-text translation. We set new state-of-the-art results for end-to-end spoken named entity recognition, and show consistent improvements on 8 language pairs for speech-to-text translation, even when competing methods use additional text data for training. On ASR, our approach enables encoder-decoder methods to benefit from pre-training for all parts of the network, and shows comparable performance to highly optimized recent methods.
Felix Wu, Kwangyoun Kim, Shinji Watanabe 0001, Kyu Jeong Han, Ryan McDonald, Kilian Q. Weinberger, Yoav Artzi
ICASSP1
2023 On the Effectiveness of Offline RL for Dialogue Response Generation
abstract
A common training technique for language models is teacher forcing (TF). TF attempts to match human language exactly, even though identical meanings can be expressed in different ways. This motivates use of sequence-level objectives for dialogue response generation. In this paper, we study the efficacy of various offline reinforcement learning (RL) methods to maximize such objectives. We present a comprehensive evaluation across multiple datasets, models, and metrics. Offline RL shows a clear performance improvement over teacher forcing while not inducing training instability or sacrificing practical training budgets.
Paloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger, Ryan McDonald
ICML2
2023 A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks
Yifan Peng 0003, Kwangyoun Kim, Felix Wu, Brian Yan, Siddhant Arora, Jiyang Tang, Suwon Shon, Prashant Sridhar, Shinji Watanabe 0001
INTERSPEECH3
2022 SLUE: New Benchmark Tasks For Spoken Language Understanding Evaluation on Natural Speech
abstract
Progress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest has been growing in higher-level spoken language understanding tasks, including using end-to-end models, but there are fewer annotated datasets for such tasks. At the same time, recent work shows the possibility of pre-training generic representations and then fine-tuning for several tasks using relatively little labeled data. We propose to create a suite of benchmark tasks for Spoken Language Understanding Evaluation (SLUE) consisting of limited-size labeled training sets and corresponding evaluation sets. This resource would allow the research community to track progress, evaluate pre-trained representations for higher-level tasks, and study open questions such as the utility of pipeline versus end-to-end approaches. We present the first phase of the SLUE benchmark suite, consisting of named entity recognition, sentiment analysis, and ASR on the corresponding datasets. We focus on naturally produced (not read or synthesized) speech, and freely available datasets. We pro-vide new transcriptions and annotations on subsets of the VoxCeleb and VoxPopuli datasets, evaluation metrics and results for baseline models, and an open-source toolkit to reproduce the baselines and evaluate new models.
Suwon Shon, Ankita Pasad, Felix Wu, Pablo Brusco, Yoav Artzi, Karen Livescu, Kyu Jeong Han
ICASSP3
2022 Performance-Efficiency Trade-Offs in Unsupervised Pre-Training for Speech Recognition
abstract
This paper is a study of performance-efficiency trade-offs in pre-trained models for automatic speech recognition (ASR). We focus on wav2vec 2.0, and formalize several architecture designs that influence both the model performance and its efficiency. Putting together all our observations, we introduce SEW-D (Squeezed and Efficient Wav2vec with Disentangled Attention), a pre-trained model architecture with significant improvements along both performance and efficiency dimensions across a variety of training setups. For example, under the 100h-960h semi-supervised setup on LibriSpeech, SEW-D achieves a 1.9x inference speedup compared to wav2vec 2.0, with a 13.5% relative reduction in word error rate. With a similar inference time, SEW reduces word error rate by 25–50% across different model sizes.
Felix Wu, Kwangyoun Kim, Kyu Jeong Han, Kilian Q. Weinberger, Yoav Artzi
ICASSP1
2022 On the Use of External Data for Spoken Named Entity Recognition
abstract
Ankita Pasad, Felix Wu, Suwon Shon, Karen Livescu, Kyu Han. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ankita Pasad, Felix Wu, Suwon Shon, Karen Livescu, Kyu Jeong Han
NAACL-HLT2
2022 E-Branchformer: Branchformer with Enhanced Merging for Speech Recognition
abstract
Conformer, combining convolution and self-attention sequentially to capture both local and global information, has shown remarkable performance and is currently regarded as the state-of-the-art for automatic speech recognition (ASR). Several other studies have explored integrating convolution and self-attention but they have not managed to match Conformer's performance. The recently introduced Branchformer achieves comparable performance to Conformer by using dedicated branches of convolution and self-attention and merging local and global context from each branch. In this paper, we propose E-Branchformer, which enhances Branchformer by applying an effective merging method and stacking additional point-wise modules. E-Branchformer sets new state-of-the-art word error rates (WERs) 1.81% and 3.65% on LibriSpeech test-clean and test-other sets without using any external training data.
Kwangyoun Kim, Felix Wu, Yifan Peng 0003, Prashant Sridhar, Kyu Jeong Han, Shinji Watanabe 0001
SLT2
2021 On Feature Normalization and Data Augmentation
abstract
The moments (a.k.a., mean and standard deviation) of latent features are often removed as noise when training image recognition models, to increase stability and reduce training time. However, in the field of image generation, the moments play a much more central role. Studies have shown that the moments extracted from instance normalization and positional normalization can roughly capture style and shape information of an image. Instead of being discarded, these moments are instrumental to the generation process. In this paper we propose Moment Exchange, an implicit data augmentation method that encourages the model to utilize the moment information also for recognition models. Specifically, we replace the moments of the learned features of one training image by those of another, and also interpolate the target labels—forcing the model to extract training signal from the moments in addition to the normalized features. As our approach is fast, operates entirely in feature space, and mixes different signals than prior methods, one can effectively combine it with existing augmentation approaches. We demonstrate its efficacy across several recognition benchmark data sets where it improves the generalization capability of highly competitive baseline networks with remarkable consistency.
Boyi Li 0001, Felix Wu, Ser-Nam Lim, Serge J. Belongie, Kilian Q. Weinberger
CVPR2
2021 Revisiting Few-sample BERT Fine-tuning
Tianyi Zhang 0007, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger, Yoav Artzi
ICLR2
2021 Making Paper Reviewing Robust to Bid Manipulation Attacks
abstract
Most computer science conferences rely on paper bidding to assign reviewers to papers. Although paper bidding enables high-quality assignments in days of unprecedented submission numbers, it also opens the door for dishonest reviewers to adversarially influence paper reviewing assignments. Anecdotal evidence suggests that some reviewers bid on papers by "friends" or colluding authors, even though these papers are outside their area of expertise, and recommend them for acceptance without considering the merit of the work. In this paper, we study the efficacy of such bid manipulation attacks and find that, indeed, they can jeopardize the integrity of the review process. We develop a novel approach for paper bidding and assignment that is much more robust against such attacks. We show empirically that our approach provides robustness even when dishonest reviewers collude, have full knowledge of the assignment system’s internal workings, and have access to the system’s inputs. In addition to being more robust, the quality of our paper review assignments is comparable to that of current, non-robust assignment approaches.
Ruihan Wu, Chuan Guo 0001, Felix Wu, Rahul Kidambi, Laurens van der Maaten, Kilian Q. Weinberger
ICML3
2021 Multi-Mode Transformer Transducer with Stochastic Future Context
abstract
Automatic speech recognition (ASR) models make fewer errors when more surrounding speech information is presented as context.Unfortunately, acquiring a larger future context leads to higher latency.There exists an inevitable trade-off between speed and accuracy.Naïvely, to fit different latency requirements, people have to store multiple models and pick the best one under the constraints.Instead, a more desirable approach is to have a single model that can dynamically adjust its latency based on different constraints, which we refer to as Multimode ASR.A Multi-mode ASR model can fulfill various latency requirements during inference -when a larger latency becomes acceptable, the model can process longer future context to achieve higher accuracy and when a latency budget is not flexible, the model can be less dependent on future context but still achieve reliable accuracy.In pursuit of Multi-mode ASR, we propose Stochastic Future Context, a simple training procedure that samples one streaming configuration in each iteration.Through extensive experiments on AISHELL-1 and Lib-riSpeech datasets, we show that a Multi-mode ASR model rivals, if not surpasses, a set of competitive streaming baselines trained with different latency budgets.
Kwangyoun Kim, Felix Wu, Prashant Sridhar, Kyu Jeong Han, Shinji Watanabe 0001
Interspeech2
2020 BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang 0007, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi
ICLR3
2019 Pay Less Attention with Lightweight and Dynamic Convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N. Dauphin, Michael Auli
ICLR1
2019 Simplifying Graph Convolutional Networks
abstract
Graph Convolutional Networks (GCNs) and their variants have experienced significant attention and have become the de facto methods for learning graph representations. GCNs derive inspiration primarily from recent deep learning approaches, and as a result, may inherit unnecessary complexity and redundant computation. In this paper, we reduce this excess complexity through successively removing nonlinearities and collapsing weight matrices between consecutive layers. We theoretically analyze the resulting linear model and show that it corresponds to a fixed low-pass filter followed by a linear classifier. Notably, our experimental evaluation demonstrates that these simplifications do not negatively impact accuracy in many downstream applications. Moreover, the resulting model scales to larger datasets, is naturally interpretable, and yields up to two orders of magnitude speedup over FastGCN.
Felix Wu, Amauri H. Souza, Tianyi Zhang 0007, Christopher Fifty, Kilian Q. Weinberger
ICML1
2019 Positional Normalization
abstract
A widely deployed method for reducing the training time of deep neural networks is to normalize activations at each layer. Although various normalization schemes have been proposed, they all follow a common theme: normalize across spatial dimensions and discard the extracted statistics. In this paper, we propose a novel normalization method that deviates from this theme. Our approach, which we refer to as Positional Normalization (PONO), normalizes exclusively across channels, which allows us to capture structural information of the input image in the first and second moments. Instead of disregarding this information, we inject it into later layers to preserve or transfer structural information in generative networks. We show that PONO significantly improves the performance of deep networks across a wide range of model architectures and image generation tasks.
Boyi Li 0001, Felix Wu, Kilian Q. Weinberger, Serge J. Belongie
NeurIPS2
2019 Multi-Beam Multi-Stream Communications for 5G and beyond Mobile User Equipment and UAV Proof of Concept Designs
abstract
Millimeter-wave (mmWave), massive multiple-input multiple-output (MIMO), are expected to play a crucial role for 5G and beyond cellular and next-generation wireless local area network (WLAN) communications. Moreover, unmanned aerial vehicles (UAVs) are also considered as an important component of next-generation networks. In this paper, we propose and present a mmWave distributed phased-arrays (DPA) architecture and proof-of-concept (PoC) designs for user equipment (UE) and unmanned aerial vehicles (UAVs) which will be used in 5G/Beyond 5G wireless communication networks. Through enabling a multi-stream multi-beam communication mode, the UE PoC achieves a peak downlink speed of more than 4 Gbps with optimized thermal distribution performance. Furthermore, based on the DPA topology, the UAV aerial base station (ABS) prototype is designed and demonstrates for the first time an aggregated peak downlink data rate of 2.2 Gbps in the real-world field tests supporting multi-user (MU) application scenarios.
Yiming Huo, Franklin Lu, Felix Wu, Xiaodai Dong
VTC Fall3
2018 Multi-Scale Dense Networks for Resource Efficient Image Classification
Gao Huang 0001, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, Kilian Q. Weinberger
ICLR4
2017 On Fairness and Calibration
abstract
The machine learning community has become increasingly concerned with the potential for bias and discrimination in predictive models. This has motivated a growing line of work on what it means for a classification procedure to be "fair." In this paper, we investigate the tension between minimizing error disparity across different population groups while maintaining calibrated probability estimates. We show that calibration is compatible only with a single error constraint (i.e. equal false-negatives rates across groups), and show that any algorithm that satisfies this relaxation is no better than randomizing a percentage of predictions for an existing classifier. These unsettling findings, which extend and generalize existing results, are empirically confirmed on several datasets.
Geoff Pleiss, Manish Raghavan, Felix Wu, Jon M. Kleinberg, Kilian Q. Weinberger
NIPS3
2015 Combination of feature engineering and ranking models for paper-author identification in KDD cup 2013
Chun-Liang Li, Yu-Chuan Su, Ting-Wei Lin, Cheng-Hao Tsai, Wei-Cheng Chang, Kuan-Hao Huang, Tzu-Ming Kuo, Shan-Wei Lin, Young-San Lin, Yu-Chen Lu, Chun-Pai Yang, Cheng-Xia Chang, Wei-Sheng Chin, Yu-Chin Juan, Hsiao-Yu Fish Tung, Jui-Pin Wang, Cheng-Kuang Wei, Felix Wu, Tu-Chun Yin, Tong Yu 0001, Yong Zhuang, Shou-De Lin, Hsuan-Tien Lin, Chih-Jen Lin
J. Mach. Learn. Res.18
2014 Effective string processing and matching for author disambiguation
Wei-Sheng Chin, Yong Zhuang, Yu-Chin Juan, Felix Wu, Hsiao-Yu Fish Tung, Tong Yu 0001, Jui-Pin Wang, Cheng-Xia Chang, Chun-Pai Yang, Wei-Cheng Chang, Kuan-Hao Huang, Tzu-Ming Kuo, Shan-Wei Lin, Young-San Lin, Yu-Chen Lu, Yu-Chuan Su, Cheng-Kuang Wei, Tu-Chun Yin, Chun-Liang Li, Ting-Wei Lin, Cheng-Hao Tsai, Shou-De Lin, Hsuan-Tien Lin, Chih-Jen Lin
J. Mach. Learn. Res.4
2013 Flickr-tag prediction using multi-modal fusion and meta information
abstract
We present our evaluation and analysis on Yahoo! Large-scale Flickr-tag Image Classification dataset. Our evaluations show that combining multi-features and different classification models, the MAP of tag prediction can be significantly improve over ordinary linear classification. Further analysis shows that some tags are given not because of the visual content but the meta information of images. Our experiments show that we can make more accurate prediction on certain tags using meta information without any training process, compared with visual content based classifiers. Combine the meta information, multi-features and multi-models fusion, we achieve significantly better performance than simple linear classification. We also evaluate the performance of various mid-level feature, and the results suggest that "Concept Bank" feature may be a promising direction for the task.
Yu-Chuan Su, Tzu-Hsuan Chiu, Guan-Long Wu, Chun-Yen Yeh, Felix Wu, Winston H. Hsu
ACM Multimedia5
2005 Should we share honeypot information for security management?
Felix Wu
Integrated Network Management1
2004 Online learning in online auctions
Avrim Blum, Atri Rudra, Felix Wu
Theor. Comput. Sci.4
2003 Online learning in online auctions
Avrim Blum, Atri Rudra, Felix Wu
SODA4
2002 Incentive-compatible online auctions for digital goods
Ziv Bar-Yossef, Kirsten Hildrum, Felix Wu
SODA3
1999 The Quantum Query Complexity of Approximating the Median and Related Statistics
abstract
Let X = (z,, , z,-,) be a sequence of n numbers.For 6 > 0, we say that 5; is an e-approximate median if the number of elements strictly less than zi and the number of elements strictly greater than zi are each less than (1 + 6):.We consider the quantum query complexity of computing an c-approximate median, given the sequence X as an oracle.We prove a lower bound of n(min{t,n}) queries for any quantum algorithm that computes an r-approximate median with any constant probability greater than l/2.We also show how an c-approximate median may be computed with 0( $ log(t) log log( $)) oracle queries, which rep resents an improvement over an earlier algorithm due to Grover [ll, 121.Thus, the lower bound we obtain is essentially optimal.The upper and the lower bound both hold in the comparison tree model as well.Our lower bound result is an application of the polynomial paradigm recently introduced to quantum complexity theory by Be& et ol.[l].The main ingredient in the proof is a polynomial degree lower bound far real multilinear polynomials that "approximate" symmetric partial boolean functions.The degree bound extends a result of Patti[15] and also immediately yields lower bounds for the problems of approximating the kth-smallest element, approximating the mean of a sequence of numbers, and approximately counting the number of ones of a boolean function.All bounds obtained come within a polylogarithmic factor of the optimal (as we show by presenting algorithms where no such optimal or near optimal algorithms were known), thus demonstrating the power of the polynomial method.
Ashwin Nayak 0001, Felix Wu
STOC2
1994 Dynamic Bandwidth Allocation Using Infinitesimal Perturbation Analysis
abstract
Advances in network management and switching technologies make dynamic bandwidth allocation of logical networks built on top of a physical network possible. Previous proposed dynamic bandwidth allocation algorithms are based on simplified network model. The analytical model is valid only under restrictive assumptions. Infinitesimal perturbation analysis, a technique which estimates the gradients of the functions in discrete event dynamic systems by passively observing the system, is used to estimate delay sensitivities under general traffic patterns. A new dynamic bandwidth allocation algorithm using on-line sensitivity estimation is proposed. Simulation results show that the approach further improves network performance. Implementation of the proposed algorithm in operational networks is also discussed.>
Felix Wu, Shau-Ming Lun
INFOCOM2