Peter Wu

dblp:44/3072 · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 10 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
abstract
Modeling complex subjective tasks in Natural Language Processing, such as recognizing emotion and morality, is considerably challenging due to significant variation in human annotations.This variation often reflects reasonable differences in semantic interpretations rather than mere noise, necessitating methods to distinguish between legitimate subjectivity and error.We address this challenge by exploring label verification in these contexts using Large Language Models (LLMs).First, we propose a simple In-Context Learning binary filtering baseline that estimates the reasonableness of a document-label pair.We then introduce the Label-in-a-Haystack setting: the query and its label(s) are included in the demonstrations shown to LLMs, which are prompted to predict the label(s) again, while receiving task-specific instructions (e.g., emotion recognition) rather than label copying.We show how the failure to copy the label(s) to the output of the LLM are task-relevant and informative.Building on this, we propose the Label-in-a-Haystack Rectification (LiaHR) framework for subjective label correction: when the model outputs diverge from the reference gold labels, we assign the generated labels to the example instead of discarding it.This approach can be integrated into annotation pipelines to enhance signal-to-noise ratios.Comprehensive analyses, human evaluations, and ecological validity studies verify the utility of LiaHR for label correction.
Georgios Chochlakis, Peter Wu, Arjun Bedi, Marcus Ma, Kristina Lerman, Shri Narayanan
EMNLP2
2025 DiffMV-ETS: Diffusion-based Multi-Voice Electromyography-to-Speech Conversion using Speaker-Independent Speech Training Targets
abstract
Electromyography (EMG) signals have been investigated for novel voice prostheses to enable speech communication with silent articulation.In this work, we propose DiffMV-ETS, a multi-voice, diffusion-based EMG-to-speech system that converts EMG signals to speech in selectable voices.We evaluate it for scenarios where no speech of the speaker wearing EMG sensors is used for training.For this purpose, we introduce EMG-VCTK, a dataset containing EMG and audio recordings of sentences from the Voice Conversion Tool Kit corpus.We compare EMG models trained with audio of the same speaker, of auxiliary speakers, and of text-to-speech systems.Experiments indicate that models retain their intelligibility and naturalness when trained with synthetic speech.DiffMV-ETS enhances the speech naturalness and similarity to unseen voices.To the best of our knowledge, this is the first work to train multi-voice EMG-to-speech systems with speaker-independent targets.
Kevin Scheck, Tom Dombeck, Zhao Ren, Peter Wu, Michael Wand 0002, Tanja Schultz
INTERSPEECH4
2024 Towards an Interpretable Representation of Speaker Identity via Perceptual Voice Qualities
abstract
Unlike other data modalities such as text and vision, speech does not lend itself to easy interpretation. While lay people can understand how to describe an image or sentence via perception, non-expert descriptions of speech often end at high-level demographic information, such as gender or age. In this paper, we propose a possible interpretable representation of speaker identity based on perceptual voice qualities (PQs). By adding gendered PQs to the pathology-focused Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) protocol, our PQ-based approach provides a perceptual latent space of the character of adult voices that is an intermediary of abstraction between high-level demographics and low-level acoustic, physical, or learned representations. Contrary to prior belief, we demonstrate that these PQs are hearable by ensembles of non-experts, and further demonstrate that the information encoded in a PQ-based representation is predictable by various speech representations.
Robert Netzorg, Bohan Yu, Andrea Guzman, Peter Wu, Luna McNulty, Gopala Krishna Anumanchipalli
ICASSP4
2024 Multimodal Segmentation for Vocal Tract Modeling
Bohan Yu, Peter Wu, Tejas S. Prabhune, Gopala Krishna Anumanchipalli
INTERSPEECH3
2024 Towards EMG-to-Speech with Necklace Form Factor
Peter Wu, Ryan Kaveh, Raghav Nautiyal, Christine Zhang, Albert Guo, Anvitha Kachinthaya, Tavish Mishra, Bohan Yu, Alan W. Black, Rikky Muller, Gopala Krishna Anumanchipalli
INTERSPEECH1
2024 Topaz: Declarative and Verifiable Authoritative DNS at CDN-Scale
abstract
Today, when a CDN nameserver receives a DNS query for a customer's domain, it decides which CDN IP to return based on servicelevel objectives such as managing load or maintaining performance, but also internal needs like split testing. Many of these decisions are made a priori by assignment systems that imperatively generate maps from DNS query to IP address(es). Unfortunately, imperative assignments obfuscate nameserver behavior, especially when different objectives conflict.
James Larisch, Tim Alberdingk Thijm, Suleman Ahmad, Peter Wu, Tom Arnfeld, Marwan Fayed
SIGCOMM4
2024 Fast, High-Quality and Parameter-Efficient Articulatory Synthesis Using Differentiable DSP
abstract
Articulatory trajectories like electromagnetic articulography (EMA) provide a low-dimensional representation of the vocal tract filter and have been used as natural, grounded features for speech synthesis. Differentiable digital signal processing (DDSP) is a parameter-efficient framework for audio synthesis. Therefore, integrating low-dimensional EMA features with DDSP can significantly enhance the computational efficiency of speech synthesis. In this paper, we propose a fast, high-quality, and parameter-efficient DDSP articulatory vocoder that can synthesize speech from EMA, F0, and loudness. We incorporate several techniques to solve the harmonics / noise imbalance problem, and add a multiresolution adversarial loss for better synthesis quality. Our model achieves a transcription word error rate (WER) of 6.67% and a mean opinion score (MOS) of 3.74, with an improvement of 1.63% and 0.16 compared to the state-of-the-art (SOTA) baseline. Our DDSP vocoder is 4.9 x faster than the baseline on CPU during inference, and can generate speech of comparable quality with only 0.4 M parameters, in contrast to the 9 M parameters required by the SOTA.
Yisi Liu, Bohan Yu, Drake Lin, Peter Wu, Cheol Jun Cho, Gopala Krishna Anumanchipalli
SLT4
2023 Unconstrained Dysfluency Modeling for Dysfluent Speech Transcription and Detection
abstract
Dysfluent speech modeling requires time-accurate and silence-aware transcription at both the word-level and phonetic-level. However, current research in dysfluency modeling primarily focuses on either transcription or detection, and the performance of each aspect remains limited. In this work, we present an unconstrained dysfluency modeling (UDM) approach that addresses both transcription and detection in an automatic and hierarchical manner. UDM eliminates the need for extensive manual annotation by providing a comprehensive solution. Furthermore, we introduce a simulated dysfluent dataset called VCTK++to enhance the capabilities of UDM in phonetic transcription. Our experimental results demonstrate the effectiveness and robustness of our proposed methods in both transcription and detection tasks.
Jiachen Lian, Carly Feng, Naasir Farooqi, Steve Li, Anshul Kashyap, Cheol Jun Cho, Peter Wu, Robert Netzorg, Tingle Li, Gopala Krishna Anumanchipalli
ASRU7
2023 Articulation GAN: Unsupervised Modeling of Articulatory Learning
abstract
Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results in the production of speech sounds through physical properties of sound propagation. We introduce the Articulatory Generator to the Generative Adversarial Network paradigm, a new unsupervised generative model of speech production/synthesis. The Articulatory Generator more closely mimics human speech production by learning to generate articulatory representations (electromagnetic articulography or EMA) in a fully unsupervised manner. A separate pre-trained physical model (ema2wav) then transforms the generated EMA representations to speech waveforms, which get sent to the Discriminator for evaluation. Articulatory analysis suggests that the network learns to control articulators in a similar manner to humans during speech production. Acoustic analysis of the outputs suggests that the network learns to generate words that are both present and absent in the training distribution. We additionally discuss implications of articulatory representations for cognitive models of human language and speech technology in general.
Gasper Begus, Alan Zhou, Peter Wu, Gopala Krishna Anumanchipalli
ICASSP3
2023 Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech
abstract
Recent self-supervised learning (SSL) models have proven to learn rich representations of speech, which can readily be utilized by diverse downstream tasks. To understand such utilities, various analyses have been done for speech SSL models to reveal which and how information is encoded in the learned representations. Although the scope of previous analyses is extensive in acoustic, phonetic, and semantic perspectives, the physical grounding by speech production has not yet received full attention. To bridge this gap, we conduct a comprehensive analysis to link speech representations to articulatory trajectories measured by electromagnetic articulography (EMA). Our analysis is based on a linear probing approach where we measure articulatory score as an average correlation of linear mapping to EMA. We analyze a set of SSL models selected from the leaderboard of the SUPERB benchmark [1] and perform further layer-wise analyses on two most successful models, Wav2Vec 2.0 [2] and HuBERT [3]. Surprisingly, representations from the recent speech SSL models are highly correlated with EMA traces (best: r =0.81), and only 5 minutes are sufficient to train a linear model with high performance (r =0.77). Our findings suggest that SSL models learn to align closely with continuous articulations, and provide a novel insight into speech SSL.
Cheol Jun Cho, Peter Wu, Abdel-rahman Mohamed, Gopala Krishna Anumanchipalli
ICASSP2
2023 A Fast and Accurate Pitch Estimation Algorithm Based on the Pseudo Wigner-Ville Distribution
abstract
Estimation of fundamental frequency (F0) in voiced segments of speech signals, also known as pitch tracking, plays a crucial role in pitch synchronous speech analysis, speech synthesis, and speech manipulation. In this paper, we capitalize on the high time and frequency resolution of the pseudo Wigner-Ville distribution (PWVD) and propose a new PWVD-based pitch estimation method. We devise an efficient algorithm to compute PWVD faster and use cepstrum-based pre-filtering to avoid cross-term interference. Evaluating our approach on databases with speech and electroglottograph (EGG) recordings yields a state-of-the-art mean absolute error (MAE) of around 4Hz. Our approach is also effective at voiced/unvoiced classification and handling sudden frequency changes.
Yisi Liu, Peter Wu, Alan W. Black, Gopala Krishna Anumanchipalli
ICASSP2
2023 Speaker-Independent Acoustic-to-Articulatory Speech Inversion
abstract
To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising inversion target, since this space captures the mechanics of speech production. To this end, we build an acoustic-to-articulatory inversion (AAI) model that leverages autoregression, adversarial training, and self supervision to generalize to unseen speakers. Our approach obtains 0.784 correlation on an electromagnetic articulography (EMA) dataset, improving the state-of-the-art by 12.5%. Additionally, we show the interpretability of these representations through directly com-paring the behavior of estimated representations with speech production behavior. Finally, we propose a resynthesis-based AAI evaluation metric that does not rely on articulatory labels, demonstrating its efficacy with an 18-speaker dataset.
Peter Wu, Cheol Jun Cho, Shinji Watanabe 0001, Louis Goldstein, Alan W. Black, Gopala Krishna Anumanchipalli
ICASSP1
2023 Deep Speech Synthesis from MRI-Based Articulatory Representations
Peter Wu, Tingle Li, Yijing Lu, Yubin Zhang, Jiachen Lian, Alan W. Black, Louis Goldstein, Shinji Watanabe 0001, Gopala Krishna Anumanchipalli
INTERSPEECH1
2022 PACS: A Dataset for Physical Audiovisual CommonSense Reasoning
Samuel Yu, Peter Wu, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency
ECCV (37)2
2022 Training Strategies for Automatic Song Writing: A Unified Framework Perspective
abstract
Automatic song writing (ASW) typically involves four tasks: lyric-to-lyric generation, melody-to-melody generation, lyric-to-melody generation, and melody-to-lyric generation. Previous works have mainly focused on individual tasks without considering the correlation between them, and thus a unified framework to solve all four tasks has not yet been explored. In this paper, we propose a unified framework following the pre-training and fine-tuning paradigm to address all four ASW tasks with one model. To alleviate the data scarcity issue of paired lyric-melody data for lyric-to-melody and melody-to-lyric generation, we adopt two pre-training stages with unpaired data. In addition, we introduce a dual transformation loss to fully utilize paired data in the fine-tuning stage to enforce the weak correlation between melody and lyrics. We also design an objective music generation evaluation metric involving the chromatic rule and a more realistic setting, which removes some strict assumptions adopted in previous works. To the best of our knowledge, this work is the first to explore ASW for pop songs in Chinese. Extensive experiments demonstrate the effectiveness of the dual transformation loss and the unified model structure encompassing all four tasks. The experimental results also show that our proposed new evaluation metric aligns better with subjective opinion scores from human listeners.
Jiatong Shi, Peter Wu, Qin Jin
ICASSP4
2022 Respect the ORIGIN!: a best-case evaluation of connection coalescing in the wild
abstract
Connection coalescing, enabled by HTTP/2, permits a client to use an existing connection to request additional resources at the connected hostname. The potential for requests to be coalesced is hindered by the practice of domain sharding introduced by HTTP/1.1, because subresources are scattered across subdomains in an effort to improve performance with additional connections. When this happens, HTTP/2 clients invoke additional DNS queries and new connections to retrieve content that is available at the same server. ORIGIN Frame is an HTTP/2 extension that can be used by servers to inform clients about other domains that are reachable on the same connection. Despite being proposed by content delivery network (CDN) operators and standardized by the IETF in 2018, the extension has no known server implementation and is supported by only one browser. In this paper, we collect and characterize a large dataset. We use that dataset to model connection coalescing and identify a least-effort set of certificate changes that maximize opportunities for clients to coalesce. We then implemented and deployed ORIGIN Frame support at a large CDN. To evaluate and validate our modeling at scale, 5000 certificates were reissued. Passive measurements were conducted on production traffic over two weeks, during which we also actively measured on the 5000 domains.
Sudheesh Singanamalla, Muhammad Talha Paracha, Suleman Ahmad, Jonathan Hoyland, Luke Valenta, Yevgen Safronov, Peter Wu, Andrew Galloni, Kurtis Heimerl, Nick Sullivan, Christopher A. Wood, Marwan Fayed
IMC7
2022 Muskits: an End-to-end Music Processing Toolkit for Singing Voice Synthesis
abstract
This paper introduces a new open-source platform named Muskits for end-to-end music processing, which mainly focuses on end-to-end singing voice synthesis (E2E-SVS). Muskits supports state-of-the-art SVS models, including RNN SVS, transformer SVS, and XiaoiceSing. The design of Muskits follows the style of widely-used speech processing toolkits, ESPnet and Kaldi, for data prepossessing, training, and recipe pipelines. To the best of our knowledge, this toolkit is the first platform that allows a fair and highly-reproducible comparison between several published works in SVS. In addition, we also demonstrate several advanced usages based on the toolkit functionalities, including multilingual training and transfer learning. This paper describes the major framework of Muskits, its functionalities, and experimental results in single-singer, multi-singer, multilingual, and transfer learning scenarios. The toolkit is publicly available at https://github.com/SJTMusicTeam/Muskits.
Jiatong Shi, Tomoki Hayashi, Yuning Wu 0001, Fangzheng Xu, Xuankai Chang, Huazhe Li, Peter Wu, Shinji Watanabe 0001, Qin Jin
INTERSPEECH9
2022 Deep Speech Synthesis from Articulatory Representations
Peter Wu, Shinji Watanabe 0001, Louis Goldstein, Alan W. Black, Gopala Krishna Anumanchipalli
INTERSPEECH1
2021 Cross-Lingual Transfer for Speech Processing Using Acoustic Language Similarity
abstract
Speech processing systems currently do not support the vast majority of languages, in part due to the lack of data in low-resource languages. Cross-lingual transfer offers a compelling way to help bridge this digital divide by incorporating high-resource data into low-resource systems. Current cross-lingual algorithms have shown success in text-based tasks and speech-related tasks over some low-resource languages. However, scaling up speech systems to support hundreds of low-resource languages remains unsolved. To help bridge this gap, we propose a language similarity approach that can efficiently identify acoustic cross-lingual transfer pairs across hundreds of languages. We demonstrate the effectiveness of our approach in language family classification, speech recognition, and speech synthesis tasks.
Peter Wu, Jiatong Shi, Yifan Zhong, Shinji Watanabe 0001, Alan W. Black
ASRU1
2021 Cross-Modal Generalization: Learning in Low Resource Modalities via Meta-Alignment
abstract
How can we generalize to a new prediction task at test time when it also uses a new modality as input? More importantly, how can we do this with as little annotated data as possible? This problem of cross-modal generalization is a new research milestone with concrete impact on real-world applications. For example, can an AI system start understanding spoken language from mostly written text? Or can it learn the visual steps of a new recipe from only text descriptions? In this work, we formalize cross-modal generalization as a learning paradigm to train a model that can (1) quickly perform new tasks (from new domains) while (2) being originally trained on a different input modality. Such a learning paradigm is crucial for generalization to low-resource modalities such as spoken speech in rare languages while utilizing a different high-resource modality such as text. One key technical challenge that makes it different from other learning paradigms such as meta-learning and domain adaptation is the presence of different source and target modalities which will require different encoders. We propose an effective solution based on meta-alignment, a novel method to align representation spaces using strongly and weakly paired cross-modal data while ensuring quick generalization to new tasks across different modalities. This approach uses key ideas from cross-modal learning and meta-learning, and presents strong results on the cross-modal generalization problem. We benchmark several approaches on 3 real-world classification tasks: few-shot recipe classification from text to images of recipes, object classification from images to audio of objects, and language classification from text to spoken speech across 100 languages spanning many rare languages. Our results demonstrate strong performance even when the new target modality has only a few (1-10) labeled samples and in the presence of noisy labels, a scenario particularly prevalent in low-resource modalities.
Paul Pu Liang, Peter Wu, Liu Ziyin 0001, Louis-Philippe Morency, Ruslan Salakhutdinov
ACM Multimedia2
2021 Oblivious DNS over HTTPS (ODoH): A Practical Privacy Enhancement to DNS
abstract
Abstract The Internet’s Domain Name System (DNS) responds to client hostname queries with corresponding IP addresses and records. Traditional DNS is unencrypted and leaks user information to on-lookers. Recent efforts to secure DNS using DNS over TLS (DoT) and DNS over HTTPS (DoH) have been gaining traction, ostensibly protecting DNS messages from third parties. However, the small number of available public large-scale DoT and DoH resolvers has reinforced DNS privacy concerns, specifically that DNS operators could use query contents and client IP addresses to link activities with identities. Oblivious DNS over HTTPS (ODoH) safeguards against these problems. In this paper we implement and deploy interoperable instantiations of the protocol, construct a corresponding formal model and analysis, and evaluate the protocols’ performance with wide-scale measurements. Results suggest that ODoH is a practical privacy-enhancing replacement for DNS.
Sudheesh Singanamalla, Suphanat Chunhapanya, Jonathan Hoyland, Marek Vavrusa, Tanya Verma, Peter Wu, Marwan Fayed, Kurtis Heimerl, Nick Sullivan, Christopher A. Wood
Proc. Priv. Enhancing Technol.6
2019 Ordinal Triplet Loss: Investigating Sleepiness Detection from Speech
Peter Wu, Sai Krishna Rallabandi, Alan W. Black, Eric Nyberg
INTERSPEECH1
2014 Protein folding estimation using Paired-Bacteria Optimizer
abstract
Protein folding estimation attracts a large attention in the area of computational biology, due to its benefits on medical research and the challenge of NP-hard objective functions. In order to simulate the protein folding procedure and estimate the structure of the protein after folding, this paper adopts a Paired-Bacteria Optimizer (PBO), which is a biologically-inspired optimization algorithm. Compared with most Evolutionary Algorithms (EAs), the computational complexity of PBO is much less. Therefore, it is suitable to be applied to solve NP-hard problem. The experimental studies is performed on several benchmark lattice protein combination. The experimental results demonstrated that PBO is able to estimate the folded protein structure with a superior convergence.
Mengshi Li, Tianyao Ji, Peter Wu, Shan He 0001, Q. Henry Wu
IEEE Congress on Evolutionary Computation3
1996 Home-study software: flexible, interactive, and distributed software for independent study
abstract
Article Free Access Share on Home-study software: flexible, interactive, and distributed software for independent study Authors: Christopher Connelly Electrotechnical Laboratory, Umezono, 1-1-4,Tsukuba City, Ibaraki, Japan and Department of Computer Science, Duke University, Durham, NC Electrotechnical Laboratory, Umezono, 1-1-4,Tsukuba City, Ibaraki, Japan and Department of Computer Science, Duke University, Durham, NCView Profile , Alan W. Biermann Department of Computer Science, Duke University, Durham, NC Department of Computer Science, Duke University, Durham, NCView Profile , David Pennock Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor and Department of Computer Science, Duke University, Durham, NC Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor and Department of Computer Science, Duke University, Durham, NCView Profile , Peter Wu Microsoft Corporation, Menlo Park, CA and Department of Computer Science, Duke University, Durham, NC Microsoft Corporation, Menlo Park, CA and Department of Computer Science, Duke University, Durham, NCView Profile Authors Info & Claims SIGCSE '96: Proceedings of the twenty-seventh SIGCSE technical symposium on Computer science educationMarch 1996Pages 63–67https://doi.org/10.1145/236452.236509Published:01 March 1996Publication History 10citation218DownloadsMetricsTotal Citations10Total Downloads218Last 12 Months30Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Christopher Connelly, Alan W. Biermann, David M. Pennock, Peter Wu
SIGCSE4
1994 Teaching a hierarchical model of computation with animation software in the first course
abstract
In a world saturated with computers, it is important that the popnlace have some understanding of what thesedevices am, how they work, what they can do, and what they cannot do.People will not intelligently
Alan W. Biermann, Amr F. Fahmy, Curry I. Guinn, David M. Pennock, Dietolf Ramm, Peter Wu
SIGCSE6