VLDB 2026 Research / reviewers in the wild / expert
Jennifer Williams 0001
dblp:99/817-1
· DBLP profile ↗
14ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0003-1410-0427ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Public perceptions of speech technology trust in the United KingdomabstractSpeech technology is now pervasive throughout the world, impacting a variety of socio-technical use-cases. Speech technology is a broad term encompassing capabilities that translate, analyse, transcribe, generate, modify, enhance, or summarise human speech. Many of the technical features and the possibility of speech data misuse are not often revealed to the users of such systems. When combined with the rapid development of AI and the plethora of use-cases where speech-based AI systems are now being applied, the consequence is that researchers, regulators, designers and government policymakers still have little understanding of the public’s perception of speech technology. Our research explores the public’s perceptions of trust in speech technology by asking people about their experiences, awareness of their rights, their susceptibility to being harmed, their expected behaviour, and ethical choices governing behavioural responsibility. We adopt a multidisciplinary lens to our work, in order to present a fuller picture of the United Kingdom (UK) public perspective through a series of socio-technical scenarios in a large-scale survey. We analysed survey responses from 1,000 participants from the UK, where people from different walks of life were asked to reflect on existing, emerging, and hypothetical speech technologies. Our socio-technical scenarios are designed to provoke and stimulate debate and discussion on principles of trust, privacy, responsibility, fairness, and transparency. We found that gender is a statistically significant factor correlated to awareness of rights and trust. We also found that awareness of rights is statistically correlated to perceptions of trust and responsible use of speech technology. By understanding the notions of responsibility in behaviour and differing perspectives of trust, our work encapsulates the current state of public acceptance of speech technology in the UK. Such an understanding has the potential to affect how regulatory and policy frameworks are developed, how the UK invests in its AI research and development ecosystem, and how speech technology that is developed within the UK might be received by global stakeholders. Jennifer Williams 0001, Tayyaba Azim, Anna-Maria Piskopani, Richard Hyde, Zack Hodari |
Comput. Speech Lang. | 1 |
| 2024 | Anonymizing Speaker Voices: Easy to Imitate, Difficult to Recognize?abstractA vastly under-explored area in speech anonymization involves characterizing how different speakers perform in voice privacy tasks. In this paper, we present a deeper analysis by creating and analyzing groups of challenging speakers categorized based on their performance in two related facets of voice anonymization evaluation: (1) speaker similarity using automatic speaker verification (ASV) and (2) human perception using a large-scale A/B listening test. We group speakers into four categories (sheep, goats, lambs, and wolves) based on their anonymization properties. We present an extension of voice anonymization evaluation by identifying speakers who are easy to imitate or difficult to recognize. This knowledge is important for trustworthy anonymization evaluation, and it has the potential to influence how evaluation datasets are created from a pool of speakers. We provide further insights on speaker influence on anonymized speech between human perception and automatic speaker similarity scoring. Jennifer Williams 0001, Karla Pizzi, Natalia A. Tomashenko, Sneha Das |
ICASSP | 1 |
| 2024 | Predicting Acute Pain Levels Implicitly from Vocal FeaturesabstractEvaluating pain in speech represents a critical challenge in high-stakes clinical scenarios, from analgesia delivery to emergency triage. Clinicians have predominantly relied on direct verbal communication of pain which is difficult for patients with communication barriers, such as those affected by stroke, autism, and learning difficulties. Many previous efforts have focused on multimodal data which does not suit all clinical applications. Our work is the first to collect a new English speech dataset wherein we have induced acute pain in adults using a cold pres-sor task protocol and recorded subjects reading sentences out loud. We report pain discrimination performance as F1 scores from binary (pain vs. no pain) and three-class (mild, moderate , severe) prediction tasks, and support our results with ex-plainable feature analysis. Our work is a step towards providing medical decision support for pain evaluation from speech to improve care across diverse and remote healthcare settings. Jennifer Williams 0001, Eike Schneiders, Henry Card, Tina Seabrooke, Beatrice Pakenham-Walsh, Tayyaba Azim, Lucy Valls-Reed, Ganesh Vigneswaran, John Robert Bautista, Rohan Chandra, Arya Farahi |
INTERSPEECH | 1 |
| 2024 | A New Approach to Voice Authenticityabstract2245 Nicolas M. Müller, Piotr Kawa, Shen Hu, Matthias Neu, Jennifer Williams 0001, Philip Sperl, Konstantin Böttinger |
INTERSPEECH | 5 |
| 2023 | Protecting Publicly Available Data With Machine Learning Shortcuts
Nicolas M. Müller, Maximilian Burgert, Pascal Debus, Jennifer Williams 0001, Philip Sperl, Konstantin Böttinger |
BMVC | 4 |
| 2023 | Privacy-Preserving Occupancy EstimationabstractIn this paper, we introduce an audio-based framework for occupancy estimation, including a new public dataset, and evaluate occupancy in a ‘cocktail party’ scenario where the party is simulated by mixing audio to produce speech with overlapping talkers (1-10 people). To estimate the number of speakers in an audio clip, we explored five different types of speech signal features and trained several versions of our model using convolutional neural networks (CNNs). Further, we adapted the framework to be privacy-preserving by making random perturbations of audio frames in order to conceal speech content and speaker identity. We show that some of our privacy-preserving features perform better at occupancy estimation than original waveforms. We analyse privacy further using two adversarial tasks: speaker recognition and speech recognition. Our privacy-preserving models can estimate the number of speakers in the simulated cocktail party clips within 1-2 persons based on a mean-square error (MSE) of 0.9-1.6 and we achieve up to 34.9% classification accuracy while preserving speech content privacy. However, it is still possible for an attacker to identify individual speakers, which motivates further work in this area. Jennifer Williams 0001, Vahid Yazdanpanah, Sebastian Stein 0001 |
ICASSP | 1 |
| 2022 | Attacker Attribution of Audio Deepfakesabstract2788 Nicolas M. Müller, Franziska Dieckmann, Jennifer Williams 0001 |
INTERSPEECH | 3 |
| 2021 | Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE ParadigmabstractWe present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not generalize well to unseen speakers or content. To alleviate this problem, we have incorporated a speaker encoder and speaker VQ codebook that learns global speaker characteristics entirely separate from the existing sub-phone codebooks. We also compare two training methods: self-supervised with global conditions and semi-supervised with speaker labels. Adding a speaker VQ component improves objective measures of speech synthesis quality (estimated MOS, speaker similarity, ASR-based intelligibility) and provides learned representations that are meaningful. Our speaker VQ codebook indices can be used in a simple speaker diarization task and perform slightly better than an x-vector baseline. Additionally, phones can be recognized from sub-phone VQ codebook indices in our semi-supervised VQ-VAE better than self-supervised with global conditions. Jennifer Williams 0001, Yi Zhao 0006, Erica Cooper, Junichi Yamagishi |
ICASSP | 1 |
| 2020 | End-to-End Signal Factorization for Speech: Identity, Content, and StyleabstractPreliminary experiments in this dissertation show that it is possible to factorize specific types of information from the speech signal in an abstract embedding space using machine learning. This information includes characteristics of the recording environment, speaking style, and speech quality. Based on these findings, a new technique is proposed to factorize multiple types of information from the speech signal simultaneously using a combination of state-of-the-art machine learning methods for speech processing. Successful speech signal factorization will lead to advances across many speech technologies, including improved speaker identification, detection of speech audio deep fakes, and controllable expression in speech synthesis. Jennifer Williams 0001 |
IJCAI | 1 |
| 2020 | Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform ReconstructionabstractVector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a speech signal without supervision. Until now, the VQ-VAE architecture has previously modeled individual types of speech features, such as only phones or only F0. This paper introduces an important extension to VQ-VAE for learning F0-related suprasegmental information simultaneously along with traditional phone features.The proposed framework uses two encoders such that the F0 trajectory and speech waveform are both input to the system, therefore two separate codebooks are learned. We used a WaveRNN vocoder as the decoder component of VQ-VAE. Our speaker-independent VQ-VAE was trained with raw speech waveforms from multi-speaker Japanese speech databases. Experimental results show that the proposed extension reduces F0 distortion of reconstructed speech for all unseen test speakers, and results in significantly higher preference scores from a listening test. We additionally conducted experiments using single-speaker Mandarin speech to demonstrate advantages of our architecture in another language which relies heavily on F0. Yi Zhao 0006, Cheng-I Lai, Jennifer Williams 0001, Erica Cooper, Junichi Yamagishi |
INTERSPEECH | 4 |
| 2020 | An Unsupervised Method to Select a Speaker Subset from Large Multi-Speaker Speech Synthesis DatasetsabstractLarge multi-speaker datasets for TTS typically contain diverse speakers, recording conditions, styles and quality of data. Although one might generally presume that more data is better, in this paper we show that a model trained on a carefully-chosen subset of speakers from LibriTTS provides significantly better quality synthetic speech than a model trained on a larger set. We propose an unsupervised methodology to find this subset by clustering per-speaker acoustic representations. Pilar Oplustil, Jennifer Williams 0001, Joanna Rownicka, Simon King 0001 |
INTERSPEECH | 2 |
| 2019 | Disentangling Style Factors from Speaker RepresentationsabstractOur goal is to separate out speaking style from speaker identity in utterance-level representations of speech such as i-vectors and x-vectors. We first show that both i-vectors and x-vectors contain information not only about speaker but also about speaking style (for one data set) or emotion (for another data set), even when projected into a low-dimensional space. To disentangle these factors, we use an autoencoder in which the latent space is split into two subspaces. The entangled information about speaker and style/emotion is pushed apart by the use of auxiliary classifiers that take one of the two latent subspaces as input and that are jointly learned with the autoencoder. We evaluate how well the latent subspaces separate the factors by using them as input to separate style/emotion classification tasks. In traditional speaker identification tasks, speaker-invariant characteristics are factorized from channel and then the channel information is ignored. Our results suggest that this so-called channel may contain exploitable information, which we refer to as style factors. Finally, we propose future work to use information theory to formalize style factors in the context of speaker identity. Jennifer Williams 0001, Simon King 0001 |
INTERSPEECH | 1 |
| 2019 | Speech Replay Detection with x-Vector Attack Embeddings and Spectral FeaturesabstractWe present our system submission to the ASVspoof 2019 Challenge Physical Access (PA) task. The objective for this challenge was to develop a countermeasure that identifies speech audio as either bona fide or intercepted and replayed. The target prediction was a value indicating that a speech segment was bona fide (positive values) or “spoofed” (negative values). Our system used convolutional neural networks (CNNs) and a representation of the speech audio that combined x-vector attack embeddings with signal processing features. The x-vector attack embeddings were created from mel-frequency cepstral coefficients (MFCCs) using a time-delay neural network (TDNN). These embeddings jointly modeled 27 different environments and 9 types of attacks from the labeled data. We also used sub-band spectral centroid magnitude coefficients (SCMCs) as features. We included an additive Gaussian noise layer during training as a way to augment the data to make our system more robust to previously unseen attack examples. We report system performance using the tandem detection cost function (tDCF) and equal error rate (EER). Our approach performed better that both of the challenge baselines. Our technique suggests that our x-vector attack embeddings can help regularize the CNN predictions even when environments or attacks are more challenging. Jennifer Williams 0001, Joanna Rownicka |
INTERSPEECH | 1 |
| 2014 | Finding Good Enough: A Task-Based Evaluation of Query Biased Summarization for Cross-Language Information RetrievalabstractIn this paper we present our task-based evaluation of query biased summarization for cross-language information retrieval (CLIR) using relevance prediction. We de-scribe our 13 summarization methods each from one of four summarization strate-gies. We show how well our methods perform using Farsi text from the CLEF 2008 shared-task, which we translated to English automtatically. We report preci-sion/recall/F1, accuracy and time-on-task. We found that different summarization methods perform optimally for different evaluation metrics, but overall query bi-ased word clouds are the best summariza-tion strategy. In our analysis, we demon-strate that using the ROUGE metric on our sentence-based summaries cannot make the same kinds of distinctions as our evalu-ation framework does. Finally, we present our recommendations for creating much-needed evaluation standards and datasets. 1 Jennifer Williams 0001, Sharon W. Tam, Wade Shen |
EMNLP | 1 |