VLDB 2026 Research / reviewers in the wild / expert
Paul Foulkes
dblp:157/7908
· DBLP profile ↗
13ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-9481-1004ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The discriminative capacity of English segments in forensic speaker comparison
Paul Foulkes, Vincent Hughes, Kayleigh Peters, Jasmine Rouse |
Speech Commun. | 1 |
| 2025 | Variability in performance across four generations of automatic speaker recognition systems
Lauren Harrington, Vincent Hughes, Philip Harrison, Paul Foulkes, Jessica Wormald, Finnian Kelly, David van der Vloed |
INTERSPEECH | 4 |
| 2025 | Evaluating the suitability of acoustic parameters for capturing breathy voice in non-pathological female speakersabstractThis study evaluates the suitability of acoustic parameters for capturing breathy voice in non-pathological female speakers. Much existing literature on the acoustic analysis of voice quality (VQ) focuses on pathological speakers and/or sustained vowel productions. The limited work on more naturalistic speech (e.g. as needed for forensic casework) focuses on male speakers. Using studio-quality recordings from the PASR dataset [1], nine phoneticians (6 male, 3 female) read a passage in their normal ('default') and breathy voice multiple times. Acoustic parameters (H1-H2, H1*-H2*, H1-A1, H1*-A1*, CPP, HNR05/15/25/35) were automatically extracted using VoiceSauce [2]. A linear mixed effects regression analysis showed significant interactions between acoustic parameters and speaker sex. In support of [3], our study confirms CPP as a relatively robust parameter for males and extends its suitability to females, with CPP and HNR particularly suitable for female breathy voice. Chloe Patman, Paul Foulkes, Kirsty McDougall |
INTERSPEECH | 2 |
| 2024 | Voice quality in telephone speech: Comparing acoustic measures between VoIP telephone and high-quality recordings
Chenzi Xu, Jessica Wormald, Paul Foulkes, Philip Harrison, Vincent Hughes, Poppy Welch, Finnian Kelly, David van der Vloed |
INTERSPEECH | 3 |
| 2024 | Analysis of forced aligner performance on L2 English speechabstractThere is growing interest in how speech technologies perform on L2 speech. Largely omitted from this discussion are tools used in the early data processing steps, such as forced aligners, that can introduce errors and biases. This study adds to the conversation and tests how well a model pre-trained for the alignment of L1 American English speech performs on L2 English speech. We test and discuss the impact of language variety, demographic factors, and segment type on the performance of the forced aligner. We also examine systematic errors encountered. Forty-five speakers representing nine L2 varieties were selected from the Speech Accent Archive and force aligned using the Montreal Forced Aligner. The phoneme-level boundary placements were manually corrected in order to assess differences between the automatic and manual alignments. Results show marked variation in the performance across language groups and segment types for the two metrics used to assess accuracy: Onset Boundary Displacement, a distance metric between the automatic and manual boundary placements, and Overlap Rate, which indicates to what extent the automatically aligned segment overlaps with the manually aligned segment. The highest accuracy on both measures was obtained for German and French, and lowest accuracy for Russian. The aligner's performance on all varieties was comparable to that on conversational American English and non-standard varieties of English. Furthermore, the percentage of boundary placements within 10 and 20ms of the corrected boundary was similar to that observed between transcribers. Apart from errors due to variety mismatch, most issues encountered in the alignment were due to issues not exclusive to L2 speech such as inaccurate orthographic transcriptions, hesitations, specific voice qualities, and background noise. The results of this study can inform the use of automatic aligners on L2 English speech and provide a baseline of potential errors and information to help the development of more robust alignment tools for further development of automatic systems using L2 English. Samantha Williams, Paul Foulkes, Vincent Hughes |
Speech Commun. | 2 |
| 2023 | Automatic speaker recognition with variation across vocal conditions: a controlled experiment with implications for forensicsabstractAutomatic Speaker Recognition (ASR) involves a complex range of processes to extract, model, and compare speaker-specific information from a pair of voice samples.Using heavily controlled recordings, this paper explores the impact of specific vocal conditions (i.e.vocal setting, disguise, accent guises) on ASR performance.When vocal conditions are matched, ASR performance is generally excellent (whisper is an exception).When conditions are mismatched, as in most forensic cases, we see an increase in discrimination and calibration error in some cases.The most problematic mismatches are those involving whisper and supralaryngeal vocal settings; these produce the greatest phonetic changes to speech.Mismatches involving high pitch also produce poor performance, although this appears to be driven by speaker-specific differences in articulatory implementation.We discuss the implications of the findings for the use of ASR in forensic casework and the interpretability of system output. Vincent Hughes, Jessica Wormald, Paul Foulkes, Philip Harrison, Finnian Kelly, David van der Vloed, Poppy Welch, Chenzi Xu |
INTERSPEECH | 3 |
| 2022 | The effect of sampling variability on systems and individual speakers in likelihood ratio-based forensic voice comparison
Bruce Xiao Wang, Vincent Hughes, Paul Foulkes |
Speech Commun. | 3 |
| 2019 | Voice as a Design Material: Sociophonetic Inspired Design Strategies in Human-Computer InteractionabstractWhile there is a renewed interest in voice user interfaces (VUI) in HCI, little attention has been paid to the design of VUI voice output beyond intelligibility and naturalness. We draw on the field of sociophonetics - the study of the social factors that influence the production and perception of speech - to highlight how current VUIs are based on a limited and homogenised set of voice outputs. We argue that current systems do not adequately consider the diversity of peoples' speech, how that diversity represents sociocultural identities, and how voices have the potential to shape user perceptions and experiences. Ultimately, as other technological developments have influenced the ideologies of language, the voice outputs of VUIs will influence the ideologies of speech. Based on our argument, we pose three design strategies for VUI voice output design - individualisation, context awareness, and diversification - to motivate new ways of conceptualising and designing these technologies. Selina Jeanne Sutton, Paul Foulkes, David S. Kirk, Shaun W. Lawson |
CHI | 2 |
| 2018 | The Individual and the System: Assessing the Stability of the Output of a Semi-automatic Forensic Voice Comparison SystemabstractISCA grants each author permission to use the paper in that author's dissertation or in institutional and public repositories Vincent Hughes, Philip Harrison, Paul Foulkes, Peter French, Colleen Kavanagh, Eugenia San Segundo |
INTERSPEECH | 3 |
| 2017 | What is the Relevant Population? Considerations for the Computation of Likelihood Ratios in Forensic Voice ComparisonabstractIn forensic voice comparison, it is essential to consider not only the similarity between samples, but also the typicality of the evidence in the relevant population. This is explicit within the likelihood ratio (LR) framework. A significant issue, however, is the definition of the relevant population. This paper explores the complexity of population selection for voice evidence. We evaluate the effects of population specificity in terms of regional background on LR output using combinations of the F1, F2, and F3 trajectories of the diphthong /aɪ/. LRs were computed using development and reference data which were regionally matched (Standard Southern British English) and mixed (general British English) relative to the test data. These conditions reflect the paradox that without knowing who the offender is, it is not possible to know the population of which he is a member. Results show that the more specific population produced stronger evidence and better system validity than the more general definition. However, as region-specific voice features (lower formants) were removed, the difference in the output from the matched and mixed systems was reduced. This shows that the effects of population selection are dependent on the sociolinguistic constraints on the feature analysed. Vincent Hughes, Paul Foulkes |
INTERSPEECH | 2 |
| 2017 | Mapping Across Feature Spaces in Forensic Voice Comparison: The Contribution of Auditory-Based Voice Quality to (Semi-)Automatic System TestingabstractISCA grants each author permission to use the paper in that author's dissertation or in institutional and public repositories Vincent Hughes, Philip Harrison, Paul Foulkes, Peter French, Colleen Kavanagh, Eugenia San Segundo |
INTERSPEECH | 3 |
| 2015 | From newcastle MOUTH to aussie ears: australians' perceptual assimilation and adaptation for newcastle UK vowelsabstractTo probe how episodic and abstract processes contribute to flexible perception of phonetically variable speech, we evaluated Australian (Aus) listeners' perception of Aus-accented vowels versus those of an unfamiliar accent: Newcastle UK (Ncl).Aus listeners first heard a round-robin story told by multiple talkers of Aus or Ncl, then categorized multi-talker tokens of 20 vowels in nonce words spoken in the Aus or Ncl accent.Categorization was variable even across Aus nonce vowels (M accuracy ranged from 21-80%).Perceptual assimilation of Ncl vowels (Aus passage/Ncl nonce) was diverse: Some were categorized very much like the corresponding Aus vowel.Some showed within-category differentiation from Aus; others were heard as a different vowel altogether.Perception of some Ncl vowels changed after Ncl passage exposure, including both positive adaptation (improved categorization: e.g., MOUTH, FLEECE, TRAP) and negative shifts (increased differentiation from the corresponding Aus vowel: e.g., NURSE, FOOT).Assimilation and adaptation patterns were largely consistent with similarities and dissimilarities between the Aus and Ncl vowel spaces.Implications of the results for episodic and abstract contributions to perceptual flexibility are discussed.We also consider the possibility that listeners perceptually adjust to other-accent vowels as a system, rather than treating each vowel as an independent entity. Catherine T. Best, Jason A. Shaw, Gerard Docherty, Bronwen G. Evans, Paul Foulkes, Jennifer Hay, Jalal Al-Tamimi, Katharine Mair, Karen E. Mulak, Sophie Wood |
INTERSPEECH | 5 |
| 2015 | The relevant population in forensic voice comparison: Effects of varying delimitations of social class and age
Vincent Hughes, Paul Foulkes |
Speech Commun. | 2 |