EDBT 2026 Demo / reviewers in the wild / expert
Vincent Hughes
dblp:49/7463
· DBLP profile ↗
19ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-4660-979XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The discriminative capacity of English segments in forensic speaker comparison
Paul Foulkes, Vincent Hughes, Kayleigh Peters, Jasmine Rouse |
Speech Commun. | 2 |
| 2025 | Variability in performance across four generations of automatic speaker recognition systems
Lauren Harrington, Vincent Hughes, Philip Harrison, Paul Foulkes, Jessica Wormald, Finnian Kelly, David van der Vloed |
INTERSPEECH | 2 |
| 2025 | Human and automatic voice comparison with regionally variable speech samples
Vincent Hughes, Carmen Llamas, Thomas Kettig |
Speech Commun. | 1 |
| 2024 | Voice quality in telephone speech: Comparing acoustic measures between VoIP telephone and high-quality recordings
Chenzi Xu, Jessica Wormald, Paul Foulkes, Philip Harrison, Vincent Hughes, Poppy Welch, Finnian Kelly, David van der Vloed |
INTERSPEECH | 5 |
| 2024 | Analysis of forced aligner performance on L2 English speechabstractThere is growing interest in how speech technologies perform on L2 speech. Largely omitted from this discussion are tools used in the early data processing steps, such as forced aligners, that can introduce errors and biases. This study adds to the conversation and tests how well a model pre-trained for the alignment of L1 American English speech performs on L2 English speech. We test and discuss the impact of language variety, demographic factors, and segment type on the performance of the forced aligner. We also examine systematic errors encountered. Forty-five speakers representing nine L2 varieties were selected from the Speech Accent Archive and force aligned using the Montreal Forced Aligner. The phoneme-level boundary placements were manually corrected in order to assess differences between the automatic and manual alignments. Results show marked variation in the performance across language groups and segment types for the two metrics used to assess accuracy: Onset Boundary Displacement, a distance metric between the automatic and manual boundary placements, and Overlap Rate, which indicates to what extent the automatically aligned segment overlaps with the manually aligned segment. The highest accuracy on both measures was obtained for German and French, and lowest accuracy for Russian. The aligner's performance on all varieties was comparable to that on conversational American English and non-standard varieties of English. Furthermore, the percentage of boundary placements within 10 and 20ms of the corrected boundary was similar to that observed between transcribers. Apart from errors due to variety mismatch, most issues encountered in the alignment were due to issues not exclusive to L2 speech such as inaccurate orthographic transcriptions, hesitations, specific voice qualities, and background noise. The results of this study can inform the use of automatic aligners on L2 English speech and provide a baseline of potential errors and information to help the development of more robust alignment tools for further development of automatic systems using L2 English. Samantha Williams, Paul Foulkes, Vincent Hughes |
Speech Commun. | 3 |
| 2023 | Evaluation of a Forensic Automatic Speaker Recognition System with Emotional Speech Recordings
Robert Essery, Philip Harrison, Vincent Hughes |
INTERSPEECH | 3 |
| 2023 | Automatic speaker recognition with variation across vocal conditions: a controlled experiment with implications for forensicsabstractAutomatic Speaker Recognition (ASR) involves a complex range of processes to extract, model, and compare speaker-specific information from a pair of voice samples.Using heavily controlled recordings, this paper explores the impact of specific vocal conditions (i.e.vocal setting, disguise, accent guises) on ASR performance.When vocal conditions are matched, ASR performance is generally excellent (whisper is an exception).When conditions are mismatched, as in most forensic cases, we see an increase in discrimination and calibration error in some cases.The most problematic mismatches are those involving whisper and supralaryngeal vocal settings; these produce the greatest phonetic changes to speech.Mismatches involving high pitch also produce poor performance, although this appears to be driven by speaker-specific differences in articulatory implementation.We discuss the implications of the findings for the use of ASR in forensic casework and the interpretability of system output. Vincent Hughes, Jessica Wormald, Paul Foulkes, Philip Harrison, Finnian Kelly, David van der Vloed, Poppy Welch, Chenzi Xu |
INTERSPEECH | 1 |
| 2023 | Automatic Speaker Recognition performance with matched and mismatched female bilingual speech data
Bryony Nuttall, Philip Harrison, Vincent Hughes |
INTERSPEECH | 3 |
| 2022 | Eliciting and evaluating likelihood ratios for speaker recognition by human listeners under forensically realistic channel-mismatched conditions
Vincent Hughes, Carmen Llamas, Thomas Kettig |
INTERSPEECH | 1 |
| 2022 | Reducing uncertainty at the score-to-LR stage in likelihood ratio-based forensic voice comparison using automatic speaker recognition systems
Bruce Xiao Wang, Vincent Hughes |
INTERSPEECH | 2 |
| 2022 | The effect of sampling variability on systems and individual speakers in likelihood ratio-based forensic voice comparison
Bruce Xiao Wang, Vincent Hughes, Paul Foulkes |
Speech Commun. | 2 |
| 2021 | A Comparison of the Accuracy of Dissen and Keshet's (2016) DeepFormants and Traditional LPC Methods for Semi-Automatic Speaker Recognition
Thomas Coy, Vincent Hughes, Philip Harrison, Amelia Jane Gully |
Interspeech | 2 |
| 2021 | System Performance as a Function of Calibration Methods, Sample Size and Sampling Variability in Likelihood Ratio-Based Forensic Voice ComparisonabstractIn data-driven forensic voice comparison, sample size is an issue which can have substantial effects on system output.Numerous calibration methods have been developed and some have been proposed as solutions to sample size issues.In this paper, we test four calibration methods (i.e.logistic regression, regularised logistic regression, Bayesian model, ELUB) under different conditions of sampling variability and sample size.Training and test scores were simulated from skewed distributions derived from real experiments, increasing sample sizes from 20 to 100 speakers for both the training and test sets.For each sample size, the experiments were replicated 100 times to test the susceptibility of different calibration methods to sampling variability.The Cllr mean and range across replications were used for evaluation.The Bayesian model and regularized logistic regression produced the most stable Cllr values when the sample size is small (i.e.20 speakers), although mean Cllr is consistently lowest using logistic regression.The ELUB calibration method generally is the least preferred as it is the most sensitive to sample size and sampling variability (mean = 0.66, range = 0.21-0.59). Bruce Xiao Wang, Vincent Hughes |
Interspeech | 2 |
| 2020 | Correlating Cepstra with Formant Frequencies: Implications for Phonetically-Informed Forensic Voice ComparisonabstractA significant question for forensic voice comparison, and for speaker recognition more generally, is the extent to which different input features capture complementary speakerspecific information.Understanding complementarity allows us to make predictions about how combining methods using different features may produce better overall performance.In forensic contexts, it is also important to be able to explain to courts what information the underlying features are actually capturing.This paper addresses these issues by examining the extent to which MFCCs and LPCCs can predict F0, F1, F2, and F3 values using data extracted from the midpoint of the vocalic portion of the hesitation marker um for 89 speakers of standard southern British English.By-speaker correlations were calculated using multiple linear regression and performance was assessed using mean rho (𝜌) values.Results show that the first two formants were more accurately predicted than F3 or F0.LPCCs consistently produced stronger correlations with the linguistic features than MFCCs, while increasing cepstral order up to 16 also increased the strength of the correlations.There was, however, considerable variability across speakers in terms of the accuracy of the predictions.We discuss the implications of these findings for forensic voice comparison. Vincent Hughes, Frantz Clermont, Philip Harrison |
INTERSPEECH | 1 |
| 2018 | The Individual and the System: Assessing the Stability of the Output of a Semi-automatic Forensic Voice Comparison SystemabstractISCA grants each author permission to use the paper in that author's dissertation or in institutional and public repositories Vincent Hughes, Philip Harrison, Paul Foulkes, Peter French, Colleen Kavanagh, Eugenia San Segundo |
INTERSPEECH | 1 |
| 2017 | What is the Relevant Population? Considerations for the Computation of Likelihood Ratios in Forensic Voice ComparisonabstractIn forensic voice comparison, it is essential to consider not only the similarity between samples, but also the typicality of the evidence in the relevant population. This is explicit within the likelihood ratio (LR) framework. A significant issue, however, is the definition of the relevant population. This paper explores the complexity of population selection for voice evidence. We evaluate the effects of population specificity in terms of regional background on LR output using combinations of the F1, F2, and F3 trajectories of the diphthong /aɪ/. LRs were computed using development and reference data which were regionally matched (Standard Southern British English) and mixed (general British English) relative to the test data. These conditions reflect the paradox that without knowing who the offender is, it is not possible to know the population of which he is a member. Results show that the more specific population produced stronger evidence and better system validity than the more general definition. However, as region-specific voice features (lower formants) were removed, the difference in the output from the matched and mixed systems was reduced. This shows that the effects of population selection are dependent on the sociolinguistic constraints on the feature analysed. Vincent Hughes, Paul Foulkes |
INTERSPEECH | 1 |
| 2017 | Mapping Across Feature Spaces in Forensic Voice Comparison: The Contribution of Auditory-Based Voice Quality to (Semi-)Automatic System TestingabstractISCA grants each author permission to use the paper in that author's dissertation or in institutional and public repositories Vincent Hughes, Philip Harrison, Paul Foulkes, Peter French, Colleen Kavanagh, Eugenia San Segundo |
INTERSPEECH | 1 |
| 2017 | Sample size and the multivariate kernel density likelihood ratio: How many speakers are enough?
Vincent Hughes |
Speech Commun. | 1 |
| 2015 | The relevant population in forensic voice comparison: Effects of varying delimitations of social class and age
Vincent Hughes, Paul Foulkes |
Speech Commun. | 1 |