Philip Harrison

dblp:69/4949 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0003-2419-2388ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 since 2021Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Variability in performance across four generations of automatic speaker recognition systems
Lauren Harrington, Vincent Hughes, Philip Harrison, Paul Foulkes, Jessica Wormald, Finnian Kelly, David van der Vloed
INTERSPEECH3
2024 Voice quality in telephone speech: Comparing acoustic measures between VoIP telephone and high-quality recordings
Chenzi Xu, Jessica Wormald, Paul Foulkes, Philip Harrison, Vincent Hughes, Poppy Welch, Finnian Kelly, David van der Vloed
INTERSPEECH4
2023 Evaluation of a Forensic Automatic Speaker Recognition System with Emotional Speech Recordings
Robert Essery, Philip Harrison, Vincent Hughes
INTERSPEECH2
2023 Automatic speaker recognition with variation across vocal conditions: a controlled experiment with implications for forensics
abstract
Automatic Speaker Recognition (ASR) involves a complex range of processes to extract, model, and compare speaker-specific information from a pair of voice samples.Using heavily controlled recordings, this paper explores the impact of specific vocal conditions (i.e.vocal setting, disguise, accent guises) on ASR performance.When vocal conditions are matched, ASR performance is generally excellent (whisper is an exception).When conditions are mismatched, as in most forensic cases, we see an increase in discrimination and calibration error in some cases.The most problematic mismatches are those involving whisper and supralaryngeal vocal settings; these produce the greatest phonetic changes to speech.Mismatches involving high pitch also produce poor performance, although this appears to be driven by speaker-specific differences in articulatory implementation.We discuss the implications of the findings for the use of ASR in forensic casework and the interpretability of system output.
Vincent Hughes, Jessica Wormald, Paul Foulkes, Philip Harrison, Finnian Kelly, David van der Vloed, Poppy Welch, Chenzi Xu
INTERSPEECH4
2023 Automatic Speaker Recognition performance with matched and mismatched female bilingual speech data
Bryony Nuttall, Philip Harrison, Vincent Hughes
INTERSPEECH2
2021 A Comparison of the Accuracy of Dissen and Keshet's (2016) DeepFormants and Traditional LPC Methods for Semi-Automatic Speaker Recognition
Thomas Coy, Vincent Hughes, Philip Harrison, Amelia Jane Gully
Interspeech3
2021 Human Spoofing Detection Performance on Degraded Speech
Camryn Terblanche, Philip Harrison, Amelia Jane Gully
Interspeech2
2020 Correlating Cepstra with Formant Frequencies: Implications for Phonetically-Informed Forensic Voice Comparison
abstract
A significant question for forensic voice comparison, and for speaker recognition more generally, is the extent to which different input features capture complementary speakerspecific information.Understanding complementarity allows us to make predictions about how combining methods using different features may produce better overall performance.In forensic contexts, it is also important to be able to explain to courts what information the underlying features are actually capturing.This paper addresses these issues by examining the extent to which MFCCs and LPCCs can predict F0, F1, F2, and F3 values using data extracted from the midpoint of the vocalic portion of the hesitation marker um for 89 speakers of standard southern British English.By-speaker correlations were calculated using multiple linear regression and performance was assessed using mean rho (𝜌) values.Results show that the first two formants were more accurately predicted than F3 or F0.LPCCs consistently produced stronger correlations with the linguistic features than MFCCs, while increasing cepstral order up to 16 also increased the strength of the correlations.There was, however, considerable variability across speakers in terms of the accuracy of the predictions.We discuss the implications of these findings for forensic voice comparison.
Vincent Hughes, Frantz Clermont, Philip Harrison
INTERSPEECH3
2018 The Individual and the System: Assessing the Stability of the Output of a Semi-automatic Forensic Voice Comparison System
abstract
ISCA grants each author permission to use the paper in that author's dissertation or in institutional and public repositories
Vincent Hughes, Philip Harrison, Paul Foulkes, Peter French, Colleen Kavanagh, Eugenia San Segundo
INTERSPEECH2
2017 Mapping Across Feature Spaces in Forensic Voice Comparison: The Contribution of Auditory-Based Voice Quality to (Semi-)Automatic System Testing
abstract
ISCA grants each author permission to use the paper in that author's dissertation or in institutional and public repositories
Vincent Hughes, Philip Harrison, Paul Foulkes, Peter French, Colleen Kavanagh, Eugenia San Segundo
INTERSPEECH2
2009 Large-scale extraction and use of knowledge from text
abstract
Many AI tasks, in particular natural language processing, require a large amount of world knowledge to create expectations, assess plausibility, and guide disambiguation. However, acquiring this world knowledge remains a formidable challenge. Building on ideas by Schubert, we have developed a system called DART (Discovery and Aggregation of Relations in Text) that extracts simple, semi-formal statements of world knowledge (e.g., "airplanes can fly", "people can drive cars") from text by abstracting from a parser's output, and we have used it to create a database of 23 million propositions of this kind. An evaluation of the DART database on two language processing tasks (parsing and textual entailment) shows that it improves performance, and a human evaluation shows that over half the facts in it are considered true or partially true, rising to 70% for facts seen with high frequency. The significance of this work is two-fold: First it has created a new, publically available knowledge resource for language processing and other data interpretation tasks, and second it provides empirical evidence of the utility of this type of knowledge, going beyond Schubert et al's earlier evaluations which were based solely on human inspection of its contents.
Peter Clark, Philip Harrison
K-CAP2
2007 Capturing and answering questions posed to a knowledge-based system
abstract
As part of the ongoing project, Project Halo, our goal is to build a system capable of answering questions posed by novice users to a formal knowledge base. In our current context, the knowledge base covers selected topics in physics, chemistry, and biology, and our question set consists of AP (advanced high-school) level examination questions. The task is challenging because the questions are linguistically complex and are often incomplete (assume unstated knowledge), and because the users do not have prior knowledge of the system's contents. Our solution involves two parts: a controlled language interface, in which users reformulate the original natural language questions in a simplified version of English, and a novel problem solver that can elaborate initially inadequate logical interpretations of a question by selecting relevant pieces of knowledge in the knowledge base. An evaluation of the work in 2006 showed that this approach is feasible and that complex, multisentence questions can be posed and answered, thus illustrating novel ways of dealing with the knowledge capture impedance between users and a formal knowledge base, while also revealing challenges that still remain.
Peter Clark, Shaw Yi Chaw, Ken Barker 0002, Vinay K. Chaudhri, Philip Harrison, James Fan, Bonnie E. John, Bruce W. Porter, Aaron Spaulding, John A. Thompson, Peter Z. Yeh
K-CAP5
1993 Using Bracketed Parses to evaluate a Grammar Checking Application
abstract
We describe a method for evaluating a grammar checking application with hand-bracketed parses.A randomly-selected set of sentences was submitted to a grammar checker in both bracketed and unbracketed formats.A comparison of the resulting error reports illuminates the relationship between the underlying performance of the parsergrammar system and the error critiques presented to the user.
Richard H. Wojcik, Philip Harrison, John Bremer
ACL2
1979 Task driven image understanding: lisp programming for vision research
abstract
The artificial intelligence aspects of computer vision make LISP appropriate for some vision re search needs. At the University of Washington an experimental image understanding system has been implemented in MACLISP. Since MACLISP supports arrays and compiles into reasonably efficient code, it has been convenient for the project. The major subsystems are the following: preprocessing, region growing, feature computation, and control. A novel aspect of the system is its control structure which consists of a "task agenda" which is processed by procedures each based on heuristics relevant to image analysis. Processing one task may create other tasks and the system is designed to suggest alternatives during the course of analyzing an image. Tasks are ordered by their heuristically-computed "priority" values. The system has been applied to the problem of segmenting aerial photographs of an airport. Design issues and possible extensions for further work are discussed.
Steven L. Tanimoto, Philip Harrison
COMPSAC2