VLDB 2026 Research / reviewers in the wild / expert
Naomi Harte
dblp:56/5068
· DBLP profile ↗
80ranked-venue papers
9as first author
29since 2021 · last 2027
0000-0002-9274-209XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 48 · 7 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | A multimodal perspective on adaptive communication: Extending the hyper- and hypo-articulation theoryabstractHuman communication is highly adaptive: speakers continuously adjust how they speak, move, and look in response to changing communicative conditions. These adaptations are central to face-to-face interaction and are increasingly relevant for speech technologies that aim to operate in natural, interactive, and multimodal settings. While adaptive speech has been extensively studied within the articulatory and acoustic domains, spoken communication is fundamentally multimodal, raising the question of how visual and bodily cues participate in adaptive behaviour. Lindblom’s classic Hyper- and Hypo-articulation (H&H) Theory provides a principled account of adaptive speech production by framing variability as a trade-off between effort and intelligibility. However, how this adaptive logic extends beyond speech to multimodal communication remains insufficiently integrated across disciplines. This review synthesises research on speech and a range of visual and bodily signals to examine how multimodal behaviours contribute to adaptive communication. We first review evidence on how environmental, listener-related, and speaker-related constraints shape adaptive responses in speech. We then extend this analysis to multimodal communication, highlighting how speakers and listeners flexibly redistribute communicative effort across modalities in response to environmental, listener-related, and speaker-internal constraints. We argue that multimodal cues participate in the same adaptive optimisation described by the H&H Theory, and that adaptation emerges from dynamic trade-off across modalities rather than from speech alone. This multimodal reinterpretation of H&H has implications for speech technologies, including audiovisual speech, synthesis, and interactive systems. • Revisit Lindblom’s Hyper–Hypo Theory from a multimodal perspective • A review of adaptive speech and multimodal communication. • Environmental, listener, and speaker constraints are jointly considered. • Adaptation emerges from trade-off across communicative modalities. • Implications are discussed for audiovisual and interactive speech systems. Delphine Charuau, Naomi Harte |
Comput. Speech Lang. | 2 |
| 2027 | Advancing listening effort-based evaluation of speech enhancement systemsabstractThis study investigates response time as a behavioral indicator related to listening effort (LE) for evaluating speech enhancement (SE) systems. English and Norwegian intelligibility matrix tests were conducted within a single-task paradigm that incorporated click-time recording (logging the precise time of all participant clicks), enabling simultaneous estimation of speech intelligibility and LE-related temporal behavior. Three temporal proxy measures for LE were examined—time per stimulus, reaction time, and word click time—across a broad range of input signal-to-noise ratios (SNRs) and for both discriminative and generative enhancement approaches. Time per stimulus showed an inverted-U pattern across SNRs, whereas reaction time and word click time exhibited monotonic behavior, providing more directly interpretable metrics for comparative evaluation. Analyses of pairwise SNR comparisons revealed that increases in our LE-related temporal measures at higher SNRs precede measurable intelligibility declines, suggesting that these temporal metrics can be more sensitive than intelligibility in this regime. Overall, the proposed framework—where LE-related measurements remain unknown to participants—offers a comprehensive and nuanced behavioral tool for SE evaluation, complementing intelligibility particularly under realistic, moderate-to-high SNR conditions. • A unified behavioral framework that jointly captures intelligibility and listening-effort-related information within a single listening task, enabling more expressive evaluation of speech enhancement systems. • Three temporal proxy measures for listening effort (time per stimulus, reaction time, and word click time) that capture complementary aspects of cognitive effort during system evaluation. • Evidence that temporal effort-related measures increase at higher signal-to-noise ratios (SNRs) before observable intelligibility declines, providing complementary insight under moderate-to-high SNR conditions. Iván López-Espejo, Femke B. Gelderblom, Tron V. Tronstad, Christoffer Blomberg Skiaker, Naomi Harte |
Comput. Speech Lang. | 5 |
| 2026 | VisG AV-HuBERT: Viseme-Guided AV-HuBERT
Aristeidis Papadopoulos, Naomi Harte |
ICPR (1) | 3 |
| 2026 | Evaluation of Co-Speech Gesture Tracking Techniques in Naturalistic Interactions
Victoria Ivanova, Naomi Harte |
LREC | 2 |
| 2026 | The use of variable length stimuli for assessing segmental distortion in TTS evaluation
Ayushi Pandey, Jens Edlund, Sébastien Le Maguer, Naomi Harte |
Comput. Speech Lang. | 4 |
| 2025 | Interpreting the Role of Visemes in Audio-Visual Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) models have surpassed their audio-only counterparts in terms of performance. However, the interpretability of AVSR systems, particularly the role of the visual modality, remains under-explored. In this paper, we apply several interpretability techniques to examine how visemes are encoded in AV-HuBERT a state-of-the-art AVSR model. First, we use t-distributed Stochastic Neighbour Embedding (t-SNE) to visualize learned features, revealing natural clustering driven by visual cues, which is further refined by the presence of audio. Then, we employ probing to show how audio contributes to refining feature representations, particularly for visemes that are visually ambiguous or under-represented. Our findings shed light on the interplay between modalities in AVSR and could point to new strategies for leveraging visual information to improve AVSR performance. Aristeidis Papadopoulos, Naomi Harte |
ASRU | 2 |
| 2025 | Uncovering the Visual Contribution in Audio-Visual Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) combines auditory and visual speech cues to enhance the accuracy and robustness of speech recognition systems. Recent advancements in AVSR have improved performance in noisy environments compared to audio-only counterparts. However, the true extent of the visual contribution, and whether AVSR systems fully exploit the available cues in the visual domain, remains unclear. This paper assesses AVSR systems from a different perspective, by considering human speech perception. We use three systems: Auto-AVSR, AVEC and AV-RelScore. We first quantify the visual contribution using effective SNR gains at 0 dB and then investigate the use of visual information in terms of its temporal distribution and word-level informativeness. We show that low WER does not guarantee high SNR gains. Our results suggest that current methods do not fully exploit visual information, and we recommend future research to report effective SNR gains alongside WERs. Zhaofeng Lin, Naomi Harte |
ICASSP | 2 |
| 2025 | Multimodal Dynamics of Hand Gestures and Pauses in Multiparty InteractionsabstractInternational audience Delphine Charuau, Naomi Harte |
INTERSPEECH | 2 |
| 2025 | Enabling the replicability of speech synthesis perceptual evaluationsabstractHow speech synthesis is evaluated is nowadays questioned. Not only have conventional listening tests as a whole been proven a poor match for modern synthesis, but more fundamentally, important information (e.g., the question asked to the listener) is frequently missing in the report of the outcome of the evaluation despite the impact on the interpretation of the test results. This can lead to uncertainty about the validity of these evaluations. To address this issue, we propose standardising the structure of any evaluation report. To facilitate this standardisation, our contribution is twofold: an open-source subjective evaluation platform; and a set of reporting guidelines. The platform is designed to enable the development of easily shareable evaluation recipes. The set of guidelines complements the platform to support researchers in reporting their evaluation choices and analysis in more detail while relying on the recipe to describe the actual evaluation process. Sébastien Le Maguer, Gwénolé Lecorvé, Damien Lolive, Naomi Harte, Juraj Simko |
INTERSPEECH | 4 |
| 2025 | Visual Cues Support Robust Turn-taking Prediction in NoiseabstractAccurate predictive turn-taking models (PTTMs) are essential for naturalistic human-robot interaction. However, little is known about their performance in noise. This study therefore explores PTTM performance in types of noise likely to be encountered once deployed. Our analyses reveal PTTMs are highly sensitive to noise. Hold/shift accuracy drops from 84% in clean speech to just 52% in 10 dB music noise. Training with noisy data enables a multimodal PTTM, which includes visual features to better exploit visual cues, with 72% accuracy in 10 dB music noise. The multimodal PTTM outperforms the audio-only PTTM across all noise types and SNRs, highlighting its ability to exploit visual cues; however, this does not always generalise to new types of noise. Analysis also reveals that successful training relies on accurate transcription, limiting the use of ASR-derived transcriptions to clean conditions. We make code publicly available for future research. Sam O'Connor Russell, Naomi Harte |
INTERSPEECH | 2 |
| 2025 | Noise-Robust Hearing Aid Voice ControlabstractAdvancing the design of robust hearing aid (HA) voice control is crucial to increase the HA use rate among hard of hearing people as well as to improve HA users' experience. In this work, we contribute towards this goal by, first, presenting a novel HA speech dataset consisting of noisy own voice captured by 2 behind-the-ear (BTE) and 1 in-ear-canal (IEC) microphones. Second, we provide baseline HA voice control results from the evaluation of light, state-of-the-art keyword spotting models utilizing different combinations of HA microphone signals. Experimental results show the benefits of exploiting bandwidth-limited bone-conducted speech (BCS) from the IEC microphone to achieve noise-robust HA voice control. Furthermore, results also demonstrate that voice control performance can be boosted by assisting BCS by the broader-bandwidth BTE microphone signals. Aiming at setting a baseline upon which the scientific community can continue to progress, the HA noisy speech dataset has been made publicly available. Iván López-Espejo, Eros Roselló, Amin Edraki, Naomi Harte, Jesper Jensen 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Language Bias in Self-Supervised Learning For Automatic Speech RecognitionabstractSelf-supervised learning (SSL) is used in deep learning to train on large datasets without the need for expensive labelling of the data. Recently, large Automatic Speech Recognition (ASR) models such as XLS-R have utilised SSL to train on over one hundred different languages simultaneously. However, deeper investigation shows that the bulk of the training data for XLS-R comes from a small number of languages. Biases learned through SSL have been shown to exist in multiple domains, but language bias in multilingual SSL ASR has not been thoroughly examined. In this paper, we utilise the Lottery Ticket Hypothesis (LTH) to identify language-specific subnetworks within XLS-R and test the performance of these subnetworks on a variety of different languages. We are able to show that when fine-tuning, XLS-R bypasses traditional linguistic knowledge and builds only on weights learned from the languages with the largest data contribution to the pretraining data. Edward Storey, Naomi Harte, Peter Bell 0001 |
SLT | 2 |
| 2024 | The limits of the Mean Opinion Score for speech synthesis evaluation
Sébastien Le Maguer, Simon King 0001, Naomi Harte |
Comput. Speech Lang. | 3 |
| 2024 | Smiling in the Face and Voice of Avatars and Robots: Evidence for a 'Smiling McGurk Effect'abstractMultisensory integration influences emotional perception, as the McGurk effect demonstrates for the communication between humans. Human physiology implicitly links the production of visual features with other modes like the audio channel: Face muscles responsible for a smiling face also stretch the vocal cords that result in a characteristic smiling voice. For artificial agents capable of multimodal expression, this linkage is modeled explicitly. In our studies, we observe the influence of visual and audio channels on the perception of the agents' emotional expression. We created videos of virtual characters and social robots either with matching or mismatching emotional expressions in the audio and visual channels. In two online studies, we measured the agents' perceived valence and arousal. Our results consistently lend support to the ‘emotional McGurk effect' hypothesis, according to which face transmits valence information, and voice transmits arousal. When dealing with dynamic virtual characters, visual information is enough to convey both valence and arousal, and thus audio expressivity need not be congruent. When dealing with robots with fixed facial expressions, however, both visual and audio information need to be present to convey the intended expression. Ilaria Torre 0002, Simon Holk, Elmira Yadollahi, Iolanda Leite, Rachel McDonnell, Naomi Harte |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | Learnable Frontends That Do Not Learn: Quantifying Sensitivity To Filterbank InitialisationabstractWhile much of modern speech and audio processing relies on deep neural networks trained using fixed audio representations, recent studies suggest great potential in acoustic frontends learnt jointly with a backend. In this study, we focus specifically on learnable filterbanks. Prior studies have reported that in frontends using learnable filterbanks initialised to a mel scale, the learned filters do not differ substantially from their initialisation. Using a Gabor-based filterbank, we investigate the sensitivity of a learnable filterbank to its initialisation using several initialisation strategies on two audio tasks: voice activity detection and bird species identification. We use the Jensen-Shannon Distance and analysis of the learned filters before and after training. We show that although performance is overall improved, the filterbanks exhibit strong sensitivity to their initialisation strategy. The limited movement from initialised values suggests that alternate optimisation strategies may allow a learnable frontend to reach better overall performance. Mark Anderson 0006, Tomi Kinnunen, Naomi Harte |
ICASSP | 3 |
| 2023 | Query Based Acoustic Summarization for Podcasts
Samantha Kotey, Rozenn Dahyot, Naomi Harte |
INTERSPEECH | 3 |
| 2023 | Sp1NY: A Quick and Flexible Speech Visualisation Tool in Python
Sébastien Le Maguer, Mark Anderson 0006, Naomi Harte |
INTERSPEECH | 3 |
| 2023 | Listener sensitivity to deviating obstruents in WaveNet
Ayushi Pandey, Jens Edlund, Sébastien Le Maguer, Naomi Harte |
INTERSPEECH | 4 |
| 2022 | Back to the Future: Extending the Blizzard Challenge 2013
Sébastien Le Maguer, Simon King 0001, Naomi Harte |
INTERSPEECH | 3 |
| 2022 | Production characteristics of obstruents in WaveNet and older TTS systems
Ayushi Pandey, Sébastien Le Maguer, Julie Carson-Berndsen, Naomi Harte |
INTERSPEECH | 4 |
| 2022 | RoomReader: A Multimodal Corpus of Online Multiparty Conversational InteractionsabstractWe present RoomReader, a corpus of multimodal, multiparty conversational interactions in which participants followed a collaborative student-tutor scenario designed to elicit spontaneous speech. The corpus was developed within the wider RoomReader Project to explore multimodal cues of conversational engagement and behavioural aspects of collaborative interaction in online environments. However, the corpus can be used to study a wide range of phenomena in online multimodal interaction. The publicly-shared corpus consists of over 8 hours of video and audio recordings from 118 participants in 30 gender-balanced sessions, in the “in-the-wild” online environment of Zoom. The recordings have been edited, synchronised, and fully transcribed. Student participants have been continuously annotated for engagement with a novel continuous scale. We provide questionnaires measuring engagement and group cohesion collected from the annotators, tutors and participants themselves. We also make a range of accompanying data available such as personality tests and behavioural assessments. The dataset and accompanying psychometrics present a rich resource enabling the exploration of a range of downstream tasks across diverse fields including linguistics and artificial intelligence. This could include the automatic detection of student engagement, analysis of group interaction and collaboration in online conversation, and the analysis of conversational behaviours in an online setting. Justine Reverdy, Sam O'Connor Russell, Louise Duquenne, Diego Garaialde, Benjamin R. Cowan, Naomi Harte |
LREC | 6 |
| 2022 | To smile or not to smile: The effect of mismatched emotional expressions in a Human-Robot cooperative taskabstractEmotional expressivity is essential for successful Human-Robot Interaction. However, robots often have different levels of expressivity in their face and voice. Here we ask whether this modality mismatch influences human behaviour and perception of the robot. Participants played a cooperative task with a robot that displayed matched and mismatched smiling expressions in the face and voice. Emotional expressivity did not influence acceptance of robot’s recommendations or subjective evaluations of the robot. However, we found that the robot had overall a higher social influence than a virtual character, and was evaluated more positively. Ilaria Torre 0002, Anna Deichler, Matthew Nicholson, Rachel McDonnell, Naomi Harte |
RO-MAN | 5 |
| 2022 | Fine Grained Spoken Document Summarization Through Text SegmentationabstractPodcast transcripts are long spoken documents of conversational dialogue. Challenging to summarize, podcasts cover a diverse range of topics, vary in length, and have uniquely different linguistic styles. Previous studies in podcast summarization have generated short, concise dialogue summaries. In contrast, we propose a method to generate long fine-grained summaries, which describe details of sub-topic narratives. Leveraging a readability formula, we curate a data subset to train a long sequence transformer for abstractive summarization. Through text segmentation, we filter the evaluation data and exclude specific segments of text. We apply the model to segmented data, producing different types of fine grained summaries. We show that appropriate filtering creates comparable results on ROUGE and serves as an alternative method to truncation. Experiments show our model outperforms previous studies on the Spotify podcast dataset when tasked with generating longer sequences of text. Samantha Kotey, Rozenn Dahyot, Naomi Harte |
SLT | 3 |
| 2022 | Taris: An online speech recognition framework with sequence to sequence neural networks for both audio-only and audio-visual speechabstractIt is widely accepted that the visual modality of speech provides complementary information to the speech recognition task, and many models have been introduced in order to make good use of the visual channel. This article develops Taris, a fully differentiable neural network model capable of decoding both audio-only and audio-visual speech in real time. We achieve this by connecting our previously proposed models AV Align and Taris, which are both end-to-end differentiable approaches to audio-visual speech integration and online speech recognition respectively. We evaluate AV Taris under the same conditions as AV Align and Taris on one of the largest publicly available audio-visual speech datasets, LRS2. Our results show that AV Taris is superior to the audio-only variant of Taris, demonstrating the utility of the visual modality to speech recognition within the real time decoding framework defined by Taris. Compared to an equivalent Transformer-based AV Align model that takes advantage of full sentences without meeting the real-time requirement, we report an absolute degradation of approximately 3% with AV Taris. As opposed to the more popular alternative for online speech recognition, namely the RNN Transducer, Taris offers a greatly simplified fully differentiable training pipeline. We speculate that AV Taris has the potential to popularise the adoption of Audio-Visual Speech Recognition (AVSR) technology and overcome the inherent limitations of the audio modality in less optimal listening conditions.1 George Sterpu, Naomi Harte |
Comput. Speech Lang. | 2 |
| 2022 | Comparison of discrete transforms for deep-neural-networks-based speech enhancementabstractAbstract In recent studies of speech enhancement, a deep‐learning model is trained to predict clean speech spectra from the known noisy spectra of speech. Rather than using the traditional discrete Fourier transform (DFT), this paper considers other well‐known transforms to generate the speech spectra for deep‐learning‐based speech enhancement. In addition to the DFT, seven different transforms were tested: discrete Cosine transform, discrete Sine transform, discrete Haar transform, discrete Hadamard transform, discrete Tchebichef transform, discrete Krawtchouk transform, and discrete Tchebichef‐Krawtchouk transform. Two deep‐learning architectures were tested: convolutional neural networks (CNN) and fully connected neural networks. Experiments were performed for the NOIZEUS database, and various speech quality and intelligibility measures were adopted for performance evaluation. The quality and intelligibility scores of the enhanced speech demonstrate that discrete Sine transformation is better suited for the front‐end processing with a CNN as it outperformed the DFT in this kind of application. The achieved results demonstrate that combining two or more existing transforms could improve the performance in specific conditions. The tested models suggest that we should not assume that the DFT is optimal in front‐end processing with deep neural networks (DNNs). On this basis, other discrete transformations should be taken into account when designing robust DNN‐based speech processing applications. Wissam A. Jassim, Naomi Harte |
IET Signal Process. | 2 |
| 2022 | Deep Multi-Scale Feature Learning for Defocus Blur EstimationabstractThis paper presents an edge-based defocus blur estimation method from a single defocused image. We first distinguish edges that lie at depth discontinuities (called depth edges, for which the blur estimate is ambiguous) from edges that lie at approximately constant depth regions (called pattern edges, for which the blur estimate is well-defined). Then, we estimate the defocus blur amount at pattern edges only, and explore an interpolation scheme based on guided filters that prevents data propagation across the detected depth edges to obtain a dense blur map with well-defined object boundaries. Both tasks (edge classification and blur estimation) are performed by deep convolutional neural networks (CNNs) that share weights to learn meaningful local features from multi-scale patches centered at edge locations. Experiments on naturally defocused images show that the proposed method presents qualitative and quantitative results that outperform state-of-the-art (SOTA) methods, with a good compromise between running time and accuracy. Ali Karaali, Naomi Harte, Cláudio R. Jung |
IEEE Trans. Image Process. | 2 |
| 2021 | Dimensional perception of a 'smiling McGurk effect'abstractMultisensory integration influences emotional perception, as the McGurk effect demonstrates for the communication between humans. Human physiology implicitly links the production of visual features with other modes like the audio channel: Face muscles responsible for a smiling face also stretch the vocal cords that results in a characteristic smiling voice. For artificial agents capable of multimodal expression, this linkage is modeled explicitly. In our study, we observe the influence of visual and audio channel on the perception of the agent’s emotional state. We created two virtual characters to control for anthropomorphic appearance. We record videos of these agents either with matching or mismatching emotional expression in the audio and visual channel. In an online study we measured the agent’s perceived valence and arousal. Our results show that a matched smiling voice and smiling face increase both dimensions of the Circumplex model of emotions: ratings of valence and arousal grow. When the channels present conflicting information, any type of smiling results in higher arousal rating, but only the visual channel increases the perceived valence. When engineers are constrained in their design choices, we suggest they should give precedence to convey the artificial agent’s emotional state through the visual channel. Ilaria Torre 0002, Simon Holk, Emma Carrigan, Iolanda Leite, Rachel McDonnell, Naomi Harte |
ACII | 6 |
| 2021 | Learning to Count Words in Fluent Speech Enables Online Speech RecognitionabstractSequence to Sequence models, in particular the Transformer, achieve state of the art results in Automatic Speech Recognition. Practical usage is however limited to cases where full utterance latency is acceptable. In this work we introduce Taris, a Transformer-based online speech recognition system aided by an auxiliary task of incremental word counting. We use the cumulative word sum to dynamically segment speech and enable its eager decoding into words. Experiments performed on the LRS2, LibriSpeech, and Aishell-1 datasets of English and Mandarin speech show that the online system performs comparable with the offline one when having a dynamic algorithmic delay of 5 segments. Furthermore, we show that the estimated segment length distribution resembles the word length distribution obtained with forced alignment, although our system does not require an exact segment-to-word equivalence. Taris introduces a negligible overhead compared to a standard Transformer, while the local relationship modelling between inputs and outputs grants invariance to sequence length by design. George Sterpu, Christian Saam, Naomi Harte |
SLT | 3 |
| 2021 | The Effect of Audio-Visual Smiles on Social Influence in a Cooperative Human-Agent Interaction TaskabstractEmotional expressivity is essential for human interactions, informing both perception and decision-making. Here, we examine whether creating an audio-visual emotional channel mismatch influences decision-making in a cooperative task with a virtual character. We created a virtual character that was either congruent in its emotional expression (smiling in the face and voice) or incongruent (smiling in only one channel). People (N = 98) evaluated the character in terms of valence and arousal in an online study; then, visitors in a museum played the “lunar survival task” with the character over three experiments (N = 597, 78, 101, respectively). Exploratory results suggest that multi-modal expressions are perceived, and reacted upon, differently than unimodal expressions, supporting previous theories of audio-visual integration. Ilaria Torre 0002, Emma Carrigan, Katarina Domijan, Rachel McDonnell, Naomi Harte |
ACM Trans. Comput. Hum. Interact. | 5 |
| 2020 | Neural Generation of Dialogue Response TimingsabstractThe timings of spoken response offsets in human dialogue have been shown to vary based on contextual elements of the dialogue.We propose neural models that simulate the distributions of these response offsets, taking into account the response turn as well as the preceding turn.The models are designed to be integrated into the pipeline of an incremental spoken dialogue system (SDS).We evaluate our models using offline experiments as well as human listening tests.We show that human listeners consider certain response timings to be more natural based on the dialogue context.The introduction of these models into SDS pipelines could increase the perceived naturalness of interactions.1 Matthew Roddy, Naomi Harte |
ACL | 2 |
| 2020 | Cogans For Unsupervised Visual Speech Adaptation To New SpeakersabstractAudio-Visual Speech Recognition (AVSR) faces the difficult task of exploiting acoustic and visual cues simultaneously. Augmenting speech with the visual channel creates its own challenges, e.g. every person has unique mouth movements, making the generalization of visual models very difficult. This factor motivates our focus on the generalization of speaker-independent (SI) AVSR systems especially in noisy environments by exploiting the visual domain. Specifically, we are the first to explore the visual adaptation of an SI-AVSR system to an unknown and unlabelled speaker. We adapt an AVSR system trained in a source domain to decode samples in a target domain without the need for labels in the target domain. For the domain adaptation of the unknown speaker, we use Coupled Generative Adversarial Networks to automatically learn a joint distribution of multi-domain images. We evaluate our character-based AVSR system on the TCD-TIMIT dataset and obtain up to a 10% average improvement with respect to its AVSR system equivalent. Adriana Fernandez-Lopez, Ali Karaali, Naomi Harte, Federico Sukno |
ICASSP | 3 |
| 2020 | Can Auditory Nerve Models Tell us What's Different About WaveNet Vocoded Speech?
Sébastien Le Maguer, Naomi Harte |
INTERSPEECH | 2 |
| 2020 | Should we Hard-Code the Recurrence Concept or Learn it Instead ? Exploring the Transformer Architecture for Audio-Visual Speech RecognitionabstractThe audio-visual speech fusion strategy AV Align has shown significant performance improvements in audio-visual speech recognition (AVSR) on the challenging LRS2 dataset. Performance improvements range between 7% and 30% depending on the noise level when leveraging the visual modality of speech in addition to the auditory one. This work presents a variant of AV Align where the recurrent Long Short-term Memory (LSTM) computation block is replaced by the more recently proposed Transformer block. We compare the two methods, discussing in greater detail their strengths and weaknesses. We find that Transformers also learn cross-modal monotonic alignments, but suffer from the same visual convergence problems as the LSTM model, calling for a deeper investigation into the dominant modality problem in machine learning. George Sterpu, Christian Saam, Naomi Harte |
INTERSPEECH | 3 |
| 2020 | Investigation of Auditory Nerve Model Based Analysis for Vocoded Speech SynthesisabstractIn recent decades, the quality of speech synthesized by computers has increased drastically. However, evaluating such systems remains a challenge as the relevant methodologies haven't evolved for more than a decade. Subjective evaluation provides a global overview of the quality, but lacks any detailed feedback. Furthermore, research in objective evaluation hasn't yet delivered any detailed analysis methodologies. Inspired by the speech intelligibility and speech quality fields, we investigate how we can use an Auditory Nerve (AN) model to improve objective evaluation of speech synthesis systems. To do so, we compare different configurations of Hidden Markov Model (HMM) and deep neural network (DNN) synthesis using two different metrics derived from spectrograms, mean-rate neurograms and fine-timing neurograms. The metrics are the Root Mean Square Error (RMSE) and the Neurogram Similarity Index Measure (NSIM). As using an AN model introduces a perceptual angle in the analysis, we also compare the different configurations using two established perceptual-based quality models: Perceptual Evaluation of Speech Quality (PESQ) and Virtual Speech Quality Objective Listener (ViSQOL). The results show ViSQOL and PESQ are not suitable to a refined analysis of speech synthesis. The results also show that comparing mean-rate neurograms using the NSIM metric is an effective alternative to the comparison of spectrograms using the RMSE. Sébastien Le Maguer, Naomi Harte |
QoMEX | 2 |
| 2020 | How to Teach DNNs to Pay Attention to the Visual Modality in Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) seeks to model, and thereby exploit, the dynamic relationship between a human voice and the corresponding mouth movements. A recently proposed multimodal fusion strategy, AV Align, based on state-of-the-art sequence to sequence neural networks, attempts to model this relationship by explicitly aligning the acoustic and visual representations of speech. This study investigates the inner workings of AV Align and visualises the audio-visual alignment patterns. Our experiments are performed on two of the largest publicly available AVSR datasets, TCD-TIMIT and LRS2. We find that AV Align learns to align acoustic and visual representations of speech at the frame level on TCD-TIMIT in a generally monotonic pattern. We also determine the cause of initially seeing no improvement over audio-only speech recognition on the more challenging LRS2. We propose a regularisation method which involves predicting lip-related Action Units from visual representations. Our regularisation method leads to better exploitation of the visual modality, with performance improvements between 7% and 30% depending on the noise level. Furthermore, we show that the alternative Watch, Listen, Attend, and Spell network is affected by the same problem as AV Align, and that our proposed approach can effectively help it learn visual representations. Our findings validate the suitability of the regularisation method to AVSR and encourage researchers to rethink the multimodal convergence problem when having one dominant modality. George Sterpu, Christian Saam, Naomi Harte |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | The Effect of Multimodal Emotional Expression and Agent Appearance on Trust in Human-Agent InteractionabstractEmotional expressivity can boost trust in human-human and human-machine interaction. As a multimodal phenomenon, previous research argued that a mismatch in the expressive channels provides evidence of joint audio-video emotional processing. However, while previous work studied this from the point of view of emotion recognition and processing, not much is known about what effect a multimodal agent would have on a human-agent interaction task. Also, agent appearance could influence this interaction too. Here we manipulated the agent’s multimodal emotional expression (”smiling face” and ”smiling voice”, or both) and agent type (photorealistic or cartoon-like virtual human) and assessed people’s trust toward this agent. We measured trust using a mixed-methods approach, combining behavioural data from a survival task, questionnaire ratings and qualitative comments. These methods gave different results: while people commented on the importance of emotional expressivity in the agent’s voice, this factor had limited influence on trusting behaviours; while people rated the cartoon-like agent on several traits higher than the photorealistic one, the agent’s style also was not the most influential feature on people’s trusting behaviour. These results highlight the contribution of a mixed-methods approach in human-machine interaction, as both explicit and implicit perception and behaviour will contribute to the success of the interaction. Ilaria Torre 0002, Emma Carrigan, Rachel McDonnell, Katarina Domijan, Killian McCabe, Naomi Harte |
MIG | 6 |
| 2018 | Voice Activity Detection Using NeurogramsabstractExisting acoustic-signal-based algorithms for Voice Activity Detection (VAD) do not perform well in the presence of noise. In this study, we propose a method to improve VAD accuracy by employing another type of signal representation which is derived from the response of the human Auditory-Nerve (AN) system. The neural responses referred to as a neurogram are simulated using a computational model of the AN system for a range of Characteristic Frequencies (CFs). Features are extracted from neurograms using the Discrete Cosine Transform (DCT), and are then trained using a Multilayer Perceptron (MLP) classifier to predict the VAD intervals. The proposed method was evaluated using the QUT-NOISE-TIMIT corpus, and the NIST scoring algorithm for VAD was employed as an accuracy measure. The proposed neural-response-based method exhibited an overall better VAD accuracy over most of the existing methods. Wissam A. Jassim, Naomi Harte |
ICASSP | 2 |
| 2018 | The Impact of Reduced Video Quality on Visual Speech RecognitionabstractSpeech recognition technology has become widespread in recent years to the point where almost anyone with a laptop or mobile device has access to it. Despite this, it still poses the problem of poor recognition in noisy environments. Audio-Visual Speech Recognition (AVSR) provides a possible solution to this problem as the visual channel is not affected by the acoustic noise. However there are other factors that could impact the performance, namely poor quality of the video footage. This aspect of the visual side of speech recognition in noise is less explored, partially due to a lack of large, publicly available, high quality audio-visual continuous-speech databases. Fortunately, these problems can now be considered more fully with the availability of datasets such as TCD-TIMIT. In this paper, we explore the impact of the following visual degradations on visual speech recognition: white Gaussian noise, JPEG compression, reduced resolution and motion blur. Experimental results show that in some cases, the recogniser can be remarkably resilient, i.e. in the case of the motion blur, while other degradations can affect the recogniser performance drastically. Laura Dungan, Ali Karaali, Naomi Harte |
ICIP | 3 |
| 2018 | Can DNNs Learn to Lipread Full Sentences?abstractFinding visual features and suitable models for lipreading tasks that are more complex than a well-constrained vocabulary has proven challenging. This paper explores state-of-the-art Deep Neural Network architectures for lipreading based on a Sequence to Sequence Recurrent Neural Network. We report results for both hand-crafted and 2D/3D Convolutional Neural Network visual front-ends, online monotonic attention, and a joint Connectionist Temporal Classification-Sequence-to-Sequence loss. The system is evaluated on the publicly available TCD-TIMIT dataset, with 59 speakers and a vocabulary of over 6000 words. Results show a major improvement on a Hidden Markov Model framework. A fuller analysis of performance across visemes demonstrates that the network is not only learning the language model, but actually learning to lipread. George Sterpu, Christian Saam, Naomi Harte |
ICIP | 3 |
| 2018 | Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNsabstractIn human conversational interactions, turn-taking exchanges can be coordinated using cues from multiple modalities. To design spoken dialog systems that can conduct fluid interactions it is desirable to incorporate cues from separate modalities into turn-taking models. We propose that there is an appropriate temporal granularity at which modalities should be modeled. We design a multiscale RNN architecture to model modalities at separate timescales in a continuous manner. Our results show that modeling linguistic and acoustic features at separate temporal rates can be beneficial for turn-taking modeling. We also show that our approach can be used to incorporate gaze features into turn-taking models. Matthew Roddy, Gabriel Skantze, Naomi Harte |
ICMI | 3 |
| 2018 | Attention-based Audio-Visual Fusion for Robust Automatic Speech RecognitionabstractAutomatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy that goes beyond simple feature concatenation and learns to automatically align the two modalities, leading to enhanced representations which increase the recognition accuracy in both clean and noisy conditions. We test our strategy on the TCD-TIMIT and LRS2 datasets, designed for large vocabulary continuous speech recognition, applying three types of noise at different power ratios. We also exploit state of the art Sequence-to-Sequence architectures, showing that our method can be easily integrated. Results show relative improvements from 7% up to 30% on TCD-TIMIT over the acoustic modality alone, depending on the acoustic noise level. We anticipate that the fusion strategy can easily generalise to many other multimodal tasks which involve correlated modalities. George Sterpu, Christian Saam, Naomi Harte |
ICMI | 3 |
| 2018 | Survival at the Museum: A Cooperation Experiment with Emotionally Expressive Virtual CharactersabstractCorrectly interpreting an interlocutor's emotional expression is paramount to a successful interaction. But what happens when one of the interlocutors is a machine? The facilitation of human-machine communication and cooperation is of growing importance as smartphones, autonomous cars, or social robots increasingly pervade human social spaces. Previous research has shown that emotionally expressive virtual characters generally elicit higher cooperation and trust than 'neutral' ones. Since emotional expressions are multi-modal, and given that virtual characters can be designed to our liking in all their components, would a mismatch in the emotion expressed in the face and voice influence people's cooperation with a virtual character? We developed a game where people had to cooperate with a virtual character in order to survive on the moon. The character's face and voice were designed to either smile or not, resulting in 4 conditions: smiling voice and face, neutral voice and face, smiling voice only (neutral face), smiling face only (neutral voice). The experiment was set up in a museum over the course of several weeks; we report preliminary results from over 500 visitors, showing that people tend to trust the virtual character in the mismatched condition with the smiling face and neutral voice more. This might be because the two channels express different aspects of an emotion, as previously suggested. Ilaria Torre 0002, Emma Carrigan, Killian McCabe, Rachel McDonnell, Naomi Harte |
ICMI | 5 |
| 2018 | Investigating Speech Features for Continuous Turn-Taking Prediction Using LSTMsabstractFor spoken dialog systems to conduct fluid conversational interactions with users, the systems must be sensitive to turn-taking cues produced by a user. Models should be designed so that effective decisions can be made as to when it is appropriate, or not, for the system to speak. Traditional end-of-turn models, where decisions are made at utterance end-points, are limited in their ability to model fast turn-switches and overlap. A more flexible approach is to model turn-taking in a continuous manner using RNNs, where the system predicts speech probability scores for discrete frames within a future window. The continuous predictions represent generalized turn-taking behaviors observed in the training data and can be applied to make decisions that are not just limited to end-of-turn detection. In this paper, we investigate optimal speech-related feature sets for making predictions at pauses and overlaps in conversation. We find that while traditional acoustic features perform well, part-of-speech features generally perform worse than word features. We show that our current models outperform previously reported baselines. Matthew Roddy, Gabriel Skantze, Naomi Harte |
INTERSPEECH | 3 |
| 2018 | Perception and prediction of speaker appeal - A single speaker study
Ailbhe Cullen, Andrew Hines, Naomi Harte |
Comput. Speech Lang. | 3 |
| 2017 | Speech emotion classification using combined neurogram and INTERSPEECH 2010 paralinguistic challenge featuresabstractRecently, increasing attention has been directed to study and identify the emotional content of a spoken utterance. This study introduces a method to improve emotion classification performance under clean and noisy environments by combining two types of features: the proposed neural‐responses‐based features and the traditional INTERSPEECH 2010 paralinguistic emotion challenge features. The neural‐responses‐based features are represented by the responses of a computational model of the auditory system for listeners with normal hearing. The model simulates the responses of an auditory‐nerve fibre with a characteristic frequency to a speech signal. The simulated responses of the model are represented by the 2D neurogram (time‐frequency representation). The neurogram image is sub‐divided into non‐overlapped blocks and the averaged value of each block is computed. The neurogram features and the traditional emotion features are combined together to form the feature vector for each speech signal. The features are trained using support vector machines to predict the emotion of speech. The performance of the proposed method is evaluated on two well‐known databases: the eNTERFACE and Berlin emotional speech data set. The results show that the proposed method performed better when compared with the classification results obtained using neurogram and INTERSPEECH features separately. Wissam A. Jassim, Raveendran Paramesran, Naomi Harte |
IET Signal Process. | 3 |
| 2016 | Introduction
Naomi Harte, Peter Jancovic, Karl-L. Schuchmann |
INTERSPEECH | 1 |
| 2016 | Poster Overview Presentations
Naomi Harte, Peter Jancovic, Karl-L. Schuchmann |
INTERSPEECH | 1 |
| 2016 | Discussion
Naomi Harte, Peter Jancovic, Karl-L. Schuchmann |
INTERSPEECH | 1 |
| 2016 | Closing Remarks
Naomi Harte, Peter Jancovic, Karl-L. Schuchmann |
INTERSPEECH | 1 |
| 2016 | YIN-Bird: Improved Pitch Tracking for Bird Vocalisations
Colm O'Reilly, Nicola M. Marples, David J. Kelly, Naomi Harte |
INTERSPEECH | 4 |
| 2016 | Bitrate classification of twice-encoded audio using objective quality featuresabstractWhen a user uploads audio files to a music streaming service, these files are subsequently re-encoded to lower bitrates to target different devices, e.g. low bitrate for mobile. To save time and bandwidth uploading files, some users encode their original files using a lossy codec. The metadata for these files cannot always be trusted as users might have encoded their files more than once. Determining the lowest bitrate of the files allows the streaming service to skip the process of encoding the files to bitrates higher than that of the uploaded files, saving on processing and storage space. This paper presents a model that uses quality predictions from ViSQOLAudio, a full reference objective audio quality metric, as features in combination with a multi-class support vector machine classifier. An experiment on twice-encoded files found that low bitrate codecs could be classified using audio quality features. The experiment also provides insights into the implications of multiple transcodes from a quality perspective. Colm Sloan, Naomi Harte, Damien Kelly, Anil C. Kokaram, Andrew Hines |
QoMEX | 2 |
| 2015 | Measuring and monitoring speech quality for voice over IP with POLQA, viSQOL and p.563abstractThere are many types of degradation which can occur in Voice over IP (VoIP) calls. Of interest in this work are degradations which occur independently of the codec, hardware or network in use. Specifically, their effect on the subjective and objec- tive quality of the speech is examined. Since no dataset suit- able for this purpose exists, a new dataset (TCD-VoIP) has been created and has been made publicly available. The dataset con- tains speech clips suffering from a range of common call qual- ity degradations, as well as a set of subjective opinion scores on the clips from 24 listeners. The performances of three ob- jective quality metrics: POLQA, ViSQOL and P.563, have been evaluated using the dataset. The results show that full reference metrics are capable of accurately predicting a variety of com- mon VoIP degradations. They also highlight the outstanding need for a wideband, single-ended, no-reference metric to mon- itor accurately speech quality for degradations common in VoIP scenarios. Andrew Hines, Eoin Gillen, Naomi Harte |
INTERSPEECH | 3 |
| 2015 | Quantifying difference in vocalizations of bird populations
Colm O'Reilly, Nicola M. Marples, David J. Kelly, Naomi Harte |
INTERSPEECH | 4 |
| 2015 | TCD-TIMIT: An Audio-Visual Corpus of Continuous SpeechabstractAutomatic audio-visual speech recognition currently lags behind its audio-only counterpart in terms of major progress. One of the reasons commonly cited by researchers is the scarcity of suitable research corpora. This paper details the creation of a new corpus designed for continuous audio-visual speech recognition research . TCD-TIMIT consists of high-quality audio and video footage of 62 speakers reading a total of 6913 phonetically rich sentences. Three of the speakers are professionally-trained lipspeakers, recorded to test the hypothesis that lipspeakers may have an advantage over regular speakers in automatic visual speech recognition systems. Video footage was recorded from two angles: straight on, and at 30°. The paper outlines the recording of footage, and the required post-processing to yield video and audio clips for each sentence. Audio, visual, and joint audio-visual baseline experiments are reported. Separate experiments were run on the lipspeaker and non-lipspeaker data, and the results compared. Visual and audio-visual baseline results on the non-lipspeakers were low overall. Results on the lipspeakers were found to be significantly higher. It is hoped that as a publicly available database, TCD-TIMIT will now help further state of the art in audio-visual speech recognition research. Naomi Harte, Eoin Gillen |
IEEE Trans. Multim. | 1 |
| 2014 | Effect of long-term ageing on i-vector speaker verificationabstractAssessing the impact of ageing on biometric systems is an important challenge. In this paper, an i-vector speaker verifi-cation framework is used to evaluate the impact of long-term ageing on state-of-the-art speaker verification. Using the Trin-ity College Dublin Speaker Ageing (TCDSA) database, it is ob-served that the performance of the i-vector system, in terms of both discrimination and calibration, degrades progressively as the absolute age difference between training and testing sam-ples increases. In the case of male speakers, the equal error rate (EER) increases from 4.61 % at an ageing difference of 0–1 years to 32.74 % at an age difference of 51–60 years. The performance of a Gaussian Mixture Model- Universal Back-ground Model (GMM-UBM) system is presented for compari-son. It is shown that while the i-vector system outperforms the GMM-UBM system, as absolute age difference increases, the performance of both degrades at a similar rate. It is concluded that long-term ageing variability is distinct from everyday inter-session variability, and therefore must be dealt with via dedi-cated compensation strategies. Finnian Kelly, Rahim Saeidi, Naomi Harte, David A. van Leeuwen |
INTERSPEECH | 3 |
| 2014 | Perceived Audio Quality for Streaming Stereo MusicabstractUsers of audio-visual streaming services expect an ever increasing quality of experience. Channel bandwidth remains a bottleneck commonly addressed with lossy compression schemes for both the video and audio streams. Anecdotal evidence suggests a strongly perceived link between bit rate and quality. This paper presents three audio quality listening experiments using the ITU MUSHRA methodology to assess a number of audio codecs typically used by streaming services. They were assessed for a range of bit rates using three presentation modes: consumer and studio quality headphones and loudspeakers. Our results indicate that with consumer quality headphones, listeners were not differentiating between codecs with bit rates greater than 48 kb/s (p>=0.228). For studio quality headphones and loudspeakers aac-lc at 128 kb/s and higher was differentiated over other codecs (p<=0.001). The results provide insights into quality of experience that will guide future development of objective audio quality metrics. Andrew Hines, Eoin Gillen, Damien Kelly, Jan Skoglund, Anil C. Kokaram, Naomi Harte |
ACM Multimedia | 6 |
| 2013 | Robustness of speech quality metrics to background noise and network degradations: Comparing ViSQOL, PESQ and POLQAabstractThe Virtual Speech Quality Objective Listener (ViSQOL) is a new objective speech quality model. It is a signal based full reference metric that uses a spectro-temporal measure of similarity between a reference and a test speech signal. ViSQOL aims to predict the overall quality of experience for the end listener whether the cause of speech quality degradation is due to ambient noise, or transmission channel degradations. This paper describes the algorithm and tests the model using two speech corpora: NOIZEUS and E4. The NOIZEUS corpus contains speech under a variety of background noise types, speech enhancement methods, and SNR levels. The E4 corpus contains voice over IP degradations including packet loss, jitter and clock drift. The results are compared with the ITU-T objective models for speech quality: PESQ and POLQA. The behaviour of the metrics are also evaluated under simulated time warp conditions. The results show that for both datasets ViSQOL performed comparably with PESQ. POLQA was shown to have lower correlation with subjective scores than the other metrics for the NOIZEUS database. Andrew Hines, Jan Skoglund, Anil C. Kokaram, Naomi Harte |
ICASSP | 4 |
| 2013 | Identifying new bird species from differences in birdsong
Naomi Harte, Sadhbh Murphy, David J. Kelly, Nicola M. Marples |
INTERSPEECH | 1 |
| 2013 | Monitoring the effects of temporal clipping on voIP speech qualityabstractThis paper presents work on a real-time temporal clipping monitoring tool for VoIP. Temporal clipping can occur as a result of voice activity detection (VAD) or echo cancellation where comfort noise in used in place of clipped speech segments. The algorithm presented will form part of a no-reference objective model for quantifying perceived speech quality in VoIP. The overall approach uses a modular design that will help pinpoint the reason for degradations in addition to quantifying their impact on speech quality. The new algorithm was tested for VAD compared over a range of thresholds and varied speech frame sizes. The results are compared to objective Mean Opinion Scores (MOS-LQO) from POLQA. The results show that the proposed algorithm can efficiently predict temporal clipping in speech and correlates well with the full reference quality predictions from POLQA. The model shows good potential for use in a real-time monitoring tool. Index Terms: temporal clipping, VAD, VoIP, POLQA 1. Andrew Hines, Jan Skoglund, Anil C. Kokaram, Naomi Harte |
INTERSPEECH | 4 |
| 2013 | Eigenageing compensation for speaker verificationabstractDealing with the effect of vocal ageing on speaker verification is an important challenge. In this paper, a new approach to improving speaker verification performance in the presence of long-term ageing is presented. Analogous to eigenchannel compensation, the proposed eigenageing compensation method operates by adapting a speaker model to a test sample based on a predetermined ageing subspace. An experimental evaluation of the new method, using the Trinity College Dublin Speaker Ageing database, demonstrates it to be very effective at reducing long-term speaker verification error rates, and shows it to compare favourably with our previous stacked classifier technique. Index Terms: speaker verification, ageing, eigenanalysis 1. Finnian Kelly, Niko Brümmer, Naomi Harte |
INTERSPEECH | 3 |
| 2013 | Auditory detectability of vocal ageing and its effect on forensic automatic speaker recognitionabstractThe comparison of non-contemporary speech samples is common in forensic speaker recognition cases. It has yet to be established however, to what extent the time interval between non-contemporary samples can increase before a problem is created for forensic automatic speaker recognition. This paper presents results of a human listener test designed to evaluate the detectability of vocal ageing over increasing intervals of up to 30 years. Subsequently, a forensic automatic speaker recognition evaluation of 15 ageing males at increasing intervals of up to 60 years is presented. It is shown that at intervals of around 10 years, the average detectability of vocal ageing by humans is just above chance. As the interval rises to 30 years, vocal ageing is detected 90% of the time. In the automatic system, vocal ageing is manifested as a drop in intra-speaker likelihood ratios (LRs) as the time interval between non-contemporary samples increases. At an interval of 30 years, LRs for the vast majority of intra-speaker comparisons fall below a value of 100 commonly interpreted as ‘moderate support’ on a verbal LR scale. Our findings indicate that at a time-lapse of 30 years, vocal ageing creates significant problems for forensic automatic speaker recognition. Finnian Kelly, Naomi Harte |
INTERSPEECH | 2 |
| 2013 | Speaker verification in score-ageing-quality classification space
Finnian Kelly, Andrzej Drygajlo, Naomi Harte |
Comput. Speech Lang. | 3 |
| 2012 | Phoneme-to-viseme Mapping for Visual Speech Recognition
Luca Cappelletta, Naomi Harte |
ICPRAM (2) | 2 |
| 2012 | Improved Speech Intelligibility with a Chimaera Hearing Aid AlgorithmabstractIt is recognised that current hearing aid fitting algorithms can corrupt fine timing cues in speech.This paper presents a fitting algorithm that aims to improve speech intelligibility, while preserving the temporal fine structure.The algorithm combines the signal envelope amplification from a standard hearing aid fitting algorithm with the fine timing information available to unaided listeners.The proposed "chimaera aid" is evaluated with computer simulated listener tests to measure its speech intelligibility for 3 sample hearing losses.In addition, the experiment demonstrates the potential application of auditory nerve models in the development of new hearing aid algorithm designs using the previously developed Neurogram Similarity Index Measure (NSIM) to predict speech intelligibility.The results predict that the new aid restores envelope without degrading fine timing information. Andrew Hines, Naomi Harte |
INTERSPEECH | 2 |
| 2012 | Compensating for Ageing and Quality variation in Speaker VerificationabstractPerforming speaker verification in the simultaneous presence of ageing progression and changing speech sample quality is an important, open problem. The issues of ageing and quality variation go hand in hand; the effect of ageing increases with time, while variations in quality are also more likely to be encountered as time passes. In this work we demonstrate the effect of ageing on speaker verification performance, and show the relationship between quality variation and verification score via a range of established quality measures. We employ a stacked classifier framework to combine the output of the baseline verification system with ageing information and quality measures. This new approach to long-term speaker verification allows for a multi-dimensional decision boundary that significantly improves upon the baseline performance. The proposed framework is evaluated on the Trinity College Dublin Speaker Ageing Database. Index Terms: speaker verification, ageing, stacked classifier Finnian Kelly, Andrzej Drygajlo, Naomi Harte |
INTERSPEECH | 3 |
| 2012 | Speech intelligibility prediction using a Neurogram Similarity Index Measure
Andrew Hines, Naomi Harte |
Speech Commun. | 2 |
| 2012 | Algorithms for the Digital Restoration of Torn FilmsabstractThis paper presents algorithms for the digital restoration of films damaged by tear. As well as causing local image data loss, a tear results in a noticeable relative shift in the frame between the regions at either side of the tear boundary. This paper describes a method for delineating the tear boundary and for correcting the displacement. This is achieved using a graph-cut segmentation framework that can be either automatic or interactive when automatic segmentation is not possible. Using temporal intensity differences to form the boundary conditions for the segmentation facilitates the robust division of the frame. The resulting segmentation map is used to calculate and correct the relative displacement using a global-motion estimation approach based on motion histograms. A high-quality restoration is obtained when a suitable missing-data treatment algorithm is used to recover any missing pixel intensities. David Corrigan, Anil C. Kokaram, Naomi Harte |
IEEE Trans. Image Process. | 3 |
| 2010 | Auditory Features Revisited for Robust Speech RecognitionabstractAuditory based front-ends for speech recognition have been compared before, but this paper focuses on two of the most promising algorithms for noise robustness in automatic speech recognition (ASR). The feature sets are Zero-Crossings with Peak Amplitudes (ZCPA) and the recently introduced Power-Law Nonlinearity and Power-Bias Subtraction (PNCC). Standard Mel-Frequency Cepstral Coefficients (MFCC) are also tested for reference. The performance of all features is reported on the TIMIT database using a HMM-based recogniser. It is found that the PNCC features outperform MFCC in clean conditions and are most robust to noise. ZCPA performance is shown to vary widely with filter bank configuration and frame length. The ZCPA performance is poor in clean conditions but is the least affected by white noise. PNCC is shown to be the most promising new feature set for robust ASR in recent years. Finnian Kelly, Naomi Harte |
ICPR | 2 |
| 2010 | Speech intelligibility from image processing
Andrew Hines, Naomi Harte |
Speech Commun. | 2 |
| 2009 | Error metrics for impaired auditory nerve responses of different phoneme groupsabstractAn auditory nerve model allows faster investigation of new signal processing algorithms for hearing aids.This paper presents a study of the degradation of auditory nerve (AN) responses at a phonetic level for a range of sensorineural hearing losses and flat audiograms.The AN model of Zilany & Bruce was used to compute responses to a diverse set of phoneme rich sentences from the TIMIT database.The characteristics of both the average discharge rate and spike timing of the responses are discussed.The experiments demonstrate that a mean absolute error metric provides a useful measure of average discharge rates but a more complex measure is required to capture spike timing response errors. Andrew Hines, Naomi Harte |
INTERSPEECH | 2 |
| 2007 | Automated Segmentation of Torn Frames using the Graph Cuts TechniqueabstractFilm Tear is a form of degradation in archived film and is the physical ripping of the film material. Tear causes displacement of a region of the degraded frame and the loss of image data along the boundary of tear. In [1], a restoration algorithm was proposed to correct the displacement in the frame introduced by the tear by estimating the global motion of the 2 regions either side of the tear. However, the algorithm depended on a user-defined segmentation to divide the frame. This paper presents a new fully-automated segmentation algorithm which divides affected frames along the tear. The algorithm employs the graph cuts optimisation technique and uses temporal intensity differences, rather than spatial gradient, to describe the boundary properties of the segmentation. Segmentations produced with the proposed algorithm agree well with the perceived correct segmentation. David Corrigan, Naomi Harte, Anil C. Kokaram |
ICIP (1) | 2 |
| 2007 | Rotation Detection using the Curl EquationabstractRotational types of motion can often be seen in video sequences. However, not a lot of research has been done to investigate rotational motion models for use in video. Analysing this unique type of motion could be very useful. For example, if the the centre of rotation of a spinning object can be efficiently identified, extraction and tracking of it can be made easier by grouping points moving at the same radial speed. It could also improve compression by recording rotation variables. In this paper, we introduce a method for finding the centre of rotation of a rotating object and a basic approach for modelling the rotation for improved image quality. The method requires an initial block based translational motion field. Daire Lennon, Naomi Harte, Anil C. Kokaram |
ICIP (1) | 2 |
| 2006 | Exploiting Voicing Cues for Contrast Enhanced Frequency Shaping of Speech for Impaired ListenersabstractThis paper investigates the use of voicing information as an additional cue in contrast enhanced frequency shaping (CEFS) of speech to improve perception in the hearing impaired. The presented work builds on an existing system combining multiband compression with contrast enhanced frequency shaping (MICEFS) to restore the auditory nerve response of a hearing impaired listener. CEFS can improve the perception of voiced segments. Hence voicing cues are used to differentiate segments for processing. Alternative processing for unvoiced segments is investigated and shown to improve neural representation of unvoiced segments compared to using MICEFS processing alone. Naomi Harte, Shahab U. Ansari, Ian C. Bruce |
ICASSP (5) | 1 |
| 2006 | Pathological Motion Detection for Robust Missing Data Treatment in Degraded Archived MediaabstractThis paper outlines an algorithm to improve the robustness of missing data treatment to pathological motion (PM). PM can cause misdiagnosis of clean image data as missing data. The proposed algorithm uses a probabilistic framework to jointly detect PM and missing data by exploiting more temporal information than is typically used for missing data detection and by exploiting the local smoothness assumption of motion fields. The results of the framework are compared to an equivalent missing data detector without PM detection and the framework is shown to prevent the misdiagnosis of missing data due to PM. David Corrigan, Naomi Harte, Anil C. Kokaram |
ICIP | 2 |
| 1999 | Discriminative spectral-temporal multiresolution features for speech recognitionabstractMulti-resolution features, which are based on the premise that there may be more cues for phonetic discrimination in a given sub-band than in another, have been shown to outperform the standard MFCC feature set for both classification and recognition tasks on the TIMIT database. This paper presents an investigation into possible strategies to extend these ideas from the spectral domain into both the spectral and temporal domains. Experimental work on the integration of segmental models, which are better at capturing the longer term phonetic correlation of a phonetic unit, into the discriminative multi-resolution framework is presented. Results are presented which show that including this supplementary temporal information offers an improvement performance for the phoneme classification task over the standard multi-resolution MFCC feature set with time derivatives appended. Possible strategies for the extension of theses techniques into the area of continuous speech recognition are discussed. Philip McMahon, Naomi Harte, Saeed Vaseghi, Paul M. McCourt |
ICASSP | 2 |
| 1999 | Combined temporal and spectral multi-resolution phonetic modelling
Paul M. McCourt, Naomi Harte, Saeed Vaseghi |
EUROSPEECH | 2 |
| 1998 | Multi-resolution cepstral features for phoneme recognition across speech sub-bandsabstractMulti-resolution sub-band cepstral features strive to exploit discriminative cues in localised regions of the spectral domain by supplementing the full bandwidth cepstral features with sub-band cepstral features derived from several levels of sub-band decomposition. Multi-resolution feature vectors, formed by concatenation of the sub-band cepstral features into an extended feature vector, are shown to yield better performance than conventional MFCCs for phoneme recognition on the TIMIT database. Possible strategies for the recombination of partial recognition scores from independent multi-resolution sub-band models are explored. By exploiting the sub-band variations in signal to noise ratio for linearly weighted recombination of the log likelihood probabilities we obtained improved phoneme recognition performance in broadband noise compared to MFCC features. This is an advantage over a purely sub-band approach using non-linear recombination which is robust only to narrow band noise. Paul M. McCourt, S. Vaseght, Naomi Harte |
ICASSP | 3 |
| 1998 | Joint recognition and segmentation using phonetically derived features and a hybrid phoneme model
Naomi Harte, Saeed Vaseghi, Ben P. Milner |
ICSLP | 1 |
| 1997 | Multi-resolution phonetic/segmental features and models for HMM-based speech recognitionabstractThis paper explores the modelling of phonetic segments of speech with multi-resolution spectral/time correlates. For spectral representation a set of multi-resolution cepstral features are proposed. Cepstral features obtained from a DCT of the log energy-spectrum over the full voice-bandwidth (100-4000 Hz) are combined with higher resolution features obtained from the DCT of upper subband (say 100-2100) and lower subband (2100-4000) halves. This approach can be extended to several levels of different resolutions. For representation of the temporal structure of speech segments or phonetic units, the conventional cepstral and dynamic cepstral features representing speech at the sub-phonetic levels, are supplemented by a set of phonetic features that describe the trajectory of speech over the duration of a phonetic unit. A conditional probability model for phonetic and sub-phonetic features is considered. Experiments demonstrate that the inclusion of the segmental features result in about 10% decrease in error rates. Saeed Vaseghi, Naomi Harte, Ben P. Milner |
ICASSP | 2 |
| 1996 | Dynamic features for segmental speech recognition
Naomi Harte, Saeed Vaseghi, Ben P. Milner |
ICSLP | 1 |