Andrew Hines

dblp:43/8694 · DBLP profile ↗
← Back
44ranked-venue papers
9as first author
20since 2021 · last 2027
0000-0001-9636-2556ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 9 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 15 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2027 The Closer, The Better: How does voice similarity affect our preference for synthetic voices?
abstract
This paper investigates the impact of voice similarity on synthetic voice preference and Quality of Experience (QoE), focusing on two key prosodic features: long-term fundamental frequency (LTF0) or characteristic pitch and speech rate (SR). By conducting subjective tests on synthetic voices with varying degrees of similarity to each participant’s own voice, the study aims to uncover whether and how voice similarity influences synthetic voice preference. Thirty-four participants were recorded reading 10 phonetically balanced Harvard sentences. From these data, we extracted LTF0 using CREPE and SR using Wav2Vec-XLSR-53 . The same features were extracted from four commercial Google TTS voices. The LTF0 and SR of the TTS voices were then matched with the voices of each participant. In addition to the matched voices, four variations were made for LTF0, while two speed variations were made for SR. Personalised listening tests were prepared and administered to each participant, comparing the matched stimulus with a variation in random order. The results show that the participants preferred the matched stimuli in most cases. In particular, this preference was stronger when voice matching was applied to female TTS voices. For male TTS voices, lower LTF0 was preferred. Moreover, matched and faster SR were preferred over slower SR. These findings suggest that voice adaptation and personalisation can enhance perceived QoE in terms of user preference.
Crisron Rudolf Lucas, Andrew Hines
Comput. Speech Lang.2
2026 Audio Made Simple: A Modern Framework for Audio Processing
abstract
This paper presents AudioSamples, an audio processing system that treats audio as a first-class data type rather than as generic numerical arrays accompanied by externally managed metadata. In many existing libraries, researchers must manually propagate sampling parameters such as sample rate, channel layout, and format across function calls, introducing cognitive overhead, silent error modes, and reproducibility risks. AudioSamples eliminates this coordination burden by embedding semantic properties directly within the audio object and automatically preserving them across processing pipelines, while requiring explicit and well-defined conversions when semantics change.
Jack Geraghty, Fatemeh Golpayegani, Andrew Hines
MMSys3
2025 Attention Models and Auditory Transduction Features for Noise Robustness
Cathal Ó Faoláin, Andrew Hines
INTERSPEECH2
2025 EgoMusic: An Egocentric Augmented Reality Glasses Dataset for Music
abstract
Although audio-augmented reality (AAR) has known applications in music, the use of wearables such as augmented reality (AR) glasses for egocentric audio data capture for music has not been investigated. Current egocentric datasets are mostly focused on speech research, neglecting music's unique demands for tasks such as real-time optimisation or assistive listening. This paper introduces EgoMusic, a multimodal dataset featuring synchronised egocentric audio-visual data captured with AR glasses during live performances, alongside studio-quality audio references. We investigate AR glasses' utility for music and baseline artificial intelligence (AI) approaches for hearing enhancement, positioning EgoMusic as the first dataset that enables research for egocentric music AAR.
Alessandro Ragano, Carl Timothy Tolentino, Kata Szita, Dan Barry, Davoud Shariat Panah, Niall Murray, Andrew Hines
ACM Multimedia7
2025 Learning to Associate: Multimodal Inference with Fully Missing Modalities
abstract
In this article, we propose Cross-Modal Association Models (C-MAMs), a novel approach for handling missing modalities during inference in multimodal learning. Unlike existing methods that modify the training process, C-MAMs generate missing modality features post-training , preserving the integrity of the original multimodal model. In this article, we: (i) formalise the problem of missing modality inference and its challenges, (ii) introduce C-MAMs as a flexible, lightweight, post-hoc solution for reconstructing missing modality embeddings, (iii) evaluate their effectiveness across diverse datasets, tasks and baseline models, and (iv) analyse the quality of the generated versus the ground-truth features to quantify the reconstruction fidelity. Experimental results show that C-MAMs significantly mitigate performance degradation due to missing modalities, in some cases fully restoring baseline performance, even when trained on 10% of the data. We conclude that post-training feature reconstruction is an effective, targeted alternative to existing methods, with broad applicability in multimodal systems.
Jack Geraghty, Andrew Hines, Fatemeh Golpayegani
ACM Trans. Intell. Syst. Technol.2
2025 Beyond Correlation: Evaluating Multimedia Quality Models With the Constrained Concordance Index
abstract
This study investigates the evaluation of multimedia quality models, focusing on the inherent uncertainties in subjective Mean Opinion Score (MOS) ratings due to factors like rater inconsistency and bias. Traditional statistical measures such as Pearson's Correlation Coefficient (PCC), Spearman's Rank Correlation Coefficient (SRCC), and Kendall's Tau (KTAU) often fail to account for these uncertainties, leading to inaccuracies in model performance assessment. We introduce the Constrained Concordance Index (CCI), a novel metric designed to overcome the limitations of existing metrics by considering the statistical significance of MOS differences and excluding comparisons where MOS confidence intervals overlap. Through comprehensive experiments across various domains including speech and image quality assessment, we demonstrate that CCI provides a more robust and accurate evaluation of instrumental quality models, especially in scenarios of low sample sizes, rater group variability, and restriction of range. Our findings suggest that incorporating rater subjectivity and focusing on statistically significant pairs can significantly enhance the evaluation framework for multimedia quality prediction models. This work not only sheds light on the overlooked aspects of subjective rating uncertainties but also proposes a methodological advancement for more reliable and accurate quality model evaluation.
Alessandro Ragano, Helard Becerra Martinez, Andrew Hines
IEEE Trans. Multim.3
2024 NOMAD: Unsupervised Learning of Perceptual Embeddings For Speech Enhancement and Non-Matching Reference Audio Quality Assessment
abstract
This paper presents NOMAD (Non-Matching Audio Distance), a differentiable perceptual similarity metric that measures the distance of a degraded signal against non-matching references. The proposed method is based on learning deep feature embeddings via a triplet loss guided by the Neurogram Similarity Index Measure (NSIM) to capture degradation intensity. During inference, the similarity score between any two audio samples is computed through Euclidean distance of their embeddings. NOMAD is fully unsupervised and can be used in general perceptual audio tasks for audio analysis e.g. quality assessment and generative tasks such as speech enhancement and speech synthesis. The proposed method is evaluated with 3 tasks. Ranking degradation intensity, predicting speech quality, and as a loss function for speech enhancement. Results indicate NOMAD outperforms other non-matching reference approaches in both ranking degradation intensity and quality assessment, exhibiting competitive performance with full-reference audio metrics. NOMAD demonstrates a promising technique that mimics human capabilities in assessing audio quality with non-matching references to learn perceptual embeddings without the need for human-generated labels.
Alessandro Ragano, Jan Skoglund, Andrew Hines
ICASSP3
2024 Reduce, Reuse, Recycle: Is Perturbed Data Better than Other Language Augmentation for Low Resource Self-Supervised Speech Models
Alessandro Ragano, Andrew Hines
INTERSPEECH3
2024 Impact of Auditory and Audiovisual Distractors on Task Performance in a VR-based Auditory Attention Task
abstract
Auditory attention is a fundamental cognitive process essential for effective communication and interaction. Auditory stimuli are often accompanied by distractors that can significantly impact task performance by means of reducing attention. This study investigates the influence of auditory and audiovisual distractors on an auditory attention task within a Virtual Reality classroom. As part of the user evaluation, participants had to listen to two short stories. They performed two tasks: (i) “DISTRACTORS”, where participants had to identify an auditory or an audiovisual stimulus and (ii) “KEYWORDS,” where participants had to identify a specific keyword in the story by pressing a button on the controller. During the experiment, the participants’ physiological (e.g. skin conductance levels, gaze data, etc.) and subjective data (i.e. questionnaires) were collected. The results revealed that the interval between the presentation of the keyword and the presentation of the distractors impacted task performance by negatively affecting auditory attention. Also, it was observed that the time spent looking at the speaker telling the story positively correlated with task performance, whereas the time spent looking at distractors was found to be negatively correlated with task performance. Finally, this study gives insights into how physiological metrics can be used to infer auditory attention in VR experiences.
Adrielle Nazar Moraes, Eoghan Hynes, Ronan Flynn, Andrew Hines, Niall Murray
ISMAR4
2024 SCOREQ: Speech Quality Assessment with Contrastive Regression
abstract
In this paper, we present SCOREQ, a novel approach for speech quality prediction. SCOREQ is a triplet loss function for contrastive regression that addresses the domain generalisation shortcoming exhibited by state of the art no-reference speech quality metrics. In the paper we: (i) illustrate the problem of L2 loss training failing at capturing the continuous nature of the mean opinion score (MOS) labels; (ii) demonstrate the lack of generalisation through a benchmarking evaluation across several speech domains; (iii) outline our approach and explore the impact of the architectural design decisions through incremental evaluation; (iv) evaluate the final model against state of the art models for a wide variety of data and domains. The results show that the lack of generalisation observed in state of the art speech quality metrics is addressed by SCOREQ. We conclude that using a triplet loss function for contrastive regression improves generalisation for speech quality prediction models but also has potential utility across a wide range of applications using regression-based predictive models.
Alessandro Ragano, Jan Skoglund, Andrew Hines
NeurIPS3
2024 Listenability Assessment for Language Learning Through the Quality of Experience Lens
abstract
The assessment of listenability, or the comprehensibility of spoken materials, is important for listening practice as it enables language learners to identify suitable learning materials. In this paper we present a new Listenability Framework using the Quality of Experience principles as a basis for developing an understanding of the listenability assessment process. The Listenability Framework promotes a holistic approach to the assessment of spoken materials, incorporating influencing factors derived from listening comprehension research and listenability assessment studies. We envision our framework as an aid in the development of automated assessment systems, as well as a means of examining existing systems for further development. We apply our framework in reviewing the current state of research on automated listenability assessment and identifying challenges and avenues for future work.
Michael Gringo Angelo Bayona, Andrew Hines, Elaine Uí Dhonnchadha
QoMEX2
2023 Audio Quality Assessment of Vinyl Music Collections Using Self-Supervised Learning
abstract
Metadata such as mean opinion score (MOS) quality ratings are critical to improve the usability and accessibility of music archive collections. Developing a non-intrusive objective quality metric that predicts MOS of archive music collections is challenging, since it requires labeling large datasets made of real-world recordings, which currently do not exist for this task. In this paper, we show that the self-supervised learning (SSL) model wav2vec 2.0 can be successfully used to predict the perceived audio quality of archive music collections. Using vinyl recordings, we evaluated wav2vec 2.0 on a new dataset of 620 tracks labeled with crowdsourcing. The proposed model shows superior performance to perceptual measures adapted from speech quality prediction. Finally, we propose a new evaluation metric called pairwise ranking accuracy (PRA) that takes into account subjective rater uncertainty by measuring the ability of an objective metric to rank pairs with high-confidence labels.
Alessandro Ragano, Emmanouil Benetos, Andrew Hines
ICASSP3
2023 A Comparison of Gender Differences and Performance Metrics in a VR-Based Auditory Selective Task
abstract
Virtual Reality has gained significant interest with the advancement of display and processing technologies in recent years. Traditionally, visual stimuli were the main point of interest in the development of new VR applications. However, audio is also key to the success of such experiences. Audio technology is responsible for making the VR environment more immersive as it is directly correlated with how we perceive the world. Our real-world environments are multimodal and multisensory. In this work, we investigate the effect of two types of distractors in an auditory selective attention task. Users were immersed in a virtual classroom with spatialised audio enabled. The task itself was divided into two subtasks. First, participants had to identify different types of auditory and audiovisual stimuli in the scene. Second, they had to focus their attention on a speaker in front of them and identify a keyword from the story each time they heard it. Furthermore, participants were divided into three groups that differ on when the distractor is presented in the second subtask: (a) before, (b) same time, and (c) after the keyword. Findings from this study show that there are significant gender differences for listeners immersed in an environment with competing sounds in recalling a story. Moreover, the time when distractors are presented significantly affected the response time, being inversely proportional to the mean response time.
Adrielle Nazar Moraes, Ronan Flynn, Andrew Hines, Niall Murray
QoMEX3
2022 Using Rater and System Metadata to Explain Variance in the VoiceMOS Challenge 2022 Dataset
Michael Chinen, Jan Skoglund, Chandan K. A. Reddy, Alessandro Ragano, Andrew Hines
INTERSPEECH5
2022 Exploring the influence of fine-tuning data on wav2vec 2.0 model for blind speech quality prediction
Helard Becerra Martinez, Alessandro Ragano, Andrew Hines
INTERSPEECH3
2022 AQP: an open modular Python platform for objective speech and audio quality metrics
abstract
Audio quality assessment has been widely researched in the signal processing area. Full-reference objective metrics (e.g., POLQA, ViSQOL) have been developed to estimate the audio quality relying only on human rating experiments. To evaluate the audio quality of novel audio processing techniques, researchers constantly need to compare objective quality metrics. Testing different implementations of the same metric and evaluating new datasets are fundamental and ongoing iterative activities. In this paper, we present AQP - an open-source, node-based, light-weight Python pipeline for audio quality assessment. AQP allows researchers to test and compare objective quality metrics helping to improve robustness, reproducibility and development speed. We introduce the platform, explain the motivations, and illustrate with examples how, using AQP, objective quality metrics can be (i) compared and benchmarked; (ii) prototyped and adapted in a modular fashion; (iii) visualised and checked for errors. The code has been shared on GitHub to encourage adoption and contributions from the community.
Jack Geraghty, Jiazheng Li 0002, Alessandro Ragano, Andrew Hines
MMSys4
2022 See hear now: is audio-visual QoE now just a fusion of audio and video metrics?
abstract
Single-modal audio/speech and video quality models have reached high levels of performance. Although traditional algorithms are still preferred for many practical applications, advances in machine learning (ML) and deep learning techniques have exceeded their performance in several scientific comparisons. However, audio-visual (AV) models have received signifi-cantly less attention and development. Despite the acknowledged challenge that multimodal interaction poses to the AV problem, traditional AV models generally rely on simple fusion techniques of individual audio and video predictions. Consequently, the impact of recent advances in single-modal quality assessment models on SOTA (state-of-the-art) AV quality models merits attention. This paper presents a revised and updated benchmark for AV quality assessment with particular focus on new speech quality metrics. Three AV datasets were used to test audio, video, and AV quality metrics. For audio and video, the best performing metrics were selected to build simple late-fusion models using their raw predictions. The fused models were then compared to the SOTA AV models. Results show that a simple fusion strategy produces accurate AV quality predictions (LCC and SCC greater than 0.90) with low error rates (RMSE lower than 0.33). These results highlight the influence of advances in speech quality for AV quality assessment.
Helard Becerra Martinez, Andrew Hines, Mylène C. Q. Farias
QoMEX2
2022 Speech quality assessment with WARP-Q: From similarity to subsequence dynamic time warp cost
abstract
Abstract Speech coding has been shown to achieve good speech quality using either waveform matching or parametric reconstruction. For very low bit rate streams, recently developed generative speech models can reconstruct high‐quality wideband speech from the bit streams of standard parametric encoders at less than 3 kb/s. Generative codecs produce high‐quality speech based on synthesising speech from a DNN and the parametric input. Existing objective speech quality models (e.g., ViSQOL and POLQA) cannot be used to accurately evaluate the quality of coded speech from generative models as they penalise based on signal differences not apparent in subjective listening test results. This paper presents WARP‐Q, a full‐reference objective speech quality metric that uses a dynamic time warping cost for MFCC representations of the signals. It is robust to low perceptual signal changes introduced by low bit rate neural vocoders. An evaluation using waveform matching, parametric, and generative neural vocoder‐based codecs as well as channel and environmental noise shows that WARP‐Q has better correlation and codec quality ranking for novel codecs compared to traditional metrics as well as the versatility of capturing other types of degradations, such as additive noise and transmission channel degradations.
Wissam A. Jassim, Jan Skoglund, Michael Chinen, Andrew Hines
IET Signal Process.4
2021 Warp-Q: Quality Prediction for Generative Neural Speech Codecs
abstract
Good speech quality has been achieved using waveform matching and parametric reconstruction coders. Recently developed very low bit rate generative codecs can reconstruct high quality wideband speech with bit streams less than 3 kb/s. These codecs use a DNN with parametric input to synthesise high quality speech outputs. Existing objective speech quality models (e.g., POLQA, ViSQOL) do not accurately predict the quality of coded speech from these generative models underestimating quality due to signal differences not highlighted in subjective listening tests. We present WARP-Q, a full-reference objective speech quality metric that uses dynamic time warping cost for MFCC speech representations. It is robust to small perceptual signal changes. Evaluation using waveform matching, parametric and generative neural vocoder based codecs as well as channel and environmental noise shows that WARP-Q has better correlation and codec quality ranking for novel codecs compared to traditional metrics in addition to versatility for general quality assessment scenarios.
Wissam A. Jassim, Jan Skoglund, Michael Chinen, Andrew Hines
ICASSP4
2021 More for Less: Non-Intrusive Speech Quality Assessment with Limited Annotations
abstract
Non-intrusive speech quality assessment is a crucial operation in multimedia applications. The scarcity of annotated data and the lack of a reference signal represent some of the main challenges for designing efficient quality assessment metrics. In this paper, we propose two multi-task models to tackle the problems above. In the first model, we first learn a feature representation with a degradation classifier on a large dataset. Then we perform MOS prediction and degradation classification simultaneously on a small dataset annotated with MOS. In the second approach, the initial stage consists of learning features with a deep clustering-based unsupervised feature representation on the large dataset. Next, we perform MOS prediction and cluster label classification simultaneously on a small dataset. The results show that the deep clustering-based model outperforms the degradation classifier-based model and the 3 baselines (autoencoder features, P.563, and SRMRnorm) on TCD-VoIP. This paper indicates that multi-task learning combined with feature representations from unlabelled data is a promising approach to deal with the lack of large MOS annotated datasets.
Alessandro Ragano, Emmanouil Benetos, Andrew Hines
QoMEX3
2020 Development of a Speech Quality Database Under Uncontrolled Conditions
abstract
INTERSPEECH 2020, Shanghai, China (held online due to coronavirus outbreak), 25-29 October 2020
Alessandro Ragano, Emmanouil Benetos, Andrew Hines
INTERSPEECH3
2020 Audio Inpainting based on Self-similarity for Sound Source Separation Applications
abstract
Sound source separation algorithms have advanced significantly in recent years but many algorithms can suffer from objectionable artefacts. The artefacts include phasiness, transient smearing, high frequency loss, unnatural sounding noise floor and reverberation to name a few. One of the main reasons for this is due to the fact that in many algorithm, individual time-frequency bins are often only attributed to one source at a time, meaning that many time-frequency bins will be set to zero for a separated source. This leads to an impressive signal to interference ratio but at the cost of natural sounding resynthesis. Here, we present a simple algorithm capable of audio inpainting based on self-similarity within the signal. The algorithm attempts to use the non-zero bin values observed in similar frames (past or future) as substitutes for the zero bin values in the current analysis frame. We present results from subjective listening tests which show a preference for the inpainted audio over the original audio produced from a simple source separation algorithm. Further, we use the Fréchet Audio Distance metric to evaluate the perceptual effect of the proposed inpainting algorithm. The results of this evaluation support the subjective test preferences.
Dan Barry, Alessandro Ragano, Andrew Hines
MMSP3
2020 How Crisp is the Crease? A Subjective Study on Web Browsing Perception of Above-The-Fold
abstract
Quality of Experience (QoE) for various types of websites has gained significant attention in recent years. In order to design and evaluate websites, a metric that can estimate a user's experienced quality robustly for diverse content is necessary. SpeedIndex (SI) has been widely adopted to estimate perceived web page loading progress. It measures the speed of rendering pixels for the webpage that is visible in the browser window. This is termed Above-The-Fold (ATF). The influence of animated content on the perception of ATF has been less comprehensively explored. In this paper, we present an experimental design and methodology to measure ATF perception for websites with and without animated elements for various page content categories. We found that pages with animated elements caused people to have more varied perceptions of ATF under different network conditions. Animated content also impacts the page load estimation accuracy of SI for websites. We discuss how the difference in the perception of ATF will impact the QoE management of web applications. We explain the necessity of revisiting the visual assessment of ATF to include the animated contents and improve the robustness of metrics like SI.
Hamed Z. Jahromi, Declan T. Delaney, Andrew Hines
NetSoft3
2020 ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric
abstract
Estimation of perceptual quality in audio and speech is possible using a variety of methods. The combined v3 release of ViSQOL and ViSQOLAudio (for speech and audio, respectively,) provides improvements upon previous versions, in terms of both design and usage. As an open source C++ library or binary with permissive licensing, ViSQOL can now be deployed beyond the research context into production usage. The feedback from internal production teams at Google has helped to improve this new release, and serves to show cases where it is most applicable, as well as to highlight limitations. The new model is benchmarked against real-world data for evaluation purposes. The trends and direction of future work is discussed.
Michael Chinen, Felicia Lim, Jan Skoglund, Nikita Gureev, Feargus O'Gorman, Andrew Hines
QoMEX6
2020 You Drive Me Crazy! Interactive QoE Assessment for Telepresence Robot Control
abstract
Telepresence robots (TPRs) are versatile, remotely controlled vehicles that enable physical presence and human-to-human interaction over a distance. Thanks to improving hardware and dropping price points, TPRs enjoy the growing interest in various industries and application domains. Still, a satisfying experience remains key for their acceptance and successful adoption, not only in terms of enabling remote communication with others, but also in terms of managing robot mobility by means of remote navigation. This paper focuses on the latter aspect of remote operation which has been hitherto neglected. We present the results of an extensive subjective study designed to systematically assess remote navigation Quality of Experience (QoE) in the context of using a TPR live over the Internet. Participants were ‘beamed’ into a remote office space and asked to perform characteristic TPR remote operation tasks (driving, turning, parking). Visual and control dimensions of their experience were systematically impaired by altering network characteristics (bandwidth, delay and packet loss rate) in a controlled fashion. Our results show that users can differentiate well between visual and navigation/control aspects of their experience. Furthermore, QoE impairment sensitivity varies with the actual task at hand.
Hamed Z. Jahromi, Ivan Bartolec, Edwin Gamboa, Andrew Hines, Raimund Schatz
QoMEX4
2020 Speech Quality Factors for Traditional and Neural-Based Low Bit Rate Vocoders
abstract
This study compares the performances of different algorithms for coding speech at low bit rates. In addition to widely deployed traditional vocoders, a selection of recently developed generative-model-based coders at different bit rates are contrasted. Performance analysis of the coded speech is evaluated for different quality aspects: accuracy of pitch periods estimation, the word error rates for automatic speech recognition, and the influence of speaker gender and coding delays. A number of performance metrics of speech samples taken from a publicly available database were compared with subjective scores. Results from subjective quality assessment do not correlate well with existing full reference speech quality metrics. The results provide valuable insights into aspects of the speech signal that will be used to develop a novel metric to accurately predict speech quality from generative-model-based coders.
Wissam A. Jassim, Jan Skoglund, Michael Chinen, Andrew Hines
QoMEX4
2020 How Deep is Your Encoder: An Analysis of Features Descriptors for an Autoencoder-Based Audio-Visual Quality Metric
abstract
The development of audio-visual quality assessment models poses a number of challenges in order to obtain accurate predictions. One of these challenges is the modelling of the complex interaction that audio and visual stimuli have and how this interaction is interpreted by human users. The No-Reference Audio-Visual Quality Metric Based on a Deep Autoencoder (NAViDAd) deals with this problem from a machine learning perspective. The metric receives two sets of audio and video features descriptors and produces a low-dimensional set of features used to predict the audio-visual quality. A basic implementation of NAViDAd was able to produce accurate predictions tested with a range of different audio-visual databases. The current work performs an ablation study on the base architecture of the metric. Several modules are removed or re-trained using different configurations to have a better understanding of the metric functionality. The results presented in this study provided important feedback that allows us to understand the real capacity of the metric's architecture and eventually develop a much better audio-visual quality metric.
Helard Becerra Martinez, Andrew Hines, Mylène C. Q. Farias
QoMEX2
2020 Evaluating the User in a Sound Localisation Task in a Virtual Reality Application
abstract
Virtual reality (VR) has proven to be a powerful tool enabling the development of immersive multimedia experiences. Initially focused on entertainment, industry and academia have begun to adapt and develop immersive applications for the healthcare domain, with opportunities in terms of condition assessment, diagnosis and intervention. In the context of immersive applications, audio, and in particular spatial audio, plays an important role on the immersion level. In order to process this information, the auditory cortex uses spatial cues encoded in the sound to provide relevant information about the distance, intensity and direction of the sound source. However, many different types of listening disorders can affect this capability. One condition, central auditory processing disorder (CAPD), significantly affects a user's ability to discriminate between different sound sources. People who suffer with this condition, are incapable of processing sounds properly, which may be stressful and frustrating when doing tasks with complex sounds or in noisy environments. This can have a significant impact on a person's quality of life. In this paper, an immersive VR spatial audio application is presented. It enables us to evaluate the ability of users to specify or localise the source of a sound. An integrated sensing system continuously collects relevant data from the user in order to fully understand how to quantify and evaluate spatial auditory skills from a quality of experience (QoE) perspective. QoE gives insight into a user's state and behaviour. To perform a detailed QoE evaluation of the listening task, implicit and explicit metrics were collected from the user. These included: self-reporting questionnaires, localisation performance, and physiological metrics. Data collected from this QoE evaluation gives an insight into a user's abilities to localise sound sources in VR, and also provides information on behaviour and effort (workload) in performing the task.
Adrielle Nazar Moraes, Ronan Flynn, Andrew Hines, Niall Murray
QoMEX3
2020 Audio Impairment Recognition using a Correlation-Based Feature Representation
abstract
Audio impairment recognition is based on finding noise in audio files and categorising the impairment type. Recently, significant performance improvement has been obtained thanks to the usage of advanced deep learning models. However, feature robustness is still an unresolved issue and it is one of the main reasons why we need powerful deep learning architectures. In the presence of a variety of musical styles, handcrafted features are less efficient in capturing audio degradation characteristics and they are prone to failure when recognising audio impairments and could mistakenly learn musical concepts rather than impairment types. In this paper, we propose a new representation of hand-crafted features that is based on the correlation of feature pairs. We experimentally compare the proposed correlation-based feature representation with a typical raw feature representation used in machine learning and we show superior performance in terms of compact feature dimensionality and improved computational speed in the test stage whilst achieving comparable accuracy.
Alessandro Ragano, Emmanouil Benetos, Andrew Hines
QoMEX3
2020 5G network slicing using SDN and NFV: A survey of taxonomy, architectures and future challenges
abstract
The increasing consumption of multimedia services and the demand of high-quality services from customers has triggered a fundamental change in how we administer networks in terms of abstraction, separation, and mapping of forwarding, control and management aspects of services. The industry and the academia are embracing 5G as the future network capable to support next generation vertical applications with different service requirements. To realize this vision in 5G network, the physical network has to be sliced into multiple isolated logical networks of varying sizes and structures which are dedicated to different types of services based on their requirements with different characteristics and requirements (e.g., a slice for massive IoT devices, smartphones or autonomous cars, etc.). Softwarization using Software-Defined Networking (SDN) and Network Function Virtualization (NFV)in 5G networks are expected to fill the void of programmable control and management of network resources. In this paper, we provide a comprehensive review and updated solutions related to 5G network slicing using SDN and NFV. Firstly, we present 5G service quality and business requirements followed by a description of 5G network softwarization and slicing paradigms including essential concepts, history and different use cases. Secondly, we provide a tutorial of 5G network slicing technology enablers including SDN, NFV, MEC, cloud/Fog computing, network hypervisors, virtual machines & containers. Thidly, we comprehensively survey different industrial initiatives and projects that are pushing forward the adoption of SDN and NFV in accelerating 5G network slicing. A comparison of various 5G architectural approaches in terms of practical implementations, technology adoptions and deployment strategies is presented. Moreover, we provide a discussion on various open source orchestrators and proof of concepts representing industrial contribution. The work also investigates the standardization efforts in 5G networks regarding network slicing and softwarization. Additionally, the article presents the management and orchestration of network slices in a single domain followed by a comprehensive survey of management and orchestration approaches in 5G network slicing across multiple domains while supporting multiple tenants. Furthermore, we highlight the future challenges and research directions regarding network softwarization and slicing using SDN and NFV in 5G networks.
Alcardo Alex Barakabitze, Arslan Ahmad, Rashid Mijumbi, Andrew Hines
Comput. Networks4
2019 A No-Reference Autoencoder Video Quality Metric
abstract
In this work, we introduce the No-reference Autoencoder VidEo (NAVE) quality metric, which is based on a deep au-toencoder machine learning technique. The metric uses a set of spatial and temporal features to estimate the overall visual quality, taking advantage of the autoencoder ability to produce a better and more compact set of features. NAVE was tested on two databases: the UnB-AVQ database and the LiveNetflix-II database. Results show that the method is able to estimate the perceived video quality with a good correlation performance and a small error, when compared to currently available no-reference and full-reference video quality objective metrics.
Helard Becerra Martinez, Mylène C. Q. Farias, Andrew Hines
ICIP3
2019 Adapting the Quality of Experience Framework for Audio Archive Evaluation
abstract
Perceived quality of historical audio material that is subjected to digitisation and restoration is typically evaluated by individual judgements or with inappropriate objective quality models. This paper presents a Quality of Experience (QoE) framework for predicting perceived audio quality of sound archives. The approach consists in adapting concepts used in QoE evaluation to digital audio archives. Limitations of current objective quality models employed in audio archives are provided and reasons why a QoE-based framework can overcome these limitations are discussed. This paper shows that applying a QoE framework to audio archives is feasible and it helps to identify the stages, stakeholders and models for a QoE centric approach.
Alessandro Ragano, Emmanouil Benetos, Andrew Hines
QoMEX3
2018 AMBIQUAL - a full reference objective quality metric for ambisonic spatial audio
abstract
Streaming spatial audio over networks requires efficient encoding techniques that compress the raw audio content without compromising quality of experience. Streaming service providers such as YouTube need a perceptually relevant objective audio quality metric to monitor users' perceived quality and spatial localization accuracy. In this paper we introduce a full reference objective spatial audio quality metric, AMBIQUAL, which assesses both Listening Quality and Localization Accuracy. In our solution both metrics are derived directly from the B-format Ambisonic audio. The metric extends and adapts the algorithm used in ViSQOLAudio, a full reference objective metric designed for assessing speech and audio quality. In particular, Listening Quality is derived from the omnidirectional channel and Localization Accuracy is derived from a weighted sum of similarity from B-format directional channels. This paper evaluates whether the proposed AMBIQUAL objective spatial audio quality metric can predict two factors: Listening Quality and Localization Accuracy by comparing its predictions with results from MUSHRA subjective listening tests. In particular, we evaluated the Listening Quality and Localization Accuracy of First and Third-Order Ambisonic audio compressed with the OPUS 1.2 codec at various bitrates (i.e. 32, 128 and 256, 512kbps respectively). The sample set for the tests comprised both recorded and synthetic audio clips with a wide range of time-frequency characteristics. To evaluate Localization Accuracy of compressed audio a number of fixed and dynamic (moving vertically and horizontally) source positions were selected for the test samples. Results showed a strong correlation (PCC=0.919; Spearman=0.882 regarding Listening Quality and PCC=0.854; Spearman=0.842 regarding Localization Accuracy) between objective quality scores derived from the B-format Ambisonic audio using AMBIQUAL and subjective scores obtained during listening MUSHRA tests. AMBIQUAL displays very promising quality assessment predictions for spatial audio. Future work will optimise the algorithm to generalise and validate it for any Higher Order Ambisonic formats.
Miroslaw Narbutt, Andrew Allen, Jan Skoglund, Michael Chinen, Andrew Hines
QoMEX5
2018 Perception and prediction of speaker appeal - A single speaker study
Ailbhe Cullen, Andrew Hines, Naomi Harte
Comput. Speech Lang.2
2017 A framework for post-stroke quality of life prediction using structured prediction
abstract
This paper presents a conceptual model that relates Quality of Life to the established Quality of Experience formation process. It uses concepts developed by the Quality of Experience community to propose an adapted framework for developing predictive models for Quality of Life. A mapping of common factors that can be applied to health related quality of life is proposed and practical challenges for modelling and applications are presented and discussed. The process of identifying and categorising factors and features is illustrated using stroke patient treatment as an example use case.
Andrew Hines, John D. Kelleher
QoMEX1
2016 Bitrate classification of twice-encoded audio using objective quality features
abstract
When a user uploads audio files to a music streaming service, these files are subsequently re-encoded to lower bitrates to target different devices, e.g. low bitrate for mobile. To save time and bandwidth uploading files, some users encode their original files using a lossy codec. The metadata for these files cannot always be trusted as users might have encoded their files more than once. Determining the lowest bitrate of the files allows the streaming service to skip the process of encoding the files to bitrates higher than that of the uploaded files, saving on processing and storage space. This paper presents a model that uses quality predictions from ViSQOLAudio, a full reference objective audio quality metric, as features in combination with a multi-class support vector machine classifier. An experiment on twice-encoded files found that low bitrate codecs could be classified using audio quality features. The experiment also provides insights into the implications of multiple transcodes from a quality perspective.
Colm Sloan, Naomi Harte, Damien Kelly, Anil C. Kokaram, Andrew Hines
QoMEX5
2015 Measuring and monitoring speech quality for voice over IP with POLQA, viSQOL and p.563
abstract
There are many types of degradation which can occur in Voice over IP (VoIP) calls. Of interest in this work are degradations which occur independently of the codec, hardware or network in use. Specifically, their effect on the subjective and objec- tive quality of the speech is examined. Since no dataset suit- able for this purpose exists, a new dataset (TCD-VoIP) has been created and has been made publicly available. The dataset con- tains speech clips suffering from a range of common call qual- ity degradations, as well as a set of subjective opinion scores on the clips from 24 listeners. The performances of three ob- jective quality metrics: POLQA, ViSQOL and P.563, have been evaluated using the dataset. The results show that full reference metrics are capable of accurately predicting a variety of com- mon VoIP degradations. They also highlight the outstanding need for a wideband, single-ended, no-reference metric to mon- itor accurately speech quality for degradations common in VoIP scenarios.
Andrew Hines, Eoin Gillen, Naomi Harte
INTERSPEECH1
2014 Perceived Audio Quality for Streaming Stereo Music
abstract
Users of audio-visual streaming services expect an ever increasing quality of experience. Channel bandwidth remains a bottleneck commonly addressed with lossy compression schemes for both the video and audio streams. Anecdotal evidence suggests a strongly perceived link between bit rate and quality. This paper presents three audio quality listening experiments using the ITU MUSHRA methodology to assess a number of audio codecs typically used by streaming services. They were assessed for a range of bit rates using three presentation modes: consumer and studio quality headphones and loudspeakers. Our results indicate that with consumer quality headphones, listeners were not differentiating between codecs with bit rates greater than 48 kb/s (p>=0.228). For studio quality headphones and loudspeakers aac-lc at 128 kb/s and higher was differentiated over other codecs (p<=0.001). The results provide insights into quality of experience that will guide future development of objective audio quality metrics.
Andrew Hines, Eoin Gillen, Damien Kelly, Jan Skoglund, Anil C. Kokaram, Naomi Harte
ACM Multimedia1
2013 Robustness of speech quality metrics to background noise and network degradations: Comparing ViSQOL, PESQ and POLQA
abstract
The Virtual Speech Quality Objective Listener (ViSQOL) is a new objective speech quality model. It is a signal based full reference metric that uses a spectro-temporal measure of similarity between a reference and a test speech signal. ViSQOL aims to predict the overall quality of experience for the end listener whether the cause of speech quality degradation is due to ambient noise, or transmission channel degradations. This paper describes the algorithm and tests the model using two speech corpora: NOIZEUS and E4. The NOIZEUS corpus contains speech under a variety of background noise types, speech enhancement methods, and SNR levels. The E4 corpus contains voice over IP degradations including packet loss, jitter and clock drift. The results are compared with the ITU-T objective models for speech quality: PESQ and POLQA. The behaviour of the metrics are also evaluated under simulated time warp conditions. The results show that for both datasets ViSQOL performed comparably with PESQ. POLQA was shown to have lower correlation with subjective scores than the other metrics for the NOIZEUS database.
Andrew Hines, Jan Skoglund, Anil C. Kokaram, Naomi Harte
ICASSP1
2013 Monitoring the effects of temporal clipping on voIP speech quality
abstract
This paper presents work on a real-time temporal clipping monitoring tool for VoIP. Temporal clipping can occur as a result of voice activity detection (VAD) or echo cancellation where comfort noise in used in place of clipped speech segments. The algorithm presented will form part of a no-reference objective model for quantifying perceived speech quality in VoIP. The overall approach uses a modular design that will help pinpoint the reason for degradations in addition to quantifying their impact on speech quality. The new algorithm was tested for VAD compared over a range of thresholds and varied speech frame sizes. The results are compared to objective Mean Opinion Scores (MOS-LQO) from POLQA. The results show that the proposed algorithm can efficiently predict temporal clipping in speech and correlates well with the full reference quality predictions from POLQA. The model shows good potential for use in a real-time monitoring tool. Index Terms: temporal clipping, VAD, VoIP, POLQA 1.
Andrew Hines, Jan Skoglund, Anil C. Kokaram, Naomi Harte
INTERSPEECH1
2012 Improved Speech Intelligibility with a Chimaera Hearing Aid Algorithm
abstract
It is recognised that current hearing aid fitting algorithms can corrupt fine timing cues in speech.This paper presents a fitting algorithm that aims to improve speech intelligibility, while preserving the temporal fine structure.The algorithm combines the signal envelope amplification from a standard hearing aid fitting algorithm with the fine timing information available to unaided listeners.The proposed "chimaera aid" is evaluated with computer simulated listener tests to measure its speech intelligibility for 3 sample hearing losses.In addition, the experiment demonstrates the potential application of auditory nerve models in the development of new hearing aid algorithm designs using the previously developed Neurogram Similarity Index Measure (NSIM) to predict speech intelligibility.The results predict that the new aid restores envelope without degrading fine timing information.
Andrew Hines, Naomi Harte
INTERSPEECH1
2012 Speech intelligibility prediction using a Neurogram Similarity Index Measure
Andrew Hines, Naomi Harte
Speech Commun.1
2010 Speech intelligibility from image processing
Andrew Hines, Naomi Harte
Speech Commun.1
2009 Error metrics for impaired auditory nerve responses of different phoneme groups
abstract
An auditory nerve model allows faster investigation of new signal processing algorithms for hearing aids.This paper presents a study of the degradation of auditory nerve (AN) responses at a phonetic level for a range of sensorineural hearing losses and flat audiograms.The AN model of Zilany & Bruce was used to compute responses to a diverse set of phoneme rich sentences from the TIMIT database.The characteristics of both the average discharge rate and spike timing of the responses are discussed.The experiments demonstrate that a mean absolute error metric provides a useful measure of average discharge rates but a more complex measure is required to capture spike timing response errors.
Andrew Hines, Naomi Harte
INTERSPEECH1