VLDB 2026 Research / reviewers in the wild / expert
Julie A. Wall
dblp:06/2667
· DBLP profile ↗
22ranked-venue papers
5as first author
11since 2021 · last 2027
0000-0001-6714-4867ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Robustness in deepfake speech detection: A survey of failure mechanisms including an experimental case studyabstractIn recent years, the emergence of disruptive deepfake technology (referring here to computer-generated speech/audio, images, video and text) has raised significant concerns surrounding the privacy, security, and credibility of digital content. Although detection systems report high benchmark performance, most are evaluated on in-domain datasets, codecs, and generative models seen during training. In real deployment, however, detectors must contend with out-of-domain instances, that is, a mismatch between the data seen during training and the conditions encountered in deployment (e.g., unseen generative methods, unfamiliar codecs, or signals degraded by transmission channels). Under these conditions, performance often degrades sharply. While robustness has been acknowledged in the literature, it remains underexplored. This work presents a robustness-focused survey and experimental study of deepfake speech detection. Rather than broadly cataloguing existing methods, we organise the literature around three factors that influence robustness to distribution shift: dataset design, feature representations, and model architecture. To complement this synthesis, we conduct a cross-dataset evaluation with Wav2Vec2, HuBERT, WavLM, and Whisper self-supervised speech models, trained on the ASVspoof 2019 LA and evaluated on the ASVspoof 2021 Deepfake dataset. The experimental results show that robustness varies considerably across models and conditions. HuBERT demonstrates comparatively stable performance across codec and dataset shifts, while Wav2Vec2 and WavLM experience substantial degradation, particularly under MP3-based compression and neural vocoder synthesis. A diagnostic analysis further reveals that robustness failures are strongly associated with codec-induced masking of synthesis artefacts and with dataset-dependent synthesis pipelines. By combining a robustness-oriented survey with cross-dataset experimental evidence, this work puts theory to practice, highlighting key factors that limit the generalisation of current detectors and outlining practical directions for improving robustness, including diversified dataset design, feature fusion, and evaluation under realistic cross-domain conditions. Harry Maltby, Julie A. Wall, Cornelius Glackin, Mansour Moniri, Iwa Salami, Nigel Cannings |
Comput. Speech Lang. | 2 |
| 2025 | Robust Deepfake Speech Algorithm Recognition: Classifying Generative Algorithms via Speaker X-Vectors and Deep LearningabstractThe rapid advancement of deepfake voice technologies has resulted in alarming cases of impersonation and deception, highlighting the urgent need for robust tools that can not only distinguish real audio from fake but also recognise the generative algorithms responsible. The ability to not only detect deepfake audio but also recognise the generative methods used is essential for forensic investigations, legal proceedings, and regulatory enforcement. Without robust and explainable detection frameworks, legal professionals and investigators lack the tools needed to effectively monitor, investigate, and prosecute cases involving deepfake misuse. In this work, we take a voice biometrics approach, shifting the focus from identifying who is speaking to identifying which algorithm is speaking. Doing so allows our approach to inherently handle unseen classes while achieving competitive performance for deepfake speech algorithm recognition. Our system leverages a voice-focused ResNet101-based x-vector extraction model and combines diverse audio features, and our experimental novel feature LFCC-HF, enhanced with Linear Discriminant Analysis and cosine similarity clustering. This approach allows for a more transparent and interpretable decision-making process by usinga single voice similarity decision boundary compared to the ensemble-based methods commonly used in the literature. Unlike previous works that rely on an ensemble of models, which convolute the decision-making process, our method achieves comparable results while using a significantly lighter-weight architecture, with our model having 14.84 M parameters compared to 95 M and 317 M parameters for Wav2Vec2 base and large. Furthermore, we demonstrate the benefits of targeted data augmentation, which, combined with feature fusion and our novel feature, improves system robustness and adaptability, increasing our F1 Score from 0.624 to 0.763, a 22.275% increase over our best single feature, and a 40.775% increase over the best ADD 2023 Track 3 baseline. Importantly, the system achieves interpretability through its back-end classification process, where decisions are based on a transparent, learned threshold for voice similarity to known voiceprints. This work offers a foundation for advancing more robust and interpretable solutions in the field of deepfake speech detection. Harry Maltby, Julie A. Wall, Cornelius Glackin, Mansour Moniri, Roman Shrestha, Nigel Cannings, Iwa Salami |
IJCNN | 2 |
| 2024 | A Frequency Bin Analysis of Distinctive Ranges Between Human and Deepfake Generated VoicesabstractDeepfake technology has advanced rapidly in recent years. The widespread availability of deepfake audio technology has raised concerns about its potential misuse for malicious purposes, and a need for more robust countermeasure systems is becoming ever more important. Here we analyse the differences between human and deepfake audio and introduce a novel audio pre-processing approach. Our analysis aims to show the specific locations in the frequency spectrum where these artefacts and distinctions between human and deepfake audio can be found. Our approach emphasises specific frequency ranges that we show are transferable across synthetic speech datasets. In doing so, we explore the use of a bespoke filter bank derived from our analysis of the WaveFake dataset to exploit commonalities across algorithms. Our filter bank was constructed based on a frequency bin analysis of the WaveFake dataset, we apply this filter bank to adjust gain/attenuation to improve the effective signal-to-noise ratio, doing so we reduce the similarities while accentuating differences. We then take a baseline performing model and experiment with improving the performance using these frequency ranges to show where these artefacts lie and if this knowledge is transferable across mel-spectrum algorithms. We show that there exist exploitable commonalities between deepfake voice generation methods that generate audio in the mel-spectrum and that artefacts are left behind in similar frequency regions. Our approach is evaluated on the ASVSpoof 2019 Logical Access dataset of which the test set contains unseen generative methods to test the efficacy of our filter bank approach and transferability. Our experiments show that there is enhanced classification performance to be gained from utilizing these transferable frequency bands where there are more artefacts and distinctions. Our highest-performing model provided a 14.75% improvement in Equal Error Rate against our baseline model. Harry Maltby, Julie A. Wall, Cornelius Glackin, Mansour Moniri, Nigel Cannings, Iwa Salami |
IJCNN | 2 |
| 2024 | A reinforcement learning recommender system using bi-clustering and Markov Decision ProcessabstractCollaborative filtering (CF) recommender systems are static in nature and does not adapt well with changing user preferences. User preferences may change after interaction with a system or after buying a product. Conventional CF clustering algorithms only identifies the distribution of patterns and hidden correlations globally. However, the impossibility of discovering local patterns by these algorithms, headed to the popularization of bi-clustering algorithms. Bi-clustering algorithms can analyze all dataset dimensions simultaneously and consequently, discover local patterns that deliver a better understanding of the underlying hidden correlations. In this paper, we modelled the recommendation problem as a sequential decision-making problem using Markov Decision Processes (MDP). To perform state representation for MDP, we first converted user-item votings matrix to a binary matrix. Then we performed bi-clustering on this binary matrix to determine a subset of similar rows and columns. A bi-cluster merging algorithm is designed to merge similar and overlapping bi-clusters. These bi-clusters are then mapped to a squared grid (SG). RL is applied on this SG to determine best policy to give recommendation to users. Start state is determined using Improved Triangle Similarity (ITR similarity measure. Reward function is computed as grid state overlapping in terms of users and items in current and prospective next state. A thorough comparative analysis was conducted, encompassing a diverse array of methodologies, including RL-based, pure Collaborative Filtering (CF), and clustering methods. The results demonstrate that our proposed method outperforms its competitors in terms of precision, recall, and optimal policy learning. Arta Iftikhar, Mustansar Ali Ghazanfar, Mubbashir Ayub, Saad Ali Alahmari, Nadeem Qazi, Julie A. Wall |
Expert Syst. Appl. | 6 |
| 2023 | Enhancing Automatic Speech Recognition Quality with a Second-Stage Speech Enhancement Generative Adversarial NetworkabstractSpeech enhancement is an essential preprocessing stage for automatic speech recognition in noisy conditions; however, the distortion caused by the denoising process may lead to degradation in automatic speech recognition performance. This paper presents a deep learning-based speech enhancement architecture to overcome this issue by applying a second-stage network that deals with distortion noise. Moreover, a signal-to-noise ratio binary classifier is implemented to activate the speech enhancement network for intrusive noise environments only, which improves the overall performance. The proposed architecture outperforms powerful models in the literature, as it improves a challenging noisy speech test set by 0.8 and 5.9% improvement in the quality and intelligibility scores, respectively. Furthermore, the architecture improves the performance of automatic speech recognition with a 13.8% reduction in the word error rate at 0dB signal-to-noise ratio. Finally, the second-stage network was proven to improve the performance of first-stage speech enhancement models, not previously seen in the training process. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
ICTAI | 2 |
| 2023 | A Deep Learning Speech Enhancement Architecture Optimised for Speech Recognition and Hearing AidsabstractWith the fast progression of the speech enhancement field after the introduction of deep learning techniques, there is a need to consider the adjustments needed to employ these techniques for real-life applications. In this work, we present an optimised deep learning speech enhancement architecture for automatic speech recognition and hearing aids, two key speech enhancement applications. A speech enhancement architecture with a signal-to-noise ratio switch is presented for automatic speech recognition systems, to avoid denoising artifacts that cause performance degradation in the case of clean or high signal-to-noise speech. Moreover, a smart speech enhancement architecture is presented for hearing aids to retain important emergency noise in the audio signal. The presented work achieved 13.9% reduction in the word error rate of an automatic speech recognition system. Additionally, the smart speech enhancement architecture resulted in 0.18 improvement in HAAQI audio quality metric. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
ICTAI | 2 |
| 2023 | Deception detection in conversations using the proximity of linguistic markersabstractDetecting the elements of deception in a conversation takes years of study and experience, and it is a skill set primarily used in law-enforcement agencies. In ever-growing business opportunities, organisations employ teleoperators to provide support and services to their large customer base, which is a potential platform for fraud. With technological advancements, it is desirable to have an automated system that spots the deceptive elements in the conversation, and provides this information to the teleoperators to better support them in their interactions. We propose the Decision Engine to detect deceptive conversation based on the proximity of linguistic markers present, which produces a deception score for a conversation and highlights the potential deceptive elements of the conversation. In collaboration with behavioural experts, we have selected ten linguistic markers that potentially indicate deception. We have built a variety of models to detect the trigger terms for selected linguistic markers without ambiguity, using either regular expressions or the BERT model. The BERT model has been trained on a conversational dataset that we collated and was labelled by our behavioural experts. The proposed Decision Engine employs the BERT model and regular expressions to detect the linguistic markers and compute the proximity features to further estimate the deception score. We evaluated the proposed approach on the Columbia-SRI-Colorado (CSC) dataset and a real-world Financial Services dataset. In addition to accuracy, we have also employed the True Positive Rate metric, with a high enough threshold to avoid any false-positive cases, which we indicate as TPRF0. The Decision Engine achieves 69% accuracy and 46% TPRF0 for the CSC dataset and 72% accuracy and 60% TPRF0 for the Financial Services dataset. In contrast, a baseline model, which uses non-proximity features achieves 67% accuracy and 32% TPRF0 for the CSC dataset and 67% accuracy and 10% TPRF0 for the Financial Services dataset. Furthermore, using the Decision Engine, the impact of the proximity of markers on the deception score has been analysed by our behavioural experts to provide insight into linguistic behaviour in relation to deception. Nikesh Bajaj, Marvin Rajwadi, Tracy Goodluck Constance, Julie A. Wall, Mansour Moniri, Thea Laird, Chris Woodruff, James Laird, Cornelius Glackin, Nigel Cannings |
Knowl. Based Syst. | 4 |
| 2022 | Two-Stage Deep Learning Approach for Speech Enhancement and Reconstruction in The Frequency and Time DomainsabstractDeep learning has recently shown promising improvement in the speech enhancement field, due to its effectiveness in eliminating noise. However, a drawback of the denoising process is the introduction of speech distortion, which negatively affects speech quality and intelligibility. In this work, we propose a deep convolutional denoising autoencoder-based speech enhancement network that is designed to have an encoder deeper than the decoder, to improve performance and decrease complexity. Furthermore, we present a two-stage learning approach, in which denoising is performed in the first frequency domain stage using magnitude spectrum as a training target; while, in the second stage, further denoising and speech reconstruction are performed in the time domain. Results show that our architecture achieves 0.22 improvement in the overall predicted mean opinion score (Covl) over state of the art speech enhancement architectures, using the Valentini dataset benchmark. Moreover, the architecture was trained using a larger dataset and tested using a mismatched test corpus, to achieve 0.7 and 6.35% improvement in Perceptual Evaluation of Speech Quality (PESQ) and Short Time Objective Intelligibility (STOI) scores, respectively, compared to the noisy speech. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
IJCNN | 2 |
| 2022 | Convolutional Recurrent Smart Speech Enhancement Architecture for Hearing Aids
Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
INTERSPEECH | 2 |
| 2021 | Resolving Ambiguity in Hedge Detection by Automatic Generation of Linguistic Rules
Tracy Goodluck Constance, Nikesh Bajaj, Marvin Rajwadi, Harry Maltby, Julie A. Wall, Mansour Moniri, Chris Woodruff, Thea Laird, James Laird, Cornelius Glackin, Nigel Cannings |
ICANN (5) | 5 |
| 2021 | Bird Audio Diarization with Faster R-CNN
Roman Shrestha, Cornelius Glackin, Julie A. Wall, Nigel Cannings |
ICANN (1) | 3 |
| 2020 | Mapping and Masking Targets Comparison using Different Deep Learning based Speech Enhancement ArchitecturesabstractMapping and Masking targets are both widely used in recent Deep Neural Network (DNN) based supervised speech enhancement. Masking targets are proved to have a positive impact on the intelligibility of the output speech, while mapping targets are found, in other studies, to generate speech with better quality. However, most of the studies are based on comparing the two approaches using the Multilayer Perceptron (MLP) architecture only. With the emergence of new architectures that outperform the MLP, a more generalized comparison is needed between mapping and masking approaches. In this paper, a complete comparison will be conducted between mapping and masking targets using four different DNN based speech enhancement architectures, to work out how the performance of the networks changes with the chosen training target. The results show that there is no perfect training target with respect to all the different speech quality evaluation metrics, and that there is a tradeoff between the denoising process and the intelligibility of the output speech. Furthermore, the generalization ability of the networks was evaluated, and it is concluded that the design of the architecture restricts the choice of the training target, because masking targets result in significant performance degradation for deep convolutional autoencoder architecture. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
IJCNN | 2 |
| 2020 | A Comparative Study of Time and Frequency Domain Approaches to Deep Learning based Speech EnhancementabstractDeep learning has recently made a breakthrough in the speech enhancement process. Some architectures are based on a time domain representation, while others operate in the frequency domain; however, the study and comparison of different networks working in time and frequency is not reported in the literature. In this paper, this comparison between time and frequency domain learning for five Deep Neural Network (DNN) based speech enhancement architectures is presented. The comparison covers the evaluation of the output speech using four objective evaluation metrics: PESQ, STOI, LSD, and SSNR increase. Furthermore, the complexity of the five networks was investigated by comparing the number of parameters and processing time for each architecture. Finally some of the factors that affect learning in time and frequency were discussed. The primary results of this paper show that fully connected based architectures generate speech with low overall perception when learning in the time domain. On the other hand, convolutional based designs give acceptable performance in both frequency and time domains. However, time domain implementations show an inferior generalization ability. Frequency domain based learning was proved to be better than time domain when the complex spectrogram is used in the training process. Additionally, feature extraction is also proved to be very effective in DNN based supervised speech enhancement, whether it is performed at the beginning, or implicitly by bottleneck layer features. Finally, it was concluded that the choice of the working domain is mainly restricted by the type and design of the architecture used. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
IJCNN | 2 |
| 2019 | Explaining Sentiment ClassificationabstractInternational audience Marvin Rajwadi, Cornelius Glackin, Julie A. Wall, Gérard Chollet, Nigel Cannings |
INTERSPEECH | 3 |
| 2018 | Convolutional Neural Networks for Phoneme RecognitionabstractThis paper presents a novel application of convolutional neural networks to phoneme recognition. Thephonetic transcription of the TIMIT speech corpus is used to label spectrogram segments for training theconvolutional neural network. A window of a fixed size slides over the spectrogram of the TIMIT utterancesand the resulting spectrogram patches are assigned to the appropriate phone class by parsing TIMIT’s phonetranscription. The convolutional neural network is the standard GoogLeNet implementation trained withstochastic gradient descent with mini batches. After training, phonetic rescoring is performed in the usual wayto map the TIMIT phone set to the smaller standard set. Benchmark results are presented for comparison toother state-of-the-art approaches. Finally, conclusions and future directions with regard to extending theapproach are discussed. Cornelius Glackin, Julie A. Wall, Gérard Chollet, Nazim Dugan, Nigel Cannings |
ICPRAM | 2 |
| 2017 | Privacy preserving encrypted phonetic search of speech dataabstractThis paper presents a strategy for enabling speech recognition to be performed in the cloud whilst preserving the privacy of users. The approach advocates a demarcation of responsibilities between the client and server-side components for performing the speech recognition task. On the client-side resides the acoustic model, which symbolically encodes the audio and encrypts the data before uploading to the server. The server-side then employs searchable encryption to enable the phonetic search of the speech content. Some preliminary results for speech encoding and searchable encryption are presented. Cornelius Glackin, Gérard Chollet, Nazim Dugan, Nigel Cannings, Julie A. Wall, Shahzaib Tahir, Indranil Ghosh Ray, Muttukrishnan Rajarajan |
ICASSP | 5 |
| 2016 | Recurrent lateral inhibitory spiking networks for speech enhancementabstractAutomatic speech recognition accuracy is affected adversely by the presence of noise. In this paper we present a novel noise removal and speech enhancement technique based on spiking neural network processing of speech data. The spiking network has a recurrent lateral topology that is biologically inspired, specifically by the inhibitory cells of the cochlear nucleus. The network can be configured for different acoustic environments and it will be demonstrated how the connectivity results in enhancement of temporal correlation between similar frequency bands and removal of uncorrelated noise sources. Demonstration of the speech enhancement capability will be provided with data taken from the TIMIT database with different levels of additive Gaussian white noise. Future directions for further development of this novel approach to noise removal and signal processing will also be discussed. Julie A. Wall, Cornelius Glackin, Nigel Cannings, Gérard Chollet, Nazim Dugan |
IJCNN | 1 |
| 2014 | REVERIE: Natural human interaction in virtual immersive environmentsabstractREVERIE (REal and Virtual Engagement in Realistic Immersive Environments [1]) targets novel research to address the demanding challenges involved with developing state-of-the-art technologies for online human interaction. The REVERIE framework enables users to meet, socialise and share experiences online by integrating cutting-edge technologies for 3D data acquisition and processing, networking, autonomy and real-time rendering. In this paper, we describe the innovative research that is showcased through the REVERIE integrated framework through richly defined use-cases which demonstrate the validity and potential for natural interaction in a virtual immersive and safe environment. Previews of the REVERIE demo and its key research components can be viewed at www.youtube.com/user/REVERIEFP7. Julie A. Wall, Ebroul Izquierdo, Lemonia Argyriou, David S. Monaghan, Noel E. O'Connor, Steven Poulakos, Aljoscha Smolic, Rufael Mekuria |
ICIP | 1 |
| 2014 | Tools for User Interaction in Immersive Environments
Noel E. O'Connor, Dimitrios S. Alexiadis, Konstantinos C. Apostolakis, Petros Daras, Ebroul Izquierdo, Yingbo Li, David S. Monaghan, Fiona M. Rivera, C. Stevens, Sigurd Van Broeck, Julie A. Wall, Haolin Wei |
MMM (2) | 11 |
| 2012 | Spiking Neural Network Model of Sound Localization Using the Interaural Intensity DifferenceabstractIn this paper, a spiking neural network (SNN) architecture to simulate the sound localization ability of the mammalian auditory pathways using the interaural intensity difference cue is presented. The lateral superior olive was the inspiration for the architecture, which required the integration of an auditory periphery (cochlea) model and a model of the medial nucleus of the trapezoid body. The SNN uses leaky integrate-and-fire excitatory and inhibitory spiking neurons, facilitating synapses and receptive fields. Experimentally derived head-related transfer function (HRTF) acoustical data from adult domestic cats were employed to train and validate the localization ability of the architecture, training used the supervised learning algorithm called the remote supervision method to determine the azimuthal angles. The experimental results demonstrate that the architecture performs best when it is localizing high-frequency sound data in agreement with the biology, and also shows a high degree of robustness when the HRTF acoustical data is corrupted by noise. Julie A. Wall, Liam McDaid, Liam P. Maguire, T. Martin McGinnity |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2011 | A comparison of sound localisation techniques using cross-correlation and spiking neural networks for mobile roboticsabstractThis paper outlines the development of a cross-correlation algorithm and a spiking neural network (SNN) for sound localisation based on real sound recorded in a noisy and dynamic environment by a mobile robot. The SNN architecture aims to simulate the sound localisation ability of the mammalian auditory pathways by exploiting the binaural cue of interaural time difference (ITD). The medial superior olive was the inspiration for the SNN architecture which required the integration of an encoding layer which produced biologically realistic spike trains, a model of the bushy cells found in the cochlear nucleus and a supervised learning algorithm. The experimental results demonstrate that biologically inspired sound localisation achieved using a SNN can compare favourably to the more classical technique of cross-correlation. Julie A. Wall, T. Martin McGinnity, Liam P. Maguire |
IJCNN | 1 |
| 2008 | Spiking neuron models of the medial and lateral superior olive for sound localisationabstractSound localisation is defined as the ability to identify the position of a sound source. The brain employs two cues to achieve this functionality for the horizontal plane, interaural time difference (ITD) by means of neurons in the medial superior olive (MSO) and interaural intensity difference (IID) by neurons of the lateral superior olive (LSO), both located in the superior olivary complex of the auditory pathway. This paper presents spiking neuron architectures of the MSO and LSO. An implementation of the Jeffress model using spiking neurons is presented as a representation of the MSO, while a spiking neuron architecture showing how neurons of the medial nucleus of the trapezoid body interact with LSO neurons to determine the azimuthal angle is discussed. Experimental results to support this work are presented. Julie A. Wall, Liam McDaid, Liam P. Maguire, T. Martin McGinnity |
IJCNN | 1 |