VLDB 2026 Research / reviewers in the wild / expert
Mansour Moniri
dblp:23/821
· DBLP profile ↗
18ranked-venue papers
0as first author
9since 2021 · last 2027
0000-0002-5564-0692ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Robustness in deepfake speech detection: A survey of failure mechanisms including an experimental case studyabstractIn recent years, the emergence of disruptive deepfake technology (referring here to computer-generated speech/audio, images, video and text) has raised significant concerns surrounding the privacy, security, and credibility of digital content. Although detection systems report high benchmark performance, most are evaluated on in-domain datasets, codecs, and generative models seen during training. In real deployment, however, detectors must contend with out-of-domain instances, that is, a mismatch between the data seen during training and the conditions encountered in deployment (e.g., unseen generative methods, unfamiliar codecs, or signals degraded by transmission channels). Under these conditions, performance often degrades sharply. While robustness has been acknowledged in the literature, it remains underexplored. This work presents a robustness-focused survey and experimental study of deepfake speech detection. Rather than broadly cataloguing existing methods, we organise the literature around three factors that influence robustness to distribution shift: dataset design, feature representations, and model architecture. To complement this synthesis, we conduct a cross-dataset evaluation with Wav2Vec2, HuBERT, WavLM, and Whisper self-supervised speech models, trained on the ASVspoof 2019 LA and evaluated on the ASVspoof 2021 Deepfake dataset. The experimental results show that robustness varies considerably across models and conditions. HuBERT demonstrates comparatively stable performance across codec and dataset shifts, while Wav2Vec2 and WavLM experience substantial degradation, particularly under MP3-based compression and neural vocoder synthesis. A diagnostic analysis further reveals that robustness failures are strongly associated with codec-induced masking of synthesis artefacts and with dataset-dependent synthesis pipelines. By combining a robustness-oriented survey with cross-dataset experimental evidence, this work puts theory to practice, highlighting key factors that limit the generalisation of current detectors and outlining practical directions for improving robustness, including diversified dataset design, feature fusion, and evaluation under realistic cross-domain conditions. Harry Maltby, Julie A. Wall, Cornelius Glackin, Mansour Moniri, Iwa Salami, Nigel Cannings |
Comput. Speech Lang. | 4 |
| 2025 | Robust Deepfake Speech Algorithm Recognition: Classifying Generative Algorithms via Speaker X-Vectors and Deep LearningabstractThe rapid advancement of deepfake voice technologies has resulted in alarming cases of impersonation and deception, highlighting the urgent need for robust tools that can not only distinguish real audio from fake but also recognise the generative algorithms responsible. The ability to not only detect deepfake audio but also recognise the generative methods used is essential for forensic investigations, legal proceedings, and regulatory enforcement. Without robust and explainable detection frameworks, legal professionals and investigators lack the tools needed to effectively monitor, investigate, and prosecute cases involving deepfake misuse. In this work, we take a voice biometrics approach, shifting the focus from identifying who is speaking to identifying which algorithm is speaking. Doing so allows our approach to inherently handle unseen classes while achieving competitive performance for deepfake speech algorithm recognition. Our system leverages a voice-focused ResNet101-based x-vector extraction model and combines diverse audio features, and our experimental novel feature LFCC-HF, enhanced with Linear Discriminant Analysis and cosine similarity clustering. This approach allows for a more transparent and interpretable decision-making process by usinga single voice similarity decision boundary compared to the ensemble-based methods commonly used in the literature. Unlike previous works that rely on an ensemble of models, which convolute the decision-making process, our method achieves comparable results while using a significantly lighter-weight architecture, with our model having 14.84 M parameters compared to 95 M and 317 M parameters for Wav2Vec2 base and large. Furthermore, we demonstrate the benefits of targeted data augmentation, which, combined with feature fusion and our novel feature, improves system robustness and adaptability, increasing our F1 Score from 0.624 to 0.763, a 22.275% increase over our best single feature, and a 40.775% increase over the best ADD 2023 Track 3 baseline. Importantly, the system achieves interpretability through its back-end classification process, where decisions are based on a transparent, learned threshold for voice similarity to known voiceprints. This work offers a foundation for advancing more robust and interpretable solutions in the field of deepfake speech detection. Harry Maltby, Julie A. Wall, Cornelius Glackin, Mansour Moniri, Roman Shrestha, Nigel Cannings, Iwa Salami |
IJCNN | 4 |
| 2024 | A Frequency Bin Analysis of Distinctive Ranges Between Human and Deepfake Generated VoicesabstractDeepfake technology has advanced rapidly in recent years. The widespread availability of deepfake audio technology has raised concerns about its potential misuse for malicious purposes, and a need for more robust countermeasure systems is becoming ever more important. Here we analyse the differences between human and deepfake audio and introduce a novel audio pre-processing approach. Our analysis aims to show the specific locations in the frequency spectrum where these artefacts and distinctions between human and deepfake audio can be found. Our approach emphasises specific frequency ranges that we show are transferable across synthetic speech datasets. In doing so, we explore the use of a bespoke filter bank derived from our analysis of the WaveFake dataset to exploit commonalities across algorithms. Our filter bank was constructed based on a frequency bin analysis of the WaveFake dataset, we apply this filter bank to adjust gain/attenuation to improve the effective signal-to-noise ratio, doing so we reduce the similarities while accentuating differences. We then take a baseline performing model and experiment with improving the performance using these frequency ranges to show where these artefacts lie and if this knowledge is transferable across mel-spectrum algorithms. We show that there exist exploitable commonalities between deepfake voice generation methods that generate audio in the mel-spectrum and that artefacts are left behind in similar frequency regions. Our approach is evaluated on the ASVSpoof 2019 Logical Access dataset of which the test set contains unseen generative methods to test the efficacy of our filter bank approach and transferability. Our experiments show that there is enhanced classification performance to be gained from utilizing these transferable frequency bands where there are more artefacts and distinctions. Our highest-performing model provided a 14.75% improvement in Equal Error Rate against our baseline model. Harry Maltby, Julie A. Wall, Cornelius Glackin, Mansour Moniri, Nigel Cannings, Iwa Salami |
IJCNN | 4 |
| 2023 | Enhancing Automatic Speech Recognition Quality with a Second-Stage Speech Enhancement Generative Adversarial NetworkabstractSpeech enhancement is an essential preprocessing stage for automatic speech recognition in noisy conditions; however, the distortion caused by the denoising process may lead to degradation in automatic speech recognition performance. This paper presents a deep learning-based speech enhancement architecture to overcome this issue by applying a second-stage network that deals with distortion noise. Moreover, a signal-to-noise ratio binary classifier is implemented to activate the speech enhancement network for intrusive noise environments only, which improves the overall performance. The proposed architecture outperforms powerful models in the literature, as it improves a challenging noisy speech test set by 0.8 and 5.9% improvement in the quality and intelligibility scores, respectively. Furthermore, the architecture improves the performance of automatic speech recognition with a 13.8% reduction in the word error rate at 0dB signal-to-noise ratio. Finally, the second-stage network was proven to improve the performance of first-stage speech enhancement models, not previously seen in the training process. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
ICTAI | 3 |
| 2023 | A Deep Learning Speech Enhancement Architecture Optimised for Speech Recognition and Hearing AidsabstractWith the fast progression of the speech enhancement field after the introduction of deep learning techniques, there is a need to consider the adjustments needed to employ these techniques for real-life applications. In this work, we present an optimised deep learning speech enhancement architecture for automatic speech recognition and hearing aids, two key speech enhancement applications. A speech enhancement architecture with a signal-to-noise ratio switch is presented for automatic speech recognition systems, to avoid denoising artifacts that cause performance degradation in the case of clean or high signal-to-noise speech. Moreover, a smart speech enhancement architecture is presented for hearing aids to retain important emergency noise in the audio signal. The presented work achieved 13.9% reduction in the word error rate of an automatic speech recognition system. Additionally, the smart speech enhancement architecture resulted in 0.18 improvement in HAAQI audio quality metric. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
ICTAI | 3 |
| 2023 | Deception detection in conversations using the proximity of linguistic markersabstractDetecting the elements of deception in a conversation takes years of study and experience, and it is a skill set primarily used in law-enforcement agencies. In ever-growing business opportunities, organisations employ teleoperators to provide support and services to their large customer base, which is a potential platform for fraud. With technological advancements, it is desirable to have an automated system that spots the deceptive elements in the conversation, and provides this information to the teleoperators to better support them in their interactions. We propose the Decision Engine to detect deceptive conversation based on the proximity of linguistic markers present, which produces a deception score for a conversation and highlights the potential deceptive elements of the conversation. In collaboration with behavioural experts, we have selected ten linguistic markers that potentially indicate deception. We have built a variety of models to detect the trigger terms for selected linguistic markers without ambiguity, using either regular expressions or the BERT model. The BERT model has been trained on a conversational dataset that we collated and was labelled by our behavioural experts. The proposed Decision Engine employs the BERT model and regular expressions to detect the linguistic markers and compute the proximity features to further estimate the deception score. We evaluated the proposed approach on the Columbia-SRI-Colorado (CSC) dataset and a real-world Financial Services dataset. In addition to accuracy, we have also employed the True Positive Rate metric, with a high enough threshold to avoid any false-positive cases, which we indicate as TPRF0. The Decision Engine achieves 69% accuracy and 46% TPRF0 for the CSC dataset and 72% accuracy and 60% TPRF0 for the Financial Services dataset. In contrast, a baseline model, which uses non-proximity features achieves 67% accuracy and 32% TPRF0 for the CSC dataset and 67% accuracy and 10% TPRF0 for the Financial Services dataset. Furthermore, using the Decision Engine, the impact of the proximity of markers on the deception score has been analysed by our behavioural experts to provide insight into linguistic behaviour in relation to deception. Nikesh Bajaj, Marvin Rajwadi, Tracy Goodluck Constance, Julie A. Wall, Mansour Moniri, Thea Laird, Chris Woodruff, James Laird, Cornelius Glackin, Nigel Cannings |
Knowl. Based Syst. | 5 |
| 2022 | Two-Stage Deep Learning Approach for Speech Enhancement and Reconstruction in The Frequency and Time DomainsabstractDeep learning has recently shown promising improvement in the speech enhancement field, due to its effectiveness in eliminating noise. However, a drawback of the denoising process is the introduction of speech distortion, which negatively affects speech quality and intelligibility. In this work, we propose a deep convolutional denoising autoencoder-based speech enhancement network that is designed to have an encoder deeper than the decoder, to improve performance and decrease complexity. Furthermore, we present a two-stage learning approach, in which denoising is performed in the first frequency domain stage using magnitude spectrum as a training target; while, in the second stage, further denoising and speech reconstruction are performed in the time domain. Results show that our architecture achieves 0.22 improvement in the overall predicted mean opinion score (Covl) over state of the art speech enhancement architectures, using the Valentini dataset benchmark. Moreover, the architecture was trained using a larger dataset and tested using a mismatched test corpus, to achieve 0.7 and 6.35% improvement in Perceptual Evaluation of Speech Quality (PESQ) and Short Time Objective Intelligibility (STOI) scores, respectively, compared to the noisy speech. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
IJCNN | 3 |
| 2022 | Convolutional Recurrent Smart Speech Enhancement Architecture for Hearing Aids
Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
INTERSPEECH | 3 |
| 2021 | Resolving Ambiguity in Hedge Detection by Automatic Generation of Linguistic Rules
Tracy Goodluck Constance, Nikesh Bajaj, Marvin Rajwadi, Harry Maltby, Julie A. Wall, Mansour Moniri, Chris Woodruff, Thea Laird, James Laird, Cornelius Glackin, Nigel Cannings |
ICANN (5) | 6 |
| 2020 | Mapping and Masking Targets Comparison using Different Deep Learning based Speech Enhancement ArchitecturesabstractMapping and Masking targets are both widely used in recent Deep Neural Network (DNN) based supervised speech enhancement. Masking targets are proved to have a positive impact on the intelligibility of the output speech, while mapping targets are found, in other studies, to generate speech with better quality. However, most of the studies are based on comparing the two approaches using the Multilayer Perceptron (MLP) architecture only. With the emergence of new architectures that outperform the MLP, a more generalized comparison is needed between mapping and masking approaches. In this paper, a complete comparison will be conducted between mapping and masking targets using four different DNN based speech enhancement architectures, to work out how the performance of the networks changes with the chosen training target. The results show that there is no perfect training target with respect to all the different speech quality evaluation metrics, and that there is a tradeoff between the denoising process and the intelligibility of the output speech. Furthermore, the generalization ability of the networks was evaluated, and it is concluded that the design of the architecture restricts the choice of the training target, because masking targets result in significant performance degradation for deep convolutional autoencoder architecture. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
IJCNN | 3 |
| 2020 | A Comparative Study of Time and Frequency Domain Approaches to Deep Learning based Speech EnhancementabstractDeep learning has recently made a breakthrough in the speech enhancement process. Some architectures are based on a time domain representation, while others operate in the frequency domain; however, the study and comparison of different networks working in time and frequency is not reported in the literature. In this paper, this comparison between time and frequency domain learning for five Deep Neural Network (DNN) based speech enhancement architectures is presented. The comparison covers the evaluation of the output speech using four objective evaluation metrics: PESQ, STOI, LSD, and SSNR increase. Furthermore, the complexity of the five networks was investigated by comparing the number of parameters and processing time for each architecture. Finally some of the factors that affect learning in time and frequency were discussed. The primary results of this paper show that fully connected based architectures generate speech with low overall perception when learning in the time domain. On the other hand, convolutional based designs give acceptable performance in both frequency and time domains. However, time domain implementations show an inferior generalization ability. Frequency domain based learning was proved to be better than time domain when the complex spectrogram is used in the training process. Additionally, feature extraction is also proved to be very effective in DNN based supervised speech enhancement, whether it is performed at the beginning, or implicitly by bottleneck layer features. Finally, it was concluded that the choice of the working domain is mainly restricted by the type and design of the architecture used. Soha A. Nossier, Julie A. Wall, Mansour Moniri, Cornelius Glackin, Nigel Cannings |
IJCNN | 3 |
| 2018 | Analytical framework for adaptive compressive sensing for target detection within wireless visual sensor networks
Salema F. Fayed, Sherin M. Youssef, Amr El-Helw, Mohammad N. Patwary, Mansour Moniri |
Multim. Tools Appl. | 5 |
| 2016 | Adaptive compressive sensing for target tracking within wireless visual sensor networks-based surveillance applications
Salema F. Fayed, Sherin M. Youssef, Amr El-Helw, Mohammad N. Patwary, Mansour Moniri |
Multim. Tools Appl. | 5 |
| 2012 | Effect of inter-camera angles on the performance of an H.264/AVC based multi-view video codecabstractThis paper investigates the effect of inter-camera angles on the performance of an H.264/AVC based multi-view video codec. To achieve this, the H.264/AVC software has been modified to support multi-view video coding using its multi-frame reference property. Results were generated using a wide baseline convergent multi-view video data set: Breakdancers. To generate a set of three synchronized multi-view videos from the same scene with different inter-camera angles, all possible three camera combinations are generated and classified according to their inter-camera angles. The resulting set of multi-view videos are coded using H.264/AVC based multi-view and simulcast video codecs at different bitrates. Results demonstrate that the multi-view video codec gives superior coding performance up to 1.2dB compared to that of simulcast coding scheme at low inter-camera angles and it deteriorates as the inter camera angles increase. Finally, a range of inter-camera angles for best use of either multi-view or simulcast coding is determined. Akbar Sheikh Akbari, Hany Said, Mansour Moniri |
PCS | 3 |
| 2011 | Evolutionary multi-objective design optimisation of energy harvesting MEMS: The case of a PiezoelectricabstractThe design and optimisation of Energy Harvesting (EH) Micro-Electromechanical-Systems (MEMS) is of particular interest in this research. The application of such devices is becoming an attractive alternative to the traditional use of batteries in wireless and body sensor networks. An evolutionary Multi-Objective Design Optimisation (DO) Framework is developed to experiment with one class of EH-MEMS, namely, Piezoelectric, using a reconstructed analytical model of the system. The application of such a Framework in this application domain is unprecedented and has already shown very promising results and in some cases it outperformed the human engineer. A thorough analysis of the results has been undertaken, which reveals interesting conclusions about the behaviour and physics of such devices. Besides, the main features of the Framework are explored enabling the enhancement of the MEMS-DO. Elhadj Benkhelifa, Mansour Moniri, Ashutosh Tiwari 0001, Alfonso G. De Rueda |
IEEE Congress on Evolutionary Computation | 2 |
| 2005 | Video combiner for multi-channel video surveillance based on finite state methodsabstractIn this paper a new technique for a multi-channel video combiner based on finite state methods is presented. The proposed technique combines the input video channels in the compressed domain using generalised finite transducers. The proposed technique operates entirely in the compressed domain thus eliminating the processing overhead, delay and picture quality degradation associated with traditional video combiner methods adapting the decompress/combine/recompress approach. Mohamed Abdel-Maguid, Mansour Moniri |
AVSS | 2 |
| 2005 | Classification of smart video surveillance systems for commercial applicationsabstractVideo surveillance has a large market as the number of installed cameras around us can show. There are immediate commercial needs for smart video surveillance systems that can make use of the existing camera network (e.g. CCTV) for more intelligent security systems and to contribute in more applications (beside or) rather than security applications. This work introduces a new classification for smart video surveillance systems depending on their commercial applications. This paper highlights different links between the research and the commercial applications. The work reported here has both research and commercial motivations. Our goals are first to define a generic model of smart video surveillance systems that can meet requirements of strong commercial applications. Our second goal is to categorize different smart video surveillance applications and to relate capabilities of computer vision algorithms to the requirement of commercial application. Mohamed H. Sedky, Mansour Moniri, Claude C. Chibelushi |
AVSS | 2 |
| 2002 | Fragile and semi-fragile image authentication based on image self-similarityabstractWe propose a new approach for fragile and semi-fragile image authentication based on image self-similarity. The authentication signature is extracted from image representation using generalized finite automata which represents the image in terms of its perceptual self-similarity. This self-similarity is more likely to be damaged due to illegitimate manipulations rather than compression algorithms such as JPEG. The proposed technique is also capable of changing the degree of fragility by setting the value of the image self-similarity parameter /spl alpha/. Furthermore, our approach features an extremely compact signature length whilst maintaining a rather simple implementation. Experimental results are presented for well known images to support these claims. Sherif Nour El-Din, Mansour Moniri |
ICIP (2) | 2 |