Haytham M. Fayek

dblp:144/1507 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-1840-7605ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Fine-Grained Traceability for Transparent ML Pipelines
Mujie Liu, Haytham M. Fayek
WWW3
2026 When to Invoke: Refining LLM Fairness with Toxicity Assessment
Jing Ren 0001, Bowen Li 0012, Ziqi Xu 0001, Renqiang Luo, Shuo Yu 0001, Xin Ye 0004, Haytham M. Fayek, Xiaodong Li 0001, Feng Xia 0001
WWW7
2026 When to Trust: A Causality-Aware Calibration Framework for Accurate Knowledge Graph Retrieval-Augmented Generation
abstract
Knowledge Graph Retrieval-Augmented Generation (KG-RAG) extends the RAG paradigm by incorporating structured knowledge from knowledge graphs, enabling Large Language Models (LLMs) to perform more precise and explainable reasoning. While KG-RAG improves factual accuracy in complex tasks, existing KG-RAG models are often severely overconfident, producing high-confidence predictions even when retrieved sub-graphs are incomplete or unreliable, which raises concerns for deployment in high-stakes domains. To address this issue, we propose Ca2KG, a Causality-aware Calibration framework for KG-RAG. Ca2KG integrates counterfactual prompting, which exposes retrieval-dependent uncertainties in knowledge quality and reasoning reliability, with a panel-based re-scoring mechanism that stabilises predictions across interventions. Extensive experiments on two complex QA datasets demonstrate that Ca2KG consistently improves calibration while maintaining or even enhancing predictive accuracy. The source code can be found at~ https://aisuko.github.io/ca2kg/.
Jing Ren 0001, Bowen Li 0012, Ziqi Xu 0001, Xikun Zhang 0002, Haytham M. Fayek, Xiaodong Li 0001
WWW5
2026 Zero-Shot Neural Network Evaluation With Sample-Wise Activation Patterns
abstract
Zero-shot proxies, also known as training-free metrics, are widely adopted to reduce the computational overhead in neural network evaluation for scenarios such as Neural Architecture Search (NAS), as they do not require any training. Existing zero-shot metrics have several limitations, including weak correlation with the true performance and poor generalisation across different networks or downstream tasks. For example, most of these metrics apply only to either convolutional neural networks (CNNs) or Transformers, but not both. To address these limitations, we propose Sample-Wise Activation Patterns (SWAP), and its derivative, SWAP-Score, a novel and highly effective zero-shot metric. SWAP-Score is broadly applicable across both architecture families and task domains, demonstrating strong predictive performance in the majority of tasks. This metric measures the expressivity of neural networks over a mini-batch of samples, showing a high correlation with the neural networks' ground-truth performance. For both CNNs and Transformers, the SWAP-Score outperforms existing zero-shot metrics across computer vision and natural language processing tasks. For instance, Spearman's correlation coefficient between the SWAP-Score and CIFAR-10 validation accuracy for DARTS CNNs is 0.93, and 0.71 for FlexiBERT Transformers on GLUE tasks. Moreover, SWAP-Score is label-independent, hence can be applied at the pre-training stage of language models to estimate their performance for downstream tasks. When applied to NAS, SWAP-empowered NAS, SWAP-NAS can achieve competitive performance using only approximately 6 and 9 minutes of GPU time, on CIFAR-10 and ImageNet respectively.
Yameng Peng, Andy Song, Haytham M. Fayek, Victor Ciesielski, Xiaojun Chang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
abstract
Model compression through post-training pruning offers a way to reduce model size and computational requirements without significantly impacting model performance.However, the effect of pruning on the fairness of LLMgenerated summaries remains unexplored, particularly for opinion summarisation where biased outputs could influence public views.In this paper, we present a comprehensive empirical analysis of opinion summarisation, examining three state-of-the-art pruning methods and various calibration sets across three opensource LLMs using four fairness metrics.Our systematic analysis reveals that pruning methods have a greater impact on fairness than calibration sets.Building on these insights, we propose High Gradient Low Activation (HGLA) pruning, which identifies and removes parameters that are redundant for input processing but influential in output generation.Our experiments demonstrate that HGLA can better maintain or even improve fairness compared to existing methods, showing promise across models and tasks where traditional methods have limitations.Our human evaluation shows HGLA-generated outputs are fairer than existing state-of-the-art pruning methods.
Nannan Huang, Haytham M. Fayek, Xiuzhen Zhang 0001
EMNLP2
2025 Narrow-Band RFI Mitigation in Synthetic Aperture Radars Using Variable Space-Frequency Filter
abstract
Radio frequency interference (RFI) in synthetic aperture radar (SAR) is a daunting challenge, affecting both sensing reliability and image quality. To ensure that SAR remains a powerful tool for Earth observation, this letter presents a 2-D variable attenuation space (azimuth)-frequency filtration (VASFF) method. This framework leverages the time-frequency characteristics of Level-0 SAR data, the RFI power profile, estimated RFI signal parameters, and the SAR antenna pattern to design a novel variable filter. Signal power localization estimates the interference source’s relative position, facilitating filter application. Simulated results, obtained using our open-source emulator, SEMUS, to generate both clean and interference-contaminated raw SAR data, demonstrate that the proposed filter achieves a 2 dB improvement over traditional notch filtering. The framework is further tested on real-life interference events on TerraSAR-X revealing previously obscured image details, validating the framework’s effectiveness.
Nermine Hendy, Akram Al-Hourani, Thomas Kraus, Maximilian Schandri, Markus Bachmann, Haytham M. Fayek
IEEE Geosci. Remote. Sens. Lett.6
2024 SWAP-NAS: Sample-Wise Activation Patterns for Ultra-fast NAS
abstract
Training-free metrics (a.k.a. zero-cost proxies) are widely used to avoid resource-intensive neural network training, especially in Neural Architecture Search (NAS). Recent studies show that existing training-free metrics have several limitations, such as limited correlation and poor generalisation across different search spaces and tasks. Hence, we propose Sample-Wise Activation Patterns and its derivative, SWAP-Score, a novel high-performance training-free metric. It measures the expressivity of networks over a batch of input samples. The SWAP-Score is strongly correlated with ground-truth performance across various search spaces and tasks, outperforming 15 existing training-free metrics on NAS-Bench-101/201/301 and TransNAS-Bench-101. The SWAP-Score can be further enhanced by regularisation, which leads to even higher correlations in cell-based search space and enables model size control during the search. For example, Spearman’s rank correlation coefficient between regularised SWAP-Score and CIFAR-100 validation accuracies on NAS-Bench-201 networks is 0.90, significantly higher than 0.80 from the second-best metric, NWOT. When integrated with an evolutionary algorithm for NAS, our SWAP-NAS achieves competitive performance on CIFAR-10 and ImageNet in approximately 6 minutes and 9 minutes of GPU time respectively.
Yameng Peng, Andy Song, Haytham M. Fayek, Victor Ciesielski, Xiaojun Chang
ICLR3
2024 A Case for Personalized Non-Player Character Companion Design
abstract
Personalized video games have the potential to provide unique and meaningful experiences for the player. User data taken from biosensors, questionnaires, or in-game performance data can infer a player’s psychological state, to which relevant game features can be adapted to enhance the player experience. This survey discusses the data types, game elements, and methods that have been used thus far to create adaptive experiences in games. The survey specifically focuses on personalized nonplayer character (NPC) companions through adaptation. Studies using performance data, affect and cognition, and self-reports to adapt companions and other NPC types are reviewed for their success in providing an enhanced experience. We then provide a motivation for a personalized companion before detailing a framework for an adaptive system that modifies companion characteristics based on the player’s state. This framework takes a human-centered approach to personalized companion design; it proposes the game elements appropriate for adaptation, the data types that suit the adaptation of the companion type, the techniques that would enable successful adaptation, and methods for companion evaluation. The framework firstly quantifies the companion’s behaviour according to recent companion design architecture. Following this, the companion’s characteristics are updated according to the player state, previous player states, evaluations of changes, and the player model. The last phase highlights how to evaluate the companion during the design stage to ensure that it is reliable in assisting the player. We suggest this as a starting point for game designers when considering how to approach a companion that aims to enhance and sustain player experience.
Emma Pretty, Haytham M. Fayek, Fabio Zambetta
Int. J. Hum. Comput. Interact.2
2024 Multimodal Measurement of Cognitive Load in a Video Game Context: A Comparative Study Between Subjective and Objective Metrics
abstract
Understanding the interaction between cognitive load and features of video game play is important for the accurate measurement and application of this psychological construct in both research and industry scenarios to enhance the player experience. A challenge within this domain is the use of measurements developed in different areas in video games without first validating or testing the reliability of these tools. We present a holistic evaluation of methods used to measure cognitive load during naturalistic gameplay of a commercially available sandbox game. This study measures electroencephalography, electromyography, heart rate, heart rate variability, electrodermal activity, and eye blink rate from players during a building task in Minecraft whilst being assisted by a virtual companion. Various models reveal that a number of physiological measures can be used as a proxy of subjective cognitive load measurement, providing further insight into the player experience during gameplay. Our analysis also reveals some discriminant validity of the subjective measurement used. These results can help inform the choice of sensor in evaluating video game features, or when needing to accurately measure cognitive load for an adaptive situation or high-risk scenario.
Emma Pretty, Renan Luigi Martins Guarese, Chloe A. Dziego, Haytham M. Fayek, Fabio Zambetta
IEEE Trans. Games4
2023 Fast Evolutionary Neural Architecture Search by Contrastive Predictor with Linear Regions
abstract
Evolutionary neural architecture search (ENAS) has emerged as a promising approach to finding high-performance neural architectures. However, widespread application has been limited by the expensive computational costs due to the nature of evolutionary algorithms. In this study, we aim to significantly reduce the computational costs of ENAS by involving a training-free performance metric. Specifically, the network performance can be estimated by the training-free metric with only a single forward pass. However, training-free metrics have their own challenges, in particular, an insufficient correlation with ground-truth performance. We adopt a Graph Convolutional Network (GCN) based contrastive predictor which can leverage the low cost of the training-free performance metric yet improve the correlation between the estimated performance and the true performance of the candidate architectures. Combining a training-free metric - the number of linear regions with the GCN-based contrastive predictor and an active learning scheme, we propose Fast-ENAS which can achieve superior search efficiency and performance on the benchmark NAS-Bench-201 and DARTS search spaces. Furthermore, with a single GPU searching on the DARTS space, Fast-ENAS requires only 0.02 (29 minutes) and 0.026 (37 minutes) GPU days to achieve test error rates of 2.50% and 24.30% on CIFAR-10 and ImageNet respectively.
Yameng Peng, Andy Song, Victor Ciesielski, Haytham M. Fayek, Xiaojun Chang
GECCO4
2023 Evoking empathy with visually impaired people through an augmented reality embodiment experience
abstract
To promote empathy with people that have disabilities, we propose a multi-sensory interactive experience that allows sighted users to embody having a visual impairment whilst using assistive technologies. The experiment involves blindfolded sighted participants interacting with a variety of sonification methods in order to locate targets and place objects in a real kitchen environment. Prior to the tests, we enquired about the perceived benefits of increasing said empathy from the blind and visually impaired (BVI) community. To test empathy, we adapted an Empathy and Sympathy Response scale to gather sighted people's self-reported and perceived empathy with the BVI community from both sighted (N = 77) and BVI people (N = 20) respectively. We re-tested sighted people's empathy after the experiment and found that their empathetic and sympathetic responses (N = 15) significantly increased. Furthermore, survey results suggest that the BVI community believes the use of these empathy-evoking embodied experiences may lead to the development of new assistive technologies.
Renan Luigi Martins Guarese, Emma Pretty, Haytham M. Fayek, Fabio Zambetta, Ron G. van Schyndel
VR3
2023 PRE-NAS: Evolutionary Neural Architecture Search With Predictor
abstract
Neural architecture search (NAS) aims to automate architecture engineering in neural networks. This often requires a high computational overhead to evaluate a number of candidate networks from the set of all possible networks in the search space. Prediction of the performance of a network can alleviate this high computational overhead by mitigating the need for evaluating every candidate network. Developing such a predictor typically requires a large number of evaluated architectures which may be difficult to obtain. We address this challenge by proposing a novel evolutionary-based NAS strategy, predictor-assisted evolutionary NAS (PRE-NAS) which can perform well even with an extremely small number of evaluated architectures. PRE-NAS leverages new evolutionary search strategies and integrates high-fidelity weight inheritance over generations. Unlike one-shot strategies, which may suffer from bias in the evaluation due to weight sharing, offspring candidates in PRE-NAS are topologically homogeneous. This circumvents bias and leads to more accurate predictions. Extensive experiments on the NAS-Bench-201 and DARTS search spaces show that PRE-NAS can outperform state-of-the-art NAS methods. With only a single GPU searching for 0.6 days, a competitive architecture can be found by PRE-NAS which achieves 2.40% and 24% test error rates on CIFAR-10 and ImageNet, respectively.
Yameng Peng, Andy Song, Victor Ciesielski, Haytham M. Fayek, Xiaojun Chang
IEEE Trans. Evol. Comput.4
2022 PRE-NAS: predictor-assisted evolutionary neural architecture search
abstract
Neural architecture search (NAS) aims to automate architecture engineering in neural networks. This often requires a high computational overhead to evaluate a number of candidate networks from the set of all possible networks in the search space during the search. Prediction of the networks' performance can alleviate this high computational overhead by mitigating the need for evaluating every candidate network. Developing such a predictor typically requires a large number of evaluated architectures which may be difficult to obtain. We address this challenge by proposing a novel evolutionary-based NAS strategy, Predictor-assisted E-NAS (PRE-NAS), which can perform well even with an extremely small number of evaluated architectures. PRE-NAS leverages new evolutionary search strategies and integrates high-fidelity weight inheritance over generations. Unlike one-shot strategies, which may suffer from bias in the evaluation due to weight sharing, offspring candidates in PRE-NAS are topologically homogeneous, which circumvents bias and leads to more accurate predictions. Extensive experiments on NAS-Bench-201 and DARTS search spaces show that PRE-NAS can outperform state-of-the-art NAS methods. With only a single GPU searching for 0.6 days, competitive architecture can be found by PRE-NAS which achieves 2.40% and 24% test error rates on CIFAR-10 and ImageNet respectively.
Yameng Peng, Andy Song, Victor Ciesielski, Haytham M. Fayek, Xiaojun Chang
GECCO4
2022 Knowledge Capture and Replay for Continual Learning
abstract
Deep neural networks model data for a task or a sequence of tasks, where the knowledge extracted from the data is encoded in the parameters and representations of the network. Extraction and utilization of these representations is vital when data is no longer available in the future, especially in a continual learning scenario. We introduce flashcards, which are visual representations that capture the encoded knowledge of a network as a recursive function of some predefined random image patterns. In a continual learning scenario, flashcards help to prevent catastrophic forgetting by consolidating the knowledge of all the previous tasks. Flashcards are required to be constructed only before learning the subsequent task, hence, they are independent of the number of tasks trained before, making them task agnostic. We demonstrate the efficacy of flashcards in capturing learned knowledge representation (as an alternative to the original data), and empirically validate on a variety of continual learning tasks: reconstruction, denoising, and task-incremental classification, using several heterogeneous (varying background and complexity) benchmark datasets. Experimental evidence indicates that: (i) flashcards as a replay strategy is task agnostic, (ii) performs better than generative replay, and (iii) is on par with episodic replay without additional memory overhead.
Saisubramaniam Gopalakrishnan, Pranshu Ranjan Singh, Haytham M. Fayek, Savitha Ramasamy, Arulmurugan Ambikapathi
WACV3
2020 Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data
abstract
Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds using both the audio and visual modalities as opposed to one or the other. We present an audiovisual fusion model that learns to recognize sounds from weakly labeled video recordings. The proposed fusion model utilizes an attention mechanism to dynamically combine the outputs of the individual audio and visual models. Experiments on the large scale sound events dataset, AudioSet, demonstrate the efficacy of the proposed model, which outperforms the single-modal models, and state-of-the-art fusion and multi-modal models. We achieve a mean Average Precision (mAP) of 46.16 on Audioset, outperforming prior state of the art by approximately +4.35 mAP (relative: 10.4%).
Haytham M. Fayek
IJCAI1
2020 Progressive learning: A deep learning framework for continual learning
Haytham M. Fayek, Lawrence Cavedon, Hong Ren Wu
Neural Networks1
2020 Temporal Reasoning via Audio Question Answering
abstract
Multimodal question answering tasks can be used as proxy tasks to study systems that can perceive and reason about the world. Answering questions about different types of input modalities stresses different aspects of reasoning such as visual reasoning, reading comprehension, story understanding, or navigation. In this article, we use the task of Audio Question Answering (AQA) to study the temporal reasoning abilities of machine learning models. To this end, we introduce the Diagnostic Audio Question Answering (DAQA) dataset comprising audio sequences of natural sound events and programmatically generated questions and answers that probe various aspects of temporal reasoning. We adapt several recent state-of-the-art methods for visual question answering to the AQA task, and use DAQA to demonstrate that they perform poorly on questions that require in-depth temporal reasoning. Finally, we propose a new model, Multiple Auxiliary Controllers for Linear Modulation (MALiMo) that extends the recent Feature-wise Linear Modulation (FiLM) model and significantly improves its temporal reasoning capabilities. We envisage DAQA to foster research on AQA and temporal reasoning and MALiMo a step towards models for AQA.
Haytham M. Fayek, Justin Johnson 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Evaluating deep learning architectures for Speech Emotion Recognition
Haytham M. Fayek, Margaret Lech, Lawrence Cavedon
Neural Networks1
2016 Modeling subjectiveness in emotion recognition with deep neural networks: Ensembles vs soft labels
abstract
Ground truth labels obtained by averaging or majority voting are commonly used to train automatic emotion classifiers. However, ground truth labels fail to encapsulate inter-annotator variability and ignore the subjectivity of emotions. In this paper, we propose two viable approaches to model the subjectiveness of emotions by incorporating inter-annotator variability, which are soft labels and model ensembling, where each model represents an annotator. Using a deep neural network that recognizes emotions in real-time from one second windows of speech spectrograms, we demonstrate that both approaches lead to consistent improvement over using ground truth labels. It is empirically shown that the performance gain of the ensemble over the baseline model could be achieved using soft labels generated from multiple annotators.
Haytham M. Fayek, Margaret Lech, Lawrence Cavedon
IJCNN1
2016 On the Correlation and Transferability of Features Between Automatic Speech Recognition and Speech Emotion Recognition
abstract
The correlation between Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER) is poorly understood. Studying such correlation may pave the way for integrating both tasks into a single system or may provide insights that can aid in advancing both systems such as improving ASR in dealing with emotional speech or embedding linguistic input into SER. In this paper, we quantify the relation between ASR and SER by studying the relevance of features learned between both tasks in deep convolutional neural networks using transfer learning. Experiments are conducted using the TIMIT and IEMOCAP databases. Results reveal an intriguing correlation between both tasks, where features learned in some layers particularly towards initial layers of the network for either task were found to be applicable to the other task with varying degree.
Haytham M. Fayek, Margaret Lech, Lawrence Cavedon
INTERSPEECH1