VLDB 2026 Research / reviewers in the wild / expert
Yehia El-khatib
dblp:22/1017 · also Yehia Elkhatib
· DBLP profile ↗
51ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0003-4639-436XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 5 since 2021Computer networks · 10 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A systematic evaluation of the potential of carbon-aware execution for scientific workflowsabstractScientific workflows are widely used to automate scientific data analysis and often involve computationally intensive processing of large datasets on compute clusters. As such, their execution tends to be long-running and resource-intensive, resulting in substantial energy consumption and, depending on the energy mix, carbon emissions. Meanwhile, a wealth of carbon-aware computing methods have been proposed, yet little work has focused specifically on scientific workflows, even though they present a substantial opportunity for carbon-aware computing because they are often significantly delay tolerant, efficiently interruptible, highly scalable and widely heterogeneous. In this study, we first exemplify the problem of carbon emissions associated with running scientific workflows, and then show the potential for carbon-aware workflow execution. For this, we estimate the carbon footprint of seven real-world Nextflow workflows executed on different cluster infrastructures using both average and marginal carbon intensity data. Furthermore, we systematically evaluate the impact of carbon-aware temporal shifting, and the pausing and resuming of the workflow. Moreover, we apply resource scaling to workflows and workflow tasks. Finally, we report the potential reduction in overall carbon emissions, with temporal shifting capable of decreasing emissions by over 80%, and resource scaling capable of decreasing emissions by 67%. Kathleen West, Youssef Moawad, Fabian Lehmann, Vasilis Bountris, Ulf Leser, Yehia El-khatib, Lauritz Thamsen |
Future Gener. Comput. Syst. | 6 |
| 2026 | Beyond Bitrate: Understanding the QoE Impact of Playback Rate and Seeking in Adaptive Video StreamingabstractQuality of Experience (QoE) is a key component in adaptive bitrate (ABR) streaming. Whilst the effects of delivery disruptions—such as changes in video quality or rebuffering—have been extensively studied, the impact of playback rate variations remains relatively unexplored. Existing work has examined the QoE impact of playback rate in isolation, without comparing it to other common ABR streaming artifacts such as video quality variations, rebuffering, or seeking. Moreover, the role of gradual playback rate transitions has not been explored. This article addresses these gaps through four large-scale subjective studies that provide a systematic QoE evaluation of playback rate. We compare the acceptability of playback rate changes against quality degradations and rebuffering, and assess rate increases relative to seeking, both of which are strategies for maintaining the desired latency in low-latency streaming. We further investigate how gradual versus instantaneous playback rate transitions affect QoE. Through our subjective studies, we identify levels of slowed-down playback that are imperceptible to users and can be combined with video quality adaptation in ABR algorithms to reduce rebuffering. Moreover, we identify imperceptible rates of speeded-up video playback, which can be used as part of catch-up mechanisms to maintain a desired latency in live streaming, offering a less detrimental alternative to seeking events, which we found to significantly degrade QoE. This work presents the first systematic comparison of playback rate variations with video quality degradation, rebuffering and seeking. The findings extend our understanding of QoE in adaptive video streaming and provide actionable design guidelines for video players to improve the user experience of streaming services. Tomasz Lyko, Yehia El-khatib, Rajiv Ramdhany, Nicholas J. P. Race |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | The IoT Whisperer: A Framework for Intelligent IoT Service Composition Through LLMsabstractManually configuring Internet of Things (IoT) work-flows often involves intricate and time-consuming processes. This complexity can lead to an increase in development time and difficulty in managing and scaling IoT systems. The challenge lies in creating efficient and deployable workflows that seamlessly integrate diverse IoT devices and services. To address these challenges, we propose an approach that leverages IoT system orchestration, workflow development, and Large Language Models (LLMs) to convert high-level natural language system descriptions into executable workflows. Utilizing the Web of Things (WoT) framework, Node-RED and LLMs, we enabled the automatic generation of workflow components from natural language descriptions, enhancing the accessibility and usability of IoT system development for non-experts. We describe an assessment framework for the evaluation of abstract workflow descriptions and use this to analyze the implementation of our approach over systems within various domains, comparing the automatically generated workflows with those produced manu-ally. Our approach provides a powerful and versatile solution for IoT system development. This approach not only simplifies the process, but also provides a foundation for future innovations in the IoT landscape, paving the way for more adaptable and user-friendly IoT solutions. Ewan Warburton, Abdessalam Elhabbash, Saad Ezzini, Yehia El-khatib |
CLOUD | 4 |
| 2025 | Exploring the Potential of Carbon-Aware Execution for Scientific WorkflowsabstractScientific workflows are widely used to automate scientific data analysis and often involve processing large quantities of data on compute clusters. As such, their execution tends to be long-running and resource intensive, leading to significant energy consumption and carbon emissions. Meanwhile, a wealth of carbon-aware computing methods have been proposed, yet little work has focused specifically on scientific workflows, even though they present a substantial opportunity for carbon-aware computing because they are inherently delay tolerant, efficiently interruptible, and highly scalable. In this study, we demonstrate the potential for carbonaware workflow execution. For this, we estimate the carbon footprint of two real-world Nextflow workflows executed on cluster infrastructure. We use a linear power model for energy consumption estimates and real-world average and marginal CI data for two regions. We evaluate the impact of carbonaware temporal shifting, pausing and resuming, and resource scaling. Our findings highlight significant potential for reducing emissions of workflows and workflow tasks. Kathleen West, Fabian Lehmann, Vasilis Bountris, Ulf Leser, Yehia El-khatib, Lauritz Thamsen |
CCGrid | 5 |
| 2025 | SecureMind: A Framework for Benchmarking Large Language Models in Memory Bug Detection and RepairabstractLarge language models (LLMs) hold great promise for automating software vulnerability detection and repair, but ensuring their correctness remains a challenge. While recent work has developed benchmarks for evaluating LLMs in bug detection and repair, existing studies rely on hand-crafted datasets that quickly become outdated. Moreover, systematic evaluation of advanced reasoning-based LLMs using chain-of-thought prompting for software security is lacking. We introduce SecureMind, an open-source framework for evaluating LLMs in vulnerability detection and repair, focusing on memory-related vulnerabilities. SecureMind provides a user-friendly Python interface for defining test plans, which automates data retrieval, preparation, and benchmarking across a wide range of metrics. Using SecureMind, we assess 10 representative LLMs, including 7 state-of-the-art reasoning models, on 16K test samples spanning 8 Common Weakness Enumeration (CWE) types related to memory safety violations. Our findings highlight the strengths and limitations of current LLMs in handling memory-related vulnerabilities. Huanting Wang, Dejice Jacob, David Kelly, Yehia El-khatib, Jeremy Singer, Zheng Wang 0001 |
ISMM | 4 |
| 2025 | xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
Jiabo Shi, Dimitrios P. Pezaros, Yehia El-khatib |
Middleware | 3 |
| 2025 | The Double-Edged Impact of User Customisation on QoE in Personalised Media ExperiencesabstractUser-driven experience customisation can potentially enhance the Quality of Experience (QoE) in personalised multimedia. Such experiences could be, for example, delivered using HTTP Adaptive Streaming - the most prominent way of consuming media over the Internet. In this paper, we present a subjective study designed to investigate the QoE impact of user-driven customisation in personalised media experiences. We offered participants the option to customise their video layout, after which we asked them to score a range of quality impairments found in HTTP Adaptive Streaming. Based on our analysis of the collected Mean Opinion Scores (MOS), we found that experience customisation impacts the QoE in a surprising way. We found that experience personalisation caused participants’ expectations to increase when it comes to QoE, as they perceived the most severe quality impairments worse than the control. This establishes the need for specialised QoE models that take into account different levels of user expectations. Tomasz Lyko, Edward Austin, Alexander Lee, Yehia El-khatib, Nicholas J. P. Race |
QoMEX | 4 |
| 2024 | Drop or Stop: Investigating the Impact of Playback Rate on QoE in Adaptive Video StreamingabstractQuality of Experience (QoE) is a crucial component of adaptive bitrate (ABR) streaming, with the effects of abrupt changes in playback quality or rebuffering, caused by delivery disruptions, being widely studied. However, the collective ABR community has a limited understanding of the effects of changes in playback rate on QoE. In this pioneering work, we investigate two aspects of playback rate fluctuations. In particular, we carry out two subjective studies to assess if a change in playback rate is more or less acceptable than a drop in video quality or a rebuffering event. Furthermore, we examine the effect of the transition in playback rate on QoE, comparing gradual and instant variations. Our subjective studies recruited 120 participants who evaluated 102 test sequences. In summary, we find that playback rate drops of 0.8-0.9 are imperceptible for most content, and rated similarly to a video quality drop to medium level. In contrast, lower playback rates of 0.6-0.7 were perceived as poorly as rebuffering events. Gradual changes in playback rate can offer better QoE, but only in limited cases depending on the content, target playback rate, as well as magnitude of change. Tomasz Lyko, Yehia El-khatib, Rajiv Ramdhany, Nicholas J. P. Race |
QoMEX | 2 |
| 2024 | Principled and automated system of systems composition using an ontological architectureabstractA distributed system’s functionality must continuously evolve, especially when environmental context changes. Such required evolution imposes unbearable complexity on system development. An alternative is to make systems able to self-adapt by opportunistically composing at runtime to generate systems of systems (SoSs) that offer value-added functionality. The success of such an approach calls for abstracting the heterogeneity of systems and enabling the programmatic construction of SoSs with minimal developer intervention. We propose a general ontology-based approach to describe distributed systems, seeking to achieve abstraction and enable runtime reasoning between systems. We also propose an architecture for systems that utilizes such ontologies to enable systems to discover and ‘understand’ each other, and potentially compose, all at runtime. We detail features of the ontology and the architecture through three contrasting case studies: one on controlling multiple systems in smart home environment, another on the management of dynamic computing clusters, and a third on autonomic connection of rescue teams. We also quantitatively evaluate the scalability and validity of our approach through experiments and simulations. Our approach enables system developers to focus on high-level SoS composition without being constrained by deployment-specific implementation details. We demonstrate the feasibility of our approach to raise the level of abstraction of SoS construction through reasoned composition at runtime. Our architecture presents a strong foundation for further work due to its generality and extensibility. Abdessalam Elhabbash, Yehia El-khatib, Vatsala Nundloll, Vicent Sanz Marco, Gordon S. Blair |
Future Gener. Comput. Syst. | 2 |
| 2023 | MARTIN: An End-to-end Microservice Architecture for Predictive Maintenance in Industry 4.0abstractThe amount of data generated in Industry 4.0 and the introduction of advanced data analytics support establishing “smart factories” and one of its crucial characteristics - predictive maintenance. Current solutions primarily focus on offline predictions and do not provide end-to-end scalable solutions. Furthermore, there is a lack of support for incremental ma-chine learning in predictive maintenance. This paper addresses these limitations by proposing MARTIN, a scalable microservice architecture for predictive maintenance that can collect, store, and analyse data, and make decisions based on the machine state. The architecture uses incremental learning as the basis for predictions. The designed system was implemented and its performance was evaluated experimentally. The results show that the solution can provide high prediction accuracy in terms of practical processing time. Abdessalam Elhabbash, Kamil Rogoda, Yehia El-khatib |
SSE | 3 |
| 2023 | NLP-based Generation of Ontological System Descriptions for Composition of Smart Home DevicesabstractWith the current rapid development of Internet of Things (IoT) technology and the widespread popularity of smart home devices, wireless technology has made it possible for IoT devices to integrate with each other in a complex system. Previous works have proposed utilizing ontological descriptions, called Holons, of IoT devices and subsystems to reason about the construction of systems. The holonic description, defined by an ontology, includes parameters, services, and properties of the IoT device. However, these previous works assume that Holon descriptions of IoT devices are already provided e.g., by vendors. This assumption requires device vendors and system engineers to manually create descriptions, which is time-consuming and error-prone given the increasing number of IoT devices that are offered in the market. This paper introduces a method for the automatic generation of Holon descriptions of IoT devices. This method uses the brand and model of a device and utilizes knowledge extraction in natural language processing to automatically generate the ontological description of the IoT device. The experimental results show that the proposed method can generate descriptions with a precision of 96.72% and a recall of 87.53% in a practically acceptable time. Yehia El-khatib, Abdessalam Elhabbash |
ICWS | 2 |
| 2023 | Information-guided Planning: An Online Approach for Partially Observable ProblemsabstractThis paper presents IB-POMCP, a novel algorithm for online planning under partial observability. Our approach enhances the decision-making process by using estimations of the world belief's entropy to guide a tree search process and surpass the limitations of planning in scenarios with sparse reward configurations. By performing what we denominate as an *information-guided planning process*, the algorithm, which incorporates a novel I-UCB function, shows significant improvements in reward and reasoning time compared to state-of-the-art baselines in several benchmark scenarios, along with theoretical convergence guarantees. Matheus Aparecido do Carmo Alves, Amokh Varma, Yehia El-khatib, Leandro Soriano Marcolino |
NeurIPS | 3 |
| 2023 | Differential QoE in Picture-in-Picture Gaming Videos: A Subjective StudyabstractVideo streaming continues to be the largest service delivered on the internet. This includes gaming videos, delivered both on-demand and live, where gaming footage is usually accompanied by a video of the player overlaid on top of the gameplay - resulting in Picture-In-Picture (PiP) content. Currently, PiP content is usually combined into a single video before being delivered to the client via technologies such as HTTP Adaptive Streaming (HAS). In this study, we investigated the QoE importance of gameplay and player elements in PiP gaming videos by varying the video quality of these elements individually. We conducted a subjective study, testing nine quality permutations based on three quality levels across three pieces of content from different gaming genres, with 30 participants recruited using an ethical crowdsourcing platform. We found that gameplay was significantly more important in terms of overall QoE, while the player element made a difference in only a few cases. Tomasz Lyko, Yehia El-khatib, Rajiv Ramdhany, Nicholas J. P. Race |
QoMEX | 2 |
| 2022 | On the Performance Benefits of Heterogeneous Virtual Network Function Execution FrameworksabstractAs the adoption of softwarized network functions (NFs) keeps growing, we evaluate the performance benefits of SDN-aware data-plane implementations when compared to diverse acceleration and process-based NFV frameworks. Typical network functions have been implemented using four alternative frameworks scenarios, an SDN-aware software switch (data-plane), a virtual machine (VM), a Data-Plane Development Kit (DPDK) NF, and a containerized NF. Results from our experiments show that the data-plane NF implementation yields much higher bandwidth and packets per second (pps) rates. The bandwidth obtained is 14% more than the user-space scenario while retaining CPU utilization. The DPDK NFs in our evaluation can process packets at a much higher rate for 64B packets, on a single CPU core, which is 7 times higher than the containerized NF implementations, also tied to a single core. Our results also show the performance gains from deploying virtual network functions on heterogeneous frameworks. Haruna Umar Adoga, Yehia El-khatib, Dimitrios P. Pezaros |
NetSoft | 2 |
| 2022 | QoE Assessment for Multi-Video Object Based MediaabstractRecent multimedia experiences using techniques such as DASH allow the streaming delivery to be adapted to suit network context. Object Based Media (OBM) provides even more flexibility as distinct media objects are streamed and combined based on user preferences, allowing the experience to be personalised for the user. As adaptation can lead to degradation, modelling and measuring Quality of Experience (QoE) are crucial to ensure a perceptibly-optimal user experience. QoE models proposed for DASH include quality-related factors from single video-object streams and hence, are unsuitable for multi-video OBM experiences. In this paper, we propose an objective method to quantify QoE for video-based OBM experiences. Our model provides different strategies to aggregate individual object QoE contributions for different OBM experience genres. We apply our model to a case study and contrast it with the QoE levels obtained using a standard QoE model for DASH. Tomasz Lyko, Yehia El-khatib, Michael Sparks, Rajiv Ramdhany, Nicholas J. P. Race |
QoMEX | 2 |
| 2022 | Transferable Knowledge for Low-Cost Decision Making in Cloud EnvironmentsabstractUsers of Infrastructure as a Service (IaaS) are increasingly overwhelmed with the wide range of providers and services offered by each provider. As such, many users select services based on description alone. An emerging alternative is to use a decision support system (DSS), which typically relies on gaining insights from observational data in order to assist a customer in making decisions regarding optimal deployment of cloud applications. The primary activity of such systems is the generation of a prediction model (e.g. using machine learning), which requires a significantly large amount of training data. However, considering the varying architectures of applications, cloud providers, and cloud offerings, this activity is not sustainable as it incurs additional time and cost to collect data to train the models. We overcome this through developing a Transfer Learning (TL) approach where knowledge (in the form of a prediction model and associated data set) gained from running an application on a particular IaaS is transferred in order to substantially reduce the overhead of building new models for the performance of new applications and/or cloud infrastructures. In this article, we present our approach and evaluate it through extensive experimentation involving three real world applications over two major public cloud providers, namely Amazon and Google. Our evaluation shows that our novel two-mode TL scheme increases overall efficiency with a factor of 60 percent reduction in the time and cost of generating a new prediction model. We test this under a number of cross-application and cross-cloud scenarios. Faiza Samreen, Gordon S. Blair, Yehia El-khatib |
IEEE Trans. Cloud Comput. | 3 |
| 2021 | Energy-Aware Placement of Device-to-Device Mediation Services in IoT Systems
Abdessalam Elhabbash, Yehia El-khatib |
ICSOC | 2 |
| 2021 | Attaining Meta-self-awareness through Assessment of Quality-of-KnowledgeabstractSelf-awareness is a crucial capability of autonomous service-based systems that enables them to self-adapt. There are different types of self-awareness whereby certain types of knowledge are captured at various levels. We argue that effective management of the trade-offs of dependability requirements can be achieved through “seamless” switching between different levels of awareness. However, the assessment of the quality of knowledge to enable dynamic switching between self-awareness levels has not been tackled yet. We propose a general architecture that exploits symbiotic simulation in order to tackle the complexity of assessing the quality of knowledge and attaining the meta-self-awareness property, wherein the system can reflect on its different levels of awareness. We conduct a thorough real-world study in the context of volunteer services. We conclude that a system made meta-self-aware using our approach achieves optimal performance by activating the most suitable awareness level. This comes at the cost of a modest computational overhead. Abdessalam Elhabbash, Rami Bahsoon, Peter Tiño, Peter R. Lewis 0001, Yehia El-khatib |
ICWS | 5 |
| 2021 | An Empirical Study of Inter-cluster Resource Orchestration within Federated Cloud ClustersabstractFederated clusters are composed of multiple independent clusters of machines interconnected by a resource management system, and possess several advantages over centralized cloud datacenter clusters including seamless provisioning of applications across large geographic regions, greater fault tolerance, and increased cluster resource utilization. However, while existing resource management systems for federated clusters are capable of improving application intra-cluster performance, they do not capture inter-cluster performance in their decision making. This is important given federated clusters must execute a wide variety of applications possessing heterogeneous system architectures, which are a impacted by unique inter-cluster performance conditions such as network latency and localized cluster resource contention. In this work we present an empirical study demonstrating how inter-cluster performance conditions negatively impact federated cluster orchestration systems. We conduct a series of micro-benchmarks under various cluster operational scenarios showing the critical importance in capturing inter-cluster performance for resource orchestration in federated clusters. From this benchmark, we determine precise limitations in existing federated orchestration, and highlight key insights to design future orchestration systems. Findings of notable interest entail different application types exhibiting innate performance affinities across various federated cluster operational conditions, and experience substantial performance degradation from even minor increases to latency (8.7x) and resource contention (12.0x) in comparison to centralized cluster architectures. Dominic Lindsay, Gingfung Yeung, Yehia El-khatib, Peter Garraghan |
JCC | 3 |
| 2020 | SDN Heading North: Towards a Declarative Intent-based Northbound InterfaceabstractThe Intent-Based Northbound Interface (NBI) offers users the ability to express what they want to achieve instead of how to achieve it, enabling improvements to network management and reducing operational costs. However, development of an Intent-Based NBI remains in its infancy. Existing solutions do not allow users to express high-level operational targets that appropriately capture business objectives, nor link these to lower-level management policies and operations. We propose an extensible Intent-based NBI framework and a higher-level declarative intent expression to enable service-oriented intents with different targets. We focus on the creation of intents and their mapping from high-level expressions to low-level policies, and consider this from the perspective of an intent developer in the context of a Cloud CDN use case. Shiyam Alalmaei, Yehia El-khatib, Mehdi Bezahaf, Matthew Broadbent, Nicholas J. P. Race |
CNSM | 2 |
| 2020 | Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model SelectionabstractDeep neural networks (DNNs) are becoming a key enabling technique for many application domains. However, on-device inference on battery-powered, resource-constrained embedding systems is often infeasible due to prohibitively long inferencing time and resource requirements of many DNNs. Offloading computation into the cloud is often unacceptable due to privacy concerns, high latency, or the lack of connectivity. Although compression algorithms often succeed in reducing inferencing times, they come at the cost of reduced accuracy. This article presents a new, alternative approach to enable efficient execution of DNNs on embedded devices. Our approach dynamically determines which DNN to use for a given input by considering the desired accuracy and inference time. It employs machine learning to develop a low-cost predictive model to quickly select a pre-trained DNN to use for a given input and the optimization constraint. We achieve this first by offline training a predictive model and then using the learned model to select a DNN model to use for new, unseen inputs. We apply our approach to two representative DNN domains: image classification and machine translation. We evaluate our approach on a Jetson TX2 embedded deep learning platform and consider a range of influential DNN models including convolutional and recurrent neural networks. For image classification, we achieve a 1.8x reduction in inference time with a 7.52% improvement in accuracy over the most capable single DNN model. For machine translation, we achieve a 1.34x reduction in inference time over the most capable single model with little impact on the quality of translation. Vicent Sanz Marco, Ben Taylor 0001, Zheng Wang 0001, Yehia El-khatib |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2019 | CadaML: A Modeling Language for Multi-Tenant Cloud Application Data ArchitecturesabstractMulti-tenancy is used for efficient resource utilization when cloud resources are shared across multiple customers. In cloud applications, the data layer is often the prime candidate for multi-tenancy, and usually comprises a combination of different cloud storage solutions such as relational and non-relational databases, and blob storage. Each of these storage types is different, requiring its own partitioning schemes to ensure tenant isolation and scalability. Current multi-tenant data architectures are implemented mainly through manual coding techniques that tend to be time consuming and error prone. As an alternative, we propose a domain-specific modeling language, CadaML, that provides concepts and notations to model a multi-tenant data architecture in an abstract way. CadaML also provides tools to validate the data architecture and automatically produce application code to implement said architecture. Assylbek Jumagaliyev, Yehia El-khatib |
CLOUD | 2 |
| 2019 | A Framework for SLO-driven Cloud Specification and BrokerageabstractThe diversity of cloud offerings motivated the proposition of cloud modelling languages (CMLs) to abstract complexities related to selection of cloud services. However, current CMLs lack the support for modelling service level objectives (SLOs) that are required for the customer applications. Consequently, we propose an application- and provider-independent SLO modelling language (SLO-ML) to enable customers to specify the required SLOs. We also sketch the architecture to realise SLO-ML. Abdessalam Elhabbash, Yehia El-khatib, Gordon S. Blair, Yuhui Lin, Adam Barker |
CCGRID | 2 |
| 2019 | Same Same, but Different: A Descriptive Intra-IaaS DifferentiationabstractUsers of cloud computing are overwhelmed with choice, even within the services offered by one provider. As such, many users select cloud services based on description alone. In this quantitative study, we investigate the services of 2 of major IaaS providers. We use 2 representative applications to obtain longitudinal observations over 7 days of the week and over different times of the day, totalling over 14,000 executions. We give evidence of significant variations of performance offered within IaaS services, calling for data-driven brokers that are able to offer automated and adaptive decision making processes with means for incorporating expressive user constraints. Yehia El-khatib, Faiza Samreen, Gordon S. Blair |
CCGRID | 1 |
| 2019 | Widening the Circle of Engagement Around Environmental Issues using Cloud-based ToolsabstractEnvironmental data are being generated and collected at unprecedented rates. However, the diversity in form and format of these environmental assets poses challenges for collaborative and reproducible science. Moreover, access constraints that surround environmental data lead to difficulty in use and interpretation of results. Cloud computing offers high potential to break down such barriers and engender collaboration, attribution, reuse, and reproducibility. In this article we review the design of the Environmental Virtual Observatory pilot (EVOp) that was conceived as a cloud-enabled virtual research space for different users interested in environmental science, ranging from domain specialists to the general public. We discuss the key technologies and processes used: a hybrid cloud infrastructure; standard service interfaces; a unified service delivery platform; and a test-driven development cycle. We also discuss the methodology by showcasing one of the exemplars developed in EVOp, stressing the importance of weaving stakeholder engagement from the beginning and throughout the process. We also briefly highlight some of the lessons learnt of working in an interdisciplinary team. Yehia El-khatib, Alastair L. Gemmell, Claudia Vitolo, Mark E. Wilkinson, Eleanor B. Mackay, Barbara J. Percy, Gordon S. Blair, Robert J. Gurney |
ICDCS | 1 |
| 2019 | A Modelling Language to Support the Evolution of Multi-tenant Cloud Data ArchitecturesabstractMulti-tenant data architectures enable efficient resource utilization in cloud applications, but are currently being implemented in industry and research using manual coding techniques that tend to be time consuming and error prone. We propose a novel domain-specific modeling language, CadaML, to automatically manage the development and evolution of cloud data architectures that (a) adopt multi-tenancy and/or (b) comprise of a combination of different storage solutions such as relational and non-relational databases, and blob storage. CadaML provides concepts and notations to support abstract modelling of a multi-tenant data architecture, and also provides tools to validate the data architecture and automatically produce application code. We rigorously evaluate CadaML through a user experiment where developers of various capabilities are asked to re-architect the data layer of an industrial business process analysis application. We observe that CadaML users required 3.5x less development time than manual coders. In addition to improved productivity, CadaML users highlighted other benefits gained in terms of reliability of generated code and usability. Assylbek Jumagaliyev, Yehia El-khatib |
MoDELS | 2 |
| 2019 | On Optimizing Backup Sharing Through Efficient VNF MigrationabstractWith the emergence of software defined networking and network function virtualization technologies, network services are expected to be offered as service function chains made out from virtual network functions that are connected to steer and process the incoming traffic. In this context, achieving the survivability of these chains against failures is a key challenge to ensure high availability and continuity of the services. A promising solution proposed in the literature is to provision backups for the virtual network functions that could be shared among multiple service chains. These backups are used in case of a failure to take over the failed functions and ensure service continuity. In this paper, we propose two solutions to efficiently place and provision the shared backups in order to ensure the survivability of the service chains against single node failures. The originality of these solutions is that they leverage the migration of virtual network functions to minimize the resources consumed by the backups. Simulation results show that, compared to existing solutions, the proposed schemes leveraging migration are able to reduce by up to 20% the amount of resources allocated for the shared backups while ensuring the survivability of the service chains. Saifeddine Aidi, Mohamed Faten Zhani, Yehia El-khatib |
NetSoft | 3 |
| 2019 | Emergent Overlays for Adaptive MANET BroadcastabstractMobile Ad-Hoc Networks (MANETs) allow distributed applications where no fixed network infrastructure is available. MANETs use wireless communication subject to faults and uncertainty, and must support efficient broadcast. Controlled flooding is suitable for highly-dynamic networks, while overlay-based broadcast is suitable for dense and more static ones. Density and mobility vary significantly over a MANET deployment area. We present the design and implementation of emergent overlays for efficient and reliable broadcast in heterogeneous MANETs. This adaptation technique allows nodes to automatically switch from controlled flooding to the use of an overlay. Interoperability protocols support the integration of both protocols in a single heterogeneous system. Coordinated adaptation policies allow regions of nodes to autonomously and collectively emerge and dissolve overlays. Our simulation of the full network stack of 600 mobile nodes shows that emergent overlays reduce energy consumption, and improve reliability and coverage compared to single protocols and to two previously-proposed adaptation techniques. Raziel Carvajal-Gomez, Yehia El-khatib, Laurent Réveillère, Etienne Rivière, Yérom-David Bromberg |
SRDS | 2 |
| 2019 | Research challenges in nextgen service orchestration
Luis Miguel Vaquero González, Félix Cuadrado, Yehia El-khatib, Jorge Bernal Bernabé, Satish Narayana Srirama, Mohamed Faten Zhani |
Future Gener. Comput. Syst. | 3 |
| 2018 | On Improving Service Chains Survivability Through Efficient Backup Provisioning
Saifeddine Aidi, Mohamed Faten Zhani, Yehia El-khatib |
CNSM | 3 |
| 2018 | Adaptive Service Deployment using In-Network Mediation
Abdessalam Elhabbash, Gordon S. Blair, Gareth Tyson, Yehia El-khatib |
CNSM | 4 |
| 2018 | Adaptive deep learning model selection on embedded systemsabstractThe recent ground-breaking advances in deep learning networks (DNNs) make them attractive for embedded systems. However, it can take a long time for DNNs to make an inference on resource-limited embedded devices. Offloading the computation into the cloud is often infeasible due to privacy concerns, high latency, or the lack of connectivity. As such, there is a critical need to find a way to effectively execute the DNN models locally on the devices. Ben Taylor 0001, Vicent Sanz Marco, Willy Wolff, Yehia El-khatib, Zheng Wang 0001 |
LCTES | 4 |
| 2017 | Fake it till you make it: Fishing for CatfishesabstractMany adult content websites incorporate social networking features. Although these are popular, they raise significant challenges, including the potential for users to "catfish", i.e., to create fake profiles to deceive other users. This paper takes an initial step towards automated catfish detection. We explore the characteristics of the different age and gender groups, identifying a number of distinctions. Through this, we train models based on user profiles and comments, via the ground truth of specially verified profiles. When applying our models for age and gender estimation to unverified profiles, 38% of profiles are classified as lying about their age, and 25% are predicted to be lying about their gender. The results suggest that women have a greater propensity to catfish than men. Our preliminary work has notable implications on operators of such online social networks, as well as users who may worry about interacting with catfishes. Walid Magdy, Yehia El-khatib, Gareth Tyson, Sagar Joglekar 0001, Nishanth Sastry |
ASONAM | 2 |
| 2017 | Navigating Diverse Data Science Learning: Critical Reflections Towards Future PracticeabstractData Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learned lessons that are of relevance to others teaching Data Science. Yehia El-khatib |
CloudCom | 1 |
| 2017 | Using DSML for Handling Multi-tenant Evolution in Cloud ApplicationsabstractMulti-tenancy is sharing a single application's resources to serve more than a single group of users (i.e. tenant). Cloud application providers are encouraged to adopt multi-tenancy as it facilitates increased resource utilization and ease of maintenance, translating into lower operational and energy costs. However, introducing multi-tenancy to a single-tenant application requires significant changes in its structure to ensure tenant isolation, configurability and extensibility. In this paper, we analyse and address the different challenges associated with evolving an application's architecture to a multi-tenant cloud deployment. We focus specifically on multi-tenant data architectures, commonly the prime candidate for consolidation and multi-tenancy. We present a Domain-Specific Modeling language (DSML) to model a multi-tenant data architecture, and automatically generate source code that handles the evolution of the application's data layer. We apply the DSML on a representative case study of a single-tenant application evolving to become a multi-tenant cloud application under two resource sharing scenarios. We evaluate the costs associated with using this DSML against the state of the art and against manual evolution, reporting specifically on the gained benefits in terms of development effort and reliability. Assylbek Jumagaliyev, Jon Whittle 0001, Yehia El-khatib |
CloudCom | 3 |
| 2017 | Charting an intent driven networkabstractThe strong divide between applications and the network control plane is desirable, but keeps the network in the dark regarding the ultimate purpose of applications and, as a result, is unable to optimize for these. An alternative approach is for applications to declare to the network their abstract desires; e.g. “I require group multicast”, or “I will run within a local domain and am latency sensitive”. Such an enriched semantic has the potential to enable the network to better fulfill application intent, while also helping optimize network resource usage across applications. We refer to this approach as intent driven networking (IDN). We sketch an incrementally-deployable design to serve as a stepping stone towards a practical realization of IDN within today's Internet. Yehia El-khatib, Geoff Coulson, Gareth Tyson |
CNSM | 1 |
| 2016 | It Bends But Would It Break? Topological Analysis of BGP Infrastructures in EuropeabstractThe Internet is often thought to be a model of resilience, due to a decentralised, organically-grown architecture. This paper puts this perception into perspective through the results of a security analysis of the Border Gateway Protocol (BGP) routing infrastructure. BGP is a fundamental Internet protocol and its intrinsic fragilities have been highlighted extensively in the literature. A seldom studied aspect is how robust the BGP infrastructure actually is as a result of nearly three decades of perpetual growth. Although global black-outs seem unlikely, local security events raise growing concerns on the robustness of the backbone. In order to better protect this critical infrastructure, it is crucial to understand its topology in the context of the weaknesses of BGP and to identify possible security scenarios. Firstly, we establish a comprehensive threat model that classifies main attack vectors, including but non limited to BGP vulnerabilities. We then construct maps of the European BGP backbone based on publicly available routing data. We analyse the topology of the backbone and establish several disruption scenarios that highlight the possible consequences of different types of attacks, for different attack capabilities. We also discuss existing mitigation and recovery strategies, and we propose improvements to enhance the robustness and resilience of the backbone. To our knowledge, this study is the first to combine a comprehensive threat analysis of BGP infrastructures withadvanced network topology considerations. We find that the BGP infrastructure is at higher risk than already understood, due to topologies that remain vulnerable to certain targeted attacks as a result of organic deployment over the years. Significant parts of the system are still uncharted territory, which warrants further investigation in this direction. Sylvain Frey, Yehia El-khatib, Awais Rashid, Karolina Follis, John Edward Vidler, Nicholas J. P. Race, Christopher Edwards |
EuroS&P | 2 |
| 2016 | Using context switches for VM scalingabstractVirtualisation encourages users to procure, relin-quish and scale resources frequently. Such fluid deployments require continuous performance assessment to inform resource allocation. In current systems Tick Accounting, Load Average, and Memory Usage are used to judge system performance and to trigger resource scaling. We argue that the readily available Context Switch (CS) counter is also an effective performance indicator. We demonstrate the efficacy of CS and show that it can also be used as base for scaling decisions. James Hadley, Utz Roedig, Yehia El-khatib |
IPCCC | 3 |
| 2016 | Daleel: Simplifying cloud instance selection using machine learningabstractDecision making in cloud environments is quite challenging due to the diversity in service offerings and pricing models, especially considering that the cloud market is an incredibly fast moving one. In addition, there are no hard and fast rules; each customer has a specific set of constraints (e.g. budget) and application requirements (e.g. minimum computational resources). Machine learning can help address some of the complicated decisions by carrying out customer-specific analytics to determine the most suitable instance type(s) and the most opportune time for starting or migrating instances. We employ machine learning techniques to develop an adaptive deployment policy, providing an optimal match between the customer demands and the available cloud service offerings. We provide an experimental study based on extensive set of job executions over a major public cloud infrastructure. Faiza Samreen, Yehia El-khatib, Matthew Rowe 0001, Gordon S. Blair |
NOMS | 2 |
| 2016 | Measurements and Analysis of a Major Adult Video PortalabstractToday, the Internet is a large multimedia delivery infrastructure, with websites such as YouTube appearing at the top of most measurement studies. However, most traffic studies have ignored an important domain: adult multimedia distribution. Whereas, traditionally, such services were provided primarily via bespoke websites, recently these have converged towards what is known as “Porn 2.0”. These services allow users to upload, view, rate, and comment on videos for free (much like YouTube). Despite their scale, we still lack even a basic understanding of their operation. This article addresses this gap by performing a large-scale study of one of the most popular Porn 2.0 websites: YouPorn. Our measurements reveal a global delivery infrastructure that we have repeatedly crawled to collect statistics (on 183k videos). We use this data to characterise the corpus, as well as to inspect popularity trends and how they relate to other features, for example, categories and ratings. To explore our discoveries further, we use a small-scale user study, highlighting key system implications. Gareth Tyson, Yehia El-khatib, Nishanth Sastry, Steve Uhlig |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2015 | Are People Really Social in Porn 2.0?
Gareth Tyson, Yehia El-khatib, Nishanth Sastry, Steve Uhlig |
ICWSM | 2 |
| 2015 | The design of a generalised approach to the programming of systems of systemsabstractThe world's computing infrastructure is increasingly differentiating into self-contained sub-systems (e.g. Internet of Things installations, clouds, VANETs, ...), which are post-hoc composed to generate value-added functionality (“systems of systems”). Today, however, such system-of-systems composition is typically carried out in an ad-hoc and infrastructure-dependent manner, with obvious associated disadvantages. In this paper, we propose a generalised system-of-systems-oriented programming approach that enables programmers to manage the composition of systems without a need for intimate knowledge of their internals, and also facilitates dynamic and spontaneous system composition, as systems discover each other opportunistically in their environment. Geoff Coulson, Gordon S. Blair, Yehia El-khatib, Andreas Mauthe |
WOWMOM | 3 |
| 2015 | Passive Network Awareness as a Means for Improved Grid Scheduling
Yehia El-khatib, Chris Edwards |
J. Grid Comput. | 1 |
| 2014 | Just Browsing?: Understanding User Journeys in Online TVabstractUnderstanding the dynamics of user interactions and the behaviour of users as they browse for content is vital for advancements in content discovery, service personalisation, and recommendation engines which ultimately improve quality of user experience. In this paper, we analyse how more than 1,100 users browse an online TV service over a period of six months. Through the use of model-based clustering, we identify distinctive groups of users with discernible browsing patterns that vary during the course of the day. Yehia El-khatib, Rebecca Killick, Mu Mu 0001, Nicholas J. P. Race |
ACM Multimedia | 1 |
| 2014 | Can SPDY really make the web faster?abstractHTTP is a successful Internet technology on top of which a lot of the web resides. However, limitations with its current specification have encouraged some to look for the next generation of HTTP. In SPDY, Google has come up with such a proposal that has growing community acceptance, especially after being adopted by the IETF HTTPbis-WG as the basis for HTTP/2.0. SPDY has the potential to greatly improve web experience with little deployment overhead, but we still lack an understanding of its true potential in different environments. This paper offers a comprehensive evaluation of SPDY's performance using extensive experiments. We identify the impact of network characteristics and website infrastructure on SPDY's potential page loading benefits, finding that these factors are decisive for an optimal SPDY deployment strategy. Through exploring such key aspects that affect SPDY, and accordingly HTTP/2.0, we feed into the wider debate regarding the impact of future protocols. Yehia El-khatib, Gareth Tyson, Michael Welzl |
Networking | 1 |
| 2014 | Dataset on usage of a live & VoD P2P IPTV serviceabstractThis paper presents a dataset of user statistics collected from a P2P multimedia service infrastructure that delivers both live and on-demand content in high quality to users via different platforms: PC/Mac, and set top boxes. The dataset covers a period of seven months starting from October 2011, exposing a total of over 94k system statistic reports from thousands of user devices at a fine granularity. Such rich data source is made available to fellow researchers to aid in developing better understanding of video delivery mechanisms, user behaviour, and programme popularity evolution. Yehia El-khatib, Mu Mu 0001, Nicholas J. P. Race |
P2P | 1 |
| 2013 | Demystifying porn 2.0: a look into a major adult video streaming websiteabstractThe Internet has evolved into a huge video delivery infrastructure, with websites such as YouTube and Netflix appearing at the top of most traffic measurement studies. However, most traffic studies have largely kept silent about an area of the Internet that (even today) is poorly understood: adult media distribution. Whereas ten years ago, such services were provided primarily via peer-to-peer file sharing and bespoke websites, recently these have converged towards what is known as ``Porn 2.0''. These popular web portals allow users to upload, view, rate and comment videos for free. Despite this, we still lack even a basic understanding of how users interact with these services. This paper seeks to address this gap by performing the first large-scale measurement study of one of the most popular Porn 2.0 websites: YouPorn. We have repeatedly crawled the website to collect statistics about 183k videos, witnessing over 60 billion views. Through this, we offer the first characterisation of this type of corpus, highlighting the nature of YouPorn's repository. We also inspect the popularity of objects and how they relate to other features such as the categories to which they belong. We find evidence for a high level of flexibility in the interests of its user base, manifested in the extremely rapid decay of content popularity over time, as well as high susceptibility to browsing order. Using a small-scale user study, we validate some of our findings and explore the infrastructure design and management implications of our observations. Gareth Tyson, Yehia El-khatib, Nishanth Sastry, Steve Uhlig |
Internet Measurement Conference | 2 |
| 2012 | A Trace-Driven Analysis of Caching in Content-Centric NetworksabstractA content-centric network is one which supports host-to-content routing, rather than the host-to-host routing of the existing Internet. This paper investigates the potential of caching data at the router-level in content-centric networks. To achieve this, two measurement sets are combined to gain an understanding of the potential caching benefits of deploying content-centric protocols over the current Internet topology. The first set of measurements is a study of the BitTorrent network, which provides detailed traces of content request patterns. This is then combined with CAIDA's ITDK Internet traces to replay the content requests over a real-world topology. Using this data, simulations are performed to measure how effective content-centric networking would have been if it were available to these consumers/providers. We find that larger cache sizes (10,000 packets) can create significant reductions in packet path lengths. On average, 2.02 hops are saved through caching (a 20% reduction), whilst also allowing 11% of data requests to be maintained within the requester's AS. Importantly, we also show that these benefits extend significantly beyond that of edge caching by allowing transit ASes to also reduce traffic. Gareth Tyson, Sebastian Kaune, Simon Miles, Yehia El-khatib, Andreas Mauthe, Adel Taweel |
ICCCN | 4 |
| 2010 | Characterising a grid site's trafficabstractGrid computing has been widely adopted for intensive high performance computing. Since grid resources are distributed over complex large-scale infrastructures, understanding grid site data traffic behaviour is important for efficient resource utilisation, performance optimisation, and the design of future grid sites as well as traffic-aware grid applications. In this paper, we study and analyse the traffic generated at a grid site in the Large Hadron Collider (LHC) Computing Grid (LCG). We find that most of the generated traffic is TCP-based and that a small set of grid applications generate significant amounts of the data. Upon analysing the different traffic metrics, we also find that the traffic exhibits long-range dependence and self-similarity. We also investigate packet-level metrics such as throughput, packet rate, round trip time (RTT) and packet loss. Our study establishes that these metrics can be well represented by Gaussian mixture models. The findings we present in this paper will enable accurate grid site traffic monitoring and potentially on-the-fly traffic modelling and prediction. It will also lead to a better understanding of grid site's traffic behaviour and contribute to more efficient grid site planning, traffic management, data transmission protocol optimisation, and data-aware grid application design. Tiejun Ma, Yehia El-khatib, Michael Mackay 0001, Christopher Edwards |
HPDC | 2 |
| 2009 | Providing Grid Schedulers with Passive Network MeasurementsabstractGrids offer the potential to carry out difficult computing tasks and achieve superior aggregate performance. However, grids are highly complex systems. They consist of heterogeneous resources on disparate hosts from various virtual organizations interconnected via a mixture of communication standards. Monitoring grid resources allows grid schedulers to adapt to changes in the status of these remote resources and the network paths between them. This is crucial to ensuring optimum performance. In this paper we introduce a distributed solution, called GridMAP, to collect network and end-host resource measurements, analyze their performance and feed these statistics and predictions back to schedulers. At this stage, we present our implementation of a passive TCP-SYN-based technique to provide GridMAP with round trip time and throughput measurements and we evaluate our approach against ping and iperf. Yehia El-khatib, Christopher Edwards, Michael Mackay 0001, Gareth Tyson |
ICCCN | 1 |
| 2009 | CompactPSH: An efficient transitive TFT incentive scheme for Peer-to-Peer NetworksabstractIncentive schemes in Peer-to-Peer (P2P) networks are necessary to discourage free-riding. One example is the Tit-for-Tat (TFT) incentive scheme, a variant of which is used in BitTorrent to encourage peers to upload. TFT uses data from local observations making it suitable for systems with direct reciprocity. This paper presents CompactPSH, an incentive scheme that works with direct and indirect reciprocity. CompactPSH allows peers to establish indirect reciprocity by finding intermediate peers, thus enabling trade with more peers and capitalizing on more resources. CompactPSH finds transitive paths while keeping the overhead of additional messages low. In a P2P file-sharing scenario based on input data from a large BitTorrent tracker, CompactPSH was found to exploit more reciprocity than TFT which enabled more chunks to be downloaded. As a consequence, peers are allowed to be stricter to fight white-washing without compromising performance. Thomas Bocek, Fabio Victora Hecht, David Hausheer, Burkhard Stiller, Yehia El-khatib |
LCN | 5 |