Tarek Zaarour

dblp:201/4834 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-9807-4109ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Intent-Based Agentic AI Framework For Data Placement in Compute Continuum
Ahmed Khalid, Ali Amin, Tarek Zaarour, Michal Sworzeniowski
ICNP3
2025 Secure Onboarding of Devices and Applications to a Smart Decentralized Ecosystem
abstract
Data Confidence Fabrics (DCFs) are emerging as a mechanism to obtain measurable trust in decentralized smart computing environments, while remote attestation (RA) is being established as a key mechanism for verifying security in distributed systems. However, both of these approaches remain underutilized in container orchestration platforms, resulting in inadequate trust guarantees, increased attack surface areas, and insufficient mechanisms for verifying the integrity of devices and applications. In this paper, we propose leveraging DCFs to integrate RA with software security practices, and utilizing this integration to securely onboard devices and containerized applications by tracing their provenance and analyzing vulner-abilities. This dual-level protection bridges device attestation measurements with measurements of application security to provide measurable and transparent confidence scores for both. The proposed approach brings trustworthiness into a distributed environment where devices are continuously monitored for malicious tampering, and applications are assessed before being run. We verified this approach on a setup representing real environments. Additionally, a machine learning model was put under test and trained on data weighted with confidence scores produced by our proposed approach. The model saw improved performance and accuracy, showing that this approach can increase the reliability of systems.
Ali Amin, Tarek Zaarour, Ahmed Khalid, Seán Óg Murphy, Utz Roedig, Cormac J. Sreenan
SMARTCOMP2
2024 Using Distributed Ledgers To Build Knowledge Graphs For Decentralized Computing Ecosystems
abstract
Knowledge graphs have proven vital for efficient data management, enhanced search capabilities, and improved decision-making in various information technology domains. However, constructing reliable knowledge graphs in decentralized ecosystems, with distributed autonomous actors, poses significant challenges related to asynchronous transmission, out-of-order knowledge-sharing, device heterogeneity, and trust issues. These challenges are also present in resource orchestration within multi-cloud edge ecosystems where multiple stakeholders must collaborate and share information to enable next-gen smart applications. In this paper, we propose a novel system design that utilizes Distributed Ledger Technology to build knowledge graphs. This approach ensures consistent and trustworthy knowledge sharing among orchestrators in a cloud-edge continuum. Our solution accommodates diverse requirements of both cloud and edge servers, allowing clients to construct complete historic graphs or build filtered sub-graphs. We deploy our solution in a multi-cloud edge environment and construct knowledge graphs representing the system state, including clusters, servers, microservices, and various resources. We validate the feasibility and performance of our solution through a real-world deployment and experiments in a smart shopping use case. Results demonstrate that the proposed solution achieves the claimed benefits with minimal or acceptable delays in comparison to traditional event streaming services.
Tarek Zaarour, Ahmed Khalid, Preeja Pradeep, Ahmed H. Zahran
CIKM1
2024 Safeguarding Blockchain and Dlts From Arbitrary and Malicious Content
abstract
Public blockchains and distributed ledger technologies (DLTs) provide reliable distributed, decentralized datasharing, with consensus mechanisms and cryptographic security, data integrity, transparency and trust. Though this transparency and open access allows malicious actors to record arbitrary information to the DLT by exploiting free-form editable text fields in request headers or transaction bodies. The authors propose an architecture that runs external agents interfaced with blockchains, capable of filtering unwanted content to mitigate this threat. Implementation of a resource description framework (RDF) based verification mechanism for blockchain input fields, and an alternate approach using natural language processing (NLP) are then described. Finally, the efficiency and effectiveness of the solution is demonstrated on a blockchain deployment.
Alan Barnett, Merry Globin, Tarek Zaarour, Sean Ahearne, Ahmed Khalid
ICNP3
2024 ELEVATE: Optimal Scheduling of Time-Sensitive Tasks on the Heterogeneous Reconfigurable Edge
abstract
Edge computing is evolving to include heterogeneous compute nodes with distinct characteristics. Graphic processing units (GPU) and field-programmable gate arrays (FPGA) can execute demanding deep learning (DL) tasks while meeting the deadlines of time-sensitive applications. However, FPGAs require reconfiguration to execute different tasks. In this paper, we first demonstrate that FPGAs can be reconfigured in real-time. Additionally, we propose ELEVATE as a novel scheduling algorithm for reconfigurable heterogeneous edge computing platforms targeting Industry 4.0 post-production quality control. ELEVATE design focusses on optimising the reconfiguration of the FPGA unit for heterogeneous quality inspection tasks. Our simulations indicate that ELEVATE reduces task waiting time by up to two orders of magnitude and achieves energy savings of up to 25 % compared to a statically configured FPGA unit.
Ingo Hoyer, Tarek Zaarour, Ahmed Khalid, Alexander Utz, Karsten Seidl, Ken Brown, Ahmed H. Zahran
ICNP2
2023 Foundation Data Space Models: Bridging the Artificial Intelligence and Data Ecosystems (Vision Paper)
abstract
Two major trends significantly changed the global Artificial Intelligence (AI) and Data landscape. Recent AI and Machine Learning developments are driving a paradigm shift to creating large task-agnostic foundation models pre-trained using web-scale data. Foundation models are then adapted to different downstream tasks via techniques such as fine-tuning. At the same time, we see a movement to the creation of large-scale data-sharing infrastructures. Data Spaces are an emerging approach to data management and sharing at the core of the European Data Strategy to provide access to high-quality data for AI. This paper brings together work on foundation models and data spaces into a holistic vision for Foundation Data Space Models. The paper highlights the data management requirements challenges for data spaces and details a high-level approach for foundation data space models together with a unified lifecycle for data spaces and foundation models. Finally, it sets out a research agenda.
Edward Curry, Tarek Zaarour, Yang Yang 0008, Mohan Timilsina, Majjed Al-Qatf, Rafiqul Haque
IEEE Big Data2
2022 SemanticPeer: A distributional semantic peer-to-peer lookup protocol for large content spaces at internet-scale
abstract
Peer-to-peer networks offer a solid foundation for wide-scale resource sharing, collaborative computing, and data distribution. Such networks have been commonly used for group communication by overlaying a publish/subscribe service atop their routing substrate. In this work, we focus on offering a group communication service that targets unstructured content such as images and videos for dissemination at internet scale. The decoupled nature of publish/subscribe systems exacerbated by the decentralized and large-scale nature of peer-to-peer networks brings about a semantic boundary between publishers and subscribers. More precisely, the large semantic space of human-level recognition creates a very large content space of object labels, attributes, and relationships. The scale of the content space makes it nearly impossible for participants to agree on a bounded set of terms for subscribers to express their exact interests. We identify an inherent limitation of peer-to-peer networks lying in the exact-match property of their key-based routing primitives. We propose an approximate matching model where participants agree on a distributional model of word meaning that maps terms to a vector space. We overcome the exact-match limitation by proposing a novel distributed lookup protocol and algorithm to construct a peer-to-peer network and route content. We replace conventional logical key spaces with a high-dimensional vector space that preserves the semantic properties of the data being mapped. Experiments show that the proposed model achieves more than 97% recall in routing accuracy, that is, locating a node responsible for storing a data item in a few routing hops. Furthermore, results also show that the network achieves over 90% recall in approximately matching two semantically related terms via rendezvous routing.
Tarek Zaarour, Edward Curry
Future Gener. Comput. Syst.1
2022 OpenPubSub: Supporting Large Semantic Content Spaces in Peer-to-Peer Publish/Subscribe Systems for the Internet of Multimedia Things
abstract
The decentralized and highly scalable nature of structured peer-to-peer networks, based on distributed hash tables (DHTs), makes them a great fit for facilitating the interaction and exchange of information between dynamic and geographically dispersed autonomous entities. The recent emergence of multimedia-based services and applications in the Internet of Things (IoT) has led to a noticeable shift in the type of data traffic generated by sensing devices from structured textual and numerical content to unstructured and bulky multimedia content. The wide semantic spectrum of human recognizable concepts that can be stemmed from multimedia data, e.g., video and audio, introduces a very large semantic content space. The scale of the content space poses a semantic boundary between data consumers and producers in large-scale peer-to-peer publish/subscribe systems. The exact-match query model of DHTs falls short when participants use different terms to describe the same semantic concepts. In this work, we present OpenPubSub, a peer-to-peer content-based approximate semantic publish/subscribe system. We propose a hybrid event routing model that combines rendezvous routing and gossiping over a structured peer-to-peer network. The network is built on the basis of a high-dimensional semantic vector space as opposed to conventional logical key spaces. We propose methods to partition the space, construct a semantic DHT via bootstrapping, perform approximate semantic lookup operations, and cluster nodes based on their shared interests. Results show that for an approximate event matching upper bound recall of 56.7%, rendezvous-based routing achieves up to 54% recall while decreasing the messaging overhead by 44%, whereas, the hybrid routing approach achieves up to 43.8% recall while decreasing the messaging overhead by 59%.
Tarek Zaarour, Anuraag Bhattacharya, Edward Curry
IEEE Internet Things J.1
2019 Adaptive Filtering of Visual Content in Distributed Publish/Subscribe Systems
abstract
Classic event matching techniques in large-scale Content-based Publish/Subscribe Systems mostly rely on predicate indexing or tree-based mechanisms for fast subscription evaluation. In the context of visual analytics, such techniques are limited in supporting subscriptions requiring expensive filtering operators over unstructured event types (i.e. images and videos). In this work, user subscriptions over visual content are answered as conjunctions of commutative Boolean filters where each filter is associated with a single high-level semantic concept that may be shared across multiple subscriptions. The shared-filter ordering problem has been previously studied in centralized data stream management systems; prior works propose approximation algorithms that achieve near-optimal cost reductions in the evaluation of overlapping queries. However, in a distributed publish/subscribe setting, even an optimal ordering of filter evaluations at brokers with high workloads can create bottlenecks and waste downstream resources. We present a distributed greedy algorithm that leverages existing routing methodologies to order and distribute the execution of filters across brokers on various dissemination paths. Experiments with several pub/sub workloads show 50% to 70% decrease in event latencies and noticeable improvements in resource utilization across the overlay.
Tarek Zaarour, Edward Curry
NCA1