VLDB 2026 Research / reviewers in the wild / expert
Mohamed Ali Kâafar
dblp:71/5612 · also Dali Kaafar, Mohamad Ali Kâafar
· DBLP profile ↗
125ranked-venue papers
7as first author
42since 2021 · last 2026
0000-0003-2714-0276ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 49 · 24 since 2021Computer networks · 40 · 7 first-author · 1 since 2021Databases, data management, data science and information retrieval · 18 · 10 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 6Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DIALEVAL: Automated Type-Theoretic Evaluation of LLM Instruction Following
Nardine Basta, Mohamed Ali Kâafar |
PAKDD (4) | 2 |
| 2026 | Forget Me, Not My Friends! Object Unlearning Based on Scene GraphsabstractMachine unlearning offers a practical technical means for fulfilling users' requests to remove personally identifiable information (PII) under ''right to be forgotten'' regulations such as GDPR and COPPA. Traditionally, unlearning is performed with the removal of entire data samples (sample unlearning) or whole features across the dataset (feature unlearning). However, when the removal request targets only certain parts of the PII, such as specific objects within a sample, these traditional unlearning approaches fall short of meeting such finer-grained unlearning requirements. To address this gap, we propose a scene graph-based object unlearning framework. This framework utilizes scene graphs, rich in semantic representation, transparently translate unlearning requests into actionable steps. The result, is the preservation of the overall semantic integrity of the generated image, bar the unlearned object. Furthermore, we develop three distinct approaches for object unlearning, grounded in the mainstream unlearning techniques of fine-tuning and model redaction. For validation, we evaluate the unlearned object's fidelity in outputs under the tasks of image reconstruction and image synthesis. Our proposed framework demonstrates improved object unlearning outcomes, with the preservation of unrequested samples in contrast to sample and feature learning methods. This work addresses critical privacy issues by increasing the granularity of targeted machine unlearning through forgetting specific object-level details without sacrificing the utility of the whole data sample or dataset feature. Chenhan Zhang, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Weiqi Wang 0003, An Liu 0002, Mohamed Ali Kâafar |
WSDM | 6 |
| 2026 | dX-Privacy for Text and the Curse of DimensionalityabstractA widely used method to ensure privacy of unstructured text data is the multidimensional Laplace mechanism for dX-privacy, which is a relaxation of differential privacy for metric spaces. We identify an intriguing peculiarity of this mechanism. When applied on a word-by-word basis, the mechanism either outputs the original word, or completely dissimilar words, and very rarely outputs semantically similar words. We investigate this observation in detail, and tie it to the fact that the distance of the nearest neighbor of a word in any word embedding model (which are high-dimensional) is much larger than the relative difference in distances to any of its two consecutive neighbors. We also show that the dot product of the multidimensional Laplace noise vector with any word embedding plays a crucial role in designating the nearest neighbor. We derive the distribution, moments and tail bounds of this dot product. We further propose a fix as a post-processing step, which satisfactorily removes the above-mentioned issue. Hassan Jameel Asghar, Robin Carpentier, Benjamin Zi Hao Zhao, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 4 |
| 2025 | Membership Inference Attack Vulnerabilities of Record Linkage ModelsabstractRecord linkage plays a crucial role in integrating health, legal, and administrative data. In domains such as healthcare, unstructured records, including clinical notes, contain rich information, making their integration valuable for applications like clinical trials. While deep learning models have improved linkage quality, their privacy risks remain under-explored. We present what is, to our knowledge, the first systematic study of membership inference attack vulnerabilities in record linkage models trained on de-identified texts. Unlike traditional classifiers, linkage models operate on record pairs, prompting a rethinking of what constitutes membership leakage. Does it occur if only a record pair was seen together during training or individually? Or if neither was seen, but the pair resembles known patterns? We introduce a black-box attack based on semantic perturbation sensitivity, requiring no access to the model's internals. Our findings expose a previously unaddressed membership inference vulnerability in record linkage models: black-box attacks, even with a simple threshold-based attack model, achieved up to 92% precision and AUC 0.79. Across record linkage models trained on MIMIC-IV and PMC-Patients datasets, we observe that perturbing training-seen phrases causes significantly larger confidence shifts (e.g., Δ s = -0.648) and higher change in predicted label (label flip from 0 = non-match to 1 = match and vice versa) (e.g., 68.75%) compared to unseen variants. These preliminary results reveal membership inference vulnerabilities in text-based linkage systems, highlighting the need for deeper investigation into privacy risks, motivating new lines of privacy defences for pairwise models. Piyumi Seneviratne, Dinusha Vatsalan, Mohamed Ali Kâafar |
CIKM | 3 |
| 2025 | CARE: Enhancing LLM Instruction Following via Dual-Agent Prompt RefinementabstractPrompt engineering is crucial for optimizing the performance of Large Language Models (LLMs), yet it remains a manual and resource-intensive process that requires multiple iterations of trial and error. Current automated prompt enhancement approaches face key challenges in preserving component relationships, managing computational requirements, and maintaining optimization traceability. This paper introduces CARE (Comprehensive Analyzer & REfiner), an LLM-based dual-agent framework that models prompt enhancement as a staged transformation pipeline with explicit validation constraints. CARE tackles three fundamental challenges: prompt decomposition with interleaved dependencies, component interference in LLM processing, and semantic drift during refinement. The framework enables reliable prompt enhancement in a single iteration through systematic component extraction and rule-based transformations. Evaluation using established benchmarks demonstrates consistent improvements across diverse LLM architectures. Nardine Basta, Benjamin Zi Hao Zhao, Muhammad Ikram 0001, Mohamed Ali Kâafar |
ECAI | 4 |
| 2025 | Bot Wars Evolved: Orchestrating Competing LLMs in a Counterstrike Against Phone Scams
Nardine Basta, Conor Atkins, Mohamed Ali Kâafar |
PAKDD (7) | 3 |
| 2025 | Can Self Supervision Rejuvenate Similarity-Based Link Prediction?
Chenhan Zhang, Weiqi Wang 0003, Zhiyi Tian, James Jian Qiao Yu, Mohamed Ali Kâafar, An Liu 0002, Shui Yu 0001 |
PAKDD (7) | 5 |
| 2025 | Deception Meets Diagnostics: Deception-based Real-Time Threat Detection in Healthcare Web SystemsabstractIncreased cloud adoption in healthcare has amplified ransomware and malware threats, accounting for $19 \%$ of global breaches in 2024. Despite this surge, the behavior of attackers exploiting healthcare systems remains under-explored in academic literature. This paper bridges that gap by deploying a scalable and stealthy deception network specifically designed for healthcare environments. The network comprises 30 real-world vulnerable healthcare web applications, mimicking domainspecific workflows across multi-cloud infrastructures, such as patient registration and billing. We leveraged ATTACK-BERT to generate semantic embeddings and applied co-regularized spectral clustering with normalized cuts to analyze multi-protocol attack traffic. Our analysis revealed nuanced attacker behaviors, including regional and protocol-specific variations, exploitation of healthcare protocols like HL7, and the use of encryption to bypass detection. A comparative sub-study further showed that attackers deliberately engage with vulnerable systems, highlighting the strategic value of deception-based defenses. By focusing on behavioral insights within healthcare-specific settings, this work lays the groundwork for integrating deception into the broader security posture of critical infrastructures. Zeeshan Zulkifl Shah, Muhammad Ikram 0001, Hassan Jameel Asghar, Mohamed Ali Kâafar |
RAID | 4 |
| 2025 | Facing the Challenge of Leveraging Untrained Humans in Malware Analysis
Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Muhammad Ikram 0001, Mohamed Ali Kâafar, Sean Lamont, Daniel Coscia |
SEC (1) | 4 |
| 2025 | Practical, Private Assurance of the Value of Collaboration via Fully Homomorphic EncryptionabstractTwo parties wish to collaborate on their datasets. However, before they reveal their datasets to each other, the parties want to have the guarantee that the collaboration would be fruitful. We look at this problem from the point of view of machine learning, where one party is promised an improvement on its prediction model by incorporating data from the other party. The parties would only wish to collaborate further if the updated model shows an improvement in accuracy. Before this is ascertained, the two parties would not want to disclose their models and datasets. In this work, we construct an interactive protocol for this problem based on the fully homomorphic encryption scheme over the Torus (TFHE) and label differential privacy, where the underlying machine learning model is a neural network. Label differential privacy is used to ensure that computations are not done entirely in the encrypted domain, which is a significant bottleneck for neural network training according to the current state-of-the-art FHE implementations. We formally prove the security of our scheme assuming honest-but-curious parties, but where one party may not have any expertise in labelling its initial dataset. Experiments show that we can obtain the output, i.e., the accuracy of the updated model, with time many orders of magnitude faster than a protocol using entirely FHE operations. Hassan Jameel Asghar, Zhigang Lu 0001, Zhongrui Zhao, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 4 |
| 2025 | Measuring, Characterizing, and Analyzing the Free Web Games EcosystemabstractWeb games, that are directly playable within web browsers, have recently garnered substantial popularity, particularly among younger demographics. The absence of a paywall for these games has raised concerns regarding the potential privacy-compromising monetization strategies. Comprehensive investigations have been carried out into domains like paid online games, multiplayer online video games, mobile gaming, and their associated privacy and security concerns, but a significant gap persists in the characterization and examination of freely accessible web games. To address these voids, our research conducts an exhaustive analysis of the web games ecosystem. We rigorously scrutinize 22 distinct web game websites with the goal of understanding player experience and addressing privacy concerns. Our methodology involves simulating player interactions across approximately 100,000 individual web games to extract insights into player behavior and privacy risks. The outcomes of our work demonstrate substantial insights into the popularity, ownership, geolocation, and game environment, including metadata diversity. It also reveals the privacy risks faced by users on these websites, encompassing various aspects, such as the presence of questionable third-party advertisements designed for revenue generation, sporadic instances of objectionable content, and the persistence of tracking mechanisms like persistent cookies. Most of all, there is also a discouraging absence of transparent privacy policy statements, a website's privacy policy statement is a legal document to protect users' rights. Also, this deficiency of clear policy information hinders users' capacity to make informed decisions about opting out and understanding the potential consequences. In short, the insights from our study highlight substantial concerns regarding prevalent privacy practices within the domain of free web game websites. Moreover, we contribute our dataset and code, which contain metadata collected from approximately 100 K web games originating from the 22 websites examined in our investigation, thereby we provide a valuable resource for further research and analysis of this less investigated genre of websites. Hina Qayyum, Muhammad Ikram 0001, Mohamed Ali Kâafar, Gareth Tyson |
IEEE Trans. Games | 3 |
| 2025 | VEH-Attack: Stealthy Tracking of Train Passengers With Side-Channel Attack on Vibration Energy Harvesting WearablesabstractVibration energy harvesting (VEH) has emerged as a viable option for mobile devices that serves the dual purpose of generating power and sensing ambient vibrations. This paper highlights the location privacy leakage resulting from unrestricted access to seemingly innocuous VEH data on mobile devices. We present VEH-Attack, a side-channel attack that exploits an inference model and VEH data patterns generated from train vibrations, enabling precise tracking of train passengers. VEH-Attack achieves an accuracy of 97% and 83.13% for VEH derived data and actual VEH data, respectively, for trip length of 6 stations with the accuracy reaching 100% for longer trip lengths. Marzieh Jalal Abadi, Sara Khalifa, Mahbub Hassan, Salil S. Kanhere, Mohamed Ali Kâafar |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Privacy Preserving Release of Mobile Sensor DataabstractSensors embedded in mobile smart devices can monitor users’ activity with high accuracy to provide a variety of services to end-users ranging from precise geolocation, health monitoring, and handwritten word recognition. However, this involves the risk of accessing and potentially disclosing sensitive information of individuals to the apps that may lead to privacy breaches. In this paper, we aim to minimize privacy leakages that may lead to user identification on mobile devices through user tracking and distinguishability while preserving the functionality of apps and services. We propose a privacy-preserving mechanism that effectively handles the sensor data fluctuations (e.g., inconsistent sensor readings while walking, sitting, and running at different times) by formulating the data as time-series modeling and forecasting. The proposed mechanism uses correlated noise-series against noise filtering attacks from an adversary, which aims to filter out the noise from the perturbed data to re-identify the original data. Unlike existing solutions, our mechanism keeps running in isolation without the interaction of a user or a service provider. We perform rigorous experiments on three benchmark datasets and show that our proposed mechanism limits user tracking and distinguishability threats to a significant extent compared to the original data while maintaining a reasonable level of utility of functionalities. In general, we show that our obfuscation mechanism reduces the user trackability threat by 60% across all the datasets while maintaining the utility loss below 0.3 Mean Absolute Error (MAE). More specifically, we observe that 80% of users achieve a 100% untrackability rate in the Swipes dataset across all noise scales. In the handwriting dataset, distinguishability is 17% for 60% of the users. Overall, our mechanism provides a utility error (MAE) of only 0.12 for 60% of users, and this increases to 0.2 for 100% users when correction thresholds are altered. Rahat Masood, Wing Yan Cheng, Dinusha Vatsalan, Deepak Mishra 0001, Hassan Jameel Asghar, Mohamed Ali Kâafar |
ARES | 6 |
| 2024 | ConvoCache: Smart Re-Use of Chatbot Responses
Conor Atkins, Ian D. Wood, Mohamed Ali Kâafar, Hassan Jameel Asghar, Nardine Basta, Michal Kepkowski |
INTERSPEECH | 3 |
| 2024 | Auditing and Attributing Behaviours of Suspicious Android Health Applications
I Wayan Budi Sentana, Muhammad Ikram 0001, Mohamed Ali Kâafar |
NSS | 4 |
| 2024 | Performance Evaluation of Quantum-Secure Symmetric Key AgreementabstractQuantum-safe public key exchange protocols face significant challenges in hardware- and software-based approaches. Quantum key distribution, which relies on specialized quantum hardware, presents a significant barrier to widespread adoption due to its high cost and limited scalability. Conversely, software-based solutions using post-quantum algorithms introduce complications, such as increased resource demands and larger cipher-texts. Furthermore, the security of these post-quantum algorithms remains relatively untested, which has led to the emerging trend of hybrid deployment, combining clas-sical and quantum-resistant techniques to hedge against potential vulnerabilities. Recently, Arqit proposed a quantum-secure symmet-ric key agreement (SKA) protocol, claiming that it is lightweight and scalable [1] to address these problems. However, their proprietary solution is not available for independent analysis. To evaluate the performance and scalability of quantum-secure SKA techniques, we develop variations of the SKA protocol using open-source and accessible components in this work. To analyze quantum-secure SKA scheme, we imple-mented an SKA technique that involves a hybrid mech-anism, leveraging secret strings distributed through a combination of existing classical and quantum public key pairs during the initial key exchange. We analyze our scheme and demonstrate that it incurs minimal performance overhead, with only 99ms for purely quantum SKA and 199ms for the hybrid version, compared to the classical SKA protocol. We also show that our scheme remains robust under various network conditions, including delays, packet losses, and bandwidth variations, maintaining small and consistent overheads. We also show that this solution is scalable, with an overhead of only one second for every additional five concurrent users. This performance improves significantly with increased computational resources-achieving a 50-60% improvement when scaling from two to four CPUs. Additionally, our security evaluations confirm that the protocol provides consistent and sufficient randomness throughout the key agreement process, ensuring quantum-resistance at every stage. Amin Rois Sinung Nugroho, Muhammad Ikram 0001, Mohamed Ali Kâafar |
SIN | 3 |
| 2024 | More Than Just a Random Number Generator! Unveiling the Security and Privacy Risks of Mobile OTP Authenticator Apps
Muhammad Ikram 0001, I Wayan Budi Sentana, Hassan Jameel Asghar, Mohamed Ali Kâafar, Michal Kepkowski |
WISE (5) | 4 |
| 2024 | On Adversarial Training with Incorrect Labels
Benjamin Zi Hao Zhao, Junda Lu 0001, Xiaowei Zhou 0003, Dinusha Vatsalan, Muhammad Ikram 0001, Mohamed Ali Kâafar |
WISE (4) | 6 |
| 2024 | DRTP: A generic Differentiated Reliable Transport Protocol
Yongmao Ren, Anmin Xu, Yifang Qin, Qinghua Wu 0004, Mohamed Ali Kâafar, Gaogang Xie |
Comput. Commun. | 9 |
| 2024 | SPGNN-API: A Transferable Graph Neural Network for Attack Paths Identification and Autonomous MitigationabstractAttack paths are the potential chain of malicious activities an attacker performs to compromise network assets and acquire privileges through exploiting network vulnerabilities. Attack path analysis helps organizations to identify new/unknown chains of attack vectors exposing critical assets, as opposed to individual attack vectors in signature-based attack analysis. Timely identification of attack paths enables proactive mitigation of threats. Nevertheless, manual analysis of complex network configurations, vulnerabilities, and security events to identify attack paths is rarely feasible. This work proposes a novel transferable graph neural network-based model for shortest path identification. The shortest path, integrated with a novel holistic model for identifying potential network vulnerabilities interactions, is then utilized to detect network attack paths. Our framework automates the risk assessment of attack paths indicating the propensity of the paths to enable the compromise of highly-critical assets (e.g., databases). The proposed framework, named SPGNN-API, incorporates automated threat mitigation through a proactive timely tuning of the network firewall rules and Zero-Trust (ZT) policies to break critical attack paths and bolster cyber defenses. Our evaluation process is twofold; evaluating the performance of the shortest path identification and assessing the attack path detection accuracy. Our results show that SPGNN-API largely outperforms the baseline model for shortest path identification with an average accuracy$\geq95$% and successfully detects 100% of the potentially compromised assets, outperforming the attack graph baseline by 47%. Houssem Jmal, Firas Ben Hmida, Nardine Basta, Muhammad Ikram 0001, Mohamed Ali Kâafar, Andy Walker |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Those Aren't Your Memories, They're Somebody Else's: Seeding Misinformation in Chat Bot Memories
Conor Atkins, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Ian D. Wood, Mohamed Ali Kâafar |
ACNS (1) | 5 |
| 2023 | An Empirical Analysis of Security and Privacy Risks in Android Cryptocurrency Wallet Apps
I Wayan Budi Sentana, Muhammad Ikram 0001, Mohamed Ali Kâafar |
ACNS | 3 |
| 2023 | Privacy-Preserving Record Linkage for Cardinality CountingabstractSeveral applications require counting the number of distinct items in the data, which is known as the cardinality counting problem. Example applications include health applications such as rare disease patients counting for adequate awareness and funding, and counting the number of cases of a new disease for outbreak detection, marketing applications such as counting the visibility reached for a new product, and cybersecurity applications such as tracking the number of unique views of social media posts. The data needed for the counting is however often personal and sensitive, and need to be processed using privacy-preserving techniques. The quality of data in different databases, for example typos, errors and variations, poses additional challenges for accurate cardinality estimation. While privacy-preserving cardinality counting has gained much attention in the recent times and a few privacy-preserving algorithms have been developed for cardinality estimation, no work has so far been done on privacy-preserving cardinality counting using record linkage techniques with fuzzy matching and provable privacy guarantees. We propose a novel privacy-preserving record linkage algorithm using unsupervised clustering techniques to link and count the cardinality of individuals in multiple datasets without compromising their privacy or identity. In addition, existing Elbow methods to find the optimal number of clusters as the cardinality are far from accurate as they do not take into account the purity and completeness of generated clusters. We propose a novel method to find the optimal number of clusters in unsupervised learning. Our experimental results on real and synthetic datasets are highly promising in terms of significantly smaller error rate of less than 0.1 with a privacy budget ϵ = 1.0 compared to the state-of-the-art fuzzy matching and clustering method. Nan Wu 0013, Dinusha Vatsalan, Mohamed Ali Kâafar, Sanath Kumar Ramesh |
AsiaCCS | 3 |
| 2023 | Exploring the Distinctive Tweeting Patterns of Toxic Twitter UsersabstractIn the pursuit of bolstering user safety, social media platforms deploy active moderation strategies, including content removal and user suspension. These measures target users engaged in discussions marked by hate speech or toxicity, often linked to specific keywords or hashtags. Nonetheless, the increasing prevalence of toxicity indicates that certain users adeptly circumvent these measures.This study examines consistently toxic users on Twitter (rebranded as X) Rather than relying on traditional methods based on specific topics or hashtags, we employ a novel approach based on patterns of toxic tweets, yielding deeper insights into their behavior.We analyzed 38 million tweets from the timelines of 12,148 Twitter users and identified the top 1,457 users who consistently exhibit toxic behavior, relying on metrics like the Gini index and Toxicity score. By comparing their posting patterns to those of non-consistently toxic users, we have uncovered distinctive temporal patterns, including contiguous activity spans, inter-tweet intervals (referred to as “Burstiness”), and churn analysis. These findings provide strong evidence for the existence of a unique tweeting pattern associated with toxic behavior on Twitter.Crucially, our methodology transcends Twitter and can be adapted to various social media platforms, facilitating the identification of consistently toxic users based on their posting behavior. This research contributes to ongoing efforts to combat online toxicity and offers insights for refining moderation strategies in the digital realm. We are committed to open research and will provide our code and data to the research community. Hina Qayyum, Muhammad Ikram 0001, Benjamin Zi Hao Zhao, Ian D. Wood, Nicolas Kourtellis, Mohamed Ali Kâafar |
IEEE Big Data | 6 |
| 2023 | On mission Twitter Profiles: A Study of Selective Toxic BehaviorabstractThe argument for persistent social media influence campaigns, often funded by malicious entities, is gaining traction. These entities utilize instrumented profiles to disseminate divisive content and disinformation, shaping public perception. Despite ample evidence of these instrumented profiles, few identification methods exist to locate them in the wild. To evade detection and appear genuine, small clusters of instrumented profiles engage in unrelated discussions, diverting attention from their true goals [34]. This strategic thematic diversity conceals their selective polarity towards certain topics and fosters public trust [49]. This study aims to characterize profiles potentially used for influence operations, termed “on-mission profiles,” relying solely on thematic content diversity within unlabeled data. Distinguishing this work is its focus on content volume and toxicity towards specific themes. Longitudinal data from 138K Twitter (rebranded as X) profiles and 293M tweets enables profiling based on theme diversity. High thematic diversity groups predominantly produce toxic content concerning specific themes, like politics, health, and news—classifying them as “on-mission” profiles. Using the identified on-mission” profiles, we design a classifier for unseen, unlabeled data. Employing a linear SVM model, we train and test it on an 80/20% split of the most diverse profiles. The classifier achieves a flawless 100% accuracy, facilitating the discovery of previously unknown “on-mission” profiles in the wild. Hina Qayyum, Muhammad Ikram 0001, Benjamin Zi Hao Zhao, Ian D. Wood, Nicolas Kourtellis, Mohamed Ali Kâafar |
IEEE Big Data | 6 |
| 2023 | Fast IDentity Online with Anonymous Credentials (FIDO-AC)
Wei-Zhu Yeoh, Michal Kepkowski, Gunnar Heide, Mohamed Ali Kâafar, Lucjan Hanzlik |
USENIX Security Symposium | 4 |
| 2023 | Unintended Memorization and Timing Attacks in Named Entity Recognition ModelsabstractNamed entity recognition models (NER), are widely used for identifying named entities (e.g., individuals, locations, and other information) in text documents. Machine learning based NER models are increasingly being applied in privacy-sensitive applications that need automatic and scalable identification of sensitive information to redact text for data sharing. In this paper, we study the setting when NER models are available as a black-box service for identifying sensitive information in user documents and show that these models are vulnerable to membership inference on their training datasets. With updated pre-trained NER models from spaCy, we demonstrate two distinct membership attacks on these models. Our first attack capitalizes on unintended memorization in the NER's underlying neural network, a phenomenon NNs are known to be vulnerable to. Our second attack leverages a timing side-channel to target NER models that maintain vocabularies constructed from the training data. We show that different functional paths of words within the training dataset in contrast to words not previously seen have measurable differences in execution time. Revealing membership status of training samples has clear privacy implications. For example, in text redaction, sensitive words or phrases to be found and removed, are at risk of being detected in the training dataset. Our experimental evaluation includes the redaction of both password and health data, presenting both security risks and a privacy/regulatory issues. This is exacerbated by results that indicate memorization after only a single phrase. We achieved a 70% AUC in our first attack on a text redaction use-case. We also show overwhelming success in the second timing attack with an 99.23% AUC. Finally we discuss potential mitigation approaches to realize the safe use of NER models in light of the presented privacy and security implications of membership inference attacks. Rana Salal Ali, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Tham Nguyen, Ian D. Wood, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 6 |
| 2023 | Local Differentially Private Fuzzy Counting in Stream Data Using Probabilistic Data StructuresabstractPrivacy-preserving estimation of counts of items in streaming data finds applications in several real-world scenarios including word auto-correction and traffic management applications. Recent works of RAPPOR [1] and Apple's count-mean sketch (CMS) algorithm [2] propose privacy preserving mechanisms for count estimation in large volumes of data using probabilistic data structures like counting Bloom filter and CMS. However, these existing methods fall short in providing a sound solution for real-time streaming data applications. Since the size of the data structure in these methods is not adaptive to the volume of the streaming data, the utility (accuracy of the count estimate) can suffer over time due to increased false positive rates. Further, the lookup operation needs to be highly efficient to answer count estimate queries in real-time. More importantly, the local Differential privacy mechanisms used in these approaches to provide privacy guarantees come at a large cost to utility (impacting the accuracy of count estimation). In this work, we propose a novel (local) Differentially private mechanism that provides high utility for the streaming data count estimation problem with similar or even lower privacy budgets while providing: a) fuzzy counting to report counts of related or similar items (for instance to account for typing errors and data variations), and b) improved querying efficiency to reduce the response time for real-time querying of counts. Our algorithm uses a combination of two probabilistic data structures Cuckoo filter and Bloom filter. We provide formal proofs for privacy and utility guarantees and present extensive experimental evaluation of our algorithm using real and synthetic English words datasets for both the exact and fuzzy counting scenarios. Our privacy preserving mechanism substantially outperforms the prior work in terms of lower querying time, significantly higher utility (accuracy of count estimation) under similar or lower privacy guarantees, at the cost of communication overhead. Dinusha Vatsalan, Raghav Bhaskar, Mohamed Ali Kâafar |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Towards a Zero-Trust Micro-segmentation Network Security Strategy: An Evaluation FrameworkabstractMicro-segmentation is an emerging security technique that separates physical networks into isolated logical micro-segments (workloads). By tying fine-grained security policies to individual workloads, it limits the attacker’s ability to move laterally through the network, even after infiltrating the perimeter defences. While micro-segmentation is proved to be effective for shrinking enterprise networks attack surface, its impact assessment is almost absent in the literature. This research is dedicated to developing an analytical framework to characterise and quantify the effectiveness of micro-segmentation on enhancing networks security. We rely on a twofold graph-feature-based framework of the network connectivity and attack graphs to evaluate the network exposure and robustness, respectively. Tracking the variations of formulated metrics values post the deployment of micro-segmentation reveals exposure reduction and robustness improvement in the range of 60% – 90%. Nardine Basta, Muhammad Ikram 0001, Mohamed Ali Kâafar, Andy Walker |
NOMS | 3 |
| 2022 | A First Look at Android Apps' Third-Party Resources Loading
Hina Qayyum, I Wayan Budi Sentana, Giang L. D. Nguyen, Muhammad Ikram 0001, Gareth Tyson, Mohamed Ali Kâafar |
NSS | 7 |
| 2022 | How Not to Handle Keys: Timing Attacks on FIDO Authenticator PrivacyabstractThis paper presents a timing attack on the FIDO2 (Fast IDentity Online) authentication protocol that allows attackers to link user accounts stored in vulnerable authenticators, a serious privacy concern. FIDO2 is a new standard specified by the FIDO industry alliance for secure token online authentication. It complements the W3C WebAuthn specification by providing means to use a USB token or other authenticator (which holds the secret authenticating material and implements FIDO protocols) as a second factor during the authentication process. From a cryptographic perspective, the protocol is a simple challenge-response where the elliptic curve digital signature algorithm is used to sign challenges. To protect the privacy of the user the token uses unique key pairs per service. To accommodate for small memory, tokens use various techniques that make use of a special parameter called a key handle sent by the service to the token with which the token can securely produce an authentication key (through generation or decryption). We identify and analyse a vulnerability in the way the processing of key handles is implemented that allows attackers to remotely link user accounts on multiple services. We show that for vulnerable authenticators there is a difference between the time it takes to process a key handle for a different service but correct authenticator, and for a different authenticator but correct service. This difference can be used to perform a timing attack allowing an adversary to link user’s accounts across services. We present several real world examples of adversaries that are in a position to execute our attack and can benefit from linking accounts. We found that two of the eight hardware authenticators we tested were vulnerable despite FIDO level 1 certification, indicating a not insignificant problem. This vulnerability cannot be easily mitigated on authenticators because, for security reasons, they usually do not allow firmware updates. In addition, we show that due to the way existing browsers implement the WebAuthn standard, the attack can be executed remotely. However, we discuss countermeasures that can be implemented by browser providers to mitigate the remote form of the attack. Michal Kepkowski, Lucjan Hanzlik, Ian D. Wood, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 4 |
| 2022 | DDCA: A Distortion Drift-Based Cost Assignment Method for Adaptive Video Steganography in the Transform DomainabstractCost assignment plays a key role in coding performance and security of video steganography. Existing cost assignment methods (for adaptive video steganography) are designed for specific transform coefficients rather than all transform coefficients. In addition, existing video steganographic frameworks do not allow Syndrome-Trellis Codes (STCs) to modify all transform coefficients in both intra-coded and inter-coded frames at the same time. To address these limitations, in this article, we first propose a novel video steganographic framework. Then, we give a theoretical analysis of distortion drift in both intra- and inter-coding procedures. Based on the analysis, we design a Distortion Drift-Based Cost Assignment method, hereafter referred to as DDCA. DDCA considers the inner-block, inter-block and inter-frame distortion costs in order to improve the coding performance and the security of stego videos when the embedding payload is fixed. We conducted extensive experiments using two video datasets to evaluate the proposed video steganographic framework and DDCA, in terms of the coding performance and the security. Our experiments show that the proposed framework outperforms three recent state-of-the-art methods, for example the coding performance and the security of stego videos can benefit from DDCA by making full use of all nonzero transform coefficients. Yi Chen 0008, Hongxia Wang 0001, Kim-Kwang Raymond Choo, Peisong He, Zoran A. Salcic, Mohamed Ali Kâafar, Xuyun Zhang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2022 | A Differentially Private Framework for Deep Learning With Convexified Loss FunctionsabstractDifferential privacy (DP) has been applied in deep learning for preserving privacy of the underlying training sets. Existing DP practice falls into three categories—objective perturbation (injecting DP noise into the objective function), gradient perturbation (injecting DP noise into the process of gradient descent) and output perturbation (injecting DP noise into the trained neural networks, scaled by the global sensitivity of the trained model parameters). They suffer from three main problems. First, conditions on objective functions limit objective perturbation in general deep learning tasks. Second, gradient perturbation does not achieve a satisfactory privacy-utility trade-off due to over-injected noise in each epoch. Third, high utility of the output perturbation method is not guaranteed because of the loose upper bound on the global sensitivity of the trained model parameters as the noise scale parameter. To address these problems, we analyse a tighter upper bound on the global sensitivity of the model parameters. Under a black-box setting, based on this global sensitivity, to control the overall noise injection, we propose a novel output perturbation framework by injecting DP noise into a randomly sampled neuron (via the exponential mechanism) at the output layer of a baseline non-private neural network trained with a convexified loss function. We empirically compare the privacy-utility trade-off, measured by accuracy loss to baseline non-private models and the privacy leakage against black-box membership inference (MI) attacks, between our framework and the open-source differentially private stochastic gradient descent (DP-SGD) approaches on six commonly used real-world datasets. The experimental evaluations show that, when the baseline models have observable privacy leakage under MI attacks, our framework achieves a better privacy-utility trade-off than existing DP-SGD implementations, given an overall privacy budget$\epsilon \leq 1$for a large number of queries. Zhigang Lu 0001, Hassan Jameel Asghar, Mohamed Ali Kâafar, Darren Webb, Peter Dickinson |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Fairness and Cost Constrained Privacy-Aware Record LinkageabstractRecord linkage algorithms match and link records from different databases that refer to the same real-world entity based on direct and/or quasi-identifiers, such as name, address, age, and gender, available in the records. Since these identifiers generally contain personal identifiable information (PII) about the entities, record linkage algorithms need to be developed with privacy constraints. Known as privacy-preserving record linkage (PPRL), many research studies have been conducted to perform the linkage on encoded and/or encrypted identifiers. Differential privacy (DP) combined with computationally efficient encoding methods, e.g. Bloom filter encoding, has been used to develop PPRL with provable privacy guarantees. The standard DP notion does not however address other constraints, among which the most important ones are fairness-bias and cost of linkage in terms of number of record pairs to be compared. In this work, we propose new notions of fairness-constrained DP and fairness and cost-constrained DP for PPRL and develop a framework for PPRL with these new notions of DP combined with Bloom filter encoding. We provide theoretical proofs for the new DP notions for fairness and cost-constrained PPRL and experimentally evaluate them on two datasets containing person-specific data. Our experimental results show that with these new notions of DP, PPRL with better performance (compared to the standard DP notion for PPRL) can be achieved with regard to privacy, cost and fairness constraints. Nan Wu 0013, Dinusha Vatsalan, Sunny Verma, Mohamed Ali Kâafar |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Do Auto-Regressive Models Protect Privacy? Inferring Fine-Grained Energy Consumption From Aggregated Model ParametersabstractWe investigate the extent to which statistical predictive models leak information about their training data. More specifically, based on the use case of household (electrical) energy consumption, we evaluate whether white-box access to auto-regressive (AR) models trained on such data together with background information, such as household energy data aggregates (e.g., monthly billing information) and publicly-available weather data, can lead to inferring fine-grained energy data of any particular household. We construct two adversarial models aiming to infer fine-grained energy consumption patterns. Both threat models use monthly billing information of target households. The second adversary has access to the AR model for a cluster of households containing the target household. Using two real-world energy datasets, we demonstrate that this adversary can apply maximuma posterioriestimation to reconstruct daily consumption of target households with significantly lower error than the first adversary, which serves as a baseline. Such fine-grained data can essentially expose private information, such as occupancy levels. Finally, we use differential privacy (DP) to alleviate the privacy concerns of the adversary in dis-aggregating energy data. Our evaluations show that differentially private model parameters offer strong privacy protection against the adversary with moderate utility, captured in terms of model fitness to the cluster. Nazim Uddin Sheikh, Hassan Jameel Asghar, Farhad Farokhi, Mohamed Ali Kâafar |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | On the (In)Feasibility of Attribute Inference Attacks on Machine Learning ModelsabstractWith an increase in low-cost machine learning APIs, advanced machine learning models may be trained on private datasets and monetized by providing them as a service. However, privacy researchers have demonstrated that these models may leak information about records in the training dataset via membership inference attacks. In this paper, we take a closer look at another inference attack reported in literature, called attribute inference, whereby an attacker tries to infer missing attributes of a partially known record used in the training dataset by accessing the machine learning model as an API. We show that even if a classification model succumbs to membership inference attacks, it is unlikely to be susceptible to attribute inference attacks. We demonstrate that this is because membership inference attacks fail to distinguish a member from a nearby non-member. We call the ability of an attacker to distinguish the two (similar) vectors as strong membership inference. We show that membership inference attacks cannot infer membership in this strong setting, and hence inferring attributes is infeasible. However, under a relaxed notion of attribute inference, called approximate attribute inference, we show that it is possible to infer attributes close to the true attributes. We verify our results on three publicly available datasets, five membership, and three attribute inference attacks reported in literature. Benjamin Zi Hao Zhao, Aviral Agrawal, Catisha Coburn, Hassan Jameel Asghar, Raghav Bhaskar, Mohamed Ali Kâafar, Darren Webb, Peter Dickinson |
EuroS&P | 6 |
| 2021 | BlockJack: Towards Improved Prevention of IP Prefix Hijacking Attacks in Inter-domain Routing via Blockchain
I Wayan Budi Sentana, Muhammad Ikram 0001, Mohamed Ali Kâafar |
SECRYPT | 3 |
| 2021 | Empirical Security and Privacy Analysis of Mobile Symptom Checking Apps on Google PlayabstractSmartphone technology has drastically improved over the past decade. These improvements have seen the creation of specialized health applications, which offer consumers a range of health-related activities such as tracking and checking symptoms of health conditions or diseases through their smartphones. We term these applications as Symptom Checking apps or simply SymptomCheckers. Due to the sensitive nature of the private data they collect, store and manage, leakage of user information could result in significant consequences. In this paper, we use a combination of techniques from both static and dynamic analysis to detect, trace and categorize security and privacy issues in 36 popular SymptomCheckers on Google Play. Our analyses reveal that SymptomCheckers request a significantly higher number of sensitive permissions and embed a higher number of third-party tracking libraries for targeted advertisements and analytics exploiting the privileged access of the SymptomCheckers in which they exist, as a mean of collecting and sharing critically sensitive data about the user and their device. We find that these are sharing the data that they collect through unencrypted plain text to the third-party advertisers and, in some cases, to malicious domains. The results reveal that the exploitation of SymptomCheckers is present in popular apps, still readily available on Google Play. I Wayan Budi Sentana, Muhammad Ikram 0001, Mohamed Ali Kâafar, Shlomo Berkovsky |
SECRYPT | 3 |
| 2021 | Trace Recovery: Inferring Fine-grained Trace of Energy Data from AggregatesabstractSmart meter data is collected and shared with different stakeholders involved in a smart grid ecosystem. The fine-grained energy data is extremely useful for grid operations and maintenance, monitoring and for market segmentation purposes. However, sharing and releasing fine-grained energy data induces explicit violations of private information of consumers (Molina-Markham et al., 2010). Service providers do then share and release aggregated statistics to preserve the privacy of consumers with data aggregation aiming at reducing the risks of individual consumption traces being revealed. In this paper, we show that an adversary can reconstruct individual traces of energy data by exploiting consistency (similar consumption patterns over time) and distinctiveness (one household’s energy consumption pattern is significantly different from that of others) properties of individual consumption load patterns. We propose an unsupervised attack framework to recover hourly energy consumption ti me-series of individual users without any prior knowledge. We pose the problem of assigning aggregated energy consumption meter readings to individuals as an assignment problem and solve it by the Hungarian algorithm (Xu et al., 2017; Kuhn, 1955). Using two real-world datasets, our empirical evaluations show that an adversary is capable of recovering over 70% of households’ energy consumption patterns with over 90% accuracy. Nazim Uddin Sheikh, Zhigang Lu 0001, Hassan Jameel Asghar, Mohamed Ali Kâafar |
SECRYPT | 4 |
| 2021 | Analyzing security issues of android mobile health and medical applicationsabstractOBJECTIVE: We conduct a first large-scale analysis of mobile health (mHealth) apps available on Google Play with the goal of providing a comprehensive view of mHealth apps' security features and gauging the associated risks for mHealth users and their data. MATERIALS AND METHODS: We designed an app collection platform that discovered and downloaded more than 20 000 mHealth apps from the Medical and Health & Fitness categories on Google Play. We performed a suite of app code and traffic measurements to highlight a range of app security flaws: certificate security, sensitive or unnecessary permission requests, malware presence, communication security, and security-related concerns raised in user reviews. RESULTS: Compared to baseline non-mHealth apps, mHealth apps generally adopt more reliable signing mechanisms and request fewer dangerous permissions. However, significant fractions of mHealth apps expose users to serious security risks. Specifically, 1.8% of mHealth apps package suspicious codes (eg, trojans), 45.0% rely on unencrypted communication, and as much as 23.0% of personal data (eg, location information and passwords) is sent on unsecured traffic. An analysis of the app reviews reveals that mHealth app users are largely unaware of the surfaced security issues. CONCLUSION: Despite being better aligned with security best practices than non-mHealth apps, mHealth apps are still far from ensuring robust security guarantees. App users, clinicians, technology developers, and policy makers alike should be cognizant of the uncovered security issues and weigh them carefully against the benefits of mHealth apps. Gioacchino Tangari, Muhammad Ikram 0001, I Wayan Budi Sentana, Kiran Ijaz, Mohamed Ali Kâafar, Shlomo Berkovsky |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | The Audio Auditor: User-Level Membership Inference in Internet of Things Voice ServicesabstractAbstract With the rapid development of deep learning techniques, the popularity of voice services implemented on various Internet of Things (IoT) devices is ever increasing. In this paper, we examine user-level membership inference in the problem space of voice services, by designing an audio auditor to verify whether a specific user had unwillingly contributed audio used to train an automatic speech recognition (ASR) model under strict black-box access. With user representation of the input audio data and their corresponding translated text, our trained auditor is effective in user-level audit. We also observe that the auditor trained on specific data can be generalized well regardless of the ASR model architecture. We validate the auditor on ASR models trained with LSTM, RNNs, and GRU algorithms on two state-of-the-art pipelines, the hybrid ASR system and the end-to-end ASR system. Finally, we conduct a real-world trial of our auditor on iPhone Siri, achieving an overall accuracy exceeding 80%. We hope the methodology developed in this paper and findings can inform privacy advocates to overhaul IoT privacy. Yuantian Miao, Minhui Xue 0001, Chao Chen 0015, Lei Pan 0002, Jun Zhang 0010, Benjamin Zi Hao Zhao, Mohamed Ali Kâafar, Yang Xiang 0001 |
Proc. Priv. Enhancing Technol. | 7 |
| 2021 | The Cost of Privacy in Asynchronous Differentially-Private Machine Learning
Farhad Farokhi, Nan Wu 0013, David B. Smith 0001, Mohamed Ali Kâafar |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | On the Resilience of Biometric Authentication Systems against Random Inputs
Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Mohamed Ali Kâafar |
NDSS | 3 |
| 2020 | The Value of Collaboration in Convex Machine Learning with Differential PrivacyabstractIn this paper, we apply machine learning to distributed private data owned by multiple data owners, entities with access to non-overlapping training datasets. We use noisy, differentially-private gradients to minimize the fitness cost of the machine learning model using stochastic gradient descent. We quantify the quality of the trained model, using the fitness cost, as a function of privacy budget and size of the distributed datasets to capture the trade-off between privacy and utility in machine learning. This way, we can predict the outcome of collaboration among privacy-aware data owners prior to executing potentially computationally-expensive machine learning algorithms. Particularly, we show that the difference between the fitness of the trained machine learning model using differentially-private gradient queries and the fitness of the trained machine model in the absence of any privacy concerns is inversely proportional to the size of the training datasets squared and the privacy budget squared. We successfully validate the performance prediction with the actual performance of the proposed privacy-aware learning algorithms, applied to: financial datasets for determining interest rates of loans using regression; and detecting credit card frauds using support vector machines. Nan Wu 0013, Farhad Farokhi, David B. Smith 0001, Mohamed Ali Kâafar |
SP | 4 |
| 2020 | Averaging Attacks on Bounded Noise-based Disclosure Control AlgorithmsabstractWe describe and evaluate an attack that reconstructs the histogram of any target attribute of a sensitive dataset which can only be queried through a specific class of real-world privacy-preserving algorithms which we call bounded perturbation algorithms. A defining property of such an algorithm is that it perturbs answers to the queries by adding zero-mean noise distributed within a bounded (possibly undisclosed) range. Other key properties of the algorithm include only allowing restricted queries (enforced via an online interface), suppressing answers to queries which are only satisfied by a small group of individuals (e.g., by returning a zero as an answer), and adding the same perturbation to two queries which are satisfied by the same set of individuals (to thwart differencing or averaging attacks). A real-world example of such an algorithm is the one deployed by the Australian Bureau of Statistics’ (ABS) online tool called TableBuilder, which allows users to create tables, graphs and maps of Australian census data [30]. We assume an attacker (say, a curious analyst) who is given oracle access to the algorithm via an interface. We describe two attacks on the algorithm. Both attacks are based on carefully constructing (different) queries that evaluate to the same answer. The first attack finds the hidden perturbation parameter r (if it is assumed not to be public knowledge). The second attack removes the noise to obtain the original answer of some (counting) query of choice. We also show how to use this attack to find the number of individuals in the dataset with a target attribute value a of any attribute A, and then for all attribute values ai ∈ A. None of the attacks presented here depend on any background information. Our attacks are a practical illustration of the (informal) fundamental law of information recovery which states that “overly accurate estimates of too many statistics completely destroys privacy” [9, 15]. Hassan Jameel Asghar, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 2 |
| 2020 | Not All Attributes are Created Equal: dX -Private Mechanisms for Linear QueriesabstractAbstract Differential privacy provides strong privacy guarantees simultaneously enabling useful insights from sensitive datasets. However, it provides the same level of protection for all elements (individuals and attributes) in the data. There are practical scenarios where some data attributes need more/less protection than others. In this paper, we consider dX -privacy, an instantiation of the privacy notion introduced in [6], which allows this flexibility by specifying a separate privacy budget for each pair of elements in the data domain. We describe a systematic procedure to tailor any existing differentially private mechanism that assumes a query set and a sensitivity vector as input into its d X -private variant, specifically focusing on linear queries. Our proposed meta procedure has broad applications as linear queries form the basis of a range of data analysis and machine learning algorithms, and the ability to define a more flexible privacy budget across the data domain results in improved privacy/utility tradeoff in these applications. We propose several d X -private mechanisms, and provide theoretical guarantees on the trade-off between utility and privacy. We also experimentally demonstrate the effectiveness of our procedure, by evaluating our proposed d X -private Laplace mechanism on both synthetic and real datasets using a set of randomly generated linear queries. Parameswaran Kamalaruban, Victor Perrier, Hassan Jameel Asghar, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 4 |
| 2020 | Measuring and Analysing the Chain of Implicit Trust: A Study of Third-party Resources LoadingabstractThe web is a tangled mass of interconnected services, whereby websites import a range of external resources from various third-party domains. The latter can also load further resources hosted on other domains. For each website, this creates a dependency chain underpinned by a form of implicit trust between the first-party and transitively connected third parties. The chain can only be loosely controlled as first-party websites often have little, if any, visibility on where these resources are loaded from. This article performs a large-scale study of dependency chains in the web to find that around 50% of first-party websites render content that they do not directly load. Although the majority (84.91%) of websites have short dependency chains (below three levels), we find websites with dependency chains exceeding 30. Using VirusTotal, we show that 1.2% of these third parties are classified as suspicious—although seemingly small, this limited set of suspicious third parties have remarkable reach into the wider ecosystem. We find that 73% of websites under-study load resources from suspicious third parties, and 24.8% of first-party webpages contain at least three third parties classified as suspicious in their dependency chain. By running sandboxed experiments, we observe a range of activities with the majority of suspicious JavaScript codes downloading malware. Muhammad Ikram 0001, Rahat Masood, Gareth Tyson, Mohamed Ali Kâafar, Noha Loizon, Roya Ensafi |
ACM Trans. Priv. Secur. | 4 |
| 2020 | Exploiting Behavioral Side Channels in Observation Resilient Cognitive Authentication SchemesabstractObservation Resilient Authentication Schemes (ORAS) are a class of shared secret challenge–response identification schemes where a user mentally computes the response via a cognitive function to authenticate herself such that eavesdroppers cannot readily extract the secret. Security evaluation of ORAS generally involves quantifying information leaked via observed challenge–response pairs. However, little work has evaluated information leaked via human behavior while interacting with these schemes. A common way to achieve observation resilience is by including a modulus operation in the cognitive function. This minimizes the information leaked about the secret due to the many-to-one map from the set of possible secrets to a given response. In this work, we show that user behavior can be used as a side channel to obtain the secret in such ORAS. Specifically, the user’s eye-movement patterns and associated timing information can deduce whether a modulus operation was performed (a fundamental design element) to leak information about the secret. We further show that the secret can still be retrieved if the deduction is erroneous, a more likely case in practice. We treat the vulnerability analytically and propose a generic attack algorithm that iteratively obtains the secret despite the “faulty” modulus information. We demonstrate the attack on five ORAS and show that the secret can be retrieved with considerably less challenge–response pairs than non-side-channel attacks (e.g., algebraic/statistical attacks). In particular, our attack is applicable on Mod10, a one-time-pad-based scheme, for which no non-side-channel attack exists. We field test our attack with a small-scale eye-tracking user study. Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Mohamed Ali Kâafar, Francesca Trevisan, Haiyue Yuan |
ACM Trans. Priv. Secur. | 3 |
| 2019 | A Decade of Mal-Activity Reporting: A Retrospective Analysis of Internet Malicious Activity BlacklistsabstractThis paper focuses on reporting of Internet malicious activity (or mal-activity in short) by public blacklists with the objective of providing a systematic characterization of what has been reported over the years, and more importantly, the evolution of reported activities. Using an initial seed of 22 blacklists, covering the period from January 2007 to June 2017, we collect more than 51 million mal-activity reports involving 662K unique IP addresses worldwide. Leveraging the Wayback Machine, antivirus (AV) tool reports and several additional public datasets (e.g., BGP Route Views and Internet registries) we enrich the data with historical meta-information including geo-locations (countries), autonomous system (AS) numbers and types of mal-activity. Furthermore, we use the initially labelled dataset of ~1.57 million mal-activities (obtained from public blacklists) to train a machine learning classifier to classify the remaining unlabeled dataset of ~44 million mal-activities obtained through additional sources. We make our unique collected dataset (and scripts used) publicly available for further research. The main contributions of the paper are a novel means of report collection, with a machine learning approach to classify reported activities, characterization of the dataset and, most importantly, temporal analysis of mal-activity reporting behavior. Inspired by P2P behavior modeling, our analysis shows that some classes of mal-activities (e.g., phishing) and a small number of mal-activity sources are persistent, suggesting that either blacklist-based prevention systems are ineffective or have unreasonably long update periods. Our analysis also indicates that resources can be better utilized by focusing on heavy mal-activity contributors, which constitute the bulk of mal-activities. Benjamin Zi Hao Zhao, Muhammad Ikram 0001, Hassan Jameel Asghar, Mohamed Ali Kâafar, Abdelberi Chaabane, Kanchana Thilakarathna |
AsiaCCS | 4 |
| 2019 | Private Continual Release of Real-Valued Data Streams
Victor Perrier, Hassan Jameel Asghar, Mohamed Ali Kâafar |
NDSS | 3 |
| 2019 | The Chain of Implicit Trust: An Analysis of the Web Third-party Resources LoadingabstractThe Web is a tangled mass of interconnected services, where websites import a range of external resources from various third-party domains. The latter can also load resources hosted on other domains. For each website, this creates a dependency chain underpinned by a form of implicit trust between the first-party and transitively connected third-parties. The chain can only be loosely controlled as first-party websites often have little, if any, visibility on where these resources are loaded from. This paper performs a large-scale study of dependency chains in the Web, to find that around 50% of first-party websites render content that they did not directly load. Although the majority (84.91%) of websites have short dependency chains (below 3 levels), we find websites with dependency chains exceeding 30. Using VirusTotal, we show that 1.2% of these third-parties are classified as suspicious - although seemingly small, this limited set of suspicious third-parties have remarkable reach into the wider ecosystem. Muhammad Ikram 0001, Rahat Masood, Gareth Tyson, Mohamed Ali Kâafar, Noha Loizon, Roya Ensafi |
WWW | 4 |
| 2019 | Fast privacy-preserving network function outsourcingabstractIn this paper, we present the design and implementation of SplitBox, a system for privacy-preserving processing of network functions outsourced to cloud middleboxes—i.e., without revealing the policies governing these functions. SplitBox is built to provide privacy for a generic network function that abstracts the functionality of a variety of network functions and associated policies, including firewalls, virtual LANs, network address translators (NATs), deep packet inspection , and load balancers. We present a scalable design aiming to provide high throughput and low latency, by distributing functionalities to a few virtual machines (VMs), while providing provably secure guarantees. We implement SplitBox inside FastClick, an extension of the Click modular router, using Intel’s DPDK to handle packet I/O. We evaluate our prototype experimentally to find its bottlenecks and stress-test its different components, vis-à-vis two widely used network functions, i.e., firewall and VLAN tagging. Our evaluation shows that, on commodity hardware, SplitBox can process packets close to line rate (i.e., 8.9Gbps) with up to 50 traversed policies. Hassan Jameel Asghar, Emiliano De Cristofaro, Guillaume Jourjon, Mohamed Ali Kâafar, Laurent Mathy, Luca Melis, Craig Russell, Mang Yu |
Comput. Networks | 4 |
| 2019 | Optimized deployment of drone base station to improve user experience in cellular networks
Hailong Huang 0001, Andrey V. Savkin, Ming Ding 0001, Mohamed Ali Kâafar |
J. Netw. Comput. Appl. | 4 |
| 2019 | uStash: A Novel Mobile Content Delivery System for Improving User QoE in Public TransportabstractMobile data traffic is growing exponentially and it is even more challenging to distribute content efficiently while users are “on the move” such as in public transport. The use of mobile devices for accessing content (e.g., videos) while commuting are both expensive and unreliable, although it is becoming common practice worldwide. Leveraging on the spatial and temporal correlation of content popularity and users' diverse network connectivity, we propose a novel content distribution system, uStash, which guarantees better QoE with regards to access delays and cost of usage. The proposed collaborative download and content stashing schemes provide the uStash provider the flexibility to control the cost of content access via cellular networks. We model the uStash system in a probabilistic framework and thereby analytically derive the optimal portions for collaborative downloading. Then, we validate the proposed models using real-life trace driven simulations. In particular, we use dataset from 22 inter-city buses running on six different routes and from a mobile VoD service provider to show that uStash reduces the cost of monthly cellular data by approximately 50 percent and the expected delay for content access by 60 percent compared to content downloaded via users' cellular network connections. Fangzhou Jiang, Kanchana Thilakarathna, Sirine Mrabet, Mohamed Ali Kâafar, Aruna Seneviratne |
IEEE Trans. Mob. Comput. | 4 |
| 2018 | Gargoyle: A Network-based Insider Attack Resilient Framework for OrganizationsabstractAnytime, Anywhere' data access model has become a widespread IT policy in organizations making insider attacks even more complicated to model, predict and deter. Here, we propose Gargoyle, a network-based insider attack resilient framework against the most complex insider threats within a pervasive computing context. Compared to existing solutions, Gargoyle evaluates the trustworthiness of an access request context through a new set of contextual attributes called Network Context Attribute (NCA). NCAs are extracted from the network traffic and include information such as the user's device capabilities, security-level, current and prior interactions with other devices, network connection status, and suspicious online activities. Retrieving such information from the user's device and its integrated sensors are challenging in terms of device performance overheads, sensor costs, availability, reliability and trustworthiness. To address these issues, Gargoyle leverages the capabilities of Software-Defined Network (SDN) for both policy enforcement and implementation. In fact, Gargoyle's SDN App can interact with the network controller to create a 'defence-in-depth' protection system. For instance, Gargoyle can automatically quarantine a suspicious data requestor in the enterprise network for further investigation or filter out an access request before engaging a data provider. Finally, instead of employing simplistic binary rules in access authorizations, Gargoyle incorporates Function-based Access Control (FBAC) and supports the customization of access policies into a set of functions (e.g., disabling copy, allowing print) depending on the perceived trustworthiness of the context. Our extensive evaluation results prove the practicality of Gargoyle with better performance metrics compared to existing solutions. Arash Shaghaghi, Salil S. Kanhere, Mohamed Ali Kâafar, Elisa Bertino, Sanjay K. Jha |
LCN | 3 |
| 2018 | Gwardar: Towards Protecting a Software-Defined Network from Malicious Network Operating SystemsabstractA Software-Defined Network (SDN) controller (aka. Network Operating System or NOS) is regarded as the brain of the network and is the single most critical element responsible to manage an SDN. Complimentary to existing solutions that aim to protect a NOS, we propose an intrusion protection system designed to protect an SDN against a controller that has been successfully compromised. Gwardar maintains a virtual replica of the data plane by intercepting the OpenFlow messages exchanged between the control and data plane. By observing the long-term flow of the packets, Gwardar learns the normal set of trajectories in the data plane for distinct packet headers. Upon detecting an unexpected packet trajectory, it starts by verifying the data plane forwarding devices by comparing the actual packet trajectories with the expected ones computed over the virtual replica. If the anomalous trajectories match the NOS instructions, Gwardar inspects the NOS itself. For this, it submits policies matching the normal set of trajectories and verifies whether the controller submits matching flow rules to the data plane and whether the network view provided to the application plane reflects the changes. Our evaluation results prove the practicality of Gwardar with a high detection accuracy in a reasonable time-frame. Arash Shaghaghi, Salil S. Kanhere, Mohamed Ali Kâafar, Sanjay K. Jha |
NCA | 3 |
| 2018 | Incognito: A Method for Obfuscating Web DataabstractUsers leave a trail of their personal data, interests, and intents while surfing or sharing information on the Web. Web data could therefore reveal some private/sensitive information about users based on inference analysis. The possible identification of information corresponding to a single individual by an inference attack holds true even if the user identifiers are encoded or removed in the Web data. Several works have been done on improving privacy of Web data through obfuscation methods~\citeHow09,Dom09,Sha05,Che14. However, these methods are neither comprehensive, generic to be applicable to any Web data, nor effective against adversarial attacks. To this end, we propose a privacy-aware obfuscation method for Web data addressing these identified drawbacks of existing methods. We use probabilistic methods to predict privacy risk of Web data that incorporates all key privacy aspects, which are uniqueness, uniformity, and linkability of Web data. The Web data with high predicted risk are then obfuscated by our method to minimize the privacy risk using semantically similar data. Our method is resistant against adversary who has knowledge about the datasets and model learned risk probabilities using differential privacy-based noise addition. Experimental study conducted on two real Web datasets validates the significance and efficacy of our method. Our results indicate that the average privacy risk reaches to 100% with a minimum of 10 sensitive Web entries, while at most 0% privacy risk could be attained with our obfuscation method at the cost of average utility loss of 64.3%. Rahat Masood, Dinusha Vatsalan, Muhammad Ikram 0001, Mohamed Ali Kâafar |
WWW | 4 |
| 2018 | Touch and You're Trapp(ck)ed: Quantifying the Uniqueness of Touch Gestures for TrackingabstractAbstract We argue that touch-based gestures on touch-screen devices enable the threat of a form of persistent and ubiquitous tracking which we call touch-based tracking. Touch-based tracking goes beyond the tracking of virtual identities and has the potential for cross-device tracking as well as identifying multiple users using the same device. We demonstrate the likelihood of touch-based tracking by focusing on touch gestures widely used to interact with touch devices such as swipes and taps.. Our objective is to quantify and measure the information carried by touch-based gestures which may lead to tracking users. For this purpose, we develop an information theoretic method that measures the amount of information about users leaked by gestures when modelled as feature vectors. Our methodology allows us to evaluate the information leaked by individual features of gestures, samples of gestures, as well as samples of combinations of gestures. Through our purpose-built app, called TouchTrack, we gather gesture samples from 89 users, and demonstrate that touch gestures contain sufficient information to uniquely identify and track users. Our results show that writing samples (on a touch pad) can reveal 73.7% of information (when measured in bits), and left swipes can reveal up to 68.6% of information. Combining different combinations of gestures results in higher uniqueness, with the combination of keystrokes, swipes and writing revealing up to 98.5% of information about users. We further show that, through our methodology, we can correctly re-identify returning users with a success rate of more than 90%. Rahat Masood, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 4 |
| 2018 | Access Types Effect on Internet Video Services and Its Implications on CDN CachingabstractVideo providers heavily rely on geographically distributed content distribution networks (CDNs) to place video content as close to users as possible, with an aim of improving video quality and avoiding single point of failure at the server side. The effectiveness of CDNs is mostly dependent on the content consumption patterns. Currently, video providers are offering access to content from different platforms (e.g., mobile devices and PC clients), which might result in distinct video content consumption patterns and finally affect the efficiency of CDN caching. Nevertheless, the access type effect on Internet videos is not well understood. In this paper, using a data set consisting of 26 million video requests of a large-scale commercial video-on-demand system, we study the effect of three main access types, i.e., proprietary software on PC clients, Web browser, and mobile apps. Several observations suggest that access types should be considered carefully in CDN design. In particular, the user engagement, user interests in content, and video popularity dynamics patterns, three important factors for video caching, vary remarkably in the three access types. Leveraging off our findings, we propose an access type-aware CDN caching system that associates a cache for each access type and also several optimizations, including partial caching of videos based on chunk-level caching, cross-platform read-only cache access, and prefiltering of the least popular videos. Trace-driven simulations demonstrate that the access type-aware CDN caching achieves high cache hit rate and, more importantly, greatly reduces the disk load that is measured by the number of cache replacement operations. Gaogang Xie, Zhenyu Li 0001, Mohamed Ali Kâafar, Qinghua Wu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | POSTER: TouchTrack: How Unique are your Touch Gestures?abstractThis paper studies a privacy threat induced by the collection and monitoring of a user's touch gestures on touchscreen devices. The threat is a new form of persistent tracking which we refer to as "touch-based tracking". It goes beyond tracking of virtual identities and has the potential for cross-device tracking as well as identifying multiple users using the same device. To demonstrate the likelihood of touch-based tracking, we propose an information theoretic method that quantifies the amount of information revealed by individual features of gestures, samples of gestures as well as samples of gesture combinations, when modelled as feature vectors. We have also developed a purpose-built app, named "TouchTrack" that collects data from users and informs them on how unique they are when interacting with their touch devices. Our results from 89 different users indicate that writing samples and left swipes can reveal 73.7% and 68.6% of user information, respectively. Combining different combinations of gestures results in higher uniqueness, with the combination of keystrokes, swipes and writing revealing up to 98.5% of information about users. We correctly re-identify returning users with a success rate of more than 90%. Rahat Masood, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Mohamed Ali Kâafar |
CCS | 4 |
| 2017 | WedgeTail: An Intrusion Prevention System for the Data Plane of Software Defined NetworksabstractNetworks are vulnerable to disruptions caused by malicious forwarding devices. The situation is likely to worsen in Software Defined Networks (SDNs) with the incompatibility of existing solutions, use of programmable soft switches and the potential of bringing down an entire network through compromised forwarding devices. In this paper, we present WedgeTail, an Intrusion Prevention System (IPS) designed to secure the SDN data plane. WedgeTail regards forwarding devices as points within a geometric space and stores the path packets take when traversing the network as trajectories. To be efficient, it prioritizes forwarding devices before inspection using an unsupervised trajectory-based sampling mechanism. For each of the forwarding device, WedgeTail computes the expected and actual trajectories of packets and 'hunts' for any forwarding device not processing packets as expected. Compared to related work, WedgeTail is also capable of distinguishing between malicious actions such as packet drop and generation. Moreover, WedgeTail employs a radically different methodology that enables detecting threats autonomously. In fact, it has no reliance on pre-defined rules by an administrator and may be easily imported to protect SDN networks with different setups, forwarding devices, and controllers. We have evaluated WedgeTail in simulated environments, and it has been capable of detecting and responding to all implanted malicious forwarding devices within a reasonable time-frame. We report on the design, implementation, and evaluation of WedgeTail in this manuscript. Arash Shaghaghi, Mohamed Ali Kâafar, Sanjay K. Jha |
AsiaCCS | 2 |
| 2017 | Quantifying the impact of adversarial evasion attacks on machine learning based android malware classifiersabstractWith the proliferation of Android-based devices, malicious apps have increasingly found their way to user devices. Many solutions for Android malware detection rely on machine learning; although effective, these are vulnerable to attacks from adversaries who wish to subvert these algorithms and allow malicious apps to evade detection. In this work, we present a statistical analysis of the impact of adversarial evasion attacks on various linear and non-linear classifiers, using a recently proposed Android malware classifier as a case study. We systematically explore the complete space of possible attacks varying in the adversary's knowledge about the classifier; our results show that it is possible to subvert linear classifiers (Support Vector Machines and Logistic Regression) by perturbing only a few features of malicious apps, with more knowledgeable adversaries degrading the classifier's detection rate from 100% to 0% and a completely blind adversary able to lower it to 12%. We show non-linear classifiers (Random Forest and Neural Network) to be more resilient to these attacks. We conclude our study with recommendations for designing classifiers to be more robust to the attacks presented in our work. Zainab Abaid, Mohamed Ali Kâafar, Sanjay K. Jha |
NCA | 2 |
| 2017 | A first look at mobile Ad-Blocking appsabstractOnline advertisers, third party trackers and analytics services are constantly tracking user activities as they access web services through their web browsers or mobile apps. While, web browser plugins disabling and blocking Ads (often associated tracking/analytics scripts), e.g. AdBlock Plus[3] have been well studied and are relatively well understood, an emerging new category of apps in the tracking mobile eco-system, referred as the mobile Ad-Blocking apps, received very little to no attention. With the recent significant increase of the number of mobile Ad-Blockers and the exponential growth of mobile Ad-Blocking apps' popularity, this paper aims to fill in the gap and study this new category of players in the mobile ad/tracking eco-system. This paper presents the first study of Android Ad-Blocking apps (or Ad-Blockers), analysing 97 Ad-Blocking mobile apps extracted from a corpus of more than 1.5 million Android apps on Google Play. While the main (declared) purpose of the apps is to block advertisements and mobile tracking services, our data analysis revealed the paradoxical presence of third-party tracking libraries and permissions to access sensitive resources on users' mobile devices, as well as the existence of embedded malware code within some mobile Ad-Blockers. We also analysed user reviews and found that even though a fraction of users raised concerns about the privacy and the actual performance of the mobile Ad-Blocking apps, most of the apps still attract a relatively high rating. Muhammad Ikram 0001, Mohamed Ali Kâafar |
NCA | 2 |
| 2017 | AC-PROT: An Access Control Model to Improve Software-Defined Networking SecurityabstractThe logically-centralized controllers have largely operated as the coordination points in software-defined networking(SDN), through which applications submit network operations to manage the global network resource. Therefore, the validity of these network operations from SDN applications are critical for the security of SDN. In this paper, we analyze the mechanism that generates network operations in SDN, and present a fine-grained access control model, called Access Control Protector(AC-PROT),that employs an attribute-based signature scheme for network applications. The simulation result demonstrates that AC-PROT can efficiently identify and reject unauthorized network operations generated by applications. Wei Wu 0027, Ren Ping Liu 0001, Wei Ni 0001, Mohamed Ali Kâafar, Xiaojing Huang 0001 |
VTC Spring | 4 |
| 2017 | Crowd-Cache: Leveraging on spatio-temporal correlation in content popularity for mobile networking in proximity
Kanchana Thilakarathna, Fangzhou Jiang, Sirine Mrabet, Mohamed Ali Kâafar, Aruna Seneviratne, Gaogang Xie |
Comput. Commun. | 4 |
| 2017 | A deep dive into location-based communities in social discovery networks
Kanchana Thilakarathna, Suranga Seneviratne, Mohamed Ali Kâafar, Aruna Seneviratne |
Comput. Commun. | 4 |
| 2017 | Towards Seamless Tracking-Free Web: Improved Detection of Trackers via One-class LearningabstractAbstract Numerous tools have been developed to aggressively block the execution of popular JavaScript programs in Web browsers. Such blocking also affects functionality of webpages and impairs user experience. As a consequence, many privacy preserving tools that have been developed to limit online tracking, often executed via JavaScript programs, may suffer from poor performance and limited uptake. A mechanism that can isolate JavaScript programs necessary for proper functioning of the website from tracking JavaScript programs would thus be useful. Through the use of a manually labelled dataset composed of 2,612 JavaScript programs, we show how current privacy preserving tools are ineffective in finding the right balance between blocking tracking JavaScript programs and allowing functional JavaScript code. To the best of our knowledge, this is the first study to assess the performance of current web privacy preserving tools in determining tracking vs. functional JavaScript programs. To improve this balance, we examine the two classes of JavaScript programs and hypothesize that tracking JavaScript programs share structural similarities that can be used to differentiate them from functional JavaScript programs. The rationale of our approach is that web developers often “borrow” and customize existing pieces of code in order to embed tracking (resp. functional) JavaScript programs into their webpages. We then propose one-class machine learning classifiers using syntactic and semantic features extracted from JavaScript programs. When trained only on samples of tracking JavaScript programs, our classifiers achieve accuracy of 99%, where the best of the privacy preserving tools achieve accuracy of 78%. The performance of our classifiers is comparable to that of traditional two-class SVM. One-class classification, where a training set of only tracking JavaScript programs is used for learning, has the advantage that it requires fewer labelled examples that can be obtained via manual inspection of public lists of well-known trackers. We further test our classifiers and several popular privacy preserving tools on a larger corpus of 4,084 websites with 135,656 JavaScript programs. The output of our best classifier on this data is between 20 to 64% different from the tools under study. We manually analyse a sample of the JavaScript programs for which our classifier is in disagreement with all other privacy preserving tools, and show that our approach is not only able to enhance user web experience by correctly classifying more functional JavaScript programs, but also discovers previously unknown tracking services. Muhammad Ikram 0001, Hassan Jameel Asghar, Mohamed Ali Kâafar, Anirban Mahanti, Balachander Krishnamurthy |
Proc. Priv. Enhancing Technol. | 3 |
| 2017 | Characterizing and Modeling User Behavior in a Large-Scale Mobile Live Streaming SystemabstractIn mobile live streaming systems, users have fairly limited interactions with streaming objects due to the constraints coming from mobile devices and the event-driven nature of live content. The constraints could lead to unique user behavior characteristics, which have yet to be explored. This paper investigates over 9 million access logs collected from the PPTV live streaming system, with an emphasis on the discrepancies that might exist when users access the live streaming catalog from mobile and nonmobile terminals. We observe a much higher likelihood of abandoning sessions by mobile users and examine the structure of abandoned sessions from the perspectives of time of day, channel content, and mobile device types. Surprisingly, we find relatively low abandonment rates during peak-load time periods and a notable impact of mobile device type (i.e., Android or iOS) on the abandonment behavior. To further capture the intrinsic characteristics of user behavior, we develop a series of models for session duration, user activity, and time dynamics of user arrivals/departures. More importantly, we relate the model parameters to physical and real-life meanings. The observations and models shed light on a video delivery system, telco-content delivery networks, and mobile applications. Zhenyu Li 0001, Mohamed Ali Kâafar, Kavé Salamatian, Gaogang Xie |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Measuring, Characterizing, and Detecting Facebook Like FarmsabstractOnline social networks offer convenient ways to reach out to large audiences. In particular, Facebook pages are increasingly used by businesses, brands, and organizations to connect with multitudes of users worldwide. As the number of likes of a page has become a de-facto measure of its popularity and profitability, an underground market of services artificially inflating page likes (“like farms ”) has emerged alongside Facebook’s official targeted advertising platform. Nonetheless, besides a few media reports, there is little work that systematically analyzes Facebook pages’ promotion methods. Aiming to fill this gap, we present a honeypot-based comparative measurement study of page likes garnered via Facebook advertising and from popular like farms. First, we analyze likes based on demographic, temporal, and social characteristics and find that some farms seem to be operated by bots and do not really try to hide the nature of their operations, while others follow a stealthier approach, mimicking regular users’ behavior. Next, we look at fraud detection algorithms currently deployed by Facebook and show that they do not work well to detect stealthy farms that spread likes over longer timespans and like popular pages to mimic regular users. To overcome their limitations, we investigate the feasibility of timeline-based detection of like farm accounts, focusing on characterizing content generated by Facebook accounts on their timelines as an indicator of genuine versus fake social activity. We analyze a wide range of features extracted from timeline posts, which we group into two main categories: lexical and non-lexical. We find that like farm accounts tend to re-share content more often, use fewer words and poorer vocabulary, and more often generate duplicate comments and likes compared to normal users. Using relevant lexical and non-lexical features, we build a classifier to detect like farms accounts that achieves a precision higher than 99% and a 93% recall. Muhammad Ikram 0001, Lucky Onwuzurike, Shehroze Farooqi, Emiliano De Cristofaro, Arik Friedman, Guillaume Jourjon, Mohamed Ali Kâafar, Zubair Shafiq |
ACM Trans. Priv. Secur. | 7 |
| 2017 | Spam Mobile Apps: Characteristics, Detection, and in the Wild AnalysisabstractThe increased popularity of smartphones has attracted a large number of developers to offer various applications for the different smartphone platforms via the respective app markets. One consequence of this popularity is that the app markets are also becoming populated with spam apps. These spam apps reduce the users’ quality of experience and increase the workload of app market operators to identify these apps and remove them. Spam apps can come in many forms such as apps not having a specific functionality, those having unrelated app descriptions or unrelated keywords, or similar apps being made available several times and across diverse categories. Market operators maintain antispam policies and apps are removed through continuous monitoring. Through a systematic crawl of a popular app market and by identifying apps that were removed over a period of time, we propose a method to detect spam apps solely using app metadata available at the time of publication. We first propose a methodology to manually label a sample of removed apps, according to a set of checkpoint heuristics that reveal the reasons behind removal. This analysis suggests that approximately 35% of the apps being removed are very likely to be spam apps. We then map the identified heuristics to several quantifiable features and show how distinguishing these features are for spam apps. We build an Adaptive Boost classifier for early identification of spam apps using only the metadata of the apps. Our classifier achieves an accuracy of over 95% with precision varying between 85% and 95% and recall varying between 38% and 98%. We further show that a limited number of features, in the range of 10--30, generated from app metadata is sufficient to achieve a satisfactory level of performance. On a set of 180,627 apps that were present at the app market during our crawl, our classifier predicts 2.7% of the apps as potential spam. Finally, we perform additional manual verification and show that human reviewers agree with 82% of our classifier predictions. Suranga Seneviratne, Aruna Seneviratne, Mohamed Ali Kâafar, Anirban Mahanti, Prasant Mohapatra |
ACM Trans. Web | 3 |
| 2016 | Gesture-Based Continuous Authentication for Wearable Devices: The Smart Glasses Use Case
Jagmohan Chauhan, Hassan Jameel Asghar, Anirban Mahanti, Mohamed Ali Kâafar |
ACNS | 4 |
| 2016 | An Analysis of the Privacy and Security Risks of Android VPN Permission-enabled Apps
Muhammad Ikram 0001, Narseo Vallina-Rodriguez, Suranga Seneviratne, Mohamed Ali Kâafar, Vern Paxson |
Internet Measurement Conference | 4 |
| 2016 | An Empirical Analysis of a Large-scale Mobile Cloud Storage Service
Zhenyu Li 0001, Xiaohui Wang 0012, Ningjing Huang, Mohamed Ali Kâafar, Zhenhua Li 0001, Jianer Zhou, Gaogang Xie, Peter Steenkiste |
Internet Measurement Conference | 4 |
| 2016 | Adwords management for third-parties in SEM: An optimisation model and the potential of TwitterabstractIn Search Engine Marketing (SEM), “third-party” partners play an important intermediate role by bridging the gap between search engines and advertisers in order to optimise advertisers' campaigns in exchange of a service fee. In this paper, we present an economic analysis of the market involving a third-party broker in Google AdWords and the broker's customers. We show that in order to optimise his profit, a third-party broker should minimise the weighted average Cost Per Click (CPC) of the portfolio of keywords attached to customer's ads while still satisfying the negotiated customer's demand. To help the broker build and manage such portfolio of keywords, we develop an optimisation framework inspired from the classical Markowitz portfolio management which integrates the customer's demand constraint and enables the broker to manage the tradeoff between return on investment and risk through a single risk aversion parameter. We then propose a method to augment the keywords portfolio with relevant keywords extracted from trending and popular topics on Twitter. Our evaluation shows that such a keywords-augmented strategy is very promising and enables the broker to achieve, on average, four folds larger return on investment than with a non-augmented strategy, while still maintaining the same level of risk. Dong Wang 0027, Zhenyu Li 0001, Gaogang Xie, Mohamed Ali Kâafar, Kavé Salamatian |
INFOCOM | 4 |
| 2016 | The Early Bird Gets the Botnet: A Markov Chain Based Early Warning System for Botnet AttacksabstractBotnet threats include a plethora of possible attacks ranging from distributed denial of service (DDoS), to drive-by-download malware distribution and spam. While for over two decades, techniques have been proposed for either improving accuracy or speeding up the detection of attacks, much of the damage is done by the time attacks are contained. In this work we take a new direction which aims to predict forthcoming attacks (i.e. before they occur), providing early warnings to network administrators who can then prepare to contain them as soon as they manifest or simply quarantine hosts. Our approach is based on modelling the Botnet infection sequence as a Markov chain with the objective of identifying behaviour that is likely to lead to attacks. We present the results of applying a Markov model to real world Botnets' data, and show that with this approach we are successfully able to predict more than 98% of attacks from a variety of Botnet families with a very low false alarm rate. Zainab Abaid, Dilip Sarkar, Mohamed Ali Kâafar, Sanjay K. Jha |
LCN | 3 |
| 2016 | TLS in the Wild: An Internet-wide Analysis of TLS-based Protocols for Electronic Communication
Ralph Holz, Johanna Amann, Olivier Mehani, Mohamed Ali Kâafar, Matthias Wachs |
NDSS | 4 |
| 2016 | Privacy-Aware Multipath Video Caching for Content-Centric NetworksabstractThe prevalence of Internet video streaming challenges the design and operation of modern networks. Content centric networking (CCN) has been proposed to address the challenges through ubiquitous in-network caching. While the expected benefits include higher performance and lower bandwidth consumption, CCN introduces new privacy issues at layer 3. This is because adversaries could infer the content consumed by others by checking cached data in routers. In this paper, we first analyze the design space to improve both caching performance and cache privacy for video delivery in CCN. In light of the observation that these two metrics need to be balanced, we propose CodingCache. It adopts network coding and random forwarding to exploit the potentials of multipath routing in CCN to improve both the diversity of cached content along different paths and the anonymity set for consumers. We evaluate CodingCache through extensive experiments based on a real-world topology and a unique data set of video access logs from a large-scale commercial video service. Our results demonstrate that, compared with the existing CCN strategies, CodingCache is able to increase the cache hit rate while also improve the use of caches across the network, together with reasonable cache privacy. Qinghua Wu 0004, Zhenyu Li 0001, Gareth Tyson, Steve Uhlig, Mohamed Ali Kâafar, Gaogang Xie |
IEEE J. Sel. Areas Commun. | 5 |
| 2016 | A differential privacy framework for matrix factorization recommender systems
Arik Friedman, Shlomo Berkovsky, Mohamed Ali Kâafar |
User Model. User Adapt. Interact. | 3 |
| 2015 | Characterizing and Predicting Viral-and-Popular Video ContentabstractThe proliferation of online video content has triggered numerous works on its evolution and popularity, as well as on the effect of social sharing on content propagation. In this paper, we focus on the observable dependencies between the virality of video content on a micro-blogging social network (in this case, Twitter) and the popularity of such content on a video distribution service (YouTube). To this end, we collected and analysed a corpus of Twitter posts containing links to YouTube clips and the corresponding video meta-data from YouTube. Our analysis highlights the unique properties of content that is both popular and viral, which allows such content to attract high number of views on YouTube and achieve fast propagation on Twitter. With this in mind, we proceed to the predictions of popular-and-viral clips and propose a framework that can, with high degree of accuracy and low amount of training data, predict videos that are likely to be popular, viral, and both. The key contribution of our work is the focus on cross-system dynamics between YouTube and Twitter. We conjecture and validate that cross-system prediction of both popularity and virality of videos is feasible, and can be performed with a reasonably high degree of accuracy. One of our key findings is that YouTube features capturing user engagement, have strong virality prediction capabilities. This findings allows to solely rely on data extracted from a video sharing service to predict popularity and virality aspects of videos. David Vallet, Shlomo Berkovsky, Sebastien Ardon, Anirban Mahanti, Mohamed Ali Kâafar |
CIKM | 5 |
| 2015 | A proactive transport mechanism with Explicit Congestion Notification for NDNabstractNamed Data Networking (NDN) shifts the communication paradigm from the quest of where the content is to what content is to be consumed. In such a new Internet architecture, transmission control mechanisms are of particular importance and have to be carefully designed to enable efficient data transmission. Existing work advocates the use of TCP-like reactive mechanisms for NDN transmission control. In this paper, we show that the statefull and adaptive forwarding properties of NDN makes proactive and efficient mechanisms for transmission control possible. We achieve this by using Explicit Congestion Notifications (ECN), which explicitly notify content consumers about network conditions through the communication path. Specifically, we propose an ECN-based proactive interest-sending rate control mechanism, which aims to achieve a high link utilisation for fast data transmission as well as a low packet dropping rate. To have a globally optimal data transmission, we further propose a smart forwarding mechanism, which locally utilises network-wide information to select the forwarding paths for individual flows. Extensive packet-level simulations in ndnSIM demonstrate that the ECN-based approach, coupled with smart forwarding, outperforms TCP-like reactive mechanisms in terms of link utilisation, packet dropping rate and flow completion time. Jianer Zhou, Qinghua Wu 0004, Zhenyu Li 0001, Mohamed Ali Kâafar, Gaogang Xie |
ICC | 4 |
| 2015 | Applying Differential Privacy to Matrix FactorizationabstractRecommender systems are increasingly becoming an integral part of on-line services. As the recommendations rely on personal user information, there is an inherent loss of privacy resulting from the use of such systems. While several works studied privacy-enhanced neighborhood-based recommendations, little attention has been paid to privacy preserving latent factor models, like those represented by matrix factorization techniques. In this paper, we address the problem of privacy preserving matrix factorization by utilizing differential privacy, a rigorous and provable privacy preserving method. We propose and study several approaches for applying differential privacy to matrix factorization, and evaluate the privacy-accuracy trade-offs offered by each approach. We show that input perturbation yields the best recommendation accuracy, while guaranteeing a solid level of privacy protection. Arnaud Berlioz, Arik Friedman, Mohamed Ali Kâafar, Roksana Boreli, Shlomo Berkovsky |
RecSys | 3 |
| 2015 | Poster: Toward Efficient and Secure Code Dissemination Protocol for the Internet of ThingsabstractCurrent Wireless Sensor Networks (WSNs) approaches do not provide an efficient and secure code dissemination function due to emerging issues of IoT applications. In this work, we adopt a multicast approach instead of the existing end-to-end or epidemic approaches. In order to enable the multicast approach, we propose an efficient/robust group key distribution scheme. We will evaluate and quantify the performance of our prototype implementation in a public testbed, while emulating several practical IoT settings, and show our security measures against known attack models. Jun Young Kim, Sanjay K. Jha, Wen Hu 0001, Hossein Shafagh, Mohamed Ali Kâafar |
SenSys | 5 |
| 2015 | Early Detection of Spam Mobile AppsabstractIncreased popularity of smartphones has attracted a large number of developers to various smartphone platforms. As a result, app markets are also populated with spam apps, which reduce the users' quality of experience and increase the workload of app market operators. Apps can be "spammy" in multiple ways including not having a specific functionality, unrelated app description or unrelated keywords and publishing similar apps several times and across diverse categories. Market operators maintain anti-spam policies and apps are removed through continuous human intervention. Through a systematic crawl of a popular app market and by identifying a set of removed apps, we propose a method to detect spam apps solely using app metadata available at the time of publication. We first propose a methodology to manually label a sample of removed apps, according to a set of checkpoint heuristics that reveal the reasons behind removal. This analysis suggests that approximately 35% of the apps being removed are very likely to be spam apps. We then map the identified heuristics to several quantifiable features and show how distinguishing these features are for spam apps. Finally, we build an Adaptive Boost classifier for early identification of spam apps using only the metadata of the apps. Our classifier achieves an accuracy over 95% with precision varying between 85%-95% and recall varying between 38%-98%. By applying the classifier on a set of apps present at the app market during our crawl, we estimate that at least 2.7% of them are spam apps. Suranga Seneviratne, Aruna Seneviratne, Mohamed Ali Kâafar, Anirban Mahanti, Prasant Mohapatra |
WWW | 3 |
| 2015 | On the Linearization of Human Identification Protocols: Attacks Based on Linear Algebra, Coding Theory, and LatticesabstractHuman identification protocols are challenge-response protocols that rely on human computational ability to reply to random challenges from the server based on a public function of a shared secret and the challenge to authenticate the human user. One security criterion for a human identification protocol is the number of challenge-response pairs the adversary needs to observe before it can deduce the secret. In order to increase this number, protocol designers have tried to construct protocols that cannot be represented as a system of linear equations or congruences. In this paper, we take a closer look at different ways from algebra, lattices, and coding theory to obtain the secret from a system of linear congruences. We then show two examples of human identification protocols from literature that can be transformed into a system of linear congruences. The resulting attack limits the number of authentication sessions these protocols can be used before secret renewal. Prior to this paper, these protocols had no known upper bound on the number of allowable sessions per secret. Hassan Jameel Asghar, Ron Steinfeld, Shujun Li 0001, Mohamed Ali Kâafar, Josef Pieprzyk |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2014 | Trace-Driven Analysis of ICN Caching Algorithms on Video-on-Demand WorkloadsabstractEven though a key driver for Information-Centric Networking (ICN) has been the rise in Internet video traffic, there has been surprisingly little work on analyzing the interplay between ICN and video ? which ICN caching strategies work well on video work- loads and how ICN helps improve video-centric quality of experience (QoE). In this work, we bridge this disconnect with a trace- driven study using 196M video requests from over 16M users on a country-wide topology with 80K routers. We evaluate a broad space of content replacement (e.g., LRU, LFU, FIFO) and content placement (e.g., leave a copy everywhere, probabilistic) strategies over a range of cache sizes. We highlight four key findings: (1) the best placement and re- placement strategies depend on the cache size and vary across improvement metrics; that said, LFU+probabilistic caching [37] is a close-to-optimal strategy overall; (2) video workloads show considerable caching-related benefits (e.g., -- 10% traffic reduction) only with very large cache sizes (≥ 100GB); (3) the improvement in video QoE is low (≥ 12%) if the content provider already has a substantial geographical presence; and (4) caches in the middle and the edge of the network, requests from highly populated regions and without content servers, and requests for popular content contribute most to the overall ICN-induced improvements in video QoE. Yi Sun 0004, Seyed Kaveh Fayaz, Vyas Sekar, Yun Jin, Mohamed Ali Kâafar, Steve Uhlig |
CoNEXT | 6 |
| 2014 | CRCache: Exploiting the correlation between content popularity and network topology information for ICN cachingabstractInformation-centric networking (ICN) is designed to decouple contents from hosts at the network layer, using in-network caching as a key feature to improve the overall performance. However, the en-route caching strategy used in many ICN implementations generally yields redundancies in the cached contents across different routers. There are also some recent works focusing on cache optimization by respectively exploiting either application layer or network layer, which we think is not sufficient to increase cache hit rate and reduce traffic. In this paper, we propose a novel caching scheme (CRCache) that utilizes a cross-layer design to cache contents in a few selected routers based on the correlation of content popularity and the network topology. Specifically, through exploiting information available at both application and network layers, CRCache aims to improve the cache hit rate and reduce the overall network traffic. We conduct a large scale and real traces-driven simulation with an underlying real Internet topology in China, and show that by using CRCache, the overall cache hit rate is increased by 62.5% and network traffic reduction is improved by at least 42% compared with recent single layer schemes. Wei Wang 0157, Yi Sun 0004, Mohamed Ali Kâafar, Jiong Jin, Jun Li 0002, Zhongcheng Li |
ICC | 4 |
| 2014 | Freeway: Adaptively Isolating the Elephant and Mice Flows on Different Transmission PathsabstractThe network resource competition of today' data enters is extremely intense between long-lived elephant flows and latency-sensitive mice flows. Achieving both goals of high throughput and low latency respectively for the two types of flows requires compromise, which recent research has not successfully solved mainly due to the transfer of elephant and mice flows on shared links without any differentiation. However, current data enters usually adopt clos-based topology, e.g. Fat-tree/VL2, so there exist multiple shortest paths between any pair of source and destination. In this paper, we leverage on this observation to propose a flow scheduling scheme, Freeway, to adaptively partition the transmission paths into low latency paths and high throughput paths respectively for the two types of flows. An algorithm is proposed to dynamically adjust the number of the two types of paths according to the real-time traffic. And based on these separated transmission paths, we propose different flow type-specific scheduling and forwarding methods to make full utilization of the bandwidth. Our simulation results show that Freeway significantly reduces the delay of mice flow by 85.8% and achieves 9.2% higher throughput compared with Hedera. Wei Wang 0157, Yi Sun 0004, Kai Zheng 0003, Mohamed Ali Kâafar, Dan Li 0001, Zhongcheng Li |
ICNP | 4 |
| 2014 | Censorship in the Wild: Analyzing Internet Filtering in SyriaabstractInternet censorship is enforced by numerous governments worldwide, however, due to the lack of publicly available information, as well as the inherent risks of performing active measurements, it is often hard for the research community to investigate censorship practices in the wild. Thus, the leak of 600GB worth of logs from 7 Blue Coat SG-9000 proxies, deployed in Syria to filter Internet traffic at a country scale, represents a unique opportunity to provide a detailed snapshot of a real-world censorship ecosystem. Abdelberi Chaabane, Terence Chen, Mathieu Cunche, Emiliano De Cristofaro, Arik Friedman, Mohamed Ali Kâafar |
Internet Measurement Conference | 6 |
| 2014 | Paying for Likes?: Understanding Facebook Like Fraud Using HoneypotsabstractFacebook pages offer an easy way to reach out to a very large audience as they can easily be promoted using Facebook's advertising platform. Recently, the number of likes of a Facebook page has become a measure of its popularity and profitability, and an underground market of services boosting page likes, aka like farms, has emerged. Some reports have suggested that like farms use a network of profiles that also like other pages to elude fraud protection algorithms, however, to the best of our knowledge, there has been no systematic analysis of Facebook pages' promotion methods. Emiliano De Cristofaro, Arik Friedman, Guillaume Jourjon, Mohamed Ali Kâafar, Zubair Shafiq |
Internet Measurement Conference | 4 |
| 2014 | On the geographic patterns of a large-scale mobile video-on-demand systemabstractThe widespread availability of smart mobile terminals along with the ever increasing bandwidth capabilities has promoted the popularity of mobile Internet video systems. Understanding the geographic features of mobile content consumption is of an extreme importance for the design and the performance optimization of a mobile video delivery system. This paper is a first step towards characterization of the geographic patterns of a large-scale commercial mobile video-on-demand (VoD) system, by measuring both uniformity and intensity of geographic interests on videos. In particular, we identify a geographical concentration effect of views for individual videos, which is however dependent on video popularity. We also analyze the temporal evolution trends of the geographic popularity which reveal distinct behavior of popular and non-popular videos. While the set of locations that contribute to most of the views of non-popular videos largely varies, the daily geographic popularity distribution of popular videos closely follows the distribution of global traffic and remains stable. We also examine the impact of content type and viewing sources on the geographic features of mobile videos consumption, and the correlation between content similarity and geographic locality. Finally, we provide insights into the implications of our findings. Zhenyu Li 0001, Gaogang Xie, Jiali Lin, Yun Jin, Mohamed Ali Kâafar, Kavé Salamatian |
INFOCOM | 5 |
| 2014 | Improving business rating predictions using graph based featuresabstractMany types of recommender systems rely on a rich ensemble of user, item, and context features when generating recommendations for users. The features can be either manually engineered or automatically extracted from the available data, such that feature engineering becomes an important step in the recommendation process. In this work, we propose to leverage graph based representation of the data in order to generate and automatically populate features. We represent the standard user-item rating matrix and some domain metadata, as graph vertices and edges. Then, we apply a suite of graph theory and network analysis metrics to the graph based data representation, to populate features that augment the original user-item ratings data. The augmented data is fed into a classifier that predicts unknown user ratings, which are used for the generation of recommendations. We evaluate the proposed methodology using the recently released Yelp business ratings dataset. Our results indicate that the automatically populated graph features allow for more accurate and robust predictions, with respect to both the variability and sparsity of ratings. Amit Tiroshi, Shlomo Berkovsky, Mohamed Ali Kâafar, David Vallet, Terence Chen, Tsvi Kuflik |
IUI | 3 |
| 2014 | Demo: Crowd-cache - popular content for freeabstractCrowd-Cache is a novel crowd-sourced content caching system which provides cheap and convenient content access for mobile users. Our system exploits both transient colocation of devices and the spatial temporal correlation of content popularity, where users in a particular location and at specific times would be likely interested in similar content. We demonstrate the feasibility of Crowd-Cache system through a prototype implementation on Android smartphones. Kanchana Thilakarathna, Fangzhou Jiang, Sirine Mrabet, Mohamed Ali Kâafar, Aruna Seneviratne, Prasant Mohapatra |
MobiSys | 4 |
| 2014 | A Closer Look at Third-Party OSN Applications: Are They Leaking Your Personal Information?
Abdelberi Chaabane, Yuan Ding 0003, Ratan Dey, Mohamed Ali Kâafar, Keith W. Ross |
PAM | 4 |
| 2014 | On the Effectiveness of Obfuscation Techniques in Online Social Networks
Terence Chen, Roksana Boreli, Mohamed Ali Kâafar, Arik Friedman |
Privacy Enhancing Technologies | 3 |
| 2014 | Graph-Based Recommendations: Make the Most Out of Social Data
Amit Tiroshi, Shlomo Berkovsky, Mohamed Ali Kâafar, David Vallet, Tsvi Kuflik |
UMAP | 3 |
| 2014 | Linking wireless devices using information contained in Wi-Fi probe requests
Mathieu Cunche, Mohamed Ali Kâafar, Roksana Boreli |
Pervasive Mob. Comput. | 2 |
| 2013 | The Where and When of Finding New Friends: Analysis of a Location-based Social Discovery Network
Terence Chen, Mohamed Ali Kâafar, Roksana Boreli |
ICWSM | 2 |
| 2013 | A genealogy of information spreading on microblogs: A Galton-Watson-based explicative modelabstractIn this paper, we study the process of information diffusion in a microblog service developing Galton-Watson with Killing (GWK) model. Microblog services offer a unique approach to online information sharing allowing microblog users to forward messages to others. We describe an information propagation as a discrete GWK process based on Galton-Watson model which models the evolution of family names. Our model explains the interaction between the topology of the social graph and the intrinsic interest of the message. We validate our model on dataset collected from Sina Weibo and Twitter microblog. Sina Weibo is a Chinese microblog web service which reached over 100 million users as for January 2011. Our Sina Weibo dataset contains over 261 thousand tweets which have retweets and 2 million retweets from 500 thousand users. Twitter dataset contains over 1.1 million tweets which have retweets and 3.3 million retweets from 4.3 million users. The results of the validation show that our proposed GWK model fits the information diffusion of microblog service very well in terms of the number of message receivers. We show that our model can be used in generating tweets load and also analyze the relationships between parameters of our model and popularity of the diffused information. To the best of our knowledge, this paper is the first to give a systemic and comprehensive analysis for the information diffusion on microblog services, to be used in tweets-like load generators while still guaranteeing popularity distribution characteristics. Dong Wang 0027, Hosung Park, Gaogang Xie, Sue B. Moon, Mohamed Ali Kâafar, Kavé Salamatian |
INFOCOM | 5 |
| 2013 | LMD: A local minimum driven and self-organized method to obtain locatorsabstractThe scalability of routing architectures for large networks is one of the biggest challenges that the Internet faces today. Greedy routing, in which each node is assigned a locator used as a distance metric, recently received increased attention from researchers and is considered as a potential solution for scalable routing. In this paper, we propose LMD - a Local Minimum Driven method to compute the topology-based locator. As opposed to previous work, our algorithm employs a quasigreedy and self-organized embedding method, which outperforms similar decentralized algorithms by up to 20% in success rate. To eliminate the negative effect of the “quasi” greedy property - transfer routes longer than the shortest routes, we introduce a two-stage routing strategy, which combines the greedy routing with source routing. The greedy routing path discovered and compressed in the first stage is then used by the following source-routing stage. Through extensive evaluations, based on synthetic topologies as well as on a snapshot of the real Internet AS topology, we show that LMD guarantees 100% delivery rate on large networks with a very low stretch. Yonggong Wang, Gaogang Xie, Mohamed Ali Kâafar, Steve Uhlig |
ISCC | 3 |
| 2013 | A Layered Secret Sharing Scheme for Automated Profile Sharing in OSN Groups
Guillaume Smith, Roksana Boreli, Mohamed Ali Kâafar |
MobiQuitous | 3 |
| 2013 | How Much Is Too Much? Leveraging Ads Audience Estimation to Evaluate Public Profile Uniqueness
Terence Chen, Abdelberi Chaabane, Pierre-Ugo Tournoux, Mohamed Ali Kâafar, Roksana Boreli |
Privacy Enhancing Technologies | 4 |
| 2013 | Holiday Pictures or Blockbuster Movies? Insights into Copyright Infringement in User Uploads to One-Click File Hosters
Tobias Lauinger, Kaan Onarlioglu, Abdelberi Chaabane, Engin Kirda, William K. Robertson, Mohamed Ali Kâafar |
RAID | 6 |
| 2013 | Cross social networks interests predictions based ongraph featuresabstractThe tremendous popularity of Online Social Networks (OSN) has led to situations, where users have their profiles spread across multiple networks. These partial profiles reflect different user characteristics, depending mainly on the nature of the network, e.g., Facebook's social vs. LinkedIn's professional focus. Combining data gathered by multiple networks may benefit individual users, and the community as a whole, as this could facilitate the provision of more accurate services and recommendations. This paper reports on an exploratory study of the process of making such recommendations using a unique multi-network dataset containing user interests across multiple domains, e.g., music, books, and movies. We represent the data using a graph model and generate recommendations using a set of features extracted from and populated by the model. We assess the contribution of various network- and domain-related features to the accuracy of the recommendations and motivate future work into automated feature selection. Amit Tiroshi, Shlomo Berkovsky, Mohamed Ali Kâafar, Terence Chen, Tsvi Kuflik |
RecSys | 3 |
| 2013 | On the Effectiveness of Dynamic Taint Analysis for Protecting against Private Information Leaks on Android-based Devices
Golam Sarwar, Olivier Mehani, Roksana Boreli, Mohamed Ali Kâafar |
SECRYPT | 4 |
| 2012 | IBTrack: An ICMP black holes trackerabstractICMP is a key protocol to exchange control and error messages over the Internet. An appropriate ICMP's processing throughout a path is therefore a key requirement both for troubleshooting operations (e.g. debugging routing problems) and for several functionnalities (e.g. Path Maximum Transmission Unit Discovery, PMTUD). Unfortunately it is common to see ICMP malfunctions, thereby causing various levels of problems. The contributions of this paper are threefold. We first introduce a taxonomy of the way routers process ICMP, which is of great help to understand for instance certain traceroute outputs. Secondly we introduce IBTrack, a tool that any user can use to automatically characterize ICMP issues within the Internet, without requiring any additional in-network assistance (e.g. there is no vantage point). Finally we validate our IBTrack tool with large scale experiments and we take advantage of this opportunity to provide some statistics on how ICMP is managed by Internet routers. Ludovic Jacquin, Vincent Roca, Mohamed Ali Kâafar, Fabrice Schuler, Jean-Louis Roch |
GLOBECOM | 3 |
| 2012 | Watching videos from everywhere: a study of the PPTV mobile VoD systemabstractIn this paper, we examine mobile users' behavior and their corresponding video viewing patterns from logs extracted from the servers of a large scale VoD system. We focus on the analysis of the main discrepancies that might exist when users access the VoD system catalog from WiFi or 3G connections. We also study factors that might impact mobile users' interests and video popularity. The users' behavior exhibits strong daily and weekly patterns, with mobile users' interests being surprisingly spread across almost all categories and video lengths, independently of the connection type. However, by examining the activity of users individually, we observed a concentration of interests and peculiar access patterns, which allows to classify the users and thus better predict their behavior. We also find a skewed video popularity distribution and then demonstrate that the popularity of a video can be predicted using its very early popularity level. We further analyzed the sources of video viewing and found that even if search engines are the dominant sources for a majority of videos, they represent less than 10% (resp. 20%) of the sources for the highly popular videos in 3G (resp. WiFi) network. We report that both the type of connections and mobile devices in use have an impact on the viewing time and the source of viewing. Based on our findings, we provide insights and recommendations that can be used to design intelligent mobile VoD systems and help improving personalized services on these platforms. Zhenyu Li 0001, Jiali Lin, Marc-Ismaël Akodjènou, Gaogang Xie, Mohamed Ali Kâafar, Yun Jin |
Internet Measurement Conference | 5 |
| 2012 | FPC: A self-organized greedy routing in scale-free networksabstractIn this paper we propose FPC - a Force-based layout and Path Compressing routing schema for scale-free network. As opposed to previous work, our algorithm employs a quasi-greedy but self-organized and configuration-free embedding method - force-based layout. In order to eliminate the negative influences of the “quasi” greedy property, we present a two-stage routing strategy, which combines the greedy routing with source routing. The greedy routing path discovered and compressed in a first stage is then used by the following source-routing stage. The detailed evaluation based on synthetic topologies as well as on a real Internet AS topology shows that: FPC guarantees 100% delivery rates on scale-free networks with an attractive low stretch (e.g. less than 1.2 on the real Internet AS topology). Yonggong Wang, Gaogang Xie, Mohamed Ali Kâafar |
ISCC | 3 |
| 2012 | You are what you like! Information leakage through users' Interests
Abdelberi Chaabane, Gergely Ács, Mohamed Ali Kâafar |
NDSS | 3 |
| 2012 | Betrayed by Your Ads! - Reconstructing User Profiles from Targeted Ads
Claude Castelluccia, Mohamed Ali Kâafar, Minh-Dung Tran |
Privacy Enhancing Technologies | 2 |
| 2012 | I know who you will meet this evening! Linking wireless devices using Wi-Fi probe requestsabstractActive service discovery in Wi-Fi involves wireless stations broadcasting their Wi-Fi fingerprint, i.e. the SSIDs of their preferred wireless networks. The content of those Wi-Fi fingerprints can reveal different types of information about the owner. We focus on the relation between the fingerprints and the links between the owners. Our hypothesis is that social links between devices owners can be identified by exploiting the information contained in the fingerprint. More specifically we propose to consider the similarity between fingerprints as a metric, with the underlying idea: similar fingerprints are likely to be linked. We first study the performances of several similarity metrics on a controlled dataset and then apply the designed classifier to a dataset collected in the wild. Finally we discuss how Wi-Fi fingerprint can reveal informations on the nature of the links between users. This study is based on a dataset collected in Sydney, Australia, composed of fingerprints corresponding to more than 8000 devices. Mathieu Cunche, Mohamed Ali Kâafar, Roksana Boreli |
WOWMOM | 2 |
| 2012 | Path similarity evaluation using Bloom filters
Benoit Donnet, Bamba Gueye, Mohamed Ali Kâafar |
Comput. Networks | 3 |
| 2011 | EphPub: Toward robust Ephemeral PublishingabstractThe increasing amount of personal and sensitive information disseminated over the Internet prompts commen-surately growing privacy concerns. Digital data often lingers indefinitely and users lose its control. This motivates the desire to restrict content availability to an expiration time set by the data owner. This paper presents and formalizes the notion of Ephemeral Publishing (EphPub), to prevent the access to expired content. We propose an efficient and robust protocol that builds on the Domain Name System (DNS) and its caching mechanism. With EphPub, sensitive content is published encrypted and the key material is distributed, in a steganographic manner, to randomly selected and independent resolvers. The availability of content is then limited by the evanescence of DNS cache entries. The EphPub protocol is transparent to existing applications, and does not rely on trusted hardware, centralized servers, or user proactive actions. We analyze its robustness and show that it incurs a negligible overhead on the DNS infrastructure. We also perform a large-scale study of the caching behavior of 900K open DNS resolvers. Finally, we propose Firefox and Thunderbird extensions that provide ephemeral publishing capabilities, as well as a command-line tool to create ephemeral files. Claude Castelluccia, Emiliano De Cristofaro, Aurélien Francillon, Mohamed Ali Kâafar |
ICNP | 4 |
| 2011 | How Unique and Traceable Are Usernames?
Daniele Perito, Claude Castelluccia, Mohamed Ali Kâafar, Pere Manils |
PETS | 3 |
| 2010 | Digging into Anonymous Traffic: A Deep Analysis of the Tor Anonymizing NetworkabstractUsers' anonymity and privacy are among the major concerns of today's Internet. Anonymizing networks are then poised to become an important service to support anonymous-driven Internet communications and consequently enhance users' privacy protection. Indeed, Tor an example of anonymizing networks based on onion routing concept attracts more and more volunteers, and is now popular among dozens of thousands of Internet users. Surprisingly, very few researches shed light on such an anonymizing network. Beyond providing global statistics on the typical usage of Tor in the wild, we show that Tor is actually being is-used, as most of the observed traffic belongs to P2P applications. In particular, we quantify the BitTorrent traffic and show that the load of the latter on the Tor network is underestimated because of encrypted BitTorrent traffic (that can go unnoticed). Furthermore, this paper provides a deep analysis of both the HTTP and BitTorrent protocols giving a complete overview of their usage. We do not only report such usage in terms of traffic size and number of connections but also depict how users behave on top of Tor. We also show that Tor usage is now diverted from the onion routing concept and that Tor exit nodes are frequently used as 1-hop SOCKS proxies, through a so-called tunneling technique. We provide an efficient method allowing an exit node to detect such an abnormal usage. Finally, we report our experience in effectively crawling bridge nodes, supposedly revealed sparingly in Tor. Abdelberi Chaabane, Pere Manils, Mohamed Ali Kâafar |
NSS | 3 |
| 2009 | Certified Internet CoordinatesabstractWe address the issue of asserting the accuracy of coordinates advertised by nodes of Internet coordinate systems during distance estimations. Indeed, some nodes may lie deliberately about their coordinates to mount various attacks against applications and overlays. Our proposed method consists in two steps: 1) establish the correctness of a node's claimed coordinate (which leverages our previous work on securing the coordinates embedding phase using a Surveyor infrastructure); and 2) issue a time limited validity certificate for each verified coordinate. Validity periods are computed based on an analysis of coordinate inter-shift times observed on PlanetLab, and shown to follow a long-tail distribution (lognormal distribution in most cases, or Weibull distribution otherwise). The effectiveness of the coordinate certification method is validated by measuring the impact of a variety of attacks on distance estimates. Mohamed Ali Kâafar, Laurent Mathy, Chadi Barakat, Kavé Salamatian, Thierry Turletti, Walid Dabbous |
ICCCN | 1 |
| 2009 | Geolocalization of proxied services and its application to fast-flux hidden serversabstractFast-flux is a redirection technique used by cyber-criminals to hide the actual location of malicious servers. Its purpose is to evade identification and prevent or, at least delay, the shutdown of these illegal servers by law enforcement. Claude Castelluccia, Mohamed Ali Kâafar, Pere Manils, Daniele Perito |
Internet Measurement Conference | 2 |
| 2009 | Detecting Triangle Inequality Violations in Internet Coordinate Systems by Supervised Learning
Yongjun Liao, Mohamed Ali Kâafar, Bamba Gueye, François Cantin, Pierre Geurts, Guy Leduc |
Networking | 2 |
| 2008 | Overlay routing using coordinate systemsabstractWe address the problem of finding indirect overlay paths that reduce the latency between pairs of nodes in an overlay. To this end we propose to rely on an Internet Coordinate System (ICS), namely Vivaldi, to estimate RTTs and help find these interesting detours. We define two initial criteria to illustrate our approach and assess their true/false positive rates. François Cantin, Bamba Gueye, Mohamed Ali Kâafar, Guy Leduc |
CoNEXT | 3 |
| 2008 | Towards a Two-Tier Internet Coordinate System to Mitigate the Impact of Triangle Inequality Violations
Mohamed Ali Kâafar, Bamba Gueye, François Cantin, Guy Leduc, Laurent Mathy |
Networking | 1 |
| 2007 | Securing internet coordinate embedding systemsabstractThis paper addresses the issue of the security of Internet Coordinate Systems,by proposing a general method for malicious behavior detection during coordinate computations. We first show that the dynamics of a node, in a coordinate system without abnormal or malicious behavior, can be modeled by a Linear State Space model and tracked by a Kalman filter. Then we show, that the obtained model can be generalized in the sense that the parameters of a filtercalibrated at a node can be used effectively to model and predict the dynamic behavior at another node, as long as the two nodes are not too far apart in the network. This leads to the proposal of a Surveyor infrastructure: Surveyor nodes are trusted, honest nodes that use each other exclusively to position themselves in the coordinate space, and are therefore immune to malicious behavior in the system.During their own coordinate embedding, other nodes can thenuse the filter parameters of a nearby Surveyor as a representation of normal, clean system behavior to detect and filter out abnormal or malicious activity. A combination of simulations and PlanetLab experiments are used to demonstrate the validity, generality, and effectiveness of the proposed approach for two representative coordinate embedding systems, namely Vivaldi and NPS. Mohamed Ali Kâafar, Laurent Mathy, Chadi Barakat, Kavé Salamatian, Thierry Turletti, Walid Dabbous |
SIGCOMM | 1 |
| 2006 | Virtual networks under attack: disrupting internet coordinate systemsabstractInternet coordinate-based systems are poised to become an important service to support overlay construction and topology-aware applications. Indeed, through network distance embedding into an appropriate geometric space, such systems allow for accurate network distance estimations with low overhead. However, coordinate systems often rely on good cooperation between nodes for correct coordination and assume that information reported by probed nodes is correct. In this paper, we identify various attacks against coordinate embedding systems and show their effectiveness on two representative positioning systems, namely Vivaldi and NPS. Our study demonstrates that these attacks can seriously disrupt the operations of these systems and therefore the virtual networks and applications relying on them for distance measurements. Through simulations of different potential scenarios where malicious nodes provide biased coordinate information and delay measurement probes, we quantify the effects of attack strategies that aim to (i) introduce disorder in the system, (ii) fool honest nodes to move far away from their correct positions and (iii) isolate particular target nodes in the system through collusion. Our findings confirm the susceptibility of the coordinate systems to such attacks. Mohamed Ali Kâafar, Laurent Mathy, Thierry Turletti, Walid Dabbous |
CoNEXT | 1 |
| 2006 | Wireless alternative best effort service: a case study of ALMabstractThe availability of interactive multimedia applications (skype, wengo, LiveCom, etc.) is going to modify the usage of future Internet. Such Internet is likely to be more heterogeneous, consisting of different islands of ad-hoc networks. This argues for the growth of interactive applications traffic at the ad-hoc networks level. Providing Quality of Service (QoS) support, in particular low loss rate and low end-to-end delay, in such networks is a challenging task. Limited bandwidth resource and high mobility are two major characteristics of such networks. Moreover, the link breakage rate is high, which leads to high loss rate in the network. In the literature, researches generally span over routing protocols and MAC Layer mechanisms [1] to guarantee few QoS aspects by introducing signalling traffic. Cyrine Mrabet, Mohamed Ali Kâafar, Farouk Kamoun |
CoNEXT | 2 |
| 2006 | A Locating-First Approach for Scalable Overlay MulticastabstractRecent proposals in multicast overlay networks have demonstrated the importance of exploiting underlying network topology data to construct efficient overlays. While they avoid virtual coordinates embedding and fixed landmarks measurements, these topology-aware proposals often rely on incremental and periodic refinements to improve each node's position in the delivery tree. We claim that there are barriers for the scalability of existing overlay multicast protocols. In fact, periodical refinement and control processes induce additional overhead and high communication cost. On the other hand, users attending a video conferencing session or an event broadcast expect an acceptable quality as soon as they join the multicast session. It is then important to overcome an efficiency problem from which almost all current overlay multicast proposals suffer. This problem is the long convergence time to reach a stabilized quality state in the overlay delivery tree. We propose a novel overlay multicast tree construction scheme, called LCC : Locate, Cluster and Conquer, designed to address the aforementioned scalability and efficiency issues. The scheme consists in two phases: a selective locating phase and an overlay construction phase. Using partial knowledge of location-information for participating nodes, the selective locating phase algorithm consists in locating the closest existing set of nodes (cluster) in the overlay for a newcomer. It allows then to avoid initially randomly-connected structures without using virtual coordinates system embedding nor fixed landmarks measurements. Then, on the basis of this locating process, the overlay construction phase consists in building and managing a topology-aware clustered hierarchical overlay. Mohamed Ali Kâafar, Thierry Turletti, Walid Dabbous |
INFOCOM | 1 |
| 2006 | A Locating-First Approach for Scalable Overlay MulticastabstractRecent proposals in multicast overlay construction have demonstrated the importance of exploiting underlying network topology. However, these topology-aware proposals often rely on incremental and periodic refinements to improve the system performance. These approaches are therefore neither scalable, as they induce high communication cost due to refinement overhead, nor efficient because long convergence time is necessary to obtain a stabilized structure. In this paper, we propose a highly scalable locating algorithm that gradually directs newcomers to their a set of their closest nodes without inducing high overhead. On the basis of this locating process, we build a robust and scalable topology-aware clustered hierarchical overlay scheme, called LCC. We conducted both simulations and PlanetLab experiments to evaluate the performance of LCC. Results show that the locating process entails modest resources in terms of time and bandwidth. Moreover, LCC demonstrates promising performance to support large scale multicast applications Mohamed Ali Kâafar, Thierry Turletti, Walid Dabbous |
IWQoS | 1 |
| 2004 | A Kerberos-Based Authentication Architecture for Wireless LANs
Mohamed Ali Kâafar, Lamia Ben Azzouz, Farouk Kamoun, Davor Males |
NETWORKING | 1 |