Mohamed Ali Kâafar

dblp:71/5612 · also Dali Kaafar, Mohamad Ali Kâafar · DBLP profile ↗
← Back
18ranked-venue papers in the field
0as first author
10since 2021 · last 2026
0000-0003-2714-0276ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 11Data Mining & Knowledge Discovery · 4Big Data, Cloud & Distributed Data Systems · 2Database Systems & Data Management · 1
YearPublicationVenuePosition
2026 DIALEVAL: Automated Type-Theoretic Evaluation of LLM Instruction Following
Nardine Basta, Mohamed Ali Kâafar
PAKDD (4)2
2026 Forget Me, Not My Friends! Object Unlearning Based on Scene Graphs
abstract
Machine unlearning offers a practical technical means for fulfilling users' requests to remove personally identifiable information (PII) under ''right to be forgotten'' regulations such as GDPR and COPPA. Traditionally, unlearning is performed with the removal of entire data samples (sample unlearning) or whole features across the dataset (feature unlearning). However, when the removal request targets only certain parts of the PII, such as specific objects within a sample, these traditional unlearning approaches fall short of meeting such finer-grained unlearning requirements. To address this gap, we propose a scene graph-based object unlearning framework. This framework utilizes scene graphs, rich in semantic representation, transparently translate unlearning requests into actionable steps. The result, is the preservation of the overall semantic integrity of the generated image, bar the unlearned object. Furthermore, we develop three distinct approaches for object unlearning, grounded in the mainstream unlearning techniques of fine-tuning and model redaction. For validation, we evaluate the unlearned object's fidelity in outputs under the tasks of image reconstruction and image synthesis. Our proposed framework demonstrates improved object unlearning outcomes, with the preservation of unrequested samples in contrast to sample and feature learning methods. This work addresses critical privacy issues by increasing the granularity of targeted machine unlearning through forgetting specific object-level details without sacrificing the utility of the whole data sample or dataset feature.
Chenhan Zhang, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Weiqi Wang 0003, An Liu 0002, Mohamed Ali Kâafar
WSDM6
2025 Membership Inference Attack Vulnerabilities of Record Linkage Models
abstract
Record linkage plays a crucial role in integrating health, legal, and administrative data. In domains such as healthcare, unstructured records, including clinical notes, contain rich information, making their integration valuable for applications like clinical trials. While deep learning models have improved linkage quality, their privacy risks remain under-explored. We present what is, to our knowledge, the first systematic study of membership inference attack vulnerabilities in record linkage models trained on de-identified texts. Unlike traditional classifiers, linkage models operate on record pairs, prompting a rethinking of what constitutes membership leakage. Does it occur if only a record pair was seen together during training or individually? Or if neither was seen, but the pair resembles known patterns? We introduce a black-box attack based on semantic perturbation sensitivity, requiring no access to the model's internals. Our findings expose a previously unaddressed membership inference vulnerability in record linkage models: black-box attacks, even with a simple threshold-based attack model, achieved up to 92% precision and AUC 0.79. Across record linkage models trained on MIMIC-IV and PMC-Patients datasets, we observe that perturbing training-seen phrases causes significantly larger confidence shifts (e.g., Δ s = -0.648) and higher change in predicted label (label flip from 0 = non-match to 1 = match and vice versa) (e.g., 68.75%) compared to unseen variants. These preliminary results reveal membership inference vulnerabilities in text-based linkage systems, highlighting the need for deeper investigation into privacy risks, motivating new lines of privacy defences for pairwise models.
Piyumi Seneviratne, Dinusha Vatsalan, Mohamed Ali Kâafar
CIKM3
2025 Bot Wars Evolved: Orchestrating Competing LLMs in a Counterstrike Against Phone Scams
Nardine Basta, Conor Atkins, Mohamed Ali Kâafar
PAKDD (7)3
2025 Can Self Supervision Rejuvenate Similarity-Based Link Prediction?
Chenhan Zhang, Weiqi Wang 0003, Zhiyi Tian, James Jian Qiao Yu, Mohamed Ali Kâafar, An Liu 0002, Shui Yu 0001
PAKDD (7)5
2024 More Than Just a Random Number Generator! Unveiling the Security and Privacy Risks of Mobile OTP Authenticator Apps
Muhammad Ikram 0001, I Wayan Budi Sentana, Hassan Jameel Asghar, Mohamed Ali Kâafar, Michal Kepkowski
WISE (5)4
2024 On Adversarial Training with Incorrect Labels
Benjamin Zi Hao Zhao, Junda Lu 0001, Xiaowei Zhou 0003, Dinusha Vatsalan, Muhammad Ikram 0001, Mohamed Ali Kâafar
WISE (4)6
2023 Exploring the Distinctive Tweeting Patterns of Toxic Twitter Users
abstract
In the pursuit of bolstering user safety, social media platforms deploy active moderation strategies, including content removal and user suspension. These measures target users engaged in discussions marked by hate speech or toxicity, often linked to specific keywords or hashtags. Nonetheless, the increasing prevalence of toxicity indicates that certain users adeptly circumvent these measures.This study examines consistently toxic users on Twitter (rebranded as X) Rather than relying on traditional methods based on specific topics or hashtags, we employ a novel approach based on patterns of toxic tweets, yielding deeper insights into their behavior.We analyzed 38 million tweets from the timelines of 12,148 Twitter users and identified the top 1,457 users who consistently exhibit toxic behavior, relying on metrics like the Gini index and Toxicity score. By comparing their posting patterns to those of non-consistently toxic users, we have uncovered distinctive temporal patterns, including contiguous activity spans, inter-tweet intervals (referred to as “Burstiness”), and churn analysis. These findings provide strong evidence for the existence of a unique tweeting pattern associated with toxic behavior on Twitter.Crucially, our methodology transcends Twitter and can be adapted to various social media platforms, facilitating the identification of consistently toxic users based on their posting behavior. This research contributes to ongoing efforts to combat online toxicity and offers insights for refining moderation strategies in the digital realm. We are committed to open research and will provide our code and data to the research community.
Hina Qayyum, Muhammad Ikram 0001, Benjamin Zi Hao Zhao, Ian D. Wood, Nicolas Kourtellis, Mohamed Ali Kâafar
IEEE Big Data6
2023 On mission Twitter Profiles: A Study of Selective Toxic Behavior
abstract
The argument for persistent social media influence campaigns, often funded by malicious entities, is gaining traction. These entities utilize instrumented profiles to disseminate divisive content and disinformation, shaping public perception. Despite ample evidence of these instrumented profiles, few identification methods exist to locate them in the wild. To evade detection and appear genuine, small clusters of instrumented profiles engage in unrelated discussions, diverting attention from their true goals [34]. This strategic thematic diversity conceals their selective polarity towards certain topics and fosters public trust [49]. This study aims to characterize profiles potentially used for influence operations, termed “on-mission profiles,” relying solely on thematic content diversity within unlabeled data. Distinguishing this work is its focus on content volume and toxicity towards specific themes. Longitudinal data from 138K Twitter (rebranded as X) profiles and 293M tweets enables profiling based on theme diversity. High thematic diversity groups predominantly produce toxic content concerning specific themes, like politics, health, and news—classifying them as “on-mission” profiles. Using the identified on-mission” profiles, we design a classifier for unseen, unlabeled data. Employing a linear SVM model, we train and test it on an 80/20% split of the most diverse profiles. The classifier achieves a flawless 100% accuracy, facilitating the discovery of previously unknown “on-mission” profiles in the wild.
Hina Qayyum, Muhammad Ikram 0001, Benjamin Zi Hao Zhao, Ian D. Wood, Nicolas Kourtellis, Mohamed Ali Kâafar
IEEE Big Data6
2023 Local Differentially Private Fuzzy Counting in Stream Data Using Probabilistic Data Structures
abstract
Privacy-preserving estimation of counts of items in streaming data finds applications in several real-world scenarios including word auto-correction and traffic management applications. Recent works of RAPPOR [1] and Apple's count-mean sketch (CMS) algorithm [2] propose privacy preserving mechanisms for count estimation in large volumes of data using probabilistic data structures like counting Bloom filter and CMS. However, these existing methods fall short in providing a sound solution for real-time streaming data applications. Since the size of the data structure in these methods is not adaptive to the volume of the streaming data, the utility (accuracy of the count estimate) can suffer over time due to increased false positive rates. Further, the lookup operation needs to be highly efficient to answer count estimate queries in real-time. More importantly, the local Differential privacy mechanisms used in these approaches to provide privacy guarantees come at a large cost to utility (impacting the accuracy of count estimation). In this work, we propose a novel (local) Differentially private mechanism that provides high utility for the streaming data count estimation problem with similar or even lower privacy budgets while providing: a) fuzzy counting to report counts of related or similar items (for instance to account for typing errors and data variations), and b) improved querying efficiency to reduce the response time for real-time querying of counts. Our algorithm uses a combination of two probabilistic data structures Cuckoo filter and Bloom filter. We provide formal proofs for privacy and utility guarantees and present extensive experimental evaluation of our algorithm using real and synthetic English words datasets for both the exact and fuzzy counting scenarios. Our privacy preserving mechanism substantially outperforms the prior work in terms of lower querying time, significantly higher utility (accuracy of count estimation) under similar or lower privacy guarantees, at the cost of communication overhead.
Dinusha Vatsalan, Raghav Bhaskar, Mohamed Ali Kâafar
IEEE Trans. Knowl. Data Eng.3
2019 The Chain of Implicit Trust: An Analysis of the Web Third-party Resources Loading
abstract
The Web is a tangled mass of interconnected services, where websites import a range of external resources from various third-party domains. The latter can also load resources hosted on other domains. For each website, this creates a dependency chain underpinned by a form of implicit trust between the first-party and transitively connected third-parties. The chain can only be loosely controlled as first-party websites often have little, if any, visibility on where these resources are loaded from. This paper performs a large-scale study of dependency chains in the Web, to find that around 50% of first-party websites render content that they did not directly load. Although the majority (84.91%) of websites have short dependency chains (below 3 levels), we find websites with dependency chains exceeding 30. Using VirusTotal, we show that 1.2% of these third-parties are classified as suspicious - although seemingly small, this limited set of suspicious third-parties have remarkable reach into the wider ecosystem.
Muhammad Ikram 0001, Rahat Masood, Gareth Tyson, Mohamed Ali Kâafar, Noha Loizon, Roya Ensafi
WWW4
2018 Incognito: A Method for Obfuscating Web Data
abstract
Users leave a trail of their personal data, interests, and intents while surfing or sharing information on the Web. Web data could therefore reveal some private/sensitive information about users based on inference analysis. The possible identification of information corresponding to a single individual by an inference attack holds true even if the user identifiers are encoded or removed in the Web data. Several works have been done on improving privacy of Web data through obfuscation methods~\citeHow09,Dom09,Sha05,Che14. However, these methods are neither comprehensive, generic to be applicable to any Web data, nor effective against adversarial attacks. To this end, we propose a privacy-aware obfuscation method for Web data addressing these identified drawbacks of existing methods. We use probabilistic methods to predict privacy risk of Web data that incorporates all key privacy aspects, which are uniqueness, uniformity, and linkability of Web data. The Web data with high predicted risk are then obfuscated by our method to minimize the privacy risk using semantically similar data. Our method is resistant against adversary who has knowledge about the datasets and model learned risk probabilities using differential privacy-based noise addition. Experimental study conducted on two real Web datasets validates the significance and efficacy of our method. Our results indicate that the average privacy risk reaches to 100% with a minimum of 10 sensitive Web entries, while at most 0% privacy risk could be attained with our obfuscation method at the cost of average utility loss of 64.3%.
Rahat Masood, Dinusha Vatsalan, Muhammad Ikram 0001, Mohamed Ali Kâafar
WWW4
2017 Spam Mobile Apps: Characteristics, Detection, and in the Wild Analysis
abstract
The increased popularity of smartphones has attracted a large number of developers to offer various applications for the different smartphone platforms via the respective app markets. One consequence of this popularity is that the app markets are also becoming populated with spam apps. These spam apps reduce the users’ quality of experience and increase the workload of app market operators to identify these apps and remove them. Spam apps can come in many forms such as apps not having a specific functionality, those having unrelated app descriptions or unrelated keywords, or similar apps being made available several times and across diverse categories. Market operators maintain antispam policies and apps are removed through continuous monitoring. Through a systematic crawl of a popular app market and by identifying apps that were removed over a period of time, we propose a method to detect spam apps solely using app metadata available at the time of publication. We first propose a methodology to manually label a sample of removed apps, according to a set of checkpoint heuristics that reveal the reasons behind removal. This analysis suggests that approximately 35% of the apps being removed are very likely to be spam apps. We then map the identified heuristics to several quantifiable features and show how distinguishing these features are for spam apps. We build an Adaptive Boost classifier for early identification of spam apps using only the metadata of the apps. Our classifier achieves an accuracy of over 95% with precision varying between 85% and 95% and recall varying between 38% and 98%. We further show that a limited number of features, in the range of 10--30, generated from app metadata is sufficient to achieve a satisfactory level of performance. On a set of 180,627 apps that were present at the app market during our crawl, our classifier predicts 2.7% of the apps as potential spam. Finally, we perform additional manual verification and show that human reviewers agree with 82% of our classifier predictions.
Suranga Seneviratne, Aruna Seneviratne, Mohamed Ali Kâafar, Anirban Mahanti, Prasant Mohapatra
ACM Trans. Web3
2015 Characterizing and Predicting Viral-and-Popular Video Content
abstract
The proliferation of online video content has triggered numerous works on its evolution and popularity, as well as on the effect of social sharing on content propagation. In this paper, we focus on the observable dependencies between the virality of video content on a micro-blogging social network (in this case, Twitter) and the popularity of such content on a video distribution service (YouTube). To this end, we collected and analysed a corpus of Twitter posts containing links to YouTube clips and the corresponding video meta-data from YouTube. Our analysis highlights the unique properties of content that is both popular and viral, which allows such content to attract high number of views on YouTube and achieve fast propagation on Twitter. With this in mind, we proceed to the predictions of popular-and-viral clips and propose a framework that can, with high degree of accuracy and low amount of training data, predict videos that are likely to be popular, viral, and both. The key contribution of our work is the focus on cross-system dynamics between YouTube and Twitter. We conjecture and validate that cross-system prediction of both popularity and virality of videos is feasible, and can be performed with a reasonably high degree of accuracy. One of our key findings is that YouTube features capturing user engagement, have strong virality prediction capabilities. This findings allows to solely rely on data extracted from a video sharing service to predict popularity and virality aspects of videos.
David Vallet, Shlomo Berkovsky, Sebastien Ardon, Anirban Mahanti, Mohamed Ali Kâafar
CIKM5
2015 Applying Differential Privacy to Matrix Factorization
abstract
Recommender systems are increasingly becoming an integral part of on-line services. As the recommendations rely on personal user information, there is an inherent loss of privacy resulting from the use of such systems. While several works studied privacy-enhanced neighborhood-based recommendations, little attention has been paid to privacy preserving latent factor models, like those represented by matrix factorization techniques. In this paper, we address the problem of privacy preserving matrix factorization by utilizing differential privacy, a rigorous and provable privacy preserving method. We propose and study several approaches for applying differential privacy to matrix factorization, and evaluate the privacy-accuracy trade-offs offered by each approach. We show that input perturbation yields the best recommendation accuracy, while guaranteeing a solid level of privacy protection.
Arnaud Berlioz, Arik Friedman, Mohamed Ali Kâafar, Roksana Boreli, Shlomo Berkovsky
RecSys3
2015 Early Detection of Spam Mobile Apps
abstract
Increased popularity of smartphones has attracted a large number of developers to various smartphone platforms. As a result, app markets are also populated with spam apps, which reduce the users' quality of experience and increase the workload of app market operators. Apps can be "spammy" in multiple ways including not having a specific functionality, unrelated app description or unrelated keywords and publishing similar apps several times and across diverse categories. Market operators maintain anti-spam policies and apps are removed through continuous human intervention. Through a systematic crawl of a popular app market and by identifying a set of removed apps, we propose a method to detect spam apps solely using app metadata available at the time of publication. We first propose a methodology to manually label a sample of removed apps, according to a set of checkpoint heuristics that reveal the reasons behind removal. This analysis suggests that approximately 35% of the apps being removed are very likely to be spam apps. We then map the identified heuristics to several quantifiable features and show how distinguishing these features are for spam apps. Finally, we build an Adaptive Boost classifier for early identification of spam apps using only the metadata of the apps. Our classifier achieves an accuracy over 95% with precision varying between 85%-95% and recall varying between 38%-98%. By applying the classifier on a set of apps present at the app market during our crawl, we estimate that at least 2.7% of them are spam apps.
Suranga Seneviratne, Aruna Seneviratne, Mohamed Ali Kâafar, Anirban Mahanti, Prasant Mohapatra
WWW3
2013 The Where and When of Finding New Friends: Analysis of a Location-based Social Discovery Network
Terence Chen, Mohamed Ali Kâafar, Roksana Boreli
ICWSM2
2013 Cross social networks interests predictions based ongraph features
abstract
The tremendous popularity of Online Social Networks (OSN) has led to situations, where users have their profiles spread across multiple networks. These partial profiles reflect different user characteristics, depending mainly on the nature of the network, e.g., Facebook's social vs. LinkedIn's professional focus. Combining data gathered by multiple networks may benefit individual users, and the community as a whole, as this could facilitate the provision of more accurate services and recommendations. This paper reports on an exploratory study of the process of making such recommendations using a unique multi-network dataset containing user interests across multiple domains, e.g., music, books, and movies. We represent the data using a graph model and generate recommendations using a set of features extracted from and populated by the model. We assess the contribution of various network- and domain-related features to the accuracy of the recommendations and motivate future work into automated feature selection.
Amit Tiroshi, Shlomo Berkovsky, Mohamed Ali Kâafar, Terence Chen, Tsvi Kuflik
RecSys3