VLDB 2026 Research / reviewers in the wild / expert
Zubair Shafiq
dblp:83/9528 · also M. Zubair Shafiq, Muhammad Zubair Shafiq
· DBLP profile ↗
112ranked-venue papers
19as first author
43since 2021 · last 2026
0000-0002-4500-9354ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 48 · 2 first-author · 31 since 2021Computer networks · 30 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 3 since 2021Systems, architecture and hardware · 10 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 10 · 1 since 2021Software engineering, systems software and programming languages · 8 · 5 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Multi-Stakeholder Vulnerability Notifications in the Ad-Tech Supply ChainabstractOnline advertising relies on a complex and opaque supply chain that involves multiple stakeholders, including advertisers, publishers, and ad-networks, each with distinct and sometimes conflicting incentives. Recent research has demonstrated the existence of ad-tech supply chain vulnerabilities such as dark pooling, where low-quality publishers bundle their ad inventory with higher-quality ones to mislead advertisers. We investigate the effectiveness of vulnerability notification campaigns aimed at mitigating dark pooling. Prior research on vulnerability notifications have primarily explored single-stakeholder contexts, leaving multi-stakeholder scenarios understudied. There is limited attention to complex multi-stakeholder supply chain ecosystems such as ad-tech supply chain, where resolving vulnerabilities often requires coordinated action across entities with misaligned incentives and interdependent roles. We address this gap by implementing the first online advertising supply chain vulnerability notification pipeline to systematically evaluate the responsiveness of various stakeholders in ad-tech supply chain, including publishers, ad-networks, and advertisers to vulnerability notifications by academics and activists. Our nine-month long automated multi-stakeholder notification study shows that notifications are an effective method for reducing dark pooling vulnerabilities in the online advertising ecosystem, especially when targeted towards ad-networks. Further, the sender reputation does not impact responses to notifications from activists and academics in a statistically different way. Overall, our research fosters industry-scale solution to combat ad inventory fraud and fosters future research on feasibility of multi-stakeholder vulnerability notifications in other supply chain ecosystems. Yash Vekaria, Rishab Nithyanand, Zubair Shafiq |
EuroS&P | 3 |
| 2026 | Understanding Data Collection, Brokerage, and Spam in the Lead Marketing EcosystemabstractThe lead marketing ecosystem enables collection, sale, and use of personal data submitted via web forms to deliver personalized quotes in high-value verticals such as insurance. Despite its scale and sensitivity of the collected data, this ecosystem remains largely unexplored by the research community. We present the first empirical study of privacy and spam risks in lead marketing, developing an end-toend measurement framework to trace data flows from data collection to consumer contact. Our setup instruments over 100 health-related lead-generation websites and monitors 200 controlled phone numbers and email addresses to understand downstream marketing practices. We observe sharing of highly personal and sensitive health information to more than 70 distinct third parties on these lead generation websites. By purchasing our own and other organic leads from three major lead platforms, we uncover deceptive brokerage practices, where consumer data is sold to unvetted buyers and often augmented or fabricated with attributes such as health status and weight. We received a total of over 8,000 telemarketing phone calls, 600 text messages, and 200 emails, where calls often began within seconds of form submission. Many campaigns relied on VoIP-based neighbor spoofing and high-frequency dialing, at times rendering phones unusable. Our experiments with phone and email opt-outs suggest phone-based opt-outs to help the most, although all were ineffective at completely stopping marketing communications. Analysis of 7,432 Better Business Bureau (BBB) complaints and reviews corroborates these findings from the consumer perspective. Overall, our results reveal a highly interconnected and non-compliant lead marketing ecosystem that aggressively monetizes sensitive consumer data. Yash Vekaria, Nurullah Demir, Konrad Kollnig, Zubair Shafiq |
SP | 4 |
| 2025 | FP-Rowhammer: DRAM-Based Device Fingerprinting
Hari Venugopalan, Kaustav Goswami 0002, Zain ul Abi Din, Jason Lowe-Power, Samuel T. King, Zubair Shafiq |
AsiaCCS | 6 |
| 2025 | Towards Characterizing and Detecting Incentivized Reviews on eCommerce PlatformsabstractCustomer reviews play an important role in rankings and visibility on e-commerce sites, and also strongly influence a customer's decision to purchase a product. Motivated by this, malicious sellers engage in incentivized review fraud to inflate their product ratings by providing customers with free products in exchange for five-star reviews, thus compromising review integrity. While there is ample prior work on fake reviews in general, there is limited prior work on incentivized review fraud. In this work, we infiltrate an underground market for fake reviews and implement a custom crawler to collect a dataset of malicious products that seek incentivized reviews. We devise and extract a set of features, and show that these are statistically significant in differentiating between benign and malicious products. Using hypothesis testing, we identify characteristics and trends exhibited by malicious products. While we are unable to achieve a high precision when we train standard machine learning models without compromising on the recall, we propose a lightweight two-phase technique that combines high-precision product classifiers with high-recall review classifiers. This technique allows us to minimize the false positives, with only a slight increase in false negatives. Finally, we also audit the effectiveness of two publicly available tools for incentivized review detection and find that they are not reliable. In summary, we contribute a new high-fidelity dataset, characterize products seeking incentivized reviews, audit existing tools for review analysis, and present a superior method for detecting review fraud. We hope that this research could be useful for e-commerce companies and other entities who have a stake in preserving opinion and review integrity online. Rajvardhan Oak, Zubair Shafiq |
ICWSM | 2 |
| 2025 | $CookieGuard: $ Characterizing and Isolating the First-Party Cookie JarabstractAs third-party cookies are being phased out or restricted by major browsers, first-party cookies are increasingly being used for web tracking. Prior work has shown that third-party scripts embedded in the main frame can access and exfiltrate first-party cookies—including those set by other third-party scripts. However, existing browser security mechanisms such as the Same-Origin Policy (SOP), Content Security Policy (CSP), and third-party storage partitioning do not prevent cross-domain access to first-party cookies in the main frame. While recent studies have begun to highlight this issue, there remains a lack of comprehensive measurement and practical defenses. In this work, we conduct the first large-scale measurement and analysis of cross-domain access to first-party cookies in the main frame for 20,000 websites. We find that 56% of the websites include third-party scripts that exfiltrate first-party cookies that they did not originally set, and 32% where such scripts overwrite or delete these first-party cookies. To mitigate potential confidentiality and integrity risks due to this lack of isolation, we propose CookieGuard, a browser-based runtime mechanism to isolate first-party cookies on a per-script-origin basis. CookieGuard blocks unauthorized cross-domain cookie operations while preserving site functionality, with only 3% of the tested websites being affected by Single Sign-On (SSO) breakage. Our work highlights the risks posed by the lack of first-party cookie isolation in the current browser security model and offers a deployable path toward stronger protection. Pouneh Nikkhah Bahrami, Aurore Fass, Zubair Shafiq |
IMC | 3 |
| 2025 | From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice CharacteristicsabstractMany devices are accessed and controlled through voice assistants today, a representative example being Echo smart speakers and other Amazon devices controlled by Alexa. These offer the convenience of accessing services through voice interactions, but also raise privacy concerns, as data can be stored and used for personalization, and voice biometric information is sensitive. Unfortunately, there remains a lack of transparency and control over the collection and use of this data. Although prior work has shown evidence of ad targeting based on data derived from voice interactions and user profiles/interests, it has so far been an open question whether voice biometric information itself is utilized for targeting. In this paper, (i) we build a general auditing methodology to answer this question for off-the-shelf commercial smart speakers, and (ii) we apply it specifically to Amazon Echo Dot. Our findings suggest that Amazon Music ad content is more strongly associated with attributes (gender and age) related to voice characteristics than would be expected by chance. This has important implications for compliance, since voice contains sensitive biometric information that is protected by several privacy regulations. Tu Le, Luca Baldesi, Athina Markopoulou, Carter T. Butts, Zubair Shafiq |
IMC | 5 |
| 2025 | FP-Inconsistent: Measurement and Analysis of Fingerprint Inconsistencies in Evasive Bot TrafficabstractBrowser fingerprinting is used for bot detection. In response, bots have started altering their fingerprints to evade detection. We conduct the first large-scale evaluation to study whether and how altering fingerprints helps bots evade detection. To systematically investigate such evasive bots, we deploy a honey site that includes two anti-bot services (DataDome and BotD) and solicit bot traffic from 20 different bot services that purport to sell ''realistic and undetectable traffic.'' Across half a million requests recorded on our honey site, we find an average evasion rate of 52.93% against DataDome and 44.56% evasion rate against BotD. Our analysis of fingerprint attributes of evasive bots shows that they indeed alter their fingerprints. Moreover, we find that the attributes of these altered fingerprints are often inconsistent with each other. We propose FP-Inconsistent, a data-driven approach to detect such inconsistencies across space (two attributes in a given browser fingerprint) and time (a single attribute at two different points in time). Our evaluation shows that our approach can reduce the evasion rate of evasive bots by 44.95%-48.11% while maintaining a true negative rate of 96.84% on traffic from real users. Hari Venugopalan, Shaoor Munir, S. Shuaib Ahmed, Tangbaihe Wang, Samuel T. King, Zubair Shafiq |
IMC | 6 |
| 2025 | "Hello, is this Anna?": Unpacking the Lifecycle of Pig-Butchering Scams
Rajvardhan Oak, Zubair Shafiq |
SOUPS | 2 |
| 2025 | Victims, Vigilantes, and Advice Givers: An Analysis of Scam-Related Discourse on Reddit
Rajvardhan Oak, Zubair Shafiq |
SOUPS | 2 |
| 2025 | Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative AI Assistants
Yash Vekaria, Aurelio Loris Canino, Jonathan Levitsky, Alex Ciechonski, Patricia Callejo, Anna Maria Mandalari, Zubair Shafiq |
USENIX Security Symposium | 7 |
| 2025 | Editors' IntroductionabstractEditors' Introduction, Issue 1 of PETS Volume 2025 Rob Jansen, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 2 |
| 2025 | Editors' IntroductionabstractEditors' Introduction, Issue 2 of PETS Volume 2025 Rob Jansen, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 2 |
| 2025 | Editors' IntroductionabstractEditors' Introduction, Issue 3 of PETS Volume 2025 Rob Jansen, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 2 |
| 2025 | Editors' IntroductionabstractEditors' Introduction, Issue 4 of PETS Volume 2025 Rob Jansen, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 2 |
| 2025 | AutoFR: Automated Filter Rule Generation for AdblockingabstractAdblocking relies on filter lists, which are manually curated and maintained by a community of filter list authors. Filter list curation is a laborious process that does not scale well to a large number of sites or over time. In this article, we introduce AutoFR, a reinforcement learning framework to fully automate the process of filter rule creation and evaluation for sites of interest. We design an algorithm based on multi-arm bandits to generate filter rules that block ads while controlling the trade-off between blocking ads and avoiding visual breakage. We test AutoFR on thousands of sites and show that it is efficient: It takes only a few minutes to generate filter rules for a site of interest. AutoFR is effective: It optimizes filter rules for a particular site that can block 86% of the ads, as compared to 87% by EasyList, while achieving comparable visual breakage. Using AutoFR as a building block, we devise three methodologies that generate filter rules across sites based on: (1) a modified version of AutoFR, (2) rule popularity, and (3) site similarity. We conduct an in-depth comparative analysis of these approaches by considering their effectiveness, efficiency, and maintainability. We demonstrate that some of them can generalize well to new sites in both controlled and live settings. We envision that AutoFR can assist the adblocking community in automatically generating and updating filter rules at scale. Hieu Le 0003, Salma Hosni Emam Mohamed Elmalaki, Athina Markopoulou, Zubair Shafiq |
ACM Trans. Priv. Secur. | 4 |
| 2024 | Blocking Tracking JavaScript at the Function GranularityabstractModern websites extensively rely on JavaScript to implement both functionality and tracking. Existing privacy-enhancing content blocking tools struggle against mixed scripts, which simultaneously implement both functionality and tracking. Blocking such scripts would break functionality, and not blocking themwould allowtracking. We propose NoT.js, a fine-grained JavaScript blocking tool that operates at the function-level granularity. NoT.js’s strengths lie in analyzing the dynamic execution context, including the call stack and calling context of each JavaScript function, and then encoding this context to build a rich graph representation. NoT.js trains a supervised machine learning classifier on a webpage’s graph representation to first detect tracking at the function-level and then automatically generates surrogate scripts that preserve functionality while removing tracking. Our evaluation of NoT.js on the top-10K websites demonstrates that it achieves high precision (94%) and recall (98%) in detecting tracking functions, outperforming the state-of-the-art while being robust against off-the-shelf JavaScript obfuscation. Fine-grained detection of tracking functions allows NoT.js to automatically generate surrogate scripts, which our evaluation shows that successfully remove tracking functions without causing major breakage. Our deployment of NoT.js shows that mixed scripts are present on 62.3% of the top-10K websites, with 70.6% of the mixed scripts being third-party that engage in tracking activities such as cookie ghostwriting. Abdul Haddi Amjad, Shaoor Munir, Zubair Shafiq, Muhammad Ali Gulzar |
CCS | 3 |
| 2024 | Understanding Underground Incentivized Review ServicesabstractWhile human factors in fraud have been studied by the HCI and security communities, most research has been directed to understanding either the victims’ perspectives or prevention strategies, and not on fraudsters, their motivations and operation techniques. Additionally, the focus has been on a narrow set of problems: phishing, spam and bullying. In this work, we seek to understand review fraud on e-commerce platforms through an HCI lens. Through surveys with real fraudsters (N=36 agents and N=38 reviewers), we uncover sophisticated recruitment, execution, and reporting mechanisms fraudsters use to scale their operation while resisting takedown attempts, including the use of AI tools like ChatGPT. We find that countermeasures that crack down on communication channels through which these services operate are effective in combating incentivized reviews. This research sheds light on the complex landscape of incentivized reviews, providing insights into the mechanics of underground services and their resilience to removal efforts. Rajvardhan Oak, Zubair Shafiq |
CHI | 2 |
| 2024 | Watching TV with the Second-Party: A First Look at Automatic Content Recognition Tracking in Smart TVsabstractSmart TVs implement a unique tracking approach called Automatic Content Recognition (ACR) to profile viewing activity of their users. ACR is a Shazam-like technology that works by periodically capturing the content displayed on a TV's screen and matching it against a content library to detect what content is being displayed at any given point in time. While prior research has investigated third-party tracking in the smart TV ecosystem, it has not looked into second-party ACR tracking that is directly conducted by the smart TV platform. In this work, we conduct a black-box audit of ACR network traffic between ACR clients on the smart TV and ACR servers. We use our auditing approach to systematically investigate whether (1) ACR tracking is agnostic to how a user watches TV (e.g., linear vs. streaming vs. HDMI), (2) privacy controls offered by smart TVs have an impact on ACR tracking, and (3) there are any differences in ACR tracking between the UK and the US. We perform a series of experiments on two major smart TV platforms: Samsung and LG. Our results show that ACR works even when the smart TV is used as a ''dumb'' external display, opting-out stops network traffic to ACR servers, and there are differences in how ACR works across the UK and the US. Gianluca Anselmi, Yash Vekaria, Alexander D'Souza, Patricia Callejo, Anna Maria Mandalari, Zubair Shafiq |
IMC | 6 |
| 2024 | The Inventory is Dark and Full of Misinformation: Understanding Ad Inventory Pooling in the Ad-Tech Supply ChainabstractAd-tech enables publishers to programmatically sell their ad inventory to millions of demand partners through a complex supply chain. The complexity and opacity of the ad-tech supply chain can be exploited by low-quality publishers (e.g., misinformation websites) to deceptively monetize their ad inventory. To combat such deception, the ad-tech industry has developed transparency standards and brand safety products. In this paper, we show that these developments still fall short of preventing deceptive monetization. Specifically, we focus on how publishers can exploit the ad-tech supply chain, subvert ad-tech transparency standards, and undermine brand safety protections by pooling their ad inventory with unrelated sites. This type of deception is referred to as "dark pooling." Our study shows that dark pooling is commonly employed by misinformation publishers on various major ad exchanges, and allows misinformation publishers to deceptively sell their ad inventory to reputable brands. Our work suggests the need for improved vetting of ad exchange supply partners, the adoption of new ad-tech transparency standards that enable end-to-end validation of the ad-tech supply chain, and the widespread deployment of independent audits like ours. Yash Vekaria, Rishab Nithyanand, Zubair Shafiq |
SP | 3 |
| 2024 | PURL: Safe and Effective Sanitization of Link Decoration
Shaoor Munir, Patrick Lee, Umar Iqbal 0002, Sandra Deepthy Siby, Zubair Shafiq |
USENIX Security Symposium | 5 |
| 2024 | Editors' IntroductionabstractFace images are a rich source of information that can be used to identify individuals and infer private information about them.To mitigate this privacy risk, anonymizations employ transformations on clear images to obfuscate sensitive information, all while retaining some utility.Albeit published with impressive claims, they sometimes are not evaluated with convincing methodology.Reversing anonymized images to resemble their real input -and even be identified by face recognition approaches -represents the strongest indicator for flawed anonymization.Some recent results indeed indicate that this is possible for some approaches.It is, however, not well understood, which approaches are reversible, and why.In this paper, we provide an exhaustive investigation in the phenomenon of face anonymization reversibility.Among other things, we find that 11 out of 15 tested face anonymizations are at least partially reversible and highlight how both reconstruction and inversion are the underlying processes that make reversal possible. Micah Sherr, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 2 |
| 2024 | Editors' IntroductionabstractIn this model, articles are published throughout the year at regular intervals, and the papers for the year are then presented at an annual conference.Reviewers can request revisions of submitted articles, which may then be revised and resubmitted in the same year.PoPETs publishes four issues per year.By enabling resubmission across these issues, PoPETs provides a high-quality peer-review process that enables authors and reviewers to work together to produce and recognize significant scholarly contributions.The PoPETs double-blind peer-review process is similar to other top-tier computer-security publications.The process includes initial review by the Editors-in-Chief for rules compliance and in-scope content, written reviews by multiple independent reviewers, author rebuttal, discussion among reviewers, and consensus decisions with disagreements resolved by the Editors-in-Chief or the Vice Chairs.The output of the review process is a set of reviews, a meta-review summarizing the reviewers' opinions after discussion (for papers that are not rejected during the first round), and one of the following decisions: Accept, Accept with Minor Micah Sherr, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 2 |
| 2024 | Editors' IntroductionabstractEditors' Introduction, Issue 3 of PETS Volume 2024 Micah Sherr, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 2 |
| 2024 | Editors' IntroductionabstractEditors' Introduction, Issue 4 of PETS Volume 2024 Micah Sherr, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 2 |
| 2023 | CookieGraph: Understanding and Detecting First-Party Tracking CookiesabstractAs third-party cookie blocking is becoming the norm in mainstream web browsers, advertisers and trackers have started to use first-party cookies for tracking. To understand this phenomenon, we conduct a differential measurement study with versus without third-party cookies. We find that first-party cookies are used to store and exfiltrate identifiers to known trackers even when third-party cookies are blocked. Shaoor Munir, Sandra Deepthy Siby, Umar Iqbal 0002, Steven Englehardt, Zubair Shafiq, Carmela Troncoso |
CCS | 5 |
| 2023 | Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker EcosystemabstractSmart speakers collect voice commands, which can be used to infer sensitive information about users. Given the potential for privacy harms, there is a need for greater transparency and control over the data collected, used, and shared by smart speaker platforms as well as third party skills supported on them. To bridge this gap, we build a framework to measure data collection, usage, and sharing by the smart speaker platforms. We apply our framework to the Amazon smart speaker ecosystem. Our results show that Amazon and third parties, including advertising and tracking services that are unique to the smart speaker ecosystem, collect smart speaker interaction data. We also find that Amazon processes smart speaker interaction data to infer user interests and uses those inferences to serve targeted ads to users. Smart speaker interaction also leads to ad targeting and as much as 30X higher bids in ad auctions, from third party advertisers. Finally, we find that Amazon's and third party skills' data practices are often not clearly disclosed in their policy documents. Umar Iqbal 0002, Pouneh Nikkhah Bahrami, Rahmadi Trimananda, Hao Cui 0004, Alexander Gamero-Garrido, Daniel J. Dubois, David R. Choffnes, Athina Markopoulou, Franziska Roesner, Zubair Shafiq |
IMC | 10 |
| 2023 | Accuracy-Privacy Trade-off in Deep Ensemble: A Membership Inference PerspectiveabstractDeep ensemble learning has been shown to improve accuracy by training multiple neural networks and averaging their outputs. Ensemble learning has also been suggested to defend against membership inference attacks that undermine privacy. In this paper, we empirically demonstrate a trade-off between these two goals, namely accuracy and privacy (in terms of membership inference attacks), in deep ensembles. Using a wide range of datasets and model architectures, we show that the effectiveness of membership inference attacks increases when ensembling improves accuracy. We analyze the impact of various factors in deep ensembles and demonstrate the root cause of the trade-off. Then, we evaluate common defenses against membership inference attacks based on regularization and differential privacy. We show that while these defenses can mitigate the effectiveness of membership inference attacks, they simultaneously degrade ensemble accuracy. We illustrate similar trade-off in more advanced and state-of-the-art ensembling techniques, such as snapshot ensembles and diversified ensemble networks. Finally, we propose a simple yet effective defense for deep ensembles to break the trade-off and, consequently, improve the accuracy and privacy, simultaneously. Shahbaz Rezaei, Zubair Shafiq, Xin Liu 0002 |
SP | 2 |
| 2023 | AutoFR: Automated Filter Rule Generation for Adblocking
Hieu Le 0003, Salma Hosni Emam Mohamed Elmalaki, Athina Markopoulou, Zubair Shafiq |
USENIX Security Symposium | 4 |
| 2023 | Blocking JavaScript Without Breaking the Web: An Empirical InvestigationabstractModern websites heavily rely on JavaScript (JS) to implement legitimate functionality as well as privacy-invasive advertising and tracking. Browser extensions such as NoScript block any script not loaded by a trusted list of endpoints, thus hoping to block privacy-invasive scripts while avoiding breaking legitimate website functionality. In this paper, we investigate whether blocking JS on the web is feasible without breaking legitimate functionality. To this end, we conduct a large-scale measurement study of JS blocking on 100K websites. We evaluate the effectiveness of different JS blocking strategies in tracking prevention and functionality breakage. Our evaluation relies on quantitative analysis of network requests and resource loads as well as manual qualitative analysis of visual breakage. First, we show that while blocking all scripts is quite effective at reducing tracking, it significantly degrades functionality on approximately two-thirds of the tested websites. Second, we show that selective blocking of a subset of scripts based on a curated list achieves a better trade-off. However, there remain approximately 15% “mixed” scripts, which essentially merge tracking and legitimate functionality and thus cannot be blocked without causing website breakage. Finally, we show that fine-grained blocking of a subset of JS methods, instead of scripts, reduces major breakage by 3.8× while providing the same level of tracking prevention. Our work highlights the promise and open challenges in fine-grained JS blocking for tracking prevention without breaking the web. Abdul Haddi Amjad, Zubair Shafiq, Muhammad Ali Gulzar |
Proc. Priv. Enhancing Technol. | 2 |
| 2023 | A Utility-Preserving Obfuscation Approach for YouTube RecommendationsabstractOnline content platforms optimize engagement by providing personalized recommendations to their users. These recommendation systems track and profile users to predict relevant content a user is likely interested in. While the personalized recommendations provide utility to users, the tracking and profiling that enables them poses a privacy issue because the platform might infer potentially sensitive user interests. There is increasing interest in building privacy-enhancing obfuscation approaches that do not rely on cooperation from online content platforms. However, existing obfuscation approaches primarily focus on enhancing privacy but at the same time they degrade the utility because obfuscation introduces unrelated recommendations. We design and implement DeHarpo, an obfuscation approach for YouTube's recommendation system that not only obfuscates a user's video watch history to protect privacy but then also denoises the video recommendations by YouTube to preserve their utility. In contrast to prior obfuscation approaches, DeHarpo adds a denoiser that makes use of a ``secret'' input (i.e., a user's actual watch history) as well as information that is also available to the adversarial recommendation system (i.e., obfuscated watch history and corresponding ``nois`` recommendations). Our large-scale evaluation of DeHarpo shows that it outperforms the state-of-the-art by a factor of 2x in terms of preserving utility for the same level of privacy, while maintaining stealthiness and robustness to de-obfuscation. Jiang Zhang 0003, Hadi Askari, Konstantinos Psounis, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 4 |
| 2022 | On the Robustness of Offensive Language ClassifiersabstractSocial media platforms are deploying machine learning based offensive language classification systems to combat hateful, racist, and other forms of offensive speech at scale.However, despite their real-world deployment, we do not yet comprehensively understand the extent to which offensive language classifiers are robust against adversarial attacks.Prior work in this space is limited to studying robustness of offensive language classifiers against primitive attacks such as misspellings and extraneous spaces.To address this gap, we systematically analyze the robustness of state-of-theart offensive language classifiers against more crafty adversarial attacks that leverage greedyand attention-based word selection and contextaware embeddings for word replacement.Our results on multiple datasets show that these crafty adversarial attacks can degrade the accuracy of offensive language classifiers by more than 50% while also being able to preserve the readability and meaning of the modified text. Jonathan Rusert, Zubair Shafiq, Padmini Srinivasan |
ACL (1) | 2 |
| 2022 | Adversarial Authorship Attribution for DeobfuscationabstractRecent advances in natural language processing have enabled powerful privacy-invasive authorship attribution.To counter authorship attribution, researchers have proposed a variety of rule-based and learning-based text obfuscation approaches.However, existing authorship obfuscation approaches do not consider the adversarial threat model.Specifically, they are not evaluated against adversarially trained authorship attributors that are aware of potential obfuscation.To fill this gap, we investigate the problem of adversarial authorship attribution for deobfuscation.We show that adversarially trained authorship attributors are able to degrade the effectiveness of existing obfuscators from 20-30% to 5-10%.We also evaluate the effectiveness of adversarial training when the attributor makes incorrect assumptions about whether and which obfuscator was used.While there is a a clear degradation in attribution accuracy, it is noteworthy that this degradation is still at or above the attribution accuracy of the attributor that is not adversarially trained at all.Our results underline the need for stronger obfuscation approaches that are resistant to deobfuscation.* This paper is third in the series.See (Mahmood et al., 2019) and (Mahmood et al., 2020) for the first two papers. Wanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini Srinivasan |
ACL (1) | 3 |
| 2022 | Stealthy Inference Attack on DNN via Cache-based Side-Channel AttacksabstractThe advancement of deep neural networks (DNNs) motivates the deployment in various domains, including image classification, disease diagnoses, voice recognition, etc. Since some tasks that DNN undertakes are very sensitive, the label information is confidential and contains a commercial value or critical privacy. This paper demonstrates that DNNs also bring a new security threat, leading to the leakage of label information of input instances for the DNN models. In particular, we leverage the cache-based side-channel attack (SCA), i.e., Flush-Reload on the DNN (victim) models, to observe the execution of computation graphs, and create a database of them for building a classifier that the attacker can use to decide the label information of (unknown) input instances for victim models. Then we deploy the cache-based SCA on the same host machine with victim models and deduce the labels with the attacker's classification model to compromise the privacy and confidentiality of victim models. We explore different settings and classification techniques to achieve a high attack success rate of stealing label information from the victim models. Additionally, we consider two attacking scenarios: binary attacking identifies specific sensitive labels and others while multi-class attacking targets recognize all classes victim DNNs provide. Last, we implement the attack on both static DNN models with identical architectures for all inputs and dynamic DNN models with an adaptation of architectures for different inputs to demonstrate the vast existence of the proposed attack, including DenseNet 121, DenseNet 169, VGG 16, VGG 19, MobileNet v1, and MobileNet v2. Our experiment exhibits that MobileNet v1 is the most vulnerable one with 99% and 75.6% attacking success rates for binary and multi-class attacking scenarios, respectively. Han Wang 0020, Syed Mahbub Hafiz, Kartik Patwari, Chen-Nee Chuah, Zubair Shafiq, Houman Homayoun |
DATE | 5 |
| 2022 | DNN Model Architecture Fingerprinting Attack on CPU-GPU Edge DevicesabstractEmbedded systems for edge computing are getting more powerful, and some are equipped with a GPU to enable on-device deep neural network (DNN) learning tasks such as image classification and object detection. Such DNN-based applications frequently deal with sensitive user data, and their architectures are considered intellectual property to be protected. We investigate a potential avenue of fingerprinting attack to identify the (running) DNN model architecture family (out of state-of-the-art DNN categories) on CPU-GPU edge devices. We exploit a stealthy analysis of aggregate system-level side-channel information such as memory, CPU, and GPU usage available at the user-space level. To the best of our knowledge, this is the first attack of its kind that does not require physical access and/or sudo access to the victim device and only collects the system traces passively, as opposed to most of the existing reverse-engineering-based DNN model architecture extraction attacks. We perform feature selection analysis and supervised machine learning-based classification to detect the model architecture. With a combination of RAM, CPU, and GPU features and a Random Forest-based classifier, our proposed attack classifies a known DNN model into its model architecture family with 99% accuracy. Also, the introduced attack is so transferable that it can detect an unknown DNN model into the right DNN architecture category with 87.2% accuracy. Our rigorous feature analysis illustrates that memory usage (RAM) is a critical feature for such fingerprinting. Furthermore, we successfully replicate this attack on two different CPU-GPU platforms and observe similar experimental results that exhibit the capability of platform portability of the attack. Also, we investigate the robustness of the proposed attack to varying background noises and a modified DNN pipeline. Besides, we exhibit that the leakage of model architecture family information from this stealthy attack can strengthen an adversarial attack against a victim DNN model by 2×. Kartik Patwari, Syed Mahbub Hafiz, Han Wang 0020, Houman Homayoun, Zubair Shafiq, Chen-Nee Chuah |
EuroS&P | 5 |
| 2022 | HARPO: Learning to Subvert Online Behavioral Advertising
Jiang Zhang 0003, Konstantinos Psounis, Zubair Shafiq |
NDSS | 4 |
| 2022 | Khaleesi: Breaker of Advertising and Tracking Request Chains
Umar Iqbal 0002, Charlie Wolfe, Charles Nguyen, Steven Englehardt, Zubair Shafiq |
USENIX Security Symposium | 5 |
| 2022 | WebGraph: Capturing Advertising and Tracking Information Flows for Robust Blocking
Sandra Deepthy Siby, Umar Iqbal 0002, Steven Englehardt, Zubair Shafiq, Carmela Troncoso |
USENIX Security Symposium | 4 |
| 2022 | FP-Radar: Longitudinal Measurement and Early Detection of Browser FingerprintingabstractAbstract Browser fingerprinting is a stateless tracking technique that aims to combine information exposed by multiple different web APIs to create a unique identifier for tracking users across the web. Over the last decade, trackers have abused several existing and newly proposed web APIs to further enhance the browser fingerprint. Existing approaches are limited to detecting a specific fingerprinting technique(s) at a particular point in time. Thus, they are unable to systematically detect novel fingerprinting techniques that abuse different web APIs. In this paper, we propose FP-R adar , a machine learning approach that leverages longitudinal measurements of web API usage on top-100K websites over the last decade for early detection of new and evolving browser fingerprinting techniques. The results show that FP-R adar is able to early detect the abuse of newly introduced properties of already known (e.g., WebGL , Sensor ) and as well as previously unknown (e.g., Gamepad , Clipboard ) APIs for browser fingerprinting. To the best of our knowledge, FP-R adar is the first to detect the abuse of the Visibility API for ephemeral fingerprinting in the wild. Pouneh Nikkhah Bahrami, Umar Iqbal 0002, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 3 |
| 2021 | Eluding ML-based Adblockers With Actionable Adversarial ExamplesabstractOnline advertisers have been quite successful in circumventing traditional adblockers that rely on manually curated rules to detect ads. As a result, adblockers have started to use machine learning (ML) classifiers for more robust detection and blocking of ads. Among these, AdGraph which leverages rich contextual information to classify ads, is arguably, the state of the art ML-based adblocker. In this paper, we present a4, a tool that intelligently crafts adversarial ads to evade AdGraph. Unlike traditional adversarial examples in the computer vision domain that can perturb any pixels (i.e., unconstrained), adversarial ads generated by a4 are actionable in the sense that they preserve the application semantics of the web page. Through a series of experiments we show that a4 can bypass AdGraph about 81% of the time, which surpasses the state-of-the-art attack by a significant margin of 145.5%, with an overhead of <20% and perturbations that are visually imperceptible in the rendered webpage. We envision that a4’s framework can be used to potentially launch adversarial attacks against other ML-based web applications. Shitong Zhu, Zhongjie Wang 0002, Shasha Li 0001, Keyu Man, Umar Iqbal 0002, Zhiyun Qian, Kevin S. Chan, Srikanth V. Krishnamurthy, Zubair Shafiq, Yu Hao 0006, Guoren Li, Zheng Zhang 0058, Xiaochen Zou |
ACSAC | 10 |
| 2021 | Through the Looking Glass: Learning to Attribute Synthetic Text Generated by Language ModelsabstractShaoor Munir, Brishna Batool, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Shaoor Munir, Brishna Batool, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar |
EACL | 3 |
| 2021 | TrackerSift: untangling mixed tracking and functional web resourcesabstractTrackers have recently started to mix tracking and functional resources to circumvent privacy-enhancing content blocking tools. Such mixed web resources put content blockers in a bind: risk breaking legitimate functionality if they act and risk missing privacy-invasive advertising and tracking if they do not. In this paper, we propose TrackerSift to progressively classify and untangle mixed web resources (that combine tracking and legitimate functionality) at multiple granularities of analysis (domain, hostname, script, and method). Using TrackerSift, we conduct a large-scale measurement study of such mixed resources on 100K websites. We find that more than 17% domains, 48% hostnames, 6% scripts, and 9% methods observed in our crawls combine tracking and legitimate functionality. While mixed web resources are prevalent across all granularities, TrackerSift is able to attribute 98% of the script-initiated network requests to either tracking or functional resources at the finest method-level granularity. Our analysis shows that mixed resources at different granularities are typically served from CDNs or as in-lined and bundled scripts, and that blocking them indeed results in breakage of legitimate functionality. Our results highlight opportunities for finer-grained content blocking to remove mixed resources without breaking legitimate functionality. Abdul Haddi Amjad, Danial Saleem, Muhammad Ali Gulzar, Zubair Shafiq, Fareed Zaffar |
Internet Measurement Conference | 4 |
| 2021 | CV-Inspector: Towards Automating Detection of Adblock Circumvention
Hieu Le 0003, Athina Markopoulou, Zubair Shafiq |
NDSS | 3 |
| 2021 | Fingerprinting the Fingerprinters: Learning to Detect Browser Fingerprinting BehaviorsabstractBrowser fingerprinting is an invasive and opaque stateless tracking technique. Browser vendors, academics, and standards bodies have long struggled to provide meaningful protections against browser fingerprinting that are both accurate and do not degrade user experience. We propose FP-Inspector, a machine learning based syntactic-semantic approach to accurately detect browser fingerprinting. We show that FP-Inspector performs well, allowing us to detect 26% more fingerprinting scripts than the state-of-the-art. We show that an API-level fingerprinting countermeasure, built upon FP-Inspector, helps reduce website breakage by a factor of 2. We use FP-Inspector to perform a measurement study of browser fingerprinting on top-100K websites. We find that browser fingerprinting is now present on more than 10% of the top-100K websites and over a quarter of the top-10K websites. We also discover previously unreported uses of JavaScript APIs by fingerprinting scripts suggesting that they are looking to exploit APIs in new and unexpected ways. Umar Iqbal 0002, Steven Englehardt, Zubair Shafiq |
SP | 3 |
| 2020 | A Girl Has A Name: Detecting Authorship ObfuscationabstractAuthorship attribution aims to identify the author of a text based on the stylometric analysis.Authorship obfuscation, on the other hand, aims to protect against authorship attribution by modifying a text's style.In this paper, we evaluate the stealthiness of state-of-the-art authorship obfuscation methods under an adversarial threat model.An obfuscator is stealthy to the extent an adversary finds it challenging to detect whether or not a text modified by the obfuscator is obfuscated -a decision that is key to the adversary interested in authorship attribution.We show that the existing authorship obfuscation methods are not stealthy as their obfuscated texts can be identified with an average F1 score of 0.87.The reason for the lack of stealthiness is that these obfuscators degrade text smoothness, as ascertained by neural language models, in a detectable manner.Our results highlight the need to develop stealthy authorship obfuscation methods that can better protect the identity of an author seeking anonymity. Asad Mahmood, Zubair Shafiq, Padmini Srinivasan |
ACL | 2 |
| 2020 | Understanding Incentivized Mobile App Installs on Google Play Storeabstract"Incentivized" advertising platforms allow mobile app developers to acquire new users by directly paying users to install and engage with mobile apps (e.g., create an account, make in-app purchases). Incentivized installs are banned by the Apple App Store and discouraged by the Google Play Store because they can manipulate app store metrics (e.g., install counts, appearance in top charts). Yet, many organizations still offer incentivized install services for Android apps. In this paper, we present the first study to understand the ecosystem of incentivized mobile app install campaigns in Android and its broader ramifications through a series of measurements. We identify incentivized install campaigns that require users to install an app and perform in-app tasks targeting manipulation of a wide variety of user engagement metrics (e.g., daily active users, user session lengths) and revenue. Our results suggest that these artificially inflated metrics can be effective in improving app store metrics as well as helping mobile app developers to attract funding from venture capitalists. Our study also indicates lax enforcement of the Google Play Store's existing policies to prevent these behaviors. It further motivates the need for stricter policing of incentivized install campaigns. Our proposed measurements can also be leveraged by the Google Play Store to identify potential policy violations. Shehroze Farooqi, Álvaro Feal, Tobias Lauinger, Damon McCoy, Zubair Shafiq, Narseo Vallina-Rodriguez |
Internet Measurement Conference | 5 |
| 2020 | FlowTrace : A Framework for Active Bandwidth Measurements Using In-band Packet Trains
Ricky Mok, Zubair Shafiq |
PAM | 3 |
| 2020 | AdGraph: A Graph-Based Approach to Ad and Tracker BlockingabstractUser demand for blocking advertising and tracking online is large and growing. Existing tools, both deployed and described in research, have proven useful, but lack either the completeness or robustness needed for a general solution. Existing detection approaches generally focus on only one aspect of advertising or tracking (e.g. URL patterns, code structure), making existing approaches susceptible to evasion.In this work we present AdGraph, a novel graph-based machine learning approach for detecting advertising and tracking resources on the web. AdGraph differs from existing approaches by building a graph representation of the HTML structure, network requests, and JavaScript behavior of a webpage, and using this unique representation to train a classifier for identifying advertising and tracking resources. Because AdGraph considers many aspects of the context a network request takes place in, it is less susceptible to the single-factor evasion techniques that flummox existing approaches.We evaluate AdGraph on the Alexa top-10K websites, and find that it is highly accurate, able to replicate the labels of human-generated filter lists with 95.33% accuracy, and can even identify many mistakes in filter lists. We implement AdGraph as a modification to Chromium. AdGraph adds only minor overhead to page loading and execution, and is actually faster than stock Chromium on 42% of websites and AdBlock Plus on 78% of websites. Overall, we conclude that AdGraph is both accurate enough and performant enough for online use, breaking comparable or fewer websites than popular filter list based approaches. Umar Iqbal 0002, Peter Snyder, Shitong Zhu, Benjamin Livshits, Zhiyun Qian, Zubair Shafiq |
SP | 6 |
| 2020 | Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header BiddingabstractAbstract Online advertising relies on trackers and data brokers to show targeted ads to users. To improve targeting, different entities in the intricately interwoven online advertising and tracking ecosystems are incentivized to share information with each other through client-side or server-side mechanisms. Inferring data sharing between entities, especially when it happens at the server-side, is an important and challenging research problem. In this paper, we introduce Kashf: a novel method to infer data sharing relationships between advertisers and trackers by studying how an advertiser’s bidding behavior changes as we manipulate the presence of trackers. We operationalize this insight by training an interpretable machine learning model that uses the presence of trackers as features to predict the bidding behavior of an advertiser. By analyzing the machine learning model, we can infer relationships between advertisers and trackers irrespective of whether data sharing occurs at the client-side or the server-side. We are able to identify several server-side data sharing relationships that are validated externally but are not detected by client-side cookie syncing. John Cook, Rishab Nithyanand, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 3 |
| 2020 | CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social NetworksabstractOnline social networks support a vibrant ecosystem of third-party apps that get access to personal information of a large number of users. Despite several recent high-profile incidents, methods to systematically detect data misuse by third-party apps on online social networks are lacking. We propose CanaryTrap to detect misuse of data shared with third-party apps. CanaryTrap associates a honeytoken to a user account and then monitors its unrecognized use via different channels after sharing it with the third-party app. We design and implement CanaryTrap to investigate misuse of data shared with third-party apps on Facebook. Specifically, we share the email address associated with a Facebook account as a honeytoken by installing a third-party app. We then monitor the received emails and use Facebook’s ad transparency tool to detect any unrecognized use of the shared honeytoken. Our deployment of CanaryTrap to monitor 1,024 Facebook apps has uncovered multiple cases of misuse of data shared with third-party apps on Facebook including ransomware, spam, and targeted advertising. Shehroze Farooqi, Maaz Bin Musa, Zubair Shafiq, Fareed Zaffar |
Proc. Priv. Enhancing Technol. | 3 |
| 2020 | The TV is Smart and Full of Trackers: Measuring Smart TV Advertising and TrackingabstractAbstract In this paper, we present a large-scale measurement study of the smart TV advertising and tracking ecosystem. First, we illuminate the network behavior of smart TVs as used in the wild by analyzing network traffic collected from residential gateways. We find that smart TVs connect to well-known and platform-specific advertising and tracking services (ATSes). Second, we design and implement software tools that systematically explore and collect traffic from the top-1000 apps on two popular smart TV platforms, Roku and Amazon Fire TV. We discover that a subset of apps communicate with a large number of ATSes, and that some ATS organizations only appear on certain platforms, showing a possible segmentation of the smart TV ATS ecosystem across platforms. Third, we evaluate the (in)effectiveness of DNS-based blocklists in preventing smart TVs from accessing ATSes. We highlight that even smart TV-specific blocklists suffer from missed ads and incur functionality breakage. Finally, we examine our Roku and Fire TV datasets for exposure of personally identifiable information (PII) and find that hundreds of apps exfiltrate PII to third parties and platform domains. We also find evidence that some apps send the advertising ID alongside static PII values, effectively eliminating the user’s ability to opt out of ad personalization. Janus Varmarken, Hieu Le 0003, Anastasia Shuba, Athina Markopoulou, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 5 |
| 2020 | Optimizing Taxi Driver Profit Efficiency: A Spatial Network-Based Markov Decision Process ApproachabstractTaxi services play an important role in the public transportation system of large cities. Improving taxi business efficiency is an important societal problem. Most of the recent analytical approaches on this topic only considered how to maximize the pickup chance, energy efficiency, or profit for the immediate next trip when recommending seeking routes, therefore may not be optimal for the overall profit over an extended period of time due to ignoring the destination choice of potential passengers. To tackle this issue, we propose a novel Spatial Network-based Markov Decision Process (SN-MDP) with a rolling horizon configuration to recommend better driving directions. Given a set of historical taxi records and the current status (e.g., road segment and time) of a vacant taxi, we find the best move for this taxi to maximize the profit in the near future. We propose statistical models to estimate the necessary time-variant parameters of SN-MDP from data to avoid competition between drivers. In addition, we take into account fuel cost to assess profit, rather than only income. A case study and several experimental evaluations on a real taxi dataset from a major city in China show that our proposed approach improves the profit efficiency by up to 13.7 percent and outperforms baseline methods in all the time slots. Xun Zhou 0001, Huigui Rong, Qun Zhang 0003, Amin Vahedian Khezerlou, Zubair Shafiq, Alex X. Liu |
IEEE Trans. Big Data | 7 |
| 2020 | Large Scale Characterization of Software Vulnerability Life CyclesabstractSoftware systems inherently contain vulnerabilities that have been exploited in the past resulting in significant revenue losses. The study of various aspects related to vulnerabilities such as their severity, rates of disclosure, exploit and patch release, and existence of common vulnerabilities in different products can help in improving the development, deployment, and maintenance process of software systems. It can also help in designing future security policies and conducting audits of past incidents. Furthermore, such an analysis can help customers to assess the security risks associated with software products of different vendors. In this paper, we conduct an exploratory measurement study of a large software vulnerability data set containing 56077 vulnerabilities disclosed since 1988 till 2013. We investigate vulnerabilities along following eight dimensions: (1) phases in the life cycle of vulnerabilities, (2) evolution of vulnerabilities over the years, (3) functionality of vulnerabilities, (4) access requirement for exploitation of vulnerabilities, (5) risk level of vulnerabilities, (6) software vendors, (7) software products, and (8) existence of common vulnerabilities in multiple software products. Our exploratory analysis uncovers several statistically significant findings that have important implications for software development and deployment. Muhammad Shahzad 0001, Zubair Shafiq, Alex X. Liu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2019 | A postmortem of suspended Twitter accounts in the 2016 U.S. presidential electionabstractSocial media sites such as Twitter have faced significant pressure to mitigate spam and abuse on their platform in the aftermath of congressional investigations into Russian interference in the 2016 U.S. presidential election. Twitter publicly acknowledged the exploitation of their platform and has since conducted aggressive cleanups to suspend the involved accounts. To shed light on Twitter's countermeasures, we conduct a postmortem analysis of about one million Twitter accounts who engaged in the 2016 U.S. presidential election but were later suspended by Twitter. To systematically analyze coordinated activities of these suspended accounts, we group them into communities based on their retweet/mention network and analyze different characteristics such as popular tweeters, domains, and hashtags. The results show that suspended and regular communities exhibit significant differences in terms of popular tweeter and hashtags. Our qualitative analysis also shows that suspended communities are heterogeneous in terms of their characteristics. We further find that accounts suspended by Twitter's new countermeasures are tightly connected to the original suspended communities. Huyen T. Le, Bob Boynton, Zubair Shafiq, Padmini Srinivasan |
ASONAM | 3 |
| 2019 | Measurement and Early Detection of Third-Party Application Abuse on TwitterabstractThird-party applications present a convenient way for attackers to orchestrate a large number of fake and compromised accounts on popular online social networks. Despite recent high-profile reports of third-party application abuse on popular online social networks, prior work lacks automated approaches for accurate and early detection of abusive applications. In this paper, we perform a longitudinal study of abusive third-party applications on Twitter that perform a variety of malicious and spam activities in violation of Twitter's Terms of Service (ToS). Our measurements spanning over a period of 16 months demonstrate an ongoing arms race between attackers continuously registering and abusing new applications and Twitter trying to detect them. We find that hundreds of thousands of abusive applications remain undetected by Twitter for several months while posting tens of millions of tweets. We propose a machine learning approach for accurate and early detection of abusive Twitter applications by analyzing their first few tweets. The evaluation shows that our machine learning approach can accurately detect abusive application with 92.7% precision and 87.0% recall by analyzing their first seven tweets. The deployment of our machine learning approach in the wild shows that attackers continue to abuse third-party applications despite Twitter's recent countermeasures targeting third-party applications. Shehroze Farooqi, Zubair Shafiq |
WWW | 2 |
| 2019 | Measuring Political Personalization of Google News SearchabstractThere is a growing concern about the extent to which algorithmic personalization limits people's exposure to diverse viewpoints, thereby creating “filter bubbles” or “echo chambers.” Prior research on web search personalization has mainly reported location-based personalization of search results. In this paper, we investigate whether web search results are personalized based on a user's browsing history, which can be inferred by search engines via third-party tracking. Specifically, we develop a “sock puppet” auditing system in which a pair of fresh browser profiles, first, visits web pages that reflect divergent political discourses and, second, executes identical politically oriented Google News searches. Comparing the search results returned by Google News for distinctly trained browser profiles, we observe statistically significant personalization that tends to reinforce the presumed partisanship. Huyen T. Le, Raven Maragh, Brian Ekdale, Andrew High, Timothy Havens, Zubair Shafiq |
WWW | 6 |
| 2019 | ShadowBlock: A Lightweight and Stealthy Adblocking BrowserabstractAs the popularity of adblocking has soared over the last few years, publishers are increasingly deploying anti-adblocking paywalls that ask users to either disable their adblockers or pay to access content. In this work we propose ShadowBlock, a new Chromium-based adblocking browser that can hide traces of adblocking activities from anti-adblockers as it removes ads from web pages. To bypass anti-adblocking paywalls, ShadowBlock takes advantage of existing filter lists used by adblockers and hides all ad elements stealthily in such a way that anti-adblocking scripts cannot detect any tampering of the ads (e.g., absence of ad elements). Specifically, ShadowBlock introduces lightweight hooks in Chromium to ensure that DOM states queried by anti-adblocking scripts are exactly as if adblocking is not employed. We implement a fully working prototype by modifying Chromium which shows great promise in terms of adblocking effectiveness and anti-adblocking circumvention but also more efficient than the state-of-the-art adblocking browser extensions. Our evaluation on Alexa top-1K websites shows that ShadowBlock successfully blocks 98.3% of all visible ads while only causing minor breakage on less than 0.6% of the websites. Most importantly, ShadowBlock is able to bypass anti-adblocking paywalls on more than 200 websites that deploy visible anti-adblocking paywalls with a 100% success rate. Our performance evaluation further shows that ShadowBlock loads pages as fast as the state-of-the-art adblocking browser extension on average. Shitong Zhu, Umar Iqbal 0002, Zhongjie Wang 0002, Zhiyun Qian, Zubair Shafiq, Weiteng Chen |
WWW | 5 |
| 2019 | A Girl Has No Name: Automated Authorship Obfuscation using Mutant-XabstractAbstract Stylometric authorship attribution aims to identify an anonymous or disputed document’s author by examining its writing style. The development of powerful machine learning based stylometric authorship attribution methods presents a serious privacy threat for individuals such as journalists and activists who wish to publish anonymously. Researchers have proposed several authorship obfuscation approaches that try to make appropriate changes (e.g. word/phrase replacements) to evade attribution while preserving semantics. Unfortunately, existing authorship obfuscation approaches are lacking because they either require some manual effort, require significant training data, or do not work for long documents. To address these limitations, we propose a genetic algorithm based random search framework called Mutant-X which can automatically obfuscate text to successfully evade attribution while keeping the semantics of the obfuscated text similar to the original text. Specifically, Mutant-X sequentially makes changes in the text using mutation and crossover techniques while being guided by a fitness function that takes into account both attribution probability and semantic relevance. While Mutant-X requires black-box knowledge of the adversary’s classifier, it does not require any additional training data and also works on documents of any length. We evaluate Mutant-X against a variety of authorship attribution methods on two different text corpora. Our results show that Mutant-X can decrease the accuracy of state-of-the-art authorship attribution methods by as much as 64% while preserving the semantics much better than existing automated authorship obfuscation approaches. While Mutant-X advances the state-of-the-art in automated authorship obfuscation, we find that it does not generalize to a stronger threat model where the adversary uses a different attribution classifier than what Mutant-X assumes. Our findings warrant the need for future research to improve the generalizability (or transferability) of automated authorship obfuscation approaches. Asad Mahmood, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar |
Proc. Priv. Enhancing Technol. | 3 |
| 2019 | No Place to Hide: Inadvertent Location Privacy Leaks on TwitterabstractAbstract There is a natural tension between the desire to share information and keep sensitive information private on online social media. Privacy seeking social media users may seek to keep their location private by avoiding the mentions of location revealing words such as points of interest (POIs), believing this to be enough. In this paper, we show that it is possible to uncover the location of a social media user’s post even when it is not geotagged and does not contain any POI information. Our proposed approach Jasoosachieves this by exploiting the shared vocabulary between users who reveal their location and those who do not. To this end, Jasoosuses a variant of the Naive Bayes algorithm to identify location revealing words or hashtags based on both temporal and atemporal perspectives. Our evaluation using tweets collected from four different states in the United States shows that Jasooscan accurately infer the locations of close to half a million tweets corresponding to more than 20,000 distinct users (i.e., more than 50% of the test users) from the four states. Our work demonstrates that location privacy leaks do occur despite due precautions by a privacy conscious user. We design and evaluate countermeasures based Jasoosto mitigate location privacy leaks. Jonathan Rusert, Osama Khalid, Dat Hong, Zubair Shafiq, Padmini Srinivasan |
Proc. Priv. Enhancing Technol. | 4 |
| 2018 | Real-time Video Quality of Experience Monitoring for HTTPS and QUICabstractThe widespread deployment of end-to-end encryption protocols such as HTTPS and QUIC has reduced the visibility for operators into traffic on their networks. Network operators need the visibility to monitor and mitigate Quality of Experience (QoE) impairments in popular applications such as video streaming. To address this problem, we propose a machine learning based approach to monitor QoE metrics for encrypted video traffic. We leverage network and transport layer information as features to train machine learning classifiers for inferring video QoE metrics such as startup delay and rebuffering events. Using our proposed approach, network operators can detect and react to encrypted video QoE impairments in real-time. We evaluate our approach for YouTube adaptive video streams using HTTPS and QUIC. The experimental evaluations show that our approach achieves up to 90% classification accuracy for HTTPS and up to 85 % classification accuracy for QUIC. M. Hammad Mazhar, Zubair Shafiq |
INFOCOM | 2 |
| 2018 | Measuring and Disrupting Anti-Adblockers Using Differential Execution Analysis
Shitong Zhu, Xunchao Hu, Zhiyun Qian, Zubair Shafiq, Heng Yin 0001 |
NDSS | 4 |
| 2018 | NoMoAds: Effective and Efficient Cross-App Mobile Ad-BlockingabstractAbstract Although advertising is a popular strategy for mobile app monetization, it is often desirable to block ads in order to improve usability, performance, privacy, and security. In this paper, we propose NoMoAds to block ads served by any app on a mobile device. NoMoAds leverages the network interface as a universal vantage point: it can intercept, inspect, and block outgoing packets from all apps on a mobile device. NoMoAds extracts features from packet headers and/or payload to train machine learning classifiers for detecting ad requests. To evaluate NoMoAds, we collect and label a new dataset using both EasyList and manually created rules. We show that NoMoAds is effective: it achieves an F-score of up to 97.8% and performs well when deployed in the wild. Furthermore, NoMoAds is able to detect mobile ads that are missed by EasyList (more than one-third of ads in our dataset). We also show that NoMoAds is efficient: it performs ad classification on a per-packet basis in real-time. To the best of our knowledge, NoMoAds is the first mobile ad-blocker to effectively and efficiently block ads served across all apps using a machine learning approach. Anastasia Shuba, Athina Markopoulou, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 3 |
| 2018 | Optimizing Internet Transit Routing for Content Delivery NetworksabstractContent delivery networks (CDNs) maintain multiple transit routes from content distribution servers to eyeball ISP networks which provide Internet connectivity to end users. Due to the dynamics of varying performance and pricing on transit routes, CDNs need to implement a transit route selection strategy to optimize performance and cost tradeoffs. In this paper, we formalize the transit routing problem using a multi-attribute objective function to simultaneously optimize end-to-end performance and cost. Our approach allows CDNs to navigate the cost and performance tradeoff in transit routing through a single control knob. We evaluate our approach using real-world measurements from CDN servers located at 19 geographically distributed Internet exchange points. Using our approach, CDNs can reduce transit costs on average by 57% without sacrificing performance. Faraz Ahmed, Zubair Shafiq, Amir R. Khakpour, Alex X. Liu |
IEEE/ACM Trans. Netw. | 2 |
| 2017 | Revisiting The American Voter on TwitterabstractThe American Voter - a seminal work in political science - uncovered the multifaceted nature of voting behavior which has been corroborated in electoral research for decades since. In this paper, we leverage The American Voter as an analysis framework in the realm of computational political science, employing the factors of party, personality, and policy to structure the analysis of public discourse on online social media during the 2016 U.S. presidential primaries. Our analysis of 50 million tweets reveals the continuing importance of these three factors; our understanding is also enriched by the application of sentiment analysis techniques. The overwhelmingly negative sentiment of conversations surrounding 10 major presidential candidates reveals more "crosstalk" from Democratic leaning users towards Republican candidates, and less vice-versa. We uncover the lack of moderation as the most discussed personality dimension during this campaign season, as the political field becomes more extreme - Clinton and Rubio are perceived as moderate, while Trump, Sanders, and Cruz are not. While the most discussed issues are foreign policy and immigration, Republicans tweet more about abortion than Democrats who tweet more about gay rights than Republicans. Finally, we illustrate the importance of multifaceted political discourse analysis by applying regression to quantify the impact of party, personality, and policy on national polls. Huyen T. Le, Bob Boynton, Yelena Mejova, Zubair Shafiq, Padmini Srinivasan |
CHI | 4 |
| 2017 | Distributed Load Balancing in Key-Value Networked CachesabstractModern web services rely on a network of distributed cache servers to efficiently deliver content to users. Load imbalance among cache servers can substantially degrade content delivery performance. Due to the skewed and dynamic nature of real-world workloads, cache servers that serve viral content experience higher load as compared to other cache servers. We propose a novel distributed load balancing protocol called Meezan to address the load imbalance among cache servers. Meezan replicates popular objects to mitigate skewness and adjusts hash space boundaries in response to load dynamics in a novel way. Our theoretical analysis shows that Meezan achieves near perfect load balancing for a wide range of operating parameters. Our trace driven simulations shows that Meezan reduces load imbalance by up to 52% as compared to prior solutions. Sikder Huq, Zubair Shafiq, Sukumar Ghosh, Amir R. Khakpour, Harkeerat Bedi |
ICDCS | 2 |
| 2017 | Accurate Detection of Automatically Spun Content via Stylometric AnalysisabstractSpammers use automated content spinning techniques to evade plagiarism detection by search engines. Text spinners help spammers in evading plagiarism detectors by automatically restructuring sentences and replacing words or phrases with their synonyms. Prior work on spun content detection relies on the knowledge about the dictionary used by the text spinning software. In this work, we propose an approach to detect spun content and its seed without needing the text spinner's dictionary. Our key idea is that text spinners introduce stylometric artifacts that can be leveraged for detecting spun documents. We implement and evaluate our proposed approach on a corpus of spun documents that are generated using a popular text spinning software. The results show that our approach can not only accurately detect whether a document is spun but also identify its source (or seed) document - all without needing the dictionary used by the text spinner. Usman Shahid, Shehroze Farooqi, Raza Ahmad, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar |
ICDM | 4 |
| 2017 | Peering vs. transit: Performance comparison of peering and transit interconnectionsabstractThe economic aspects of peering and transit interconnections between ISPs have been extensively studied in prior literature. Prior research primarily focuses on the economic issues associated with establishing peering and transit connectivity among ISPs to model interconnection strategies. Performance analysis, on the other hand, while understood intuitively, has not been empirically quantified and incorporated in such models. To fill this gap, we conduct a large scale measurement based performance comparison of peering and transit interconnection strategies. We use JavaScript to conduct application layer latency measurements between 510K clients in 900 access ISPs and multi-homed CDN servers located at 33 IXPs around the world. Overall, we find that peering paths outperformed transit paths for 91% Autonomous Systems (ASes) in our data. Peering paths have smaller propagation delays as compared to transit paths for more than 95% ASes. Peering paths outperform transit paths in terms of propagation delay due to shorter path lengths. Peering paths also have smaller queueing delays as compared to transit paths for more than 50% ASes. Zubair Shafiq, Harkeerat Bedi, Amir R. Khakpour |
ICNP | 2 |
| 2017 | Suffering from buffering? Detecting QoE impairments in live video streamsabstractFueled by increasing network bandwidth and decreasing costs, the popularity of over-the-top large-scale live video streaming has dramatically increased over the last few years. In this paper, we present a measurement study of adaptive bitrate video streaming for a large-scale live event. Using server-side logs from a commercial content delivery network, we study live video delivery for the annual Academy Awards event that was streamed by hundreds of thousands of viewers in the United States. We analyze the relationship between Quality-of-Experience (QoE) and user engagement. We first study the impact of buffering, average bitrate, and bitrate fluctuations on user engagement. To account for interdependencies among QoE metrics and other confounding factors, we use quasi-experiments to quantify the causal impact of different QoE metrics on user engagement. We further design and implement a Principal Component Analysis (PCA) based technique to detect live video QoE impairments in real-time. We then use Hampel filters to detect QoE impairments and report 92% accuracy with 20% improvement in true positive rate as compared to baselines. Our approach allows content providers to detect and mitigate QoE impairments on the fly instead of relying on post-hoc analysis. Zubair Shafiq, Harkeerat Bedi, Amir R. Khakpour |
ICNP | 2 |
| 2017 | Multipath TCP traffic diversion attacks and countermeasuresabstractMultipath TCP (MPTCP) is an IETF standardized suite of TCP extensions that allow two endpoints to simultaneously use multiple paths between them. In this paper, we report vulnerabilities in MPTCP that arise because of cross-path interactions between MPTCP subflows. First, an attacker eavesdropping one MPTCP subflow can infer throughput of other subflows. Second, an attacker can inject forged MPTCP packets to change priorities of any MPTCP subflow. We present two attacks to exploit these vulnerabilities. In the connection hijack attack, an attacker takes full control of the MPTCP connection by suspending the subflows he has no access to. In the traffic diversion attack, an attacker diverts traffic from one path to other paths. Proposed vulnerabilities fixes, changes to MPTCP specification, provide the guarantees that MPTCP is at least as secure as TCP and the original MPTCP. We validate attacks and prevention mechanism, using MPTCP Linux implementation (v0.91), on a real-network testbed. Ali Munir, Zhiyun Qian, Zubair Shafiq, Alex X. Liu, Franck Le |
ICNP | 3 |
| 2017 | Scalable News Slant Measurement Using Twitter
Huyen T. Le, Zubair Shafiq, Padmini Srinivasan |
ICWSM | 2 |
| 2017 | Measuring and mitigating oauth access token abuse by collusion networksabstractWe uncover a thriving ecosystem of large-scale reputation manipulation services on Facebook that leverage the principle of collusion. Collusion networks collect OAuth access tokens from colluding members and abuse them to provide fake likes or comments to their members. We carry out a comprehensive measurement study to understand how these collusion networks exploit popular third-party Facebook applications with weak security settings to retrieve OAuth access tokens. We infiltrate popular collusion networks using honeypots and identify more than one million colluding Facebook accounts by "milking" these collusion networks. We disclose our findings to Facebook and collaborate with them to implement a series of countermeasures that mitigate OAuth access token abuse without sacrificing application platform usability for third-party developers. These countermeasures remained in place until April 2017, after which Facebook implemented a set of unrelated changes in its infrastructure to counter collusion networks. We are the first to report and effectively mitigate large-scale OAuth access token abuse in the wild. Shehroze Farooqi, Fareed Zaffar, Nektarios Leontiadis, Zubair Shafiq |
Internet Measurement Conference | 4 |
| 2017 | The ad wars: retrospective measurement and analysis of anti-adblock filter listsabstractThe increasing popularity of adblockers has prompted online publishers to retaliate against adblock users by deploying anti-adblock scripts, which detect adblock users and bar them from accessing content unless they disable their adblocker. To circumvent anti-adblockers, adblockers rely on manually curated anti-adblock filter lists for removing anti-adblock scripts. Anti-adblock filter lists currently rely on informal crowdsourced feedback from users to add/remove filter list rules. In this paper, we present the first comprehensive study of anti-adblock filter lists to analyze their effectiveness against anti-adblockers. Specifically, we compare and contrast the evolution of two popular anti-adblock filter lists. We show that these filter lists are implemented very differently even though they currently have a comparable number of filter list rules. We then use the Internet Archive's Wayback Machine to conduct a retrospective coverage analysis of these filter lists on Alexa top-5K websites over the span of last five years. We find that the coverage of these filter lists has considerably improved since 2014 and they detect anti-adblockers on about 9% of Alexa top-5K websites. To improve filter list coverage and speedup addition of new filter rules, we also design and implement a machine learning based method to automatically detect anti-adblock scripts using static JavaScript code analysis. Umar Iqbal 0002, Zubair Shafiq, Zhiyun Qian |
Internet Measurement Conference | 2 |
| 2017 | Detecting Anti Ad-blockers in the WildabstractAbstract The rise of ad-blockers is viewed as an economic threat by online publishers who primarily rely on online advertising to monetize their services. To address this threat, publishers have started to retaliate by employing anti ad-blockers, which scout for ad-block users and react to them by pushing users to whitelist the website or disable ad-blockers altogether. The clash between ad-blockers and anti ad-blockers has resulted in a new arms race on the Web. In this paper, we present an automated machine learning based approach to identify anti ad-blockers that detect and react to ad-block users. The approach is promising with precision of 94.8% and recall of 93.1%. Our automated approach allows us to conduct a large-scale measurement study of anti ad-blockers on Alexa top-100K websites. We identify 686 websites that make visible changes to their page content in response to ad-block detection. We characterize the spectrum of different strategies used by anti ad-blockers. We find that a majority of publishers use fairly simple first-party anti ad-block scripts. However, we also note the use of third-party anti ad-block services that use more sophisticated tactics to detect and respond to ad-blockers. Muhammad Haris Mughees, Zhiyun Qian, Zubair Shafiq |
Proc. Priv. Enhancing Technol. | 3 |
| 2017 | Measuring, Characterizing, and Detecting Facebook Like FarmsabstractOnline social networks offer convenient ways to reach out to large audiences. In particular, Facebook pages are increasingly used by businesses, brands, and organizations to connect with multitudes of users worldwide. As the number of likes of a page has become a de-facto measure of its popularity and profitability, an underground market of services artificially inflating page likes (“like farms ”) has emerged alongside Facebook’s official targeted advertising platform. Nonetheless, besides a few media reports, there is little work that systematically analyzes Facebook pages’ promotion methods. Aiming to fill this gap, we present a honeypot-based comparative measurement study of page likes garnered via Facebook advertising and from popular like farms. First, we analyze likes based on demographic, temporal, and social characteristics and find that some farms seem to be operated by bots and do not really try to hide the nature of their operations, while others follow a stealthier approach, mimicking regular users’ behavior. Next, we look at fraud detection algorithms currently deployed by Facebook and show that they do not work well to detect stealthy farms that spread likes over longer timespans and like popular pages to mimic regular users. To overcome their limitations, we investigate the feasibility of timeline-based detection of like farm accounts, focusing on characterizing content generated by Facebook accounts on their timelines as an indicator of genuine versus fake social activity. We analyze a wide range of features extracted from timeline posts, which we group into two main categories: lexical and non-lexical. We find that like farm accounts tend to re-share content more often, use fewer words and poorer vocabulary, and more often generate duplicate comments and likes compared to normal users. Using relevant lexical and non-lexical features, we build a classifier to detect like farms accounts that achieves a precision higher than 99% and a 93% recall. Muhammad Ikram 0001, Lucky Onwuzurike, Shehroze Farooqi, Emiliano De Cristofaro, Arik Friedman, Guillaume Jourjon, Mohamed Ali Kâafar, Zubair Shafiq |
ACM Trans. Priv. Secur. | 8 |
| 2017 | A Traffic Flow Approach to Early Detection of Gathering Events: Comprehensive ResultsabstractGiven a spatial field and the traffic flow between neighboring locations, the early detection of gathering events ( edge ) problem aims to discover and localize a set of most likely gathering events. It is important for city planners to identify emerging gathering events that might cause public safety or sustainability concerns. However, it is challenging to solve the edge problem due to numerous candidate gathering footprints in a spatial field and the nontrivial task of balancing pattern quality and computational efficiency. Prior solutions to model the edge problem lack the ability to describe the dynamic flow of traffic and the potential gathering destinations because they rely on static or undirected footprints. In our recent work, we modeled the footprint of a gathering event as a Gathering Graph (G-Graph), where the root of the directed acyclic G-Graph is the potential destination and the directed edges represent the most likely paths traffic takes to move toward the destination. We also proposed an efficient algorithm called SmartEdge to discover the most likely nonoverlapping G-Graphs in the given spatial field. However, it is challenging to perform a systematic performance study of the proposed algorithm, due to unavailability of the ground truth of gathering events. In this article, we introduce an event simulation mechanism, which makes it possible to conduct a comprehensive performance study of the SmartEdge algorithm. We measure the quality of the detected patterns, in a systematic way, in terms of timeliness and location accuracy. The results show that, on average, the SmartEdge algorithm is able to detect patterns within a grid cell away (less than 500 meters) of the simulated events and detect patterns of the simulated events as early as 10 minutes prior to the first arrival to the gathering event. Amin Vahedian Khezerlou, Xun Zhou 0001, Lufan Li, Zubair Shafiq, Alex X. Liu, Fan Zhang 0019 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2016 | The Rich and the Poor: A Markov Decision Process Approach to Optimizing Taxi Driver Revenue EfficiencyabstractTaxi services play an important role in the public transportation system of large cities. Improving taxi business efficiency is an important societal problem since it could improve the income of the drivers and reduce gas emissions and fuel consumption. The recent research on seeking strategies may not be optimal for the overall revenue over an extended period of time as they ignored the important impact of passengers' destinations on future passenger seeking. To address these issues, this paper investigates how to increase the revenue efficiency (revenue per unit time) of taxi drivers, and models the passenger seeking process as a Markov Decision Process (MDP). For each one-hour time slot, we learn a different set of parameters for the MDP from data and find the best move for a vacant taxi to maximize the total revenue in that time slot. A case study and several experimental evaluations on a real dataset from a major city in China show that our proposed approach improves the revenue efficiency of inexperienced drivers by up to 15% and outperforms a baseline method in all the time slots. Huigui Rong, Xun Zhou 0001, Zubair Shafiq, Alex X. Liu |
CIKM | 4 |
| 2016 | Malware Slums: Measurement and Analysis of Malware on Traffic ExchangesabstractAuto-surf and manual-surf traffic exchanges are an increasingly popular way of artificially generating website traffic. Previous research in this area has focused on the makeup, usage, and monetization of underground traffic exchanges. In this paper, we analyze the role of traffic exchanges as a vector for malware propagation. We conduct a measurement study of nine auto-surf and manual-surf traffic exchanges over several months. We present a first of its kind analysis of the different types of malware that are propagated through these traffic exchanges. We find that more than 26% of the URLs surfed on traffic exchanges contain malicious content. We further analyze different categories of malware encountered on traffic exchanges, including blacklisted domains, malicious JavaScript, malicious Flash, and malicious shortened URLs. Salman Yousaf, Umar Iqbal 0002, Shehroze Farooqi, Raza Ahmad, Zubair Shafiq, Fareed Zaffar |
DSN | 5 |
| 2016 | A traffic flow approach to early detection of gathering eventsabstractGiven a spatial field and the traffic flow between neighboring locations, the early detection of gathering events (edge) problem aims to discover and localize a set of most likely gathering events. It is important for city planners to identify emerging gathering events which might cause public safety or sustainability concerns. However, it is challenging to solve the edge problem due to numerous candidate gathering footprints in a spatial field and the non-trivial task to balance pattern quality and computational efficiency. Prior solutions to model the edge problem lack the ability to describe the dynamic flow of traffic and the potential gathering destinations because they rely on static or undirected footprints. In contrast, in this paper, we model the footprint of a gathering event as a Gathering directed acyclic Graph (G-Graph), where the root of the G-Graph is the potential destination and the directed edges represent the most likely paths traffic takes to move towards the destination. We also proposed an efficient algorithm called SmartEdge to discover the most likely non-overlapping G-Graphs in the given spatial field. Our analysis shows that the proposed G-Graph model and the SmartEdge algorithm have the ability to efficiently and effectively capture important gathering events from real-world human mobility data. Our experimental evaluations show that SmartEdge saves 50% computation time over the baseline algorithm. Xun Zhou 0001, Amin Vahedian Khezerlou, Alex X. Liu, Zubair Shafiq, Fan Zhang 0019 |
SIGSPATIAL/GIS | 4 |
| 2016 | The Internet is for Porn: Measurement and Analysis of Online Adult TrafficabstractAdult (or pornographic) websites attract a large number of visitors and account for a substantial fraction of the global Internet traffic. However, little is known about the makeup and characteristics of online adult traffic. In this paper, we present the first large-scale measurement study of online adult traffic using HTTP logs collected from a major commercial content delivery network. Our data set contains approximately 323 terabytes worth of traffic from 80 million users, and includes traffic from several dozen major adult websites and their users in four different continents. We analyze several characteristics of online adult traffic including content and traffic composition, device type composition, temporal dynamics, content popularity, content injection, and user engagement. Our analysis reveals several unique characteristics of online adult traffic. We also analyze implications of our findings on adult content delivery. Our findings suggest several content delivery and cache performance optimizations for adult traffic, e.g., modifications to website design, content delivery, cache placement strategies, and cache storage configurations. Faraz Ahmed, Zubair Shafiq, Alex X. Liu |
ICDCS | 2 |
| 2016 | Optimizing Internet transit routing for content delivery networksabstractContent Distribution Networks (CDNs) maintain multiple transit routes from content distribution servers to eyeball ISP networks which provide Internet connectivity to end users. Due to the dynamics of varying performance and pricing on transit routes, CDNs need to implement a transit route selection strategy to optimize performance and cost tradeoffs. In this paper, we formalize the transit routing problem using a multi-attribute objective function to simultaneously optimize end-to-end performance and cost. Our approach allows CDNs to navigate the cost and performance tradeoff in transit routing through a single control knob. We evaluate our approach using real-world measurements from CDN servers located at 19 geographically distributed IXPs. Using our approach, CDNs can reduce transit costs on average by 57% without sacrificing performance. Faraz Ahmed, Zubair Shafiq, Amir R. Khakpour, Alex X. Liu |
ICNP | 2 |
| 2016 | Characterizing caching workload of a large commercial Content Delivery NetworkabstractContent Delivery Networks (CDNs) have emerged as a dominant mechanism to deliver content over the Internet. Despite their importance, to our best knowledge, large-scale analysis of CDN cache performance is lacking in prior literature. A CDN serves many content publishers simultaneously and thus has unique workload characteristics; it typically deals with extremely large content volume and high content diversity from multiple content publishers. CDNs also have unique performance metrics; other than hit ratio, CDNs also need to minimize network and disk load on cache servers. In this paper, we present measurement and analysis of caching workload at a large commercial CDN. Using detailed logs from four geographically distributed CDN cache servers, we analyze over 600 million content requests accounting for more than 1.3 petabytes worth of traffic. We analyze CDN workload from a wide range of perspectives, including request composition, size, popularity, and temporal dynamics. Using real-world logs, we also evaluate cache replacement algorithms, including two enhancements designed based on our CDN workload analysis: N-hit and content-aware caching. The results show that these enhancements achieve substantial performance gains in terms of cache hit ratio, disk load, and origin traffic volume. Zubair Shafiq, Amir R. Khakpour, Alex X. Liu |
INFOCOM | 1 |
| 2016 | QoE Analysis of a Large-Scale Live Video Streaming EventabstractStreaming video has received a lot of attention from industry and academia. In this work, we study the characteristics and challenges associated with large-scale live video delivery. Using logs from a commercial Content Delivery Network (CDN), we study live video delivery for a major entertainment event that was streamed by hundreds of thousands of viewers in North America. We analyze Quality-of-Experience (QoE) for the event and note that a significant number of users suffer QoE impairments. As a consequence of QoE impairments, these users exhibit lower engagement metrics. Zubair Shafiq, Amir R. Khakpour |
SIGMETRICS | 2 |
| 2016 | Special Issue on Mobile Traffic Analytics
Marco Fiore 0001, Zubair Shafiq, Zbigniew Smoreda, Razvan Stanica, Roberto Trasarti |
Comput. Commun. | 2 |
| 2016 | Characterizing and Optimizing Cellular Network Performance During Crowded EventsabstractDuring crowded events, cellular networks face voice and data traffic volumes that are often orders of magnitude higher than what they face during routine days. Despite the use of portable base stations for temporarily increasing communication capacity and free Wi-Fi access points for offloading Internet traffic from cellular base stations, crowded events still present significant challenges for cellular network operators looking to reduce dropped call events and improve Internet speeds. For an effective cellular network design, management, and optimization, it is crucial to understand how cellular network performance degrades during crowded events, what causes this degradation, and how practical mitigation schemes would perform in real-life crowded events. This paper makes a first step toward this end by characterizing the operational performance of a tier-1 cellular network in the U.S. during two high-profile crowded events in 2012. We illustrate how the changes in population distribution, user behavior, and application workload during crowded events result in significant voice and data performance degradation, including more than two orders of magnitude increase in connection failures. Our findings suggest two mechanisms that can improve performance without resorting to costly infrastructure changes: radio resource allocation tuning and opportunistic connection sharing. Using trace-driven simulations, we show that more aggressive release of radio resources via 1–2 s shorter radio resource control timeouts as compared with routine days helps to achieve better tradeoff between wasted radio resources, energy consumption, and delay during crowded events, and opportunistic connection sharing can reduce connection failures by 95% when employed by a small number of devices in each cell sector. Zubair Shafiq, Lusheng Ji, Alex X. Liu, Jeffrey Pang, Shobha Venkataraman, Jia Wang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2015 | Geospatial and Temporal Dynamics of Application Usage in Cellular Data NetworksabstractSignificant geospatial and temporal correlations, in terms of traffic volume and application access, exist in cellular network usage as shown in recent studies on cellular network measurement. Such geospatial and temporal correlation patterns provide local optimization opportunities to cellular network operators for handling the explosive growth in the traffic volume observed in recent years. To the best of our knowledge, in this paper, we provide the first fine-grained joint characterization of the geospatial and temporal dynamics of application usage in a 3G cellular data network. Our analysis is based on two simultaneously collected traces from the radio access network (containing location records) and the core network (containing traffic records) of a tier-1 cellular network in the United States. To better understand the application usage in our data, we first cluster cell locations based on their application distributions and then study the geospatial and temporal dynamics of application usage across different geographical regions. The results of our measurement study present cellular network operators with fine-grained insights that can be leveraged to tune network parameter settings for better network performance and user experience. Zubair Shafiq, Lusheng Ji, Alex X. Liu, Jeffrey Pang, Jia Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2014 | Breaching IM session privacy using causalityabstractThe breach of privacy in encrypted instant messenger (IM) service is a serious threat to user anonymity. Performance of previous de-anonymization strategies was limited to 65%. We perform network de-anonymization by taking advantage of the cause-effect relationship between sent and received packet streams and demonstrate this approach on a data set of Yahoo! IM service traffic traces. An investigation of various measures of causality shows that IM networks can be breached with a hit rate of 99%. A KCI Causality based approach alone can provide a true positive rate of about 97%. Individual performances of Granger, Zhang and IGCI causality are limited owing to the very low SNR of packet traces and variable network delays. Saad Saleh, Mamoon Raja, Muhammad Shahnawaz, Muhammad Usman Ilyas, Khawar Khurshid, Zubair Shafiq, Alex X. Liu, Hayder Radha, Shirish S. Karande |
GLOBECOM | 6 |
| 2014 | Paying for Likes?: Understanding Facebook Like Fraud Using HoneypotsabstractFacebook pages offer an easy way to reach out to a very large audience as they can easily be promoted using Facebook's advertising platform. Recently, the number of likes of a Facebook page has become a measure of its popularity and profitability, and an underground market of services boosting page likes, aka like farms, has emerged. Some reports have suggested that like farms use a network of profiles that also like other pages to elude fraud protection algorithms, however, to the best of our knowledge, there has been no systematic analysis of Facebook pages' promotion methods. Emiliano De Cristofaro, Arik Friedman, Guillaume Jourjon, Mohamed Ali Kâafar, Zubair Shafiq |
Internet Measurement Conference | 5 |
| 2014 | Understanding the impact of network dynamics on mobile video user engagementabstractMobile network operators have a significant interest in the performance of streaming video on their networks because network dynamics directly influence the Quality of Experience (QoE). However, unlike video service providers, network operators are not privy to the client- or server-side logs typically used to measure key video performance metrics, such as user engagement. To address this limitation, this paper presents the first large-scale study characterizing the impact of cellular network performance on mobile video user engagement from the perspective of a network operator. Our study on a month-long anonymized data set from a major cellular network makes two main contributions. First, we quantify the effect that 31 different network factors have on user behavior in mobile video. Our results provide network operators direct guidance on how to improve user engagement --- for example, improving mean signal-to-interference ratio by 1 dB reduces the likelihood of video abandonment by 2%. Second, we model the complex relationships between these factors and video abandonment, enabling operators to monitor mobile video user engagement in real-time. Our model can predict whether a user completely downloads a video with more than 87% accuracy by observing only the initial 10 seconds of video streaming sessions. Moreover, our model achieves significantly better accuracy than prior models that require client- or server-side logs, yet we only use standard radio network statistics and/or TCP/IP headers available to network operators. Zubair Shafiq, Jeffrey Erman, Lusheng Ji, Alex X. Liu, Jeffrey Pang, Jia Wang 0001 |
SIGMETRICS | 1 |
| 2014 | Revisiting caching in content delivery networksabstractContent Delivery Networks (CDNs) differ from other caching systems in terms of both workload characteristics and performance metrics. However, there has been little prior work on large-scale measurement and characterization of content requests and caching performance in CDNs. For workload characteristics, CDNs deal with extremely large content volume, high content diversity, and strong temporal dynamics. For performance metrics, other than hit ratio, CDNs also need to minimize the disk operations and the volume of traffic from origin servers. In this paper, we conduct a large-scale measurement study to characterize the content request patterns using real-world data from a commercial CDN provider. Zubair Shafiq, Alex X. Liu, Amir R. Khakpour |
SIGMETRICS | 1 |
| 2013 | Is news sharing on Twitter ideologically biased?abstractIn this paper we explore effects of perceived ideology of news outlets on consumption and sharing of news in Twitter. Selective exposure theory suggests that when given access to a broad range of information, people will tend to consume and share news that confirms their existing beliefs and biases. We find that users share news in similar ways regardless of outlet or perceived ideology of outlet, and that as a user shares more news content, they tend to quickly include outlets with opposing viewpoints. This suggests that while perceived ideology does not inspire most Twitter users to treat liberal or conservative news outlets differently, it is a factor in their news consumption and sharing. Specifically, users in our sample who sent multiple tweets tended to increase the ideological diversity in news they shared within two or three tweets, and users' information diversity increased as their number of tweets sent increased. Jonathan Scott Morgan, Cliff Lampe, Zubair Shafiq |
CSCW | 3 |
| 2013 | Cross-path inference attacks on multipath TCPabstractMultipath TCP (MPTCP) allows the concurrent use of multiple paths between two end points, and as such holds great promise for improving application performance. However, in this paper, we report a newly discovered class of attacks on MPTCP that may jeopardize and hamper its wide-scale adoption. The attacks stem from the interdependence between the multiple subflows in an MPTCP connection. MPTCP congestion control algorithms are designed to achieve resource pooling and fairness with single-path TCP users at shared bottlenecks. Therefore, multiple MPTCP subflows are inherently coupled with each other, resulting in potential side-channels that can be exploited to infer cross-path properties. In particular, an ISP monitoring one or more paths used by an MPTCP connection can infer sensitive and proprietary information (e.g., level of network congestion, end-to-end TCP throughput, packet loss, network delay) about its competitors. Since the side-channel information enabled by the coupling among the subflows in an MPTCP connection results directly from the design goals of MPTCP congestion control algorithms, it is not obvious how to circumvent this attack easily. We believe our findings provide insights that can be used to guide future security-related research on MPTCP and other similar multipath extensions. Zubair Shafiq, Franck Le, Mudhakar Srivatsa, Alex X. Liu |
HotNets | 1 |
| 2013 | Who are you talking to? Breaching privacy in encrypted IM networksabstractWe present a novel attack on relayed instant messaging (IM) traffic that allows an attacker to infer who's talking to whom with high accuracy. This attack only requires collection of packet header traces between users and IM servers for a short time period, where each packet in the trace goes from a user to an IM server or vice-versa. The specific goal of the attack is to accurately identify a candidate set of top-k users with whom a given user possibly talked to, while using only the information available in packet header traces (packet payloads cannot be used because they are mostly encrypted). Towards this end, we propose a wavelet-based scheme, called COmmunication Link De-anonymization (COLD), and evaluate its effectiveness using a real-world Yahoo! Messenger data set. The results of our experiments show that COLD achieves a hit rate of more than 90% for a candidate set size of 10. For slightly larger candidate set size of 20, COLD achieves almost 100% hit rate. In contrast, a baseline method using time series correlation could only achieve less than 5% hit rate for similar candidate set sizes. Muhammad Usman Ilyas, Zubair Shafiq, Alex X. Liu, Hayder Radha |
ICNP | 2 |
| 2013 | A first look at cellular network performance during crowded eventsabstractDuring crowded events, cellular networks face voice and data traffic volumes that are often orders of magnitude higher than what they face during routine days. Despite the use of portable base stations for temporarily increasing communication capacity and free Wi-Fi access points for offloading Internet traffic from cellular base stations, crowded events still present significant challenges for cellular network operators looking to reduce dropped call events and improve Internet speeds. For effective cellular network design, management, and optimization, it is crucial to understand how cellular network performance degrades during crowded events, what causes this degradation, and how practical mitigation schemes would perform in real-life crowded events. This paper makes a first step towards this end by characterizing the operational performance of a tier-1 cellular network in the United States during two high-profile crowded events in 2012. We illustrate how the changes in population distribution, user behavior, and application workload during crowded events result in significant voice and data performance degradation, including more than two orders of magnitude increase in connection failures. Our findings suggest two mechanisms that can improve performance without resorting to costly infrastructure changes: radio resource allocation tuning and opportunistic connection sharing. Using trace-driven simulations, we show that more aggressive release of radio resources via 1-2 seconds shorter RRC timeouts as compared to routine days helps to achieve better tradeoff between wasted radio resources, energy consumption, and delay during crowded events; and opportunistic connection sharing can reduce connection failures by 95% when employed by a small number of devices in each cell sector. Zubair Shafiq, Lusheng Ji, Alex X. Liu, Jeffrey Pang, Shobha Venkataraman, Jia Wang 0001 |
SIGMETRICS | 1 |
| 2013 | A Distributed Algorithm for Identifying Information Hubs in Social NetworksabstractThis paper addresses the problem of identifying the top-k information hubs in a social network. Identifying top-k information hubs is crucial for many applications such as advertising in social networks where advertisers are interested in identifying hubs to whom free samples can be given. Existing solutions are centralized and require time stamped information about pair-wise user interactions and can only be used by social network owners as only they have access to such data. Existing distributed algorithms suffer from poor accuracy. In this paper, we propose a new algorithm to identify information hubs that preserves user privacy. Our method can identify hubs without requiring a central entity to access the complete friendship graph. We achieve this by fully distributing the computation using the Kempe-McSherry algorithm, while addressing user privacy concerns. We evaluate the effectiveness of our proposed technique using three real-world data set; The first two are Facebook data sets containing about 6 million users and more than 40 million friendship links. The third data set is from Twitter and comprises of a little over 2 million users. The results of our analysis show that our algorithm is up to 50% more accurate than existing algorithms. Results also show that the proposed algorithm can estimate the rank of the top-k information hubs users more accurately than existing approaches. Muhammad Usman Ilyas, Zubair Shafiq, Alex X. Liu, Hayder Radha |
IEEE J. Sel. Areas Commun. | 2 |
| 2013 | Identifying Leaders and Followers in Online Social NetworksabstractIdentifying leaders and followers in online social networks is important for various applications in many domains such as advertisement, community health campaigns, administrative science, and even politics. In this paper, we study the problem of identifying leaders and followers in online social networks using user interaction information. We propose a new model, called the Longitudinal User Centered Influence (LUCI) model, that takes as input user interaction information and clusters users into four categories: introvert leaders, extrovert leaders, followers, and neutrals. To validate our model, we first apply it to a data set collected from an online social network called Everything2. Our experimental results show that our LUCI model achieves an average classification accuracy of up to 90.3% in classifying users as leaders and followers, where the ground truth is based on the labeled roles of users. Second, we apply our LUCI model on a data set collected from Facebook consisting of interactions among more than 3 million users over the duration of one year. However, we do not have ground truth data for Facebook users. Therefore, we analyze several important topological properties of the friendship graph for different user categories. Our experimental results show that different user categories exhibit different topological characteristics in the friendship graph and these observed characteristics are in accordance with the expected ones based on the general definition of the four roles. Zubair Shafiq, Muhammad Usman Ilyas, Alex X. Liu, Hayder Radha |
IEEE J. Sel. Areas Commun. | 1 |
| 2013 | Large-Scale Measurement and Characterization of Cellular Machine-to-Machine TrafficabstractCellular network-based machine-to-machine (M2M) communication is fast becoming a market-changing force for a wide spectrum of businesses and applications such as telematics, smart metering, point-of-sale terminals, and home security and automation systems. In this paper, we aim to answer the following important question: Does traffic generated by M2M devices impose new requirements and challenges for cellular network design and management? To answer this question, we take a first look at the characteristics of M2M traffic and compare it to traditional smartphone traffic. We have conducted our measurement analysis using a week-long traffic trace collected from a tier-1 cellular network in the US. We characterize M2M traffic from a wide range of perspectives, including temporal dynamics, device mobility, application usage, and network performance. Our experimental results show that M2M traffic exhibits significantly different patterns than smartphone traffic in multiple aspects. For instance, M2M devices have a much larger ratio of uplink-to-downlink traffic volume, their traffic typically exhibits different diurnal patterns, they are more likely to generate synchronized traffic resulting in bursty aggregate traffic volumes, and are less mobile compared to smartphones. On the other hand, we also find that M2M devices are generally competing with smartphones for network resources in co-located geographical regions. These and other findings suggest that better protocol design, more careful spectrum allocation, and modified pricing schemes may be needed to accommodate the rise of M2M devices. Zubair Shafiq, Lusheng Ji, Alex X. Liu, Jeffrey Pang, Jia Wang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2012 | A semantics aware approach to automated reverse engineering unknown protocolsabstractExtracting the protocol message format specifications of unknown applications from network traces is important for a variety of applications such as application protocol parsing, vulnerability discovery, and system integration. In this paper, we propose ProDecoder, a network trace based protocol message format inference system that exploits the semantics of protocol messages without the executable code of application protocols. ProDecoder is based on the key insight that the n-grams of protocol traces exhibit highly skewed frequency distribution that can be leveraged for accurate protocol message format inference. In ProDecoder, we first discover the latent relationship among n-grams by first grouping protocol messages with the same semantics and then inferring message formats by keyword based clustering and cluster sequence alignment. We implemented and evaluated ProDecoder to infer message format specifications of SMB (a binary protocol) and SMTP (a textual protocol). Our experimental results show that ProDecoder accurately parses and infers SMB protocol with 100% precision and recall. For SMTP, ProDecoder achieves approximately 95% precision and recall. Yipeng Wang 0001, Xiao-chun Yun, Zubair Shafiq, Alex X. Liu, Danfeng Yao, Yongzheng Zhang 0002, Li Guo 0001 |
ICNP | 3 |
| 2012 | A large scale exploratory analysis of software vulnerability life cyclesabstractSoftware systems inherently contain vulnerabilities that have been exploited in the past resulting in significant revenue losses. The study of vulnerability life cycles can help in the development, deployment, and maintenance of software systems. It can also help in designing future security policies and conducting audits of past incidents. Furthermore, such an analysis can help customers to assess the security risks associated with software products of different vendors. In this paper, we conduct an exploratory measurement study of a large software vulnerability data set containing 46310 vulnerabilities disclosed since 1988 till 2011. We investigate vulnerabilities along following seven dimensions: (1) phases in the life cycle of vulnerabilities, (2) evolution of vulnerabilities over the years, (3) functionality of vulnerabilities, (4) access requirement for exploitation of vulnerabilities, (5) risk level of vulnerabilities, (6) software vendors, and (7) software products. Our exploratory analysis uncovers several statistically significant findings that have important implications for software development and deployment. Muhammad Shahzad 0001, Zubair Shafiq, Alex X. Liu |
ICSE | 2 |
| 2012 | Characterizing geospatial dynamics of application usage in a 3G cellular data networkabstractRecent studies on cellular network measurement have provided the evidence that significant geospatial correlations, in terms of traffic volume and application access, exist in cellular network usage. Such geospatial correlation patterns provide local optimization opportunities to cellular network operators for handling the explosive growth in the traffic volume observed in recent years. To the best of our knowledge, in this paper, we provide the first fine-grained characterization of the geospatial dynamics of application usage in a 3G cellular data network. Our analysis is based on two simultaneously collected traces from the radio access network (containing location records) and the core network (containing traffic records) of a tier-1 cellular network in the United States. To better understand the application usage in our data, we first cluster cell locations based on their application distributions and then study the geospatial dynamics of application usage across different geographical regions. The results of our measurement study present cellular network operators with fine-grained insights that can be leveraged to tune network parameter settings. Zubair Shafiq, Lusheng Ji, Alex X. Liu, Jeffrey Pang, Jia Wang 0001 |
INFOCOM | 1 |
| 2012 | A first look at cellular machine-to-machine traffic: large scale measurement and characterizationabstractCellular network based Machine-to-Machine (M2M) communication is fast becoming a market-changing force for a wide spectrum of businesses and applications such as telematics, smart metering, point-of-sale terminals, and home security and automation systems. In this paper, we aim to answer the following important question: Does traffic generated by M2M devices impose new requirements and challenges for cellular network design and management? To answer this question, we take a first look at the characteristics of M2M traffic and compare it with traditional smartphone traffic. We have conducted our measurement analysis using a week-long traffic trace collected from a tier-1 cellular network in the United States. We characterize M2M traffic from a wide range of perspectives, including temporal dynamics, device mobility, application usage, and network performance. Zubair Shafiq, Lusheng Ji, Alex X. Liu, Jeffrey Pang, Jia Wang 0001 |
SIGMETRICS | 1 |
| 2011 | A distributed and privacy preserving algorithm for identifying information hubs in social networksabstractThis paper addresses the problem of identifying the top-k information hubs in a social network. Identifying top-k information hubs is crucial for many applications such as advertising in social networks where advertisers are interested in identifying hubs to whom free samples can be given. Existing solutions are centralized and require time stamped information about pair-wise user interactions and can only be used by social network owners as only they have access to such data. Existing distributed and privacy preserving algorithms suffer from poor accuracy. In this paper, we propose a new algorithm to identify information hubs that preserves user privacy. The intuition is that highly connected users tend to have more interactions with their neighbors than less connected users. Our method can identify hubs without requiring a central entity to access the complete friendship graph. We achieve this by fully distributing the computation using the Kempe-McSherry algorithm to address user privacy concerns. To the best of our knowledge, the proposed algorithm represents an arguably first attempt that (1) uses friendship graphs (instead of interaction graphs), (2) employs a truly distributed method over friendship graphs, and (3) maintains user privacy by not requiring them to disclose their friend associations and interactions, for identifying information hubs in social networks. We evaluate the effectiveness of our proposed technique using a real-world Facebook data set containing about 3.1 million users and more than 23 million friendship links. The results of our experiments show that our algorithm is 50% more accurate than existing distributed algorithms. Results also show that the proposed algorithm can estimate the rank of the top-k information hubs users more accurately than existing approaches. Muhammad Usman Ilyas, Zubair Shafiq, Alex X. Liu, Hayder Radha |
INFOCOM | 2 |
| 2011 | A Random Walk Approach to Modeling the Dynamics of the Blogosphere
Zubair Shafiq, Alex X. Liu |
Networking (1) | 1 |
| 2011 | Characterizing and modeling internet traffic dynamics of cellular devicesabstractUnderstanding Internet traffic dynamics in large cellular networks is important for network design, troubleshooting, performance evaluation, and optimization. In this paper, we present the results from our study, which is based upon a week-long aggregated flow level mobile device traffic data collected from a major cellular operator's core network. In this study, we measure and characterize the spatial and temporal dynamics of mobile Internet traffic. We distinguish our study from other related work by conducting the measurement at a larger scale and exploring mobile data traffic patterns along two new dimensions -- device types and applications that generate such traffic patterns. Based on the findings of our measurement analysis, we propose a Zipf-like model to capture the volume distribution of application traffic and a Markov model to capture the volume dynamics of aggregate Internet traffic. We further customize our models for different device types using an unsupervised clustering algorithm to improve prediction accuracy. Zubair Shafiq, Lusheng Ji, Alex X. Liu, Jia Wang 0001 |
SIGMETRICS | 1 |
| 2009 | Evolvable malwareabstractThe concept of artificial evolution has been applied to numerous real world applications in different domains. In this paper, we use this concept in the domain of virology to evolve computer viruses. We call this domain as "Evolvable Malware". To this end, we propose an evolutionary framework that consists of three modules: (1) a code analyzer that generates a high-level genotype representation of a virus from its machine code, (2) a genetic algorithm that uses the standard selection, cross-over and mutation operators to evolve viruses, and (3) the code generator converts the genotype of a newly evolved virus to its machinelevel code. In this paper, we validate the notion of evolution in viruses on a well-known virus family, called Bagle. The results of our proof-of-concept study show that we have successfully evolved new viruses-previously unknown and known-variants of Bagle-starting from a random population of individuals. To the best of our knowledge, this is the first empirical work on evolution of computer viruses. In future, we want to improve this proof-of-concept framework into a full-blown virus evolution engine. Sadia Noreen, Shafaq Murtaza, Zubair Shafiq, Muddassar Farooq |
GECCO | 3 |
| 2009 | Are evolutionary rule learning algorithms appropriate for malware detection?abstractIn this paper, we evaluate the performance of ten well-known evolutionary and non-evolutionary rule learning algorithms. The comparative study is performed on a real-world classification problem of detecting malicious executables. The executable dataset, used in this study, consists of a total of 189 attributes which are statically extracted from the executables of Microsoft Windows operating system. In our study, we evaluate the performance of rule learning algorithms with respect to four metrics: (1) classification accuracy, (2) the number of rules in the developed rule set, (3) the comprehensibility of the generated rules, and (4) the processing overhead of the rule learning process. The results of our study highlight important shortcomings in evolutionary rule learning classifiers that render them infeasible for deployment in a real-world malware detection system. Zubair Shafiq, S. Momina Tabish, Muddassar Farooq |
GECCO | 1 |
| 2009 | On the Inefficient Use of Entropy for Anomaly Detection
Mobin Javed, Ayesha Binte Ashfaq, Zubair Shafiq, Syed Ali Khayam |
RAID | 3 |
| 2009 | Using Formal Grammar and Genetic Operators to Evolve Malware
Sadia Noreen, Shafaq Murtaza, Zubair Shafiq, Muddassar Farooq |
RAID | 3 |
| 2009 | PE-Miner: Mining Structural Information to Detect Malicious Executables in Realtime
Zubair Shafiq, S. Momina Tabish, Fauzan Mirza, Muddassar Farooq |
RAID | 1 |
| 2009 | Fuzzy case-based reasoning for facial expression recognition
Aasia Khanum, Muid Mufti, Muhammad Younus Javed, Zubair Shafiq |
Fuzzy Sets Syst. | 4 |
| 2008 | Embedded Malware Detection Using Markov n-Grams
Zubair Shafiq, Syed Ali Khayam, Muddassar Farooq |
DIMVA | 1 |
| 2008 | Improving accuracy of immune-inspired malware detectors by using intelligent featuresabstractIn this paper, we show that a Bio-inspired classifier's accuracy can be dramatically improved if it operates on intelligent features. We propose a novel set of intelligent features for the well-known problem of malware portscan detection. We compare the performance of three well-known Bio-inspired classifiers operating on the proposed intelligent features: (1) Real Valued Negative Selection (RVNS) based on the adaptive immune system; (2) Dendritic Cell Algorithm (DCA) based on the innate immune system; and (3) Adaptive Neuro Fuzzy Inference System (ANFIS). To empirically evaluate the improvements provided by the intelligent features, we use a network traffic dataset collected on diverse endpoints for a period of 12 months. The endpoints' traffic is infected with well-known malware. For unbiased performance comparison, we also include a machine learning algorithm, Support Vector Machine (SVM), and two state-of-the-art statistical malware detectors, Rate-Limiting (RL) and Maximum-Entropy (ME). To the best of our knowledge, this is the first study in which RVNS and DCA are not only compared with each other but also with several other classifiers on a comprehensive real-world dataset. The experimental results indicate that our proposed features significantly improve the TP rate and FP rate of both RVNS and DCA. Zubair Shafiq, Syed Ali Khayam, Muddassar Farooq |
GECCO | 1 |
| 2007 | Extended thymus action for improving response of AIS based NID system against malicious trafficabstractArtificial immune systems (AISs) are being increasingly utilized to develop network intrusion detection (NID) systems. The fundamental reason for their success in NID is their ability to learn normal behavior of a network system and then differentiate it from an anomalous behavior. As a result, they can detect a majority of innovative attacks. In comparison, classical signature based systems fail to detect innovative attacks. Light Weight Intrusion Detection System (LISYS) provides the basic framework for AIS based NID systems. This framework has been improved incrementally, including incorporation of thymus action, since it was first developed. In this paper, we have extended the basic thymus action model, which provides immature detectors with multiple chances to develop tolerization to normal. However, AIS is prone to successful attacks by malicious traffic which appears similar to the normal traffic. This results in high number of false positives. In this paper, we present a mathematical model of malicious traffic for TCP-SYN flood based distributed denial of services (DDoS) attacks. This model is used to generate different sets of malicious traffic. These sets are used for performance comparison of the proposed extended thymus action with the simple thymus action model. The results of our experiments demonstrate that the extended model has significantly reduced the number of false positives. Zubair Shafiq, Mehrin Kiani, Bisma Hashmi, Muddassar Farooq |
IEEE Congress on Evolutionary Computation | 1 |
| 2007 | Extended thymus action for reducing false positives in ais based network intrusion detection systemsabstractOne of the major problems faced by anomaly based Network Intrusion Detection (NID) systems is the high number of false positives. False positives refer to the false detection of normal behavior as malicious behavior. Artificial Immune Systems (AISs) also fall under the category of anomaly based-NID systems. AIS presented in this paper is as a victim-end filter, consisting of detectors distributed on the network, which distinguishes normal traffic from malicious traffic. In this work, we focus on TCP-SYN flood based Distributed Denial of Services (DDoS) attacks. Light Weight Intrusion Detection System (LISYS) provides the basic framework for AIS based NID systems. AISs normally utilize the negative selection algorithm in thymus action to tolerize the detectors to normal traffic so they may not detect normal traffic as malicious traffic. We propose and implement `extended thymus action' model to improve this characteristic of AIS. Results verify that our model significantly reduces false positives which is a major concern in anomaly-based NID systems. Zubair Shafiq, Mehrin Kiani, Bisma Hashmi, Muddassar Farooq |
GECCO | 1 |