VLDB 2026 Research / reviewers in the wild / expert
Fareed Zaffar
dblp:59/3605 · also Muhammad Fareed Zaffar
· DBLP profile ↗
36ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-1270-5513ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 12 · 1 first-author · 4 since 2021Computer networks · 6 · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | (Mis-)Informed Consent: Predatory Apps and the Exploitation of Populations with Limited Literacy
Muhammad Muneeb Pervez, Muhammad Qasim Atiq Ullah, Ibrahim Ahmed Khan, Roshnik Rahat, Fareed Zaffar, Rashid Tahir, Talal Rahwan, Yasir Zaki |
WWW | 5 |
| 2025 | Multitask-Bench: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuningabstractRecent breakthroughs in Large Language Models (LLMs) have led to their adoption across a wide range of tasks, ranging from code generation to machine translation and sentiment analysis, etc. Red teaming/Safety alignment efforts show that fine-tuning models on benign (non-harmful) data could compromise safety. However, it remains unclear to what extent this phenomenon is influenced by different variables, including fine-tuning task, model calibrations, etc. This paper explores the task-wise safety degradation due to fine-tuning on downstream tasks such as summarization, code generation, translation, and classification across various calibration. Our results reveal that: 1) Fine-tuning LLMs for code generation and translation leads to the highest degradation in safety guardrails. 2) LLMs generally have weaker guardrails for translation and classification, with 73-92% of harmful prompts answered, across baseline and other calibrations, falling into one of two concern categories. 3) Current solutions, including guards and safety tuning datasets, lack cross-task robustness. To address these issues, we developed a new multitask safety dataset effectively reducing attack success rates across a range of tasks without compromising the model’s overall helpfulness. Our work underscores the need for generalized alignment measures to ensure safer and more robust models. Essa Jan, Nouar AlDahoul, Moiz Ali, Fareed Zaffar, Yasir Zaki |
COLING | 5 |
| 2025 | SCALPEL: Structured Content Access Logging and Pruning for Efficient LayoutsabstractContainerizing scientific workflows helps ensure their reproducibility. Including all data required for deterministic re-execution aids the process. In data-intensive climate science and other high-performance computing domains, workflows routinely process large-scale data archives through parallel frameworks such as Dask. Packaging these archives inflates container images, with vast swaths of data that are never accessed by the application driving up transfer costs and hindering deployment. We present SCALPEL, a framework for semantic carving - selectively retaining only the data an application actually consumes during analysis, while excluding data that is accessed but not used.We provide two complementary carving modes, each operating at a different level of observability: (1) value-level carving, which traces data access at the resolution of individual rows (in tabular datasets) or specific multidimensional elements (in array datasets), enabling deterministic re-execution with precisely the same inputs; and (2) partition/chunk-level carving, a configurable alternative that tracks data access at the coarser granularity of table partitions or array chunks, facilitating flexible re-execution that can accommodate additional inputs within these broader segments. By interposing on HDF5 I/O operations and Dask’s execution layer, both widely adopted technologies for large-scale data storage and parallel computation, we capture application-specific data access patterns, achieving reductions in container size by several orders of magnitude. We demonstrate these reductions in data-intensive scientific workflows, including precipitation-driven climate modeling and other geophysical workloads. Raffay Atiq, Ashish Gehani, Tanu Malik, Fareed Zaffar |
eScience | 4 |
| 2023 | SoK: A Tale of Reduction, Security, and Correctness - Evaluating Program Debloating Paradigms and Their Compositions
Muaz Ali, M. Faraz Karim, Ayesha Naeem, Rukhshan Haroon, Huzaifah Nadeem, Waseem Sabir, Fahad Shaon, Fareed Zaffar, Vinod Yegneswaran, Ashish Gehani, Sazzadur Rahaman |
ESORICS (4) | 10 |
| 2022 | PACED: Provenance-based Automated Container Escape DetectionabstractThe security of container-based microservices relies heavily on the isolation of operating system resources that is provided by namespaces. However, vulnerabilities exist in the isolation of containers that may be exploited by attackers to gain access to the host. These are commonly referred to as container escape attacks. While prior work has identified vulnerabilities in namespace isolation, no general container escape detection and warning system has been presented. We present Paced, a novel, realtime system to detect container-escape attacks. We define what constitutes a cross-namespace event and how such events can be used to detect a container escape attack. We develop a provenance-based approach to isolate cross-namespace events and propose a rule—privileged_flow—to detect attacks on Docker and Kubernetes environments. We evaluate our detection method on a suite of contemporary CVEs with container escape exploits, bad container configurations, and benchmarks. Paced achieves near-perfect accuracy with no false negatives. We release our implementation and datasets as free, open-source software. Mashal Abbas, Shahpar Khan, Abdul Monum, Fareed Zaffar, Rashid Tahir, David M. Eyers, Hassaan Irshad, Ashish Gehani, Vinod Yegneswaran, Thomas Pasquier |
IC2E | 4 |
| 2022 | To Block or Not to Block: Accelerating Mobile Web Pages On-The-Fly Through JavaScript ClassificationabstractThe increasing complexity of JavaScript (JS) in modern mobile web pages has become a performance bottleneck for low-end mobile phone users, especially in developing regions. In this paper we propose SlimWeb, a novel approach that automatically derives lightweight versions of mobile web pages on-the-fly by eliminating non-essential JavaScript that does not impact the core page content and interactive functionality. SlimWeb consists of a JavaScript classification service powered by a supervised Machine Learning (ML) model that provides insights into each JavaScript element embedded in a web page. SlimWeb aims to improve the web browsing experience by predicting the class of each element, such that essential elements are preserved and non-essential elements are blocked by the browsers using the service. We motivate SlimWeb’s core design via a preference survey where 306 users overwhelmingly preferred having faster page load times over fetching various categories of non-essential JavaScript. We evaluate SlimWeb across 500 popular web pages in a developing region on real cellular networks, along with a user experience study with 20 real-world users and a usage willingness survey of 588 users. Evaluation results show that SlimWeb achieves 50% reduction in page load time compared to the original pages, and more than 30% reduction compared to competing solutions, while achieving high similarity scores to the original pages measured via a qualitative evaluation study with 62 users. SlimWeb improves the overall user experience metric (defined by Google Lighthouse combining first contentful paint, time to interactive, speed index) by more than 60% compared to the original pages, while maintaining 90-100% of the visual and functional components of most pages. Moumena Chaqfeh, Waleed Hashmi, Patrick Inshuti, Manesha Ramesh, Matteo Varvello, Lakshminarayanan Subramanian, Fareed Zaffar, Yasir Zaki |
ICTD | 8 |
| 2022 | Are Proactive Interventions for Reddit Communities Feasible?
Hussam Habib, Maaz Bin Musa, Fareed Zaffar, Rishab Nithyanand |
ICWSM | 3 |
| 2022 | Trimmer: Context-Specific Code ReductionabstractWe present Trimmer, a state-of-the-art tool for reducing code size. Trimmer reduces code sizes by specializing programs with respect to constant inputs provided by developers. The static data can be provided as command-line options or through configuration files. The constants define the features that must be retained, which in turn determine the features that are unused in a specific deployment (and can therefore be removed). Trimmer includes sophisticated compiler transformations for input specialization, supports precise yet efficient context-sensitive inter-procedural constant propagation, and introduces a custom loop unroller. Trimmer is easy-to-use and extensively parameterized. We discuss how Trimmer can be configured by developers to explicitly trade analysis precision and specialization time. We also provide a high-level description of Trimmer’s static analysis passes. The source code is publicly available at: https://github.com/ashish-gehani/Trimmer. A video demonstration can be found here: https://youtu.be/6pAuJ68INnI. Aatira Anum Ahmad, Mubashir Anwar, Hashim Sharif, Ashish Gehani, Fareed Zaffar |
ASE | 5 |
| 2022 | Forensic Analysis of Configuration-based Attacks
Muhammad Adil Inam, Wajih Ul Hassan, Ali Ahad, Adam Bates 0001, Rashid Tahir, Tianyin Xu, Fareed Zaffar |
NDSS | 7 |
| 2022 | An internet of secure and private things: A service-oriented architectureabstractLow-cost networked IoT devices are fast becoming commonplace. From implanted medical devices to motion-activated surveillance cameras and from driverless smart cars to voice-operated home management systems, IoT devices continue to permeate further and deeper into our lives. However, this widespread adoption of IoT devices has also given rise to a wide range of security, privacy and trust issues that are unique to the IoT ecosystem. Conventional solutions are ill-suited for the IoT domain due to limited resources, network dynamics, and evolving trust boundaries. Hence, novel mechanisms are needed to address the specific challenges of the IoT landscape. To this end, we propose a user-centric cloud-based service that allows device owners to have fine-grained control over what kind and how much data is shared through their IoT devices. Our scheme builds on top of Intel Software Guard Extensions (SGX) to instantiate secure virtual clones (shadows) of actual devices in the cloud, substantially reducing the attack surface for IoT networks. Furthermore, a scalable infrastructure in the cloud allows us to deploy sophisticated policy enforcement and data scrubbing mechanisms on a per application basis giving users explicit control over data sharing. The presented approach requires little effort on part of device vendors and users as the service provider handles the bulk of the work. We demonstrate the effectiveness of our approach empirically by implementing the service on SGX hardware and deploying advanced data cleansing policies on device-generated data. Ahmad Showail, Rashid Tahir, Fareed Zaffar, Muhammad Haris Noor, Mohammed Alkhatib |
Comput. Secur. | 3 |
| 2022 | Analyzing the Feasibility and Generalizability of Fingerprinting Internet of Things DevicesabstractAbstract In recent years, we have seen rapid growth in the use and adoption of Internet of Things (IoT) devices. However, some loT devices are sensitive in nature, and simply knowing what devices a user owns can have security and privacy implications. Researchers have, therefore, looked at fingerprinting loT devices and their activities from encrypted network traffic. In this paper, we analyze the feasibility of fingerprinting IoT devices and evaluate the robustness of such fingerprinting approach across multiple independent datasets — collected under different settings. We show that not only is it possible to effectively fingerprint 188 loT devices (with over 97% accuracy), but also to do so even with multiple instances of the same make-and-model device. We also analyze the extent to which temporal, spatial and data-collection-methodology differences impact fingerprinting accuracy. Our analysis sheds light on features that are more robust against varying conditions. Lastly, we comprehensively analyze the performance of our approach under an open-world setting and propose ways in which an adversary can enhance their odds of inferring additional information about unseen devices (e.g., similar devices manufactured by the same company). Dilawer Ahmed, Anupam Das 0001, Fareed Zaffar |
Proc. Priv. Enhancing Technol. | 3 |
| 2022 | Trimmer: An Automated System for Configuration-Based Software DebloatingabstractSoftware bloat has negative implications for security, reliability, and performance. To counter bloat, we proposeTrimmer, a static analysis-based system for pruning unused functionality.Trimmerremoves code that is unused with respect to user-provided command-line arguments and application-specific configuration files.Trimmeruses concrete memory tracking and a custom inter-procedural constant propagation analysis that facilitates dead code elimination. Our system supports both context-sensitive and context-insensitive constant propagation. We show that context-sensitive constant propagation is important for effective software pruning in most applications. We introducesparse constant propagationthat performs constant propagation only for configuration-hosting variables and show that it performs better (higher code size reductions) compared to constant propagation for all program variables. Overall, our results show thatTrimmerreduces binary sizes for real-world programs with reasonable analysis times. Across 20 evaluated programs, we observe a mean binary size reduction of 22.7 percent and a maximum reduction of 62.7 percent. For 5 programs, we observe performance speedups ranging from 5 to 53 percent. Moreover, we show that winnowing software applications can reduce the program attack surface by removing code that contains exploitable vulnerabilities. We find that debloating usingTrimmerremoves CVEs in 4 applications. Aatira Anum Ahmad, Abdul Rafae Noor, Hashim Sharif, Usama Hameed, Shoaib Asif, Mubashir Anwar, Ashish Gehani, Fareed Zaffar, Junaid Haroon Siddiqui |
IEEE Trans. Software Eng. | 8 |
| 2021 | Accelerating Fourier and Number Theoretic Transforms using Tensor Cores and Warp ShufflesabstractThe discrete Fourier transform (DFT) and its specialized case, the number theoretic transform (NTT), are two important mathematical tools having applications in several areas of science and engineering. However, despite their usefulness and utility, their adoption continues to be a challenge as computing the DFT of a signal can be a time-consuming and expensive operation. To speed things up, fast Fourier transform (FFT) algorithms, which are reduced-complexity formulations for computing the DFT of a sequence, have been proposed and implemented for traditional processors and their corresponding instruction sets. With the rise of GPUs, NVIDIA introduced its own FFT computation library called cuFFT, which leverages the power of GPUs to compute the DFT. However, as this paper demonstrates, there is a lot of room for improvement to accelerate the FFT and NTT algorithms on modern GPUs by utilizing specialized operations and architectural advancements. In particular, we present four major types of optimizations that leverage tensor cores and the warp-shuffle instruction. Through extensive evaluations, we show that our approach consistently outperforms existing GPU-based implementations with a speedup of up to 4× for NTT and a speed of up to 1.5× for FFT. Sultan Durrani, Muhammad Saad Chughtai, Mert Hidayetoglu, Rashid Tahir, Abdul Dakkak, Lawrence Rauchwerger, Fareed Zaffar, Wen-Mei W. Hwu |
PACT | 7 |
| 2021 | Seeing is Believing: Exploring Perceptual Differences in DeepFake VideosabstractWith AI on the boom, DeepFakes have emerged as a tool with a massive potential for abuse. The hyper-realistic imagery of these manipulated videos coupled with the expedited delivery models of social media platforms gives deception, propaganda, and disinformation an entirely new meaning. Hence, raising awareness about DeepFakes and how to accurately flag them has become imperative. However, given differences in human cognition and perception, this is not straightforward. In this paper, we perform an investigative user study and also analyze existing AI detection algorithms from the literature to demystify the unknowns that are at play behind the scenes when detecting DeepFakes. Based on our findings, we design a customized training program to improve detection and evaluate on a treatment group of low-literate population, which is most vulnerable to DeepFakes. Our results suggest that, while DeepFakes are becoming imperceptible, contextualized education and training can help raise awareness and improve detection. Rashid Tahir, Brishna Batool, Hira Jamshed, Mahnoor Jameel, Mubashir Anwar, Muhammad Adeel Zaffar, Fareed Zaffar |
CHI | 8 |
| 2021 | Through the Looking Glass: Learning to Attribute Synthetic Text Generated by Language ModelsabstractShaoor Munir, Brishna Batool, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Shaoor Munir, Brishna Batool, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar |
EACL | 5 |
| 2021 | TrackerSift: untangling mixed tracking and functional web resourcesabstractTrackers have recently started to mix tracking and functional resources to circumvent privacy-enhancing content blocking tools. Such mixed web resources put content blockers in a bind: risk breaking legitimate functionality if they act and risk missing privacy-invasive advertising and tracking if they do not. In this paper, we propose TrackerSift to progressively classify and untangle mixed web resources (that combine tracking and legitimate functionality) at multiple granularities of analysis (domain, hostname, script, and method). Using TrackerSift, we conduct a large-scale measurement study of such mixed resources on 100K websites. We find that more than 17% domains, 48% hostnames, 6% scripts, and 9% methods observed in our crawls combine tracking and legitimate functionality. While mixed web resources are prevalent across all granularities, TrackerSift is able to attribute 98% of the script-initiated network requests to either tracking or functional resources at the finest method-level granularity. Our analysis shows that mixed resources at different granularities are typically served from CDNs or as in-lined and bundled scripts, and that blocking them indeed results in breakage of legitimate functionality. Our results highlight opportunities for finer-grained content blocking to remove mixed resources without breaking legitimate functionality. Abdul Haddi Amjad, Danial Saleem, Muhammad Ali Gulzar, Zubair Shafiq, Fareed Zaffar |
Internet Measurement Conference | 5 |
| 2020 | Phishcasting: Deep Learning for Time Series Forecasting of Phishing AttacksabstractPhishing attacks remain pervasive and continue to be a source of significant monetary loss, identity theft, and malware. One of the challenges is that in most organizational settings, the detection paradigm is inherently about identifying and reacting to threats in real-time, as they are unfolding. As a way to complement these efforts with greater foresight, we introduce the idea of phishcasting — forecasting of phishing threat levels weeks or months into the future. Given that phishing attack volume time series data is noisy and devoid of traditional seasonal and cyclical trends, we extend the time series forecasting framework to utilize multiple time series, auxiliary information and alternate representations. We also introduce CoT-Net, a flexible, end-to-end CNN-LSTM based deep learning method for forecasting of complex phishing attack volume time series. CoT-Net uses time series embeddings to uncover correlations between organizational attack patterns within and across industry sectors. Using a publicly available test bed featuring multiple organizations’ attack volume over time, we find CoT-Net to outperform most state-of-the-art time series forecasting methods. By showing that phishcasting might be possible and practical, our work has important proactive implications for cybersecurity. Syed Hasan Amin Mahmood, Syed Mustafa Ali Abbasi, Ahmed Abbasi, Fareed Zaffar |
ISI | 4 |
| 2020 | Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the WebabstractData generated by web crawlers has formed the basis for much of our current understanding of the Internet. However, not all crawlers are created equal and crawlers generally find themselves trading off between computational overhead, developer effort, data accuracy, and completeness. Therefore, the choice of crawler has a critical impact on the data generated and knowledge inferred from it. In this paper, we conduct a systematic study of the trade-offs presented by different crawlers and the impact that these can have on various types of measurement studies. We make the following contributions: First, we conduct a survey of all research published since 2015 in the premier security and Internet measurement venues to identify and verify the repeatability of crawling methodologies deployed for different problem domains and publication venues. Next, we conduct a qualitative evaluation of a subset of all crawling tools identified in our survey. This evaluation allows us to draw conclusions about the suitability of each tool for specific types of data gathering. Finally, we present a methodology and a measurement framework to empirically highlight the differences between crawlers and how the choice of crawler can impact our understanding of the web. Suleman Ahmad, Muhammad Daniyal Pirwani Dar, Fareed Zaffar, Narseo Vallina-Rodriguez, Rishab Nithyanand |
WWW | 3 |
| 2020 | CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social NetworksabstractOnline social networks support a vibrant ecosystem of third-party apps that get access to personal information of a large number of users. Despite several recent high-profile incidents, methods to systematically detect data misuse by third-party apps on online social networks are lacking. We propose CanaryTrap to detect misuse of data shared with third-party apps. CanaryTrap associates a honeytoken to a user account and then monitors its unrecognized use via different channels after sharing it with the third-party app. We design and implement CanaryTrap to investigate misuse of data shared with third-party apps on Facebook. Specifically, we share the email address associated with a Facebook account as a honeytoken by installing a third-party app. We then monitor the received emails and use Facebook’s ad transparency tool to detect any unrecognized use of the shared honeytoken. Our deployment of CanaryTrap to monitor 1,024 Facebook apps has uncovered multiple cases of misuse of data shared with third-party apps on Facebook including ransomware, spam, and targeted advertising. Shehroze Farooqi, Maaz Bin Musa, Zubair Shafiq, Fareed Zaffar |
Proc. Priv. Enhancing Technol. | 4 |
| 2019 | Bringing the kid back into YouTube kids: detecting inappropriate content on video streaming platformsabstractWith the advent of child-centric content-sharing platforms, such as YouTube Kids, thousands of children, from all age groups are consuming gigabytes of content on a daily basis. With PBS Kids, Disney Jr. and countless others joining in the fray, this consumption of video data stands to grow further in quantity and diversity. However, it has been observed increasingly that content unsuitable for children often slips through the cracks and lands on such platforms. To investigate this phenomenon in more detail, we collect a first of its kind dataset of inappropriate videos hosted on such children-focused apps and platforms. Alarmingly, our study finds that there is a noticeable percentage of such videos currently being watched by kids with some inappropriate videos having millions of views already. To address this problem, we develop a deep learning architecture that can flag such videos and report them. Our results show that the proposed system can be successfully applied to various types of animations, cartoons and CGI videos to detect any inappropriate content within them. Rashid Tahir, Mohammad Hammas Saeed, Shiza Ali, Fareed Zaffar, Christo Wilson |
ASONAM | 5 |
| 2019 | The Browsers Strike Back: Countering Cryptojacking and Parasitic Miners on the WebabstractWith the recent boom in the cryptocurrency market, hackers have been on the lookout to find novel ways of commandeering users' machine for covert and stealthy mining operations. In an attempt to expose such under-the-hood practices, this paper explores the issue of browser cryptojacking, whereby miners are secretly deployed inside browser code without the knowledge of the user. To this end, we analyze the top 50k websites from Alexa and find a noticeable percentage of sites that are indulging in this exploitative exercise often using heavily obfuscated code. Furthermore, mining prevention plug-ins, such as NoMiner, fail to flag such cleverly concealed instances. Hence, we propose a machine learning solution based on hardware-assisted profiling of browser code in real-time. A fine-grained micro-architectural footprint allows us to classify mining applications with >99% accuracy and even flags them if the mining code has been heavily obfuscated or encrypted. We build our own browser extension and show that it outperforms other plug-ins. The proposed design has negligible overhead on the user's machine and works for all standard off-the-shelf CPUs. Rashid Tahir, Sultan Durrani, Mohammad Hammas Saeed, Fareed Zaffar, Muhammad Saqib Ilyas |
INFOCOM | 5 |
| 2019 | Quantity vs. Quality: Evaluating User Interest Profiles Using Ad Preference Managers
Muhammad Ahmad Bashir, Maryam Shahid, Fareed Zaffar, Christo Wilson |
NDSS | 4 |
| 2019 | A Girl Has No Name: Automated Authorship Obfuscation using Mutant-XabstractAbstract Stylometric authorship attribution aims to identify an anonymous or disputed document’s author by examining its writing style. The development of powerful machine learning based stylometric authorship attribution methods presents a serious privacy threat for individuals such as journalists and activists who wish to publish anonymously. Researchers have proposed several authorship obfuscation approaches that try to make appropriate changes (e.g. word/phrase replacements) to evade attribution while preserving semantics. Unfortunately, existing authorship obfuscation approaches are lacking because they either require some manual effort, require significant training data, or do not work for long documents. To address these limitations, we propose a genetic algorithm based random search framework called Mutant-X which can automatically obfuscate text to successfully evade attribution while keeping the semantics of the obfuscated text similar to the original text. Specifically, Mutant-X sequentially makes changes in the text using mutation and crossover techniques while being guided by a fitness function that takes into account both attribution probability and semantic relevance. While Mutant-X requires black-box knowledge of the adversary’s classifier, it does not require any additional training data and also works on documents of any length. We evaluate Mutant-X against a variety of authorship attribution methods on two different text corpora. Our results show that Mutant-X can decrease the accuracy of state-of-the-art authorship attribution methods by as much as 64% while preserving the semantics much better than existing automated authorship obfuscation approaches. While Mutant-X advances the state-of-the-art in automated authorship obfuscation, we find that it does not generalize to a stronger threat model where the adversary uses a different attribution classifier than what Mutant-X assumes. Our findings warrant the need for future research to improve the generalizability (or transferability) of automated authorship obfuscation approaches. Asad Mahmood, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar |
Proc. Priv. Enhancing Technol. | 5 |
| 2018 | It's All in the Name: Why Some URLs are More Vulnerable to TyposquattingabstractTyposquatting is a blackhat practice that relies on human error and low-cost domain registrations to hijack legitimate traffic from well-established websites. The technique is typically used for phishing, driving traffic towards competitors or disseminating indecent or malicious content and as such remains a concern for businesses. We take a fresh new look at this well-studied phenomenon to explore why some URLs are more vulnerable to typing mistakes than others. We explore the relationship between human hand anatomy, keyboard layouts and typing mistakes using various URL datasets. We create an extensive user-centric typographical model and compute a Hardness Quotient (likelihood of mistyping) for each URL using a quantitative measure for finger and hand effort. Furthermore, our model predicts the most likely typos for each URL which can then be defensively registered. Cross-validation against actual URL and DNS datasets suggests that this is a meaningful and effective defense mechanism. Rashid Tahir, Ali Raza 0003, Jehangir Kazi, Fareed Zaffar, Chris Kanich, Matthew Caesar 0001 |
INFOCOM | 5 |
| 2018 | TRIMMER: application specialization for code debloatingabstractWith the proliferation of new hardware architectures and ever-evolving user requirements, the software stack is becoming increasingly bloated. In practice, only a limited subset of the supported functionality is utilized in a particular usage context, thereby presenting an opportunity to eliminate unused features. In the past, program specialization has been proposed as a mechanism for enabling automatic software debloating. In this work, we show how existing program specialization techniques lack the analyses required for providing code simplification for real-world programs. We present an approach that uses stronger analysis techniques to take advantage of constant configuration data, thereby enabling more effective debloating. We developed Trimmer, an application specialization tool that leverages user-provided configuration data to specialize an application to its deployment context. The specialization process attempts to eliminate the application functionality that is unused in the user-defined context. Our evaluation demonstrates Trimmer can effectively reduce code bloat. For 13 applications spanning various domains, we observe a mean binary size reduction of 21% and a maximum reduction of 75%. We also show specialization reduces the surface for code-reuse attacks by reducing the number of exploitable gadgets. For the evaluated programs, we observe a 20% mean reduction in the total gadget count and a maximum reduction of 87%. Hashim Sharif, Muhammad Abubakar, Ashish Gehani, Fareed Zaffar |
ASE | 4 |
| 2018 | Using SGX-Based Virtual Clones for IoT SecurityabstractWidespread permeation of IoT devices into our daily lives has created a diverse spectrum of security and privacy concerns unique to the IoT ecosystem. Conventional host and network security mechanisms fail to address these issues due to resource constraints, ad-hoc network models and vendor-centric data collection and sharing policies. Hence, there is a need to redesign the IoT infrastructure to secure both the device and the data. To this end, we propose a design where users are in the driving seat, devices are less exposed and data sharing models are flexible and fine-grained. Our proposal comprises hardware-secured data banks based on Intel Software Guard Extensions (SGX) to house the data in clouds without the need to trust the cloud provider. Virtual clones (shadows) of devices running on top of these data banks serve as competent proxies of actual IoT devices hiding away device weaknesses. The proposed infrastructure is scalable and robust and serves as a good first step for the community to build on and improve. Rashid Tahir, Ali Raza 0003, Fareed Zaffar, Faizan Ul Ghani, Mubeen Zulfiqar |
NCA | 3 |
| 2018 | Detecting and Defending Against Certificate Attacks with Origin-Bound CAPTCHAs
Adil Ahmad, Vinod Yegneswaran, Fareed Zaffar |
SecureComm (2) | 5 |
| 2017 | An Anomaly Detection Fabric for Clouds Based on Collaborative VM CommunitiesabstractThe vast attack surface of clouds presents a challenge in deploying scalable and effective defenses. Traditional security mechanisms, which work from inside the VM fail to provide strong protection as attackers can bypass them easily. The only available option is to provide security from the layer below the VM i.e., the hypervisor. Previous works that attempt to secure VMs from "outside" either incur substantial space or compute overheads making them slow and impractical or require modifications to the OS or the application codebase. To address these issues, we propose an anomaly detection fabric for clouds based on system call monitoring, which compresses the stream of system calls at their source making the system scalable and near real-time. Our system requires no modifications to the guest OS or the application making it ideal for the data center setting. Additionally, for robust and early detection of threats, we leverage the notion of VM/container communities that share information about attacks in their early stages to provide immunity to the entire deployment. We make certain aspects of the system flexible so that vendors can tune metrics to offer customized protection to clients based on their workload types. Detailed evaluation on a prototype implementation on KVM substantiates our claims. Rashid Tahir, Matthew Caesar 0001, Ali Raza 0003, Mazhar Naqvi, Fareed Zaffar |
CCGrid | 5 |
| 2017 | Accurate Detection of Automatically Spun Content via Stylometric AnalysisabstractSpammers use automated content spinning techniques to evade plagiarism detection by search engines. Text spinners help spammers in evading plagiarism detectors by automatically restructuring sentences and replacing words or phrases with their synonyms. Prior work on spun content detection relies on the knowledge about the dictionary used by the text spinning software. In this work, we propose an approach to detect spun content and its seed without needing the text spinner's dictionary. Our key idea is that text spinners introduce stylometric artifacts that can be leveraged for detecting spun documents. We implement and evaluate our proposed approach on a corpus of spun documents that are generated using a popular text spinning software. The results show that our approach can not only accurately detect whether a document is spun but also identify its source (or seed) document - all without needing the dictionary used by the text spinner. Usman Shahid, Shehroze Farooqi, Raza Ahmad, Zubair Shafiq, Padmini Srinivasan, Fareed Zaffar |
ICDM | 6 |
| 2017 | Measuring and mitigating oauth access token abuse by collusion networksabstractWe uncover a thriving ecosystem of large-scale reputation manipulation services on Facebook that leverage the principle of collusion. Collusion networks collect OAuth access tokens from colluding members and abuse them to provide fake likes or comments to their members. We carry out a comprehensive measurement study to understand how these collusion networks exploit popular third-party Facebook applications with weak security settings to retrieve OAuth access tokens. We infiltrate popular collusion networks using honeypots and identify more than one million colluding Facebook accounts by "milking" these collusion networks. We disclose our findings to Facebook and collaborate with them to implement a series of countermeasures that mitigate OAuth access token abuse without sacrificing application platform usability for third-party developers. These countermeasures remained in place until April 2017, after which Facebook implemented a set of unrelated changes in its infrastructure to counter collusion networks. We are the first to report and effectively mitigate large-scale OAuth access token abuse in the wild. Shehroze Farooqi, Fareed Zaffar, Nektarios Leontiadis, Zubair Shafiq |
Internet Measurement Conference | 2 |
| 2017 | Mining on Someone Else's Dime: Mitigating Covert Mining Operations in Clouds and Enterprises
Rashid Tahir, Muhammad Huzaifa, Anupam Das 0001, Mohammad Ahmad, Carl A. Gunter, Fareed Zaffar, Matthew Caesar 0001, Nikita Borisov |
RAID | 6 |
| 2016 | Malware Slums: Measurement and Analysis of Malware on Traffic ExchangesabstractAuto-surf and manual-surf traffic exchanges are an increasingly popular way of artificially generating website traffic. Previous research in this area has focused on the makeup, usage, and monetization of underground traffic exchanges. In this paper, we analyze the role of traffic exchanges as a vector for malware propagation. We conduct a measurement study of nine auto-surf and manual-surf traffic exchanges over several months. We present a first of its kind analysis of the different types of malware that are propagated through these traffic exchanges. We find that more than 26% of the URLs surfed on traffic exchanges contain malicious content. We further analyze different categories of malware encountered on traffic exchanges, including blacklisted domains, malicious JavaScript, malicious Flash, and malicious shortened URLs. Salman Yousaf, Umar Iqbal 0002, Shehroze Farooqi, Raza Ahmad, Zubair Shafiq, Fareed Zaffar |
DSN | 6 |
| 2016 | Sneak-Peek: High speed covert channels in data center networksabstractWith the advent of big data, modern businesses face an increasing need to store and process large volumes of sensitive customer information on the cloud. In these environments, resources are shared across a multitude of mutually untrusting tenants increasing propensity for data leakage. This problem stands to grow further in severity with increasing use of clouds in all aspects of our daily lives and the recent spate of high-profile data exfiltration attacks are evidence. To highlight this serious issue, we present a novel and highspeed network-based covert channel that is robust and circumvents a broad set of security mechanisms currently deployed by cloud vendors. We successfully test our channel on numerous network environments, including commercial clouds such as EC2 and Azure. Using an information theoretic model of the channel, we derive an upper bound on the maximum information rate and propose an optimal coding scheme. Our adaptive decoding algorithm caters to the cross traffic in the channel and maintains high bit rates and extremely low error rates. Finally, we discuss several effective avenues for mitigation of the aforementioned channel and provide insights into how data exfiltration can be prevented in such shared environments. Rashid Tahir, Mohammad Taha Khan, Xun Gong 0001, AmirEmad Ghassami, Hasanat Kazmi, Matthew Caesar 0001, Fareed Zaffar, Negar Kiyavash |
INFOCOM | 8 |
| 2016 | To route or to secure: Tradeoffs in ICNs over MANETsabstractInformation-Centric Networks (ICNs) operating over Mobile Ad hoc Networks (MANETs) are challenged by the node churn, evolving topologies, and limited resources of the underlying network. The complex interplay of publishers, subscribers, and brokers brings with it a corresponding set of security concerns, where precisely-defined trust boundaries are needed to guarantee the confidentiality and integrity of all data objects in the ecosystem. Building a practical framework that can service users efficiently requires understanding the motivations and actions of the participants. We explore several tradeoffs between efficiency and the security of data objects in such environments, using ICEMAN - a real-wold implementation of an ICN that operates on MANETs. Since our findings are based on an actual system, they have significant implications for building efficient ICNs that have security designed in at the outset (rather than added later when options may be limited). We empirically establish that there is a strong interplay between the need to have more specific information for efficient routing and the need to ensure trust and confidentiality in such a decentralized system. Hasanat Kazmi, Hasnain Lakhani, Ashish Gehani, Rashid Tahir, Fareed Zaffar |
NCA | 5 |
| 2014 | Covert channels in online rogue-like gamesabstractCovert channels allow two parties to exchange secret data in the presence of adversaries without disclosing the fact that there is any secret data in their communications. We propose and implement EEDGE, an improved method for steganography in mazes that builds upon the work done by Lee et al; and has a significantly higher embedding capacity. We apply EEDGE to the setting of online rogue-like games, which have randomly generated mazes as the levels for players; and show that this can be used to successfully create an efficient, error-free, high bit-rate covert channel. Hasnain Lakhani, Fareed Zaffar |
ICC | 2 |
| 2005 | Paranoid: A Global Secure File Access Control SystemabstractThe Paranoid file system is an encrypted, secure, global file system with user managed access control. The system provides efficient peer-to-peer application transparent file sharing. This paper presents the design, implementation and evaluation of the Paranoid file system and its access-control architecture. The system lets users grant safe, selective, UNIX-like, file access to peer groups across administrative boundaries. Files are kept encrypted and access control translates into key management. The system uses a novel transformation key scheme to effect access revocation. The file system works seamlessly with existing applications through the use of interposition agents. The interposition agents provide a layer of indirection making it possible to implement transparent remote file access and data encryption/decryption without any kernel modifications. System performance evaluations show that encryption and remote file-access overheads are small, demonstrating that the Paranoid system is practical Fareed Zaffar, Gershon Kedem, Ashish Gehani |
ACSAC | 1 |