VLDB 2026 Research / reviewers in the wild / expert
Ben Y. Zhao
dblp:z/BenYZhao
· DBLP profile ↗
147ranked-venue papers
3as first author
22since 2021 · last 2025
0009-0003-8909-0494ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 63 · 2 first-author · 3 since 2021Security and privacy · 27 · 14 since 2021Databases, data management, data science and information retrieval · 22Systems, architecture and hardware · 17 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 15 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 13Artificial intelligence and machine learning · 5 · 3 since 2021Software engineering, systems software and programming languages · 5Theory of computation · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial MislabelingabstractToday's text-to-image generative models are trained on millions of images sourced from the Internet, each paired with a detailed caption produced by Vision-Language Models (VLMs). This part of the training pipeline is critical for supplying the models with large volumes of high-quality image-caption pairs during training. However, recent work suggests that VLMs are vulnerable to stealthy adversarial attacks, where adversarial perturbations are added to images to mislead the VLMs into producing incorrect captions. Stanley Wu, Ronik Bhaskar, Anna Yoo Jeong Ha, Shawn Shan, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 6 |
| 2025 | Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI CrawlersabstractThe success of generative AI relies heavily on training on data scraped through extensive crawling of the Internet, a practice that has raised significant copyright, privacy, and ethical concerns. While few measures are designed to resist a resource-rich adversary determined to scrape a site, crawlers can be impacted by a range of existing tools such as robots.txt, NoAI meta tags, and active crawler blocking by reverse proxies. In this work, we seek to understand the ability and efficacy of today's networking tools to protect content creators against AI-related crawling. For targeted populations like human artists, do they have the technical knowledge and agency to utilize crawler blocking tools such as robots.txt, and can such tools be effective? Using large scale measurements and a targeted user study of 203 professional artists, we find strong demand for tools like robots.txt, but significantly constrained by critical hurdles in technical awareness, agency in deploying them, and limited efficacy against unresponsive crawlers. We further test and evaluate network level crawler blockers provided by reverse proxies. Despite relatively limited deployment today, they offer stronger protections against AI crawlers, but still come with their own set of limitations. Enze Liu 0001, Elisa Luo, Shawn Shan, Geoffrey M. Voelker, Ben Y. Zhao, Stefan Savage |
IMC | 5 |
| 2024 | Understanding Implosion in Text-to-Image Generative ModelsabstractRecent works show that text-to-image generative models are surprisingly vulnerable to a variety of poisoning attacks. Empirical results find that these models can be corrupted by altering associations between individual text prompts and associated visual features. Furthermore, a number of concurrent poisoning attacks can induce "model implosion," where the model becomes unable to produce meaningful images for unpoisoned prompts. These intriguing findings highlight the absence of an intuitive framework to understand poisoning attacks on these models. In this work, we establish the first analytical framework on robustness of image generative models to poisoning attacks, by modeling and analyzing the behavior of the cross-attention mechanism in latent diffusion models. We model cross-attention training as an abstract problem of "supervised graph alignment" and formally quantify the impact of training data by the hardness of alignment, measured by an Alignment Difficulty (AD) metric. The higher the AD, the harder the alignment. We prove that AD increases with the number of individual prompts (or concepts) poisoned. As AD grows, the alignment task becomes increasingly difficult, yielding highly distorted outcomes that frequently map meaningful text prompts to undefined or meaningless visual representations. As a result, the generative model implodes and outputs random, incoherent images at large. We validate our analytical framework through extensive experiments, and we confirm and explain the unexpected (and unexplained) effect of model implosion while producing new, unforeseen insights. Our work provides a useful tool for studying poisoning attacks against diffusion models and their defenses. Wenxin Ding, Cathy Yuanchen Li, Shawn Shan, Ben Y. Zhao, Hai-Tao Zheng 0002 |
CCS | 4 |
| 2024 | Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?abstractThe advent of generative AI images has completely disrupted the art world. Distinguishing AI generated images from human art is a challenging problem whose impact is growing over time. A failure to address this problem allows bad actors to defraud individuals paying a premium for human art and companies whose stated policies forbid AI imagery. It is also critical for content owners to establish copyright, and for model trainers interested in curating training data in order to avoid potential model collapse. There are several different approaches to distinguishing human art from AI images, including classifiers trained by supervised learning, research tools targeting diffusion models, and identification by professional artists using their knowledge of artistic techniques. In this paper, we seek to understand how well these approaches can perform against today's modern generative models in both benign and adversarial settings. We curate real human art across 7 styles, generate matching images from 5 generative models, and apply 8 detectors (5 automated detectors and 3 different human groups including 180 crowdworkers, 3800+ professional artists, and 13 expert artists experienced at detecting AI). Both Hive and expert artists do very well, but make mistakes in different ways (Hive is weaker against adversarial perturbations while Expert artists produce higher false positives). We believe these weaknesses will persist, and argue that a combination of human and automated detectors provides the best combination of accuracy and robustness. Anna Yoo Jeong Ha, Josephine Passananti, Ronik Bhaskar, Shawn Shan, Reid Southen, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 7 |
| 2024 | Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative ModelsabstractTrained on billions of images, diffusion-based text-to-image models seem impervious to traditional data poisoning attacks, which typically require poison samples approaching 20% of the training set. In this paper, we show that state-of-the-art text-to-image generative models are in fact highly vulnerable to poisoning attacks. Our work is driven by two key insights. First, while diffusion models are trained on billions of samples, the number of training samples associated with a specific concept or prompt is generally on the order of thousands. This suggests that these models will be vulnerable to prompt-specific poisoning attacks that corrupt a model’s ability to respond to specific targeted prompts. Second, poison samples can be carefully crafted to maximize poison potency to ensure success with very few samples.We introduce Nightshade, a prompt-specific poisoning attack optimized for potency that can completely control the output of a prompt in Stable Diffusion’s newest model (SDXL) with less than 100 poisoned training samples. Nightshade also generates stealthy poison images that look visually identical to their benign counterparts, and produces poison effects that "bleed through" to related concepts. More importantly, a moderate number of Nightshade attacks on independent prompts can destabilize a model and disable its ability to generate images for any and all prompts. Finally, we propose the use of Nightshade and similar tools as a defense for content owners against web scrapers that ignore opt-out/do-not-crawl directives, and discuss potential implications for both model trainers and content owners. Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng 0001, Ben Y. Zhao |
SP | 6 |
| 2024 | Can Virtual Reality Protect Users from Keystroke Inference Attacks?
Zhuolin Yang 0001, Zain Sarwar, Iris Hwang, Ronik Bhaskar, Ben Y. Zhao, Haitao Zheng 0001 |
USENIX Security Symposium | 5 |
| 2024 | Data Isotopes for Data Provenance in DNNsabstractToday, creators of data-hungry deep neural networks (DNNs) scour the Internet for training fodder, leaving users with little control over or knowledge of when their data, and in particular their images, are used to train models. To empower users to counteract unwanted use of their images, we design, implement and evaluate a practical system that enables users to detect if their data was used to train a DNN model for image classification. We show how users can create special images we call isotopes, which introduce ``spurious features'' into DNNs during training. With only query access to a model and no knowledge of the model-training process, nor control of the data labels, a user can apply statistical hypothesis testing to detect if the model learned these spurious features by training on the user's images. Isotopes can be viewed as an application of a particular type of data poisoning. In contrast to backdoors and other poisoning attacks, our purpose is not to cause misclassification but rather to create tell-tale changes in confidence scores output by the model that reveal the presence of isotopes in the training data. Isotopes thus turn DNNs' vulnerability to memorization and spurious correlations into a tool for data provenance. Our results confirm efficacy in multiple image classification settings, detecting and distinguishing between hundreds of isotopes with high accuracy. We further show that our system works on public ML-as-a-service platforms and larger models such as ImageNet, can use physical objects in images instead of digital marks, and remains robust against several adaptive countermeasures. Emily Wenger, Xiuyu Li, Ben Y. Zhao, Vitaly Shmatikov |
Proc. Priv. Enhancing Technol. | 3 |
| 2023 | SoK: Anti-Facial Recognition TechnologyabstractThe rapid adoption of facial recognition (FR) technology by both government and commercial entities in recent years has raised concerns about civil liberties and privacy. In response, a broad suite of so-called "anti-facial recognition" (AFR) tools has been developed to help users avoid unwanted facial recognition. The set of AFR tools proposed in the last few years is wide-ranging and rapidly evolving, necessitating a step back to consider the broader design space of AFR systems and long-term challenges. This paper aims to fill that gap and provides the first comprehensive analysis of the AFR research landscape. Using the operational stages of FR systems as a starting point, we create a systematic framework for analyzing the benefits and tradeoffs of different AFR approaches. We then consider both technical and social challenges facing AFR tools and propose directions for future research in this field. Emily Wenger, Shawn Shan, Haitao Zheng 0001, Ben Y. Zhao |
SP | 4 |
| 2023 | Glaze: Protecting Artists from Style Mimicry by Text-to-Image Models
Shawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng 0001, Rana Hanocka, Ben Y. Zhao |
USENIX Security Symposium | 6 |
| 2023 | Towards a General Video-based Keystroke Inference Attack
Zhuolin Yang 0001, Yuxin Chen 0001, Zain Sarwar, Hadleigh Schwartz, Ben Y. Zhao, Haitao Zheng 0001 |
USENIX Security Symposium | 5 |
| 2023 | "My face, my rules": Enabling Personalized Protection Against Unacceptable Face EditingabstractToday, face editing is widely used to refine/alter photos in both professional and recreational settings. Yet it is also used to modify (and repost) existing online photos for cyberbullying. Our work considers an important open question: 'How can we support the collaborative use of face editing on social platforms while protecting against unacceptable edits and reposts by others?' This is challenging because, as our user study shows, users vary widely in their definition of what edits are (un)acceptable. Any global filter policy deployed by social platforms is unlikely to address the needs of all users, but hinders social interactions enabled by photo editing. Instead, we argue that face edit protection policies should be implemented by social platforms based on individual user preferences. When posting an original photo online, a user can choose to specify the types of face edits (dis)allowed on the photo. Social platforms use these per-photo edit policies to moderate future photo uploads, i.e., edited photos containing modifications that violate the original photo's policy are either blocked or shelved for user approval. Realizing this personalized protection, however, faces two immediate challenges: (1) how to accurately recognize specific modifications, if any, contained in a photo; and (2) how to associate an edited photo with its original photo (and thus the edit policy). We show that these challenges can be addressed by combining highly efficient hashing based image search and scalable semantic image comparison, and build a prototype protector (Alethia) covering nine edit types. Evaluations using IRB-approved user studies and data-driven experiments (on 839K face photos) show that Alethia accurately recognizes edited photos that violate user policies and induces a feeling of protection to study participants. This demonstrates the initial feasibility of personalized face edit protection. We also discuss current limitations and future directions to push the concept forward. Zhujun Xiao, Jenna Cryan, Yuanshun Yao, Yi Hong Gordon Cheo, Yuanchao Shu, Stefan Saroiu, Ben Y. Zhao, Haitao Zheng 0001 |
Proc. Priv. Enhancing Technol. | 7 |
| 2023 | Trimming Mobile Applications for Bandwidth-Challenged Networks in Developing RegionsabstractDespite continuous efforts to build and update mobile network infrastructure, mobile devices in developing regions continue to be constrained by limited bandwidth. Unfortunately, this coincides with a period of unprecedented growth in the sizes of mobile applications. Thus it is becoming prohibitively expensive for users in developing regions to download and update mobile apps critical to their economic and educational development. Unchecked, these trends can further contribute to a large and growing global digital divide. Our goal is to better understand the source of this rapid growth in mobile app code size, whether it is reflective of new functionality, and identify steps that can be taken to make existing mobile apps more friendly to bandwidth constrained mobile networks. We hypothesize that much of this growth in mobile apps is due to poor resource/code management, and do not reflect proportional increases in functionality. Our hypothesis is partially validated by mini-programs, apps with extremely small footprints gaining popularity in Chinese mobile platforms. Here, we use functionally equivalent pairs of mini-programs and Android apps to identify potential sources of “bloat,” i.e., inefficient uses of code or resources that contribute to large package sizes. We analyze a large sample of popular Android apps and quantify instances of code and resource bloat. We develop techniques for automated code and resource trimming, and successfully validate them on a large set of Android apps. We hope our results will lead to continued efforts to streamline mobile apps, making them easier to access and maintain for users in developing regions. Qinge Xie, Qingyuan Gong, Xinlei He 0001, Yang Chen 0001, Xin Wang 0002, Haitao Zheng 0001, Ben Y. Zhao |
IEEE Trans. Mob. Comput. | 7 |
| 2022 | Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN ModelsabstractServer breaches are an unfortunate reality on today's Internet. In the context of deep neural network (DNN) models, they are particularly harmful, because a leaked model gives an attacker "white-box'' access to generate adversarial examples, a threat model that has no practical robust defenses. For practitioners who have invested years and millions into proprietary DNNs, e.g. medical imaging, this seems like an inevitable disaster looming on the horizon. Shawn Shan, Wenxin Ding, Emily Wenger, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 5 |
| 2022 | Understanding Robust Learning through the Lens of Representation SimilaritiesabstractRepresentation learning, \textit{i.e.} the generation of representations useful for downstream applications, is a task of fundamental importance that underlies much of the success of deep neural networks (DNNs). Recently, \emph{robustness to adversarial examples} has emerged as a desirable property for DNNs, spurring the development of robust training methods that account for adversarialexamples. In this paper, we aim to understand how the properties of representations learned by robust training differ from those obtained from standard, non-robust training. This is critical to diagnosing numerous salient pitfalls in robust networks, such as, degradation of performance on benign inputs, poor generalization of robustness, and increase in over-fitting. We utilize a powerful set of tools known as representation similarity metrics, across 3 vision datasets, to obtain layer-wise comparisons between robust and non-robust DNNs with different architectures, training procedures and adversarial constraints. Our experiments highlight hitherto unseen properties of robust representations that we posit underlie the behavioral differences of robust networks. We discover a lack of specialization in robust networks' representations along with a disappearance of `block structure'. We also find overfitting during robust training largely impacts deeper layers. These, along with other findings, suggest ways forward for the design and training of better robust networks. Christian Cianfarani, Arjun Nitin Bhagoji, Vikash Sehwag, Ben Y. Zhao, Haitao Zheng 0001, Prateek Mittal |
NeurIPS | 4 |
| 2022 | Finding Naturally Occurring Physical Backdoors in Image DatasetsabstractExtensive literature on backdoor poison attacks has studied attacks and defenses for backdoors using “digital trigger patterns.” In contrast, “physical backdoors” use physical objects as triggers, have only recently been identified, and are qualitatively different enough to resist most defenses targeting digital trigger backdoors. Research on physical backdoors is limited by access to large datasets containing real images of physical objects co-located with misclassification targets. Building these datasets is time- and labor-intensive.This work seeks to address the challenge of accessibility for research on physical backdoor attacks. We hypothesize that there may be naturally occurring physically co-located objects already present in popular datasets such as ImageNet. Once identified, a careful relabeling of these data can transform them into training samples for physical backdoor attacks. We propose a method to scalably identify these subsets of potential triggers in existing datasets, along with the specific classes they can poison. We call these naturally occurring trigger-class subsets natural backdoor datasets. Our techniques successfully identify natural backdoors in widely-available datasets, and produce models behaviorally equivalent to those trained on manually curated datasets. We release our code to allow the research community to create their own datasets for research on physical backdoor attacks. Emily Wenger, Roma Bhattacharjee, Arjun Nitin Bhagoji, Josephine Passananti, Emilio Andere, Haitao Zheng 0001, Ben Y. Zhao |
NeurIPS | 7 |
| 2022 | Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box Attacks
Huiying Li 0001, Shawn Shan, Emily Wenger, Jiayun Zhang, Haitao Zheng 0001, Ben Y. Zhao |
USENIX Security Symposium | 6 |
| 2022 | Poison Forensics: Traceback of Data Poisoning Attacks in Neural Networks
Shawn Shan, Arjun Nitin Bhagoji, Haitao Zheng 0001, Ben Y. Zhao |
USENIX Security Symposium | 4 |
| 2021 | "Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real WorldabstractAdvances in deep learning have introduced a new wave of voice synthesis tools, capable of producing audio that sounds as if spoken by a target speaker. If successful, such tools in the wrong hands will enable a range of powerful attacks against both humans and software systems (aka machines). This paper documents efforts and findings from a comprehensive experimental study on the impact of deep-learning based speech synthesis attacks on both human listeners and machines such as speaker recognition and voice-signin systems. We find that both humans and machines can be reliably fooled by synthetic speech, and that existing defenses against synthesized speech fall short. These findings highlight the need to raise awareness and develop new protections against synthetic speech for both humans and machines. Emily Wenger, Max Bronckers, Christian Cianfarani, Jenna Cryan, Angela Sha, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 7 |
| 2021 | User Authentication via Electrical Muscle StimulationabstractWe propose a novel modality for active biometric authentication: electrical muscle stimulation (EMS). To explore this, we engineered an interactive system, which we call ElectricAuth, that stimulates the user’s forearm muscles with a sequence of electrical impulses (i.e., EMS challenge) and measures the user’s involuntary finger movements (i.e., response to the challenge). ElectricAuth leverages EMS’s intersubject variability, where the same electrical stimulation results in different movements in different users because everybody’s physiology is unique (e.g., differences in bone and muscular structure, skin resistance and composition, etc.). As such, ElectricAuth allows users to login without memorizing passwords or PINs. Yuxin Chen 0001, Zhuolin Yang 0001, Ruben Abbou, Pedro Lopes 0001, Ben Y. Zhao, Haitao Zheng 0001 |
CHI | 5 |
| 2021 | Backdoor Attacks Against Deep Learning Systems in the Physical WorldabstractBackdoor attacks embed hidden malicious behaviors into deep learning models, which only activate and cause misclassifications on model inputs containing a specific "trigger." Existing works on backdoor attacks and defenses, however, mostly focus on digital attacks that apply digitally generated patterns as triggers. A critical question remains unanswered: "can backdoor attacks succeed using physical objects as triggers, thus making them a credible threat against deep learning systems in the real world?"We conduct a detailed empirical study to explore this question for facial recognition, a critical deep learning task. Using 7 physical objects as triggers, we collect a custom dataset of 3205 images of 10 volunteers and use it to study the feasibility of "physical" backdoor attacks under a variety of real-world conditions. Our study reveals two key findings. First, physical backdoor attacks can be highly successful if they are carefully configured to overcome the constraints imposed by physical objects. In particular, the placement of successful triggers is largely constrained by the target model’s dependence on key facial features. Second, four of today’s state-of-the-art defenses against (digital) backdoors are ineffective against physical backdoors, because the use of physical objects breaks core assumptions used to construct these defenses.Our study confirms that (physical) backdoor attacks are not a hypothetical phenomenon but rather pose a serious real-world threat to critical classification tasks. We need new and more robust defenses against backdoors in the physical world. Emily Wenger, Josephine Passananti, Arjun Nitin Bhagoji, Yuanshun Yao, Haitao Zheng 0001, Ben Y. Zhao |
CVPR | 6 |
| 2021 | Towards Performance Clarity of Edge Video Analytics
Zhujun Xiao, Zhengxu Xia, Haitao Zheng 0001, Ben Y. Zhao, Junchen Jiang |
SEC | 4 |
| 2021 | On Migratory Behavior in Video ConsumptionabstractToday’s video streaming market is crowded with various content providers (CPs). For individual CPs, understanding user behavior, in particular how users migrate among different CPs, is crucial for improving users’ on-site experience and the CP’s chance of success. In this article, we take a data-driven approach to analyze and model user migration behavior in video streaming, i.e., users switching content provider during active sessions. Based on a large ISP dataset over two months (6 major content providers, 3.8 million users, and 315 million video requests), we study common migration patterns and reasons of migration. We find that migratory behavior is prevalent: 66% of users switch CPs with an average switching frequency of 13%. In addition, migration behaviors are highly diverse: regardless large or small CPs, they all have dedicated groups of users who like to switch to them for certain types of videos. Regarding reasons of migration, we find CP service quality rarely causes migration, while a few popular videos play a bigger role. Nearly 60% of cross-site migrations are landed to 0.14% top videos. Finally, we validate our findings by building an accurate regression model to predict user migration frequency as well as user survey, and discuss the implications of our results to CPs. Huan Yan 0003, Haohao Fu, Yong Li 0008, Tzu-Heng Lin, Gang Wang 0011, Haitao Zheng 0001, Depeng Jin, Ben Y. Zhao |
IEEE Trans. Netw. Serv. Manag. | 8 |
| 2020 | Gotta Catch'Em All: Using Honeypots to Catch Adversarial Attacks on Neural NetworksabstractDeep neural networks (DNN) are known to be vulnerable to adversarial attacks. Numerous efforts either try to patch weaknesses in trained models, or try to make it difficult or costly to compute adversarial examples that exploit them. In our work, we explore a new "honeypot" approach to protect DNN models. We intentionally inject trapdoors, honeypot weaknesses in the classification manifold that attract attackers searching for adversarial examples. Attackers' optimization algorithms gravitate towards trapdoors, leading them to produce attacks similar to trapdoors in the feature space. Our defense then identifies attacks by comparing neuron activation signatures of inputs to those of trapdoors. Shawn Shan, Emily Wenger, Bolun Wang, Bo Li 0026, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 6 |
| 2020 | Wearable Microphone JammingabstractWe engineered a wearable microphone jammer that is capable of disabling microphones in its user's surroundings, including hidden microphones. Our device is based on a recent exploit that leverages the fact that when exposed to ultrasonic noise, commodity microphones will leak the noise into the audible range. Yuxin Chen 0001, Huiying Li 0001, Shan-Yuan Teng, Steven Nagels, Zhijing Li 0001, Pedro Lopes 0001, Ben Y. Zhao, Haitao Zheng 0001 |
CHI | 7 |
| 2020 | Detecting Gender Stereotypes: Lexicon vs. Supervised Learning MethodsabstractBiases in language influence how we interact with each other and society at large. Language affirming gender stereotypes is often observed in various contexts today, from recommendation letters and Wikipedia entries to fiction novels and movie dialogue. Yet to date, there is little agreement on the methodology to quantify gender stereotypes in natural language (specifically the English language). Common methodology (including those adopted by companies tasked with detecting gender bias) rely on a lexicon approach largely based on the original BSRI study from 1974. Jenna Cryan, Shiliang Tang, Xinyi Zhang 0003, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao |
CHI | 6 |
| 2020 | Deep Graph Convolutional Networks for Incident-Driven Traffic Speed PredictionabstractAccurate traffic speed prediction is an important and challenging topic for transportation planning. Previous studies on traffic speed prediction predominately used spatio-temporal and context features for prediction. However, they have not made good use of the impact of traffic incidents. In this work, we aim to make use of the information of incidents to achieve a better prediction of traffic speed. Our incident-driven prediction framework consists of three processes. First, we propose a critical incident discovery method to discover traffic incidents with high impact on traffic speed. Second, we design a binary classifier, which uses deep learning methods to extract the latent incident impact features. Combining above methods, we propose a Deep Incident-Aware Graph Convolutional Network (DIGC-Net) to effectively incorporate traffic incident, spatio-temporal, periodic and context features for traffic speed prediction. We conduct experiments using two real-world traffic datasets of San Francisco and New York City. The results demonstrate the superior performance of our model compared with the competing benchmarks. Qinge Xie, Tiancheng Guo, Yang Chen 0001, Yu Xiao 0001, Xin Wang 0002, Ben Y. Zhao |
CIKM | 6 |
| 2020 | Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors
Yanzi Zhu, Zhujun Xiao, Yuxin Chen 0001, Zhijing Li 0001, Max Liu, Ben Y. Zhao, Haitao Zheng 0001 |
NDSS | 6 |
| 2020 | Fawkes: Protecting Privacy against Unauthorized Deep Learning Models
Shawn Shan, Emily Wenger, Jiayun Zhang, Huiying Li 0001, Haitao Zheng 0001, Ben Y. Zhao |
USENIX Security Symposium | 6 |
| 2019 | Latent Backdoor Attacks on Deep Neural NetworksabstractRecent work proposed the concept of backdoor attacks on deep neural networks (DNNs), where misclassification rules are hidden inside normal models, only to be triggered by very specific inputs. However, these "traditional" backdoors assume a context where users train their own models from scratch, which rarely occurs in practice. Instead, users typically customize "Teacher" models already pretrained by providers like Google, through a process called transfer learning. This customization process introduces significant changes to models and disrupts hidden backdoors, greatly reducing the actual impact of backdoors in practice. In this paper, we describe latent backdoors, a more powerful and stealthy variant of backdoor attacks that functions under transfer learning. Latent backdoors are incomplete backdoors embedded into a "Teacher" model, and automatically inherited by multiple "Student" models through transfer learning. If any Student models include the label targeted by the backdoor, then its customization process completes the backdoor and makes it active. We show that latent backdoors can be quite effective in a variety of application contexts, and validate its practicality through real-world attacks against traffic sign recognition, iris identification of volunteers, and facial recognition of public figures (politicians). Finally, we evaluate 4 potential defenses, and find that only one is effective in disrupting latent backdoors, but might incur a cost in classification accuracy as tradeoff. Yuanshun Yao, Huiying Li 0001, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 4 |
| 2019 | Scaling Deep Learning Models for Spectrum Anomaly DetectionabstractSpectrum management in cellular networks is a challenging task that will only increase in difficulty as complexity grows in hardware, configurations, and new access technology (e.g. LTE for IoT devices). Wireless providers need robust and flexible tools to monitor and detect faults and misbehavior in physical spectrum usage, and to deploy them at scale. In this paper, we explore the design of such a system by building deep neural network (DNN) models1 to capture spectrum usage patterns and use them as baselines to detect spectrum usage anomalies resulting from faults and misuse. Using detailed LTE spectrum measurements, we show that the key challenge facing this design is model scalability, i.e. how to train and deploy DNN models at a large number of static and mobile observers located throughout the network. We address this challenge by building context-agnostic models for spectrum usage and applying transfer learning to minimize training time and dataset constraints. The end result is a practical DNN model that can be easily deployed on both mobile and static observers, enabling timely detection of spectrum anomalies across LTE networks. Zhijing Li 0001, Zhujun Xiao, Bolun Wang, Ben Y. Zhao, Haitao Zheng 0001 |
MobiHoc | 4 |
| 2019 | Safely and automatically updating in-network ACL configurations with intent languageabstractIn-network Access Control List (ACL) is an important technique in ensuring network-wide connectivity and security. As cloud-scale WANs today constantly evolve in size and complexity, in-network ACL rules are becoming increasingly more complex. This presents a great challenge to the updating process of ACL configurations: network operators are frequently required to update "tangled" ACL rules across thousands of devices to meet diverse business requirements, and even a single ACL misconfiguration may lead to network disruptions. Such increasing challenges call for an automated system to improve the efficiency and correctness of ACL updates. This paper presents Jinjing, a system that aids Alibaba's network operators in automatically and correctly updating ACL configurations in Alibaba's global WAN. Jinjing allows the operators to express in a declarative language, named LAI, their update intent (e.g., ACL migration and traffic control). Then, Jinjing automatically synthesizes ACL update plans that satisfy their intent. At the heart of Jinjing, we develop a set of novel verification and synthesis techniques to rigorously guarantee the correctness of update plans. In Alibaba, our operators have used Jinjing to efficiently update their ACLs and have thus prevented significant service downtime. Bingchuan Tian, Xinyi Zhang 0003, Ennan Zhai, Hongqiang Harry Liu, Qiaobo Ye, Chunsheng Wang, Zhiming Ji, Yihong Sang, Ming Zhang 0005, Chen Tian 0001, Haitao Zheng 0001, Ben Y. Zhao |
SIGCOMM | 14 |
| 2019 | Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksabstractLack of transparency in deep neural networks (DNNs) make them susceptible to backdoor attacks, where hidden associations or triggers override normal classification to produce unexpected results. For example, a model with a backdoor always identifies a face as Bill Gates if a specific symbol is present in the input. Backdoors can stay hidden indefinitely until activated by an input, and present a serious security risk to many security or safety related applications, e.g. biometric authentication systems or self-driving cars. We present the first robust and generalizable detection and mitigation system for DNN backdoor attacks. Our techniques identify backdoors and reconstruct possible triggers. We identify multiple mitigation techniques via input filters, neuron pruning and unlearning. We demonstrate their efficacy via extensive experiments on a variety of DNNs, against two types of backdoor injection methods identified by prior work. Our techniques also prove robust against a number of variants of the backdoor attack. Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 0001, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao |
IEEE Symposium on Security and Privacy | 7 |
| 2018 | DeepCredit: Exploiting User Cickstream for Loan Risk Prediction in P2P Lending
Zhi Yang 0001, Ben Y. Zhao, Yafei Dai |
ICWSM | 4 |
| 2018 | Predictive Analysis in Network Function Virtualization
Zhijing Li 0001, Zihui Ge, Ajay Mahimkar, Jia Wang 0001, Ben Y. Zhao, Haitao Zheng 0001, Joanne Emmons, Laura Ogden |
Internet Measurement Conference | 5 |
| 2018 | With Great Training Comes Great Vulnerability: Practical Attacks against Transfer Learning
Bolun Wang, Yuanshun Yao, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao |
USENIX Security Symposium | 5 |
| 2018 | Understanding Motivations behind Inaccurate Check-insabstractCheck-in data from social networks provide researchers a unique opportunity to model human dynamics at scale. However, it is unclear how indicative these check-in traces are of real human mobility. Prior work showed that significant amounts of Foursquare check-ins did not match with the physical mobility patterns of users, and suggested that misrepresented check-ins were incentivized by external rewards provided by the system. In this paper, our goal is to understand the root cause of inaccurate check-in data, by studying the validity of check-in traces in social media platforms without external rewards for check-ins. We conduct a data-driven analysis using an empirical check-in data trace of more than 276,000 users from WeChat Moments, with matching traces of their physical mobility. We develop a set of hypotheses on the underlying motivations behind people's inaccurate check-ins, and validate them using a detailed survey study. Our analysis reveals that there are surprisingly high amount of inaccurate check-ins even in the absence of rewards: 43% of total check-ins are inaccurate and 61% of survey participants report they have misrepresented their check-ins. We also find that inaccurate check-ins are often a result of user interface design as well as for convenience, self-advertisement and self-presentation. Fengli Xu, Guozhen Zhang 0001, Zhilong Chen, Jiaxin Huang 0001, Yong Li 0008, Diyi Yang, Ben Y. Zhao |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2018 | Ghost Riders: Sybil Attacks on Crowdsourced Mobile Mapping Services
Gang Wang 0011, Bolun Wang, Tianyi Wang 0001, Ana Nika, Haitao Zheng 0001, Ben Y. Zhao |
IEEE/ACM Trans. Netw. | 6 |
| 2017 | Automated Crowdturfing Attacks and Defenses in Online Review SystemsabstractMalicious crowdsourcing forums are gaining traction as sources of spreading misinformation online, but are limited by the costs of hiring and managing human workers. In this paper, we identify a new class of attacks that leverage deep learning language models (Recurrent Neural Networks or RNNs) to automate the generation of fake online reviews for products and services. Not only are these attacks cheap and therefore more scalable, but they can control rate of content output to eliminate the signature burstiness that makes crowdsourced campaigns easy to detect. Yuanshun Yao, Bimal Viswanath, Jenna Cryan, Haitao Zheng 0001, Ben Y. Zhao |
CCS | 5 |
| 2017 | On Migratory Behavior in Video ConsumptionabstractToday's video streaming market is crowded with various content providers (CPs). For individual CPs, understanding user behavior, in particular how users migrate among different CPs, is crucial for improving users' on-site experience and the CP's chance of success. In this paper, we take a data-driven approach to analyze and model user migration behavior in video streaming, i.e., users switching content provider during active sessions. Based on a large ISP dataset over two months (6 major content providers, 3.8 million users, and 315 million video requests), we study common migration patterns and reasons of migration. We find that migratory behavior is prevalent: 66% of users switch CPs with an average switching frequency of 13%. In addition, migration behaviors are highly diverse: regardless large or small CPs, they all have dedicated groups of users who like to switch to them for certain types of videos. Regarding reasons of migration, we find CP service quality rarely causes migration, while a few popular videos play a bigger role. Nearly 60% of cross-site migrations are landed to 0.14% top videos. Finally, we validate our findings by building an accurate regression model to predict user migration frequency, and discuss the implications of our results to CPs. Huan Yan 0003, Tzu-Heng Lin, Gang Wang 0011, Yong Li 0008, Haitao Zheng 0001, Depeng Jin, Ben Y. Zhao |
CIKM | 7 |
| 2017 | Echo Chambers in Investment Discussion Boards
Shiliang Tang, Qingyun Liu 0003, Megan McQueen, Scott Counts, Apurv Jain, Haitao Zheng 0001, Ben Y. Zhao |
ICWSM | 7 |
| 2017 | A First Look at User Switching Behaviors Over Multiple Video Content Providers
Huan Yan 0003, Tzu-Heng Lin, Gang Wang 0011, Yong Li 0008, Haitao Zheng 0001, Depeng Jin, Ben Y. Zhao |
ICWSM | 7 |
| 2017 | Cold Hard E-Cash: Friends and Vendors in the Venmo Digital Payments System
Xinyi Zhang 0003, Shiliang Tang, Yun Zhao 0001, Gang Wang 0011, Haitao Zheng 0001, Ben Y. Zhao |
ICWSM | 6 |
| 2017 | Complexity vs. performance: empirical analysis of machine learning as a serviceabstractMachine learning classifiers are basic research tools used in numerous types of network analysis and modeling. To reduce the need for domain expertise and costs of running local ML classifiers, network researchers can instead rely on centralized Machine Learning as a Service (MLaaS) platforms. Yuanshun Yao, Zhujun Xiao, Bolun Wang, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 6 |
| 2017 | Object Recognition and Navigation using a Single Networking DeviceabstractTomorrow's autonomous mobile devices need accurate, robust and real-time sensing of their operating environment. Today's solutions fall short. Vision or acoustic-based techniques are vulnerable against challenging lighting conditions or background noise, while more robust laser or RF solutions require either bulky expensive hardware or tight coordination between multiple devices. This paper describes the design, implementation and evaluation of Ulysses, a practical environmental imaging system using colocated 60GHz radios on a single mobile device. Unlike alternatives that require specialized hardware, Ulysses reuses low-cost commodity networking chipsets available today. Ulysses' new imaging approach leverages RF beamforming, operates on specular (direct) reflection, and integrates the device's movement trajectory with sensing. Ulysses also includes a navigation component that uses the same 60GHz radios to compute "safety regions" where devices can move freely without collision, and to compute optimal paths for imaging within safety regions. Using our implementation of a small robotic car prototype, our experimental results show that Ulysses images objects meters away with cm-level precision, and provides accurate estimates of objects' surface materials. Yanzi Zhu, Yuanshun Yao, Ben Y. Zhao, Haitao Zheng 0001 |
MobiSys | 3 |
| 2017 | Identifying Value in Crowdsourced Wireless Signal MeasurementsabstractWhile crowdsourcing is an attractive approach to collect large-scale wireless measurements, understanding the quality and variance of the resulting data is difficult. Our work analyzes the quality of crowdsourced cellular signal measurements in the context of basestation localization, using large international public datasets (419M signal measurements and 1M cells) and corresponding ground truth values. Performing localization using raw received signal strength (RSS) data produces poor results and very high variance. Applying supervised learning improves results moderately, but variance remains high. Instead, we propose feature clustering, a novel application of unsupervised learning to detect hidden correlation between measurement instances, their features, and localization accuracy. Our results identify RSS standard deviation and RSS-weighted dispersion mean as key features that correlate with highly predictive measurement samples for both sparse and dense measurements respectively. Finally, we show how optimizing crowdsourcing measurements for these two features dramatically improves localization accuracy and reduces variance. Zhijing Li 0001, Ana Nika, Xinyi Zhang 0003, Yanzi Zhu, Yuanshun Yao, Ben Y. Zhao, Haitao Zheng 0001 |
WWW | 6 |
| 2017 | Gender Bias in the Job Market: A Longitudinal AnalysisabstractFor millions of workers, online job listings provide the first point of contact to potential employers. As a result, job listings and their word choices can significantly affect the makeup of the responding applicant pool. Here, we study the effects of potentially gender-biased terminology in job listings, and their impact on job applicants, using a large historical corpus of 17 million listings on LinkedIn spanning 10 years. We develop algorithms to detect and quantify gender bias, validate them using external tools, and use them to quantify job listing bias over time. We then perform a user survey over two user populations (N 1=469 , N 2=273 ) to validate our findings and to quantify the end-to-end impact of such bias on applicant decisions. Our findings show gender-bias has decreased significantly over the last 10 years. More surprisingly, we find that impact of gender bias in listings is dwarfed by our respondents' inherent bias towards specific job types. Shiliang Tang, Xinyi Zhang 0003, Jenna Cryan, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2017 | Value and Misinformation in Collaborative Investing PlatformsabstractIt is often difficult to separate the highly capable “experts” from the average worker in crowdsourced systems. This is especially true for challenge application domains that require extensive domain knowledge. The problem of stock analysis is one such domain, where even the highly paid, well-educated domain experts are prone to make mistakes. As an extremely challenging problem space, the “wisdom of the crowds” property that many crowdsourced applications rely on may not hold. In this article, we study the problem of evaluating and identifying experts in the context of SeekingAlpha and StockTwits, two crowdsourced investment services that have recently begun to encroach on a space dominated for decades by large investment banks. We seek to understand the quality and impact of content on collaborative investment platforms, by empirically analyzing complete datasets of SeekingAlpha articles (9 years) and StockTwits messages (4 years). We develop sentiment analysis tools and correlate contributed content to the historical performance of relevant stocks. While SeekingAlpha articles and StockTwits messages provide minimal correlation to stock performance in aggregate, a subset of experts contribute more valuable (predictive) content. We show that these authors can be easily identified by user interactions, and investments based on their analysis significantly outperform broader markets. This effectively shows that even in challenging application domains, there is a secondary or indirect wisdom of the crowds. Finally, we conduct a user survey that sheds light on users’ views of SeekingAlpha content and stock manipulation. We also devote efforts to identify potential manipulation of stocks by detecting authors controlling multiple identities. Tianyi Wang 0001, Gang Wang 0011, Bolun Wang, Divya Sambasivan, Zengbin Zhang, Xing Li 0001, Haitao Zheng 0001, Ben Y. Zhao |
ACM Trans. Web | 8 |
| 2017 | Clickstream User Behavior ModelsabstractThe next generation of Internet services is driven by users and user-generated content. The complex nature of user behavior makes it highly challenging to manage and secure online services. On one hand, service providers cannot effectively prevent attackers from creating large numbers of fake identities to disseminate unwanted content (e.g., spam). On the other hand, abusive behavior from real users also poses significant threats (e.g., cyberbullying). In this article, we propose clickstream models to characterize user behavior in large online services. By analyzing clickstream traces (i.e., sequences of click events from users), we seek to achieve two goals: (1) detection: to capture distinct user groups for the detection of malicious accounts, and (2) understanding: to extract semantic information from user groups to understand the captured behavior. To achieve these goals, we build two related systems. The first one is a semisupervised system to detect malicious user accounts (Sybils). The core idea is to build a clickstream similarity graph where each node is a user and an edge captures the similarity of two users’ clickstreams. Based on this graph, we propose a coloring scheme to identify groups of malicious accounts without relying on a large labeled dataset. We validate the system using ground-truth clickstream traces of 16,000 real and Sybil users from Renren, a large Chinese social network. The second system is an unsupervised system that aims to capture and understand the fine-grained user behavior. Instead of binary classification (malicious or benign), this model identifies the natural groups of user behavior and automatically extracts features to interpret their semantic meanings. Applying this system to Renren and another online social network, Whisper (100K users), we help service providers identify unexpected user behaviors and even predict users’ future actions. Both systems received positive feedback from our industrial collaborators including Renren, LinkedIn, and Whisper after testing on their internal clickstream data. Gang Wang 0011, Xinyi Zhang 0003, Shiliang Tang, Christo Wilson, Haitao Zheng 0001, Ben Y. Zhao |
ACM Trans. Web | 6 |
| 2016 | Unsupervised Clickstream Clustering for User Behavior AnalysisabstractOnline services are increasingly dependent on user participation. Whether it's online social networks or crowdsourcing services, understanding user behavior is important yet challenging. In this paper, we build an unsupervised system to capture dominating user behaviors from clickstream data (traces of users' click events), and visualize the detected behaviors in an intuitive manner. Our system identifies "clusters" of similar users by partitioning a similarity graph (nodes are users; edges are weighted by clickstream similarity). The partitioning process leverages iterative feature pruning to capture the natural hierarchy within user clusters and produce intuitive features for visualizing and understanding captured user behaviors. For evaluation, we present case studies on two large-scale clickstream traces (142 million events) from real social networks. Our system effectively identifies previously unknown behaviors, e.g., dormant users, hostile chatters. Also, our user study shows people can easily interpret identified behaviors using our visualization tool. Gang Wang 0011, Xinyi Zhang 0003, Shiliang Tang, Haitao Zheng 0001, Ben Y. Zhao |
CHI | 5 |
| 2016 | Trimming the Smartphone Network StackabstractNetwork transmissions are the cornerstone of most mobile apps today, and a main contributor to energy consumption. We use a componentized energy model to quantify energy use by device, and observe significant energy consumption by the CPU in network operations. We assert that optimizing network operations in the CPU can produce significant energy savings, and explore the impact of two potential approaches: one-copy data moves and offloading the network stack to the basestation. Yanzi Zhu, Yibo Zhu 0001, Ana Nika, Ben Y. Zhao, Haitao Zheng 0001 |
HotNets | 4 |
| 2016 | "Will Check-in for Badges": Understanding Bias and Misbehavior on Location-Based Social Networks
Gang Wang 0011, Sarita Yardi Schoenebeck, Haitao Zheng 0001, Ben Y. Zhao |
ICWSM | 4 |
| 2016 | Network Growth and Link Prediction Through an Empirical Lens
Qingyun Liu 0003, Shiliang Tang, Xinyi Zhang 0003, Xiaohan Zhao, Ben Y. Zhao, Haitao Zheng 0001 |
Internet Measurement Conference | 5 |
| 2016 | Anatomy of a Personalized Livestreaming System
Bolun Wang, Xinyi Zhang 0003, Gang Wang 0011, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 5 |
| 2016 | On the performance of cloud storage applications with global measurementabstractIn recent years, Dropbox, Google, and Microsoft have been competing in the market of consumer cloud storage (CCS) services. While once the key comparative metric, storage capacity per user has outgrown the needs of most users. Today, third-party applications based on CCS's RESTful Web APIs are becoming a primary way for users to utilize their expanded storage resources. Unfortunately, there is very little visibility into the performance of these Web APIs, even though they are primary determinants of the end user experience on these storage applications. In this paper, we report results from a comprehensive measurement study of the Web APIs of five popular CCS providers. Our results reveal significant differences and limitations in API performance, which result in performance bottlenecks visible to the user through the storage application. We analyze the underlying system designs of the five providers' Web APIs, and present the performance implications of their different design choices. Our research provides practical guidance for service providers to optimize their API performance, for developers to improve the experience of third-party applications, and for users to pick appropriate services that best match their requirements. Guangyuan Wu, Fangming Liu, Haowen Tang, Keke Huang, Qixia Zhang, Zhenhua Li 0001, Ben Y. Zhao, Hai Jin 0001 |
IWQoS | 7 |
| 2016 | Defending against Sybil Devices in Crowdsourced Mapping ServicesabstractReal-time crowdsourced maps such as Waze provide timely updates on traffic, congestion, accidents and points of interest. In this paper, we demonstrate how lack of strong location authentication allows creation of software-based Sybil devices that expose crowdsourced map systems to a variety of security and privacy attacks. Our experiments show that a single Sybil device with limited resources can cause havoc on Waze, reporting false congestion and accidents and automatically rerouting user traffic. More importantly, we describe techniques to generate Sybil devices at scale, creating armies of virtual vehicles capable of remotely tracking precise movements for large user populations while avoiding detection. We propose a new approach to defend against Sybil devices based on co-location edges, authenticated records that attest to the one-time physical co-location of a pair of devices. Over time, co-location edges combine to form large proximity graphs that attest to physical interactions between devices, allowing scalable detection of virtual vehicles. We demonstrate the efficacy of this approach using large-scale simulations, and discuss how they can be used to dramatically reduce the impact of attacks against crowdsourced mapping services. Gang Wang 0011, Bolun Wang, Tianyi Wang 0001, Ana Nika, Haitao Zheng 0001, Ben Y. Zhao |
MobiSys | 6 |
| 2016 | Exploring Cross-Application Cellular Traffic Optimization with Baidu TrafficGuard
Zhenhua Li 0001, Weiwei Wang 0002, Tianyin Xu, Xiang-Yang Li 0001, Yunhao Liu 0001, Christo Wilson, Ben Y. Zhao |
NSDI | 8 |
| 2016 | Empirical Validation of Commodity Spectrum MonitoringabstractWe describe our efforts to empirically validate a distributed spectrum monitoring system built on commodity smartphones and embedded low-cost spectrum sensors. This system enables real-time spectrum sensing, identifies and locates active transmitters, and generates alarm events when detecting anomalous transmitters. To evaluate the feasibility of such a platform, we perform detailed experiments using a prototype hardware platform using smartphones and RTL dongles. We identify multiple sources of error in the sensing results and the end-user overhead (i.e. smartphone energy draw). We propose and implement a variety of techniques to identify and overcome errors and uncertainty in the data, and to reduce energy consumption. Our work demonstrates the basic viability of user-driven spectrum monitoring on commodity devices. Ana Nika, Zhijing Li 0001, Yanzi Zhu, Yibo Zhu 0001, Ben Y. Zhao, Haitao Zheng 0001 |
SenSys | 5 |
| 2016 | The power of comments: fostering social interactions in microblog networks
Tianyi Wang 0001, Yang Chen 0001, Bolun Wang, Gang Wang 0011, Xing Li 0001, Haitao Zheng 0001, Ben Y. Zhao |
Frontiers Comput. Sci. | 8 |
| 2016 | Understanding and Predicting Data Hotspots in Cellular Networks
Ana Nika, Asad Ismail, Ben Y. Zhao, Sabrina Gaito, Gian Paolo Rossi 0001, Haitao Zheng 0001 |
Mob. Networks Appl. | 3 |
| 2015 | Crowds on Wall Street: Extracting Value from Collaborative Investing PlatformsabstractIn crowdsourced systems, it is often difficult to separate the highly capable "experts" from the average worker. In this paper, we study the problem of evaluating and identifying experts in the context of SeekingAlpha and StockTwits, two crowdsourced investment services that are encroaching on a space dominated for decades by large investment banks. We seek to understand the quality and impact of content on collaborative investment platforms, by empirically analyzing complete datasets of SeekingAlpha articles (9 years) and StockTwits messages (4 years). We develop sentiment analysis tools and correlate contributed content to the historical performance of relevant stocks. While SeekingAlpha articles and StockTwits messages provide minimal correlation to stock performance in aggregate, a subset of experts contribute more valuable (predictive) content. We show that these authors can be easily identified by user interactions, and investments using their analysis significantly outperform broader markets. Finally, we conduct a user survey that sheds light on users views of SeekingAlpha content and stock manipulation. Gang Wang 0011, Tianyi Wang 0001, Bolun Wang, Divya Sambasivan, Zengbin Zhang, Haitao Zheng 0001, Ben Y. Zhao |
CSCW | 7 |
| 2015 | Uncovering User Interaction Dynamics in Online Social Networks
Zhi Yang 0001, Jilong Xue, Christo Wilson, Ben Y. Zhao, Yafei Dai |
ICWSM | 4 |
| 2015 | Reusing 60GHz Radios for Mobile Radar ImagingabstractThe future of mobile computing involves autonomous drones, robots and vehicles. To accurately sense their surroundings in a variety of scenarios, these mobile computers require a robust environmental mapping system. One attractive approach is to reuse millimeterwave communication hardware in these devices, e.g. 60GHz networking chipset, and capture signals reflected by the target surface. The devices can also move while collecting reflection signals, creating a large synthetic aperture radar (SAR) for high-precision RF imaging. Our experimental measurements, however, show that this approach provides poor precision in practice, as imaging results are highly sensitive to device positioning errors that translate into phase errors. We address this challenge by proposing a new 60GHz imaging algorithm, {\em RSS Series Analysis}, which images an object using only RSS measurements recorded along the device's trajectory. In addition to object location, our algorithm can discover a rich set of object surface properties at high precision, including object surface orientation, curvature, boundaries, and surface material. We tested our system on a variety of common household objects (between 5cm--30cm in width). Results show that it achieves high accuracy (cm level) in a variety of dimensions, and is highly robust against noises in device position and trajectory tracking. We believe that this is the first practical mobile imaging system (re)using 60GHz networking devices, and provides a basic primitive towards the construction of detailed environmental mapping systems. Yanzi Zhu, Yibo Zhu 0001, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 3 |
| 2015 | Packet-Level Telemetry in Large Datacenter NetworksabstractDebugging faults in complex networks often requires capturing and analyzing traffic at the packet level. In this task, datacenter networks (DCNs) present unique challenges with their scale, traffic volume, and diversity of faults. To troubleshoot faults in a timely manner, DCN administrators must a) identify affected packets inside large volume of traffic; b) track them across multiple network components; c) analyze traffic traces for fault patterns; and d) test or confirm potential causes. To our knowledge, no tool today can achieve both the specificity and scale required for this task. Yibo Zhu 0001, Nanxi Kang, Jiaxin Cao, Albert G. Greenberg, Guohan Lu, Ratul Mahajan, David A. Maltz, Ming Zhang 0005, Ben Y. Zhao, Haitao Zheng 0001 |
SIGCOMM | 10 |
| 2015 | Energy and Performance of Smartphone Radio Bundling in Outdoor EnvironmentsabstractMost of today's mobile devices come equipped with both cellular LTE and WiFi wireless radios, making radio bundling (simultaneous data transfers over multiple interfaces) both appealing and practical. Despite recent studies documenting the benefits of radio bundling with MPTCP, many fundamental questions remain about potential gains from radio bundling, or the relationship between performance and energy consumption in these scenarios. In this study, we seek to answer these questions using extensive measurements to empirically characterize both energy and performance for radio bundling approaches. In doing so, we quantify potential gains of bundling using MPTCP versus an ideal protocol. We study the links between traffic partitioning and bundling performance, and use a novel componentized energy model to quantify the energy consumed by CPUs (and radios) during traffic management. Our results show that MPTCP achieves only a fraction of the total performance gain possible, and that its energy-agnostic design leads to considerable power consumption by the CPU. We conclude that not only there is room for improved bundling performance, but an energy-aware bundling protocol is likely to achieve a much better tradeoff between performance and power consumption. Ana Nika, Yibo Zhu 0001, Ning Ding 0004, Abhilash Jindal, Y. Charlie Hu, Ben Y. Zhao, Haitao Zheng 0001 |
WWW | 7 |
| 2015 | Practical Conflict Graphs in the WildabstractToday, most spectrum allocation algorithms use conflict graphs to capture interference conditions. The use of conflict graphs, however, is often questioned by the wireless community for two reasons. First, building accurate conflict graphs requires significant overhead, and hence does not scale to outdoor networks. Second, conflict graphs cannot properly capture accumulative interference. In this paper, we use large-scale measurement data as ground truth to understand how severe these problems are and whether they can be overcome. We build “practical” conflict graphs using measurement-calibrated propagation models, which remove the need for exhaustive signal measurements by interpolating signal strengths using calibrated models. Calibrated models are imperfect, and we study the impact of their errors on multiple steps in the process, from calibrating propagation models, predicting signal strengths, to building conflict graphs. At each step, we analyze the introduction, propagation, and final impact of errors by comparing each intermediate result to its ground-truth counterpart. Our work produces several findings. Calibrated propagation models generate location-dependent prediction errors, ultimately producing conservative conflict graphs. While these “estimated conflict graphs” lower spectrum utilization, their conservative nature improves reliability by reducing the impact of accumulative interference. Finally, we propose a graph augmentation technique to address remaining accumulative interference. Zengbin Zhang, Gang Wang 0011, Xiaoxiao Yu, Ben Y. Zhao, Haitao Zheng 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2014 | Link and Triadic Closure Delay: Temporal Metrics for Social Network Dynamics
Matteo Zignani, Sabrina Gaito, Gian Paolo Rossi 0001, Xiaohan Zhao, Haitao Zheng 0001, Ben Y. Zhao |
ICWSM | 6 |
| 2014 | Whispers in the dark: analysis of an anonymous social networkabstractSocial interactions and interpersonal communication has undergone significant changes in recent years. Increasing awareness of privacy issues and events such as the Snowden disclosures have led to the rapid growth of a new generation of anonymous social networks and messaging applications. By removing traditional concepts of strong identities and social links, these services encourage communication between strangers, and allow users to express themselves without fear of bullying or retaliation. Gang Wang 0011, Bolun Wang, Tianyi Wang 0001, Ana Nika, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 6 |
| 2014 | Demystifying 60GHz outdoor picocellsabstractMobile network traffic is set to explode in our near future, driven by the growth of bandwidth-hungry media applications. Current capacity solutions, including buying spectrum, WiFi offloading, and LTE picocells, are unlikely to supply the orders-of-magnitude bandwidth increase we need. In this paper, we explore a dramatically different alternative in the form of 60GHz mmwave picocells with highly directional links. While industry is investigating other mmwave bands (e.g. 28GHz to avoid oxygen absorption), we prefer the unlicensed 60GHz band with highly directional, short-range links (~100m). 60GHz links truly reap the spatial reuse benefits of small cells while delivering high per-user data rates and leveraging efforts on indoor 60GHz PHY technology and standards. Using extensive measurements on off-the-shelf 60GHz radios and system-level simulations, we explore the feasibility of 60GHz picocells by characterizing range, attenuation due to reflections, sensitivity to movement and blockage, and interference in typical urban environments. Our results dispel some common myths, and show that there are no fundamental physical barriers to high-capacity 60GHz outdoor picocells. We conclude by identifying open challenges and associated research opportunities. Yibo Zhu 0001, Zengbin Zhang, Zhinus Marzi, Chris Nelson, Upamanyu Madhow, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 6 |
| 2014 | Cutting the cord: a robust wireless facilities network for data centersabstractToday's network control and management traffic are limited by their reliance on existing data networks. Fate sharing in this context is highly undesirable, since control traffic has very different availability and traffic delivery requirements. In this paper, we explore the feasibility of building a dedicated wireless facilities network for data centers. We propose Angora, a low-latency facilities network using low-cost, 60GHz beamforming radios that provides robust paths decoupled from the wired network, and flexibility to adapt to workloads and network dynamics. We describe our solutions to address challenges in link coordination, link interference and network failures. Our testbed measurements and simulation results show that Angora enables large number of low-latency control paths to run concurrently, while providing low latency end-to-end message delivery with high tolerance for radio and rack failures. Yibo Zhu 0001, Zengbin Zhang, Amin Vahdat, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 6 |
| 2014 | Understanding data hotspots in cellular networksabstractThe unprecedented growth in mobile data usage is posing significant challenges to cellular operators. One key challenge is how to provide quality of service to subscribers when their residing cell is experiencing a significant amount of traffic, i.e. becoming a traffic hotspot. In this paper, we perform an empirical study on data hotspots in today's cellular networks using a 9-week cellular dataset with 734K+ users and 5327 cell sites. Our analysis examines in details static and dynamic characteristics, predictability, and causes of data hotspots, and their correlation with call hotspots. We believe the understanding of these key issues will lead to more efficient and responsive resource management and thus better QoS provision in cellular networks. To the best of our knowledge, our work is the first to characterize in detail traffic hotspots in today's cellular networks using real data. Ana Nika, Asad Ismail, Ben Y. Zhao, Sabrina Gaito, Gian Paolo Rossi 0001, Haitao Zheng 0001 |
QSHINE | 3 |
| 2014 | Man vs. Machine: Practical Adversarial Detection of Malicious Crowdsourcing Workers
Gang Wang 0011, Tianyi Wang 0001, Haitao Zheng 0001, Ben Y. Zhao |
USENIX Security Symposium | 4 |
| 2014 | Uncovering social network Sybils in the wildabstractSybil accounts are fake identities created to unfairly increase the power or resources of a single malicious user. Researchers have long known about the existence of Sybil accounts in online communities such as file-sharing systems, but they have not been able to perform large-scale measurements to detect them or measure their activities. In this article, we describe our efforts to detect, characterize, and understand Sybil account activity in the Renren Online Social Network (OSN). We use ground truth provided by Renren Inc. to build measurement-based Sybil detectors and deploy them on Renren to detect more than 100,000 Sybil accounts. Using our full dataset of 650,000 Sybils, we examine several aspects of Sybil behavior. First, we study their link creation behavior and find that contrary to prior conjecture, Sybils in OSNs do not form tight-knit communities. Next, we examine the fine-grained behaviors of Sybils on Renren using clickstream data. Third, we investigate behind-the-scenes collusion between large groups of Sybils. Our results reveal that Sybils with no explicit social ties still act in concert to launch attacks. Finally, we investigate enhanced techniques to identify stealthy Sybils. In summary, our study advances the understanding of Sybil behavior on OSNs and shows that Sybils can effectively avoid existing community-based Sybil detectors. We hope that our results will foster new research on Sybil detection that is based on novel types of Sybil features. Zhi Yang 0001, Christo Wilson, Xiao Wang 0018, Tingting Gao, Ben Y. Zhao, Yafei Dai |
ACM Trans. Knowl. Discov. Data | 5 |
| 2014 | Preserving Location Privacy in Geosocial ApplicationsabstractUsing geosocial applications, such as FourSquare, millions of people interact with their surroundings through their friends and their recommendations. Without adequate privacy protection, however, these systems can be easily misused, for example, to track users or target them for home invasion. In this paper, we introduce LocX, a novel alternative that provides significantly improved location privacy without adding uncertainty into query results or relying on strong assumptions about server security. Our key insight is to apply secure user-specific, distance-preserving coordinate transformations to all location data shared with the server. The friends of a user share this user's secrets so they can apply the same transformation. This allows all location queries to be evaluated correctly by the server, but our privacy mechanisms guarantee that servers are unable to see or infer the actual location data from the transformed data or from the data access. We show that LocX provides privacy even against a powerful adversary model, and we use prototype measurements to show that it provides privacy with very little performance overhead, making it suitable for today's mobile devices. Krishna P. N. Puttaswamy, Troy Steinbauer, Divyakant Agrawal, Amr El Abbadi, Christopher Krügel, Ben Y. Zhao |
IEEE Trans. Mob. Comput. | 7 |
| 2013 | On the validity of geosocial mobility tracesabstractMobile networking researchers have long searched for large-scale, fine-grained traces of human movement, which have remained elusive for both privacy and logistical reasons. Recently, researchers have begun to focus on geosocial mobility traces, e.g. Foursquare checkin traces, because of their availability and scale. But are we conceding correctness in our zeal for data? In this paper, we take initial steps towards quantifying the value of geosocial datasets using a large ground truth dataset gathered from a user study. By comparing GPS traces against Foursquare checkins, we find that a large portion of visited locations is missing from checkins, and most checkin events are either forged or superfluous events. We characterize extraneous checkins, describe possible techniques for their detection, and show that both extraneous and missing checkins introduce significant errors into applications driven by these traces. Zengbin Zhang, Xiaohan Zhao, Gang Wang 0011, Yu Su 0001, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao |
HotNets | 8 |
| 2013 | Follow the green: growth and dynamics in twitter follower marketsabstractThe users of microblogging services, such as Twitter, use the count of followers of an account as a measure of its reputation or influence. For those unwilling or unable to attract followers naturally, a growing industry of "Twitter follower markets" provides followers for sale. Some markets use fake accounts to boost the follower count of their customers, while others rely on a pyramid scheme to turn non-paying customers into followers for each other, and into followers for paying customers. In this paper, we present a detailed study of Twitter follower markets, report in detail on both the static and dynamic properties of customers of these markets, and develop and evaluate multiple techniques for detecting these activities. We show that our detection system is robust and reliable, and can detect a significant number of customers in the wild. Gianluca Stringhini, Gang Wang 0011, Manuel Egele, Christopher Krügel, Giovanni Vigna, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 7 |
| 2013 | Efficient Batched Synchronization in Dropbox-Like Cloud Storage Services
Zhenhua Li 0001, Christo Wilson, Zhefu Jiang, Yao Liu 0001, Ben Y. Zhao, Cheng Jin 0008, Zhi-Li Zhang, Yafei Dai |
Middleware | 5 |
| 2013 | Social Turing Tests: Crowdsourcing Sybil Detection
Gang Wang 0011, Manish Mohanlal, Christo Wilson, Xiao Wang 0018, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao |
NDSS | 7 |
| 2013 | Characterizing and detecting malicious crowdsourcingabstractPopular Internet services in recent years have shown that remarkable things can be achieved by harnessing the power of the masses. However, crowd-sourcing systems also pose a real challenge to existing security mechanisms deployed to protect Internet services, particularly those tools that identify malicious activity by detecting activities of automated programs such as CAPTCHAs. Tianyi Wang 0001, Gang Wang 0011, Xing Li 0001, Haitao Zheng 0001, Ben Y. Zhao |
SIGCOMM | 5 |
| 2013 | Practical conflict graphs for dynamic spectrum distributionabstractMost spectrum distribution proposals today develop their allocation algorithms that use conflict graphs to capture interference relationships. The use of conflict graphs, however, is often questioned by the wireless community because of two issues. First, building conflict graphs requires significant overhead and hence generally does not scale to outdoor networks, and second, the resulting conflict graphs do not capture accumulative interference. In this paper, we use large-scale measurement data as ground truth to understand just how severe these issues are in practice, and whether they can be overcome. We build "practical" conflict graphs using measurement-calibrated propagation models, which remove the need for exhaustive signal measurements by interpolating signal strengths using calibrated models. These propagation models are imperfect, and we study the impact of their errors by tracing the impact on multiple steps in the process, from calibrating propagation models to predicting signal strength and building conflict graphs. At each step, we analyze the introduction, propagation and final impact of errors, by comparing each intermediate result to its ground truth counterpart generated from measurements. Our work produces several findings. Calibrated propagation models generate location-dependent prediction errors, ultimately producing conservative conflict graphs. While these "estimated conflict graphs" lose some spectrum utilization, their conservative nature improves reliability by reducing the impact of accumulative interference. Finally, we propose a graph augmentation technique that addresses any remaining accumulative interference, the last missing piece in a practical spectrum distribution system using measurement-calibrated conflict graphs. Zengbin Zhang, Gang Wang 0011, Xiaoxiao Yu, Ben Y. Zhao, Haitao Zheng 0001 |
SIGMETRICS | 5 |
| 2013 | You Are How You Click: Clickstream Analysis for Sybil Detection
Gang Wang 0011, Tristan Konolige, Christo Wilson, Xiao Wang 0018, Haitao Zheng 0001, Ben Y. Zhao |
USENIX Security Symposium | 6 |
| 2013 | Wisdom in the social crowd: an analysis of quoraabstractEfforts such as Wikipedia have shown the ability of user communities to collect, organize and curate information on the Internet. Recently, a number of question and answer (Q&A) sites have successfully built large growing knowledge repositories, each driven by a wide range of questions and answers from its users community. While sites like Yahoo Answers have stalled and begun to shrink, one site still going strong is Quora, a rapidly growing service that augments a regular Q&A system with social links between users. Despite its success, however, little is known about what drives Quora's growth, and how it continues to connect visitors and experts to the right questions as it grows. Gang Wang 0011, Konark Gill, Manish Mohanlal, Haitao Zheng 0001, Ben Y. Zhao |
WWW | 5 |
| 2013 | On the Embeddability of Random Walk DistancesabstractAnalysis of large graphs is critical to the ongoing growth of search engines and social networks. One class of queries centers around node affinity, often quantified by random-walk distances between node pairs, including hitting time, commute time, and personalized PageRank (PPR). Despite the potential of these "metrics," they are rarely, if ever, used in practice, largely due to extremely high computational costs. In this paper, we investigate methods to scalably and efficiently compute random-walk distances, by "embedding" graphs and distances into points and distances in geometric coordinate spaces. We show that while existing graph coordinate systems (GCS) can accurately estimate shortest path distances, they produce significant errors when embedding random-walk distances. Based on our observations, we propose a new graph embedding system that explicitly accounts for per-node graph properties that affect random walk. Extensive experiments on a range of graphs show that our new approach can accurately estimate both symmetric and asymmetric random-walk distances. Once a graph is embedded, our system can answer queries between any two nodes in 8 microseconds, orders of magnitude faster than existing methods. Finally, we show that our system produces estimates that can replace ground truth in applications with minimal impact on application output. Xiaohan Zhao, Adelbert Chang, Atish Das Sarma, Haitao Zheng 0001, Ben Y. Zhao |
Proc. VLDB Endow. | 5 |
| 2013 | Measurement-Based Design of Roadside Content Delivery SystemsabstractWith today's ubiquity of thin computing devices, mobile users are accustomed to having rich location-aware information at their fingertips, such as restaurant menus, shopping mall maps, movie showtimes, and trailers. However, delivering rich content is challenging, particularly for highly mobile users in vehicles. Technologies such as cellular-3G provide limited bandwidth at significant costs. In contrast, providers can cheaply and easily deploy a small number of WiFi infostations that quickly deliver large content to vehicles passing by for future offline browsing. While several projects have proposed systems for disseminating content via roadside infostations, most use simplified models and simulations to guide their design for scalability. Many suspect that scalability with increasing vehicle density is the major challenge for infostations, but few if any have studied the performance of these systems via real measurements. Intuitively, per-vehicle throughput for unicast infostations degrades with the number of vehicles near the infostation, while broadcast infostations are unreliable, and lack rate adaptation. In this work, we collect over 200 h of detailed highway measurements with a fleet of WiFi-enabled vehicles. We use analysis of these results to explore the design space of WiFi infostations, in order to determine whether unicast or broadcast should be used to build high-throughput infostations that scale with device density. Our measurement results demonstrate the limitations of both approaches. Our insights lead to Starfish, a high-bandwidth and scalable infostation system that incorporates device-to-device data scavenging, where nearby vehicles share data received from the infostation. Data scavenging increases dissemination throughput by a factor of 2-6, allowing both broadcast and unicast throughput to scale with device density. Vinod Kone, Haitao Zheng 0001, Antony I. T. Rowstron, Greg O'Shea, Ben Y. Zhao |
IEEE Trans. Mob. Comput. | 5 |
| 2013 | Understanding latent interactions in online social networksabstractPopular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions , that is, passive actions, such as profile browsing, that cannot be observed by traditional measurement techniques. In this article, we seek a deeper understanding of both active and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 220 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed, publicly viewable visitor logs for each user profile. We capture detailed histories of profile visits over a period of 90 days for users in the Peking University Renren network and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than active events, are nonreciprocal in nature, and that profile popularity is correlated with page views of content rather than with quantity of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior and compare their structural properties, evolution, community structure, and mixing times against those of both active interaction graphs and social graphs. Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Wenpeng Sha, Peng Huang 0005, Yafei Dai, Ben Y. Zhao |
ACM Trans. Web | 7 |
| 2012 | Multi-scale dynamics in a massive online social networkabstractData confidentiality policies at major social network providers have severely limited researchers' access to large-scale datasets. The biggest impact has been on the study of network dynamics, where researchers have studied citation graphs and content-sharing networks, but few have analyzed detailed dynamics in the massive social networks that dominate the web today. In this paper, we present results of analyzing detailed dynamics in a large Chinese social network, covering a period of 2 years when the network grew from its first user to 19 million users and 199 million edges. Rather than validate a single model of network dynamics, we analyze dynamics at different granularities (per-user, per-community, and network-wide) to determine how much, if any, users are influenced by dynamics processes at different scales. We observe independent predictable processes at each level, and find that the growth of communities has moderate and sustained impact on users. In contrast, we find that significant events such as network merge events have a strong but short-lived impact on users, and they are quickly eclipsed by the continuous arrival of new users. Xiaohan Zhao, Alessandra Sala, Christo Wilson, Xiao Wang 0018, Sabrina Gaito, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 7 |
| 2012 | Enforcing dynamic spectrum access with spectrum permitsabstractDynamic spectrum access is a maturing technology that allows next generation wireless devices to make highly efficient use of wireless spectrum. Spectrum can be allocated on an on-demand basis for a given geographic location, time duration and frequency range. However, a major obstacle to adoption remains. There are no effective solutions to protect licensed users from spectrum misuse, where users transmit without properly licensing spectrum, and in doing so, interfere and disrupt legitimate flows to whom the spectrum is assigned. Given the flexibility of today's cognitive radios, an application can easily transmit on frequencies outside of its allocated range, either accidentally due to misconfiguration, or intentionally to avoid spectrum licensing costs. In this paper, we propose a system to secure dynamic spectrum transmissions, where authorized users embed secure spectrum permits into data transmissions, thus enabling patrolling trusted devices to detect devices transmitting without authorization. We focus our attention on the development of spectrum permits, and describe Gelato, a spectrum misuse detection system that minimizes both hardware costs and performance overhead on legitimate data transmissions. Lei Yang 0020, Zengbin Zhang, Ben Y. Zhao, Christopher Krügel, Haitao Zheng 0001 |
MobiHoc | 3 |
| 2012 | Mirror mirror on the ceiling: flexible wireless links for data centersabstractModern data centers are massive, and support a range of distributed applications across potentially hundreds of server racks. As their utilization and bandwidth needs continue to grow, traditional methods of augmenting bandwidth have proven complex and costly in time and resources. Recent measurements show that data center traffic is often limited by congestion loss caused by short traffic bursts. Thus an attractive alternative to adding physical bandwidth is to augment wired links with wireless links in the 60 GHz band. Zengbin Zhang, Yibo Zhu 0001, Saipriya Kumar, Amin Vahdat, Ben Y. Zhao, Haitao Zheng 0001 |
SIGCOMM | 7 |
| 2012 | Serf and turf: crowdturfing for fun and profitabstractPopular Internet services in recent years have shown that remarkable things can be achieved by harnessing the power of the masses using crowd-sourcing systems. However, crowd-sourcing systems can also pose a real challenge to existing security mechanisms deployed to protect Internet services. Many of these security techniques rely on the assumption that malicious activity is generated automatically by automated programs. Thus they would perform poorly or be easily bypassed when attacks are generated by real users working in a crowd-sourcing system. Through measurements, we have found surprising evidence showing that not only do malicious crowd-sourcing systems exist, but they are rapidly growing in both user base and total revenue. We describe in this paper a significant effort to study and understand these "crowdturfing" systems in today's Internet. We use detailed crawls to extract data about the size and operational structure of these crowdturfing systems. We analyze details of campaigns offered and performed in these sites, and evaluate their end-to-end effectiveness by running active, benign campaigns of our own. Finally, we study and compare the source of workers on crowdturfing sites in different countries. Our results suggest that campaigns on these systems are highly effective at reaching users, and their continuing growth poses a concrete threat to online communities both in the US and elsewhere. Gang Wang 0011, Christo Wilson, Xiaohan Zhao, Yibo Zhu 0001, Manish Mohanlal, Haitao Zheng 0001, Ben Y. Zhao |
WWW | 7 |
| 2012 | The effectiveness of opportunistic spectrum access: a measurement studyabstractDynamic spectrum access networks are designed to allow today's bandwidth-hungry “secondary devices” to share spectrum allocated to legacy devices, or “primary users.” The success of this wireless communication model relies on the availability of unused spectrum and the ability of secondary devices to utilize spectrum without disrupting transmissions of primary users. While recent measurement studies have shown that there is sufficient underutilized spectrum available, little is known about whether secondary devices can efficiently make use of available spectrum while minimizing disruptions to primary users. In this paper, we present the first comprehensive study on the presence of “usable” spectrum in opportunistic spectrum access systems, and whether sufficient spectrum can be extracted by secondary devices to support traditional networking applications. We use for our study fine-grain usage traces of a wide spectrum range (20 MHz–6 GHz) taken at four locations in Germany, the Netherlands, and Santa Barbara, CA. Our study shows that on average, 54% of spectrum is never used and 26% is only partially used. Surprisingly, in this 26% of partially used spectrum, secondary devices can utilize very little spectrum using conservative access policies to minimize interference with primary users. Even assuming an optimal access scheme and extensive statistical knowledge of primary-user access patterns, a user can only extract between 20%–30% of the total available spectrum. To provide better spectrum availability, we propose frequency bundling, where secondary devices build reliable channels by combining multiple unreliable frequencies into virtual frequency bundles. Analyzing our traces, we find that there is little correlation of spectrum availability across channels, and that bundling random channels together can provide sustained periods of reliable transmission with only short interruptions. Vinod Kone, Lei Yang 0020, Xue Yang 0007, Ben Y. Zhao, Haitao Zheng 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2012 | Beyond Social Graphs: User Interactions in Online Social Networks and their ImplicationsabstractSocial networks are popular platforms for interaction, communication, and collaboration between friends. Researchers have recently proposed an emerging class of applications that leverage relationships from social networks to improve security and performance in applications such as email, Web browsing, and overlay routing. While these applications often cite social network connectivity statistics to support their designs, researchers in psychology and sociology have repeatedly cast doubt on the practice of inferring meaningful relationships from social network connections alone. This leads to the question: “Are social links valid indicators of real user interaction? If not, then how can we quantify these factors to form a more accurate model for evaluating socially enhanced applications?” In this article, we address this question through a detailed study of user interactions in the Facebook social network. We propose the use of “interaction graphs” to impart meaning to online social links by quantifying user interactions. We analyze interaction graphs derived from Facebook user traces and show that they exhibit significantly lower levels of the “small-world” properties present in their social graph counterparts. This means that these graphs have fewer “supernodes” with extremely high degree, and overall graph diameter increases significantly as a result. To quantify the impact of our observations, we use both types of graphs to validate several well-known social-based applications that rely on graph properties to infuse new functionality into Internet applications, including Reliable Email (RE), SybilGuard, and the weighted cascade influence maximization algorithm. The results reveal new insights into each of these systems, and confirm our hypothesis that to obtain realistic and accurate results, ongoing research on social network applications studies of social applications should use real indicators of user interactions in lieu of social graphs. Christo Wilson, Alessandra Sala, Krishna P. N. Puttaswamy, Ben Y. Zhao |
ACM Trans. Web | 4 |
| 2011 | Silverline: toward data confidentiality in storage-intensive cloud applicationsabstractBy offering high availability and elastic access to resources, third-party cloud infrastructures such as Amazon EC2 are revolutionizing the way today's businesses operate. Unfortunately, taking advantage of their benefits requires businesses to accept a number of serious risks to data security. Factors such as software bugs, operator errors and external attacks can all compromise the confidentiality of sensitive application data on external clouds, by making them vulnerable to unauthorized access by malicious parties. Krishna P. N. Puttaswamy, Christopher Krügel, Ben Y. Zhao |
SoCC | 3 |
| 2011 | Efficient shortest paths on massive social graphsabstractAnalysis of large networks is a critical component of many of today’s application environments. The arrival of massive network graphs with hundreds of millions of nodes, e.g. social graphs, presents a unique challenge to graph analysis applications. Most of these applications rely on computing dista Xiaohan Zhao, Alessandra Sala, Haitao Zheng 0001, Ben Y. Zhao |
CollaborateCom | 4 |
| 2011 | 3D beamforming for wireless data centersabstractContrary to prior assumptions, recent measurements show that data center traffic is not constrained by network bisection bandwidth, but is instead prone to congestion loss caused by short traffic bursts. Compared to the cost and complexity of modifying data center architectures, a much more attractive option is to augment wired links with flexible wireless links in the 60 GHz band. Current proposals, however, are severely constrained by two factors. First, 60 GHz wireless links are limited by line-of-sight, and can be blocked by even small obstacles between the endpoints. Second, even beamforming links leak power, and potential interference will severely limit concurrent transmissions in dense data centers. In this paper, we explore the feasibility of a new wireless primitive for data centers, 3D beamforming. We explore the design space, and show how bouncing 60 GHz wireless links off reflective ceilings can address both link blockage and link interference, thus improving link range and number of current transmissions in the data center. Weile Zhang, Lei Yang 0020, Zengbin Zhang, Ben Y. Zhao, Haitao Zheng 0001 |
HotNets | 5 |
| 2011 | BTLab: A System-Centric, Data-Driven Analysis and Measurement Platform for BitTorrent ClientsabstractWe present BTLab, a distributed platform to measure and analyze the differences between BitTorrent clients. Due to extensibility, and a certain vagueness in the BitTorrent specification, many clients diverge in some aspects from each other. Most research to date disregarded the effects of these differences. BTLab allows us to create and control BitTorrent swarms, composed of hundreds of clients of our choice. We use captured network traffic to measure the performance and uncover the reasons for observed differences. For our experiments, we deployed BTLab on a cluster and on Planetlab and selected four popular BitTorrent clients. Our analysis reveals flaws in piece selection and connection management algorithms that adversely affect the performance of some clients. Martin Szydlowski, Ben Y. Zhao, Engin Kirda, Christopher Krügel |
ICCCN | 2 |
| 2011 | Sharing graphs using differentially private graph modelsabstractContinuing success of research on social and computer networks requires open access to realistic measurement datasets. While these datasets can be shared, generally in the form of social or Internet graphs, doing so often risks exposing sensitive user data to the public. Unfortunately, current techniques to improve privacy on graphs only target specific attacks, and have been proven to be vulnerable against powerful de-anonymization attacks. Alessandra Sala, Xiaohan Zhao, Christo Wilson, Haitao Zheng 0001, Ben Y. Zhao |
Internet Measurement Conference | 5 |
| 2011 | Uncovering social network sybils in the wildabstractSybil accounts are fake identities created to unfairly increase the power or resources of a single user. Researchers have long known about the existence of Sybil accounts in online communities such as file-sharing systems, but have not been able to perform large scale measurements to detect them or measure their activities. In this paper, we describe our efforts to detect, characterize and understand Sybil account activity in the Renren online social network (OSN). We use ground truth provided by Renren Inc. to build measurement based Sybil account detectors, and deploy them on Renren to detect over 100,000 Sybil accounts. We study these Sybil accounts, as well as an additional 560,000 Sybil accounts caught by Renren, and analyze their link creation behavior. Most interestingly, we find that contrary to prior conjecture, Sybil accounts in OSNs do not form tight-knit communities. Instead, they integrate into the social graph just like normal users. Using link creation timestamps, we verify that the large majority of links between Sybil accounts are created accidentally, unbeknownst to the attacker. Overall, only a very small portion of Sybil accounts are connected to other Sybils with social links. Our study shows that existing Sybil defenses are unlikely to succeed in today's OSNs, and we must design new techniques to effectively detect and defend against Sybil attacks. Zhi Yang 0001, Christo Wilson, Xiao Wang 0018, Tingting Gao, Ben Y. Zhao, Yafei Dai |
Internet Measurement Conference | 5 |
| 2011 | Scaling Microblogging Services with Divergent Traffic Demands
Tianyin Xu, Yang Chen 0001, Lei Jiao 0002, Ben Y. Zhao, Pan Hui 0001, Xiaoming Fu 0001 |
Middleware | 4 |
| 2011 | I am the antenna: accurate outdoor AP location using smartphonesabstractToday's WiFi access points (APs) are ubiquitous, and provide critical connectivity for a wide range of mobile networking devices. Many management tasks, e.g. optimizing AP placement and detecting rogue APs, require a user to efficiently determine the location of wireless APs. Unlike prior localization techniques that require either specialized equipment or extensive outdoor measurements, we propose a way to locate APs in real-time using commodity smartphones. Our insight is that by rotating a wireless receiver (smartphone) around a signal-blocking obstacle (the user's body), we can effectively emulate the sensitivity and functionality of a directional antenna. Our measurements show that we can detect these signal strength artifacts on multiple smartphone platforms for a variety of outdoor environments. We develop a model for detecting signal dips caused by blocking obstacles, and use it to produce a directional analysis technique that accurately predicts the direction of the AP, along with an associated confidence value. The result is Borealis, a system that provides accurate directional guidance and leads users to a desired AP after a few measurements. Detailed measurements show that Borealis is significantly more accurate than other real-time localization systems, and is nearly as accurate as offline approaches using extensive wireless measurements. Zengbin Zhang, Weile Zhang, Yuanyang Zhang, Gang Wang 0011, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 6 |
| 2011 | The Impact of Infostation Density on Vehicular Data DisseminationabstractVehicle-to-Vehicle and Vehicle-to-Roadside communications are going to become an indispensable part of the modern day automotive experience. For people on the move, vehicular networks can provide critical network connectivity and access to real-time information. Infostations play a vital role in these networks by acting as gateways to the Internet and by extending network connectivity. In this context, an important question is “What is the minimum number of infostations that need to be deployed in an area in order to support vehicular applications?” Optimizing infostation density is vital to understanding and reducing the cost of deployment and management. In this paper, we examine the required infostation density in a highway scenario using different data dissemination models. We start from a simple analysis that captures the required density under idealized assumptions. These models are validated by an event-driven simulator. We then run detailed QualNet simulations on both controlled and realistic vehicular traces to observe the information density trends in practical environments, and consequently propose techniques to improve dissemination performance and reduce the required infostation density. Vinod Kone, Haitao Zheng 0001, Antony I. T. Rowstron, Ben Y. Zhao |
Mob. Networks Appl. | 4 |
| 2011 | Protector: A Probabilistic Failure Detector for Cost-Effective Peer-to-Peer StorageabstractMaintaining a given level of data redundancy is a fundamental requirement of peer-to-peer (P2P) storage systems—to ensure desired data availability, additional replicas must be created when peers fail. Since the majority of failures in P2P networks are transient (i.e., peers return with data intact), an intelligent system can reduce significant replication costs by not replicating data following transient failures. Reliably distinguishing permanent and transient failures, however, is a challenging task, because peers are unresponsive to probes in both cases. In this paper, we propose Protector, an algorithm that enables efficient replication policies by estimating the number of “remaining replicas” for each object, including those temporarily unavailable due to transient failures. Protector dramatically improves detection accuracy by exploiting two opportunities. First, it leverages failure patterns to predict the likelihood that a peer (and the data it hosts) has permanently failed given its current downtime. Second, it detects replication level across groups of replicas (or fragments), thereby balancing false positives for some peers against false negatives for others. Extensive simulations based on both synthetic and real traces show that Protector closely approximates the performance of a perfect “oracle” failure detector, and significantly outperforms time-out-based detectors using a wide range of parameters. Finally, we design, implement and deploy an efficient P2P storage system called AmazingStore by combining Protector with structured P2P overlays. Our experience proves that Protector enables efficient long-term data maintenance in P2P storage systems. Zhi Yang 0001, Ben Y. Zhao, Wei Chen 0013, Yafei Dai |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2010 | Detecting and characterizing social spam campaignsabstractOnline social networks (OSNs) are exceptionally useful collaboration and communication tools for millions of users and their friends. Unfortunately, in the wrong hands, they are also extremely effective tools for executing spam campaigns and spreading malware. Christo Wilson, Zhichun Li, Yan Chen 0004, Ben Y. Zhao |
CCS | 6 |
| 2010 | Exploiting locality of interest in online social networksabstractOnline Social Networks (OSN) are fun, popular, and socially significant. An integral part of their success is the immense size of their global user base. To provide a consistent service to all users, Facebook, the world’s largest OSN, is heavily dependent on centralized U.S. data centers, which renders service outside of the U.S. sluggish and wasteful of Internet bandwidth. In this paper, we investigate the detailed causes of these two problems and identify mitigation opportunities. Because details of Facebook’s service remain proprietary, we treat the OSN as a black box and reverse engineer its operation from publicly available traces. We find that contrary to current wisdom, OSN state is amenable to partitioning and that its fine grained distribution and processing can significantly improve performance without loss in service consistency. Through simulations of reconstructed Facebook traffic over measured Internet paths, we show that user requests can be processed 79% faster and use 91% less bandwidth. We conclude that the partitioning of OSN state is an attractive scaling strategy for Facebook and other OSN services. Mike P. Wittie, Veljko Pejovic, Lara B. Deek, Kevin C. Almeroth, Ben Y. Zhao |
CoNEXT | 5 |
| 2010 | Detecting and characterizing social spam campaignsabstractOnline social networks (OSNs) are popular collaboration and communication tools for millions of users and their friends. Unfortunately, in the wrong hands, they are also effective tools for executing spam campaigns and spreading malware. Intuitively, a user is more likely to respond to a message from a Facebook friend than from a stranger, thus making social spam a more effective distribution mechanism than traditional email. In fact, existing evidence shows malicious entities are already attempting to compromise OSN account credentials to support these "high-return" spam campaigns. In this paper, we present an initial study to quantify and characterize spam campaigns launched using accounts on online social networks. We study a large anonymized dataset of asynchronous "wall" messages between Facebook users. We analyze all wall messages received by roughly 3.5 million Facebook users (more than 187 million messages in all), and use a set of automated techniques to detect and characterize coordinated spam campaigns. Our system detected roughly 200,000 malicious wall posts with embed- ded URLs, originating from more than 57,000 user accounts. We find that more than 70% of all malicious wall posts advertise phishing sites. We also study the characteristics of malicious accounts, and see that more than 97% are compromised accounts, rather than "fake" accounts created solely for the purpose of spamming. Finally, we observe that, when adjusted to the local time of the sender, spamming dominates actual wall post activity in the early morning hours, when normal users are asleep. Christo Wilson, Zhichun Li, Yan Chen 0004, Ben Y. Zhao |
Internet Measurement Conference | 6 |
| 2010 | Understanding latent interactions in online social networksabstractPopular online social networks (OSNs) like Facebook and Twitter are changing the way users communicate and interact with the Internet. A deep understanding of user interactions in OSNs can provide important insights into questions of human social behavior, and into the design of social platforms and applications. However, recent studies have shown that a majority of user interactions on OSNs are latent interactions, passive actions such as profile browsing that cannot be observed by traditional measurement techniques. In this paper, we seek a deeper understanding of both visible and latent user interactions in OSNs. For quantifiable data on latent user interactions, we perform a detailed measurement study on Renren, the largest OSN in China with more than 150 million users to date. All friendship links in Renren are public, allowing us to exhaustively crawl a connected graph component of 42 million users and 1.66 billion social links in 2009. Renren also keeps detailed visitor logs for each user profile, and counters for each photo and diary/blog entry. We capture detailed histories of profile visits over a period of 90 days for more than 61,000 users in the Peking University Renren network, and use statistics of profile visits to study issues of user profile popularity, reciprocity of profile visits, and the impact of content updates on user popularity. We find that latent interactions are much more prevalent and frequent than visible events, non-reciprocal in nature, and that profile popularity are uncorrelated with the frequency of content updates. Finally, we construct latent interaction graphs as models of user browsing behavior, and compare their structural properties against those of both visible interaction graphs and social graphs. Jing Jiang 0005, Christo Wilson, Xiao Wang 0018, Peng Huang 0005, Wenpeng Sha, Yafei Dai, Ben Y. Zhao |
Internet Measurement Conference | 7 |
| 2010 | On the feasibility of effective opportunistic spectrum accessabstractDynamic spectrum access networks are designed to allow today's bandwidth hungry "secondary devices" to share spectrum allocated to legacy devices, or "primary users." The success of this wireless communication model relies on the availability of unused spectrum, and the ability of secondary devices to utilize spectrum without disrupting transmissions of primary users. While recent measurement studies have shown that there is sufficient underutilized spectrum available, little is known about whether secondary devices can efficiently make use of available spectrum while minimizing disruptions to primary users. Vinod Kone, Lei Yang 0020, Xue Yang 0007, Ben Y. Zhao, Haitao Zheng 0001 |
Internet Measurement Conference | 4 |
| 2010 | The spaces between us: setting and maintaining boundaries in wireless spectrum accessabstractGuardbands are designed to insulate transmissions on adjacent frequencies from mutual interference. As more devices in a given area are packed into orthogonal wireless channels, choosing the right guardband size to minimize cross-channel interference becomes critical to network performance. Using both WiFi and GNU radio experiments, we show that the traditional "one-size-fits-all" approach to guardband assignment is ineffective, and can produce throughput degradation up to 80%. We find that ideal guardband values vary across different network configurations, and across different links in the same network. We argue that guardband values should be set based on network conditions and adapt to changes over time. Lei Yang 0020, Ben Y. Zhao, Haitao Zheng 0001 |
MobiCom | 2 |
| 2010 | Supporting Demanding Wireless Applications with Frequency-agile Radios
Lei Yang 0020, Lili Cao, Ben Y. Zhao, Haitao Zheng 0001 |
NSDI | 4 |
| 2010 | Brief announcement: revisiting the power-law degree distribution for social graph analysisabstractThe study of complex networks led to the belief that the connectivity of network nodes generally follows a Power-law distribution. In this work, we show that modeling large-scale online social networks using a Power-law distribution produces significant fitting errors. We propose the use of a more accurate node degree distribution model based on the Pareto-Lognormal distribution. Using large datasets gathered from Facebook, we show that the Power-law curve produces a significant over-estimation of the number of high degree nodes, leading researchers to erroneous designs for a number of social applications and systems, including shortest-path prediction, community detection, and influence maximization. We provide a formal proof of the error reduction using the Pareto-Lognormal distribution, which we envision will have strong implications on the correctness of social systems and applications. Alessandra Sala, Haitao Zheng 0001, Ben Y. Zhao, Sabrina Gaito, Gian Paolo Rossi 0001 |
PODC | 3 |
| 2010 | Coexistence-Aware Scheduling for Wireless System-on-a-Chip DevicesabstractToday's mobile devices support many wireless technologies to achieve ubiquitous connectivity. Economic and energy constraints, however, are driving the industry to implement multiple technologies into a single radio. This system-on-a-chip architecture leads to competition among networks when devices toggle across different technologies to communicate with multiple networks. In this paper, we study the impact of such network competition using a representative scenario where devices split their time between WiMAX and WiFi connections. We show that competition with WiMAX significantly lowers WiFi's throughput, but this performance degradation is largely unnecessary, and can be attributed to the fact that WiMAX's transmission scheduling does not consider competing networks. We propose PACT, a new coexistence-aware WiMAX scheduling policy that cooperates with WiFi links hosted by its users without compromising its own transmission requirements. We derive PACT's design using an analytical model of network competition, and apply it to design practical WiMAX scheduling algorithms for various traffic classes. We evaluate PACT using OPNET's realistic models for WiFi and WiMAX. Using real network topologies, our experiment results show that PACT significantly improves WiFi performance by up to 17 fold without affecting the WiMAX user experience. Lei Yang 0020, Vinod Kone, Xue Yang 0007, York Liu, Ben Y. Zhao, Haitao Zheng 0001 |
SECON | 5 |
| 2010 | Measurement-calibrated graph models for social network experimentsabstractAccess to realistic, complex graph datasets is critical to research on social networking systems and applications. Simulations on graph data provide critical evaluation of new systems and applications ranging from community detection to spam filtering and social web search. Due to the high time and resource costs of gathering real graph datasets through direct measurements, researchers are anonymizing and sharing a small number of valuable datasets with the community. However, performing experiments using shared real datasets faces three key disadvantages: concerns that graphs can be de-anonymized to reveal private information, increasing costs of distributing large datasets, and that a small number of available social graphs limits the statistical confidence in the results. Alessandra Sala, Lili Cao, Christo Wilson, Robert Zablit, Haitao Zheng 0001, Ben Y. Zhao |
WWW | 6 |
| 2009 | StarClique: guaranteeing user privacy in social networks against intersection attacksabstractBuilding on the popularity of online social networks (OSNs) such as Facebook, social content-sharing applications allow users to form communities around shared interests. Millions of users worldwide use them to share recommendations on everything from music and books to resources on the web. However, their increasing popularity is beginning to attract the attention of malicious attackers. As social network credentials become valued targets of phishing attacks and social worms, attackers look to leverage compromised accounts for further financial gain. Krishna P. N. Puttaswamy, Alessandra Sala, Ben Y. Zhao |
CoNEXT | 3 |
| 2009 | User interactions in social networks and their implicationsabstractSocial networks are popular platforms for interaction, communication and collaboration between friends. Researchers have recently proposed an emerging class of applications that leverage relationships from social networks to improve security and performance in applications such as email, web browsing and overlay routing. While these applications often cite social network connectivity statistics to support their designs, researchers in psychology and sociology have repeatedly cast doubt on the practice of inferring meaningful relationships from social network connections alone. Christo Wilson, Bryce Boe, Alessandra Sala, Krishna P. N. Puttaswamy, Ben Y. Zhao |
EuroSys | 5 |
| 2009 | Rome: Performance and Anonymity using Route MeshesabstractDeployed anonymous networks such as Tor focus on delivering messages through end-to-end paths with high anonymity. Selection of routers in the anonymous path construction is either performed randomly, or relies on self-described resource availability at routers, making systems vulnerable to low-resource attacks. In this paper, we investigate an alternative router and path selection mechanism for constructing efficient end-to-end paths with low loss of path anonymity. We propose a novel construct called a "route mesh," and a dynamic programming algorithm that determines optimal-latency paths from many random samples using only a small number of end-to-end measurements. We prove analytically that our path search algorithm finds the optimal path, and requires exponentially lower number of measurements compared to a standard measurement approach. In addition, our analysis shows that route meshes incur only a small loss in anonymity for its users. Krishna P. N. Puttaswamy, Alessandra Sala, Ömer Egecioglu, Ben Y. Zhao |
INFOCOM | 4 |
| 2009 | Peer-exchange schemes to handle mismatch in peer-to-peer systems
Tongqing Qiu, Edward Chan, Mao Ye 0010, Guihai Chen, Ben Y. Zhao |
J. Supercomput. | 5 |
| 2009 | Securing Structured Overlays against Identity AttacksabstractStructured overlay networks can greatly simplify data storage and management for a variety of distributed applications. Despite their attractive features, these overlays remain vulnerable to the Identity attack, where malicious nodes assume control of application components by intercepting and hijacking key-based routing requests. Attackers can assume arbitrary application roles such as storage node for a given file, or return falsified contents of an online shopper's shopping cart. In this paper, we define a generalized form of the Identity attack, and propose a lightweight detection and tracking system that protects applications by redirecting traffic away from attackers. We describe how this attack can be amplified by a Sybil or Eclipse attack, and analyze the costs of performing such an attack. Finally, we present measurements of a deployed overlay that show our techniques to be significantly more lightweight than prior techniques, and highly effective at detecting and avoiding both single node and colluding attacks under a variety of conditions. Krishna P. N. Puttaswamy, Haitao Zheng 0001, Ben Y. Zhao |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2008 | Protecting anonymity in dynamic peer-to-peer networksabstractPeer-to-peer anonymous networks offer the resources to support todaypsilas Internet applications. In todaypsilas dynamic networks, the key challenge to these systems arises from node dynamics and failures that disrupt anonymous routing paths, forcing them to be frequently rebuilt. Not only do these path rebuilds interrupt application sessions, but they also leak information to logging attacks such as the predecessor attack, leading to significant degradation of anonymity over long sessions. In this paper, we propose Bluemoon, a new anonymous protocol that provides strong resilience against the predecessor attack through the use of persistent anonymous links called hooks. When chained together, these links create robust anonymous paths that avoid path disruptions and rebuilds across node failures. Through detailed analysis, we show that relative to prior approaches, Bluemoon provides significantly stronger resistance against predecessor attacks. Finally, we implement and deploy a prototype on both local and Internet-scale network testbeds, and show that it provides high throughput even in high-load environments such as PlanetLab. Krishna P. N. Puttaswamy, Alessandra Sala, Christo Wilson, Ben Y. Zhao |
ICNP | 4 |
| 2008 | Searching for Rare Objects Using Index ReplicationabstractSearching for objects is a fundamental problem for popular peer-to-peer file-sharing networks that contribute to much of the traffic on today's Internet. While existing protocols can effectively locate highly popular files, studies show that they fail to locate a significant portion of existing files in the network. High recall for these "rare" objects would drastically improve the user experience, and make these networks the ideal distribution infrastructure for user-generated content such as home videos and photo albums. In this paper, we examine simple techniques that can improve search recall for rare objects while minimizing the overhead incurred by participating peers. We propose several strategies for multi-hop index replication, and demonstrate their effectiveness and efficiency through both analysis and simulation. We further evaluate our simple techniques using detailed traces from a real Gnutella network, and show that they improve the performance of these overlays by orders of magnitude in both lookup success and overhead. Krishna P. N. Puttaswamy, Alessandra Sala, Ben Y. Zhao |
INFOCOM | 3 |
| 2008 | Towards Reliable Reputations for Dynamic Networked SystemsabstractA new generation of distributed systems and applications rely on the cooperation of diverse user populations motivated by self-interest. While they can utilize "reputation systems" to reduce selfish behaviors that disrupt or manipulate the network for personal gain, current reputations face a key challenge in large dynamic networks: vulnerability to peer collusion. In this paper, we propose to dramatically improve the accuracy of reputation systems with the use of a statistical metric that measures the "reliability" of a peer's reputation taking into account collusion-like behavior. Trace-driven simulations on P2P network traffic show that our reliability metric drastically improves system performance. We also apply our metric to 18,000 randomly selected eBay user reputation profiles, and surprisingly discover numerous users with collusion-like behaviors worthy of additional investigation. Gayatri Swamynathan, Ben Y. Zhao, Kevin C. Almeroth, Sreenivasa Rao Jammalamadaka |
SRDS | 2 |
| 2008 | Probabilistic Failure Detection for Efficient Distributed Storage MaintenanceabstractDistributed storage systems often use data replication to mask failures and guarantee high data availability. Node failures can be transient or permanent. While the system must generate new replicas to replace replica lost to permanent failures, it can save significant replication costs by not replicating following transient faults. Given the unpredictability of network dynamics, however, distinguishing permanent and transient failures is extremely difficult. Traditional timeout approaches are difficult to tune and can introduce unnecessary replication. In this paper, we propose Protector, an algorithm that addresses this problem using network-wide statistical prediction. Our algorithm drastically improves prediction accuracy by making predictions across aggregate replica groups instead of single nodes. These estimates of the number of "live replicas" can guide efficient data replication policies. We prove that given data on node down times and the probability of permanent failures, the estimate given by our algorithm is more accurate than all alternatives. We describe two ways to obtain the failure probability function driven by models or traces. We conduct extensive simulations based both on synthetic and real traces, and show that Protector closely approximates the performance of a perfect "oracle" failure detector, while significantly outperforming timeout-based detectors using a wide range of parameters. Zhi Yang 0001, Wei Chen 0013, Ben Y. Zhao, Yafei Dai |
SRDS | 4 |
| 2008 | Exploring the feasibility of proactive reputationsabstractAbstract Reputation mechanisms help peers in a peer‐to‐peer system avoid unreliable or malicious peers. In application‐level networks, however, short peer lifetimes mean reputations are often generated from a small number of past transactions. These reputation values are less ‘reliable’, and more vulnerable to bad‐mouthing or collusion attacks. We address this issue by introducing proactive reputations, a first‐hand history of transactions initiated to augment incomplete or short‐term reputation values. We present several mechanisms to generate proactive reputations, along with a statistical similarity metric to measure their effectiveness. Copyright © 2007 John Wiley & Sons, Ltd. Gayatri Swamynathan, Ben Y. Zhao, Kevin C. Almeroth |
Concurr. Comput. Pract. Exp. | 2 |
| 2008 | Special Issue: Recent Advances in Peer-to-Peer Systems and SecurityabstractPeer-to-peer systems have revolutionized the way we store 1, 2, disseminate 3, 4 and share 5, 6 content. In recent years, researchers have examined numerous aspects of these innovative architectures and proposed numerous protocols and applications. While peer-to-peer research has impacted a number of areas, including theory, networking, databases and distributed systems, recent work has focused more on the issues of practical, deployable systems, measurements, and issues of security and privacy in peer-to-peer applications. This issue brings together recent work on several projects focused on building robust, practical peer-to-peer systems. The work described has also been presented at the 5th international workshop on peer-to-peer systems (IPTPS 2006), an annual forum for researchers and practitioners of large distributed systems, held in Santa Barbara, California, and attended by more than 80 participants. The goal of the workshop was to examine peer-to-peer technologies, applications and systems, and to identify key research issues and challenges that lie ahead. In the context of this workshop, peer-to-peer systems were characterized as being decentralized, self-organizing distributed systems, in which all or most communication is symmetric. This issue includes five papers, two of which 7, 8 present novel approaches to enhance the efficiency of current file-sharing networks, while the other three 9-11 examine security-related issues including reputations, incentives and misbehaving users. In the first paper, Epema et al. 7 propose Tribler, a new content-sharing protocol that leverages social relationships to enhance performance. While Tribler leverages BitTorrent's downloading mechanism, it is primarily a social network, where users obtain persistent, long-term identities, and self-organize according to their preferences and operational history. By leveraging shared interests, Tribler encourages users to behave altruistically, and shows that with willing helpers, downloaders can dramatically reduce their download latencies. The second paper 8 proposes the novel application of a probabilistic counting algorithm to hybrid search. To estimate whether a file is popular, the authors propose that super-peers count its occurrence by applying the Flajolet–Martin algorithm, where each member of a set flips a fair coin up to N times, and counts the number of heads seen before the first tail. The maximum value from all members is an approximate count of the set size. The remaining papers discuss issues closely related to reputations and user behavior. In the first paper, Swamynathan et al. 11 examine one of the challenges of applying reputation systems to peer-to-peer networks, that of short session histories. She proposes a proactive reputation scheme that allows peers to actively probe and test unknown peers with sample requests, all while making the probes indistinguishable from normal transactional requests and remaining anonymous. Second, Lian et al. 9 show using measurements that in the presence of colluding users, global reputation systems such as EigenTrust can produce non-ideal results. They propose instead a multi-trust reputation approach that bridges the gap between personal trust and global trust, and shows how it addresses the issues of collusion behavior found in current peer-to-peer networks. Finally, Liogkas et al. 10 explore the effectiveness of incentive and enforcement mechanisms in the popular BitTorrent network by implementing and evaluating three different exploits. They evaluate their strategies on both private and public BitTorrent networks, and present some surprising results. Ben Y. Zhao |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Multi-channel Jamming Attacks using Cognitive RadiosabstractTo improve spectrum efficiency, future wireless devices will use cognitive radios to dynamically access spectrum. While offering great flexibility and software-reconfigurability, unsecured cognitive radios can be easily manipulated to attack legacy and future wireless networks. In this paper, we explore the feasibility and impact of cognitive radio based jamming attacks on 802.11 networks. We show that attackers can utilize cognitive radios' fast channel switching capability to amplify their jamming impact across multiple channels using a single radio. We also examine the impact of hardware channel switching delays and jamming duration on the impact of jamming. Ashwin Sampath, Hui Dai, Haitao Zheng 0001, Ben Y. Zhao |
ICCCN | 4 |
| 2007 | Deploying Video-on-Demand Services on Cable NetworksabstractEfficient video-on-demand (VoD) is a highly desired service for media and telecom providers. VoD allows subscribers to view any item in a large media catalog nearly-instantaneously. However, systems that provide this services currently require large amounts of centralized resources and significant bandwidth to accommodate their subscribers. Hardware requirements become more substantial as the service providers increase the catalog size or number of subscribers. In this paper, we describe how cable companies can leverage deployed hardware in a peer- to-peer architecture to provide an efficient alternative We propose a distributed VoD system, and use real measurements from a deployed VoD system to evaluate different design decisions. Our results show that with minor changes, currently deployed cable infrastructures can support a video-on-demand system that scales to a large number of users and catalog size with low centralized resources. Matthew S. Allen, Ben Y. Zhao, Richard Wolski |
ICDCS | 2 |
| 2007 | An Empirical Study of Collusion Behavior in the Maze P2P File-Sharing SystemabstractPeer-to-peer networks often use incentive policies to encourage cooperation between nodes. Such systems are generally susceptible to collusion by groups of users in order to gain unfair advantages over others. While techniques have been proposed to combat Web spam collusion, there are few measurements of real collusion in deployed systems. In this paper, we report analysis and measurement results of user collusion in Maze, a large-scale peer-to-peer file sharing system with a non-net-zero point-based incentive policy. We search for colluding behavior by examining complete user logs, and incrementally refine a set of collusion detectors to identify common collusion patterns. We find collusion patterns similar to those found in Web spamming. We evaluate how proposed reputation systems would perform on the Maze system. Our results can help guide the design of more robust incentive schemes. Qiao Lian, Zheng Zhang 0001, Mao Yang 0004, Ben Y. Zhao, Yafei Dai, Xiaoming Li 0001 |
ICDCS | 4 |
| 2007 | Towards Location-aware Topology in both Unstructured and Structured P2P SystemsabstractA self-organizing peer-to-peer system is built upon an application level overlay, whose topology is independent of underlying physical network. A well-routed message path in such systems may result in a long delay and excessive traffic due to the mismatch between logical and physical networks. In order to solve this problem, we present a family of Peer-exchange Routing Optimization Protocols (PROP) to reconstruct the overlay. It includes two policies: PROP- G for generic condition and PROP-0 for optimized one. Both theoretical analysis and simulation experiments show that these two protocols greatly reduce the average latency of the overlay and achieve a location-aware topology with low overhead. Their overall performance can be further improved if combined with other recent approaches. Specifically, PROP-G can be easily applied to both structured and unstructured systems without the loss of their primary characteristics, such as efficient routing and anonymity. PROP- O, on the other hand, is more efficient, especially in a heterogeneous environment where nodes have different processing capabilities. Tongqing Qiu, Guihai Chen, Mao Ye 0010, Edward Chan, Ben Y. Zhao |
ICPP | 5 |
| 2007 | Fairness Attacks in the Explicit Control ProtocolabstractProtocols such as the explicit control protocol (XCP) use explicit router feedback to guide endpoint transmission rates for near-optimal capacity utilization and fairness. However, non-cooperative end hosts can manipulate and ignore feedback to either obtain unfair advantages over cooperative hosts, or perform denial-of-service attacks on intervening network links. In this paper we explore the methodology behind, and construct working examples of different attack vectors on XCP, including both cheating senders and receivers. Through detailed simulations in ns, we show that misbehaving users can dominate bandwidth allocation on shared links, and our strategies allow them to successfully allocate bandwidth by either sharing or selfishly competing for the bottleneck bandwidth capacity. Christo Wilson, Chris Coakley, Ben Y. Zhao |
IWQoS | 3 |
| 2007 | QUORUM: quality of service routing in wireless mesh networksabstractWireless Mesh Networks (WMNs) can provide seamless broadband connectivity to network users, with the advantage of low setup and maintenance costs. To support next-generation applications with real-time requirements, however, these networks must provide improved Quality of Service guarantees. Most current mesh network routing protocols are adapted from MANET protocols, and do not optimize for mesh network properties. In this paper, we propose QUORUM (QUality Of service RoUting in wireless Mesh networks ), a routing protocol optimized for WMNs that addresses these drawbacks. QUORUM integrates a novel end-to-end packet delay estimation mechanism with stability-aware routing policies, allowing it to more accurately follow QoS requirements while minimizing misbehavior of selfish nodes. Vinod Kone, Sudipto Das, Ben Y. Zhao, Haitao Zheng 0001 |
QSHINE | 3 |
| 2007 | QUORUM - Quality of Service in Wireless Mesh Networks
Vinod Kone, Sudipto Das, Ben Y. Zhao, Haitao Zheng 0001 |
Mob. Networks Appl. | 3 |
| 2006 | Parallelizing Skyline Queries for Scalable Distribution
Caijie Zhang, Ben Y. Zhao, Divyakant Agrawal, Amr El Abbadi |
EDBT | 4 |
| 2006 | Understanding user behavior in large-scale video-on-demand systemsabstractVideo-on-demand over IP (VOD) is one of the best-known examples of "next-generation" Internet applications cited as a goal by networking and multimedia researchers. Without empirical data, researchers have generally relied on simulated models to drive their design and developmental efforts. In this paper, we present one of the first measurement studies of a large VOD system, using data covering 219 days and more than 150,000 users in a VOD system deployed by China Telecom. Our study focuses on user behavior, content access patterns, and their implications on the design of multimedia streaming systems. Our results also show that when used to model the user-arrival rate, the traditional Poisson model is conservative and overestimates the probability of large arrival groups. We introduce a modified Poisson distribution that more accurately models our observations. We also observe a surprising result, that video session lengths has a weak inverse correlation with the video's popularity. Finally, we gain better understanding of the sources of video popularity through analysis of a number of internal and external factors. Ben Y. Zhao |
EuroSys | 3 |
| 2006 | Determining model accuracy of network traces
Almudena Konrad, Ben Y. Zhao, Anthony D. Joseph |
J. Comput. Syst. Sci. | 2 |
| 2006 | Utilization and fairness in spectrum assignment for opportunistic spectrum access
Haitao Zheng 0001, Ben Y. Zhao |
Mob. Networks Appl. | 3 |
| 2005 | Z-Ring: Fast Prefix Routing via a Low Maintenance Membership ProtocolabstractIn this paper, we introduce Z-ring, a fast prefix routing protocol for peer-to-peer overlay networks. Z-ring incorporates cost-efficient membership protocol to achieve fast routing with small maintenance cost. Z-ring achieves routing in logGN steps, where N is the network size and G is the size of a group that can be maintained by a membership protocol with low cost. With G=4096, it translates to one-hop routing for intranet environments (N<4096), two-hop routing for mid-scale internet applications (N<16 million), and three-hop routing for ultra-large Internet applications (N<64 billion). Z-ring maintains good routing success rate under churn and low maintenance cost even at large network size. Its modularized use of the membership protocol also makes it adaptive to dynamic and wide-range network size changes. Qiao Lian, Wei Chen 0013, Zheng Zhang 0001, Shaomei Wu, Ben Y. Zhao |
ICNP | 5 |
| 2005 | Cashmere: Resilient Anonymous Routing
Li Zhuang, Ben Y. Zhao, Antony I. T. Rowstron |
NSDI | 3 |
| 2004 | Tapestry: a resilient global-scale overlay for service deploymentabstractWe present Tapestry, a peer-to-peer overlay routing infrastructure offering efficient, scalable, location-independent routing of messages directly to nearby copies of an object or service using only localized resources. Tapestry supports a generic decentralized object location and routing applications programming interface using a self-repairing, soft-state-based routing layer. The paper presents the Tapestry architecture, algorithms, and implementation. It explores the behavior of a Tapestry deployment on PlanetLab, a global testbed of approximately 100 machines. Experimental results show that Tapestry exhibits stable behavior and performance as an overlay, despite the instability of the underlying network layers. Several widely distributed applications have been implemented on Tapestry, illustrating its utility as a deployment infrastructure. Ben Y. Zhao, Ling Huang 0001, Jeremy Stribling, Sean C. Rhea, Anthony D. Joseph, John Kubiatowicz |
IEEE J. Sel. Areas Commun. | 1 |
| 2004 | Distributed Object Location in a Dynamic Network
Kirsten Hildrum, John Kubiatowicz, Satish Rao, Ben Y. Zhao |
Theory Comput. Syst. | 4 |
| 2003 | Pond: The OceanStore Prototype
Sean C. Rhea, Patrick R. Eaton, Dennis Geels, Hakim Weatherspoon, Ben Y. Zhao, John Kubiatowicz |
FAST | 5 |
| 2003 | Exploiting Routing Redundancy via Structured Peer-to-Peer OverlaysabstractStructured peer-to-peer overlays provide a natural infrastructure for resilient routing via efficient fault detection and precomputation of backup paths. These overlays can respond to faults in a few hundred milliseconds by rapidly shifting between alternate routes. In this paper, we present two adaptive mechanisms for structured overlays and illustrate their operation in the context of Tapestry, a fault-resilient overlay from Berkeley. We also describe a transparent, protocol-independent traffic redirection mechanism that tunnels legacy application traffic through overlays. Our measurements of a Tapestry prototype show it to be a highly responsive routing service, effective at circumventing a range of failures while incurring reasonable cost in maintenance bandwidth and additional routing latency. Ben Y. Zhao, Ling Huang 0001, Jeremy Stribling, Anthony D. Joseph, John Kubiatowicz |
ICNP | 1 |
| 2003 | Approximate Object Location and Spam Filtering on Peer-to-Peer Systems
Li Zhuang, Ben Y. Zhao, Ling Huang 0001, Anthony D. Joseph, John Kubiatowicz |
Middleware | 3 |
| 2003 | A Markov-Based Channel Model Algorithm for Wireless Networks
Almudena Konrad, Ben Y. Zhao, Anthony D. Joseph, Reiner Ludwig |
Wirel. Networks | 2 |
| 2002 | Distributed object location in a dynamic networkabstractModern networking applications replicate data and services widely, leading to a need for location-independent routing -- the ability to route queries directly to objects using names independent of the objects' physical locations. Two important properties of a routing infrastructure are routing locality and rapid adaptation to arriving and departing nodes. We show how these two properties can be efficiently achieved for certain network topologies. To do this, we present a new distributed algorithm that can solve the nearest-neighbor problem for these networks. We describe our solution in the context of Tapestry, an overlay network infrastructure that employs techniques proposed by Plaxton, Rajaraman, and Richa [14]. Kirsten Hildrum, John Kubiatowicz, Satish Rao, Ben Y. Zhao |
SPAA | 4 |
| 2002 | An Architecture for Secure Wide-Area Service Discovery
Todd D. Hodes, Steven E. Czerwinski, Ben Y. Zhao, Anthony D. Joseph, Randy H. Katz |
Wirel. Networks | 3 |
| 2001 | A Markov-based channel model algorithm for wireless networksabstractTechniques for modeling and simulating channel conditions play an essential role in understanding network protocol and application behavior. In [11], we demonstrated that inaccurate modeling using a traditional analytical model yielded significant errors in error control protocol parameters choices. In this paper, we demonstrate that time-varying effects on wireless channels result in wireless traces which exhibit non-stationary behavior over small window sizes. We then present an algorithm that divides traces into stationary components in order to provide analytical channel models that, relative to traditional approaches, more accurately represent characteristics such as burstiness, statistical distribution of errors, and packet loss processes. Our algorithm also generates artificial traces with the same statistical characteristics as actual collected network traces. For validation, we develop a channel model for the circuit-switched data service in GSM and show that it: (1) more closely approximates GSM channel characteristics than a traditional Gilbert model and (2) generates artificial traces that closely match collected traces' statistics. Using these traces in a simulator environment enables future protocol and application testing under different controlled and repeatable conditions. Almudena Konrad, Ben Y. Zhao, Anthony D. Joseph, Reiner Ludwig |
MSWiM | 2 |
| 2001 | Bayeux: an architecture for scalable and fault-tolerant wide-area data disseminationabstractThe demand for streaming multimedia applications is growing at an incr edible rate. In this paper, we propose Bayeux, an efficient application-level multicast system that scales to arbitrarily large receiver groups while tolerating failures in routers and network links. Bayeux also includes specific mechanisms for load-balancing across replicate root nodes and more efficient bandwidth consumption. Our simulation results indicate that Bayeux maintains these properties while keeping transmission overhead low. To achieve these properties, Bayeux leverages the architecture of Tapestry, a fault-tolerant, wide-area overlay routing and location network. Shelley Zhuang, Ben Y. Zhao, Anthony D. Joseph, Randy H. Katz, John Kubiatowicz |
NOSSDAV | 2 |
| 2001 | The Ninja architecture for robust Internet-scale systems and services
Steve D. Gribble, Matt Welsh, J. Robert von Behren, Eric A. Brewer, David E. Culler, Nikita Borisov, Steven E. Czerwinski, Ramakrishna Gummadi, Jon R. Hill, Anthony D. Joseph, Randy H. Katz, Z. Morley Mao, Steven J. Ross, Ben Y. Zhao |
Comput. Networks | 14 |
| 2000 | OceanStore: An Architecture for Global-Scale Persistent StorageabstractOceanStore is a utility infrastructure designed to span the globe and provide continuous access to persistent information. Since this infrastructure is comprised of untrusted servers, data is protected through redundancy and cryptographic techniques. To improve performance, data is allowed to be cached anywhere, anytime. Additionally, monitoring of usage patterns allows adaptation to regional outages and denial of service attacks; monitoring also enhances performance through pro-active movement of data. A prototype implementation is currently under development. John Kubiatowicz, David Bindel, Yan Chen 0004, Steven E. Czerwinski, Patrick R. Eaton, Dennis Geels, Ramakrishna Gummadi, Sean C. Rhea, Hakim Weatherspoon, Westley Weimer, Chris Wells, Ben Y. Zhao |
ASPLOS | 12 |
| 1999 | An Architecture for a Secure Service Discovery ServiceabstractThe widespread deployment of inexpensive communications technology, computational resources in the networking infrastructure, and network-enabled end devices poses an interesting problem for end users: how to locate a particular network service or device out of hundreds of thousands of accessible services and devices.This paper presents the architecture and implementation of a secure Service Discovery Service (SDS).Service providers use the SDS to advertise complex descriptions of available or already running services, while clients use the SDS to compose complex queries for locating these services.Service descriptions and queries use the extensible Markup Language (XML) to encode such factors as cost, performance, location, and device-or service-specific capabilities.The SDS provides a highlyavailable, fault-tolerant, incrementally scalable service for locating services in the wide-area.Security is a core component of the SDS and, where necessary, communications are both encrypted and authenticated.Furthermore, the SDS uses an hybrid access control list and capability system to control access to service information. Steven E. Czerwinski, Ben Y. Zhao, Todd D. Hodes, Anthony D. Joseph, Randy H. Katz |
MobiCom | 2 |