Haitao Zheng 0001

dblp:43/4261 · also Heather Zheng · DBLP profile ↗
← Back
107ranked-venue papers
0as first author
21since 2021 · last 2025
0000-0002-5918-2940ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 51 · 2 since 2021Security and privacy · 23 · 12 since 2021Databases, data management, data science and information retrieval · 14Human-computer interaction and ubiquitous computing · 12 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11Artificial intelligence and machine learning · 6 · 5 since 2021Systems, architecture and hardware · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial Mislabeling
abstract
Today's text-to-image generative models are trained on millions of images sourced from the Internet, each paired with a detailed caption produced by Vision-Language Models (VLMs). This part of the training pipeline is critical for supplying the models with large volumes of high-quality image-caption pairs during training. However, recent work suggests that VLMs are vulnerable to stealthy adversarial attacks, where adversarial perturbations are added to images to mislead the VLMs into producing incorrect captions.
Stanley Wu, Ronik Bhaskar, Anna Yoo Jeong Ha, Shawn Shan, Haitao Zheng 0001, Ben Y. Zhao
CCS5
2024 Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?
abstract
The advent of generative AI images has completely disrupted the art world. Distinguishing AI generated images from human art is a challenging problem whose impact is growing over time. A failure to address this problem allows bad actors to defraud individuals paying a premium for human art and companies whose stated policies forbid AI imagery. It is also critical for content owners to establish copyright, and for model trainers interested in curating training data in order to avoid potential model collapse. There are several different approaches to distinguishing human art from AI images, including classifiers trained by supervised learning, research tools targeting diffusion models, and identification by professional artists using their knowledge of artistic techniques. In this paper, we seek to understand how well these approaches can perform against today's modern generative models in both benign and adversarial settings. We curate real human art across 7 styles, generate matching images from 5 generative models, and apply 8 detectors (5 automated detectors and 3 different human groups including 180 crowdworkers, 3800+ professional artists, and 13 expert artists experienced at detecting AI). Both Hive and expert artists do very well, but make mistakes in different ways (Hive is weaker against adversarial perturbations while Expert artists produce higher false positives). We believe these weaknesses will persist, and argue that a combination of human and automated detectors provides the best combination of accuracy and robustness.
Anna Yoo Jeong Ha, Josephine Passananti, Ronik Bhaskar, Shawn Shan, Reid Southen, Haitao Zheng 0001, Ben Y. Zhao
CCS6
2024 Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
abstract
Trained on billions of images, diffusion-based text-to-image models seem impervious to traditional data poisoning attacks, which typically require poison samples approaching 20% of the training set. In this paper, we show that state-of-the-art text-to-image generative models are in fact highly vulnerable to poisoning attacks. Our work is driven by two key insights. First, while diffusion models are trained on billions of samples, the number of training samples associated with a specific concept or prompt is generally on the order of thousands. This suggests that these models will be vulnerable to prompt-specific poisoning attacks that corrupt a model’s ability to respond to specific targeted prompts. Second, poison samples can be carefully crafted to maximize poison potency to ensure success with very few samples.We introduce Nightshade, a prompt-specific poisoning attack optimized for potency that can completely control the output of a prompt in Stable Diffusion’s newest model (SDXL) with less than 100 poisoned training samples. Nightshade also generates stealthy poison images that look visually identical to their benign counterparts, and produces poison effects that "bleed through" to related concepts. More importantly, a moderate number of Nightshade attacks on independent prompts can destabilize a model and disable its ability to generate images for any and all prompts. Finally, we propose the use of Nightshade and similar tools as a defense for content owners against web scrapers that ignore opt-out/do-not-crawl directives, and discuss potential implications for both model trainers and content owners.
Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng 0001, Ben Y. Zhao
SP5
2024 Can Virtual Reality Protect Users from Keystroke Inference Attacks?
Zhuolin Yang 0001, Zain Sarwar, Iris Hwang, Ronik Bhaskar, Ben Y. Zhao, Haitao Zheng 0001
USENIX Security Symposium6
2023 Characterizing the Optimal 0-1 Loss for Multi-class Classification with a Test-time Attacker
abstract
Finding classifiers robust to adversarial examples is critical for their safe deployment. Determining the robustness of the best possible classifier under a given threat model for a fixed data distribution and comparing it to that achieved by state-of-the-art training methods is thus an important diagnostic tool. In this paper, we find achievable information-theoretic lower bounds on robust loss in the presence of a test-time attacker for *multi-class classifiers on any discrete dataset*. We provide a general framework for finding the optimal $0-1$ loss that revolves around the construction of a conflict hypergraph from the data and adversarial constraints. The prohibitive cost of this formulation in practice leads us to formulate other variants of the attacker-classifier game that more efficiently determine the range of the optimal loss. Our valuation shows, for the first time, an analysis of the gap to optimal robustness for classifiers in the multi-class setting on benchmark datasets.
Sihui Dai, Wenxin Ding, Arjun Nitin Bhagoji, Daniel Cullina, Haitao Zheng 0001, Ben Zhao, Prateek Mittal
NeurIPS5
2023 SoK: Anti-Facial Recognition Technology
abstract
The rapid adoption of facial recognition (FR) technology by both government and commercial entities in recent years has raised concerns about civil liberties and privacy. In response, a broad suite of so-called "anti-facial recognition" (AFR) tools has been developed to help users avoid unwanted facial recognition. The set of AFR tools proposed in the last few years is wide-ranging and rapidly evolving, necessitating a step back to consider the broader design space of AFR systems and long-term challenges. This paper aims to fill that gap and provides the first comprehensive analysis of the AFR research landscape. Using the operational stages of FR systems as a starting point, we create a systematic framework for analyzing the benefits and tradeoffs of different AFR approaches. We then consider both technical and social challenges facing AFR tools and propose directions for future research in this field.
Emily Wenger, Shawn Shan, Haitao Zheng 0001, Ben Y. Zhao
SP3
2023 Glaze: Protecting Artists from Style Mimicry by Text-to-Image Models
Shawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng 0001, Rana Hanocka, Ben Y. Zhao
USENIX Security Symposium4
2023 Towards a General Video-based Keystroke Inference Attack
Zhuolin Yang 0001, Yuxin Chen 0001, Zain Sarwar, Hadleigh Schwartz, Ben Y. Zhao, Haitao Zheng 0001
USENIX Security Symposium6
2023 "My face, my rules": Enabling Personalized Protection Against Unacceptable Face Editing
abstract
Today, face editing is widely used to refine/alter photos in both professional and recreational settings. Yet it is also used to modify (and repost) existing online photos for cyberbullying. Our work considers an important open question: 'How can we support the collaborative use of face editing on social platforms while protecting against unacceptable edits and reposts by others?' This is challenging because, as our user study shows, users vary widely in their definition of what edits are (un)acceptable. Any global filter policy deployed by social platforms is unlikely to address the needs of all users, but hinders social interactions enabled by photo editing. Instead, we argue that face edit protection policies should be implemented by social platforms based on individual user preferences. When posting an original photo online, a user can choose to specify the types of face edits (dis)allowed on the photo. Social platforms use these per-photo edit policies to moderate future photo uploads, i.e., edited photos containing modifications that violate the original photo's policy are either blocked or shelved for user approval. Realizing this personalized protection, however, faces two immediate challenges: (1) how to accurately recognize specific modifications, if any, contained in a photo; and (2) how to associate an edited photo with its original photo (and thus the edit policy). We show that these challenges can be addressed by combining highly efficient hashing based image search and scalable semantic image comparison, and build a prototype protector (Alethia) covering nine edit types. Evaluations using IRB-approved user studies and data-driven experiments (on 839K face photos) show that Alethia accurately recognizes edited photos that violate user policies and induces a feeling of protection to study participants. This demonstrates the initial feasibility of personalized face edit protection. We also discuss current limitations and future directions to push the concept forward.
Zhujun Xiao, Jenna Cryan, Yuanshun Yao, Yi Hong Gordon Cheo, Yuanchao Shu, Stefan Saroiu, Ben Y. Zhao, Haitao Zheng 0001
Proc. Priv. Enhancing Technol.8
2023 Trimming Mobile Applications for Bandwidth-Challenged Networks in Developing Regions
abstract
Despite continuous efforts to build and update mobile network infrastructure, mobile devices in developing regions continue to be constrained by limited bandwidth. Unfortunately, this coincides with a period of unprecedented growth in the sizes of mobile applications. Thus it is becoming prohibitively expensive for users in developing regions to download and update mobile apps critical to their economic and educational development. Unchecked, these trends can further contribute to a large and growing global digital divide. Our goal is to better understand the source of this rapid growth in mobile app code size, whether it is reflective of new functionality, and identify steps that can be taken to make existing mobile apps more friendly to bandwidth constrained mobile networks. We hypothesize that much of this growth in mobile apps is due to poor resource/code management, and do not reflect proportional increases in functionality. Our hypothesis is partially validated by mini-programs, apps with extremely small footprints gaining popularity in Chinese mobile platforms. Here, we use functionally equivalent pairs of mini-programs and Android apps to identify potential sources of “bloat,” i.e., inefficient uses of code or resources that contribute to large package sizes. We analyze a large sample of popular Android apps and quantify instances of code and resource bloat. We develop techniques for automated code and resource trimming, and successfully validate them on a large set of Android apps. We hope our results will lead to continued efforts to streamline mobile apps, making them easier to access and maintain for users in developing regions.
Qinge Xie, Qingyuan Gong, Xinlei He 0001, Yang Chen 0001, Xin Wang 0002, Haitao Zheng 0001, Ben Y. Zhao
IEEE Trans. Mob. Comput.6
2022 Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN Models
abstract
Server breaches are an unfortunate reality on today's Internet. In the context of deep neural network (DNN) models, they are particularly harmful, because a leaked model gives an attacker "white-box'' access to generate adversarial examples, a threat model that has no practical robust defenses. For practitioners who have invested years and millions into proprietary DNNs, e.g. medical imaging, this seems like an inevitable disaster looming on the horizon.
Shawn Shan, Wenxin Ding, Emily Wenger, Haitao Zheng 0001, Ben Y. Zhao
CCS4
2022 Understanding Robust Learning through the Lens of Representation Similarities
abstract
Representation learning, \textit{i.e.} the generation of representations useful for downstream applications, is a task of fundamental importance that underlies much of the success of deep neural networks (DNNs). Recently, \emph{robustness to adversarial examples} has emerged as a desirable property for DNNs, spurring the development of robust training methods that account for adversarialexamples. In this paper, we aim to understand how the properties of representations learned by robust training differ from those obtained from standard, non-robust training. This is critical to diagnosing numerous salient pitfalls in robust networks, such as, degradation of performance on benign inputs, poor generalization of robustness, and increase in over-fitting. We utilize a powerful set of tools known as representation similarity metrics, across 3 vision datasets, to obtain layer-wise comparisons between robust and non-robust DNNs with different architectures, training procedures and adversarial constraints. Our experiments highlight hitherto unseen properties of robust representations that we posit underlie the behavioral differences of robust networks. We discover a lack of specialization in robust networks' representations along with a disappearance of `block structure'. We also find overfitting during robust training largely impacts deeper layers. These, along with other findings, suggest ways forward for the design and training of better robust networks.
Christian Cianfarani, Arjun Nitin Bhagoji, Vikash Sehwag, Ben Y. Zhao, Haitao Zheng 0001, Prateek Mittal
NeurIPS5
2022 Finding Naturally Occurring Physical Backdoors in Image Datasets
abstract
Extensive literature on backdoor poison attacks has studied attacks and defenses for backdoors using “digital trigger patterns.” In contrast, “physical backdoors” use physical objects as triggers, have only recently been identified, and are qualitatively different enough to resist most defenses targeting digital trigger backdoors. Research on physical backdoors is limited by access to large datasets containing real images of physical objects co-located with misclassification targets. Building these datasets is time- and labor-intensive.This work seeks to address the challenge of accessibility for research on physical backdoor attacks. We hypothesize that there may be naturally occurring physically co-located objects already present in popular datasets such as ImageNet. Once identified, a careful relabeling of these data can transform them into training samples for physical backdoor attacks. We propose a method to scalably identify these subsets of potential triggers in existing datasets, along with the specific classes they can poison. We call these naturally occurring trigger-class subsets natural backdoor datasets. Our techniques successfully identify natural backdoors in widely-available datasets, and produce models behaviorally equivalent to those trained on manually curated datasets. We release our code to allow the research community to create their own datasets for research on physical backdoor attacks.
Emily Wenger, Roma Bhattacharjee, Arjun Nitin Bhagoji, Josephine Passananti, Emilio Andere, Haitao Zheng 0001, Ben Y. Zhao
NeurIPS6
2022 Blacklight: Scalable Defense for Neural Networks against Query-Based Black-Box Attacks
Huiying Li 0001, Shawn Shan, Emily Wenger, Jiayun Zhang, Haitao Zheng 0001, Ben Y. Zhao
USENIX Security Symposium5
2022 Poison Forensics: Traceback of Data Poisoning Attacks in Neural Networks
Shawn Shan, Arjun Nitin Bhagoji, Haitao Zheng 0001, Ben Y. Zhao
USENIX Security Symposium3
2021 "Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World
abstract
Advances in deep learning have introduced a new wave of voice synthesis tools, capable of producing audio that sounds as if spoken by a target speaker. If successful, such tools in the wrong hands will enable a range of powerful attacks against both humans and software systems (aka machines). This paper documents efforts and findings from a comprehensive experimental study on the impact of deep-learning based speech synthesis attacks on both human listeners and machines such as speaker recognition and voice-signin systems. We find that both humans and machines can be reliably fooled by synthetic speech, and that existing defenses against synthesized speech fall short. These findings highlight the need to raise awareness and develop new protections against synthetic speech for both humans and machines.
Emily Wenger, Max Bronckers, Christian Cianfarani, Jenna Cryan, Angela Sha, Haitao Zheng 0001, Ben Y. Zhao
CCS6
2021 User Authentication via Electrical Muscle Stimulation
abstract
We propose a novel modality for active biometric authentication: electrical muscle stimulation (EMS). To explore this, we engineered an interactive system, which we call ElectricAuth, that stimulates the user’s forearm muscles with a sequence of electrical impulses (i.e., EMS challenge) and measures the user’s involuntary finger movements (i.e., response to the challenge). ElectricAuth leverages EMS’s intersubject variability, where the same electrical stimulation results in different movements in different users because everybody’s physiology is unique (e.g., differences in bone and muscular structure, skin resistance and composition, etc.). As such, ElectricAuth allows users to login without memorizing passwords or PINs.
Yuxin Chen 0001, Zhuolin Yang 0001, Ruben Abbou, Pedro Lopes 0001, Ben Y. Zhao, Haitao Zheng 0001
CHI6
2021 Backdoor Attacks Against Deep Learning Systems in the Physical World
abstract
Backdoor attacks embed hidden malicious behaviors into deep learning models, which only activate and cause misclassifications on model inputs containing a specific "trigger." Existing works on backdoor attacks and defenses, however, mostly focus on digital attacks that apply digitally generated patterns as triggers. A critical question remains unanswered: "can backdoor attacks succeed using physical objects as triggers, thus making them a credible threat against deep learning systems in the real world?"We conduct a detailed empirical study to explore this question for facial recognition, a critical deep learning task. Using 7 physical objects as triggers, we collect a custom dataset of 3205 images of 10 volunteers and use it to study the feasibility of "physical" backdoor attacks under a variety of real-world conditions. Our study reveals two key findings. First, physical backdoor attacks can be highly successful if they are carefully configured to overcome the constraints imposed by physical objects. In particular, the placement of successful triggers is largely constrained by the target model’s dependence on key facial features. Second, four of today’s state-of-the-art defenses against (digital) backdoors are ineffective against physical backdoors, because the use of physical objects breaks core assumptions used to construct these defenses.Our study confirms that (physical) backdoor attacks are not a hypothetical phenomenon but rather pose a serious real-world threat to critical classification tasks. We need new and more robust defenses against backdoors in the physical world.
Emily Wenger, Josephine Passananti, Arjun Nitin Bhagoji, Yuanshun Yao, Haitao Zheng 0001, Ben Y. Zhao
CVPR5
2021 Towards Performance Clarity of Edge Video Analytics
Zhujun Xiao, Zhengxu Xia, Haitao Zheng 0001, Ben Y. Zhao, Junchen Jiang
SEC3
2021 Understanding the Effect of Bias in Deep Anomaly Detection
abstract
Anomaly detection presents a unique challenge in machine learning, due to the scarcity of labeled anomaly data. Recent work attempts to mitigate such problems by augmenting training of deep anomaly detection models with additional labeled anomaly samples. However, the labeled data often does not align with the target distribution and introduces harmful bias to the trained model. In this paper, we aim to understand the effect of a biased anomaly set on anomaly detection. Concretely, we view anomaly detection as a supervised learning task where the objective is to optimize the recall at a given false positive rate. We formally study the relative scoring bias of an anomaly detector, defined as the difference in performance with respect to a baseline anomaly detector. We establish the first finite sample rates for estimating the relative scoring bias for deep anomaly detection, and empirically validate our theoretical results on both synthetic and real-world datasets. We also provide an extensive empirical study on how a biased training anomaly set affects the anomaly score function and therefore the detection performance on different anomaly classes. Our study demonstrates scenarios in which the biased anomaly set can be useful or problematic, and provides a solid benchmark for future research.
Ziyu Ye, Yuxin Chen 0001, Haitao Zheng 0001
IJCAI3
2021 On Migratory Behavior in Video Consumption
abstract
Today’s video streaming market is crowded with various content providers (CPs). For individual CPs, understanding user behavior, in particular how users migrate among different CPs, is crucial for improving users’ on-site experience and the CP’s chance of success. In this article, we take a data-driven approach to analyze and model user migration behavior in video streaming, i.e., users switching content provider during active sessions. Based on a large ISP dataset over two months (6 major content providers, 3.8 million users, and 315 million video requests), we study common migration patterns and reasons of migration. We find that migratory behavior is prevalent: 66% of users switch CPs with an average switching frequency of 13%. In addition, migration behaviors are highly diverse: regardless large or small CPs, they all have dedicated groups of users who like to switch to them for certain types of videos. Regarding reasons of migration, we find CP service quality rarely causes migration, while a few popular videos play a bigger role. Nearly 60% of cross-site migrations are landed to 0.14% top videos. Finally, we validate our findings by building an accurate regression model to predict user migration frequency as well as user survey, and discuss the implications of our results to CPs.
Huan Yan 0003, Haohao Fu, Yong Li 0008, Tzu-Heng Lin, Gang Wang 0011, Haitao Zheng 0001, Depeng Jin, Ben Y. Zhao
IEEE Trans. Netw. Serv. Manag.6
2020 Gotta Catch'Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks
abstract
Deep neural networks (DNN) are known to be vulnerable to adversarial attacks. Numerous efforts either try to patch weaknesses in trained models, or try to make it difficult or costly to compute adversarial examples that exploit them. In our work, we explore a new "honeypot" approach to protect DNN models. We intentionally inject trapdoors, honeypot weaknesses in the classification manifold that attract attackers searching for adversarial examples. Attackers' optimization algorithms gravitate towards trapdoors, leading them to produce attacks similar to trapdoors in the feature space. Our defense then identifies attacks by comparing neuron activation signatures of inputs to those of trapdoors.
Shawn Shan, Emily Wenger, Bolun Wang, Bo Li 0026, Haitao Zheng 0001, Ben Y. Zhao
CCS5
2020 Wearable Microphone Jamming
abstract
We engineered a wearable microphone jammer that is capable of disabling microphones in its user's surroundings, including hidden microphones. Our device is based on a recent exploit that leverages the fact that when exposed to ultrasonic noise, commodity microphones will leak the noise into the audible range.
Yuxin Chen 0001, Huiying Li 0001, Shan-Yuan Teng, Steven Nagels, Zhijing Li 0001, Pedro Lopes 0001, Ben Y. Zhao, Haitao Zheng 0001
CHI8
2020 Detecting Gender Stereotypes: Lexicon vs. Supervised Learning Methods
abstract
Biases in language influence how we interact with each other and society at large. Language affirming gender stereotypes is often observed in various contexts today, from recommendation letters and Wikipedia entries to fiction novels and movie dialogue. Yet to date, there is little agreement on the methodology to quantify gender stereotypes in natural language (specifically the English language). Common methodology (including those adopted by companies tasked with detecting gender bias) rely on a lexicon approach largely based on the original BSRI study from 1974.
Jenna Cryan, Shiliang Tang, Xinyi Zhang 0003, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao
CHI5
2020 Et Tu Alexa? When Commodity WiFi Devices Turn into Adversarial Motion Sensors
Yanzi Zhu, Zhujun Xiao, Yuxin Chen 0001, Zhijing Li 0001, Max Liu, Ben Y. Zhao, Haitao Zheng 0001
NDSS7
2020 Fawkes: Protecting Privacy against Unauthorized Deep Learning Models
Shawn Shan, Emily Wenger, Jiayun Zhang, Huiying Li 0001, Haitao Zheng 0001, Ben Y. Zhao
USENIX Security Symposium5
2019 Latent Backdoor Attacks on Deep Neural Networks
abstract
Recent work proposed the concept of backdoor attacks on deep neural networks (DNNs), where misclassification rules are hidden inside normal models, only to be triggered by very specific inputs. However, these "traditional" backdoors assume a context where users train their own models from scratch, which rarely occurs in practice. Instead, users typically customize "Teacher" models already pretrained by providers like Google, through a process called transfer learning. This customization process introduces significant changes to models and disrupts hidden backdoors, greatly reducing the actual impact of backdoors in practice. In this paper, we describe latent backdoors, a more powerful and stealthy variant of backdoor attacks that functions under transfer learning. Latent backdoors are incomplete backdoors embedded into a "Teacher" model, and automatically inherited by multiple "Student" models through transfer learning. If any Student models include the label targeted by the backdoor, then its customization process completes the backdoor and makes it active. We show that latent backdoors can be quite effective in a variety of application contexts, and validate its practicality through real-world attacks against traffic sign recognition, iris identification of volunteers, and facial recognition of public figures (politicians). Finally, we evaluate 4 potential defenses, and find that only one is effective in disrupting latent backdoors, but might incur a cost in classification accuracy as tradeoff.
Yuanshun Yao, Huiying Li 0001, Haitao Zheng 0001, Ben Y. Zhao
CCS3
2019 Scaling Deep Learning Models for Spectrum Anomaly Detection
abstract
Spectrum management in cellular networks is a challenging task that will only increase in difficulty as complexity grows in hardware, configurations, and new access technology (e.g. LTE for IoT devices). Wireless providers need robust and flexible tools to monitor and detect faults and misbehavior in physical spectrum usage, and to deploy them at scale. In this paper, we explore the design of such a system by building deep neural network (DNN) models1 to capture spectrum usage patterns and use them as baselines to detect spectrum usage anomalies resulting from faults and misuse. Using detailed LTE spectrum measurements, we show that the key challenge facing this design is model scalability, i.e. how to train and deploy DNN models at a large number of static and mobile observers located throughout the network. We address this challenge by building context-agnostic models for spectrum usage and applying transfer learning to minimize training time and dataset constraints. The end result is a practical DNN model that can be easily deployed on both mobile and static observers, enabling timely detection of spectrum anomalies across LTE networks.
Zhijing Li 0001, Zhujun Xiao, Bolun Wang, Ben Y. Zhao, Haitao Zheng 0001
MobiHoc5
2019 Safely and automatically updating in-network ACL configurations with intent language
abstract
In-network Access Control List (ACL) is an important technique in ensuring network-wide connectivity and security. As cloud-scale WANs today constantly evolve in size and complexity, in-network ACL rules are becoming increasingly more complex. This presents a great challenge to the updating process of ACL configurations: network operators are frequently required to update "tangled" ACL rules across thousands of devices to meet diverse business requirements, and even a single ACL misconfiguration may lead to network disruptions. Such increasing challenges call for an automated system to improve the efficiency and correctness of ACL updates. This paper presents Jinjing, a system that aids Alibaba's network operators in automatically and correctly updating ACL configurations in Alibaba's global WAN. Jinjing allows the operators to express in a declarative language, named LAI, their update intent (e.g., ACL migration and traffic control). Then, Jinjing automatically synthesizes ACL update plans that satisfy their intent. At the heart of Jinjing, we develop a set of novel verification and synthesis techniques to rigorously guarantee the correctness of update plans. In Alibaba, our operators have used Jinjing to efficiently update their ACLs and have thus prevented significant service downtime.
Bingchuan Tian, Xinyi Zhang 0003, Ennan Zhai, Hongqiang Harry Liu, Qiaobo Ye, Chunsheng Wang, Zhiming Ji, Yihong Sang, Ming Zhang 0005, Chen Tian 0001, Haitao Zheng 0001, Ben Y. Zhao
SIGCOMM13
2019 Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks
abstract
Lack of transparency in deep neural networks (DNNs) make them susceptible to backdoor attacks, where hidden associations or triggers override normal classification to produce unexpected results. For example, a model with a backdoor always identifies a face as Bill Gates if a specific symbol is present in the input. Backdoors can stay hidden indefinitely until activated by an input, and present a serious security risk to many security or safety related applications, e.g. biometric authentication systems or self-driving cars. We present the first robust and generalizable detection and mitigation system for DNN backdoor attacks. Our techniques identify backdoors and reconstruct possible triggers. We identify multiple mitigation techniques via input filters, neuron pruning and unlearning. We demonstrate their efficacy via extensive experiments on a variety of DNNs, against two types of backdoor injection methods identified by prior work. Our techniques also prove robust against a number of variants of the backdoor attack.
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 0001, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao
IEEE Symposium on Security and Privacy6
2018 Predictive Analysis in Network Function Virtualization
Zhijing Li 0001, Zihui Ge, Ajay Mahimkar, Jia Wang 0001, Ben Y. Zhao, Haitao Zheng 0001, Joanne Emmons, Laura Ogden
Internet Measurement Conference6
2018 With Great Training Comes Great Vulnerability: Practical Attacks against Transfer Learning
Bolun Wang, Yuanshun Yao, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao
USENIX Security Symposium4
2018 Ghost Riders: Sybil Attacks on Crowdsourced Mobile Mapping Services
Gang Wang 0011, Bolun Wang, Tianyi Wang 0001, Ana Nika, Haitao Zheng 0001, Ben Y. Zhao
IEEE/ACM Trans. Netw.5
2017 Automated Crowdturfing Attacks and Defenses in Online Review Systems
abstract
Malicious crowdsourcing forums are gaining traction as sources of spreading misinformation online, but are limited by the costs of hiring and managing human workers. In this paper, we identify a new class of attacks that leverage deep learning language models (Recurrent Neural Networks or RNNs) to automate the generation of fake online reviews for products and services. Not only are these attacks cheap and therefore more scalable, but they can control rate of content output to eliminate the signature burstiness that makes crowdsourced campaigns easy to detect.
Yuanshun Yao, Bimal Viswanath, Jenna Cryan, Haitao Zheng 0001, Ben Y. Zhao
CCS4
2017 On Migratory Behavior in Video Consumption
abstract
Today's video streaming market is crowded with various content providers (CPs). For individual CPs, understanding user behavior, in particular how users migrate among different CPs, is crucial for improving users' on-site experience and the CP's chance of success. In this paper, we take a data-driven approach to analyze and model user migration behavior in video streaming, i.e., users switching content provider during active sessions. Based on a large ISP dataset over two months (6 major content providers, 3.8 million users, and 315 million video requests), we study common migration patterns and reasons of migration. We find that migratory behavior is prevalent: 66% of users switch CPs with an average switching frequency of 13%. In addition, migration behaviors are highly diverse: regardless large or small CPs, they all have dedicated groups of users who like to switch to them for certain types of videos. Regarding reasons of migration, we find CP service quality rarely causes migration, while a few popular videos play a bigger role. Nearly 60% of cross-site migrations are landed to 0.14% top videos. Finally, we validate our findings by building an accurate regression model to predict user migration frequency, and discuss the implications of our results to CPs.
Huan Yan 0003, Tzu-Heng Lin, Gang Wang 0011, Yong Li 0008, Haitao Zheng 0001, Depeng Jin, Ben Y. Zhao
CIKM5
2017 Echo Chambers in Investment Discussion Boards
Shiliang Tang, Qingyun Liu 0003, Megan McQueen, Scott Counts, Apurv Jain, Haitao Zheng 0001, Ben Y. Zhao
ICWSM6
2017 A First Look at User Switching Behaviors Over Multiple Video Content Providers
Huan Yan 0003, Tzu-Heng Lin, Gang Wang 0011, Yong Li 0008, Haitao Zheng 0001, Depeng Jin, Ben Y. Zhao
ICWSM5
2017 Cold Hard E-Cash: Friends and Vendors in the Venmo Digital Payments System
Xinyi Zhang 0003, Shiliang Tang, Yun Zhao 0001, Gang Wang 0011, Haitao Zheng 0001, Ben Y. Zhao
ICWSM5
2017 Complexity vs. performance: empirical analysis of machine learning as a service
abstract
Machine learning classifiers are basic research tools used in numerous types of network analysis and modeling. To reduce the need for domain expertise and costs of running local ML classifiers, network researchers can instead rely on centralized Machine Learning as a Service (MLaaS) platforms.
Yuanshun Yao, Zhujun Xiao, Bolun Wang, Bimal Viswanath, Haitao Zheng 0001, Ben Y. Zhao
Internet Measurement Conference5
2017 Object Recognition and Navigation using a Single Networking Device
abstract
Tomorrow's autonomous mobile devices need accurate, robust and real-time sensing of their operating environment. Today's solutions fall short. Vision or acoustic-based techniques are vulnerable against challenging lighting conditions or background noise, while more robust laser or RF solutions require either bulky expensive hardware or tight coordination between multiple devices. This paper describes the design, implementation and evaluation of Ulysses, a practical environmental imaging system using colocated 60GHz radios on a single mobile device. Unlike alternatives that require specialized hardware, Ulysses reuses low-cost commodity networking chipsets available today. Ulysses' new imaging approach leverages RF beamforming, operates on specular (direct) reflection, and integrates the device's movement trajectory with sensing. Ulysses also includes a navigation component that uses the same 60GHz radios to compute "safety regions" where devices can move freely without collision, and to compute optimal paths for imaging within safety regions. Using our implementation of a small robotic car prototype, our experimental results show that Ulysses images objects meters away with cm-level precision, and provides accurate estimates of objects' surface materials.
Yanzi Zhu, Yuanshun Yao, Ben Y. Zhao, Haitao Zheng 0001
MobiSys4
2017 Identifying Value in Crowdsourced Wireless Signal Measurements
abstract
While crowdsourcing is an attractive approach to collect large-scale wireless measurements, understanding the quality and variance of the resulting data is difficult. Our work analyzes the quality of crowdsourced cellular signal measurements in the context of basestation localization, using large international public datasets (419M signal measurements and 1M cells) and corresponding ground truth values. Performing localization using raw received signal strength (RSS) data produces poor results and very high variance. Applying supervised learning improves results moderately, but variance remains high. Instead, we propose feature clustering, a novel application of unsupervised learning to detect hidden correlation between measurement instances, their features, and localization accuracy. Our results identify RSS standard deviation and RSS-weighted dispersion mean as key features that correlate with highly predictive measurement samples for both sparse and dense measurements respectively. Finally, we show how optimizing crowdsourcing measurements for these two features dramatically improves localization accuracy and reduces variance.
Zhijing Li 0001, Ana Nika, Xinyi Zhang 0003, Yanzi Zhu, Yuanshun Yao, Ben Y. Zhao, Haitao Zheng 0001
WWW7
2017 Gender Bias in the Job Market: A Longitudinal Analysis
abstract
For millions of workers, online job listings provide the first point of contact to potential employers. As a result, job listings and their word choices can significantly affect the makeup of the responding applicant pool. Here, we study the effects of potentially gender-biased terminology in job listings, and their impact on job applicants, using a large historical corpus of 17 million listings on LinkedIn spanning 10 years. We develop algorithms to detect and quantify gender bias, validate them using external tools, and use them to quantify job listing bias over time. We then perform a user survey over two user populations (N 1=469 , N 2=273 ) to validate our findings and to quantify the end-to-end impact of such bias on applicant decisions. Our findings show gender-bias has decreased significantly over the last 10 years. More surprisingly, we find that impact of gender bias in listings is dwarfed by our respondents' inherent bias towards specific job types.
Shiliang Tang, Xinyi Zhang 0003, Jenna Cryan, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao
Proc. ACM Hum. Comput. Interact.5
2017 Value and Misinformation in Collaborative Investing Platforms
abstract
It is often difficult to separate the highly capable “experts” from the average worker in crowdsourced systems. This is especially true for challenge application domains that require extensive domain knowledge. The problem of stock analysis is one such domain, where even the highly paid, well-educated domain experts are prone to make mistakes. As an extremely challenging problem space, the “wisdom of the crowds” property that many crowdsourced applications rely on may not hold. In this article, we study the problem of evaluating and identifying experts in the context of SeekingAlpha and StockTwits, two crowdsourced investment services that have recently begun to encroach on a space dominated for decades by large investment banks. We seek to understand the quality and impact of content on collaborative investment platforms, by empirically analyzing complete datasets of SeekingAlpha articles (9 years) and StockTwits messages (4 years). We develop sentiment analysis tools and correlate contributed content to the historical performance of relevant stocks. While SeekingAlpha articles and StockTwits messages provide minimal correlation to stock performance in aggregate, a subset of experts contribute more valuable (predictive) content. We show that these authors can be easily identified by user interactions, and investments based on their analysis significantly outperform broader markets. This effectively shows that even in challenging application domains, there is a secondary or indirect wisdom of the crowds. Finally, we conduct a user survey that sheds light on users’ views of SeekingAlpha content and stock manipulation. We also devote efforts to identify potential manipulation of stocks by detecting authors controlling multiple identities.
Tianyi Wang 0001, Gang Wang 0011, Bolun Wang, Divya Sambasivan, Zengbin Zhang, Xing Li 0001, Haitao Zheng 0001, Ben Y. Zhao
ACM Trans. Web7
2017 Clickstream User Behavior Models
abstract
The next generation of Internet services is driven by users and user-generated content. The complex nature of user behavior makes it highly challenging to manage and secure online services. On one hand, service providers cannot effectively prevent attackers from creating large numbers of fake identities to disseminate unwanted content (e.g., spam). On the other hand, abusive behavior from real users also poses significant threats (e.g., cyberbullying). In this article, we propose clickstream models to characterize user behavior in large online services. By analyzing clickstream traces (i.e., sequences of click events from users), we seek to achieve two goals: (1) detection: to capture distinct user groups for the detection of malicious accounts, and (2) understanding: to extract semantic information from user groups to understand the captured behavior. To achieve these goals, we build two related systems. The first one is a semisupervised system to detect malicious user accounts (Sybils). The core idea is to build a clickstream similarity graph where each node is a user and an edge captures the similarity of two users’ clickstreams. Based on this graph, we propose a coloring scheme to identify groups of malicious accounts without relying on a large labeled dataset. We validate the system using ground-truth clickstream traces of 16,000 real and Sybil users from Renren, a large Chinese social network. The second system is an unsupervised system that aims to capture and understand the fine-grained user behavior. Instead of binary classification (malicious or benign), this model identifies the natural groups of user behavior and automatically extracts features to interpret their semantic meanings. Applying this system to Renren and another online social network, Whisper (100K users), we help service providers identify unexpected user behaviors and even predict users’ future actions. Both systems received positive feedback from our industrial collaborators including Renren, LinkedIn, and Whisper after testing on their internal clickstream data.
Gang Wang 0011, Xinyi Zhang 0003, Shiliang Tang, Christo Wilson, Haitao Zheng 0001, Ben Y. Zhao
ACM Trans. Web5
2016 Unsupervised Clickstream Clustering for User Behavior Analysis
abstract
Online services are increasingly dependent on user participation. Whether it's online social networks or crowdsourcing services, understanding user behavior is important yet challenging. In this paper, we build an unsupervised system to capture dominating user behaviors from clickstream data (traces of users' click events), and visualize the detected behaviors in an intuitive manner. Our system identifies "clusters" of similar users by partitioning a similarity graph (nodes are users; edges are weighted by clickstream similarity). The partitioning process leverages iterative feature pruning to capture the natural hierarchy within user clusters and produce intuitive features for visualizing and understanding captured user behaviors. For evaluation, we present case studies on two large-scale clickstream traces (142 million events) from real social networks. Our system effectively identifies previously unknown behaviors, e.g., dormant users, hostile chatters. Also, our user study shows people can easily interpret identified behaviors using our visualization tool.
Gang Wang 0011, Xinyi Zhang 0003, Shiliang Tang, Haitao Zheng 0001, Ben Y. Zhao
CHI4
2016 Trimming the Smartphone Network Stack
abstract
Network transmissions are the cornerstone of most mobile apps today, and a main contributor to energy consumption. We use a componentized energy model to quantify energy use by device, and observe significant energy consumption by the CPU in network operations. We assert that optimizing network operations in the CPU can produce significant energy savings, and explore the impact of two potential approaches: one-copy data moves and offloading the network stack to the basestation.
Yanzi Zhu, Yibo Zhu 0001, Ana Nika, Ben Y. Zhao, Haitao Zheng 0001
HotNets5
2016 "Will Check-in for Badges": Understanding Bias and Misbehavior on Location-Based Social Networks
Gang Wang 0011, Sarita Yardi Schoenebeck, Haitao Zheng 0001, Ben Y. Zhao
ICWSM3
2016 Network Growth and Link Prediction Through an Empirical Lens
Qingyun Liu 0003, Shiliang Tang, Xinyi Zhang 0003, Xiaohan Zhao, Ben Y. Zhao, Haitao Zheng 0001
Internet Measurement Conference6
2016 Anatomy of a Personalized Livestreaming System
Bolun Wang, Xinyi Zhang 0003, Gang Wang 0011, Haitao Zheng 0001, Ben Y. Zhao
Internet Measurement Conference4
2016 Defending against Sybil Devices in Crowdsourced Mapping Services
abstract
Real-time crowdsourced maps such as Waze provide timely updates on traffic, congestion, accidents and points of interest. In this paper, we demonstrate how lack of strong location authentication allows creation of software-based Sybil devices that expose crowdsourced map systems to a variety of security and privacy attacks. Our experiments show that a single Sybil device with limited resources can cause havoc on Waze, reporting false congestion and accidents and automatically rerouting user traffic. More importantly, we describe techniques to generate Sybil devices at scale, creating armies of virtual vehicles capable of remotely tracking precise movements for large user populations while avoiding detection. We propose a new approach to defend against Sybil devices based on co-location edges, authenticated records that attest to the one-time physical co-location of a pair of devices. Over time, co-location edges combine to form large proximity graphs that attest to physical interactions between devices, allowing scalable detection of virtual vehicles. We demonstrate the efficacy of this approach using large-scale simulations, and discuss how they can be used to dramatically reduce the impact of attacks against crowdsourced mapping services.
Gang Wang 0011, Bolun Wang, Tianyi Wang 0001, Ana Nika, Haitao Zheng 0001, Ben Y. Zhao
MobiSys5
2016 Empirical Validation of Commodity Spectrum Monitoring
abstract
We describe our efforts to empirically validate a distributed spectrum monitoring system built on commodity smartphones and embedded low-cost spectrum sensors. This system enables real-time spectrum sensing, identifies and locates active transmitters, and generates alarm events when detecting anomalous transmitters. To evaluate the feasibility of such a platform, we perform detailed experiments using a prototype hardware platform using smartphones and RTL dongles. We identify multiple sources of error in the sensing results and the end-user overhead (i.e. smartphone energy draw). We propose and implement a variety of techniques to identify and overcome errors and uncertainty in the data, and to reduce energy consumption. Our work demonstrates the basic viability of user-driven spectrum monitoring on commodity devices.
Ana Nika, Zhijing Li 0001, Yanzi Zhu, Yibo Zhu 0001, Ben Y. Zhao, Haitao Zheng 0001
SenSys7
2016 The power of comments: fostering social interactions in microblog networks
Tianyi Wang 0001, Yang Chen 0001, Bolun Wang, Gang Wang 0011, Xing Li 0001, Haitao Zheng 0001, Ben Y. Zhao
Frontiers Comput. Sci.7
2016 Understanding and Predicting Data Hotspots in Cellular Networks
Ana Nika, Asad Ismail, Ben Y. Zhao, Sabrina Gaito, Gian Paolo Rossi 0001, Haitao Zheng 0001
Mob. Networks Appl.6
2015 Crowds on Wall Street: Extracting Value from Collaborative Investing Platforms
abstract
In crowdsourced systems, it is often difficult to separate the highly capable "experts" from the average worker. In this paper, we study the problem of evaluating and identifying experts in the context of SeekingAlpha and StockTwits, two crowdsourced investment services that are encroaching on a space dominated for decades by large investment banks. We seek to understand the quality and impact of content on collaborative investment platforms, by empirically analyzing complete datasets of SeekingAlpha articles (9 years) and StockTwits messages (4 years). We develop sentiment analysis tools and correlate contributed content to the historical performance of relevant stocks. While SeekingAlpha articles and StockTwits messages provide minimal correlation to stock performance in aggregate, a subset of experts contribute more valuable (predictive) content. We show that these authors can be easily identified by user interactions, and investments using their analysis significantly outperform broader markets. Finally, we conduct a user survey that sheds light on users views of SeekingAlpha content and stock manipulation.
Gang Wang 0011, Tianyi Wang 0001, Bolun Wang, Divya Sambasivan, Zengbin Zhang, Haitao Zheng 0001, Ben Y. Zhao
CSCW6
2015 Interference Analysis for mm-Wave Picocells
abstract
Millimeter (mm) wave picocellular networks are a promising approach for delivering the 1000-fold capacity increase required to keep up with projected demand for wireless data: the available bandwidth is orders of magnitude larger than that in existing cellular systems, and the small carrier wavelength enables the realization of highly directive antenna arrays in compact form factor, thus drastically increasing spatial reuse. In this paper, we carry out an interference analysis for mm wave picocells in an urban canyon, accounting for the geometry associated with the sparse multipath characteristic of this band. While we make some modeling simplifications, our analysis provides a strong indication of the very large capacity, of the order of Terabits/sec per km, provided by such networks, using system bandwidths of the order of a few GHz.
Zhinus Marzi, Upamanyu Madhow, Haitao Zheng 0001
GLOBECOM3
2015 Reusing 60GHz Radios for Mobile Radar Imaging
abstract
The future of mobile computing involves autonomous drones, robots and vehicles. To accurately sense their surroundings in a variety of scenarios, these mobile computers require a robust environmental mapping system. One attractive approach is to reuse millimeterwave communication hardware in these devices, e.g. 60GHz networking chipset, and capture signals reflected by the target surface. The devices can also move while collecting reflection signals, creating a large synthetic aperture radar (SAR) for high-precision RF imaging. Our experimental measurements, however, show that this approach provides poor precision in practice, as imaging results are highly sensitive to device positioning errors that translate into phase errors. We address this challenge by proposing a new 60GHz imaging algorithm, {\em RSS Series Analysis}, which images an object using only RSS measurements recorded along the device's trajectory. In addition to object location, our algorithm can discover a rich set of object surface properties at high precision, including object surface orientation, curvature, boundaries, and surface material. We tested our system on a variety of common household objects (between 5cm--30cm in width). Results show that it achieves high accuracy (cm level) in a variety of dimensions, and is highly robust against noises in device position and trajectory tracking. We believe that this is the first practical mobile imaging system (re)using 60GHz networking devices, and provides a basic primitive towards the construction of detailed environmental mapping systems.
Yanzi Zhu, Yibo Zhu 0001, Ben Y. Zhao, Haitao Zheng 0001
MobiCom4
2015 Packet-Level Telemetry in Large Datacenter Networks
abstract
Debugging faults in complex networks often requires capturing and analyzing traffic at the packet level. In this task, datacenter networks (DCNs) present unique challenges with their scale, traffic volume, and diversity of faults. To troubleshoot faults in a timely manner, DCN administrators must a) identify affected packets inside large volume of traffic; b) track them across multiple network components; c) analyze traffic traces for fault patterns; and d) test or confirm potential causes. To our knowledge, no tool today can achieve both the specificity and scale required for this task.
Yibo Zhu 0001, Nanxi Kang, Jiaxin Cao, Albert G. Greenberg, Guohan Lu, Ratul Mahajan, David A. Maltz, Ming Zhang 0005, Ben Y. Zhao, Haitao Zheng 0001
SIGCOMM11
2015 Energy and Performance of Smartphone Radio Bundling in Outdoor Environments
abstract
Most of today's mobile devices come equipped with both cellular LTE and WiFi wireless radios, making radio bundling (simultaneous data transfers over multiple interfaces) both appealing and practical. Despite recent studies documenting the benefits of radio bundling with MPTCP, many fundamental questions remain about potential gains from radio bundling, or the relationship between performance and energy consumption in these scenarios. In this study, we seek to answer these questions using extensive measurements to empirically characterize both energy and performance for radio bundling approaches. In doing so, we quantify potential gains of bundling using MPTCP versus an ideal protocol. We study the links between traffic partitioning and bundling performance, and use a novel componentized energy model to quantify the energy consumed by CPUs (and radios) during traffic management. Our results show that MPTCP achieves only a fraction of the total performance gain possible, and that its energy-agnostic design leads to considerable power consumption by the CPU. We conclude that not only there is room for improved bundling performance, but an energy-aware bundling protocol is likely to achieve a much better tradeoff between performance and power consumption.
Ana Nika, Yibo Zhu 0001, Ning Ding 0004, Abhilash Jindal, Y. Charlie Hu, Ben Y. Zhao, Haitao Zheng 0001
WWW8
2015 Practical Conflict Graphs in the Wild
abstract
Today, most spectrum allocation algorithms use conflict graphs to capture interference conditions. The use of conflict graphs, however, is often questioned by the wireless community for two reasons. First, building accurate conflict graphs requires significant overhead, and hence does not scale to outdoor networks. Second, conflict graphs cannot properly capture accumulative interference. In this paper, we use large-scale measurement data as ground truth to understand how severe these problems are and whether they can be overcome. We build “practical” conflict graphs using measurement-calibrated propagation models, which remove the need for exhaustive signal measurements by interpolating signal strengths using calibrated models. Calibrated models are imperfect, and we study the impact of their errors on multiple steps in the process, from calibrating propagation models, predicting signal strengths, to building conflict graphs. At each step, we analyze the introduction, propagation, and final impact of errors by comparing each intermediate result to its ground-truth counterpart. Our work produces several findings. Calibrated propagation models generate location-dependent prediction errors, ultimately producing conservative conflict graphs. While these “estimated conflict graphs” lower spectrum utilization, their conservative nature improves reliability by reducing the impact of accumulative interference. Finally, we propose a graph augmentation technique to address remaining accumulative interference.
Zengbin Zhang, Gang Wang 0011, Xiaoxiao Yu, Ben Y. Zhao, Haitao Zheng 0001
IEEE/ACM Trans. Netw.6
2014 Link and Triadic Closure Delay: Temporal Metrics for Social Network Dynamics
Matteo Zignani, Sabrina Gaito, Gian Paolo Rossi 0001, Xiaohan Zhao, Haitao Zheng 0001, Ben Y. Zhao
ICWSM5
2014 Whispers in the dark: analysis of an anonymous social network
abstract
Social interactions and interpersonal communication has undergone significant changes in recent years. Increasing awareness of privacy issues and events such as the Snowden disclosures have led to the rapid growth of a new generation of anonymous social networks and messaging applications. By removing traditional concepts of strong identities and social links, these services encourage communication between strangers, and allow users to express themselves without fear of bullying or retaliation.
Gang Wang 0011, Bolun Wang, Tianyi Wang 0001, Ana Nika, Haitao Zheng 0001, Ben Y. Zhao
Internet Measurement Conference5
2014 Demystifying 60GHz outdoor picocells
abstract
Mobile network traffic is set to explode in our near future, driven by the growth of bandwidth-hungry media applications. Current capacity solutions, including buying spectrum, WiFi offloading, and LTE picocells, are unlikely to supply the orders-of-magnitude bandwidth increase we need. In this paper, we explore a dramatically different alternative in the form of 60GHz mmwave picocells with highly directional links. While industry is investigating other mmwave bands (e.g. 28GHz to avoid oxygen absorption), we prefer the unlicensed 60GHz band with highly directional, short-range links (~100m). 60GHz links truly reap the spatial reuse benefits of small cells while delivering high per-user data rates and leveraging efforts on indoor 60GHz PHY technology and standards. Using extensive measurements on off-the-shelf 60GHz radios and system-level simulations, we explore the feasibility of 60GHz picocells by characterizing range, attenuation due to reflections, sensitivity to movement and blockage, and interference in typical urban environments. Our results dispel some common myths, and show that there are no fundamental physical barriers to high-capacity 60GHz outdoor picocells. We conclude by identifying open challenges and associated research opportunities.
Yibo Zhu 0001, Zengbin Zhang, Zhinus Marzi, Chris Nelson, Upamanyu Madhow, Ben Y. Zhao, Haitao Zheng 0001
MobiCom7
2014 Cutting the cord: a robust wireless facilities network for data centers
abstract
Today's network control and management traffic are limited by their reliance on existing data networks. Fate sharing in this context is highly undesirable, since control traffic has very different availability and traffic delivery requirements. In this paper, we explore the feasibility of building a dedicated wireless facilities network for data centers. We propose Angora, a low-latency facilities network using low-cost, 60GHz beamforming radios that provides robust paths decoupled from the wired network, and flexibility to adapt to workloads and network dynamics. We describe our solutions to address challenges in link coordination, link interference and network failures. Our testbed measurements and simulation results show that Angora enables large number of low-latency control paths to run concurrently, while providing low latency end-to-end message delivery with high tolerance for radio and rack failures.
Yibo Zhu 0001, Zengbin Zhang, Amin Vahdat, Ben Y. Zhao, Haitao Zheng 0001
MobiCom7
2014 Understanding data hotspots in cellular networks
abstract
The unprecedented growth in mobile data usage is posing significant challenges to cellular operators. One key challenge is how to provide quality of service to subscribers when their residing cell is experiencing a significant amount of traffic, i.e. becoming a traffic hotspot. In this paper, we perform an empirical study on data hotspots in today's cellular networks using a 9-week cellular dataset with 734K+ users and 5327 cell sites. Our analysis examines in details static and dynamic characteristics, predictability, and causes of data hotspots, and their correlation with call hotspots. We believe the understanding of these key issues will lead to more efficient and responsive resource management and thus better QoS provision in cellular networks. To the best of our knowledge, our work is the first to characterize in detail traffic hotspots in today's cellular networks using real data.
Ana Nika, Asad Ismail, Ben Y. Zhao, Sabrina Gaito, Gian Paolo Rossi 0001, Haitao Zheng 0001
QSHINE6
2014 Man vs. Machine: Practical Adversarial Detection of Malicious Crowdsourcing Workers
Gang Wang 0011, Tianyi Wang 0001, Haitao Zheng 0001, Ben Y. Zhao
USENIX Security Symposium3
2013 On the validity of geosocial mobility traces
abstract
Mobile networking researchers have long searched for large-scale, fine-grained traces of human movement, which have remained elusive for both privacy and logistical reasons. Recently, researchers have begun to focus on geosocial mobility traces, e.g. Foursquare checkin traces, because of their availability and scale. But are we conceding correctness in our zeal for data? In this paper, we take initial steps towards quantifying the value of geosocial datasets using a large ground truth dataset gathered from a user study. By comparing GPS traces against Foursquare checkins, we find that a large portion of visited locations is missing from checkins, and most checkin events are either forged or superfluous events. We characterize extraneous checkins, describe possible techniques for their detection, and show that both extraneous and missing checkins introduce significant errors into applications driven by these traces.
Zengbin Zhang, Xiaohan Zhao, Gang Wang 0011, Yu Su 0001, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao
HotNets7
2013 Follow the green: growth and dynamics in twitter follower markets
abstract
The users of microblogging services, such as Twitter, use the count of followers of an account as a measure of its reputation or influence. For those unwilling or unable to attract followers naturally, a growing industry of "Twitter follower markets" provides followers for sale. Some markets use fake accounts to boost the follower count of their customers, while others rely on a pyramid scheme to turn non-paying customers into followers for each other, and into followers for paying customers. In this paper, we present a detailed study of Twitter follower markets, report in detail on both the static and dynamic properties of customers of these markets, and develop and evaluate multiple techniques for detecting these activities. We show that our detection system is robust and reliable, and can detect a significant number of customers in the wild.
Gianluca Stringhini, Gang Wang 0011, Manuel Egele, Christopher Krügel, Giovanni Vigna, Haitao Zheng 0001, Ben Y. Zhao
Internet Measurement Conference6
2013 Social Turing Tests: Crowdsourcing Sybil Detection
Gang Wang 0011, Manish Mohanlal, Christo Wilson, Xiao Wang 0018, Miriam J. Metzger, Haitao Zheng 0001, Ben Y. Zhao
NDSS6
2013 Characterizing and detecting malicious crowdsourcing
abstract
Popular Internet services in recent years have shown that remarkable things can be achieved by harnessing the power of the masses. However, crowd-sourcing systems also pose a real challenge to existing security mechanisms deployed to protect Internet services, particularly those tools that identify malicious activity by detecting activities of automated programs such as CAPTCHAs.
Tianyi Wang 0001, Gang Wang 0011, Xing Li 0001, Haitao Zheng 0001, Ben Y. Zhao
SIGCOMM4
2013 Practical conflict graphs for dynamic spectrum distribution
abstract
Most spectrum distribution proposals today develop their allocation algorithms that use conflict graphs to capture interference relationships. The use of conflict graphs, however, is often questioned by the wireless community because of two issues. First, building conflict graphs requires significant overhead and hence generally does not scale to outdoor networks, and second, the resulting conflict graphs do not capture accumulative interference. In this paper, we use large-scale measurement data as ground truth to understand just how severe these issues are in practice, and whether they can be overcome. We build "practical" conflict graphs using measurement-calibrated propagation models, which remove the need for exhaustive signal measurements by interpolating signal strengths using calibrated models. These propagation models are imperfect, and we study the impact of their errors by tracing the impact on multiple steps in the process, from calibrating propagation models to predicting signal strength and building conflict graphs. At each step, we analyze the introduction, propagation and final impact of errors, by comparing each intermediate result to its ground truth counterpart generated from measurements. Our work produces several findings. Calibrated propagation models generate location-dependent prediction errors, ultimately producing conservative conflict graphs. While these "estimated conflict graphs" lose some spectrum utilization, their conservative nature improves reliability by reducing the impact of accumulative interference. Finally, we propose a graph augmentation technique that addresses any remaining accumulative interference, the last missing piece in a practical spectrum distribution system using measurement-calibrated conflict graphs.
Zengbin Zhang, Gang Wang 0011, Xiaoxiao Yu, Ben Y. Zhao, Haitao Zheng 0001
SIGMETRICS6
2013 You Are How You Click: Clickstream Analysis for Sybil Detection
Gang Wang 0011, Tristan Konolige, Christo Wilson, Xiao Wang 0018, Haitao Zheng 0001, Ben Y. Zhao
USENIX Security Symposium5
2013 Wisdom in the social crowd: an analysis of quora
abstract
Efforts such as Wikipedia have shown the ability of user communities to collect, organize and curate information on the Internet. Recently, a number of question and answer (Q&A) sites have successfully built large growing knowledge repositories, each driven by a wide range of questions and answers from its users community. While sites like Yahoo Answers have stalled and begun to shrink, one site still going strong is Quora, a rapidly growing service that augments a regular Q&A system with social links between users. Despite its success, however, little is known about what drives Quora's growth, and how it continues to connect visitors and experts to the right questions as it grows.
Gang Wang 0011, Konark Gill, Manish Mohanlal, Haitao Zheng 0001, Ben Y. Zhao
WWW4
2013 On the Embeddability of Random Walk Distances
abstract
Analysis of large graphs is critical to the ongoing growth of search engines and social networks. One class of queries centers around node affinity, often quantified by random-walk distances between node pairs, including hitting time, commute time, and personalized PageRank (PPR). Despite the potential of these "metrics," they are rarely, if ever, used in practice, largely due to extremely high computational costs. In this paper, we investigate methods to scalably and efficiently compute random-walk distances, by "embedding" graphs and distances into points and distances in geometric coordinate spaces. We show that while existing graph coordinate systems (GCS) can accurately estimate shortest path distances, they produce significant errors when embedding random-walk distances. Based on our observations, we propose a new graph embedding system that explicitly accounts for per-node graph properties that affect random walk. Extensive experiments on a range of graphs show that our new approach can accurately estimate both symmetric and asymmetric random-walk distances. Once a graph is embedded, our system can answer queries between any two nodes in 8 microseconds, orders of magnitude faster than existing methods. Finally, we show that our system produces estimates that can replace ground truth in applications with minimal impact on application output.
Xiaohan Zhao, Adelbert Chang, Atish Das Sarma, Haitao Zheng 0001, Ben Y. Zhao
Proc. VLDB Endow.4
2013 Measurement-Based Design of Roadside Content Delivery Systems
abstract
With today's ubiquity of thin computing devices, mobile users are accustomed to having rich location-aware information at their fingertips, such as restaurant menus, shopping mall maps, movie showtimes, and trailers. However, delivering rich content is challenging, particularly for highly mobile users in vehicles. Technologies such as cellular-3G provide limited bandwidth at significant costs. In contrast, providers can cheaply and easily deploy a small number of WiFi infostations that quickly deliver large content to vehicles passing by for future offline browsing. While several projects have proposed systems for disseminating content via roadside infostations, most use simplified models and simulations to guide their design for scalability. Many suspect that scalability with increasing vehicle density is the major challenge for infostations, but few if any have studied the performance of these systems via real measurements. Intuitively, per-vehicle throughput for unicast infostations degrades with the number of vehicles near the infostation, while broadcast infostations are unreliable, and lack rate adaptation. In this work, we collect over 200 h of detailed highway measurements with a fleet of WiFi-enabled vehicles. We use analysis of these results to explore the design space of WiFi infostations, in order to determine whether unicast or broadcast should be used to build high-throughput infostations that scale with device density. Our measurement results demonstrate the limitations of both approaches. Our insights lead to Starfish, a high-bandwidth and scalable infostation system that incorporates device-to-device data scavenging, where nearby vehicles share data received from the infostation. Data scavenging increases dissemination throughput by a factor of 2-6, allowing both broadcast and unicast throughput to scale with device density.
Vinod Kone, Haitao Zheng 0001, Antony I. T. Rowstron, Greg O'Shea, Ben Y. Zhao
IEEE Trans. Mob. Comput.2
2012 Multi-scale dynamics in a massive online social network
abstract
Data confidentiality policies at major social network providers have severely limited researchers' access to large-scale datasets. The biggest impact has been on the study of network dynamics, where researchers have studied citation graphs and content-sharing networks, but few have analyzed detailed dynamics in the massive social networks that dominate the web today. In this paper, we present results of analyzing detailed dynamics in a large Chinese social network, covering a period of 2 years when the network grew from its first user to 19 million users and 199 million edges. Rather than validate a single model of network dynamics, we analyze dynamics at different granularities (per-user, per-community, and network-wide) to determine how much, if any, users are influenced by dynamics processes at different scales. We observe independent predictable processes at each level, and find that the growth of communities has moderate and sustained impact on users. In contrast, we find that significant events such as network merge events have a strong but short-lived impact on users, and they are quickly eclipsed by the continuous arrival of new users.
Xiaohan Zhao, Alessandra Sala, Christo Wilson, Xiao Wang 0018, Sabrina Gaito, Haitao Zheng 0001, Ben Y. Zhao
Internet Measurement Conference6
2012 Enforcing dynamic spectrum access with spectrum permits
abstract
Dynamic spectrum access is a maturing technology that allows next generation wireless devices to make highly efficient use of wireless spectrum. Spectrum can be allocated on an on-demand basis for a given geographic location, time duration and frequency range. However, a major obstacle to adoption remains. There are no effective solutions to protect licensed users from spectrum misuse, where users transmit without properly licensing spectrum, and in doing so, interfere and disrupt legitimate flows to whom the spectrum is assigned. Given the flexibility of today's cognitive radios, an application can easily transmit on frequencies outside of its allocated range, either accidentally due to misconfiguration, or intentionally to avoid spectrum licensing costs. In this paper, we propose a system to secure dynamic spectrum transmissions, where authorized users embed secure spectrum permits into data transmissions, thus enabling patrolling trusted devices to detect devices transmitting without authorization. We focus our attention on the development of spectrum permits, and describe Gelato, a spectrum misuse detection system that minimizes both hardware costs and performance overhead on legitimate data transmissions.
Lei Yang 0020, Zengbin Zhang, Ben Y. Zhao, Christopher Krügel, Haitao Zheng 0001
MobiHoc5
2012 Mirror mirror on the ceiling: flexible wireless links for data centers
abstract
Modern data centers are massive, and support a range of distributed applications across potentially hundreds of server racks. As their utilization and bandwidth needs continue to grow, traditional methods of augmenting bandwidth have proven complex and costly in time and resources. Recent measurements show that data center traffic is often limited by congestion loss caused by short traffic bursts. Thus an attractive alternative to adding physical bandwidth is to augment wired links with wireless links in the 60 GHz band.
Zengbin Zhang, Yibo Zhu 0001, Saipriya Kumar, Amin Vahdat, Ben Y. Zhao, Haitao Zheng 0001
SIGCOMM8
2012 Serf and turf: crowdturfing for fun and profit
abstract
Popular Internet services in recent years have shown that remarkable things can be achieved by harnessing the power of the masses using crowd-sourcing systems. However, crowd-sourcing systems can also pose a real challenge to existing security mechanisms deployed to protect Internet services. Many of these security techniques rely on the assumption that malicious activity is generated automatically by automated programs. Thus they would perform poorly or be easily bypassed when attacks are generated by real users working in a crowd-sourcing system. Through measurements, we have found surprising evidence showing that not only do malicious crowd-sourcing systems exist, but they are rapidly growing in both user base and total revenue. We describe in this paper a significant effort to study and understand these "crowdturfing" systems in today's Internet. We use detailed crawls to extract data about the size and operational structure of these crowdturfing systems. We analyze details of campaigns offered and performed in these sites, and evaluate their end-to-end effectiveness by running active, benign campaigns of our own. Finally, we study and compare the source of workers on crowdturfing sites in different countries. Our results suggest that campaigns on these systems are highly effective at reaching users, and their continuing growth poses a concrete threat to online communities both in the US and elsewhere.
Gang Wang 0011, Christo Wilson, Xiaohan Zhao, Yibo Zhu 0001, Manish Mohanlal, Haitao Zheng 0001, Ben Y. Zhao
WWW6
2012 Balancing Reliability and Utilization in Dynamic Spectrum Access
abstract
Future wireless networks will dynamically access spectrum to maximize its utilization. Conventional design of dynamic spectrum access focuses on maximizing spectrum utilization, but faces the problem of degraded reliability due to unregulated demands and access behaviors. Without providing proper reliability guarantee, dynamic spectrum access is unacceptable to many infrastructure networks and services. In this paper, we propose SPARTA, a new architecture for dynamic spectrum access that balances access reliability and spectrum utilization. SPARTA includes two complementary techniques: proactive admission control performed by a central entity to determine the set of wireless nodes to be supported with only statistical information of their spectrum demands, and online adaptation performed by admitted wireless nodes to adjust their instantaneous spectrum usage to time-varying demand. Using both theoretical analysis and simulation, we show that SPARTA fulfills the reliability requirements while dynamically multiplexing spectrum demands to improve utilization. Compared to conventional solutions, SPARTA improves spectrum utilization by 80%-200%. Finally, SPARTA also allows service providers to explore the tradeoff between utilization and reliability to make the best use of the spectrum. To our best knowledge, our work is the first to identify and address such a tradeoff.
Lili Cao, Haitao Zheng 0001
IEEE/ACM Trans. Netw.2
2012 The effectiveness of opportunistic spectrum access: a measurement study
abstract
Dynamic spectrum access networks are designed to allow today's bandwidth-hungry “secondary devices” to share spectrum allocated to legacy devices, or “primary users.” The success of this wireless communication model relies on the availability of unused spectrum and the ability of secondary devices to utilize spectrum without disrupting transmissions of primary users. While recent measurement studies have shown that there is sufficient underutilized spectrum available, little is known about whether secondary devices can efficiently make use of available spectrum while minimizing disruptions to primary users. In this paper, we present the first comprehensive study on the presence of “usable” spectrum in opportunistic spectrum access systems, and whether sufficient spectrum can be extracted by secondary devices to support traditional networking applications. We use for our study fine-grain usage traces of a wide spectrum range (20 MHz–6 GHz) taken at four locations in Germany, the Netherlands, and Santa Barbara, CA. Our study shows that on average, 54% of spectrum is never used and 26% is only partially used. Surprisingly, in this 26% of partially used spectrum, secondary devices can utilize very little spectrum using conservative access policies to minimize interference with primary users. Even assuming an optimal access scheme and extensive statistical knowledge of primary-user access patterns, a user can only extract between 20%–30% of the total available spectrum. To provide better spectrum availability, we propose frequency bundling, where secondary devices build reliable channels by combining multiple unreliable frequencies into virtual frequency bundles. Analyzing our traces, we find that there is little correlation of spectrum availability across channels, and that bundling random channels together can provide sustained periods of reliable transmission with only short interruptions.
Vinod Kone, Lei Yang 0020, Xue Yang 0007, Ben Y. Zhao, Haitao Zheng 0001
IEEE/ACM Trans. Netw.5
2011 Efficient shortest paths on massive social graphs
abstract
Analysis of large networks is a critical component of many of today’s application environments. The arrival of massive network graphs with hundreds of millions of nodes, e.g. social graphs, presents a unique challenge to graph analysis applications. Most of these applications rely on computing dista
Xiaohan Zhao, Alessandra Sala, Haitao Zheng 0001, Ben Y. Zhao
CollaborateCom3
2011 3D beamforming for wireless data centers
abstract
Contrary to prior assumptions, recent measurements show that data center traffic is not constrained by network bisection bandwidth, but is instead prone to congestion loss caused by short traffic bursts. Compared to the cost and complexity of modifying data center architectures, a much more attractive option is to augment wired links with flexible wireless links in the 60 GHz band. Current proposals, however, are severely constrained by two factors. First, 60 GHz wireless links are limited by line-of-sight, and can be blocked by even small obstacles between the endpoints. Second, even beamforming links leak power, and potential interference will severely limit concurrent transmissions in dense data centers. In this paper, we explore the feasibility of a new wireless primitive for data centers, 3D beamforming. We explore the design space, and show how bouncing 60 GHz wireless links off reflective ceilings can address both link blockage and link interference, thus improving link range and number of current transmissions in the data center.
Weile Zhang, Lei Yang 0020, Zengbin Zhang, Ben Y. Zhao, Haitao Zheng 0001
HotNets6
2011 Sharing graphs using differentially private graph models
abstract
Continuing success of research on social and computer networks requires open access to realistic measurement datasets. While these datasets can be shared, generally in the form of social or Internet graphs, doing so often risks exposing sensitive user data to the public. Unfortunately, current techniques to improve privacy on graphs only target specific attacks, and have been proven to be vulnerable against powerful de-anonymization attacks.
Alessandra Sala, Xiaohan Zhao, Christo Wilson, Haitao Zheng 0001, Ben Y. Zhao
Internet Measurement Conference4
2011 To preempt or not: Tackling bid and time-based cheating in online spectrum auctions
abstract
Online spectrum auctions offer ample flexibility for bidders to request and obtain spectrum on-the-fly. Such flexibility, however, opens up new vulnerabilities to bidder manipulation. Aside from rigging their bids, selfish bidders can falsely report their arrival time to game the system and obtain unfair advantage over others. Such time-based cheating is easy to perform yet produces severe damage to auction performance. We propose Topaz, a truthful online spectrum auction design that distributes spectrum efficiently while discouraging bidders from misreporting their bids or time report. Topaz makes three key contributions. First, Topaz applies a 3D bin packing mechanism to distribute spectrum across time, space and frequency, exploiting spatial and time reuse to improve allocation efficiency. Second, Topaz enforces truthfulness using a novel temporal-smoothed critical value based pricing. Capturing the temporal and spatial dependency among bidders who arrive subsequently, this pricing effectively diminishes gain from bid and/or time-cheating. Finally, Topaz offers a “scalable” winner preemption to address the uncertainty of future arrivals at each decision time, which significantly boosts auction revenue. We analytically prove Topaz's truthfulness, which does not require any knowledge of bidder behavior, or an optimal spectrum allocation to enforce truthfulness. Using empirical arrival and bidding models, we perform simulations to demonstrate the efficiency of Topaz. We show that proper winner preemption improves auction revenue by 45-65% at a minimum cost of spectrum utilization.
Lara B. Deek, Kevin C. Almeroth, Haitao Zheng 0001
INFOCOM4
2011 I am the antenna: accurate outdoor AP location using smartphones
abstract
Today's WiFi access points (APs) are ubiquitous, and provide critical connectivity for a wide range of mobile networking devices. Many management tasks, e.g. optimizing AP placement and detecting rogue APs, require a user to efficiently determine the location of wireless APs. Unlike prior localization techniques that require either specialized equipment or extensive outdoor measurements, we propose a way to locate APs in real-time using commodity smartphones. Our insight is that by rotating a wireless receiver (smartphone) around a signal-blocking obstacle (the user's body), we can effectively emulate the sensitivity and functionality of a directional antenna. Our measurements show that we can detect these signal strength artifacts on multiple smartphone platforms for a variety of outdoor environments. We develop a model for detecting signal dips caused by blocking obstacles, and use it to produce a directional analysis technique that accurately predicts the direction of the AP, along with an associated confidence value. The result is Borealis, a system that provides accurate directional guidance and leads users to a desired AP after a few measurements. Detailed measurements show that Borealis is significantly more accurate than other real-time localization systems, and is nearly as accurate as offline approaches using extensive wireless measurements.
Zengbin Zhang, Weile Zhang, Yuanyang Zhang, Gang Wang 0011, Ben Y. Zhao, Haitao Zheng 0001
MobiCom7
2011 The Impact of Infostation Density on Vehicular Data Dissemination
abstract
Vehicle-to-Vehicle and Vehicle-to-Roadside communications are going to become an indispensable part of the modern day automotive experience. For people on the move, vehicular networks can provide critical network connectivity and access to real-time information. Infostations play a vital role in these networks by acting as gateways to the Internet and by extending network connectivity. In this context, an important question is “What is the minimum number of infostations that need to be deployed in an area in order to support vehicular applications?” Optimizing infostation density is vital to understanding and reducing the cost of deployment and management. In this paper, we examine the required infostation density in a highway scenario using different data dissemination models. We start from a simple analysis that captures the required density under idealized assumptions. These models are validated by an event-driven simulator. We then run detailed QualNet simulations on both controlled and realistic vehicular traces to observe the information density trends in practical environments, and consequently propose techniques to improve dissemination performance and reduce the required infostation density.
Vinod Kone, Haitao Zheng 0001, Antony I. T. Rowstron, Ben Y. Zhao
Mob. Networks Appl.2
2010 On the feasibility of effective opportunistic spectrum access
abstract
Dynamic spectrum access networks are designed to allow today's bandwidth hungry "secondary devices" to share spectrum allocated to legacy devices, or "primary users." The success of this wireless communication model relies on the availability of unused spectrum, and the ability of secondary devices to utilize spectrum without disrupting transmissions of primary users. While recent measurement studies have shown that there is sufficient underutilized spectrum available, little is known about whether secondary devices can efficiently make use of available spectrum while minimizing disruptions to primary users.
Vinod Kone, Lei Yang 0020, Xue Yang 0007, Ben Y. Zhao, Haitao Zheng 0001
Internet Measurement Conference5
2010 The spaces between us: setting and maintaining boundaries in wireless spectrum access
abstract
Guardbands are designed to insulate transmissions on adjacent frequencies from mutual interference. As more devices in a given area are packed into orthogonal wireless channels, choosing the right guardband size to minimize cross-channel interference becomes critical to network performance. Using both WiFi and GNU radio experiments, we show that the traditional "one-size-fits-all" approach to guardband assignment is ineffective, and can produce throughput degradation up to 80%. We find that ideal guardband values vary across different network configurations, and across different links in the same network. We argue that guardband values should be set based on network conditions and adapt to changes over time.
Lei Yang 0020, Ben Y. Zhao, Haitao Zheng 0001
MobiCom3
2010 Breaking bidder collusion in large-scale spectrum auctions
abstract
Dynamic spectrum auction is an effective solution to provide spectrum on-demand to many small wireless networks. As the number of participants grows, bidder collusion becomes a serious threat. In this paper, we study bidder collusion in large-scale spectrum auctions, investigating its impact on auction outcomes. We found that the nature of the complex interference constraints among bidders provides a fertile breeding ground for colluders, causing significant damage in auction efficiency and revenue. In particular, collusion group of small size plays a dominant role since it is easy to form and hard to be detected.We propose Athena, a new collusion-resistant auction framework for large-scale dynamic spectrum auction. Athena implements a soft collusion resistance, allowing the auctioneer to exploit the tradeoff between the level of collusion resistance and the cost of achieving such level of resistance. Unlike existing solutions, Athena enables spectrum reuse across bidders, achieves soft collusion resistance against any form of collusive bidding strategy, maintains provable revenue guarantee, and does so with polynomial-time complexity. To provide a comprehensive evaluation, we first analytically prove Athena's collusion resistance and revenue guarantee (under any bids), and then experimentally verify our analytical conclusions using empirical bid distributions.
Haitao Zheng 0001
MobiHoc2
2010 Supporting Demanding Wireless Applications with Frequency-agile Radios
Lei Yang 0020, Lili Cao, Ben Y. Zhao, Haitao Zheng 0001
NSDI5
2010 Brief announcement: revisiting the power-law degree distribution for social graph analysis
abstract
The study of complex networks led to the belief that the connectivity of network nodes generally follows a Power-law distribution. In this work, we show that modeling large-scale online social networks using a Power-law distribution produces significant fitting errors. We propose the use of a more accurate node degree distribution model based on the Pareto-Lognormal distribution. Using large datasets gathered from Facebook, we show that the Power-law curve produces a significant over-estimation of the number of high degree nodes, leading researchers to erroneous designs for a number of social applications and systems, including shortest-path prediction, community detection, and influence maximization. We provide a formal proof of the error reduction using the Pareto-Lognormal distribution, which we envision will have strong implications on the correctness of social systems and applications.
Alessandra Sala, Haitao Zheng 0001, Ben Y. Zhao, Sabrina Gaito, Gian Paolo Rossi 0001
PODC2
2010 Coexistence-Aware Scheduling for Wireless System-on-a-Chip Devices
abstract
Today's mobile devices support many wireless technologies to achieve ubiquitous connectivity. Economic and energy constraints, however, are driving the industry to implement multiple technologies into a single radio. This system-on-a-chip architecture leads to competition among networks when devices toggle across different technologies to communicate with multiple networks. In this paper, we study the impact of such network competition using a representative scenario where devices split their time between WiMAX and WiFi connections. We show that competition with WiMAX significantly lowers WiFi's throughput, but this performance degradation is largely unnecessary, and can be attributed to the fact that WiMAX's transmission scheduling does not consider competing networks. We propose PACT, a new coexistence-aware WiMAX scheduling policy that cooperates with WiFi links hosted by its users without compromising its own transmission requirements. We derive PACT's design using an analytical model of network competition, and apply it to design practical WiMAX scheduling algorithms for various traffic classes. We evaluate PACT using OPNET's realistic models for WiFi and WiMAX. Using real network topologies, our experiment results show that PACT significantly improves WiFi performance by up to 17 fold without affecting the WiMAX user experience.
Lei Yang 0020, Vinod Kone, Xue Yang 0007, York Liu, Ben Y. Zhao, Haitao Zheng 0001
SECON6
2010 Measurement-calibrated graph models for social network experiments
abstract
Access to realistic, complex graph datasets is critical to research on social networking systems and applications. Simulations on graph data provide critical evaluation of new systems and applications ranging from community detection to spam filtering and social web search. Due to the high time and resource costs of gathering real graph datasets through direct measurements, researchers are anonymizing and sharing a small number of valuable datasets with the community. However, performing experiments using shared real datasets faces three key disadvantages: concerns that graphs can be de-anonymized to reveal private information, increasing costs of distributing large datasets, and that a small number of available social graphs limits the statistical confidence in the results.
Alessandra Sala, Lili Cao, Christo Wilson, Robert Zablit, Haitao Zheng 0001, Ben Y. Zhao
WWW5
2009 TRUST: A General Framework for Truthful Double Spectrum Auctions
abstract
We design truthful double spectrum auctions where multiple parties can trade spectrum based on their individual needs. Open, market-based spectrum trading motivates existing spectrum owners (as sellers) to lease their selected idle spectrum to new spectrum users, and provides new users (as buyers) the spectrum they desperately need. The most significant challenge is how to make the auction economic-robust (truthful in particular) while enabling spectrum reuse to improve spectrum utilization. Unfortunately, existing designs either do not consider spectrum reuse or become untruthful when applied to double spectrum auctions. We address this challenge by proposing TRUST, a general framework for truthful double spectrum auctions. TRUST takes as input any reusability-driven spectrum allocation algorithm, and applies a novel winner determination and pricing mechanism to achieve truthfulness and other economic properties while significantly improving spectrum utilization. To our best knowledge, TRUST is the first solution for truthful double spectrum auctions that enable spectrum reuse. Our results show that economic factors introduce a tradeoff between spectrum efficiency and economic robustness. TRUST makes an important contribution on enabling spectrum reuse to minimize such tradeoff.
Haitao Zheng 0001
INFOCOM2
2009 Preamble design for non-contiguous spectrum usage in cognitive radio networks
abstract
Cognitive radios can significantly improve spectrum efficiency by using locally available spectrum. The efficiency, however, depends heavily on their transceiver design. In particular, being able to use non-contiguously aligned spectrum bands simultaneously is a critical requirement. Prior work in this area requires a control channel so that transmitter/receiver pairs can synchronize on their spectrum usage patterns. However, this approach can suffer from high cost and control congestion. In this paper, we propose an in-band solution for informing receivers the spectrum usage patterns. By judiciously designing packet preambles, we embed the spectrum usage patterns in each data packet. Using the legacy 802.11 preamble structure, we focus on choosing the appropriate preamble sequences to maintain reliable packet detection in the presence of noise and interference. We verify our design using simulation and show that it can lead to reliable packet transmissions comparable to those of contiguous spectrum usage. We also identify the impact of interference on our design and propose refinements to choose the preamble sequence using information on the interference.
Shulan Feng, Haitao Zheng 0001, Jinnan Liu, Philipp Zhang
WCNC2
2009 Securing Structured Overlays against Identity Attacks
abstract
Structured overlay networks can greatly simplify data storage and management for a variety of distributed applications. Despite their attractive features, these overlays remain vulnerable to the Identity attack, where malicious nodes assume control of application components by intercepting and hijacking key-based routing requests. Attackers can assume arbitrary application roles such as storage node for a given file, or return falsified contents of an online shopper's shopping cart. In this paper, we define a generalized form of the Identity attack, and propose a lightweight detection and tracking system that protects applications by redirecting traffic away from attackers. We describe how this attack can be amplified by a Sybil or Eclipse attack, and analyze the costs of performing such an attack. Finally, we present measurements of a deployed overlay that show our techniques to be significantly more lightweight than prior techniques, and highly effective at detecting and avoiding both single node and colluding attacks under a variety of conditions.
Krishna P. N. Puttaswamy, Haitao Zheng 0001, Ben Y. Zhao
IEEE Trans. Parallel Distributed Syst.2
2008 Stable and Efficient Spectrum Access in Next Generation Dynamic Spectrum Networks
abstract
Future wireless infrastructure networks will dynamically access spectrum for maximum utilization. However, the fundamental challenge is how to provide stable spectrum access required for most applications. Using dynamic spectrum access, each node's spectrum usage is inherently unpredictable and unstable. We propose to address this challenge by integrating interference-aware statistical admission control with stability- driven spectrum allocation. Specifically, we propose to proactively regulate nodes' spectrum demand to allow efficient statistical multiplexing while minimizing outages. Admitted nodes coordinate to adapt instantaneous spectrum allocation to match time- varying demand. While the optimization problem is NP-hard, we develop computational-efficient algorithms with strong analytical guarantees. Experimental results show that the proposed approach can provide stable spectrum usage while improving its utilization by 80-100% compared to conventional solutions.
Lili Cao, Haitao Zheng 0001
INFOCOM2
2008 eBay in the Sky: strategy-proof wireless spectrum auctions
abstract
Market-driven dynamic spectrum auctions can drastically improve the spectrum availability for wireless networks struggling to obtain additional spectrum. However, they face significant challenges due to the fear of market manipulation. A truthful or strategy-proof spectrum auction eliminates the fear by enforcing players to bid their true valuations of the spectrum. Hence bidders can avoid the expensive overhead of strategizing over others and the auctioneer can maximize its revenue by assigning spectrum to bidders who value it the most. Conventional truthful designs, however, either fail or become computationally intractable when applied to spectrum auctions. In this paper, we propose VERITAS, a truthful and computationally-efficient spectrum auction to support an eBay-like dynamic spectrum market. VERITAS makes an important contribution of maintaining truthfulness while maximizing spectrum utilization. We show analytically that VERITAS is truthful, efficient, and has a polynomial complexity of O(n3k) when n bidders compete for k spectrum bands. Simulation results show that VERITAS outperforms the extensions of conventional truthful designs by up to 200% in spectrum utilization. Finally, VERITAS supports diverse bidding formats and enables the auctioneer to reconfigure allocations for multiple market objectives.
Sorabh Gandhi, Subhash Suri, Haitao Zheng 0001
MobiCom4
2008 Malware in IEEE 802.11 Wireless Networks
Brett Stone-Gross, Christo Wilson, Kevin C. Almeroth, Elizabeth M. Belding, Haitao Zheng 0001, Konstantina Papagiannaki
PAM5
2008 Towards real-time dynamic spectrum auctions
Sorabh Gandhi, Chiranjeeb Buragohain, Lili Cao, Haitao Zheng 0001, Subhash Suri
Comput. Networks4
2008 Distributed Rule-Regulated Spectrum Sharing
abstract
Dynamic spectrum access is a promising technique to use spectrum efficiently. Without being restricted to any prefixed spectrum bands, nodes choose operating spectrum on-demand. Such flexibility, however, makes efficient and fair spectrum access in large-scale networks a great challenge. Prior work in this area focused on explicit coordination where nodes communicate with peers to modify local spectrum allocation, and may heavily stress the communication resource. In this paper, we introduce a distributed spectrum management architecture where nodes share spectrum resource fairly by making independent actions following spectrum rules. We present five spectrum rules to regulate node behavior and maximize system fairness and spectrum utilization, and analyze the associated complexity and overhead. We show analytically and experimentally that the proposed rule-based approach achieves similar performance with the explicit coordination approach, while significantly reducing communication cost.
Lili Cao, Haitao Zheng 0001
IEEE J. Sel. Areas Commun.2
2008 Understanding the Power of Distributed Coordination for Dynamic Spectrum Management
Lili Cao, Haitao Zheng 0001
Mob. Networks Appl.2
2007 Multi-channel Jamming Attacks using Cognitive Radios
abstract
To improve spectrum efficiency, future wireless devices will use cognitive radios to dynamically access spectrum. While offering great flexibility and software-reconfigurability, unsecured cognitive radios can be easily manipulated to attack legacy and future wireless networks. In this paper, we explore the feasibility and impact of cognitive radio based jamming attacks on 802.11 networks. We show that attackers can utilize cognitive radios' fast channel switching capability to amplify their jamming impact across multiple channels using a single radio. We also examine the impact of hardware channel switching delays and jamming duration on the impact of jamming.
Ashwin Sampath, Hui Dai, Haitao Zheng 0001, Ben Y. Zhao
ICCCN3
2007 QUORUM: quality of service routing in wireless mesh networks
abstract
Wireless Mesh Networks (WMNs) can provide seamless broadband connectivity to network users, with the advantage of low setup and maintenance costs. To support next-generation applications with real-time requirements, however, these networks must provide improved Quality of Service guarantees. Most current mesh network routing protocols are adapted from MANET protocols, and do not optimize for mesh network properties. In this paper, we propose QUORUM (QUality Of service RoUting in wireless Mesh networks ), a routing protocol optimized for WMNs that addresses these drawbacks. QUORUM integrates a novel end-to-end packet delay estimation mechanism with stability-aware routing policies, allowing it to more accurately follow QoS requirements while minimizing misbehavior of selfish nodes.
Vinod Kone, Sudipto Das, Ben Y. Zhao, Haitao Zheng 0001
QSHINE4
2007 QUORUM - Quality of Service in Wireless Mesh Networks
Vinod Kone, Sudipto Das, Ben Y. Zhao, Haitao Zheng 0001
Mob. Networks Appl.4
2006 Utilization and fairness in spectrum assignment for opportunistic spectrum access
Haitao Zheng 0001, Ben Y. Zhao
Mob. Networks Appl.2
2005 Distributed spectrum allocation via local bargaining
abstract
Abstract — In this paper, we present an adaptive and dis-tributed approach to spectrum allocation in mobile ad-hoc networks. We propose a local bargaining approach where users affected by the mobility event self-organize into bargaining groups and adapt their spectrum assignment to approximate a new optimal assignment. The number of computations required to adapt to topology changes can be significantly reduced com-pared to that of the conventional topology-based optimizations that ignore the prior assignment. In particular, we propose a Fairness Bargaining with Feed Poverty to improve fairness in spectrum assignment and derive a theoretical lower bound on the minimum assignment each user can get from bargaining for certain network configurations. Such bound can be utilized to guide the bargaining process. We also show that the difference between the proposed bargaining approach and the true optimal approach is upper-bounded. Experimental results demonstrate that the proposed bargaining approach provides similar perfor-mance as the topology-based optimization but with more than 50 % of reduction in complexity. I.
Lili Cao, Haitao Zheng 0001
SECON2