Zhenhua Li 0001

dblp:61/1951-1 · DBLP profile ↗
← Back
124ranked-venue papers
15as first author
51since 2021 · last 2026
0000-0001-7286-122XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 70 · 7 first-author · 33 since 2021Systems, architecture and hardware · 29 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 since 2021Security and privacy · 8 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
abstract
The rapid scaling of large language models (LLMs) has significantly increased GPU memory pressure, which is further aggravated by training optimization techniques such as virtual pipeline and recomputation that disrupt tensor lifespans and introduce considerable memory fragmentation. Such fragmentation stems from the use of online GPU memory allocators in popular deep learning frameworks like PyTorch, which disregard tensor lifespans. As a result, this inefficiency can waste as much as 43% of memory and trigger out-of-memory errors, undermining the effectiveness of optimization methods.
Zixiao Huang 0001, Hao Lin 0005, Chunyang Zhu, Yueran Tang, Quanlu Zhang, Zhenhua Li 0001, Shengen Yan, Zhenhua Zhu 0002, Guohao Dai 0001, Yu Wang 0002
EuroSys8
2026 PatternInsight: An Online Approach to Complex Pattern Detection Over Mobile Data Streams
abstract
Today's mobile applications oftentimes need to detect user-defined complex patterns (e.g., the mysterious “phantom traffic jam”) over data streams to support decision making. It is achieved by continuously creating candidate instances that have partially matched a pattern, and meanwhile aggregating common instances (across patterns) for efficiency enhancement. Existing aggregation approaches are taken in a straightforward or intuitive manner, incurring an exponential solution space and thus having to be executed offline. This paper explores how to significantly accelerate aggregation so as to make pattern detection online executable, even suited to the emerging serverless runtime that involves complicated state synchronizations among distributed cloud functions. By comprehensively investigating a wide variety of mobile data streams, we note the existence of a latent hierarchical cluster structure among complex patterns (in terms of their instance similarities), which can be utilized to quickly aggregate common instances without going through the exponential solution space. To extract the latent information, we devise a content-aware structural entropy minimization algorithm to properly determine intra-cluster patterns, together with a lightweight differential compensation mechanism to maintain those inter-cluster “residual” relations among patterns. Evaluations on real-world vehicle and sensor network data streams illustrate that the resulting approach, dubbed PatternInsight, saves the aggregation time by 10× to 50× and reduces the instance size by 40%.
Yuyang Ren, Zhenhua Li 0001, Fei Xu 0009, Yunhao Liu 0001, Guihai Chen
IEEE Trans. Mob. Comput.3
2026 vSoC: Efficient and Debug-Friendly Virtual System-on-Chip for Mobile Emulation
abstract
Emerging heavy-load mobile apps like UHD video and AR/VR access diverse high-throughput hardware devices, e.g., video codecs and cameras. However, today's mobile emulators exhibit poor performance when emulating these devices. We pinpoint the major reason to be the discrepancy between the guest's (system-on-chip) and host's (PC or cloud server) memory architectures for these devices, which makes the shared virtual memory (SVM) architecture of mobile emulators highly inefficient. To address this, we introduce vSoC, the first virtual mobile SoC featuring aunifiedSVM framework that enables efficient and secure data sharing among virtual devices, as well as anintelligentprefetch engine that effectively eliminates the vast majority of coherence maintenance overhead. While vSoC addresses runtime performance issues, app developers face a severe debugging challenge due to the inability of traditional tools to capture complete system states. Therefore, we devise an adaptive VM (virtual machine) snapshot-based approach that dynamically selects the optimal resource loading strategy to make vSoC debug-friendly. Compared to state-of-the-art emulators, vSoC brings 1.8–9.0$\times$frame rates, 35%-62% lower motion-tophoton latency, and 6.5–14.4$\times$bug reproduction rates for heavyload apps. It is applicable to a variety of scenarios like end-user high-performance emulation and cloud/web-based rendering.
Jiaxing Qiu, Zhenhua Li 0001, Feng Qian 0001, Yunhao Liu 0001, Hongwei Hu
IEEE Trans. Mob. Comput.3
2026 AirCloak: An App-Transparent Traffic Cloaking Middleware Against Wireless Fingerprinting Attack
Huafeng Bian, Jianfeng Li 0006, Haodan Luo, Xiaobo Ma 0001, Zhenhua Li 0001, Jigang Wang, Wei Wang 0012
IEEE Trans. Netw.6
2025 Democratizing the Cryptocurrency Ecosystem by Just-In-Time Transformation of Mining Programs
abstract
Democracy is crucial to a cryptocurrency ecosystem, as the diversity of miners (farms, personal computers, web clients, or even cloud functions) underlays the credibility of the cryptocurrency. Among miners, web clients used to be the vast majority, e.g., 50M+ as of March 2018. As time went on, however, cryptomining was gradually monopolized by mining farms with dedicated hardware (e.g., ASICs), and web clients scaled down to ∼0.1M. To suppress mining farms, certain cryptocurrencies (like Monero) adopted new mining algorithms such as RandomX whose execution relies on general-purpose hardware architectures. Unfortunately, this further impairs web-based cryptomining as web clients cannot provide the desired architecture support to these algorithms. This paper explores how to revive software democracy of efficient web-based crypto-mining, using a novel program transformation technique termed Vectra. Vectra employs just-in-time (JIT) transformations of mining programs for web architectures; it effectively identifies and merges isomorphic instructions upon execution. Vectra ensures correct transformations based on symbolic constraints of the instructions. Real-world deployments show that Vectra reduces WASM instructions by about 7× and achieves a 3× –16× speedup for web cryptomining in diverse execution environments like PCs, mobile phones, and serverless platforms, which translates to a high (69%–274%) return-on-investment (ROI) for common users.
Wei Liu 0148, Zhenhua Li 0001, Feng Qian 0001, Feiyu Jin, Hao Lin 0005, Yannan Zheng, Xiaokang Qin, Tianyin Xu
ASE2
2025 Mitigating Scalability Walls of RDMA-based Container Networks
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Feng Qian 0001, Tianyin Xu, Yunhao Liu 0001, Yu Guan 0005, Shuhong Zhu, Hongfei Xu, Lanlan Xi, Ennan Zhai
NSDI3
2025 Dissecting and Streamlining the Interactive Loop of Mobile Cloud Gaming
Yang Li 0092, Jiaxing Qiu, Hongyi Wang 0009, Zhenhua Li 0001, Feng Qian 0001, Jing Yang 0052, Hao Lin 0005, Yunhao Liu 0001, Xiaokang Qin, Tianyin Xu
NSDI4
2025 Fast and Synchronous Crash Consistency with Metadata Write-Once File System
Yanqi Pan, Wen Xia, Xiangyu Zou, Zhenhua Li 0001, Chentao Wu
OSDI6
2025 SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model Training
abstract
The flexibility and portability characteristics have made containers a popular serverless environment for large model training in recent years. Unfortunately, these advantages render the network support for containerized large model training extremely challenging, due to the high dynamics of containers, the complex interplay between underlay and overlay networks, and the stringent requirements on failure detection and localization. Existing data center network debugging tools, which rely on comprehensive or opportunistic monitoring, are either inefficient or inaccurate in this setting.
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Tianyin Xu, Yunhao Liu 0001, Jiakang Li, Shuhong Zhu, Xue Li 0024, Hongfei Xu, Ennan Zhai
SIGCOMM3
2025 NIER: Practical Neural-enhanced Low-bitrate Video Conferencing
abstract
We present NIER, a video conferencing system that can adaptively maintain a low bitrate (e.g., 10–100 Kbps) with reasonable visual quality while being robust to packet losses. We use key-point-based deep image animation (DIA) as a key building block and address a series of networking and system challenges to make NIER practical. Our evaluations show that NIER significantly outperforms the baseline solutions.
Anlan Zhang, Yuming Hu, Chendong Wang, Yu Liu 0096, Zejun Zhang 0002, Haoyu Gong, Ahmad Hassan 0004, Shichang Xu, Zhenhua Li 0001, Bo Han 0001, Feng Qian 0001
SIGCOMM9
2025 Falcon: A Universal Text-Only Membership Inference Attack Framework Against In-Context Learning
abstract
Membership inference attacks (MIAs) against in-context learning (ICL) serve as essential tools for privacy risk assessment and intellectual property safeguarding due to the use of small, private datasets for adaptation. However, most MIAs against language models require unrealistic, internal access or risk triggering built-in security mechanisms. In this paper, we propose Falcon (Flexible Attack on Language Context via ObfuscatioN), the first task-aware MIA framework against text-only model APIs. Falcon fully exploits the complexity of text obfuscation techniques and leverages the model’s discrepancies in reconstructing obfuscated texts from seen versus unseen data as a strong membership signal, successfully bypassing application constraints and LLM safeguarding mechanisms. Through extensive experiments on six widely used LLMs, including four open-source models (Llama-2, Llama-3, Qwen-2.5, Ministral) and two commercial models (GPT-3.5, and GPT-4o-mini), five datasets from various domains for tasks including question answering, text classification and summarization, Falcon generally achieves over 95% attack success rates, significantly outperforming existing methods. An in-depth analysis of the impact of model scale shows that Falcon exploits a capacity-induced vulnerability, indicating that models with higher capabilities are more susceptible to our attack. Additionally, we explore three defense methods, highlighting role validation as a potential mechanism for safeguarding LLM privacy. We have open-sourced Falcon’s modular, extensible codebase to support future research.
Haitao Su, Zhenhua Li 0001, Yuan Zhou 0007
IEEE Trans. Inf. Forensics Secur.3
2025 A Four-Year Retrospective of Mobile Access Bandwidth Evolution: The Inspiring, the Frustrating, and the Fluctuating
abstract
Recent advances in mobile technologies (like WiFi 6 and 5G) do not seem to deliver the promised access bandwidth. To effectively characterize mobile access bandwidth in the wild, we work with a major commercial mobile bandwidth testing app to conduct a long-term (2020-2023) and large-scale (involving 4.76M users) measurement study in China, based on coarse-grained general statistics and fine-grained sampling diagnostics. Our study presents distinct facts as to WiFi, 5G, and 4G: in the past few years, the average WiFi download bandwidth exhibits a considerable rise (by 119.7% ), the average 5G download bandwidth constantly decreases (by a total of 20.2% ) despite the enormous infrastructure investments, while the average 4G download bandwidth first declines (by 22.1% ) and then increases (by 22.5% ). The situations of upload bandwidths are generally similar to those of download bandwidths, except that 5G upload bandwidths manifestN-shaped$(\nearrow \searrow \nearrow )$fluctuations. Our cross-layer and cross-technology analysis reveals a variety of impact factors as well as their complicated interplay as the root causes, such as the bottlenecks in underlying infrastructure (e.g., communication devices and wired Internet access), the traffic offloading from one access technology to another, the influence of the COVID-19 pandemic, and the side effects of aggressively migrating radio resources from 4G to 5G. With the longitudinal, holistic picture of today's mobile access bandwidth, we finally provide multifold practical implications on closing the technology gaps.
Zhenhua Li 0001, Xinlei Yang, Jing Yang 0052, Xingyao Li, Hao Lin 0005, Feng Qian 0001, Yunhao Liu 0001, Zhi Liao, Daqiang Hu
IEEE Trans. Mob. Comput.1
2025 A Five-Year Retrospective of Cellular Reliability Evolution: The Encouraging, Disappointing, and Further Enhancements
abstract
With recent advances on cellular technologies pushing the boundary of cellular performance, cellular reliability has become a key concern of their adoption and deployment. To fully understand cellular reliability, we work with a major Android phone vendor, Xiaomi, to conduct a long-term (2020-2024) and large-scale (involving 123M users) measurement study in China, with coarse-grained general statistics and fine-grained sampling diagnostics. Our measurement reveals contrasting evolution trends of cellular failures in different stages of the data connection: in the past five years, failures after connection establishment decrease remarkably (by 29%), while failures during connection setup exhibit a sharp increase (by 38%). Our analysis illustrates that the contrast stems from the joint impact of multiple stakeholders, including ISPs’ increasing deployment of 5G base stations, 5G infrastructure upgrade from NSA (Non-Standalone) to SA (Standalone) mode, software defects coming from Android’s adaptation to new cellular technologies, and so forth. Our work provides actionable insights for improving cellular reliability at scale. More importantly, we have built on our insights to develop enhancements that effectively address cellular reliability issues with remarkable real-world impact—our optimizations have reduced 38% cellular connection failures for 5G phones and 31% failure recovery time across all phones.
Yunhao Liu 0001, Hongyi Wang 0009, Yang Li 0092, Zhenhua Li 0001, Guoquan Zhang, Lei Yang 0025
IEEE Trans. Netw.4
2024 Rethinking Process Management for Interactive Mobile Systems
abstract
Modern mobile systems are featured by their increasing interactivity with users, which however is accompanied by a severe side effect---users constantly suffer from slow UI responsiveness (SUR). To date, the community have limited understandings of this issue for the challenges of comprehensively measuring SUR events on massive mobile devices. As a major Android phone vendor, in this paper we close the knowledge gap by conducting the first large-scale, long-term measurement study on SUR with 47M devices. Our study identifies the critical factors that lead to SUR from the perspectives of device, system, application, and app market. Most importantly, we note that the largest root cause lies in the wide existence of "hogging" apps, which persistently occupy an unreasonable amount of system resources by leveraging the optimistic design of Android process management. We have built on the insights to remodel Android process states by fully considering their time-sensitive transitions and the actual behaviors of processes, with remarkable real-world impact---the occurrences of SUR are reduced by 60%, together with 10.7% saving of battery consumption.
Jianwei Zheng 0003, Zhenhua Li 0001, Feng Qian 0001, Wei Liu 0148, Hao Lin 0005, Yunhao Liu 0001, Tianyin Xu, Nan Zhang 0018, Cang Zhang
MobiCom2
2024 MuV2: Scaling up Multi-user Mobile Volumetric Video Streaming via Content Hybridization and Sharing
abstract
Volumetric videos offer a unique interactive experience and have the potential to enhance social virtual reality and telepresence. Streaming volumetric videos to multiple users remains a challenge due to its tremendous requirements of network and computation resources. In this paper, we develop MuV2, an edge-assisted multi-user mobile volumetric video streaming system to support important use cases such as tens of students simultaneously consuming volumetric content in a classroom. MuV2 achieves high scalability and good streaming quality through three orthogonal designs: hybridizing direct streaming of 3D volumetric content with remote rendering, dynamically sharing edge-transcoded views across users, and multiplexing encoding tasks of multiple transcoding sessions into a limited number of hardware encoders on the edge. MuV2 then integrates the three designs into a holistic optimization framework. We fully implement MuV2 and experimentally demonstrate that MuV2 can deliver high-quality volumetric videos to over 30 concurrent untethered mobile devices with a single WiFi access point and a commodity edge server.
Yu Liu 0096, Puqi Zhou, Zejun Zhang 0002, Anlan Zhang, Bo Han 0001, Zhenhua Li 0001, Feng Qian 0001
MobiCom6
2024 vSoC: Efficient Virtual System-on-Chip on Heterogeneous Hardware
abstract
Emerging mobile apps such as UHD video and AR/VR access diverse high-throughput hardware devices, e.g., video codecs, cameras, and image processors. However, today's mobile emulators exhibit poor performance when emulating these devices. We pinpoint the major reason to be the discrepancy between the guest's and host's memory architectures for hardware devices, i.e., the mobile guest's centralized memory on a system-on-chip (SoC) versus the PC/server's separated memory modules on individual hardware. Such a discrepancy makes the shared virtual memory (SVM) architecture of mobile emulators highly inefficient.
Jiaxing Qiu, Yang Li 0092, Zhenhua Li 0001, Feng Qian 0001, Hao Lin 0005, Haitao Su, Yunhao Liu 0001, Tianyin Xu
SOSP4
2024 WiseCam: A Systematic Approach to Intelligent Pan-Tilt Cameras for Moving Object Tracking
abstract
With the desired functionality of moving object tracking, wireless pan-tilt cameras are able to play critical roles in a growing diversity of surveillance environments. However, today's pan-tilt cameras oftentimes underperform when tracking frequently moving objects like humans – they are prone to lose sight of objects and bring about excessive mechanical rotations that are especially detrimental to those energy-constrained outdoor scenarios. The ineffectiveness and high cost of all state-of-the-art tracking approaches are rooted in their adherence to the industry's simplicity principle, which leads to their stateless nature, performing gimbal rotations based only on the latest object detection. To address the issues, we design and implement WiseCam that wisely tunes the pan-tilt cameras to minimize mechanical rotation costs while maintaining long-term object tracking. This systematic tracking approach also tackles issues of motion-rotation speed gap and scattered moving objects, which is universally applicable to complex tracking scenarios. We examine the performance of WiseCam by experiments on two types of pan-tilt cameras with different motors. Results show that it significantly outperforms the state-of-the-art tracking approaches on both tracking duration and power consumption.
Jinlong E, Fangshuo Han, Lin He 0004, Wei Xu 0057, Zhenhua Li 0001, Yunpeng Chai, Yunhao Liu 0001
IEEE Trans. Mob. Comput.5
2024 Aging or Glitching? What Leads to Poor Android Responsiveness and What Can We Do About It?
abstract
Almost all Android users have ever experienced poor responsiveness, including the common frame dropping events—slow rendering (SR) and frozen frames (FF), as well as the uncommon Application Not Responding (ANR) and System Not Responding (SNR) that directly disrupt user experience. This work takes two complementary approaches,controlled benchmarkingandin-the-wild crowdsourcing, to comprehensively understand their prevalence, characteristics, and root causes, which turn out to be significantly different from common understandings and prior studies. We find that SR, FF, ANR, and SNR all occur prevalently on all the studied hardware models of Android phones, and better hardware does not seem to relieve ANR/SNR. Most surprisingly, they are oftentimes ascribed to defective software design that incurs substantial resource overuse—lightweight apps can experience severe SR/FF events due toredundant UI rendering, and the most ANR/SNR events stem from Android's aggressive implementation ofwrite amplification mitigation. In fact, the former can be effectively overcome by simplifying the apps' UI hierarchy, and we design a practical approach to address almost all ($>$99%) of the latter while only decreasing 3% of the data write speed with large-scale deployment. We have released our measurement code/data to the research community.
Hao Lin 0005, Cai Liu, Zhenhua Li 0001, Feng Qian 0001, Yunhao Liu 0001
IEEE Trans. Mob. Comput.3
2024 All-Sky Autonomous Computing in UAV Swarm
abstract
Unmanned aerial vehicles (UAVs) play an essential role in emergency cases and adverse environments for applications like disaster detection and mine exploration. To process the massive volume of sensing data generated by various sensory payloads in these applications, existing works either compress deep learning (DL) models to conduct onboard computing, or offload raw data back to the resourceful ground station with the help of relay UAVs due to base station damage. However, the former sacrifices the inference accuracy of DL models (up to 10% accuracy loss), while the latter achieves high accuracy at the cost of significant latency, due to limited wireless communication resources in the multi-hop transmission. To address the problem, exploiting the resources of the UAV swarm including both task UAVs and relay UAVs, we build up anall-skyautonomous computing (ASAP) system to autonomously conduct collaborative computing in the swarm, to achieve both high accuracy and low latency of sensing data processing. In detail, we first propose a novel UAV swarm-native collaborative computing architecture, considering the general hierarchy and clustering structure of UAV swarms, as well as the characteristic of DL model execution. We then design an elastic efficient task scheduler to allocate computing tasks for UAVs, and update the scheduling scheme online when some UAVs are unavailable, with the aid of a lightweight and accurate DL inference performance predictor. Finally, we design an adaptive inter-UAV data compressor, to adapt to the limited and dynamic communication resources between UAVs. Experiment results on 24 airborne computers and five real-world UAVs show that, the proposed system can perform collaborative computing in a timely manner and effectively deal with situations when some UAVs become unavailable.
Yuben Qu, Chao Dong 0001, Haipeng Dai 0001, Zhenhua Li 0001, Lei Zhang 0038, Qihui Wu 0001, Song Guo 0001
IEEE Trans. Mob. Comput.5
2024 LSTAloc: A Driver-Oriented Incentive Mechanism for Mobility-on-Demand Vehicular Crowdsensing Market
abstract
With the popularity of Mobility-on-Demand (MOD) vehicles, a new market called MOD-Vehicular-Crowdsensing (MOVE-CS) was introduced for drivers to earn more by collecting road data. Unfortunately, MOVE-CS failed after two years of operation. To identify the root cause, we survey 581 drivers and reveal its simple incentive model based on blindly competitive rewards. This model brings most drivers few yields, resulting in their withdrawals. In contrast, a similar market termed MOD-Human-Crowdsensing (MOMAN-CS) remains successful thanks to a complex model based on exclusively customized rewards. Hence, we wonder whether MOVE-CS can be resurrected by learning from MOMAN-CS. Despite considerable similarity, we can hardly apply the incentive model of MOMAN-CS to MOVE-CS, since MOD drivers are also concerned with passenger missions that dominate their earnings. To this end, we analyze a large-scale dataset of 12,493 MOD vehicles, finding that drivers have explicit preference for short-term, immediate gains as well as implicit rationality in pursuit of long-term, stable profits. Therefore, we design a novel driver-oriented incentive mechanism for MOVE-CS, calledLSTAloc, at the heart of which lies a spatial-temporal differentiation-aware task allocation scheme empowered by submodular optimization. Applied to the dataset, our design would essentially benefit both the drivers and platform to incentivize MOD vehicular crowdsensing efficiently, thus possessing the potential to resurrect MOVE-CS.
Chaocan Xiang, Wenhui Cheng, Chi Lin 0001, Xinglin Zhang 0001, Daibo Liu, Zhenhua Li 0001
IEEE Trans. Mob. Comput.7
2024 Trinity: High-Performance and Reliable Mobile Emulation through Graphics Projection
abstract
Mobile emulation, which creates full-fledged software mobile devices on a physical PC/server, is pivotal to the mobile ecosystem. Unfortunately, existing mobile emulators perform poorly on graphics-intensive apps in terms of efficiency and compatibility. To address this, we introduce graphics projection , a novel graphics virtualization mechanism that adds a small-size projection space inside the guest memory, which processes graphics operations involving control contexts and resource handles without host interactions. While enhancing performance, the decoupled and asynchronous guest/host control flows introduced by graphics projection can significantly complicate emulators’ reliability issue diagnosis when faced with a variety of uncommon or non-standard app behaviors in the wild, hindering practical deployment in production. To overcome this drawback, we develop an automatic reliability issue analysis pipeline that distills the critical code paths across the guest and host control flows by runtime quarantine and state introspection. The resulting new Android emulator, dubbed Trinity, exhibits an average of 97% native hardware performance and 99.3% reliable app support, in some cases outperforming other emulators by more than an order of magnitude.
Hao Lin 0005, Zhenhua Li 0001, Yunhao Liu 0001, Feng Qian 0001, Tianyin Xu, Xiaokang Qin
ACM Trans. Comput. Syst.2
2024 Automating Cloud Deployment for Real-Time Online Foundation Model Inference
abstract
Deep neural network (DNN) foundation models are currently exhibiting high prediction accuracy and strong adaptability to broad tasks with remarkably large model scales. They are increasingly becoming the backend support of DNN-driven real-time online services, e.g., Siri and Instagram. Such services require low-latency and cost-efficiency for quality-of-service and commercial competitiveness. When deployed in a cloud environment, these services call for an appropriate selection of cloud configurations (i.e., specific types of VM instances), as well as a considerate device placement plan that places the operations of the model to multiple GPUs via model parallelism for cost-efficiency. Currently, the deployment mainly relies on service providers’ manual efforts, which is not only onerous but also far from satisfactory oftentimes due to the huge joint search space of cloud configurations and device placement plans (for a same service, a poor deployment can incur significantly more costs by tens of times). In this paper, we attempt to efficiently automate the cloud deployment for real-time foundation model inference with minimum costs under the constraint of acceptably low latency. This attempt is enabled by 1) jointly leveraging the Bayesian Optimization and Deep Reinforcement Learning to adaptively unearth the (nearly) optimal cloud configuration and device placement with limited search time, and 2) enhancing the cost-efficiency of the deployment based on the probing-informed block multiplexing mechanism and Tensor Algebra SuperOptimizer. We implement a prototype system based on TensorFlow, conduct extensive experiments on top of Microsoft Azure, and demonstrate the generality and scalability of our solution. Results show that for lightweight DNN models and foundation models, our solution essentially saves inference costs by up to 15% and 47% with 57% and 38% lower search overheads respectively, compared with non-trivial baselines.
Yang Li 0092, Zhenhua Li 0001, Zhenhua Han, Quanlu Zhang, Xiaobo Ma 0001
IEEE/ACM Trans. Netw.2
2024 Website Fingerprinting on Encrypted Proxies: A Flow-Context-Aware Approach and Countermeasures
abstract
Website fingerprinting (WFP) could infer which websites a user is accessing via an encrypted proxy by passively inspecting the traffic characteristics of accessing different websites between the user and the proxy. Designing WFP attacks is crucial for understanding potential vulnerabilities of encrypted proxies, which guides the design of defensive measures against WFP. In this paper, we design a novel WFP attack against (popular) encrypted proxies that relay connections between the user and the proxy individually (e.g., Shadowsocks, V2Ray), and accordingly implement lightweight countermeasures to effectively defend against the attack. The attack features flow-context-aware and is both accurate and immediately deployable, because it fully considers the obstacle (dubbed training-testing asymmetry) that fundamentally limits the practicability of WFP and addresses the obstacle with built-in spatial-temporal flow correlation mechanism. We implement the countermeasure as middleboxes installed on both the client and server sides of encrypted proxies, without altering any existing infrastructures for compatibility. The middleboxes can obfuscate a website’s flow regularities across different visits. Large-scale experiments in real-world scenarios demonstrate that the WFP attack can generally achieve a detection rate above 98.8% with a false positive rate below 0.2%. The countermeasure forces the attack’s false positive rate to be above 0.2 and true positive rate to be below 0.9 with just five persistent TCP connections while introducing very limited bandwidth overhead (e.g., 0.49%) and almost-zero additional network latency.
Xiaobo Ma 0001, Jian Qu, Mawei Shi, Bingyu An, Jianfeng Li 0006, Xiapu Luo, Junjie Zhang 0004, Zhenhua Li 0001, Xiaohong Guan
IEEE/ACM Trans. Netw.8
2024 On Smartly Scanning of the Internet of Things
abstract
Cyber search engines, such as Shodan and Censys, have gained popularity due to their strong capability of indexing the Internet of Things (IoT). They actively scan and fingerprint IoT devices for unearthing IP-device mapping. Because of the large address space of the Internet and the mapping’s mutative nature, efficiently tracking the evolution of IP-device mapping with a limited budget of scans is essential for building timely cyber search engines. An intuitive solution is to use reinforcement learning to schedule more scans to networks with high churn rates of IP-device mapping. However, such an intuitive solution has never been systematically studied. In this paper, we take the first step toward demystifying this problem based on our experiences in maintaining a global IoT scanning platform. Inspired by the measurement study of large-scale real-world IoT scan records, we land reinforcement learning onto a system capable of smartly scanning IoT devices in a principled way. We disclose key parameters affecting the effectiveness of different scanning strategies, and real-world experiments demonstrate that our system can scan up to around 40 times as many IP-device mapping mutations as random/sequential scanning.
Jian Qu, Xiaobo Ma 0001, Wenmao Liu, Hongqing Sang, Jianfeng Li 0006, Lei Xue 0001, Xiapu Luo, Zhenhua Li 0001, Xiaohong Guan
IEEE/ACM Trans. Netw.8
2024 Who Should We Blame for Android App Crashes? An In-Depth Study at Scale and Practical Resolutions
abstract
Android system has been widely deployed in energy-constrained IoT devices for many practical applications, such as smart phone, smart home, healthcare, fitness, and beacons. However, Android users oftentimes suffer from app crashes, which directly disrupt user experience and could lead to data loss. Till now, the community have limited understanding of their prevalence, characteristics, and root causes. In this article, we make an in-depth study of the crash events regarding ten very popular apps of different genres, based on fine-grained system-level traces crowd-sourced from 93 million Android devices. We find that app crashes occur prevalently on the various hardware models studied, and better hardware does not seem to essentially relieve the problem. Most importantly, we unravel multi-fold root causes of app crashes, and pinpoint that the most crashes stem from the subtle yet crucial inconsistency between app developers’ supposed memory/process management model and Android’s actual implementations. We design practical approaches to addressing the inconsistency; after large-scale deployment, they reduce 40.4% of the app crashes with negligible system overhead. In addition, we summarize important lessons learned from this study, and have released our measurement code/data to the community.
Liangyi Gong, Hao Lin 0005, Daibo Liu, Lanqi Yang, Hongyi Wang 0009, Jiaxing Qiu, Zhenhua Li 0001, Feng Qian 0001
ACM Trans. Sens. Networks7
2023 WiseCam: Wisely Tuning Wireless Pan-Tilt Cameras for Cost-Effective Moving Object Tracking
Jinlong E, Lin He 0004, Zhenhua Li 0001, Yunhao Liu 0001
INFOCOM3
2023 ParliRobo: Participant Lightweight AI Robots for Massively Multiplayer Online Games (MMOGs)
abstract
Recent years have witnessed the profound influence of AI technologies on computer gaming. While grandmaster-level AI robots have largely come true for complex games based on heavy back-end support, in practice many game developers crave for participant AI robots (PARs) that behave like average-level humans with inexpensive infrastructures. Unfortunately, to date there has not been a satisfactory solution that registers large-scale use. In this work, we attempt to develop practical PARs (dubbed ParliRobo) showing acceptably humanoid behaviors with well affordable infrastructures under a challenging scenario-a 3D-FPS (first-person shooter) mobile MMOG with real-time interaction requirements. Based on comprehensive real-world explorations, we eventually enable our attempt through a novel ?transform and polish" methodology. It achieves ultralight implementations of the core system components by non-intuitive yet principled approaches, and meanwhile carefully fixes the probable side effect incurred on user perceptions. Evaluation results from large-scale deployment indicate the close resemblance (96% on average) in biofidelity metrics between ParliRobo and human players; moreover, in 73% mini Turing tests ParliRobo cannot be distinguished from human players.
Jianwei Zheng 0003, Changnan Xiao, Zhenhua Li 0001, Feng Qian 0001, Wei Liu 0148
ACM Multimedia4
2023 SkipStreaming: Pinpointing User-Perceived Redundancy in Correlated Web Video Streaming through the Lens of Scenes
abstract
When streaming over the web, correlated videos (e.g., a series of TV episodes) appear to bear considerable redundant clips, mostly included in the intros, outros, recaps, and commercial breaks, leading to a waste of network traffic and playback time. Mainstream video content providers have taken various measures to identify these clips, but often result in unexpected and undesirable user experiences. In this paper, we conduct a large-scale, crowdsourced study to demystify the root causes of poor experiences. Driven by the findings, we propose to reconsider the problem from a novel perspective of scenes without going through the excessive video frames, which pays special attention to how the contents of correlated videos are organized during video production. To enable this idea, we design efficient approaches to the separation of video scenes and the identification of visual redundancy. We build an open-source system to embody our design, which achieves fast (e.g., taking ~38 seconds to process a 45-minute video using a common commodity server) and accurate (incurring only 770-ms deviation on average) redundancy recognition on representative workloads.
Wei Liu 0148, Xinlei Yang, Zhenhua Li 0001, Feng Qian 0001
ACM Multimedia3
2023 Virtual Device Farms for Mobile App Testing at Scale: A Pursuit for Fidelity, Efficiency, and Accessibility
abstract
Virtual devices based on device emulation have been widely used in lab research of mobile app testing for their efficiency and low cost. However, it remains controversial to use virtual devices for app testing in industry, given the inherent difficulties of high-fidelity emulation across diverse mobile systems and devices. Hence, mobile app companies still rely on physical device farms or services like AWS Device Farm.
Hao Lin 0005, Jiaxing Qiu, Hongyi Wang 0009, Zhenhua Li 0001, Liangyi Gong, Yunhao Liu 0001, Feng Qian 0001, Zhao Zhang 0001, Tianyin Xu
MobiCom4
2023 An Input-Agnostic Hierarchical Deep Learning Framework for Traffic Fingerprinting
Jian Qu, Xiaobo Ma 0001, Jianfeng Li 0006, Xiapu Luo, Lei Xue 0001, Junjie Zhang 0004, Zhenhua Li 0001, Xiaohong Guan
USENIX Security Symposium7
2023 Visual-Aware Testing and Debugging for Web Performance Optimization
abstract
Web performance optimization services, or web performance optimizers (WPOs), play a critical role in today’s web ecosystem by improving page load speed and saving network traffic. However, WPOs are known for introducing visual distortions that disrupt the users’ web experience. Unfortunately, visual distortions are hard to analyze, test, and debug, due to their subjective measure, dynamic content, and sophisticated WPO implementations.
Xinlei Yang, Wei Liu 0148, Hao Lin 0005, Zhenhua Li 0001, Feng Qian 0001, Xianlong Wang 0003, Yunhao Liu 0001, Tianyin Xu
WWW4
2023 Memory-efficient Transformer-based network model for Traveling Salesman Problem
Minghao Zhao 0001, Lei Yuan 0005, Yang Yu 0001, Zhenhua Li 0001
Neural Networks5
2023 Fast Uplink Bandwidth Testing for Internet Users
abstract
Access bandwidth measurement is crucial to emerging Internet applications for network-aware content delivery. However, today’s bandwidth testing services (BTSes) are slow and costly—the tests take a long time to run, consume a great deal of data usage, and usually require large-scale test server deployments. The inefficiency and high cost of BTSes root in their methodologies that use excessive temporal/spatial redundancies for combating noises in Internet measurement. In particular, compared to downlink BTSes, uplink BTSes are subject to more severe performance problems and technical challenges. This paper presents FastUpBTS to make uplink BTS fast and cheap while maintaining high accuracy. The key idea is to strategically accommodate and exploit the noise rather than repetitively and exhaustively suppress the impact of noise. This is achieved by a novel statistical sampling framework termed fuzzy rejection sampling. We build FastUpBTS as an end-to-end BTS that implements fuzzy rejection sampling based on memorization-reinforced throughput denoising, data-driven server selection, and informed multi-homing support. Our evaluation shows that with only 30 test servers, FastUpBTS achieves the same level of accuracy compared to the state-of-the-art BTS (SpeedTest.net) that deploys ~16,000 servers. Most importantly, FastUpBTS makes bandwidth tests$5.4\times $faster and$6.8\times $more data-efficient.
Zhenhua Li 0001, Xingyao Li, Xinlei Yang, Xianlong Wang 0003, Feng Qian 0001, Yunhao Liu 0001
IEEE/ACM Trans. Netw.1
2022 Landing Reinforcement Learning onto Smart Scanning of The Internet of Things
abstract
Cyber search engines, such as Shodan and Censys, have gained popularity due to their strong capability of indexing the Internet of Things (IoT). They actively scan and fingerprint IoT devices for unearthing IP-device mapping. Because of the large address space of the Internet and the mapping’s mutative nature, efficiently tracking the evolution of IP-device mapping with a limited budget of scans is essential for building timely cyber search engines. An intuitive solution is to use reinforcement learning to schedule more scans to networks with high churn rates of IP-device mapping. However, such an intuitive solution has never been systematically studied. In this paper, we take the first step toward demystifying this problem based on our experiences in maintaining a global IoT scanning platform. Inspired by the measurement study of large-scale real-world IoT scan records, we land reinforcement learning onto a system capable of smartly scanning IoT devices in a principled way. We disclose key parameters affecting the effectiveness of different scanning strategies, and find that our system would achieve growing advantages with the proliferation of IoT devices.
Jian Qu, Xiaobo Ma 0001, Wenmao Liu, Hongqing Sang, Jianfeng Li 0006, Lei Xue 0001, Xiapu Luo, Zhenhua Li 0001, Xiaohong Guan
INFOCOM8
2022 xNet: Improving Expressiveness and Granularity for Network Modeling with Graph Neural Networks
abstract
Today’s network is notorious for its complexity and uncertainty. Network operators often rely on network models to achieve efficient network planning, operation, and optimization. The network model is responsible for understanding the complex relationships between the network performance metrics (e.g., latency) and the network characteristics (e.g., traffic). However, we still lack a systematic approach to developing accurate and lightweight network models that are aware of the impact of network configurations (i.e., expressiveness) and provide fine-grained flow-level temporal predictions (i.e., granularity).In this paper, we propose xNet, a data-driven network modeling framework based on graph neural networks (GNN). Unlike the previous proposals, xNet is not a dedicated network model designed for specific network scenarios with constraint considerations. On the contrary, xNet provides a general approach to modeling the network characteristics of concern with relation graph representations and configurable GNN blocks. xNet learns the state transition function between time steps and rolls it out to obtain the full fine-grained prediction trajectory. We implement and instantiate xNet with three use cases. The experiment results show that xNet can accurately predict different performance metrics while achieving over two orders of magnitude of speedup compared with the conventional packet-level simulator.
Mowei Wang, Linbo Hui, Yong Cui 0001, Ru Liang, Zhenhua Li 0001
INFOCOM5
2022 Learning Buffer Management Policies for Shared Memory Switches
abstract
Today’s network switches often use on-chip shared memory to improve buffer efficiency and absorb bursty traffic. Current buffer management practices usually rely on simple heuristics and have unrealistic assumptions about the traffic pattern, since developing a buffer management policy suited for every scenario is infeasible. We show that modern machine learning techniques can be of essential help to learn efficient policies automatically.In this paper, we propose Neural Dynamic Threshold (NDT) that uses deep reinforcement learning (RL) to learn buffer management policies without human instructions except for a high-level objective. To tackle the high complexity and scale of the buffer management problem, we develop two domain-specific techniques upon off-the-shelf deep RL solutions. First, we design a scalable RL model by leveraging the permutation symmetry of the switch ports. Second, we use a two-level control mechanism to achieve efficient training and decision-making. The buffer allocation is directly controlled by a low-level heuristic during the decision interval, while the RL agent only decides the high-level control factor according to the traffic density. Testbed and simulation experiments demonstrate that NDT generalizes well and outperforms hand-tuned heuristic policies even on workloads for which it was not explicitly trained.
Mowei Wang, Sijiang Huang, Yong Cui 0001, Wendong Wang 0003, Zhenhua Li 0001
INFOCOM5
2022 A Comparative Approach to Resurrecting the Market of MOD Vehicular Crowdsensing
abstract
With the popularity of Mobility-on-Demand (MOD) vehicles, a new market called MOD-Vehicular-Crowdsensing (MOVE-CS) was introduced for drivers to earn more by collecting road data. Unfortunately, MOVE-CS failed after two years of operation. To identify the root cause, we survey 581 drivers and reveal its simple operation model based on blindly competitive rewards. This model brings most drivers few yields, resulting in their withdrawals. In contrast, a similar market termed MOD-Human-Crowdsensing (MOMAN-CS) remains successful thanks to a complex model based on exclusively customized rewards. Hence, we wonder whether MOVE-CS can be resurrected by learning from MOMAN-CS. Despite considerable similarity, we can hardly apply the operation model of MOMAN-CS to MOVE-CS, since drivers are also concerned with passenger missions that dominate their earnings. To this end, we analyze a large-scale dataset of 12,493 MOD vehicles, finding that drivers have explicit preference for short-term, immediate gains as well as implicit rationality in pursuit of long-term, stable profits. Therefore, we design a novel operation model for MOVE-CS, at the heart of which lies a spatial-temporal differentiation-aware task recommendation scheme empowered by submodular optimization. Applied to the dataset, our design would essentially benefit both the drivers and platform, thus possessing the potential to resurrect MOVE-CS.
Chaocan Xiang, Suining He, Yuben Qu, Zhenhua Li 0001, Liangyi Gong, Chao Chen 0004
INFOCOM6
2022 Experience: practical indoor localization for malls
abstract
We report our experiences of developing, deploying, and evaluating MLoc, a smartphone-based indoor localization system for malls. MLoc uses Bluetooth Low Energy RSSI and geomagnetic field strength as fingerprints. We develop efficient approaches for large-scale, outsourced training data collection. We also design robust online algorithms for localizing and tracking users' positions in complex malls. Since 2018, MLoc has been deployed in 7 cities in China, and used by more than 1 million customers. We conduct extensive evaluations at 35 malls in 7 cities, covering 152K m2 mall areas with a total walking distance of 215 km (1,100 km training data). MLoc yields a median location tracking error of 2.4m. We further characterize the behaviors of MLoc's customers (472K users visiting 12 malls), and demonstrate that MLoc is a promising marketing platform through a promotion event. The e-coupons delivered through MLoc yield an overall conversion rate of 22%. To facilitate future research on mobile sensing and indoor localization, we have released a large dataset (43 GB at the time when this paper was published) that contains IMU, BLE, GMF readings, and the localization ground truth collected by trained testers from 37 shopping malls.
Yuming Hu, Feng Qian 0001, Zhimeng Yin 0001, Zhenhua Li 0001, Yeqiang Han
MobiCom4
2022 Trinity: High-Performance Mobile Emulation through Graphics Projection
Hao Lin 0005, Zhenhua Li 0001, Chengen Huang, Yunhao Liu 0001, Feng Qian 0001, Liangyi Gong, Tianyin Xu
OSDI3
2022 Mobile access bandwidth in practice: measurement, analysis, and implications
abstract
Recent advances in mobile technologies such as 5G and WiFi 6E do not seem to deliver the promised mobile access bandwidth. To effectively characterize mobile access bandwidth in the wild, we work with a major commercial mobile bandwidth testing app to analyze mobile access bandwidths of 3.54M end users in China, based on fine-grained measurement and diagnostic information. Our analysis presents a surprising and frustrating fact---in the past two years, the average WiFi bandwidth remains largely unchanged, while the average 4G/5G bandwidth decreases remarkably. Our analysis further reveals the root causes---the bottlenecks in the underlying infrastructure (e.g., devices and wired Internet access) and side effects of aggressively migrating radio resources from 4G to 5G---with implications on closing the technology gaps. Additionally, our analysis provides insights on building ultra-fast, ultra-light bandwidth testing services (BTSes) at scale. Our new design dramatically reduces the test time of the commercial BTS from 10 seconds to 1 second on average, with a 15× reduction on the backend cost.
Xinlei Yang, Hao Lin 0005, Zhenhua Li 0001, Feng Qian 0001, Xingyao Li, Zhiming He, Xianlong Wang 0003, Yunhao Liu 0001, Zhi Liao, Daqiang Hu, Tianyin Xu
SIGCOMM3
2022 Exploring Potential and Feasibility of Binary Code Sharing in Mobile Computing
abstract
While tremendous growing mobile apps offer users rich services and functionalities, they also bring significant performance and energy issues. Code sharing is promising to address these issues, but existing application-level code sharing is rather restrictive. This paper develops the a transparent machine code sharing for mobile devices, and presents its design, implementation, and deployment. SnapCode enables machine code sharing across a wide variety of commercial off-the-shelf Android devices. By sharing and running machine code, SnapCode can offer significant speed-ups: an average speed-up of 9.9X for one-time trial apps, and up to 120X in apps’ regular uses. In addition, it can save more than 80 percent energy consumption.
Chao Wu 0002, Lan Zhang 0002, Zhenhua Li 0001, Qiushi Li 0002, Yaoxue Zhang
IEEE Trans. Cloud Comput.3
2022 Overlay-Based Android Malware Detection at Market Scales: Systematically Adapting to the New Technological Landscape
abstract
Androidoverlayenables one app to draw over other apps by creating an extraViewlayer atop the hostView, which nevertheless can be exploited by malicious apps (malware) to attack users. To combat this threat, prior countermeasures concentrate on restricting the capabilities of overlays at the OS level while sacrificing overlays’ usability; recently, the overlay mechanism has been substantially updated to prevent a variety of attacks, which however can still be evaded by considerable adversaries. To address these shortcomings, a more pragmatic approach is to enableearly detectionof overlay-based malware during the app market review process, so that all the capabilities of overlays can stay unchanged. For this purpose, in this paper we first conduct a large-scale comparative study of overlay characteristics in benign and malicious apps, and then implement the OverlayChecker system to automatically detect overlay-based malware for one of the world’s largest Android app stores. In particular, we have made systematic efforts in feature engineering, UI exploration, emulation architecture, and run-time environment, thus maintaining high detection accuracy (97 percent precision and 97 percent recall) and short per-app scan time ($\sim$1.7 minutes) with only two commodity servers, under an intensive workload of$\sim$10K newly submitted apps per day.
Liangyi Gong, Zhenhua Li 0001, Hongyi Wang 0009, Hao Lin 0005, Xiaobo Ma 0001, Yunhao Liu 0001
IEEE Trans. Mob. Comput.2
2022 Inferring Hidden IoT Devices and User Interactions via Spatial-Temporal Traffic Fingerprinting
abstract
With the popularization of Internet of Things (IoT) devices in smart home and industry fields, a huge number of IoT devices are connected to the Internet. However, what devices are connected to a network may not be known by the Internet Service Provider (ISP), since many IoT devices are placed within small networks (e.g., home networks) and are hidden behind network address translation (NAT). Without pinpointing IoT devices in a network, it is unlikely for the ISP to appropriately configure security policies and effectively manage the network. Additionally, inferring fine-grained user interactions of IoT devices is also an interesting yet unresolved problem. In this paper, we design an efficient and scalable system via spatial-temporal traffic fingerprinting from an ISP’s perspective in consideration of practical issues like learning-testing asymmetry. Our system can accurately identify typical IoT devices in a network, with the additional capability of identifying what devices are hidden behind NAT and the number of each type of device that share the same IP address. Our system can also detect user interactions and meanwhile identify their (concurrent) number through a multi-output regression model. Through extensive evaluation, we demonstrate that the system can generally identify IoT devices with an F1-Score above 0.999, and estimate the number of the same type of IoT device behind NAT with an average error below 5%. By studying 29 user interactions of 7 devices, we show that our system is promising in detecting user interactions.
Xiaobo Ma 0001, Jian Qu, Jianfeng Li 0006, John C. S. Lui, Zhenhua Li 0001, Wenmao Liu, Xiaohong Guan
IEEE/ACM Trans. Netw.5
2022 WebAssembly-based Delta Sync for Cloud Storage Services
abstract
Delta synchronization (sync) is crucial to the network-level efficiency of cloud storage services, especially when handling large files with small increments. Practical delta sync techniques are, however, only available for PC clients and mobile apps, but not web browsers—the most pervasive and OS-independent access method. To bridge this gap, prior work concentrates on either reversing the delta sync protocol or utilizing the native client, all striving around the tradeoffs among efficiency, applicability, and usability and thus forming an “impossible triangle.” Recently, we note the advent of WebAssembly (WASM) , a portable binary instruction format that is efficient in both encoding size and load time. In principle, the unique advantages of WASM can make web-based applications enjoy near-native runtime speed without significant cloud-side or client-side changes. Thus, we implement a straightforward WASM-based delta sync solution, WASMrsync, finding its quasi-asynchronous working manner and conventional In-situ Separate Memory Allocation greatly increase sync time and memory usage. To address them, we strategically devise sync-async code decoupling and streaming compilation, together with Informed In-place File Construction. The resulting solution, WASMrsync+, achieves comparable sync time as the state-of-the-art (most efficient) solution with nearly only half of memory usage, letting the “impossible triangle” reach a reconciliation.
Jianwei Zheng 0003, Zhenhua Li 0001, Yuanhui Qiu, Hao Lin 0005, Yang Li 0092, Yunhao Liu 0001
ACM Trans. Storage2
2022 NetSync: A Network Adaptive and Deduplication-Inspired Delta Synchronization Approach for Cloud Storage Services
abstract
Delta sync (synchronization) is a key bandwidth-saving technique for cloud storage services. The representative delta sync utility,rsync, matches data chunks by sliding a search window byte-by-byte to maximize the redundancy detection for bandwidth efficiency. However, it is difficult for this process to cater to the forthcoming high-bandwidth cloud storage services which require lightweight delta sync that can well support large files. Moreover,rsyncemploys invariant chunking and compression methods during the sync process, making it unable to cater to services from various network environments which require the sync approach to perform well under different network conditions. Inspired by the Content-Defined Chunking (CDC) technique used in data deduplication, we propose NetSync, a network adaptive and CDC-based lightweight delta sync approach with less computing and protocol (metadata) overheads than the state-of-the-art delta sync approaches. Besides, NetSync can choose appropriate compressing and chunking strategies for different network conditions. The key idea of NetSync is (1) to simplify the process of chunk matching by proposing a fast weak hash called FastFP that is piggybacked on the rolling hashes from CDC, and redesigning the delta sync protocol by exploiting deduplication locality and weak/strong hash properties; (2) to minimize the sync time by adaptively choosing chunking parameters and compression methods according to the current network conditions. Our evaluation results driven by both benchmark and real-world datasets suggest NetSync performs$2\times$2×–$10\times$10×faster and supports$30\%$30%–$80\%$80%more clients than the state-of-the-artrsync-based WebR2sync+ and deduplication-based approach.
Wen Xia, Can Wei, Zhenhua Li 0001, Xuan Wang 0002, Xiangyu Zou
IEEE Trans. Parallel Distributed Syst.3
2022 UFC2: User-Friendly Collaborative Cloud
abstract
This article studies how today's cloud storage services support collaborative file editing. As a tradeoff for transparency and user-friendliness, they do not ask collaborators to use version control systems but instead implement their own heuristics for handling conflicts, which however often lead to unexpected and undesired experiences. With specialized measurements and reverse engineering, we unravel a number of their design and implementation issues as the root causes of poor experiences. Driven by the findings, we propose to reconsider the collaboration support of cloud storage services from a novel perspective ofoperationswithout using any locks. To enable this idea, we design intelligent and efficient approaches to the inference and transformation of users’ editing operations, as well as optimizations to the maintenance of files’ historical versions and the update of individual files. We build an open-source system UFC2 (User-Friendly Collaborative Cloud) to embody our design, which can avoid most (98%) conflicts with little (2%) overhead.
Minghao Zhao 0001, Zhenhua Li 0001, Wei Liu 0148, Xingyao Li
IEEE Trans. Parallel Distributed Syst.2
2021 Enabling Conflict-free Collaborations with Cloud Storage Services
abstract
Cloud storage services (e.g., Dropbox) have become pervasive in not only simple file sharing but also advanced collaborative file editing (collaboration for short). Using Dropbox for collaboration is much easier than SVN and Git, thus greatly facilitating common users. In practice, however, many Dropbox users are perplexed by unexpected collaboration conflicts, which severely impair their experiences. Through various benchmark experiments, we unveil the two root causes of collaboration conflicts: 1) Dropbox never locks an edited file during collaboration; 2) Dropbox only guarantees eventual data consistency among the collaborators, significantly aggravating the probability of conflicts. In this paper, we attempt to enable conflict-free collaborations with Dropbox-like cloud storage services. This attempt is empowered by three key findings and measures. First, although the end-to-end sync delay is unpredictable due to eventual consistency, we can always track the latest version of an edited file by actively resorting to the cloud via certain web APIs. Second, although all application-level data is encrypted in Dropbox, we can roughly deduce the sync status from traffic statistics. Third, applying a couple of useful mechanisms (e.g., distributed architecture and data lock) learned from Git, we can effectively and efficiently avoid collaboration conflicts-of course, this requires re-implementing Git mechanisms in cloud storage services with minimum overhead and user interference. Integrating above efforts, we build the ConflictReaper system capable of helping users automatically avoid almost all collaboration conflicts with affordable network and computation overhead.
Minghao Zhao 0001, Zhenhua Li 0001
ICPADS3
2021 A nationwide census on wifi security threats: prevalence, riskiness, and the economics
abstract
Carrying over 75% of the last-mile mobile Internet traffic, WiFi has inevitably become an enticing target for various security threats. In this work, we characterize a wide variety of real-world WiFi threats at an unprecedented scale, involving 19 million WiFi access points (APs) mostly located in China, by deploying a crowdsourced security checking system on 14 million mobile devices in the wild. Leveraging the collected data, we reveal the landscape of nationwide WiFi threats for the first time. We find that the prevalence, riskiness, and breakdown of WiFi threats deviate significantly from common understandings and prior studies. In particular, we detect attacks at around 4% of all WiFi APs, uncover that most WiFi attacks are driven by an underground economy, and provide strong evidence of web analytics platforms being the bottleneck of its monetization chain. Further, we provide insightful guidance for defending against WiFi attacks at scale, and some of our efforts have already yielded real-world impact---effectively disrupted the WiFi attack ecosystem.
Hao Lin 0005, Zhenhua Li 0001, Feng Qian 0001, Qi Alfred Chen, Zhiyun Qian, Wei Liu 0148, Liangyi Gong, Yunhao Liu 0001
MobiCom3
2021 Fast and Light Bandwidth Testing for Internet Users
Xinlei Yang, Xianlong Wang 0003, Zhenhua Li 0001, Yunhao Liu 0001, Feng Qian 0001, Liangyi Gong, Tianyin Xu
NSDI3
2021 A nationwide study on cellular reliability: measurement, analysis, and enhancements
abstract
With recent advances on cellular technologies (such as 5G) that push the boundary of cellular performance, cellular reliability has become a key concern of cellular technology adoption and deployment. However, this fundamental concern has never been addressed due to the challenges of measuring cellular reliability on mobile devices and the cost of conducting large-scale measurements. This paper closes the knowledge gap by presenting the first large-scale, in-depth study on cellular reliability with more than 70 million Android phones across 34 different hardware models. Our study identifies the critical factors that affect cellular reliability and clears up misleading intuitions indicated by common wisdom. In particular, our study pinpoints that software reliability defects are among the main root causes of cellular data connection failures. Our work provides actionable insights for improving cellular reliability at scale. More importantly, we have built on our insights to develop enhancements that effectively address cellular reliability issues with remarkable real-world impact---our optimizations on Android's cellular implementations have effectively reduced 40% cellular connection failures for 5G phones and 36% failure duration across all phones.
Yang Li 0092, Hao Lin 0005, Zhenhua Li 0001, Yunhao Liu 0001, Feng Qian 0001, Liangyi Gong, Xianlong Xin, Tianyin Xu
SIGCOMM3
2021 Systematically Landing Machine Learning onto Market-Scale Mobile Malware Detection
abstract
Despite being crucial to today's mobile ecosystem, app markets have meanwhile become a natural, convenient malware delivery channel as they actually “lend credibility” to malicious apps. In the past few years, machine learning (ML) techniques have been widely explored for automated, robust malware detection, but till now we have not seen an ML-based malware detection solution applied at market scales. To systematically understand the real-world challenges, we conduct a collaborative study with T-Market, a popular Android app market that offers us large-scale ground-truth data. Our study illustrates that the key to successfully developing such systems is multifold, including feature selection and encoding, feature engineering and exposure, app analysis speed and efficacy, developer and user engagement, as well as ML model evolution. Failure in any of the above aspects could lead to the “wooden barrel effect” of the whole system. This article presents our judicious design choices and first-hand deployment experiences in building a practical ML-powered malware detection system. It has been operational at T-Market, using a single commodity server to check ~12K apps every day, and has achieved an overall precision of 98.9 percent and recall of 98.1 percent with an average per-app scan time of 0.9 minutes.
Liangyi Gong, Hao Lin 0005, Zhenhua Li 0001, Feng Qian 0001, Yang Li 0092, Xiaobo Ma 0001, Yunhao Liu 0001
IEEE Trans. Parallel Distributed Syst.3
2020 Experiences of landing machine learning onto market-scale mobile malware detection
abstract
App markets, being crucial and critical for today's mobile ecosystem, have also become a natural malware delivery channel since they actually "lend credibility" to malicious apps. In the past decade, machine learning (ML) techniques have been explored for automated, robust malware detection. Unfortunately, to date, we have yet to see an ML-based malware detection solution deployed at market scales. To better understand the real-world challenges, we conduct a collaborative study with a major Android app market (T-Market) offering us large-scale ground-truth data. Our study shows that the key to successfully developing such systems is manifold, including feature selection/engineering, app analysis speed, developer engagement, and model evolution. Failure in any of the above aspects would lead to the "wooden barrel effect" of the entire system. We discuss our careful design choices as well as our first-hand deployment experiences in building such an ML-powered malware detection system. We implement our design and examine its effectiveness in the T-Market for over one year, using a single commodity server to vet ~ 10K apps every day. The evaluation results show that this design achieves an overall precision of 98% and recall of 96% with an average per-app scan time of 1.3 minutes.
Liangyi Gong, Zhenhua Li 0001, Feng Qian 0001, Qi Alfred Chen, Zhiyun Qian, Hao Lin 0005, Yunhao Liu 0001
EuroSys2
2020 Lock-Free Collaboration Support for Cloud Storage Services with Operation Inference and Transformation
Minghao Zhao 0001, Zhenhua Li 0001, Ennan Zhai, Feng Qian 0001, Yunhao Liu 0001, Tianyin Xu
FAST3
2020 Automating Cloud Deployment for Deep Learning Inference of Real-time Online Services
abstract
Real-time online services using pre-trained deep neural network (DNN) models, e.g., Siri and Instagram, require low-latency and cost-efficiency for quality-of-service and commercial competitiveness. When deployed in a cloud environment, such services call for an appropriate selection of cloud configurations (i.e., specific types of VM instances), as well as a considerate device placement plan that places the operations of a DNN model to multiple computation devices like GPUs and CPUs. Currently, the deployment mainly relies on service providers' manual efforts, which is not only onerous but also far from satisfactory oftentimes (for a same service, a poor deployment can incur significantly more costs by tens of times). In this paper, we attempt to automate the cloud deployment for real-time online DNN inference with minimum costs under the constraint of acceptably low latency. This attempt is enabled by jointly leveraging the Bayesian Optimization and Deep Reinforcement Learning to adaptively unearth the (nearly) optimal cloud configuration and device placement with limited search time. We implement a prototype system of our solution based on TensorFlow and conduct extensive experiments on top of Microsoft Azure. The results show that our solution essentially outperforms the nontrivial baselines in terms of inference speed and cost-efficiency.
Yang Li 0092, Zhenhua Han, Quanlu Zhang, Zhenhua Li 0001, Haisheng Tan
INFOCOM4
2020 Pinpointing Hidden IoT Devices via Spatial-temporal Traffic Fingerprinting
abstract
With the popularization of Internet of Things (IoT) devices in smart home and industry fields, a huge number of IoT devices are connected to the Internet. However, what devices are connected to a network may not be known by the Internet Service Provider (ISP), since many IoT devices are placed within small networks (e.g., home networks) and are hidden behind network address translation (NAT). Without pinpointing IoT devices in a network, it is unlikely for the ISP to appropriately configure security policies and effectively manage the network. In this paper, we design an efficient and scalable system via spatial-temporal traffic fingerprinting. Our system can accurately identify typical IoT devices in a network, with the additional capability of identifying what devices are hidden behind NAT and how many they are. Through extensive evaluation, we demonstrate that the system can generally identify IoT devices with an F-Score above 0.999, and estimate the number of the same type of IoT device behind NAT with an average error below 5%. We also perform small-scale (labor-intensive) experiments to show that our system is promising in detecting user-IoT interactions.
Xiaobo Ma 0001, Jian Qu, Jianfeng Li 0006, John C. S. Lui, Zhenhua Li 0001, Xiaohong Guan
INFOCOM5
2020 Experience: aging or glitching? why does android stop responding and what can we do about it?
abstract
Almost every Android user has unsatisfying experiences regarding responsiveness, in particular Application Not Responding (ANR) and System Not Responding (SNR) that directly disrupt user experience. Unfortunately, the community have limited understanding of the prevalence, characteristics, and root causes of unresponsiveness. In this paper, we make an in-depth study of ANR and SNR at scale based on fine-grained system-level traces crowdsourced from 30,000 Android systems. We find that ANR and SNR occur prevalently on all the studied 15 hardware models, and better hardware does not seem to relieve the problem. Moreover, as Android evolves from version 7.0 to 9.0, there are fewer ANR events but more SNR events. Most importantly, we uncover multifold root causes of ANR and SNR and pinpoint the largest inefficiency which roots in Android's flawed implementation of Write Amplification Mitigation (WAM). We design a practical approach to eliminating this largest root cause; after large-scale deployment, it reduces almost all (>99%) ANR and SNR caused by WAM while only decreasing 3% of the data write speed. In addition, we document important lessons we have learned from this study, and have also released our measurement code/data to the research community.
Hao Lin 0005, Cai Liu, Zhenhua Li 0001, Feng Qian 0001, Yunhao Liu 0001, Nian Xiang Sun, Tianyin Xu
MobiCom4
2020 Randomized Security Patrolling for Link Flooding Attack Detection
abstract
With the advancement of large-scale coordinated attacks, the adversary is shifting away from traditional distributed denial of service (DDoS) attacks against servers to sophisticated DDoS attacks against Internet infrastructures. Link flooding attacks (LFAs) are such powerful attacks against Internet links. Employing network measurement techniques, the defender could detect the link under attack. However, given the large number of Internet links, the defender can only monitor a subset of the links simultaneously, whereas any link might be attacked. Therefore, it remains challenging to practically deploy detection methods. This paper addresses this challenge from a game-theoretic perspective, and proposes a randomized approach (like security patrolling) to optimize LFA detection strategies. Specifically, we formulate the LFA detection problem as a Stackelberg security game, and design randomized detection strategies in consideration of the adversary's behavior, where best and quantal response models are leveraged to characterize the adversary's behavior. We employ a series of techniques to solve the nonlinear and nonconvex NP-hard optimization problems for finding the equilibrium. The experimental results demonstrate the necessity of handling LFAs from a game-theoretic perspective and the effectiveness of our solutions. We believe our study is a significant step forward in formally understanding LFA detection strategies.
Xiaobo Ma 0001, Bo An 0001, Mengchen Zhao, Xiapu Luo, Lei Xue 0001, Zhenhua Li 0001, Tony T. N. Miu, Xiaohong Guan
IEEE Trans. Dependable Secur. Comput.6
2020 Look-Aside at Your Own Risk: Privacy Implications of DNSSEC Look-Aside Validation
abstract
The Domain Name System Security Extension (DNSSEC) leverages public-key cryptography to provide data integrity, source authentication, and denial of existence for DNS responses. To complement DNSSEC operations, DNSSEC Look-aside Validation (DLV) is designed for alternative off-path validation. Although DNS privacy attracts a lot of attention, the privacy implications of DLV are not fully investigated and understood. In this paper, we take a first in-depth look into DLV, highlighting its lax specifications and privacy implications. By performing extensive experiments over datasets of domain names under comprehensive experimental settings, our findings firmly confirm the privacy leakages caused by DLV. We discover that a large number of domains that should not be sent to DLV servers are being leaked. We explore the root causes, including the lax specifications of DLV. We also propose two approaches to fix the privacy leakages. Our approaches require trivial modifications to the existing DNS standards, and we demonstrate their cost in terms of latency and communication.
David Mohaisen, Zhongshu Gu, Kui Ren 0001, Zhenhua Li 0001, Charles A. Kamhoua, Laurent Njilla, DaeHun Nyang
IEEE Trans. Dependable Secur. Comput.4
2020 HyCloud: Tweaking Hybrid Cloud Storage Services for Cost-Efficient Filesystem Hosting
abstract
Today's cloud storage infrastructures typically provide two distinct types of services for hosting files: object storage like Amazon S3 and filesystem storage like Amazon EFS. In practice, a cloud storage user often desires the advantages of both-efficient filesystem operations with a low unit storage price. An intuitive approach to achieving this goal is to combine the two types of services, e.g., by hosting large files in S3 and small files together with directory structures in EFS. Unfortunately, our benchmark experiments indicate that the clients' download performance for large files becomes a severe system bottleneck. In this article, we attempt to address the bottleneck with little overhead by carefully tweaking the usages of S3 and EFS. Guided by two key observations, we design and implement an open-source system called HyCloud. It automatically invokes the data APIs of S3 and EFS on behalf of users, and intelligently schedules the data transfer among S3, EFS and the clients in a distributed manner. Real-world evaluations demonstrate that the unit storage price of HyCloud is close to that of S3, and the filesystem operations are executed as quickly as in EFS in most times (sometimes even more quickly than in EFS).
Jinlong E, Yong Cui 0001, Zhenhua Li 0001, Mingkang Ruan, Ennan Zhai
IEEE/ACM Trans. Netw.3
2020 Understanding the Ecosystem and Addressing the Fundamental Concerns of Commercial MVNO
abstract
Recent years have witnessed the rapid growth of mobile virtual network operators (MVNOs), which operate on top of existing cellular infrastructures of base carriers, while offering cheaper or more flexible data plans compared to those of the base carriers. In this paper, we present a two-year measurement study towards understanding various fundamental aspects of today's MVNO ecosystem, including its architecture, customers, performance, economics, and the complex interplay with the base carrier. Our study focuses on a large commercial MVNO with one million customers, operating atop a nation-wide base carrier. Our measurements clarify several key concerns raised by MVNO customers, such as inaccurate billing and potential performance discrimination with the base carrier. We also leverage big data analytics, statistical modeling, and machine learning to address the MVNO's key concerns with regard to data usage prediction, data plan reselling, customer churn mitigation, and billing delay reduction. Our proposed techniques can help achieve higher revenues and improved services for commercial MVNOs.
Yang Li 0092, Jianwei Zheng 0003, Zhenhua Li 0001, Yunhao Liu 0001, Feng Qian 0001, Sen Bai, Yao Liu 0001, Xianlong Xin
IEEE/ACM Trans. Netw.3
2019 The Cask Effect of Multi-source Content Delivery: Measurement and Mitigation
abstract
With the explosive growth of Internet traffic, multi-source content delivery has been introduced for improving the performance and quality-of-experience (QoE) of Internet services. Upgrading from single-source content delivery to multi-source content delivery, however, may not always lead to a better performance. Instead, a decline in terms of delivery speed often occurs. By conducting a comprehensive study, we show that the underlying reason of this counter-intuitive phenomenon is actually due to the cask effect of data sources at both macro and micro level. Specifically, at the macro level, data sources with different types are highly heterogeneous in terms of delivery performance, which means data sources with certain types are particularly easy to become the "short boards". At the micro level, for the data sources chosen by a client, the high diversity of participation time (DPT) of the sources could impair the acceleration effect. Motivated by the above findings, we design MDR (Multi-source Delivery Redirector), a middleware that contains two optimizations to improve the acceleration effect. One is the feature-greedy selection algorithm which can avoid selecting data sources with inferior types, and the other is the DPT-driven shuffle strategy which can avoid using unstable data sources. Simulation-based experiments show that the MDR outperforms existing approaches in terms of overall downloading performance.
Minghao Zhao 0001, Xinlei Yang, Zhenhua Li 0001, Yao Liu 0001, Zhenyu Li 0001, Yunhao Liu 0001
ICDCS4
2019 HyCloud: Tweaking Hybrid Cloud Storage Services for Cost-Efficient Filesystem Hosting
abstract
Today's cloud storage infrastructures typically provide two distinct types of services for hosting files: object storage like Amazon S3 and filesystem storage like Amazon EFS. The former supports simple, flat object operations with a low unit storage price, while the latter supports complex, hierarchical filesystem operations with a high unit storage price. In practice, however, a cloud storage user often desires the advantages of both-efficient filesystem operations with a low unit storage price. An intuitive approach to achieving this goal is to combine the two types of services, e.g., by hosting large files in S3 and small files together with directory structures in EFS. Unfortunately, our benchmark experiments indicate that the clients' download performance for large files becomes a severe system bottleneck. In this paper, we attempt to address the bottleneck with little overhead by carefully tweaking the usages of S3 and EFS. This attempt is enabled by two key observations. First, since S3 and EFS have the same unit network-traffic price and the data transfer between S3 and EFS is free of charge, we can employ EFS as a relay for the clients' quickly downloading large files. Second, noticing that significant similarity exists between the files hosted at the cloud and its users, in most times we can convert large-size file downloads into small-size file synchronizations (through delta encoding and data compression). Guided by the observations, we design and implement an open-source system called HyCloud. It automatically invokes the data APIs of S3 and EFS on behalf of users, and handles the data transfer among S3, EFS and the clients. Real-world evaluations demonstrate that the unit storage price of HyCloud is close to that of S3, and the filesystem operations are executed as quickly as in EFS in most times (sometimes even more quickly than in EFS).
Jinlong E, Yong Cui 0001, Mingkang Ruan, Zhenhua Li 0001, Ennan Zhai
INFOCOM4
2019 Companion Paper for
abstract
This artifact includes source code, scripts and datasets required to reproduce the experimental figures in the evaluation of the MM'18 paper, which is entitled "MiniView Layout for Bandwidth-Efficient 360-Degree Video". The artifact reports the comparison results among the standard cube layout (CUBE), the equi-angular layout (EAC), and the MiniView layout (MVL) in terms of compressed video size, visual quality of views and decoding and rendering time.
Mengbai Xiao, Shuoqian Wang, Chao Zhou 0004, Li Liu 0045, Zhenhua Li 0001, Yao Liu 0001, Songqing Chen, Lucile Sassatelli, Gwendal Simon
ACM Multimedia5
2019 Mobile Gaming on Personal Computers with Direct Android Emulation
abstract
Playing Android games on Windows x86 PCs has gained enormous popularity in recent years, and the de facto solution is to use mobile emulators built with the AOVB (Android-x86 On VirtualBox) architecture. When playing heavy 3D Android games with AOVB, however, users often suffer unsatisfactory smoothness due to the considerable overhead of full virtualization. This paper presents DAOW, a game-oriented Android emulator implementing the idea of direct Android emulation, which eliminates the overhead of full virtualization by directly executing Android app binaries on top of x86-based Windows. Based on pragmatic, efficient instruction rewriting and syscall emulation, DAOW offers foreign Android binaries direct access to the domestic PC hardware through Windows kernel interfaces, achieving nearly native hardware performance. Moreover, it leverages graphics and security techniques to enhance user experiences and prevent cheating in gaming. As of late 2018, DAOW has been adopted by over 50 million PC users to run thousands of heavy 3D Android games. Compared with AOVB, DAOW improves the smoothness by 21% on average, decreases the game startup time by 48%, and reduces the memory usage by 22%.
Zhenhua Li 0001, Yunhao Liu 0001, Hai Long, Yuanchao Huang, Jiaming He, Tianyin Xu, Ennan Zhai
MobiCom2
2019 Demo: Mobile Gaming on Personal Computers with Direct Android Emulation
abstract
Playing Android games with Windows x86 PCs is now popular, and the common solution is to use mobile emulators built with the AOVB (Android-x86 On VirtualBox) architecture. Nevertheless, running heavy 3D Android games on AOVB incurs considerable overhead of full virtualization, thus often leading to unsatisfactory smoothness. To tackle this issue, we present DAOW, a commercial game-oriented Android emulator implementing the idea of direct Android emulation, which eliminates the overhead of full virtualization by providing foreign Android binaries with direct access to the domestic PC hardware through Windows kernel interfaces. In this demo, we will demonstrate that DAOW essentially outperforms traditional AOVB-based emulators in terms of running smoothness, game startup time, and memory usage.
Xinlei Yang, Zhenhua Li 0001, Yunhao Liu 0001, Guoyang Du, Ziwen Wu, Tianyin Xu, Ennan Zhai
MobiCom3
2019 Understanding Fileless Attacks on Linux-based IoT Devices with HoneyCloud
abstract
With the wide adoption, Linux-based IoT devices have emerged as one primary target of today's cyber attacks. Traditional malware-based attacks can quickly spread across these devices, but they are well-understood threats with effective defense techniques such as malware fingerprinting and community-based fingerprint sharing. Recently, fileless attacks---attacks that do not rely on malware files---have been increasing on Linux-based IoT devices, and posing significant threats to the security and privacy of IoT systems. Little has been known in terms of their characteristics and attack vectors, which hinders research and development efforts to defend against them. In this paper, we present our endeavor in understanding fileless attacks on Linux-based IoT devices in the wild. Over a span of twelve months, we deploy 4 hardware IoT honeypots and 108 specially designed software IoT honeypots, and successfully attract a wide variety of real-world IoT attacks. We present our measurement study on these attacks, with a focus on fileless attacks, including the prevalence, exploits, environments, and impacts. Our study further leads to multi-fold insights towards actionable defense strategies that can be adopted by IoT vendors and end users.
Fan Dang 0001, Zhenhua Li 0001, Yunhao Liu 0001, Ennan Zhai, Qi Alfred Chen, Tianyin Xu, Yan Chen 0004
MobiSys2
2019 An In-depth Study of Commercial MVNO: Measurement and Optimization
abstract
Recent years have witnessed the rapid growth of mobile virtual network operators (MVNOs), which operate on top of the existing cellular infrastructures of base carriers while offering cheaper or more flexible data plans compared to those of the base carriers. In this paper, we present a nearly two-year measurement study towards understanding various key aspects of today's MVNO ecosystem, including its architecture, performance, economics, customers, and the complex interplay with the base carrier. Our study focuses on a large commercial MVNO with \reviseabout 1 million customers, operating atop a nation-wide base carrier. Our measurements clarify several key concerns raised by MVNO customers, such as inaccurate billing and potential performance discrimination with the base carrier. We also leverage big data analytics and machine learning to optimize an MVNO's key businesses such as data plan reselling and customer churn mitigation. Our proposed techniques can help achieve %will lead to higher revenues and improved services for commercial MVNOs.
Ao Xiao, Yunhao Liu 0001, Yang Li 0092, Feng Qian 0001, Zhenhua Li 0001, Sen Bai, Yao Liu 0001, Tianyin Xu, Xianlong Xin
MobiSys5
2019 Understanding and Detecting Overlay-based Android Malware at Market Scales
abstract
As a key UI feature of Android, overlay enables one app to draw over other apps by creating an extra View layer on top of the host View. While greatly facilitating user interactions with multiple apps at the same time, it is often exploited by malicious apps (malware) to attack users. To combat this threat, prior countermeasures concentrate on restricting the capabilities of overlays at the OS level, while barely seeing adoption by Android due to the concern of sacrificing overlays' usability. To address this dilemma, a more pragmatic approach is to enable the early detection of overlay-based malware at the app market level during the app review process, so that all the capabilities of overlays can stay unchanged. Unfortunately, little has been known about the feasibility and effectiveness of this approach for lack of understanding of malicious overlays in the wild. To fill this gap, in this paper we perform the first large-scale comparative study of overlay characteristics in benign and malicious apps using static and dynamic analyses. Our results reveal a set of suspicious overlay properties strongly correlated with the malice of apps, including several novel features. Guided by the study insights, we build OverlayChecker, a system that is able to automatically detect overlay-based malware at market scales. OverlayChecker has been adopted by one of the world's largest Android app stores to check around 10K newly submitted apps per day. It can efficiently (within 2 minutes per app) detect nearly all (96%) overlay-based malware using a single commodity server.
Yuxuan Yan, Zhenhua Li 0001, Qi Alfred Chen, Christo Wilson, Tianyin Xu, Ennan Zhai, Yong Li 0008, Yunhao Liu 0001
MobiSys2
2019 Detecting wormhole attacks in 3D wireless ad hoc networks via 3D forbidden substructures
Sen Bai, Yunhao Liu 0001, Zhenhua Li 0001, Xin Bai 0004
Comput. Networks3
2019 Pricing Data Tampering in Automated Fare Collection with NFC-Equipped Smartphones
abstract
Automated Fare Collection (AFC) systems have been globally deployed for decades, particularly in the public transportation network where the transit fee is calculated based on the length of the trip (a.k.a., distance-based pricing AFC systems). Although most messages of AFC systems are insecurely transferred in plaintext, system operators did not pay much attention to this vulnerability, since the AFC network is basically isolated from the public network (e.g., the Internet)-there is no way of exploiting such a vulnerability from the outside of the AFC network. Nevertheless, in recent years, the advent of Near Field Communication (NFC)-equipped smartphones has opened up a channel to invade into the AFC network from the mobile Internet, i.e., by Host-based Card Emulation (HCE) over NFC-equipped smartphones. In this paper, we identify a novel paradigm of attacks, called LessPay, against modern distance-based pricing AFC systems, enabling users to pay much less than what they are supposed to be charged. The identified attack has two important properties: 1) it is invisible to AFC system operators because the attack never causes any inconsistency in the back-end database of the operators; and 2) it can be scalable to affect a large number of users (e.g., 10,000) by only requiring a moderate-sized AFC card pool (e.g., containing 150 cards). To evaluate the efficacy of the attack, we developed an HCE app to launch the LessPay attack; and the real-world experiments demonstrate not only the feasibility of the LessPay attack (with 97.6 percent success rate) but also its low cost in terms of bandwidth and computation. Finally, we propose, implement and evaluate four types of countermeasures, and present security analysis and comparison of these countermeasures on defending against the LessPay attack.
Fan Dang 0001, Ennan Zhai, Zhenhua Li 0001, David Mohaisen, Kaigui Bian, Qingfu Wen, Mo Li 0001
IEEE Trans. Mob. Comput.3
2019 Robust Light-Weight Magnetic-Based Door Event Detection with Smartphones
abstract
Doors as densely-deployed natural landmarks play an important role in improving indoor positioning systems. However, the state-of-the-art door event detection works are based on either vision or infrastructure, thus incurring non-trivial device or management cost. To address these problems, we present a Light-weight Magnetic-based Door Event Detection method, called LMDD. It leverages built-in magnetic sensors of common smartphones to achieve infrastructure-free door event detection. After analyzing the special features of sensors' readings changes caused by the door, we design LMDD scheme with three main components, including data acquisition, events identification and events denoising. Moreover, an improved and robust door event detection framework based on a majority-voting model is proposed to fuse multiple-dimensional sensing data from non-magnetic built-in sensors. We have implemented a prototype of LMDD on Android-based platform. Experimental results show that LMDD with only magnetic sensor achieves door event detection accuracy of around 80 percent on average, ranging from 70 to 87 percent in various typical indoor environments. The enhanced LMDD based on the fusion of heterogeneous sensors can achieve a much higher door event detection accuracy of 90 percent on average.
Liangyi Gong, Chaocan Xiang, Zhenhua Li 0001, Chen Qian 0001, Panlong Yang
IEEE Trans. Mob. Comput.4
2019 On Designing Distributed Auction Mechanisms for Wireless Spectrum Allocation
abstract
Auctions are believed to be effective methods to solve the problem of wireless spectrum allocation. Existing spectrum auction mechanisms are all centralized and suffer from several critical drawbacks of the centralized systems, which motivates the design of distributed spectrum auction mechanisms. However, extending a centralized spectrum auction to a distributed one broadens the strategy space of agents from one dimension (bid) to three dimensions (bid, communication, and computation), and thus cannot be solved by traditional approaches from mechanism design. In this paper, we propose two distributed spectrum auction mechanisms, namely distributed VCG and FAITH. Distributed VCG implements the celebrated Vickrey-Clarke-Groves mechanism in a distributed fashion to achieve optimal social welfare, at the cost of exponential communication overhead. In contrast, FAITH achieves sub-optimal social welfare with tractable computation and communication overhead. We prove that both of the two proposed mechanisms achieve faithfulness, i.e., the agents' individual utilities are maximized, if they follow the intended strategies. Besides, we extend FAITH to adapt to dynamic scenarios where agents can arrive or depart at any time, without violating the property of faithfulness. We implement distributed VCG and FAITH, and evaluate their performance in various setups. Evaluation results show that distributed VCG results in optimal allocation, while FAITH is more efficient in computation and communication.
Shuo Yang 0001, Dan Peng, Tong Meng, Fan Wu 0006, Guihai Chen, Shaojie Tang 0001, Zhenhua Li 0001, Tie Luo 0001
IEEE Trans. Mob. Comput.7
2019 TailCutter: Wisely Cutting Tail Latency in Cloud CDNs Under Cost Constraints
abstract
Cloud computing platforms enable applications to offer low-latency services to users by deploying data storage in multiple geo-distributed data centers. In this paper, through benchmark measurements on Amazon AWS and Microsoft Azure together with an analysis of a large-scale dataset collected from a major cloud CDN provider, we identify the high tail latency problem in cloud CDNs, which can substantially undermine the efficacy of cloud CDNs. One crucial idea to reduce the tail latency is to send requests in parallel to multiple clouds in cloud CDNs. However, since application providers often have a budget for using cloud services, deciding how many chunks to download from each cloud and when to download chunks in a cost-efficient manner still remain as open problems in our concerned scenario. To address the problem, we present TailCutter, a workload scheduling framework that aims at optimizing the tail latency while meeting cost constraints given by application providers. Specifically, we formulate the tail latency minimization (TLM) problem in cloud CDNs and design the receding horizon control based maximum tail minimization algorithm (RHC-based MTMA) to efficiently solve the TLM problem in practice. We implement TailCutter across multiple data centers of Amazon AWS and Microsoft Azure. Extensive evaluations using a large-scale real-world data trace (collected from a major ISP) illustrate that TailCutter can reduce up to 58.9% of the 100th-percentile user-perceived latency, as compared with alternative solutions under the cost constraint.
Yong Cui 0001, Ningwei Dai, Zeqi Lai, Minming Li, Zhenhua Li 0001, Yuming Hu, Kui Ren 0001, Yuchi Chen
IEEE/ACM Trans. Netw.5
2018 FBSleuth: Fake Base Station Forensics via Radio Frequency Fingerprinting
abstract
Fake base station (FBS) crime is a type of wireless communication crime that has appeared recently. The key to enforcing the laws on regulating FBS based crime is not only to arrest but also to convict criminals effectively. Much work on FBS discovering, localization, and tracking can assist the arresting, but the problem of collecting evidence accurately to support a proper conviction has not been addressed yet.
Zhou Zhuang, Xiaoyu Ji 0001, Taimin Zhang, Juchuan Zhang, Wenyuan Xu 0001, Zhenhua Li 0001, Yunhao Liu 0001
AsiaCCS6
2018 Towards Web-based Delta Synchronization for Cloud Storage Services
Zhenhua Li 0001, Ennan Zhai, Tianyin Xu, Yang Li 0092, Yunhao Liu 0001, Quanlu Zhang, Yao Liu 0001
FAST2
2018 H2Cloud: Maintaining the Whole Filesystem in an Object Storage Cloud
abstract
Object storage clouds (e.g., Amazon S3) have become extremely popular due to their highly usable interface and cost-effectiveness. They are, therefore, widely used by various applications (e.g., Dropbox) to host user data. However, because object storage clouds are flat and lack the concept of a directory, it becomes necessary to maintain file meta-data and directory structure in a separate index cloud. This paper investigates the possibility of using a single object storage cloud to efficiently host the whole filesystem for users, including both the file content and directories, while avoiding meta-data loss caused by index cloud failures. We design a novel data structure, Hierarchical Hash (or H2), to natively enable the efficient mapping from filesystem operations to object-level operations. Based on H2, we implement a prototype system, H2Cloud, that can maintain large filesystems of users in an object storage cloud and support fast directory operations. Both theoretical analysis and real-world experiments confirm the efficacy of our solution: H2Cloud achieves faster directory operations than OpenStack Swift by orders of magnitude, and has similar performance to Dropbox but yet does not need a separate index cloud.
Minghao Zhao 0001, Zhenhua Li 0001, Ennan Zhai, Gareth Tyson, Chen Qian 0001, Zhenyu Li 0001, Leiyu Zhao
ICPP2
2018 Minimizing the Cask Effect of Multi-Source Content Delivery
abstract
This paper reveals the performance anomaly (i.e., the decline of delivery speed) when the client upgrades a task from single-source content delivery to multi-source content delivery. This anomaly is mainly caused by two aspects: (1) data sources with different types vary greatly in terms of acceleration reward (AR), and data sources with certain types are particularly easy to become inferior; (2) When the data sources remain fixed for a period of time, the large diversity of participant time (DPT) of data sources disturb the acceleration and the data sources with less participant time are inferior. Combing these insights, we figure out that the multi-source content delivery is limited by the so-called cask effect, i.e., the acceleration effect mainly depends on the inferior data sources.
Zhenhua Li 0001, Zhenyu Li 0001, Tianyin Xu, Ennan Zhai, Yao Liu 0001, Minghao Zhao 0001, Yunhao Liu 0001
IWQoS2
2018 MiniView Layout for Bandwidth-Efficient 360-Degree Video
abstract
With the recent increase in popularity of VR devices, 360-degree video has become increasingly popular. As more users experience this new medium, it will likely see further increases in popularity as users experience its greater immersiveness compared to traditional video streams. 360-degree video streams must encode the omnidirectional view, and, with current encoding techniques, these views require significantly higher bandwidth than traditional video streams. These larger bandwidth requirements comprise the main barrier toward wider adoption by video streaming services.
Mengbai Xiao, Shuoqian Wang, Chao Zhou 0004, Li Liu 0045, Zhenhua Li 0001, Yao Liu 0001, Songqing Chen
ACM Multimedia5
2018 Mobile Content Hosting Infrastructure in China: A View from a Cellular ISP
Zhenyu Li 0001, Donghui Yang, Zhenhua Li 0001, Chunjing Han, Gaogang Xie
PAM3
2018 On the Effectiveness of Offset Projections for 360-Degree Video Streaming
abstract
A new generation of video streaming technology, 360-degree video, promises greater immersiveness than standard video streams. This level of immersiveness is similar to that produced by virtual reality devices—users can control the field of view using head movements rather than needing to manipulate external devices. Although 360-degree video could revolutionize the streaming experience, its large-scale adoption is hindered by a number of factors: 360-degree video streams have larger bandwidth requirements and require faster responsiveness to user inputs, and users may be more sensitive to lower quality streams. In this article, we review standard approaches toward 360-degree video encoding and compare these to families of approaches that distort the spherical surface to allow oriented concentrations of the 360-degree view. We refer to these distorted projections as offset projections. Our measurement studies show that most types of offset projections produce rendered views with better quality than their nonoffset equivalents when view orientations are within 40 or 50 degrees of the offset orientation. Offset projections complicate adaptive 360-degree video streaming because they require a combination of bitrate and view orientation adaptations. We estimate that this combination of streaming adaptation in two dimensions can cause over 57% extra segments to be downloaded compared to an ideal downloading strategy, wasting 20% of the total downloading bandwidth.
Chao Zhou 0004, Zhenhua Li 0001, Joe Osgood, Yao Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2018 CoCloud: Enabling Efficient Cross-Cloud File Collaboration Based on Inefficient Web APIs
abstract
Cloud storage services such as Dropbox have been widely used for file collaboration among multiple users. However, this desirable functionality is yet restricted to the “walled-garden” of each service. At present, the only feasible approach to cross-cloud file collaboration seems to be using web APIs, whose performance is known to be highly unstable and unpredictable. Now that using inefficient web APIs is inevitable, in this paper we attempt to achieve sound user-perceived performance for cross-cloud file collaboration. This attempt is enabled by two key observations from real-world measurements. First, for each cloud, we are always able to deploy one or several nearby (client) proxies which can efficiently access the web APIs. Second, during file collaboration, significant similarity exists among different versions of a file. This can be exploited to substantially reduce inter-proxy traffic and thus shorten the data sync time. Guided by the observations, we design and implement an open-source prototype system called CoCloud. Currently, it supports file collaboration among four popular cloud storage services in the US and China. Its performance is well acceptable to users under representative workloads, even approaching or exceeding that of intra-cloud collaboration in many cases.
Jinlong E, Yong Cui 0001, Peng Wang 0037, Zhenhua Li 0001, Chaokun Zhang
IEEE Trans. Parallel Distributed Syst.4
2018 On the Synchronization Bottleneck of OpenStack Swift-Like Cloud Storage Systems
abstract
As one type of the most popular cloud storage services, OpenStack Swift and its follow-up systems replicate each object across multiple storage nodes and leverageobject sync protocolsto achieve high reliability andeventual consistency. The performance of object sync protocols heavily relies on two key parameters:$r$(number of replicas for each object) and$n$(number of objects hosted by each storage node). In existing tutorials and demos, the configurations are usually$r=3$and$n<1,000$by default, and the sync process seems to perform well. However, we discover in data-intensive scenarios, e.g., when$r>3$and$n\gg 1,000$, the sync process is significantly delayed and produces massive network overhead, referred to as thesync bottleneck problem. By reviewing the source code of OpenStack Swift, we find that its object sync protocol utilizes a fairly simple and network-intensive approach to check the consistency among replicas of objects. Hence in a sync round, the number of exchanged hash values per node is$\Theta (n\times r)$. To tackle the problem, we propose a lightweight and practical object sync protocol,LightSync, which not only remarkably reduces the sync overhead, but also preserves high reliability and eventual consistency. LightSync derives this capability from three novel building blocks: 1)Hashing of Hashes, which aggregates all the$h$hash values of each data partition into a single but representative hash value with the Merkle tree; 2)Circular Hash Checking, which checks the consistency of different partition replicas by only sending the aggregated hash value to the clockwise neighbor; and 3)Failed Neighbor Handling, which properly detects and handles node failures with moderate overhead to effectively strengthen the robustness of LightSync. The design of LightSync offers provable guarantee on reducing the per-node network overhead from$\Theta (n\times r)$to$\Theta (\frac{n}{h})$. Furthermore, we have implemented LightSync as an open-source patch and adopted it to OpenStack Swift, thus reducing the sync delay by up to 879$\times$and the network overhead by up to 47.5$\times$.
Mingkang Ruan, Thierry Titcheu Chekam, Ennan Zhai, Zhenhua Li 0001, Yao Liu 0001, Jinlong E, Yong Cui 0001, Hong Xu 0001
IEEE Trans. Parallel Distributed Syst.4
2017 Unbiased Sampling of Social Media Networks for Well-connected Subgraphs
abstract
Sampling social graphs is critical for studying things like information diffusion. However, it is often necessary to laboriously obtain unbiased and well-connected datasets because existing survey algorithms are unable to generate well-connected samples, and current random-walk based unbiased sampling algorithms adopt rejection sampling, which heavily undermines performance. This paper proposes a novel random-walk based algorithm which implements Unbiased Sampling using Dummy Edges (USDE). It injects dummy edges between nodes, on which the walkers would otherwise experience excessive rejections before moving out from such nodes. We propose a rejection probability estimation algorithm to facilitate the construction of dummy edges and the computation of moving probabilities. Finally, we apply USDE in two real-life social media: Twitter and Sina Weibo. The results demonstrate that USDE generates well-connected samples, and outperforms existing approaches in terms of sampling efficiency and quality of samples.
Dong Wang 0027, Zhenyu Li 0001, Gareth Tyson, Zhenhua Li 0001, Gaogang Xie
ASONAM4
2017 Practical Web-based Delta Synchronization for Cloud Storage Services
Zhenhua Li 0001, Ennan Zhai, Tianyin Xu
HotStorage2
2017 DeltaCFS: Boosting Delta Sync for Cloud Storage Services by Learning from NFS
abstract
Cloud storage services, such as Dropbox, iCloud Drive, Google Drive, and Microsoft OneDrive, have greatly facilitated users' synchronizing files across heterogeneous devices. Among them, Dropbox-like services are particularly beneficial owing to the delta sync functionality that strives towards greater network-level efficiency. However, when delta sync trades computation overhead for network-traffic saving, the tradeoff could be highly unfavorable under some typical workloads. We refer to this problem as the abuse of delta sync. To address this problem, we propose DeltaCFS, a novel file sync framework for cloud storage services by learning from the design of conventional NFS (Network File System). Specifically, we combine delta sync with NFS-like file RPC in an adaptive manner, thus significantly cutting computation overhead on both the client and server sides while preserving the network-level efficiency. DeltaCFS also enables a neat design for guaranteeing causal consistency and fine-grained version control of files. In our FUSE-based prototype system (which is open-source), DeltaCFS outperforms Dropbox by generating up to 11x less data transfer and up to 100x less computation overhead under concerned workloads.
Quanlu Zhang, Zhenhua Li 0001, Zhi Yang 0001, Shenglong Li, Shouyang Li, Yangze Guo, Yafei Dai
ICDCS2
2017 Large-scale invisible attack on AFC systems with NFC-equipped smartphones
abstract
Automated Fare Collection (AFC) systems have been globally deployed for decades, particularly in public transportation. Although the transaction messages of AFC systems are mostly transferred in plaintext, which is obviously insecure, system operators do not need to pay much attention to this issue, since the AFC network is well isolated from public network (e.g., the Internet). Nevertheless, in recent years, the advent of Near Field Communication (NFC)-equipped smartphones has bridged the gap between the AFC network and the Internet through Host-based Card Emulation (HCE). Motivated by this fact, we design and practice a novel paradigm of attack on modern distance-based pricing AFC systems, enabling users to pay much less than actually required. Our constructed attack has two important properties: 1) it is invisible to AFC system operators because the attack never causes any inconsistency in the backend database of the operators; and 2) it can be scalable to large number of users (e.g., 10,000) by maintaining a moderate-sized AFC card pool (e.g., containing 150 cards). Based upon this constructed attack, we developed an HCE app, named LessPay. Our real-world experiments on LessPay demonstrate not only the feasibility of our attack (with 97.6% success rate), but also its low-overhead in terms of bandwidth and computation.
Fan Dang 0001, Zhenhua Li 0001, Ennan Zhai, David Mohaisen, Qingfu Wen, Mo Li 0001
INFOCOM3
2017 CoCloud: Enabling efficient cross-cloud file collaboration based on inefficient web APIs
abstract
Cloud storage services such as Dropbox have been widely used for file collaboration among multiple users. However, this desirable functionality is yet restricted to the “walled-garden” of each service. At present, the only effective approach to cross-cloud file collaboration seems to be using web APIs, whose performance is known to be highly unstable and unpredictable. Now that using inefficient web APIs is inevitable, in this paper we attempt to achieve sound user-perceived performance for cross-cloud file collaboration. This attempt is enabled by two key observations from real-world measurements. First, for each cloud, we are always able to deploy one or several nearby (client) proxies which can efficiently access the web APIs. Second, during file collaboration, significant similarity exists among different versions of a file. This can be exploited to substantially reduce inter-proxy traffic and thus shorten the data sync time. Guided by the observations, we design and implement an open-source prototype system called CoCloud. Currently, it supports file collaboration among four popular cloud storage services in the US and China. Its performance is well acceptable to users under representative workloads, even approaching or exceeding intra-cloud performance in many cases.
Jinlong E, Yong Cui 0001, Peng Wang 0037, Zhenhua Li 0001, Chaokun Zhang
INFOCOM4
2017 A Measurement Study of Oculus 360 Degree Video Streaming
abstract
360 degree video is anew generation of video streaming technology that promises greater immersiveness than standard video streams. This level of immersiveness is similar to that produced by virtual reality devices -- users can control the field of view using head movements rather than needing to manipulate external devices. Although 360 degree video could revolutionize streaming technology, large scale adoption is hindered by a number of factors. 360 degree video streams have larger bandwidth requirements, require faster responsiveness to user inputs, and users may be more sensitive to lower quality streams.; [email protected] this paper, we review standard approaches toward 360 degree video encoding and compare these to a new, as yet unpublished, approach by Oculus which we refer to as the offset cubic projection. Compared to the standard cubic encoding, the offset cube encodes a distorted version of the spherical surface, devoting more information (i.e., pixels) to the view in a chosen direction. We estimate that the offset cube representation can produce better or similar visual quality while using less than 50% pixels under reasonable assumptions about user behavior, resulting in 5.6% to 16.4% average savings in video bitrate. During 360 degree video streaming, Oculus uses a combination of quality level adaptation and view orientation adaptation. We estimate that this combination of streaming adaptation in two dimensions can cause over 57% extra segments to be downloaded compared to an ideal downloading strategy, wasting 20% of the total downloading bandwidth.
Chao Zhou 0004, Zhenhua Li 0001, Yao Liu 0001
MMSys2
2017 FBS-Radar: Uncovering Fake Base Stations at Scale in the Wild
Zhenhua Li 0001, Weiwei Wang 0002, Christo Wilson, Chen Qian 0001, Taeho Jung, Lan Zhang 0002, Kebin Liu 0001, Xiang-Yang Li 0001, Yunhao Liu 0001
NDSS1
2017 Montage: Combine Frames with Movement Continuity for Realtime Multi-User Tracking
abstract
In this work, we design and develop Montage for real-time multi-user formation tracking and localization by off-the-shelf smartphones. Montage achieves submeter-level tracking accuracy by integrating temporal and spatial constraints from user movement vectorestimation and distance measuring. In Montage, we designed a suite of novel techniques to surmount a variety of challenges in real-time tracking, without infrastructure and fingerprints, and without any a priori user-specific (e.g., stride-length and phoneplacement) or site-specific (e.g., digitalized map) knowledge: (1) a coded audio tone to support multi-user tracking with minimal latency, in the presence of high noise, multi-path effect, and Doppler Shift, (2) an innovative stride-length and walking direction estimation method without a priori knowledge of user and site, and (3) a vector-based multi-user tracking scheme which connects successive localization snapshots to refine users' locations and generate continuous moving traces. We implemented, deployed, and evaluated Montage in both outdoor and indoor environment. Our experimental results (847 traces from 15 users) show that the stride-length estimated by Montage over all users has error within 9cm, and the moving-direction estimated by Montage is within 20 degrees. For real-time tracking, Montage provides meter-second-level formation tracking accuracy with off-the-shelf mobile phones.
Lan Zhang 0002, Kebin Liu 0001, Yonghang Jiang, Xiang-Yang Li 0001, Yunhao Liu 0001, Panlong Yang, Zhenhua Li 0001
IEEE Trans. Mob. Comput.7
2016 An Empirical Analysis of a Large-scale Mobile Cloud Storage Service
Zhenyu Li 0001, Xiaohui Wang 0012, Ningjing Huang, Mohamed Ali Kâafar, Zhenhua Li 0001, Jianer Zhou, Gaogang Xie, Peter Steenkiste
Internet Measurement Conference5
2016 On the synchronization bottleneck of OpenStack Swift-like cloud storage systems
abstract
As one type of the most popular cloud storage services, OpenStack Swift and its follow-up systems replicate each data object across multiple storage nodes and leverage object sync protocols to achieve high availability and eventual consistency. The performance of object sync protocols heavily relies on two key parameters: r (number of replicas for each object) and η (number of objects hosted by each storage node). In existing tutorials and demos, the configurations are usually r = 3 and n3 and n ≫ 1000, the object sync process is significantly delayed and produces massive network overhead. This phenomenon is referred to as the sync bottleneck problem. Then, to explore the root cause, we review the source code of OpenStack Swift and find that its object sync protocol utilizes a fairly simple and network-intensive approach to check the consistency among replicas of objects. In particular, each storage node is required to periodically multicast the hash values of all its hosted objects to all the other replica nodes. Thus in a sync round, the number of exchanged hash values per node is Θ(n×r). Further, to tackle the problem, we propose a lightweight object sync protocol called LightSync. It remarkably reduces the sync overhead by using two novel building blocks: 1) Hashing of Hashes, which aggregates all the h hash values of each data partition into a single but representative hash value with the Merkle tree; 2) Circular Hash Checking, which checks the consistency of different partition replicas by only sending the aggregated hash value to the clockwise neighbor. Its design provably reduces the per-node network overhead from Θ(n×r) to Θ(n/h). In addition, we have implemented LightSync as an open-source patch and adopted it to OpenStack Swift, thus reducing sync delay by up to 28.8× and network overhead by up to 14.2×.
Thierry Titcheu Chekam, Ennan Zhai, Zhenhua Li 0001, Yong Cui 0001, Kui Ren 0001
INFOCOM3
2016 TailCutter: Wisely cutting tail latency in cloud CDN under cost constraints
abstract
Cloud computing platforms enable applications to offer low latency access to user data by offering storage services in several geographically distributed data centers. In this paper, we identify the high tail latency problem in cloud CDN via analyzing a large-scale dataset collected from 783,944 users in a major cloud CDN. We find that the data downloading latency in cloud CDN is highly variable, which may significantly degrade the user experience of applications. To address the problem, we present TailCutter, a workload scheduling mechanism that aims at optimizing the tail latency while meeting the cost constraint given by application providers. We further design the Maximum Tail Minimization Algorithm (MTMA) working in TailCutter mechanism to optimally solve the Tail Latency Minimization (TLM) problem in polynomial time. We implement TailCutter across data centers of Amazon S3 and Microsoft Azure. Our extensive evaluation using large-scale real world data traces shows that TailCutter can reduce up to 68% 99th percentile user-perceived latency in comparison with alternative solutions under cost constraints.
Zeqi Lai, Yong Cui 0001, Minming Li, Zhenhua Li 0001, Ningwei Dai, Yuchi Chen
INFOCOM4
2016 On the performance of cloud storage applications with global measurement
abstract
In recent years, Dropbox, Google, and Microsoft have been competing in the market of consumer cloud storage (CCS) services. While once the key comparative metric, storage capacity per user has outgrown the needs of most users. Today, third-party applications based on CCS's RESTful Web APIs are becoming a primary way for users to utilize their expanded storage resources. Unfortunately, there is very little visibility into the performance of these Web APIs, even though they are primary determinants of the end user experience on these storage applications. In this paper, we report results from a comprehensive measurement study of the Web APIs of five popular CCS providers. Our results reveal significant differences and limitations in API performance, which result in performance bottlenecks visible to the user through the storage application. We analyze the underlying system designs of the five providers' Web APIs, and present the performance implications of their different design choices. Our research provides practical guidance for service providers to optimize their API performance, for developers to improve the experience of third-party applications, and for users to pick appropriate services that best match their requirements.
Guangyuan Wu, Fangming Liu, Haowen Tang, Keke Huang, Qixia Zhang, Zhenhua Li 0001, Ben Y. Zhao, Hai Jin 0001
IWQoS6
2016 Exploring Cross-Application Cellular Traffic Optimization with Baidu TrafficGuard
Zhenhua Li 0001, Weiwei Wang 0002, Tianyin Xu, Xiang-Yang Li 0001, Yunhao Liu 0001, Christo Wilson, Ben Y. Zhao
NSDI1
2016 GoCAD: GPU-Assisted Online Content-Adaptive Display Power Saving for Mobile Devices in Internet Streaming
abstract
During Internet streaming, a significant portion of the battery power is always consumed by the display panel on mobile devices. To reduce the display power consumption, backlight scaling, a scheme that intelligently dims the backlight has been proposed. To maintain perceived video appearance in backlight scaling, a computationally intensive luminance compensation process is required. However, this step, if performed by the CPU as existing schemes suggest, could easily offset the power savings gained from backlight scaling. Furthermore, computing the optimal backlight scaling values requires per-frame luminance information, which is typically too energy intensive for mobile devices to compute. Thus, existing schemes require such information to be available in advance. And such an offline approach makes these schemes impractical. To address these challenges, in this paper, we design and implement GoCAD, a GPU-assisted Online Content-Adaptive Display power saving scheme for mobile devices in Internet streaming sessions. In GoCAD, we employ the mobile device's GPU rather than the CPU to reduce power consumption during the luminance compensation phase. Furthermore, we compute the optimal backlight scaling values for small batches of video frames in an online fashion using a dynamic programming algorithm. Lastly, we make novel use of the widely available video storyboard, a pre-computed set of thumbnails associated with a video, to intelligently decide whether or not to apply our backlight scaling scheme for a given video. For example, when the GPU power consumption would offset the savings from dimming the backlight, no backlight scaling is conducted. To evaluate the performance of GoCAD, we implement a prototype within an Android application and use a Monsoon power monitor to measure the real power consumption. Experiments are conducted on more than 460 randomly selected YouTube videos. Results show that GoCAD can effectively produce power savings without affecting rendered video quality.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Songqing Chen
WWW7
2016 Resisting Tag Spam by Leveraging Implicit User Behaviors
abstract
Tagging systems are vulnerable to tag spam attacks. However, defending against tag spam has been challenging in practice, since adversaries can easily launch spam attacks in various ways and scales. To deeply understand users' tagging behaviors and explore more effective defense, this paper first conducts measurement experiments on public datasets of two representative tagging systems: Del.icio.us and CiteULike. Our key finding is that a significant fraction of correct tag-resource annotations are contributed by a small number of implicit similarity cliques, where users annotate common resources with similar tags. Guided by the above finding, we propose a new service, called Spam-Resistance-as-a-Service (or SRaaS), to effectively defend against heterogeneous tag spam attacks even at very large scales. At the heart of SRaaS is a novel reputation assessment protocol, whose design leverages the implicit similarity cliques coupled with the social networks inherent to typical tagging systems. With such a design, SRaaS manages to offer provable guarantees on diminishing the influence of tag spam attacks. We build an SRaaS prototype and evaluate it using a large-scale spam-oriented research dataset (which is much more polluted by tag spam than Del.icio.us and CiteULike datasets). Our evaluational results demonstrate that SRaaS outperforms existing tag spam defenses deployed in real-world systems, while introducing low overhead.
Ennan Zhai, Zhenhua Li 0001, Zhenyu Li 0001, Fan Wu 0006, Guihai Chen
Proc. VLDB Endow.2
2016 Content-Adaptive Display Power Saving for Internet Video Applications on Mobile Devices
abstract
Backlight scaling is a technique proposed to reduce the display panel power consumption by strategically dimming the backlight. However, for mobile video applications, a computationally intensive luminance compensation step must be performed in combination with backlight scaling to maintain the perceived appearance of video frames. This step, if done by the Central Processing Unit (CPU), could easily offset the power savings via backlight dimming. Furthermore, computing the backlight scaling values requires per-frame luminance information, which is typically too energy intensive to compute on mobile devices. In this article, we propose Content-Adaptive Display (CAD) for two typical Internet mobile video applications: video streaming and real-time video communication. CAD uses the mobile device’s Graphics Processing Unit (GPU) rather than the CPU to perform luminance compensation at reduced power consumption. For video streaming where video frames are available in advance, we compute the backlight scaling schedule using a more efficient dynamic programming algorithm than existing work. For real-time video communication where video frames are generated on the fly, we propose a greedy algorithm to determine the backlight scaling at runtime. We implement CAD in one video streaming application and one real-time video call application on the Android platform and use a Monsoon power meter to measure the real power consumption. Experiment results show that CAD can save more than 10% overall power consumption for up to 55.7% videos during video streaming and up to 31.0% overall power consumption in real-time video calls.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Lei Guo 0004, Songqing Chen
ACM Trans. Multim. Comput. Commun. Appl.7
2015 Do Twin Clouds Make Smoothness for Transoceanic Video Telephony?
abstract
Transoceanic video telephony (TVT) over the Internet is challenging due to 1) longer round-trip delay, 2) larger number of relay hops, and 3) higher packet loss rate. Real-world measurements of Skype, Face time, and QQ confirm that their TVT service quality is mostly unsatisfactory. Recently, when using We Chat to make transoceanic video calls, we are fortunate to find that it achieves stably smooth TVT. To explore how this is possible, we conduct in-depth measurements of We Chat data flow. In particular, we discover that the service provider of We Chat deploys a novel, specially designed "twin clouds" based architecture to deliver transoceanic (UDP) packets. Thus, data delivery between two callers is no longer point-to-point (used by Skype, Face time, and QQ) over the best-effort Internet. Instead, transoceanic video packets are delivered through the privileged backbone formed by twin clouds, which greatly reduces the round-trip delay, number of relay hops, and packet loss rate. Besides, whenever a packet is found lost, multiple duplicate packets are instantly sent to aggressively make up for the loss. On the other hand, we notice two-fold shortcomings of twin clouds. First, due to the sophisticated resource provisioning inside the twin clouds, the video start up time is considerably extended. Second, due to the high cost of deploying twin clouds, the capacity of the privileged backbone is limited and sometimes in shortage, and thus We Chat has to deliver data via a detour path with degraded performance. Ultimately, we believe that the twin clouds based data delivery solution will arouse a new direction of Internet video telephony research while still deserves optimization efforts.
Zhenhua Li 0001, Yao Liu 0001, Zhi-Li Zhang
ICPP2
2015 LMDD: Light-Weight Magnetic-Based Door Detection with Your Smartphone
abstract
Doors are important landmarks for indoor positioning systems. Hence an accurate and light-weight door detection approach is highly desired. The state-of-the-art solutions are either vision based or infrastructure based, which incur nontrivial device or management cost. This paper presents a novel approach, Light-weight Magnetic-based Door Detection (LMDD), which only relies on the information from built-in sensors of a smartphone. LMDD detects a door by analyzing the change of magnetic signal and extracting special features caused by doors. It is light-weight in both computation and infrastructure cost. We have implemented a prototype of LMDD that has been installed on various Android phones. Experimental results show that LMDD achieves door detection accuracy of 74% in average, ranging from 66% to 85% in various typical environments such as offices, classrooms, residential houses, and a hospital.
Chen Qian 0001, Liangyi Gong, Zhenhua Li 0001, Yunhao Liu 0001
ICPP4
2015 Offline Downloading in China: A Comparative Study
abstract
Although Internet access has become more ubiquitous in recent years, most users in China still suffer from low-quality connections, especially when downloading large files. To address this issue, hundreds of millions of China's users have resorted to technologies that allow for ``offline downloading'', where a proxy is employed to pre-download the user's requested file and then deliver the file at her convenience.
Zhenhua Li 0001, Christo Wilson, Tianyin Xu, Yao Liu 0001, Yinlong Wang
Internet Measurement Conference1
2015 Content-adaptive display power saving in internet mobile streaming
abstract
Backlight scaling is a technique proposed to reduce the display panel power consumption by strategically dimming the backlight. However, for Internet streaming to mobile devices, a computationally intensive luminance compensation step must be performed in combination with backlight scaling to maintain the perceived appearance of video frames. This step, if done by the CPU, could easily offset the power savings via backlight dimming. Furthermore, computing the backlight scaling values requires per-frame luminance information, which is typically too energy intensive to compute on mobile devices.
Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Songqing Chen
NOSSDAV7
2015 A Quantitative Study of Video Duplicate Levels in YouTube
Yao Liu 0001, Sam Blasiak, Weijun Xiao, Zhenhua Li 0001, Songqing Chen
PAM4
2015 Towards Cost-Effective Cloud Downloading with Tencent Big Data
Zhenhua Li 0001, Zhi-Yuan Ji, Roger Zimmermann
J. Comput. Sci. Technol.1
2015 Accurate DNS query characteristics estimation via active probing
Xiaobo Ma 0001, Junjie Zhang 0004, Zhenhua Li 0001, Jianfeng Li 0006, Xiaohong Guan, John C. S. Lui, Don Towsley
J. Netw. Comput. Appl.3
2015 CHARM: A Cost-Efficient Multi-Cloud Data Hosting Scheme with High Availability
abstract
Nowadays, more and more enterprises and organizations are hosting their data into the cloud, in order to reduce the IT maintenance cost and enhance the data reliability. However, facing the numerous cloud vendors as well as their heterogenous pricing policies, customers maywell be perplexed with which cloud(s) are suitable for storing their data and what hosting strategy is cheaper.The general status quo is that customers usually put their data into a single cloud (which is subject to the vendor lock-in risk) and then simply trust to luck. Based on comprehensive analysis of various state-of-the-art cloud vendors, this paper proposes a novel data hosting scheme (named CHARM) which integrates two key functions desired. The first is selecting several suitable clouds and an appropriate redundancy strategy to store data with minimized monetary cost and guaranteed availability. The second is triggering a transition process to re-distribute data according to the variations of data access pattern and pricing of clouds. We evaluate the performance of CHARM using both trace-driven simulations and prototype experiments. The results show that compared with the major existing schemes, CHARM not only saves around 20 percent of monetary cost but also exhibits sound adaptability to data and price adjustments.
Quanlu Zhang, Shenglong Li, Zhenhua Li 0001, Yuanjian Xing, Zhi Yang 0001, Yafei Dai
IEEE Trans. Cloud Comput.3
2014 Towards Network-level Efficiency for Cloud Storage Services
abstract
Cloud storage services such as Dropbox, Google Drive, and Microsoft OneDrive provide users with a convenient and reliable way to store and share data from anywhere, on any device, and at any time. The cornerstone of these services is the data synchronization (sync) operation which automatically maps the changes in users' local filesystems to the cloud via a series of network communications in a timely manner. If not designed properly, however, the tremendous amount of data sync traffic can potentially cause (financial) pains to both service providers and users.
Zhenhua Li 0001, Cheng Jin 0008, Tianyin Xu, Christo Wilson, Yao Liu 0001, Linsong Cheng, Yunhao Liu 0001, Yafei Dai, Zhi-Li Zhang
Internet Measurement Conference1
2014 PACOM: Parasitic anonymous communication in the BitTorrent network
Jianming Lv, Tieying Zhang, Zhenhua Li 0001, Xueqi Cheng 0001
Comput. Networks3
2013 Efficient Batched Synchronization in Dropbox-Like Cloud Storage Services
Zhenhua Li 0001, Christo Wilson, Zhefu Jiang, Yao Liu 0001, Ben Y. Zhao, Cheng Jin 0008, Zhi-Li Zhang, Yafei Dai
Middleware1
2013 Challenges, Designs, and Performances of Large-Scale Open-P2SP Content Distribution
abstract
Content distribution on today's Internet operates primarily in two modes: server-based and peer-to-peer (P2P). To leverage the advantages of both modes while circumventing their key limitations, a third mode: peer-to-server/peer (P2SP) has emerged in recent years. Although P2SP can provide efficient hybrid server-P2P content distribution, P2SP generally works in a closed manner by only utilizing its private owned servers to accelerate its private organized peer swarms. Consequently, P2SP still has its limitations in both content abundance and server bandwidth. To this end, the fourth mode (or says a generalized mode of P2SP) has appeared as "open-P2SP" that integrates various third-party servers, contents, and data transfer protocols all over the Internet into a large, open, and federated P2SP platform. In this paper, based on a large-scale commercial open-P2SP system named "QQXuanfeng" , we investigate the key challenging problems, practical designs and real-world performances of open-P2SP. Such "white-box" study of open-P2SP provides solid experiences and helpful heuristics to the designers of similar systems.
Zhenhua Li 0001, Fuchen Wang, Yunhao Liu 0001, Zhi-Li Zhang, Yafei Dai
IEEE Trans. Parallel Distributed Syst.1
2012 Maximizing the bandwidth multiplier effect for hybrid cloud-P2P content distribution
abstract
Hybrid cloud-P2P content distribution (“CloudP2P”) provides a promising alternative to the conventional cloud-based or peer-to-peer (P2P)-based large-scale content distribution. It addresses the potential limitations of these two conventional approaches while inheriting their advantages. A key strength of CloudP2P lies in the so-called bandwidth multiplier effect: by appropriately allocating a small portion of cloud (server) bandwidth Sito a peer swarm i (consisting of users interested in the same content) to seed the content, the users in the peer swarm - with an aggregate download bandwidth Di- can then distribute the content among themselves; we refer to the ratio Di/Sias the bandwidth multiplier (for peer swarm i). A major problem in the design of a CloudP2P content distribution system is therefore how to allocate cloud (server) bandwidth to peer swarms so as to maximize the overall bandwidth multiplier effect of the system. In this paper, using real-world measurements, we identify the key factors that affect the bandwidth multipliers of peer swarms and thus construct a fine-grained performance model for addressing the optimal bandwidth allocation problem (OBAP). Then we develop a fast-convergent iterative algorithm to solve OBAP. Both trace-driven simulations and prototype implementation confirm the efficacy of our solution.
Zhenhua Li 0001, Tieying Zhang, Zhi-Li Zhang, Yafei Dai
IWQoS1
2012 Cloud transcoder: bridging the format and resolution gap between internet videos and mobile devices
abstract
Despite the increasing popularity, Internet video streaming to mobile devices is still challenging. In particular, there has been a format and resolution "gap" between Internet videos and mobile devices, so mobile users have high demand on video transcoding to facilitate their specific devices. However, as a computation-intensive work, video transcoding is greatly challenged by the limited battery capacity of mobile devices. In this paper we propose and implement "Cloud Transcoder", which utilizes an intermediate cloud platform to bridge the "gap" via its special and practical designs. Specifically, Cloud Transcoder only requires the user to upload a video request rather than the video content. After getting the video request, Cloud Transcoder downloads the original video from the Internet, transcodes it on the user's demand, and transfers the transcoded video back to the user with a high data rate via the intra-cloud data transfer acceleration. Therefore, the mobile device only consumes energy in the last step - fast retrieving the transcoded video from the cloud. Running logs of our real-deployed system confirm the efficacy of Cloud Transcoder.
Zhenhua Li 0001, Fuchen Wang, Zhi-Li Zhang, Yafei Dai
NOSSDAV1
2012 Providing hierarchical lookup service for P2P-VoD systems
abstract
Supporting random jump in P2P-VoD systems requires efficient lookup for the “best” suppliers, where “best” means the suppliers should meet two requirements: content match and network quality match . Most studies use a DHT-based method to provide content lookup; however, these methods are neither able to meet the network quality requirements nor suitable for VoD streaming due to the large overhead. In this paper, we propose Mediacoop, a novel hierarchical lookup scheme combining both content and quality match to provide random jumps for P2P-VoD systems. It exploits the play position to efficiently locate the candidate suppliers with required data (content match), and performs refined lookup within the candidates to meet quality match. Theoretical analysis and simulation results show that Mediacoop is able to achieve lower jump latency and control overhead than the typical DHT-based method. Moreover, we implement Mediacoop in a BitTorrent-like P2P-VoD system called CoolFish and make optimizations for such “total cache” applications. The implementation and evaluation in CoolFish show that Mediacoop is able to improve user experiences, especially the jump latency, which verifies the practicability of our design.
Tieying Zhang, Xueqi Cheng 0001, Jianming Lv, Zhenhua Li 0001, Weisong Shi
ACM Trans. Multim. Comput. Commun. Appl.4
2011 A White-Box Empirical Study of P2P-VoD Systems: Several Unconventional New Findings
abstract
P2P-VoD systems have gained tremendous popularity in recent years. While existing research is mostly based on theoretical or conventional assumptions, it is particularly valuable to understand and examine how these assumptions work in realistic environments, so as to set up a solid foundation for mechanism design and optimization possibilities. In this paper, we present a comprehensive measurement study of CoolFish, a real-world P2P-VoD system. Our measurement provides several new findings which are different from the traditional assumptions or observations: the access pattern does not match Poisson distribution; session time does not have positive correlation with movie popularity; jump frequency does not have a negative correlation with movie popularity as assumed in previous studies. We analyze the reasons for these results and provide suggestions for the further study of P2P-VoD services.
Tieying Zhang, Zhenhua Li 0001, Huawei Shen, Xueqi Cheng 0001
ICCCN2
2011 RELookup: Providing Resilient and Efficient Lookup Service for P2P-VoD Streaming
abstract
For P2P-VoD streaming, an effective lookup algorithm for appropriate data suppliers is required to support the user's operation of random jump on the video. Existing lookup algorithms mainly adopt a centralized, flooding based, or DHT-based method. Facing the highly dynamic Internet environments, the centralized method incurs a single point of failure, the flooding-based method lacks scalability, and the DHT-based method is not resilient. Motivated by these problems, we propose a novel lookup algorithm, named "RELookup", which places peers on a resilient super node-based overlay and meanwhile utilizes the play point distance to efficiently locate candidate data suppliers. Besides, deliberate measures (i.e., special design of message format and node state) have been taken to reduce the coordination costs between super nodes to very little. Results of trace-driven simulations confirm the effectiveness of our proposed RELookup algorithm.
Xu Zhang 0006, Zhenhua Li 0001, Tieying Zhang, Liangpeng He, Guihai Chen
ICPADS2
2011 Cloud download: using cloud utilities to achieve high-quality content distribution for unpopular videos
abstract
Video content distribution dominates the Internet traffic. The state-of-the-art techniques generally work well in distributing popular videos, but do not provide satisfactory content distribution service for unpopular videos due to low data health or low data transfer rate. In recent years, the worldwide deployment of cloud utilities provides us with a novel perspective to consider the above problem. We propose and implement the cloud download scheme, which achieves high-quality video content distribution by using cloud utilities to guarantee the data health and enhance the data transfer rate. Specifically, a user sends his video request to the cloud which subsequently downloads the video from the Internet and stores it in the cloud cache. Then the user can usually retrieve his requested video (whether popular or unpopular) from the cloud with high data rate in any place at any time, via the intra-cloud data transfer acceleration. Running logs of our real deployed commercial system (named VideoCloud) confirm the effectiveness and efficiency of cloud download. The users' average data transfer rate of unpopular videos exceeds 1.6 Mbps, and over 80% of the data transfer rates are more than 300 Kbps which is the basic playback rate of online videos. Our study provides practical experiences and valuable heuristics for making use of cloud utilities to achieve efficient Internet services.
Zhenhua Li 0001, Yafei Dai
ACM Multimedia2
2011 Stability-Optimal Grouping Strategy of Peer-to-Peer Systems
abstract
When applied in high-churn Internet environments, P2P systems face a dilemma: although most participants are too unstable, a P2P system requires sufficient stable peers to provide satisfactory core services. Thus, determining how to leverage unstable nodes seems to be the only choice. Our primary idea is to group unstable nodes together in order to form an adequate number of stable service groups. Focusing on this topic, our main findings are three-fold: 1) A general analytical model to investigate the grouping process of P2P systems is established, in which the stability-scalability trade-off problem is paid special attention to. 2) We formalize the target of grouping as the Maximum Stability Grouping (MSG) problem. It proves to be not only NP-hard, but also infeasible; therefore, we restrict it to a feasible Homogeneous MSG (H-MSG) problem and deduce its optimal solution under the stochastic model. 3) We propose a homogeneous grouping strategy to fulfill the optimal solution. Comprehensive simulations have been performed on generated data sets and real-world traces from a P2P storage system and a P2P streaming system. Results show that our grouping strategy effectively captures the stability-scalability trade-off: besides excellent stability, it gains much higher stable service capacity, with acceptable loss in scalability.
Zhenhua Li 0001, Jie Wu 0001, Tieying Zhang, Guihai Chen, Yafei Dai
IEEE Trans. Parallel Distributed Syst.1
2010 Multi-task Downloading for P2P-VoD: An Empirical Perspective
abstract
For current P2P-VoD systems, three fundamental problems exit in user experience: exceedingly large startup delay, long jump latency, and poor playback continuity. These problems primarily stem from lack of media data. In this paper, we propose Multi-Task Downloading with Bandwidth Control (MTD(BC)), an efficient and practical mechanism to prefetch media data. In MTD, a user can download multiple videos in parallel with its current viewing, which significantly decreases video switching delays. However, MTD brings a serious problem: downloading "other" tasks could impede the playback performance of the current viewing, especially in low-bandwidth network. This problem is solved through our design of bandwidth control. To our knowledge, we are the first to propose MTD with bandwidth control for P2P-VoD and conduct empirical evaluations in the real-world system. The running results show that MTD(BC) achieves better streaming quality than the traditional method. In particular, our mechanism reduces 75% of startup delay and 36% of jump latency in low-bandwidth network with high system scalability.
Tieying Zhang, Zhenhua Li 0001, Xueqi Cheng 0001, Xianghui Sun
ICPADS2
2010 PeerDedupe: Insights into the Peer-Assisted Sampling Deduplication
abstract
As the digital data rapidly inflates to a world-wide storage crisis, data deduplication is showing its increasingly prominent function in data storage. Driven by the problems behind the mainstream server-side deduplication schemes, recently there has been a tendency of introducing peer-assisted methods into the deduplication systems. However, this topic is still quite vague at present and lacks thorough research. In this paper, we conduct in-depth and quantitative investigation on the peer-assisted deduplication. Through measurements we observe that the inter-peer duplication accounts for a large proportion of the total duplication, and exhibits strong peer locality. Then based on our observations, we propose PeerDedupe, a novel peer-assisted sampling deduplication approach. Experiments show that PeerDedupe can remove over 98% duplication with each peer coordinating with no more than 5 other peers, and it requires much less server RAM usage than the existing works.
Yuanjian Xing, Zhenhua Li 0001, Yafei Dai
Peer-to-Peer Computing2
2010 On the source switching problem of Peer-to-Peer streaming
Zhenhua Li 0001, Jiannong Cao 0001, Guihai Chen, Yan Liu 0004
J. Parallel Distributed Comput.1
2009 On Maximum Stability with Enhanced Scalability in High-Churn DHT Deployment
abstract
When applied in a commercial deployment, DHT-based P2P protocols face a dilemma: although most real-world participants are so unstable that the maintenance overhead is prohibitively high, they must be effectively utilized due to the lack of stable participants. Thus, determining how to leverage unstable nodes to enhance system scalability and then maximize stability in high-churn scenarios becomes a substantial problem. This paper focuses on this topic, and our main findings are two folds: (1) we propose a homogeneous grouping scheme for scalability enhancement. Besides extending system storage capacity by admitting all nodes, it clusters homogeneous nodes together, deploys the inter- and intra-group connections distinctively, and tunes the number of groups, which aims to facilitate search efficiency; (2) we further look into how to maximize stability under this scheme, which is formulated as the problem Maximum Stability of Grouping. It not only proves to be NP-hard, but also infeasible; therefore, we propose an approximated grouping approach and reduce it to an optimization problem that proves to be feasible. Simulation results exhibit that our grouping strategy effectively captures the stability-scalability tradeoff. Based on our proposed measurement metrics, it doubles the storage capacity of so-called GiantOnly strategy by incurring slightly more churn and search latency, and is about four times as stable as Chord with equal capacity and mild improvement in search efficiency.
Zhenhua Li 0001, Guihai Chen, Jie Wu 0001
ICPP2
2008 Fast Source Switching for Gossip-Based Peer-to-Peer Streaming
abstract
In this paper we consider gossip-based peer-to-peer streaming applications where multiple sources exist and they work serially. More specifically, we tackle the problem of fast source switching to minimize the startup delay of the new source. We model the source switch process and formulate it into an optimization problem. Then we propose a practical greedy algorithm that can approximate the optimal solution by properly interleaving the data delivery of the old source and the new source. We perform simulations on various real-trace overlay topologies to demonstrate the effectiveness of our algorithm. The simulation results show that our proposed algorithm outperforms the normal source switch algorithm by reducing the source switch time by 20%-30% without bringing extra communication overhead, and the reduction ratio tends to increase when the network scale expands.
Zhenhua Li 0001, Jiannong Cao 0001, Guihai Chen, Yan Liu 0004
ICPP1
2008 ContinuStreaming: Achieving high playback continuity of Gossip-based Peer-to-Peer streaming
abstract
Gossip-based peer-to-peer (P2P) streaming has been proved to be an effective and resilient method to stream qualified media contents in dynamic and heterogeneous network environments. Because of the intrinsic randomness of gossiping, some data segments cannot be disseminated to every node in time, which seriously affects the media playback continuity. In this paper we describe the design of ContinuStreaming, a gossip-based P2P streaming system which can maintain high resilience and low overhead while bring a novel and critical property - full coverage of the data dissemination. With the help of DHT, data segments which are likely to be missed by the gossip-based data scheduling can be quickly fetched by the on-demand data retrieval so as to guarantee continuous playback. We discuss the results of both theoretical analysis and comprehensive simulations on various real-trace overlay topologies to demonstrate the effectiveness of our system. Simulation results show that ContinuStreaming outperforms the existing representative gossip-based P2P streaming systems by increasing the playback continuity very close to 1.0 with only 4% or less extra overhead.
Zhenhua Li 0001, Jiannong Cao 0001, Guihai Chen
IPDPS1
2007 A semantic overlay network for unstructured peer-to-peer protocols
abstract
Peer-to-peer computing has become a popular networking paradigm for file sharing, distributed computing, collaborative working, etc. The widely used unstructured peer-to-peer protocols mainly face two problems affecting their working efficiency: 1) inefficient flooding-based search, 2) topology mismatch between the overlay network and its underlying network. In this paper, we propose to organize nodes into a semantic overlay network called CON which is composed of special interest groups based on nodes' contents. CON guides the content search with semantic information so that it avoids most of the flooding cost. In order to alleviate the mismatch problem, nodes in CON initialize links according to their underlay proximity. Simulation results show that our mechanism efficiently increases the query success rate and reduces the traffic cost and query latency. We also compare CON with the similar work, which illuminates that CON performs better in many aspects.
Zhenhua Li 0001, Guihai Chen
ICPADS2