Tianyin Xu

dblp:90/7566 · DBLP profile ↗
← Back
102ranked-venue papers
10as first author
57since 2021 · last 2026
0000-0003-4443-8170ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 37 · 4 first-author · 27 since 2021Computer networks · 34 · 3 first-author · 11 since 2021Systems, architecture and hardware · 32 · 2 first-author · 21 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Who Watches the Watchers? On the Reliability of Softwarizing Cloud Application Management
Jiawei Tyler Gu, Yiming Su, Bogdan Alexandru Stoica, Xudong Sun 0013, William X. Zheng, Akond Ashfaque Ur Rahman, Chen Wang 0039, Tianyin Xu
NSDI10
2026 EROICA: Online Performance Troubleshooting for Large-scale Model Training
Yu Guan 0005, Zhiyu Yin, Sheng Cheng 0002, Chaojie Yang, Kun Qian 0004, Tianyin Xu, Yang Zhang 0102, Yong Li 0008, Dennis Cai, Ennan Zhai
NSDI7
2025 CXLfork: Fast Remote Fork over CXL Fabrics
abstract
The shared and distributed memory capabilities of the emerging Compute Express Link (CXL) interconnect urge us to rethink the traditional interfaces of system software. In this paper, we explore one such interface: remote fork using CXL-attached shared memory for cluster-wide process cloning. We present CXLfork, a remote fork interface that realizes close to zero-serialization, zero-copy process cloning across nodes over CXL fabrics. CXLfork utilizes globally-shared CXL memory for cluster-wide deduplication of process states. It also enables fine-grained control of state tiering between local and CXL memory. We use CXLfork to develop CXL-porter, an efficient horizontal autoscaler for serverless functions deployed on CXL fabrics. CXLfork minimizes cold-start overhead without sacrificing local memory. CXLfork attains restore latency close to that of a local fork, outperforming state-of-practice by 2.26x on average, and reducing local memory consumption by 87% on average.
Chloe Alverti, Stratos Psomadakis, Burak Ocalan, Shashwat Jaiswal, Tianyin Xu, Josep Torrellas
ASPLOS (2)5
2025 M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory Systems
abstract
CXL has emerged as a promising memory interface that can cost-effectively expand the capacity and bandwidth of a memory system, complementing the traditional DDR interface. However, CXL DRAM presents 2-3x longer access latency than DDR DRAM, forming a tiered-memory system that demands an effective and efficient page-migration solution. Although many page-migration solutions have been proposed for past tiered-memory systems, they have achieved limited success. To tackle the challenge of managing tiered-memory systems, this work first presents a CXL-driven profiling solution to precisely and transparently count the number of accesses to every 4KB page and 64B word in CXL DRAM. Second, using the profiling solution, this work uncovers that (1) widely used CPU-driven page-migration solutions often identify warm pages as hot pages, and (2) certain applications have sparse hot pages, where only a small percentage of words in each of these pages are frequently accessed. Besides, this work demonstrates that the performance overhead of identifying hot pages is sometimes high enough to degrade application performance. Lastly, this work presents M5, a platform designed to facilitate the development of effective CXL-driven page-migration solutions, providing hardware-based hot-page and hot-word trackers in the CXL controller. On average, M5 can identify 47% hotter pages and offer 14% higher performance than the best CPU-driven page-migration solution, even with a simple policy.
Jongyul Kim 0001, Zeduo Yu, Jiyuan Zhang 0003, Siyuan Chai 0001, Michael Jaemin Kim, Hwayong Nam, Jaehyun Park 0006, Eojin Na, Ren Wang 0001, Jung Ho Ahn, Tianyin Xu, Nam Sung Kim
ASPLOS (2)13
2025 Multi-Grained Specifications for Distributed System Model Checking and Verification
abstract
This paper presents our experience specifying and verifying the correctness of ZooKeeper, a complex and evolving distributed coordination system. We use TLA+ to model finegrained behaviors of ZooKeeper and use the TLC model checker to verify its correctness properties; we also check conformance between the model and code. The fundamental challenge is to balance the granularity of specifications and the scalability of model checking---fine-grained specifications lead to state-space explosion, while coarse-grained specifications introduce model-code gaps. To address this challenge, we write specifications with different granularities for composable modules, and compose them into mixed-grained specifications based on specific scenarios. For example, to verify code changes, we compose fine-grained specifications of changed modules and coarse-grained specifications that abstract away details of unchanged code with preserved interactions. We show that writing multi-grained specifications is a viable practice and can cope with model-code gaps without untenable state space, especially for evolving software where changes are typically local and incremental. We detected six severe bugs that violate five types of invariants and verified their code fixes; the fixes have been merged to ZooKeeper. We also improve the protocol design to make it easy to implement correctly.
Lingzhi Ouyang, Xudong Sun 0013, Ruize Tang, Yu Huang 0002, Madhav Jivrajani, Xiaoxing Ma, Tianyin Xu
EuroSys7
2025 Rethinking Tiered Storage: Talk to File Systems, Not Device Drivers
abstract
Different storage technologies motivate the development of specialized file systems tailored to specific device types. A tiered file system aggregates such device types into a single file system. We argue that the current practice of developing tiered file systems tends to lag behind that of device-specific file systems because, inherently, developers are burdened with addressing multiple device types simultaneously, rather than specializing. We propose to solve this problem using Mux, a new tiered file system that accesses different device types indirectly through device-specific file systems, rather than directly through device drivers. Despite introducing an additional indirection layer, we show that Mux significantly outperforms Strata, a research tiered file system, because it utilizes specialized production-ready file systems. Compared with direct access to per-device file systems (with no tiering), Mux adds a worst-case read latency overhead of 6.6% to 87.3%, and a write throughout overhead of 1.6% to 3.5% across devices. We contend that Mux's separation of tiering and specialization concerns enables progressive evolution and flexible integration of heterogeneous storage devices.
Jiyuan Zhang 0003, Jongyul Kim 0001, Chloe Alverti, Peizhe Liu, Weiwei Jia 0001, Tianyin Xu
HotOS6
2025 Concord: Rethinking Distributed Coherence for Software Caches in Serverless Environments
abstract
Costly accesses to global storage substantially limit the performance of serverless functions. To mitigate this overhead, data can be cached in the memory of the nodes where functions are executed. Existing caching schemes either (1) restrict a data item to be cached in a single node, causing frequent remote reads or (2) allow a data item to be cached in multiple nodes concurrently, adding substantial overhead to maintain cache coherence. Unfortunately, current approaches are suboptimal for the access patterns present in serverless workloads, which are characterized by frequent reads to small data items, strong temporal locality, and a small number of nodes that concurrently execute functions of the same application. Driven by these insights, we propose Concord, a distributed software caching system tailored to serverless environments. Concord allows multiple copies of the same data item to be cached in different nodes concurrently, allowing each cache to satisfy local reads. To maintain coherence across software caches, Concord proposes a directory-based distributed coherence protocol. The protocol is inspired by hardware cache coherence, and is enhanced to minimize coherence traffic, reduce contention points, and be robust to node failures and frequent coherence domain changes. Further, with the Concord coherence protocol, we unlock two new capabilities in serverless environments: transactional storage accesses and transparent data-aware function placement. Compared to state-of-the-art severless caching schemes, Concord running on a 16-node cluster speeds-up execution by 2.4 × and improves throughput by 1.7 ×, while using only 6.2MB of otherwise idle application memory (i.e., 4.8% of the total application memory).
Jovan Stojkovic, Chloe Alverti, Alan Andrade, Nikoleta Iliakopoulou, Hubertus Franke, Tianyin Xu, Josep Torrellas
HPCA6
2025 ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
abstract
Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench.
Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri
ICML33
2025 Large Language Models as Configuration Validators
abstract
Misconfigurations are major causes of software failures. Existing practices rely on developer-written rules or test cases to validate configuration values, which are expensive. Machine learning (ML) for configuration validation is considered a promising direction, but has been facing challenges such as the need of large-scale field data and system-specific models. Recent advances in Large Language Models (LLMs) show promise in addressing some of the long-lasting limitations of ML-based configuration validation. We present the first analysis on the feasibility and effectiveness of using LLMs for configuration validation. We empirically evaluate LLMs as configuration validators by developing a generic LLM-based configuration validation framework, named Ciri. Ciri employs effective prompt engineering with few-shot learning based on both valid configuration and misconfiguration data. Ciri checks outputs from LLMs when producing results, addressing hallucination and nondeterminism of LLMs. We evaluate Ciri's validation effectiveness on eight popular LLMs using configuration data of ten widely deployed open-source systems. Our analysis (1) confirms the potential of using LLMs for configuration validation, (2) explores design space of LLM-based validators like Ciri, and (3) reveals open challenges such as ineffectiveness in detecting certain types of misconfigurations and biases towards popular configuration parameters.
Xinyu Lian, Yinfang Chen, Runxiang Cheng, Jie Huang 0009, Parth Thakkar, Minjia Zhang, Tianyin Xu
ICSE7
2025 Fidelity of Cloud Emulators: The Imitation Game of Testing Cloud-Based Software
abstract
Modern software projects have been increasingly using cloud services as important components. The cloud-based programming practice greatly simplifies software development by harvesting cloud benefits (e.g., high availability and elasticity). However, it imposes new challenges for software testing and analysis, due to opaqueness of cloud backends and monetary cost of invoking cloud services for continuous integration and deployment. As a result, cloud emulators are developed for offline development and testing, before online testing and deployment. This paper presents a systematic analysis of cloud emulators from the perspective of cloud-based software testing. Our goal is to (1) understand the discrepancies introduced by cloud emulation with regard to software quality assurance and deployment safety and (2) address inevitable gaps between emulated and real cloud services. The analysis results are concerning. Among 255 APIs of five cloud services from Azure and Amazon Web Services (AWS), we detected discrepant behavior between the emulated and real services in 94 (37 %) of the APIs. These discrepancies lead to inconsistent testing results, threatening deployment safety, introducing false alarms, and creating debuggability issues. The root causes are diverse, including accidental implementation defects and essential emulation challenges. We discuss potential solutions and develop a practical mitigation technique to address discrepancies of cloud emulators for software testing.
Anna Mazhar, Saad Sher Alam, William X. Zheng, Yinfang Chen, Suman Nath, Tianyin Xu
ICSE6
2025 An Empirical Study of Production Incidents in Generative AI Cloud Services
abstract
The ever-increasing demand for generative artificial intelligence (GenAI) has motivated cloud-based GenAI services such as Azure OpenAI Service. Like any large-scale cloud service, failures are inevitable in cloud-based GenAI services, resulting in user dissatisfaction and significant monetary losses. However, GenAI cloud services, featured by their massive parameter scales, hardware demands, and usage patterns, present unique challenges, including generated content quality issues and privacy concerns, compared to traditional cloud services. To understand the production reliability of GenAI cloud services, we analyzed production incidents from Microsoft spanning in the past four years. Our study (1) presents the general characteristics of GenAI cloud service incidents at different stages of the incident life cycle; (2) identifies the symptoms and impacts of these incidents on GenAI cloud service quality and availability; (3) uncovers why these incidents occurred and how they were resolved; (4) discusses open research challenges in terms of incident detection, triage, and mitigation, and sheds light on potential solutions.
Haoran Yan, Yinfang Chen, Minghua Ma, Ming Wen 0001, Shan Lu 0001, Shenglin Zhang, Tianyin Xu, Rujia Wang, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Chaoyun Zhang, Dongmei Zhang 0001
ISSRE7
2025 Democratizing the Cryptocurrency Ecosystem by Just-In-Time Transformation of Mining Programs
abstract
Democracy is crucial to a cryptocurrency ecosystem, as the diversity of miners (farms, personal computers, web clients, or even cloud functions) underlays the credibility of the cryptocurrency. Among miners, web clients used to be the vast majority, e.g., 50M+ as of March 2018. As time went on, however, cryptomining was gradually monopolized by mining farms with dedicated hardware (e.g., ASICs), and web clients scaled down to ∼0.1M. To suppress mining farms, certain cryptocurrencies (like Monero) adopted new mining algorithms such as RandomX whose execution relies on general-purpose hardware architectures. Unfortunately, this further impairs web-based cryptomining as web clients cannot provide the desired architecture support to these algorithms. This paper explores how to revive software democracy of efficient web-based crypto-mining, using a novel program transformation technique termed Vectra. Vectra employs just-in-time (JIT) transformations of mining programs for web architectures; it effectively identifies and merges isomorphic instructions upon execution. Vectra ensures correct transformations based on symbolic constraints of the instructions. Real-world deployments show that Vectra reduces WASM instructions by about 7× and achieves a 3× –16× speedup for web cryptomining in diverse execution environments like PCs, mobile phones, and serverless platforms, which translates to a high (69%–274%) return-on-investment (ROI) for common users.
Wei Liu 0148, Zhenhua Li 0001, Feng Qian 0001, Feiyu Jin, Hao Lin 0005, Yannan Zheng, Xiaokang Qin, Tianyin Xu
ASE9
2025 DebCovDiff: Differential Testing of Coverage Measurement Tools on Real-World Projects
abstract
Measuring code coverage is a critical practice in software testing. Incorrect or misleading coverage information reported by automatic tools can increase the software development cost and lead to negative consequences especially for safety-critical software. Ensuring the correctness of coverage measurement tools is therefore important. Prior studies have applied various techniques to find bugs in Gcov and LLVM-cov, the two most widely used coverage tools for C/C++. However, those studies had two limiting factors. First, they used only small, often synthetic, programs, potentially missing bugs in real-world scenarios. Second, they focused only on basic line coverage, neglecting advanced metrics that are both more complex to implement and commonly required for safety-critical software.This paper presents the first empirical study of coverage measurement tools for real-world projects. We implement DebCovDiff, a testing framework that takes Debian packages as the input programs and performs differential testing of Gcov and LLVM-cov, for line coverage and two advanced coverage metrics. We design robust differential oracles to (1) filter out discrepancies arising from subtle differences in the tool output presentation, (2) overcome the nondeterministic nature of certain packages, and (3) support advanced coverage metrics. From results on 47 Debian packages, we identify 34 new bugs, including 2 crashing bugs and 32 deeper bugs that produce wrong coverage reports.
Jinghao Jia, Erkai Yu, Darko Marinov, Tianyin Xu
ASE5
2025 Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
abstract
The effectiveness of LLMs has triggered an exponential rise in their deployment, imposing substantial demands on inference clusters.Such clusters often handle numerous concurrent queries for different LLM downstream tasks.To handle multi-task settings with vast LLM parameter counts, Low-Rank Adaptation (LoRA) enables task-specific fine-tuning while sharing most of the base LLM model across tasks.Hence, it supports concurrent task serving with reduced memory requirements.However, existing designs face inefficiencies: they overlook workload heterogeneity, impose high CPU-GPU link bandwidth from frequent adapter loading, and suffer from head-of-line blocking in their schedulers.To address these challenges, we present Chameleon, a novel LLM serving system optimized for many-adapter environments.Chameleon introduces two new ideas: adapter caching and adapteraware scheduling.First, Chameleon caches popular adapters in GPU memory, minimizing adapter loading times.For caching, it uses otherwise idle GPU memory, avoiding extra memory costs.Second, Chameleon uses a non-preemptive multi-queue scheduler to efficiently account for workload heterogeneity.In this way, Chameleon simultaneously prevents head of line blocking and starvation.Under high loads, Chameleon reduces the P99 and P50 TTFT latencies by 80.7% and 48.1%, respectively, over a state-of-the-art baseline, while improving the throughput by 1.5×.
Nikoleta Iliakopoulou, Jovan Stojkovic, Chloe Alverti, Tianyin Xu, Hubertus Franke, Josep Torrellas
MICRO4
2025 STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
abstract
In cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurations are reported to be more frequent. The demand for autonomous, AI-driven reliability engineering continues to grow, as existing human-in-the-loop practices can hardly keep up with the scale of modern clouds. This paper presents STRATUS, an LLM-based multi-agent system for realizing autonomous Site Reliability Engineering (SRE) of cloud services. STRATUS consists of multiple specialized agents (e.g., for failure detection, diagnosis, mitigation), organized in a state machine to assist system-level safety reasoning and enforcement. We formalize a key safety specification of agentic SRE systems like STRATUS, termed Transactional No-Regression (TNR), which enables safe exploration and iteration. We show that TNR can effectively improve autonomous failure mitigation. STRATUS significantly outperforms state-of-the-art SRE agents in terms of success rate of failure mitigation problems in AIOpsLab and ITBench (two SRE benchmark suites), by at least 1.5 times across various models. STRATUS shows a promising path toward practical deployment of agentic systems for cloud reliability.
Yinfang Chen, Jackson Clark, Yiming Su, Noah Zheutlin, Bhavya, Rohan R. Arora, Yu Deng 0004, Saurabh Jha, Tianyin Xu
NeurIPS10
2025 Mitigating Scalability Walls of RDMA-based Container Networks
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Feng Qian 0001, Tianyin Xu, Yunhao Liu 0001, Yu Guan 0005, Shuhong Zhu, Hongfei Xu, Lanlan Xi, Ennan Zhai
NSDI5
2025 Dissecting and Streamlining the Interactive Loop of Mobile Cloud Gaming
Yang Li 0092, Jiaxing Qiu, Hongyi Wang 0009, Zhenhua Li 0001, Feng Qian 0001, Jing Yang 0052, Hao Lin 0005, Yunhao Liu 0001, Xiaokang Qin, Tianyin Xu
NSDI11
2025 EMT: An OS Framework for New Memory Translation Architectures
Siyuan Chai 0001, Jiyuan Zhang 0003, Jongyul Kim 0001, Fan Chung Graham, Jovan Stojkovic, Weiwei Jia 0001, Dimitrios Skarlatos 0002, Josep Torrellas, Tianyin Xu
OSDI10
2025 SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model Training
abstract
The flexibility and portability characteristics have made containers a popular serverless environment for large model training in recent years. Unfortunately, these advantages render the network support for containerized large model training extremely challenging, due to the high dynamics of containers, the complex interplay between underlay and overlay networks, and the stringent requirements on failure detection and localization. Existing data center network debugging tools, which rely on comprehensive or opportunistic monitoring, are either inefficient or inaccurate in this setting.
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Tianyin Xu, Yunhao Liu 0001, Jiakang Li, Shuhong Zhu, Xue Li 0024, Hongfei Xu, Ennan Zhai
SIGCOMM4
2025 Rex: Closing the language-verifier gap with safe and usable kernel extensions
Jinghao Jia, Ruowen Qin, Milo Craun, Egor Lukiyanov, Ayush Bansal, Minh Phan, Michael V. Le, Hubertus Franke, Hani Jamjoom, Tianyin Xu, Dan Williams 0001
USENIX ATC10
2024 Direct Memory Translation for Virtualized Clouds
abstract
Virtual memory translation has become a key performance bottleneck of memory-intensive workloads in virtualized cloud environments. On the x86 architecture, a nested translation needs to sequentially fetch up to 24 page table entries (PTEs). This paper presents Direct Memory Translation (DMT), a hardware-software extension for x86-based virtual memory that minimizes translation overhead while maintaining backward compatibility with x86. In DMT, the OS manages last-level PTEs in a contiguous physical memory region, termed Translation Entry Areas (TEAs). DMT establishes a direct mapping from each virtual page in a Virtual Memory Area (VMA) to the corresponding PTE in a TEA. Since processes manage memory with a handful of major VMAs, the mapping can be maintained per VMA and effectively stored in a few dedicated registers. DMT further optimizes virtualized memory translation via guest-host cooperation by directly allocating guest TEAs in physical memory, bypassing intermediate virtualization layers. DMT is inherently scalable---it takes one, two, and three memory references in native, virtualized, and nested virtualized setups. Its scalability enables hardware-assisted translation for nested virtualization. Our evaluation shows that DMT significantly speeds up page walks by an average of 1.58x (1.65x with THP) in a virtualized setup, resulting in 1.20x (1.14x with THP) speedup of application execution on average.
Jiyuan Zhang 0003, Weiwei Jia 0001, Siyuan Chai 0001, Peizhe Liu, Jongyul Kim 0001, Tianyin Xu
ASPLOS (2)6
2024 Automatic Root Cause Analysis via Large Language Models for Cloud Incidents
abstract
Ensuring the reliability and availability of cloud services necessitates efficient root cause analysis (RCA) for cloud incidents. Traditional RCA methods, which rely on manual investigations of data sources such as logs and traces, are often laborious, error-prone, and challenging for on-call engineers. In this paper, we introduce RCACopilot, an innovative on-call system empowered by the large language model for automating RCA of cloud incidents. RCACopilot matches incoming incidents to corresponding incident handlers based on their alert types, aggregates the critical runtime diagnostic information, predicts the incident's root cause category, and provides an explanatory narrative. We evaluate RCACopilot using a real-world dataset consisting of a year's worth of incidents from Microsoft. Our evaluation demonstrates that RCACopilot achieves RCA accuracy up to 0.766. Furthermore, the diagnostic information collection component of RCACopilot has been successfully in use at Microsoft for over four years.
Yinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang 0006, Xin Gao 0017, Liu Shi, Yunjie Cao, Xuedong Gao, Ming Wen 0001, Jun Zeng 0006, Supriyo Ghosh, Xuchao Zhang, Chaoyun Zhang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Tianyin Xu
EuroSys18
2024 EcoFaaS: Rethinking the Design of Serverless Environments for Energy Efficiency
abstract
While serverless computing is increasingly popular, its energy and power consumption behavior is hardly explored. In this work, we perform a thorough characterization of the serverless environment and observe that it poses a set of challenges not effectively handled by existing energy-management schemes. Short serverless functions execute in opaque virtualized sandboxes, are idle for a large fraction of their invocation time, context switch frequently, and are co-located in a highly dynamic manner with many other functions of diverse properties. These features are a radical shift from more traditional application environments and require a new approach to manage energy and power. Driven by these insights, we design EcoFaaS, the first energy management framework for serverless environments. EcoFaaS takes a user-provided end-to-end application Service Level Objective (SLO). It then splits the SLO into per-function deadlines that minimize the total energy consumption. Based on the computed deadlines, EcoFaaS sets the optimal per-invocation core frequency using a prediction algorithm. The algorithm performs a fine-grained analysis of the execution time of each invocation, while taking into account the specific invocation inputs. To maximize efficiency, EcoFaaS splits the cores in a server into multiple Core Pools, where all the cores in a pool run at the same frequency and are controlled by a single scheduler. EcoFaaS dynamically changes the sizes and frequencies of the pools based on the current system state. We implement EcoFaaS on two open-source serverless platforms (OpenWhisk and KNative) and evaluate it using diverse serverless applications. Compared to state-of-the-art energy-management systems, EcoFaaS reduces the total energy consumption of serverless clusters by $42 \%$ while simultaneously reducing the tail latency by $34.8 \%$.
Jovan Stojkovic, Nikoleta Iliakopoulou, Tianyin Xu, Hubertus Franke, Josep Torrellas
ISCA3
2024 Rethinking Process Management for Interactive Mobile Systems
abstract
Modern mobile systems are featured by their increasing interactivity with users, which however is accompanied by a severe side effect---users constantly suffer from slow UI responsiveness (SUR). To date, the community have limited understandings of this issue for the challenges of comprehensively measuring SUR events on massive mobile devices. As a major Android phone vendor, in this paper we close the knowledge gap by conducting the first large-scale, long-term measurement study on SUR with 47M devices. Our study identifies the critical factors that lead to SUR from the perspectives of device, system, application, and app market. Most importantly, we note that the largest root cause lies in the wide existence of "hogging" apps, which persistently occupy an unreasonable amount of system resources by leveraging the optimistic design of Android process management. We have built on the insights to remodel Android process states by fully considering their time-sensitive transitions and the actual behaviors of processes, with remarkable real-world impact---the occurrences of SUR are reduced by 60%, together with 10.7% saving of battery consumption.
Jianwei Zheng 0003, Zhenhua Li 0001, Feng Qian 0001, Wei Liu 0148, Hao Lin 0005, Yunhao Liu 0001, Tianyin Xu, Nan Zhang 0018, Cang Zhang
MobiCom7
2024 Anvil: Verifying Liveness of Cluster Management Controllers
Xudong Sun 0013, Jiawei Tyler Gu, Zicheng Ma, Tej Chajed, Jon Howell, Andrea Lattuada 0001, Oded Padon, Lalith Suresh 0001, Adriana Szekeres, Tianyin Xu
OSDI11
2024 vSoC: Efficient Virtual System-on-Chip on Heterogeneous Hardware
abstract
Emerging mobile apps such as UHD video and AR/VR access diverse high-throughput hardware devices, e.g., video codecs, cameras, and image processors. However, today's mobile emulators exhibit poor performance when emulating these devices. We pinpoint the major reason to be the discrepancy between the guest's and host's memory architectures for hardware devices, i.e., the mobile guest's centralized memory on a system-on-chip (SoC) versus the PC/server's separated memory modules on individual hardware. Such a discrepancy makes the shared virtual memory (SVM) architecture of mobile emulators highly inefficient.
Jiaxing Qiu, Yang Li 0092, Zhenhua Li 0001, Feng Qian 0001, Hao Lin 0005, Haitao Su, Yunhao Liu 0001, Tianyin Xu
SOSP11
2024 Fast (Trapless) Kernel Probes Everywhere
Jinghao Jia, Michael V. Le, Salman Ahmed 0001, Dan Williams 0001, Hani Jamjoom, Tianyin Xu
USENIX ATC6
2024 Trinity: High-Performance and Reliable Mobile Emulation through Graphics Projection
abstract
Mobile emulation, which creates full-fledged software mobile devices on a physical PC/server, is pivotal to the mobile ecosystem. Unfortunately, existing mobile emulators perform poorly on graphics-intensive apps in terms of efficiency and compatibility. To address this, we introduce graphics projection , a novel graphics virtualization mechanism that adds a small-size projection space inside the guest memory, which processes graphics operations involving control contexts and resource handles without host interactions. While enhancing performance, the decoupled and asynchronous guest/host control flows introduced by graphics projection can significantly complicate emulators’ reliability issue diagnosis when faced with a variety of uncommon or non-standard app behaviors in the wild, hindering practical deployment in production. To overcome this drawback, we develop an automatic reliability issue analysis pipeline that distills the critical code paths across the guest and host control flows by runtime quarantine and state introspection. The resulting new Android emulator, dubbed Trinity, exhibits an average of 97% native hardware performance and 99.3% reliable app support, in some cases outperforming other emulators by more than an order of magnitude.
Hao Lin 0005, Zhenhua Li 0001, Yunhao Liu 0001, Feng Qian 0001, Tianyin Xu, Xiaokang Qin
ACM Trans. Comput. Syst.6
2023 HugeGPT: Storing Guest Page Tables on Host Huge Pages to Accelerate Address Translation
abstract
Expensive page table walks triggered by frequent TLB misses have incurred major performance bottlenecks for data-intensive workloads that are dominated by memory accesses with weak locality. Since it is hard to reduce TLB misses for such workloads, reducing page table walk overhead (i.e., the overhead of each TLB miss) is an increasingly important direction for improving application performance. The direction is more compelling for workloads running in virtual machines (VMs). In virtualized environments, each TLB miss triggers a two-dimensional page table walk, which has a significantly higher overhead than that on native systems. This paper presents HugeGPT, a software approach to reducing two-dimensional page table walk overhead in virtualized environments. HugeGPT ensures that page tables used in guest systems are physically held in the huge pages formed in the host. This brings two-fold benefits: 1) the number of steps walking down the host page table is reduced; 2) the misses of page walk caches incurred by accessing the leaf nodes on host page tables can be eliminated. Extensive evaluation based on the prototype implementation and diverse real-world applications shows that HugeGPT can efficiently reduce address translation overhead and improve application performance in virtualized clouds.
Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Yiming Du, Xiaoning Ding, Tianyin Xu
PACT6
2023 Towards GPU Memory Efficiency for Distributed Training at Scale
abstract
The scale of deep learning models has grown tremendously in recent years. State-of-the-art models have reached billions of parameters and terabyte-scale model sizes. Training of these models demands memory bandwidth and capacity that can only be accommodated distributively over hundreds to thousands of GPUs. However, large-scale distributed training suffers from GPU memory inefficiency, such as memory under-utilization and out-of-memory events (OOMs). There is a lack of understanding of actual GPU memory behavior of distributed training on terabyte-size models, which hinders the development of effective solutions to such inefficiency. In this paper, we present a systematic analysis of GPU memory behavior of large-scale distributed training jobs in production at Meta. Our analysis is based on offline training jobs of multi-terabyte Deep Learning Recommendation Models from one of Meta's largest production clusters. We measure GPU memory inefficiency; characterize GPU memory utilization, and provide fine-grained GPU memory usage analysis. We further show how to build on the understanding to develop a practical GPU provisioning system in production.
Runxiang Cheng, Chris Cai, Selman Yilmaz, Malay Bag, Mrinmoy Ghosh, Tianyin Xu
SoCC7
2023 Fail through the Cracks: Cross-System Interaction Failures in Modern Cloud Systems
abstract
Modern cloud systems are orchestrations of independent and interacting (sub-)systems, each specializing in important services (e.g., data processing, storage, resource management, etc.). Hence, cloud system reliability is affected not only by the reliability of each individual system, but also by the interplay between these systems. We observe that many recent production incidents of cloud systems are manifested through interactions across the system boundaries. However, there is a lack of systematic understanding of this emerging mode of failures, which we term as cross-system interaction failures (or CSI failures). This hinders the development of better design, integration practices, and new tooling.
Lilia Tang, Chaitanya Bhandari, Yongle Zhang 0007, Anna Karanika, Shuyang Ji, Indranil Gupta, Tianyin Xu
EuroSys7
2023 Kernel extension verification is untenable
abstract
The emergence of verified eBPF bytecode is ushering in a new era of safe kernel extensions. In this paper, we argue that eBPF's verifier---the source of its safety guarantees---has become a liability. In addition to the well-known bugs and vulnerabilities stemming from the complexity and ad hoc nature of the in-kernel verifier, we highlight a concerning trend in which escape hatches to unsafe kernel functions (in the form of helper functions) are being introduced to bypass verifier-imposed limitations on expressiveness, unfortunately also bypassing its safety guarantees. We propose safe kernel extension frameworks using a balance of not just static but also lightweight runtime techniques. We describe a design centered around kernel extensions in safe Rust that will eliminate the need of the in-kernel verifier, improve expressiveness, allow for reduced escape hatches, and ultimately improve the safety of kernel extensions.
Jinghao Jia, Raj Sahu, Adam Oswald, Dan Williams 0001, Michael V. Le, Tianyin Xu
HotOS6
2023 Memory-Efficient Hashed Page Tables
abstract
Conventional radix-tree page tables have scalability challenges, as address translation following a TLB miss potentially requires multiple memory accesses in sequence. An alternative is hashed page tables (HPTs) where, conceptually, address translation needs only one memory access. Traditionally, HPTs have been shunned due to high costs of handling conflicts and other limitations. However, recent advances have made HPTs compelling. Still, a major issue in HPT designs is their requirement for substantial contiguous physical memory.This paper addresses this problem. To minimize HPTs’ contiguous memory needs, it introduces the Logical to Physical (L2P) Table and the use of Dynamically-Changing Chunk Sizes. These techniques break down the HPT into discontiguous physical-memory chunks. In addition, the paper also introduces two techniques that minimize HPTs’ total memory needs and, indirectly, reduce the memory contiguity requirements. These techniques are In-place Page Table Resizing and Per-way Resizing. We call our complete design Memory-Efficient HPTs (ME-HPTs). Compared to state-of-the-art HPTs, ME-HPTs: (i) reduce the contiguous memory allocation needs by 92% on average, and (ii) improve the performance by 8.9% on average. For the two most demanding workloads, the contiguous memory requirements decrease from 64MB to 1MB. In addition, compared to state-of-the-art radix-tree page tables, ME-HPTs achieve an average speedup of 1.23× (without huge pages) and 1.28× (with huge pages).
Jovan Stojkovic, Namrata Mantri, Dimitrios Skarlatos 0002, Tianyin Xu, Josep Torrellas
HPCA4
2023 SpecFaaS: Accelerating Serverless Applications with Speculative Function Execution
abstract
Serverless computing has emerged as a popular cloud computing paradigm. Serverless environments are convenient to users and efficient for cloud providers. However, they can induce substantial application execution overheads, especially in applications with many functions.In this paper, we propose to accelerate serverless applications with a novel approach based on software-supported speculative execution of functions. Our proposal is termed Speculative Function-as-a-Service (SpecFaaS). It is inspired by out-of-order execution in modern processors, and is grounded in a characterization analysis of FaaS applications. In SpecFaaS, functions in an application are executed early, speculatively, before their control and data dependences are resolved. Control dependences are predicted like in pipeline branch prediction, and data dependences are speculatively satisfied with memoization. With this support, the execution of downstream functions is overlapped with that of upstream functions, substantially reducing the end-to-end execution time of applications. We prototype SpecFaaS on Apache OpenWhisk, an open-source serverless computing platform. For a set of applications in a warmed-up environment, SpecFaaS attains an average speedup of 4.6×. Further, on average, the application throughput increases by 3.9× and the tail latency decreases by 58.7%.
Jovan Stojkovic, Tianyin Xu, Hubertus Franke, Josep Torrellas
HPCA2
2023 Test Selection for Unified Regression Testing
abstract
Today's software failures have two dominating root causes: code bugs and misconfigurations. To combat failure-inducing software changes, unified regression testing (URT) is needed to synergistically test the changed code and all changed production configurations for deployment reliability. However, URT could incur high cost, as it needs to run a large number of tests under multiple configurations. Regression test selection (RTS) can reduce regression testing cost. Unfortunately, no existing RTS technique reasons about code and configuration changes collectively. We introduce Unified Regression Test Selection (uRTS) to effectively reduce the cost of URT. uRTS supports project changes on 1) code only, 2) configurations only, and 3) both code and configurations. It selects regular tests and configuration tests with a unified selection algorithm. The uRTS algorithm analyzes code and configuration dependencies of each test across runs and across configurations. uRTS provides the same safety guarantee as the state-of-the-art RTS while selecting fewer tests and, more importantly, reducing the end-to-end testing time. We implemented uRTS on top of Ekstazi (a RTS tool for code changes) and Ctest (a configuration testing framework). We evaluate uRTS on hundreds of code revisions and dozens of configurations of five large projects. The results show that uRTS reduces the end-to-end testing time, on average, by 3.64X compared to executing all tests and 1.87X compared to a competitive reference solution that directly extends RTS for URT.
Shuai Wang 0035, Xinyu Lian, Darko Marinov, Tianyin Xu
ICSE4
2023 MXFaaS: Resource Sharing in Serverless Environments for Parallelism and Efficiency
abstract
Although serverless computing is a popular paradigm, current serverless environments have high overheads. Recently, it has been shown that serverless workloads frequently exhibit bursts of invocations of the same function. Such pattern is not handled well in current platforms. Supporting it efficiently can speed-up serverless execution substantially.
Jovan Stojkovic, Tianyin Xu, Hubertus Franke, Josep Torrellas
ISCA2
2023 Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices
abstract
The ever-growing demands for memory with larger capacity and higher bandwidth have driven recent innovations on memory expansion and disaggregation technologies based on Compute eXpress Link (CXL). Especially, CXL-based memory expansion technology has recently gained notable attention for its ability not only to economically expand memory capacity and bandwidth but also to decouple memory technologies from a specific memory interface of the CPU. However, since CXL memory devices have not been widely available, they have been emulated using DDR memory in a remote NUMA node. In this paper, for the first time, we comprehensively evaluate a true CXL-ready system based on the latest 4th-generation Intel Xeon CPU with three CXL memory devices from different manufacturers. Specifically, we run a set of microbenchmarks not only to compare the performance of true CXL memory with that of emulated CXL memory but also to analyze the complex interplay between the CPU and CXL memory in depth. This reveals important differences between emulated CXL memory and true CXL memory, some of which will compel researchers to revisit the analyses and proposals from recent work. Next, we identify opportunities for memory-bandwidth-intensive applications to benefit from the use of CXL memory. Lastly, we propose a CXL-memory-aware dynamic page allocation policy, Caption to more efficiently use CXL memory as a bandwidth expander. We demonstrate that Caption can automatically converge to an empirically favorable percentage of pages allocated to CXL memory, which improves the performance of memory-bandwidth-intensive applications by up to 24% when compared to the default page allocation policy designed for traditional NUMA systems.
Zeduo Yu, Reese Kuper, Chihun Song, Jinghan Huang 0001, Houxiang Ji, Siddharth Agarwal, Jiaqi Lou, Ipoom Jeong, Ren Wang 0001, Jung Ho Ahn, Tianyin Xu, Nam Sung Kim
MICRO13
2023 Virtual Device Farms for Mobile App Testing at Scale: A Pursuit for Fidelity, Efficiency, and Accessibility
abstract
Virtual devices based on device emulation have been widely used in lab research of mobile app testing for their efficiency and low cost. However, it remains controversial to use virtual devices for app testing in industry, given the inherent difficulties of high-fidelity emulation across diverse mobile systems and devices. Hence, mobile app companies still rely on physical device farms or services like AWS Device Farm.
Hao Lin 0005, Jiaxing Qiu, Hongyi Wang 0009, Zhenhua Li 0001, Liangyi Gong, Yunhao Liu 0001, Feng Qian 0001, Zhao Zhang 0001, Tianyin Xu
MobiCom11
2023 Push-Button Reliability Testing for Cloud-Backed Applications with Rainmaker
Yinfang Chen, Xudong Sun 0013, Suman Nath, Tianyin Xu
NSDI5
2023 Relational Debugging - Pinpointing Root Causes of Performance Problems
Xiang Ren 0003, Sitao Wang, Zhuqi Jin, David Lion, Adrian Chiu, Tianyin Xu, Ding Yuan 0004
OSDI6
2023 Acto: Automatic End-to-End Testing for Operation Correctness of Cloud System Management
abstract
Cloud systems are increasingly being managed by operation programs termed operators, which automate tedious, human-based operations. Operators of modern management platforms like Kubernetes, Twine, and ECS implement declarative interfaces based on the state-reconciliation principle. An operation declares a desired system state and the operator automatically reconciles the system to that declared state.
Jiawei Tyler Gu, Xudong Sun 0013, Yuxuan Jiang 0016, Chen Wang 0039, Mandana Vaziri, Owolabi Legunsen, Tianyin Xu
SOSP8
2023 Visual-Aware Testing and Debugging for Web Performance Optimization
abstract
Web performance optimization services, or web performance optimizers (WPOs), play a critical role in today’s web ecosystem by improving page load speed and saving network traffic. However, WPOs are known for introducing visual distortions that disrupt the users’ web experience. Unfortunately, visual distortions are hard to analyze, test, and debug, due to their subjective measure, dynamic content, and sophisticated WPO implementations.
Xinlei Yang, Wei Liu 0148, Hao Lin 0005, Zhenhua Li 0001, Feng Qian 0001, Xianlong Wang 0003, Yunhao Liu 0001, Tianyin Xu
WWW8
2022 Parallel virtualized memory translation with nested elastic cuckoo page tables
abstract
A major reason why nested or virtualized address translations are slow is because current systems organize page tables in a multi-level tree that is accessed in a sequential manner. A nested translation may potentially require up to twenty-four sequential memory accesses. To address this problem, this paper presents the first page table design that supports parallel nested address translation. The design is based on using hashed page tables (HPTs) for both guest and host. However, directly extending a native HPT design to a nested environment leads to minor gains. Instead, our design solves a new set of challenges that appear in nested environments. Our scheme eliminates all but three of the potentially twenty-four sequential steps of a nested translation—while judiciously limiting the number of parallel memory accesses issued to avoid over-consuming cache bandwidth. As a result, compared to conventional nested radix tables, our design speeds-up the execution of a set of applications by an average of 1.19x (for 4KB pages) and 1.24x (when huge pages are used). In addition, we also show a migration path from current nested radix page tables to our design.
Jovan Stojkovic, Dimitrios Skarlatos 0002, Apostolos Kokolis, Tianyin Xu, Josep Torrellas
ASPLOS4
2022 Verified programs can party: optimizing kernel extensions via post-verification merging
abstract
Operating system (OS) extensions are more popular than ever. For example, Linux BPF is marketed as a "superpower" that allows user programs to be downloaded into the kernel, verified to be safe and executed at kernel hook points. So, BPF extensions have high performance and are often placed at performance-critical paths for tracing and filtering.
Hsuan-Chi Kuo, Kai-Hsun Chen, Yicheng Lu, Dan Williams 0001, Sibin Mohan, Tianyin Xu
EuroSys6
2022 Forensic Analysis of Configuration-based Attacks
Muhammad Adil Inam, Wajih Ul Hassan, Ali Ahad, Adam Bates 0001, Rashid Tahir, Tianyin Xu, Fareed Zaffar
NDSS6
2022 Automatic Reliability Testing For Cluster Management Controllers
Xudong Sun 0013, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan, Ramnatthan Alagappan, Michael Gasch, Lalith Suresh 0001, Tianyin Xu
OSDI8
2022 Trinity: High-Performance Mobile Emulation through Graphics Projection
Hao Lin 0005, Zhenhua Li 0001, Chengen Huang, Yunhao Liu 0001, Feng Qian 0001, Liangyi Gong, Tianyin Xu
OSDI8
2022 Mobile access bandwidth in practice: measurement, analysis, and implications
abstract
Recent advances in mobile technologies such as 5G and WiFi 6E do not seem to deliver the promised mobile access bandwidth. To effectively characterize mobile access bandwidth in the wild, we work with a major commercial mobile bandwidth testing app to analyze mobile access bandwidths of 3.54M end users in China, based on fine-grained measurement and diagnostic information. Our analysis presents a surprising and frustrating fact---in the past two years, the average WiFi bandwidth remains largely unchanged, while the average 4G/5G bandwidth decreases remarkably. Our analysis further reveals the root causes---the bottlenecks in the underlying infrastructure (e.g., devices and wired Internet access) and side effects of aggressively migrating radio resources from 4G to 5G---with implications on closing the technology gaps. Additionally, our analysis provides insights on building ultra-fast, ultra-light bandwidth testing services (BTSes) at scale. Our new design dramatically reduces the test time of the commercial BTS from 10 seconds to 1 second on average, with a 15× reduction on the backend cost.
Xinlei Yang, Hao Lin 0005, Zhenhua Li 0001, Feng Qian 0001, Xingyao Li, Zhiming He, Xianlong Wang 0003, Yunhao Liu 0001, Zhi Liao, Daqiang Hu, Tianyin Xu
SIGCOMM12
2022 DepFast: Orchestrating Code of Quorum Systems
Xuhao Luo, Weihai Shen, Shuai Mu 0001, Tianyin Xu
USENIX ATC4
2021 Reasoning about modern datacenter infrastructures using partial histories
abstract
Modern datacenter infrastructures are increasingly architected as a cluster of loosely coupled services. The cluster states are typically maintained in a logically centralized, strongly consistent data store (e.g., ZooKeeper, Chubby and etcd), while the services learn about the evolving state by reading from the data store, or via a stream of notifications. However, it is challenging to ensure services are correct, even in the presence of failures, networking issues, and the inherent asynchrony of the distributed system. In this paper, we identify that partial histories can be used to effectively reason about correctness for individual services in such distributed infrastructure systems. That is, individual services make decisions based on observing only a subset of changes to the world around them. We show that partial histories, when applied to distributed infrastructures, have immense explanatory power and utility over the state of the art. We discuss the implications of partial histories and sketch tooling for reasoning about distributed infrastructure systems.
Xudong Sun 0013, Lalith Suresh 0001, Aishwarya Ganesan, Ramnatthan Alagappan, Michael Gasch, Lilia Tang, Tianyin Xu
HotOS7
2021 Fail-slow fault tolerance needs programming support
abstract
The need for fail-slow fault tolerance in modern distributed systems is highlighted by the increasingly reported fail-slow hardware/software components that lead to poor performance system-wide. We argue that fail-slow fault tolerance not only needs new distributed protocol designs, but also desires programming support for implementing and verifying fail-slow fault-tolerant code. Our observation is that the inability of tolerating fail-slow faults in existing distributed systems is often rooted in the implementations and is difficult to understand and debug. We designed the Dependably Fast Library (DepFast) for implementing fail-slow tolerant distributed systems. DepFast provides expressive interfaces for taking control of possible fail-slow points in the program to prevent unexpected slowness propagation once and for all. We use DepFast to implement a distributed replicated state machine (RSM) and show that it can tolerate various types of fail-slow faults that affect existing RSM implementations.
Andrew Yoo, Yuanli Wang, Ritesh Sinha, Shuai Mu 0001, Tianyin Xu
HotOS5
2021 An Evolutionary Study of Configuration Design and Implementation in Cloud Systems
abstract
Many techniques were proposed for detecting software misconfigurations in cloud systems and for diagnosing unintended behavior caused by such misconfigurations. Detection and diagnosis are steps in the right direction: misconfigurations cause many costly failures and severe performance issues. But, we argue that continued focus on detection and diagnosis is symptomatic of a more serious problem: configuration design and implementation are not yet first-class software engineering endeavors in cloud systems. Little is known about how and why developers evolve configuration design and implementation, and the challenges that they face in doing so. This paper presents a source-code level study of the evolution of configuration design and implementation in cloud systems. Our goal is to understand the rationale and developer practices for revising initial configuration design/implementation decisions, especially in response to consequences of misconfigurations. To this end, we studied 1178 configuration-related commits from a 2.5 year version-control history of four large-scale, actively-maintained open-source cloud systems (HDFS, HBase, Spark, and Cassandra). We derive new insights into the software configuration engineering process. Our results motivate new techniques for proactively reducing misconfigurations by improving the configuration design and implementation process in cloud systems. We highlight a number of future research directions.
Yuanliang Zhang, Haochen He, Owolabi Legunsen, Shanshan Li 0001, Wei Dong 0006, Tianyin Xu
ICSE6
2021 Test-case prioritization for configuration testing
abstract
Configuration changes are among the dominant causes of failures of large-scale software system deployment. Given the velocity of configuration changes, typically at the scale of hundreds to thousands of times daily in modern cloud systems, checking these configuration changes is critical to prevent failures due to misconfigurations. Recent work has proposed configuration testing, Ctest, a technique that tests configuration changes together with the code that uses the changed configurations. Ctest can automatically generate a large number of ctests that can effectively detect misconfigurations, including those that are hard to detect by traditional techniques. However, running ctests can take a long time to detect misconfigurations. Inspired by traditional test-case prioritization (TCP) that aims to reorder test executions to speed up detection of regression code faults, we propose to apply TCP to reorder ctests to speed up detection of misconfigurations. We extensively evaluate a total of 84 traditional and novel ctest-specific TCP techniques. The experimental results on five widely used cloud projects demonstrate that TCP can substantially speed up misconfiguration detection. Our study provides guidelines for applying TCP to configuration testing in practice.
Runxiang Cheng, Lingming Zhang 0001, Darko Marinov, Tianyin Xu
ISSTA4
2021 Fast and Light Bandwidth Testing for Internet Users
Xinlei Yang, Xianlong Wang 0003, Zhenhua Li 0001, Yunhao Liu 0001, Feng Qian 0001, Liangyi Gong, Tianyin Xu
NSDI8
2021 A nationwide study on cellular reliability: measurement, analysis, and enhancements
abstract
With recent advances on cellular technologies (such as 5G) that push the boundary of cellular performance, cellular reliability has become a key concern of cellular technology adoption and deployment. However, this fundamental concern has never been addressed due to the challenges of measuring cellular reliability on mobile devices and the cost of conducting large-scale measurements. This paper closes the knowledge gap by presenting the first large-scale, in-depth study on cellular reliability with more than 70 million Android phones across 34 different hardware models. Our study identifies the critical factors that affect cellular reliability and clears up misleading intuitions indicated by common wisdom. In particular, our study pinpoints that software reliability defects are among the main root causes of cellular data connection failures. Our work provides actionable insights for improving cellular reliability at scale. More importantly, we have built on our insights to develop enhancements that effectively address cellular reliability issues with remarkable real-world impact---our optimizations on Android's cellular implementations have effectively reduced 40% cellular connection failures for 5G phones and 36% failure duration across all phones.
Yang Li 0092, Hao Lin 0005, Zhenhua Li 0001, Yunhao Liu 0001, Feng Qian 0001, Liangyi Gong, Xianlong Xin, Tianyin Xu
SIGCOMM8
2021 Vet: identifying and avoiding UI exploration tarpits
abstract
Despite over a decade of research, it is still challenging for mobile UI testing tools to achieve satisfactory effectiveness, especially on industrial apps with rich features and large code bases. Our experiences suggest that existing mobile UI testing tools are prone to exploration tarpits, where the tools get stuck with a small fraction of app functionalities for an extensive amount of time. For example, a tool logs out an app at early stages without being able to log back in, and since then the tool gets stuck with exploring the app’s pre-login functionalities (i.e., exploration tarpits) instead of its main functionalities. While tool vendors/users can manually hardcode rules for the tools to avoid specific exploration tarpits, these rules can hardly generalize, being fragile in face of diverted testing environments, fast app iterations, and the demand of batch testing product lines. To identify and resolve exploration tarpits, we propose VET, a general approach including a supporting system for the given specific Android UI testing tool on the given specific app under test (AUT). VET runs the tool on the AUT for some time and records UI traces, based on which VET identifies exploration tarpits by recognizing their patterns in the UI traces. VET then pinpoints the actions (e.g., clicking logout) or the screens that lead to or exhibit exploration tarpits. In subsequent test runs, VET guides the testing tool to prevent or recover from exploration tarpits. From our evaluation with state-of-the-art Android UI testing tools on popular industrial apps, VET identifies exploration tarpits that cost up to 98.6% testing time budget. These exploration tarpits reveal not only limitations in UI exploration strategies but also defects in tool implementations. VET automatically addresses the identified exploration tarpits, enabling each evaluated tool to achieve higher code coverage and improve crash-triggering capabilities.
Wei Yang 0013, Tianyin Xu, Tao Xie 0001
ESEC/SIGSOFT FSE3
2021 Static detection of silent misconfigurations with deep interaction analysis
abstract
The behavior of large systems is guided by their configurations: users set parameters in the configuration file to dictate which corresponding part of the system code is executed. However, it is often the case that, although some parameters are set in the configuration file, they do not influence the system runtime behavior, thus failing to meet the user’s intent. Moreover, such misconfigurations rarely lead to an error message or raising an exception. We introduce the notion of silent misconfigurations which are prohibitively hard to identify due to (1) lack of feedback and (2) complex interactions between configurations and code. This paper presents ConfigX, the first tool for the detection of silent misconfigurations. The main challenge is to understand the complex interactions between configurations and the code that they affected. Our goal is to derive a specification describing non-trivial interactions between the configuration parameters that lead to silent misconfigurations. To this end, ConfigX uses static analysis to determine which parts of the system code are associated with configuration parameters. ConfigX then infers the connections between configuration parameters by analyzing their associated code blocks. We design customized control- and data-flow analysis to derive a specification of configurations. Additionally, we conduct reachability analysis to eliminate spurious rules to reduce false positives. Upon evaluation on five real-world datasets across three widely-used systems, Apache, vsftpd, and PostgreSQL, ConfigX detected more than 2200 silent misconfigurations. We additionally conducted a user study where we ran ConfigX on misconfigurations reported on user forums by real-world users. ConfigX easily detected issues and suggested repairs for those misconfigurations. Our solutions were accepted and confirmed in the interaction with the users, who originally posted the problems.
Jialu Zhang 0002, Ruzica Piskac, Ennan Zhai, Tianyin Xu
Proc. ACM Program. Lang.4
2020 Elastic Cuckoo Page Tables: Rethinking Virtual Memory Translation for Parallelism
abstract
The unprecedented growth in the memory needs of emerging memory-intensive workloads has made virtual memory translation a major performance bottleneck. To address this problem, this paper introduces Elastic Cuckoo Page Tables, a novel page table design that transforms the sequential pointer-chasing operation used by conventional multi-level radix page tables into fully-parallel look-ups. The resulting design harvests, for the first time, the benefits of memory level parallelism for address translation. Elastic cuckoo page tables use Elastic Cuckoo Hashing, a novel extension of cuckoo hashing that supports efficient page table resizing. Elastic cuckoo page tables efficiently resolve hash collisions, provide process-private page tables, support multiple page sizes and page sharing among processes, and dynamically adapt page table sizes to meet application requirements. We evaluate elastic cuckoo page tables with full-system simulations of an 8-core processor using a set of graph analytics, bioinformatics, HPC, and system workloads. Elastic cuckoo page tables reduce the address translation overhead by an average of 41% over conventional radix page tables. The result is a 3-18% speed-up in application execution.
Dimitrios Skarlatos 0002, Apostolos Kokolis, Tianyin Xu, Josep Torrellas
ASPLOS3
2020 Lock-Free Collaboration Support for Cloud Storage Services with Operation Inference and Transformation
Minghao Zhao 0001, Zhenhua Li 0001, Ennan Zhai, Feng Qian 0001, Yunhao Liu 0001, Tianyin Xu
FAST8
2020 Draco: Architectural and Operating System Support for System Call Security
abstract
System call checking is extensively used to protect the operating system kernel from user attacks. However, existing solutions such as Seccomp execute lengthy rule-based checking programs against system calls and their arguments, leading to substantial execution overhead.To minimize checking overhead, this paper proposes Draco, a new architecture that caches system call IDs and argument values after they have been checked and validated. System calls are first looked-up in a special cache and, on a hit, skip all checks. We present both a software and a hardware implementation of Draco. The latter introduces a System Call Lookaside Buffer (SLB) to keep recently-validated system calls, and a System Call Target Buffer to preload the SLB in advance. In our evaluation, we find that the average execution time of macro and micro benchmarks with conventional Seccomp checking is 1.14× and 1.25× higher, respectively, than on an insecure baseline that performs no security checks. With our software Draco, the average execution time reduces to 1.10× and 1.18× higher, respectively, than on the insecure baseline. With our hardware Draco, the execution time is within 1% of the insecure baseline.
Dimitrios Skarlatos 0002, Qingrong Chen, Jianyan Chen, Tianyin Xu, Josep Torrellas
MICRO4
2020 Experience: aging or glitching? why does android stop responding and what can we do about it?
abstract
Almost every Android user has unsatisfying experiences regarding responsiveness, in particular Application Not Responding (ANR) and System Not Responding (SNR) that directly disrupt user experience. Unfortunately, the community have limited understanding of the prevalence, characteristics, and root causes of unresponsiveness. In this paper, we make an in-depth study of ANR and SNR at scale based on fine-grained system-level traces crowdsourced from 30,000 Android systems. We find that ANR and SNR occur prevalently on all the studied 15 hardware models, and better hardware does not seem to relieve the problem. Moreover, as Android evolves from version 7.0 to 9.0, there are fewer ANR events but more SNR events. Most importantly, we uncover multifold root causes of ANR and SNR and pinpoint the largest inefficiency which roots in Android's flawed implementation of Write Amplification Mitigation (WAM). We design a practical approach to eliminating this largest root cause; after large-scale deployment, it reduces almost all (>99%) ANR and SNR caused by WAM while only decreasing 3% of the data write speed. In addition, we document important lessons we have learned from this study, and have also released our measurement code/data to the research community.
Hao Lin 0005, Cai Liu, Zhenhua Li 0001, Feng Qian 0001, Yunhao Liu 0001, Nian Xiang Sun, Tianyin Xu
MobiCom8
2020 Testing Configuration Changes in Context to Prevent Production Failures
Xudong Sun 0013, Runxiang Cheng, Jianyan Chen, Elaine Ang, Owolabi Legunsen, Tianyin Xu
OSDI6
2020 Live forensics for HPC systems: a case study on distributed storage systems
abstract
Large-scale high-performance computing systems frequently experience a wide range of failure modes, such as reliability failures (e.g., hang or crash), and resource overload-related failures (e.g., congestion collapse), impacting systems and applications. Despite the adverse effects of these failures, current systems do not provide methodologies for proactively detecting, localizing, and diagnosing failures. We present Kaleidoscope, a near real-time failure detection and diagnosis framework, consisting of of hierarchical domain-guided machine learning models that identify the failing components, the corresponding failure mode, and point to the most likely cause indicative of the failure in near real-time (within one minute of failure occurrence). Kaleidoscope has been deployed on Blue Waters supercomputer and evaluated with more than two years of production telemetry data. Our evaluation shows that Kaleidoscope successfully localized 99.3% and pinpointed the root causes of 95.8% of 843 real-world production issues, with less than 0.01% runtime overhead.
Saurabh Jha, Shengkun Cui, Subho S. Banerjee, Tianyin Xu, Jeremy Enos, Michael T. Showerman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SC4
2020 Understanding and discovering software configuration dependencies in cloud and datacenter systems
abstract
A large percentage of real-world software configuration issues, such as misconfigurations, involve multiple interdependent configuration parameters. However, existing techniques and tools either do not consider dependencies among configuration parameters— termed configuration dependencies—or rely on one or two dependency types and code patterns as input. Without rigorous understanding of configuration dependencies, it is hard to deal with many resulting configuration issues.
Qingrong Chen, Teng Wang 0004, Owolabi Legunsen, Shanshan Li 0001, Tianyin Xu
ESEC/SIGSOFT FSE5
2019 Towards Continuous Access Control Validation and Forensics
abstract
Access control is often reported to be "profoundly broken" in real-world practices due to prevalent policy misconfigurations introduced by system administrators (sysadmins). Given the dynamics of resource and data sharing, access control policies need to be continuously updated. Unfortunately, to err is human-sysadmins often make mistakes such as over-granting privileges when changing access control policies. With today's limited tooling support for continuous validation, such mistakes can stay unnoticed for a long time until eventually being exploited by attackers, causing catastrophic security incidents. We present P-DIFF, a practical tool for monitoring access control behavior to help sysadmins early detect unintended access control policy changes and perform postmortem forensic analysis upon security attacks. P-DIFF continuously monitors access logs and infers access control policies from them. To handle the challenge of policy evolution, we devise a novel time-changing decision tree to effectively represent access control policy changes, coupled with a new learning algorithm to infer the tree from access logs. P-DIFF provides sysadmins with the inferred policies and detected changes to assist the following two tasks: (1) validating whether the access control changes are intended or not; (2) pinpointing the historical changes responsible for a given security attack. We evaluate P-DIFF with a variety of datasets collected from five real-world systems, including two from industrial companies. P-DIFF can detect 86%-100% of access control policy changes with an average precision of 89%. For forensic analysis, P-DIFF can pinpoint the root-cause change that permits the target access in 85%-98% of the evaluated cases.
Chengcheng Xiang, Yudong Wu, Bingyu Shen 0002, Mingyao Shen, Haochen Huang, Tianyin Xu, Yuanyuan Zhou 0001, Cindy Moore, Xinxin Jin, Tianwei Sheng
CCS6
2019 Mobile Gaming on Personal Computers with Direct Android Emulation
abstract
Playing Android games on Windows x86 PCs has gained enormous popularity in recent years, and the de facto solution is to use mobile emulators built with the AOVB (Android-x86 On VirtualBox) architecture. When playing heavy 3D Android games with AOVB, however, users often suffer unsatisfactory smoothness due to the considerable overhead of full virtualization. This paper presents DAOW, a game-oriented Android emulator implementing the idea of direct Android emulation, which eliminates the overhead of full virtualization by directly executing Android app binaries on top of x86-based Windows. Based on pragmatic, efficient instruction rewriting and syscall emulation, DAOW offers foreign Android binaries direct access to the domestic PC hardware through Windows kernel interfaces, achieving nearly native hardware performance. Moreover, it leverages graphics and security techniques to enhance user experiences and prevent cheating in gaming. As of late 2018, DAOW has been adopted by over 50 million PC users to run thousands of heavy 3D Android games. Compared with AOVB, DAOW improves the smoothness by 21% on average, decreases the game startup time by 48%, and reduces the memory usage by 22%.
Zhenhua Li 0001, Yunhao Liu 0001, Hai Long, Yuanchao Huang, Jiaming He, Tianyin Xu, Ennan Zhai
MobiCom7
2019 Demo: Mobile Gaming on Personal Computers with Direct Android Emulation
abstract
Playing Android games with Windows x86 PCs is now popular, and the common solution is to use mobile emulators built with the AOVB (Android-x86 On VirtualBox) architecture. Nevertheless, running heavy 3D Android games on AOVB incurs considerable overhead of full virtualization, thus often leading to unsatisfactory smoothness. To tackle this issue, we present DAOW, a commercial game-oriented Android emulator implementing the idea of direct Android emulation, which eliminates the overhead of full virtualization by providing foreign Android binaries with direct access to the domestic PC hardware through Windows kernel interfaces. In this demo, we will demonstrate that DAOW essentially outperforms traditional AOVB-based emulators in terms of running smoothness, game startup time, and memory usage.
Xinlei Yang, Zhenhua Li 0001, Yunhao Liu 0001, Guoyang Du, Ziwen Wu, Tianyin Xu, Ennan Zhai
MobiCom8
2019 Understanding Fileless Attacks on Linux-based IoT Devices with HoneyCloud
abstract
With the wide adoption, Linux-based IoT devices have emerged as one primary target of today's cyber attacks. Traditional malware-based attacks can quickly spread across these devices, but they are well-understood threats with effective defense techniques such as malware fingerprinting and community-based fingerprint sharing. Recently, fileless attacks---attacks that do not rely on malware files---have been increasing on Linux-based IoT devices, and posing significant threats to the security and privacy of IoT systems. Little has been known in terms of their characteristics and attack vectors, which hinders research and development efforts to defend against them. In this paper, we present our endeavor in understanding fileless attacks on Linux-based IoT devices in the wild. Over a span of twelve months, we deploy 4 hardware IoT honeypots and 108 specially designed software IoT honeypots, and successfully attract a wide variety of real-world IoT attacks. We present our measurement study on these attacks, with a focus on fileless attacks, including the prevalence, exploits, environments, and impacts. Our study further leads to multi-fold insights towards actionable defense strategies that can be adopted by IoT vendors and end users.
Fan Dang 0001, Zhenhua Li 0001, Yunhao Liu 0001, Ennan Zhai, Qi Alfred Chen, Tianyin Xu, Yan Chen 0004
MobiSys6
2019 An In-depth Study of Commercial MVNO: Measurement and Optimization
abstract
Recent years have witnessed the rapid growth of mobile virtual network operators (MVNOs), which operate on top of the existing cellular infrastructures of base carriers while offering cheaper or more flexible data plans compared to those of the base carriers. In this paper, we present a nearly two-year measurement study towards understanding various key aspects of today's MVNO ecosystem, including its architecture, performance, economics, customers, and the complex interplay with the base carrier. Our study focuses on a large commercial MVNO with \reviseabout 1 million customers, operating atop a nation-wide base carrier. Our measurements clarify several key concerns raised by MVNO customers, such as inaccurate billing and potential performance discrimination with the base carrier. We also leverage big data analytics and machine learning to optimize an MVNO's key businesses such as data plan reselling and customer churn mitigation. Our proposed techniques can help achieve %will lead to higher revenues and improved services for commercial MVNOs.
Ao Xiao, Yunhao Liu 0001, Yang Li 0092, Feng Qian 0001, Zhenhua Li 0001, Sen Bai, Yao Liu 0001, Tianyin Xu, Xianlong Xin
MobiSys8
2019 Understanding and Detecting Overlay-based Android Malware at Market Scales
abstract
As a key UI feature of Android, overlay enables one app to draw over other apps by creating an extra View layer on top of the host View. While greatly facilitating user interactions with multiple apps at the same time, it is often exploited by malicious apps (malware) to attack users. To combat this threat, prior countermeasures concentrate on restricting the capabilities of overlays at the OS level, while barely seeing adoption by Android due to the concern of sacrificing overlays' usability. To address this dilemma, a more pragmatic approach is to enable the early detection of overlay-based malware at the app market level during the app review process, so that all the capabilities of overlays can stay unchanged. Unfortunately, little has been known about the feasibility and effectiveness of this approach for lack of understanding of malicious overlays in the wild. To fill this gap, in this paper we perform the first large-scale comparative study of overlay characteristics in benign and malicious apps using static and dynamic analyses. Our results reveal a set of suspicious overlay properties strongly correlated with the malice of apps, including several novel features. Guided by the study insights, we build OverlayChecker, a system that is able to automatically detect overlay-based malware at market scales. OverlayChecker has been adopted by one of the world's largest Android app stores to check around 10K newly submitted apps per day. It can efficiently (within 2 minutes per app) detect nearly all (96%) overlay-based malware using a single commodity server.
Yuxuan Yan, Zhenhua Li 0001, Qi Alfred Chen, Christo Wilson, Tianyin Xu, Ennan Zhai, Yong Li 0008, Yunhao Liu 0001
MobiSys5
2019 Taiji: managing global user traffic for large-scale internet services at the edge
abstract
We present Taiji, a new system for managing user traffic for large-scale Internet services that accomplishes two goals: 1) balancing the utilization of data centers and 2) minimizing network latency of user requests.
Tianyin Xu, Kaushik Veeraraghavan, Andrew Newell, Sonia Margulis, Pol Mauri Ruiz, Justin Meza, Kiryong Ha, Shruti Padmanabha, Kevin Cole, Dmitri Perelman
SOSP2
2018 Towards Web-based Delta Synchronization for Cloud Storage Services
Zhenhua Li 0001, Ennan Zhai, Tianyin Xu, Yang Li 0092, Yunhao Liu 0001, Quanlu Zhang, Yao Liu 0001
FAST4
2018 A Large Scale Study of Data Center Network Reliability
Justin Meza, Tianyin Xu, Kaushik Veeraraghavan, Onur Mutlu
Internet Measurement Conference2
2018 Minimizing the Cask Effect of Multi-Source Content Delivery
abstract
This paper reveals the performance anomaly (i.e., the decline of delivery speed) when the client upgrades a task from single-source content delivery to multi-source content delivery. This anomaly is mainly caused by two aspects: (1) data sources with different types vary greatly in terms of acceleration reward (AR), and data sources with certain types are particularly easy to become inferior; (2) When the data sources remain fixed for a period of time, the large diversity of participant time (DPT) of data sources disturb the acceleration and the data sources with less participant time are inferior. Combing these insights, we figure out that the multi-source content delivery is limited by the so-called cask effect, i.e., the acceleration effect mainly depends on the inferior data sources.
Zhenhua Li 0001, Zhenyu Li 0001, Tianyin Xu, Ennan Zhai, Yao Liu 0001, Minghao Zhao 0001, Yunhao Liu 0001
IWQoS4
2018 Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently
Kaushik Veeraraghavan, Justin Meza, Scott Michelson, Sankaralingam Panneerselvam, Alex Gyori, Sonia Margulis, Daniel Obenshain, Shruti Padmanabha, Ashish Shah, Yee Jiun Song, Tianyin Xu
OSDI12
2017 How Do System Administrators Resolve Access-Denied Issues in the Real World?
abstract
The efficacy of access control largely depends on how system administrators (sysadmins) resolve access-denied issues. A correct resolution should only permit the expected access, while maintaining the protection against illegal access. However, anecdotal evidence suggests that correct resolutions are occasional---sysadmins often grant too much access (known as security misconfigurations) to allow the denied access, posing severe security risks. This paper presents a quantitative study on real-world practices of resolving access-denied issues, with a particular focus on how and why security misconfigurations are introduced during problem solving. We characterize the real-world security misconfigurations introduced in the field, and show that many of these misconfigurations were the results of trial-and-error practices commonly adopted by sysadmins to work around access denials. We argue that the lack of adequate feedback information is one fundamental reason that prevents sysadmins from developing precise understanding and thus induces trial and error. Our study on access-denied messages shows that many of today's software systems miss the opportunities for providing adequate feedback information, imposing unnecessary obstacles to correct resolutions.
Tianyin Xu, Han Min Naing, Le Lu 0003, Yuanyuan Zhou 0001
CHI1
2017 Practical Web-based Delta Synchronization for Cloud Storage Services
Zhenhua Li 0001, Ennan Zhai, Tianyin Xu
HotStorage4
2017 Early Detection of Configuration Errors to Reduce Failure Damage
Tianyin Xu, Xinxin Jin, Peng Huang 0005, Yuanyuan Zhou 0001, Shan Lu 0001, Shankar Pasupathy
USENIX ATC1
2017 Robust Large-Scale Spectrum Auctions against False-Name Bids
abstract
Auction is a promising approach for dynamic spectrum access in cognitive radio networks. Existing auction mechanisms are mainly strategy-proof to stimulate bidders to reveal their valuations of spectrum truthfully. However, they can suffer significantly from a new cheating pattern, named false-name bids, where a bidder can manipulate the auction by submitting bids under multiple fictitious names. We show such false-name bid cheating is easy to make but difficult to detect in dynamic spectrum auctions. To address this issue, we propose ALETHEIA, a novel flexible, false-name-proof auction framework for large-scale dynamic spectrum access. ALETHEIA not only guarantees strategy-proofness but also resists false-name bids. Moreover, ALETHEIA enables spectrum reuse across a large number of bidders, to improve spectrum utilization. Following that, we extend ALETHEIA to its general version that supports more practical and flexible auction, where bidders accept the spectrum allocation under their partial satisfactions. Theoretical analysis and simulation results show that ALETHEIA achieves both high spectrum redistribution efficiency and auction efficiency.
Qinhui Wang, Bin Tang 0002, Tianyin Xu, Song Guo 0001, Sanglu Lu, Weihua Zhuang
IEEE Trans. Mob. Comput.4
2016 NChecker: saving mobile app developers from network disruptions
abstract
Most of today's mobile apps rely on the underlying networks to deliver key functions such as web browsing, file synchronization, and social networking. Compared to desktop-based networks, mobile networks are much more dynamic with frequent connectivity disruptions, network type switches, and quality changes, posing unique programming challenges for mobile app developers.
Xinxin Jin, Peng Huang 0005, Tianyin Xu, Yuanyuan Zhou 0001
EuroSys3
2016 DefDroid: Towards a More Defensive Mobile OS Against Disruptive App Behavior
abstract
The mobile app market is enjoying an explosive growth. Many people, including teenagers, set about developing mobile apps. Unfortunately, due to developers' inexperience and the unique mobile programming paradigms, a growing number of immature apps are released to users. Despite having useful functionalities, these apps exhibit disruptive behaviors that are inconsiderate to the mobile system as a whole, e.g.,, retrying network connections too aggressively, waking up the device too frequently, or holding resources for unnecessarily long. These behaviors adversely affect other apps running on the same device and frustrate users with battery drain, excessive cellular data consumption, storage overuse, etc.
Peng Huang 0005, Tianyin Xu, Xinxin Jin, Yuanyuan Zhou 0001
MobiSys2
2016 Exploring Cross-Application Cellular Traffic Optimization with Baidu TrafficGuard
Zhenhua Li 0001, Weiwei Wang 0002, Tianyin Xu, Xiang-Yang Li 0001, Yunhao Liu 0001, Christo Wilson, Ben Y. Zhao
NSDI3
2016 Early Detection of Configuration Errors to Reduce Failure Damage
Tianyin Xu, Xinxin Jin, Peng Huang 0005, Yuanyuan Zhou 0001, Shan Lu 0001, Shankar Pasupathy
OSDI1
2016 Optimizing Cost for Online Social Networks on Geo-Distributed Clouds
abstract
Geo-distributed clouds provide an intriguing platform to deploy online social network (OSN) services. To leverage the potential of clouds, a major concern of OSN providers is optimizing the monetary cost spent in using cloud resources while considering other important requirements, including providing satisfactory quality of service (QoS) and data availability to OSN users. In this paper, we study the problem of cost optimization for the dynamic OSN on multiple geo-distributed clouds over consecutive time periods while meeting predefined QoS and data availability requirements. We model the cost, the QoS, as well as the data availability of the OSN, formulate the problem, and design an algorithm named${\tt cosplay}$. We carry out extensive experiments with a large-scale real-world Twitter trace over 10 geo-distributed clouds all across the US. Our results show that, while always ensuring the QoS and the data availability as required,${\tt cosplay}$can reduce much more one-time cost than the state-of-the-art methods, and it can also significantly reduce the accumulative cost when continuously evaluated over 48 months, with OSN dynamics comparable to real-world cases.
Lei Jiao 0002, Jun Li 0001, Tianyin Xu, Xiaoming Fu 0001
IEEE/ACM Trans. Netw.3
2015 Offline Downloading in China: A Comparative Study
abstract
Although Internet access has become more ubiquitous in recent years, most users in China still suffer from low-quality connections, especially when downloading large files. To address this issue, hundreds of millions of China's users have resorted to technologies that allow for ``offline downloading'', where a proxy is employed to pre-download the user's requested file and then deliver the file at her convenience.
Zhenhua Li 0001, Christo Wilson, Tianyin Xu, Yao Liu 0001, Yinlong Wang
Internet Measurement Conference3
2015 ALETHEIA: Robust Large-Scale Spectrum Auctions against False-name Bids
abstract
Auction is a promising approach for dynamic spectrum access in Cognitive Radio Networks. Existing auction mechanisms are mainly proposed to be strategy-proof to stimulate bidders to reveal their valuations of spectrum truthfully. However, they would suffer significantly from a new cheating pattern named false-name bids, where a bidder can manipulate the auction by submitting bids under multiple fictitious names. We show such false-name bid cheating is easy to make but hard to be detected in dynamic spectrum auctions. To resolve this issue, we propose ALETHEIA, a novel flexible, false-name-proof auction framework for large-scale dynamic spectrum access. ALETHEIA has the following important features: (1) it not only guarantees strategy-proofness but also resists false-name bids, (2) it enables spectrum reuse across a large number of bidders, (3) it provides the bidders the flexibility of diverse demand formats, and (4) it incurs low computational overhead. Simulation results show that ALETHEIA achieves both high spectrum redistribution efficiency and auction efficiency.
Qinhui Wang, Bin Tang 0002, Tianyin Xu, Song Guo 0001, Sanglu Lu, Weihua Zhuang
MobiHoc4
2015 Hey, you have given me too many knobs!: understanding and dealing with over-designed configuration in system software
abstract
Configuration problems are not only prevalent, but also severely impair the reliability of today's system software. One fundamental reason is the ever-increasing complexity of configuration, reflected by the large number of configuration parameters ("knobs"). With hundreds of knobs, configuring system software to ensure high reliability and performance becomes a daunting, error-prone task. This paper makes a first step in understanding a fundamental question of configuration design: "do users really need so many knobs?" To provide the quantitatively answer, we study the configuration settings of real-world users, including thousands of customers of a commercial storage system (Storage-A), and hundreds of users of two widely-used open-source system software projects. Our study reveals a series of interesting findings to motivate software architects and developers to be more cautious and disciplined in configuration design. Motivated by these findings, we provide a few concrete, practical guidelines which can significantly reduce the configuration space. Take Storage-A as an example, the guidelines can remove 51.9% of its parameters and simplify 19.7% of the remaining ones with little impact on existing users. Also, we study the existing configuration navigation methods in the context of "too many knobs" to understand their effectiveness in dealing with the over-designed configuration, and to provide practices for building navigation support in system software.
Tianyin Xu, Xuepeng Fan, Yuanyuan Zhou 0001, Shankar Pasupathy, Rukma Talwadker
ESEC/SIGSOFT FSE1
2014 EnCore: exploiting system environment and correlation information for misconfiguration detection
abstract
As software systems become more complex and configurable, failures due to misconfigurations are becoming a critical problem. Such failures often have serious functionality, security and financial consequences. Further, diagnosis and remediation for such failures require reasoning across the software stack and its operating environment, making it difficult and costly. We present a framework and tool called EnCore to automatically detect software misconfigurations. EnCore takes into account two important factors that are unexploited before: the interaction between the configuration settings and the executing environment, as well as the rich correlations between configuration entries. We embrace the emerging trend of viewing systems as data, and exploit this to extract information about the execution environment in which a configuration setting is used. EnCore learns configuration rules from a given set of sample configurations. With training data enriched with the execution context of configurations, EnCore is able to learn a broad set of configuration anomalies that spans the entire system. EnCore is effective in detecting both injected errors and known real-world problems - it finds 37 new misconfigurations in Amazon EC2 public images and 24 new configuration problems in a commercial private cloud. By systematically exploiting environment information and by learning correlation rules across multiple configuration settings, EnCore detects 1.6x to 3.5x more misconfiguration anomalies than previous approaches.
Lakshminarayanan Renganarayana, Xiaolan Zhang 0001, Niyu Ge, Vasanth Bala, Tianyin Xu, Yuanyuan Zhou 0001
ASPLOS6
2014 Towards Network-level Efficiency for Cloud Storage Services
abstract
Cloud storage services such as Dropbox, Google Drive, and Microsoft OneDrive provide users with a convenient and reliable way to store and share data from anywhere, on any device, and at any time. The cornerstone of these services is the data synchronization (sync) operation which automatically maps the changes in users' local filesystems to the cloud via a series of network communications in a timely manner. If not designed properly, however, the tremendous amount of data sync traffic can potentially cause (financial) pains to both service providers and users.
Zhenhua Li 0001, Cheng Jin 0008, Tianyin Xu, Christo Wilson, Yao Liu 0001, Linsong Cheng, Yunhao Liu 0001, Yafei Dai, Zhi-Li Zhang
Internet Measurement Conference3
2013 Do not blame users for misconfigurations
abstract
Similar to software bugs, configuration errors are also one of the major causes of today's system failures. Many configuration issues manifest themselves in ways similar to software bugs such as crashes, hangs, silent failures. It leaves users clueless and forced to report to developers for technical support, wasting not only users' but also developers' precious time and effort. Unfortunately, unlike software bugs, many software developers take a much less active, responsible role in handling configuration errors because "they are users' faults."
Tianyin Xu, Peng Huang 0005, Tianwei Sheng, Ding Yuan 0004, Yuanyuan Zhou 0001, Shankar Pasupathy
SOSP1
2012 Cost optimization for Online Social Networks on geo-distributed clouds
abstract
Geo-distributed IaaS (Infrastructure-as-a-Service) clouds provide an intriguing platform to deploy Online Social Network (OSN) services. To leverage the potential of clouds, a major task of OSN providers is optimizing the monetary cost spent on cloud resource utilization while providing satisfactory Quality of Service (QoS) to OSN users. We thus study the problem of cost optimization for the dynamic OSN on multiple geo-distributed clouds over consecutive time periods, with its QoS meeting the pre-defined requirement. We model the QoS as well as the cost of an OSN, formulate the problem, and design a solution named cosplay. Our experiments with a large-scale Twitter trace show that, while always ensuring the QoS as required, cosplay can achieve superior one-time cost reduction compared with the state of the art, and can also reduce the accumulative cost significantly when continuously evaluated over 48 months with dynamics comparable to real-world OSNs.
Lei Jiao 0002, Jun Li 0001, Tianyin Xu, Xiaoming Fu 0001
ICNP3
2012 CloudGPS: A scalable and ISP-friendly server selection scheme in cloud computing environments
abstract
In order to minimize user perceived latency while ensuring high data availability, cloud applications desire to select servers from one of the multiple data centers (i.e., server clusters) in different geographical locations, which are able to provide desired services with low latency and low cost. This paper presents CloudGPS, a new server selection scheme of the cloud computing environment that achieves high scalability and ISP-friendliness. CloudGPS proposes a configurable global performance function that allows Internet service providers (ISPs) and cloud service providers (CSPs) to leverage the cost in terms of inter-domain transit traffic and the quality of service in terms of network latency. CloudGPS bounds the overall burden to be linear with the number of end users. Moreover, compared with traditional approaches, CloudGPS significantly reduces network distance measurement cost (i.e., from O(N) to O(1) for each end user in an application using N data centers). Furthermore, CloudGPS achieves ISP-friendliness by significantly decreasing inter-domain transit traffic.
Cong Ding 0001, Yang Chen 0001, Tianyin Xu, Xiaoming Fu 0001
IWQoS3
2012 DOTA: A Double Truthful Auction for spectrum allocation in dynamic spectrum access
abstract
Spectrum auctions have been proposed as an effective approach to fairly and efficiently trade the scarce spectrum resource among wireless users. The most significant challenge of the auction design to provide economic robustness, particularly truthfulness, under the local-dependent interference constraints. However, existing designs either do not consider spectrum reuse or are based on the impractical assumption that each user requests at most one channel. In this paper, we address this problem by proposing DOTA, a DOuble Truthful Auction for dynamic spectrum access. DOTA is economic-robust in terms of truthfulness, individual rationality, and no-deficit. It achieves improved utilization by exploiting spectrum reuse as well as dealing with the interference constraints. Moreover, DOTA minimizes the network transaction overhead and provides flexible channel bidding including range bidding and strict bidding.
Qinhui Wang, Tianyin Xu, Sanglu Lu, Song Guo 0001
WCNC3
2011 Scaling Microblogging Services with Divergent Traffic Demands
Tianyin Xu, Yang Chen 0001, Lei Jiao 0002, Ben Y. Zhao, Pan Hui 0001, Xiaoming Fu 0001
Middleware1
2011 An approximate truthfulness motivated spectrum auction for dynamic spectrum access
abstract
Secondary Spectrum Auction (SSA) has been proposed as an effective approach to design spectrum sharing mechanism for dynamic spectrum access. However, due to the location-constrained spectrum interference among users, it is a great challenge to provide truthful auction with maximized spectrum utilization. Most previous SSA designs either fail in addressing truthfulness or cause loss on spectrum utilization. In this paper, we focus on providing truthful SSA with maximized spectrum utilization. In order to minimize the computational overhead involved in addressing location-constrained interference, we leverage the truthfulness by introducing approximate truthfulness. Moreover, we define a general spectrum auction model using linear programming. Based on this model, we further propose ETEX, a sealed-bid auction mechanism with approximate truthfulness. Theoretical analysis confirms that ETEX is able to achieve truthfulness in expectation with polynomial complexity. Extensive experimental results show that ETEX outperforms most popular truthful spectrum auctions in terms of social welfare, spectrum utilization and user satisfaction.
Qinhui Wang, Tianyin Xu, Sanglu Lu
WCNC3
2010 Improving Prediction Accuracy of Matrix Factorization Based Network Coordinate Systems
abstract
Network Coordinate (NC) systems provide a lightweight and useful way for scalable Internet distance prediction while serving as an important component in many Peer-to-Peer applications. Most of the existing NC systems utilize the Euclidean distance model, which is largely impaired by the persistent occurrence of Triangle Inequality Violation (TIV) in the Internet. Recently, matrix factorization (MF) based NC systems, which can completely remove the TIV constraint, provide an alternative approach towards better prediction accuracy. Phoenix, an NC system based on the MF model, well explores the advantage of the MF model and becomes the most accurate NC system so far. However, through experimental study, we find that the prediction accuracy of Phoenix for short links is significantly worse compared with the overall prediction accuracy. Based on the observation, we propose a new NC system, named Pancake, aiming at reducing the high prediction error for short links. By introducing a two-level architecture, Pancake achieves much higher prediction accuracy than other selected existing NC systems. Through extensive experiments, we demonstrate that Pancake reduces the 90th percentile relative error by up to 25.37% from Phoenix. Moreover, Pancake converges very fast and is robust to different dimension values. For further insights, we study the performance of Pancake using a dynamic data set in addition to the widely used aggregate data sets. With varying RTTs over time, Pancake outperforms other NC systems consistently.
Yang Chen 0001, Xiaoming Fu 0001, Tianyin Xu
ICCCN4
2010 APEX: A personalization framework to improve quality of experience for DVD-like functions in P2P VoD applications
abstract
The requirement for supporting DVD-like functions raises new challenges to the design of P2P VoD systems. The uncertainty of frequent user DVD-like interactivity makes it difficult to ensure user perceived Quality of Experience (QoE) for real-time streaming services over distributed self-organized P2P overlay networks. Most existing solutions are based on the unreasonable assumption that all the users in P2P VoD systems have the same preference. Few attention has been paid to personalization, which accommodates the differences between users. In this paper, we present a video model which characterizes the personalization information for users' contents and preferences. Based on this model, we develop APEX, a practical personalization framework for P2P VoD applications. APEX makes the personalization practical by using a hybrid architecture which leverages the offline pattern mining on the server side and online collaborative filtering on the peer side. Furthermore, APEX helps peers to personalize navigation, prefetching and membership management, aiming at improving QoE for DVD-like functions by reducing response latency and optimizing content sharing. Both theoretical analysis and comprehensive simulations show that APEX outperforms most existing schemes in terms of accumulated hit ratio, response latency, and searching efficiency.
Tianyin Xu, Qinhui Wang, Sanglu Lu, Xiaoming Fu 0001
IWQoS1
2010 Twittering by cuckoo: decentralized and socio-aware online microblogging services
abstract
Online microblogging services, as exemplified by Twitter, have become immensely popular during the latest years. However, current microblogging systems severely suffer from performance bottlenecks and malicious attacks due to the centralized architecture. As a result, centralized microblogging systems may threaten the scalability, reliability as well as availability of the offered services, not to mention the high operational and maintenance cost.
Tianyin Xu, Yang Chen 0001, Xiaoming Fu 0001, Pan Hui 0001
SIGCOMM1
2009 A Parameter-Based Scheme for Service Composition in Pervasive Computing Environment
abstract
Pervasive computing, the new computing paradigm aiming at providing services anywhere at anytime, poses great challenges on dynamic service composition. Existing service composition methods can hardly meet the requirements of dynamic characteristic and heterogeneity in pervasive computing environment. In this paper, we propose a parameter-based service model to accurately describe pervasive services. Based on the model, pervasive services are aggregated in a two-layer graph according to both semantic and syntactic information of the input and output parameters. Moreover, we design a novel service composition scheme to accomplish the user task while satisfy the QoS requirements. Both theoretical analysis and simulation experiments show that this service composition mechanism is effective in pervasive environment.
Zhenghui Wang, Tianyin Xu, Zhuzhong Qian, Sanglu Lu
CISIS2
2009 Supporting VCR-Like Operations in Derivative Tree-Based P2P Streaming Systems
abstract
Supporting user interactivity in peer-to-peer streaming systems is challenging. VCR-like operations, such as random seek, pause, fast forward and rewind, require timely P2P overlay topology adjustment and appropriate bandwidth resource re-allocation. If not handled properly, the dynamics caused by user interactivity may severely deteriorate users' perceived video quality, e.g., longer start-up delay, frequent playback freezing, or blackout altogether. In this paper, we propose a derivative tree-based overlay management scheme to support user interactivity in P2P streaming system. Derivative tree takes advantage of well organized buffer overlapping to support asynchronous user requests while brings high resilience to the impact of VCR-like operations. A session discovery service is introduced to quickly locate parent peer. We show that the overhead of VCR-like operations in derivative-tree based scheme is O(log(N)), where N is the number of sessions. Simulation experiments further demonstrate the efficiency of the proposed scheme.
Tianyin Xu, Sanglu Lu, Mounir Hamdi
ICC1
2009 Prediction-Based Prefetching to Support VCR-like Operations in Gossip-Based P2P VoD Systems
abstract
Supporting free VCR-like operations in P2P VoD streaming systems is challenging. The uncertainty of frequent VCR operations makes it difficult to provide high quality realtime streaming services over distributed self-organized P2P overlay networks. Recently, prefetching has emerged as a promising approach to smooth the streaming quality. However, how to efficiently and effectively prefetch suitable segments is still an open issue. In this paper, we propose PREP, a PREdiction-based Prefetching scheme to support VCR-like operations over gossip-based P2P on-demand streaming systems. By employing the reinforcement learning technique, PREP transforms users' streaming service procedure into a set of abstract states and presents an online prediction model to predict a user's VCR behavior via analyzing the large volumes of user viewing logs collected on the tracker. We further present a distributed data scheduling algorithm to proactively prefetch segments according to the predicted VCR behavior. Moreover, PREP takes advantage of the inherent peer collaboration of gossip protocol to optimize the response latency. Through comprehensive simulations, we demonstrate the efficiency of PREP by gaining the accumulated hit ratio close to 75% while reducing the response latency close to 70% with only less than 15% extra stress on the server side.
Tianyin Xu, Weiwei Wang 0002, Sanglu Lu, Yang Gao 0001
ICPADS1
2009 SkipStream: A Clustered Skip Graph Based On-demand Streaming Scheme over Ubiquitous Environments
abstract
Providing continuous on-demand streaming services with VCR functionality over ubiquitous environments is challenging due to the stringent QoS requirements of streaming service as well as the dynamic characteristics of both underlying network and user behavior. In this paper, we propose SkipStream, a skip graph based peer-to-peer (P2P) on-demand streaming scheme with VCR support to address the above challenges. In the design of SkipStream, we first group users into a set of disjoint clusters in accordance with their playback offset and further organize the resulted clusters into a skip graph based overlay network. In addition, we present a distributed on-demand streaming scheduling mechanism to minimize the impact of VCR operations and balance system load among nodes adaptively. The average search latency of SkipStream is O(log(N)) where N is the number of disjoint clusters. We also evaluate the performance of SkipStream via extensive simulations. Experimental results show that SkipStream outperforms early skip list based scheme DSL by reducing the search latency 20%-60% in average case and over 50% in worst case.
Tianyin Xu, Sanglu Lu, Daoxu Chen
ICPP2