Lexiang Huang

dblp:210/2490 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0009-5044-9117ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Workload Intelligence: Workload-Aware IaaS abstraction for Cloud Efficiency
abstract
Today, cloud workloads are largely opaque to the cloud platform. Typically, the only information the platform receives is the virtual machine (VM) type and possibly a decoration to the type (e.g., the VM is evictable). Similarly, workloads receive minimal information from the platform; generally, only telemetry from their VMs or occasional signals (e.g., just before a VM is evicted). The narrow interface between workloads and platforms has several drawbacks: (1) a surge in VM types and decorations in public cloud platforms complicates customer selection; (2) key workload characteristics (e.g., low availability requirements) are often unspecified, hindering platform customization for optimized resource usage and cost savings; and (3) workloads may be unaware of potential optimizations or lack sufficient time to react to platform events. To resolve these issues and improve cloud efficiency, we propose Workload Intelligence (WI), a framework for enabling dynamic bi-directional communication between cloud workloads and cloud platform.
Lexiang Huang, Anjaly Parayil, Xiaoting Qin, Chetan Bansal, Jovan Stojkovic, Pantea Zardoshti, Pulkit A. Misra, Eli Cortez, Raphael Ghelman, Íñigo Goiri, Saravan Rajmohan, Jim Kleewein, Rodrigo Fonseca, Timothy Zhu, Ricardo Bianchini
SC1
2025 Towards Workload-aware Cloud Efficiency: A Large-scale Empirical Study of Cloud Workload Characteristics
abstract
Cloud providers introduce features and optimizations to improve efficiency and reliability, such as Spot VMs, Harvest VMs, oversubscription, and auto-scaling. To use these effectively, it's important to understand workload characteristics. However, workload characterization can be complex and difficult to scale manually due to multiple signals involved. In this study, we conduct the first large-scale empirical study of first-party workloads at Microsoft to understand their characteristics. Through this empirical study, we aim to answer the following questions: (1) What are the critical workload characteristics that impact efficiency and reliability on cloud platforms? (2) How do these characteristics vary across different workloads? (3) How can cloud platforms leverage these insights to efficiently characterize all workloads at scale? This study provides a deeper understanding of workload characteristics and their impact on cloud performance, which can aid in optimizing cloud services and identifies potential areas for future research.
Anjaly Parayil, Xiaoting Qin, Íñigo Goiri, Lexiang Huang, Timothy Zhu, Chetan Bansal
ICPE5
2022 Metastable Failures in the Wild
Lexiang Huang, Matt Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak, Rebecca Isaacs, Abutalib Aghayev, Timothy Zhu, Aleksey Charapko
OSDI1
2021 tprof: Performance profiling via structural aggregation and automated analysis of distributed systems traces
abstract
The traditional approach for performance debugging relies upon performance profilers (e.g., gprof, VTune) that provide average function runtime information. These aggregate statistics help identify slow regions affecting the entire workload, but they are ill-suited for identifying slow regions that only impact a fraction of the workload, such as tail latency effects. This paper takes a new approach to performance profiling by utilizing distributed tracing systems (e.g., Dapper, Zipkin, Jaeger). Since traces provide detailed timing information on a per-request basis, it is possible to group and aggregate tracing data in many different ways to identify the slow parts of the system. Our new approach to trace aggregation uses the structure embedded within traces to hierarchically group similar traces and calculate increasingly detailed aggregate statistics based on how the traces are grouped. We also develop an automated tool for analyzing the hierarchy of statistics to identify the most likely performance issues. Our case study across two complex distributed systems illustrates how our tool is able to find multiple performance issues that lead to 10x and 28x performance improvements in terms of average and tail latency, respectively. Our comparison with a state-of-the-art industry tool shows that our tool can pinpoint performance slowdowns more accurately than current approaches.
Lexiang Huang, Timothy Zhu
SoCC1
2020 Sum-GDoF of 2-User Interference Channel With Limited Cooperation Under Finite Precision CSIT
abstract
The Generalized Degrees of Freedom (GDoF) of the two user interference channel are characterized for all parameter regimes under the assumption of finite precision channel state information at the transmitters (CSIT), when a limited amount of (half-duplex or full-duplex) cooperation is allowed between the transmitters in the form of π DoF of shared messages. In all cases, the number of over-the-air bits that each cooperation bit buys is shown to be equal to either 0, 1, 1/2 or 1/3. The most interesting aspect of the result is the 1/3 slope, which appears only under finite precision CSIT and strong interference, and as such has not been encountered in previous studies that invariably assumed perfect CSIT. Indeed, the achievability and converse for the parameter regimes with 1/3 slope are the most challenging aspects of this work. In particular, the converse relies on non-trivial applications of Aligned Images bounds.
Junge Wang, Bofeng Yuan, Lexiang Huang, Syed Ali Jafar
IEEE Trans. Inf. Theory3
2019 GDoF of Interference Channel with Limited Cooperation under Finite Precision CSIT
abstract
The Generalized Degrees of Freedom (GDoF) of the two user interference channel are characterized for all parameter regimes under the assumption of finite precision channel state information at the transmitters (CSIT), when a limited amount of cooperation is allowed between the transmitters in the form of π DoF of shared messages. In all cases, the number of over-the-air bits that each cooperation bit buys is shown to be equal to either 0, 1, 1/2 or 1/3.
Junge Wang, Syed Ali Jafar, Bofeng Yuan, Lexiang Huang
GLOBECOM4
2019 Joint Downlink Scheduling for File Placement and Delivery in Cache-Assisted Wireless Networks With Finite File Lifetime
abstract
In this paper, downlink transmission scheduling of popular files is optimized with the assistance of wireless cache nodes. Specifically, the requests of each file, which is further divided into a number of segments, are modeled as a Poisson point process within its finite lifetime. Two downlink transmission modes are considered: 1) the base station reactively multicasts the file segments to the requesting users and selected cache nodes and (2) the base station proactively multicasts some file segments to the selected cache nodes without requests. The cache nodes with decoded file segments can help to offload the traffic via other spectrum. Without the proactive multicast, we formulate the downlink transmission resource minimization as a dynamic programming problem with random stage number, which can be approximated via a finite-horizon Markov decision process (MDP) with fixed stage number. To address the prohibitively huge state space, we propose a low-complexity scheduling policy by linearly approximating the value functions of the MDP, where the bound on the approximation error is derived. Moreover, we propose a learning-based algorithm to evaluate the approximated value functions for unknown geographical distribution of requesting users. Finally, given the above reactive multicast policy, a proactive multicast policy is introduced to exploit the temporal diversity of shadowing effect. It is shown by simulation that the proposed low-complexity reactive multicast policy can significantly reduce the resource consumption at the base station, and the proactive multicast policy can further improve the performance.
Bojie Li, Lexiang Huang, Rui Wang 0007
IEEE Trans. Commun.2
2018 Cellular Offloading via Downlink Cache Placement
abstract
In this paper, the downlink file transmission within a finite lifetime is optimized with the assistance of wireless cache nodes. Specifically, the number of requests within the lifetime of one file is modeled as a Poisson point process. The base station multicasts files to downlink users and the selected the cache nodes, so that the cache nodes can help to forward the files in the next file request. Thus we formulate the downlink transmission as a Markov decision process (MDP) with random number of stages, where transmission power and time on each transmission are the control policy. Due to random number of file transmissions, we first proposed a revised Bellman's equation, where the optimal control policy can be derived. In order to address the prohibitively huge state space, we also introduce a low-complexity sub-optimal solution based on a linear approximation of the value function. The approximated value function can be calculated analytically, so that conventional numerical value iteration can be eliminated. Moreover, the gap between the approximated value function and the real value function is bounded analytically. It is shown by simulation that, with the approximated MDP approach, the proposed algorithm can significantly reduce the resource consumption at the base station.
Bojie Li, Lexiang Huang, Rui Wang 0007
ICC2