Xinkai Wang 0003

dblp:45/2185-3 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0003-3764-8065ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 7 first-author · 13 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM Serving
abstract
Generative AI, especially LLM, is driving a fundamental shift in software paradigms, prompting cloud providers to build more efficient serving infrastructures. To meet the computational demands of emerging software, modern CPU processors are integrating Accelerator Units (AU) in the pipeline to accelerate key operations, such as Intel AMX for matrix multiplication. Current practices that dedicate AU-enabled CPU exclusively to LLM serving lead to significant resource waste and inferior efficiency. To this end, sharing AU-enabled CPU with general workloads is necessary to harvest redundant resources and improve platform performance-per-watt. However, perfectly sharing AU can be challenging since they introduce three-dimensional variations: variable usage patterns, compulsory frequency interferences, and dissimilar resource bounds. Existing resource managers are oblivious to complex Accelerator Unit Variations (AUV), resulting in performance and efficiency degradations of up to 50 % in shared environments. Therefore, this paper introduces AUM, a novel AU-aware resource manager designed to handle AUV and maximize the efficiency of shared processors. AUM has two cooperative components with three stages for three-dimensional AUV. The background profiler characterizes the usage, frequency, and resource information into a discrete model, guiding the runtime controller to analyze usage-aware requirements, select frequency-aware divisions, and make bound-aware resource decisions. Through extensive evaluations on production AU-enabled CPUs, we show that AUM improves CPU efficiency by$4.7-8.8 \%$while maintaining high-performance AU applications by reducing SLO violations by$\mathbf{7 - 1 1 \%}$compared with state-of-the-art resource managers.
Xinkai Wang 0003, Chao Li 0009, Yiming Zhuansun, Jinyang Guo 0001, Xiaofeng Hou, Jing Wang 0055, Weigao Chen, Liping Zhang 0013, Minyi Guo
HPCA1
2026 LocMore: Locating More Bursty Latency-critical Jobs on Resource-constrained Nodes
Xiaofeng Hou, Xinkai Wang 0003, Jiacheng Liu 0001, Chao Li 0009, Minyi Guo
IWQoS4
2026 StarkServe: A Framework for Elastic Serverless LLM Inference at the Extreme Edge
Xiaofeng Hou, Jiacheng Liu 0001, Xinkai Wang 0003, Chao Li 0009, Minyi Guo
IWQoS5
2026 DGS: A GPU-based Adaptive Graph Sampling Framework
abstract
Graph sampling plays a critical role in graph learning applications, notably within Graph Neural Networks (GNNs). Typically, the performance of GPU-based graph sampling is determined by the efficiency of sampling kernels. Different sampling methods excel under different conditions, and no single method consistently outperforms others in all scenarios. As sampling applications become increasingly complex, graph-related sparse operations can dominate the computational workload, with performance heavily influenced by storage formats. In this article, we propose DGS, a GPU-based graph sampling framework that can detach the kernel implementation from computation logic. In addition to sampling kernels, DGS jointly optimizes sparse graph kernels. It can adaptively switch between different execution strategies based on various inputs. Experiments show that DGS outperforms current state-of-the-art GPU sampling frameworks, achieving speedups ranging from 1.1× to 92.0×. This adaptability and performance improvement establish DGS as a highly effective and efficient solution for diverse graph sampling scenarios.
Junyi Mei, Shixuan Sun, Chao Li 0009, Xinkai Wang 0003, Xiaofeng Hou, Minyi Guo, Yongchao Liu 0004, Chuntao Hong
ACM Trans. Archit. Code Optim.4
2026 Enabling Learning-Based Efficiency Optimizer With Shadow Cycles in Resource-Constrained Autonomous Embedded Systems
abstract
The emerging trend of autonomous embedded systems (AES) is promising to minimize human intervention in critical tasks. In the pursuit of maximal per-watt performance, the complex hardware and software of AES require intelligent energy efficiency optimizers (EO), and the stochastic runtime variances require continuous EO. However, deploying the desirable ondevice EO causes severe performance slowdown due to contention on limited computing power with the AES pipeline. We find that there are ignored and underutilized heterogeneous resources within AES for costly EO, which results from unbalanced accelerator behaviors and misaligned parallel inference executions. We experimentally and theoretically analyze theShadow Cycleswithin the realistic autonomous Bird’s Eye View pipeline on commercial embedded platforms, categorizing them into vertical and horizontal types with distinct properties.In this paper, we introduceSHEEO+, a continuous and intelligent energy efficiency optimizer that utilizes ignored heterogeneous shadow cycles. It achieves continuous and lightweight AES monitoring with the observation module, as well as intelligent and efficient AES power management with the optimization module. On the one hand,SHEEO+ observes both the internal runtime status and external environment variance with portable interfaces to capture shadow cycles and real-time states. On the other hand,SHEEO+ optimizes power configurations per iteration based on deep reinforcement learning (DRL) methods. It tailors DRL for two types of shadow cycles and invocates optimization processes based on resource availability. To extensively evaluateSHEEO+, we implement a prototype and deploy it on realistic edge platforms. The evaluation results show thatSHEEO+ utilizes up to 74.2% shadow cycles and achieves up to 18.6% energy efficiency improvements compared to state-of-the-art energy efficiency optimizers with negligible deployment overheads.
Xinkai Wang 0003, Chao Li 0009, Xiaofeng Hou, Jing Wang 0055, Minyi Guo, Yaqian Zhao
IEEE Trans. Computers1
2025 Accelerating Large-Scale Out-of-GPU-Core GNN Training with Two-Level Historical Caching
Jing Wang 0055, Taolei Wang, Juntao Huang, Xinkai Wang 0003, Marius Kreutzer, Chao Li 0009, Minyi Guo
APPT5
2025 AsymServe: Demystifying and Optimizing LLM Serving Efficiency on CPU Acceleration Units
Xinkai Wang 0003, Yiming Zhuansun, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo
APPT1
2025 EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in Datacenters
abstract
The complexity of online applications is rapidly increasing, bringing more sophisticated performance anomalies in today's cloud datacenter. To fully understand application behaviors, we should obtain both inter-service communication data via RPC-level tracing and intra-service execution traces via application-level tracing to precisely reason about event causality. However, the average time overhead of existing intra-service tracing schemes on the traced applications is generally about 5-10%, possibly reaching 18% in the worst case. To realize practical intra-service tracing in shared and stressed datacenters, one must achieve extreme tracing efficiency with an overhead at the per-mille level.
Xinkai Wang 0003, Xiaofeng Hou, Chao Li 0009, Yuancheng Li 0001, Du Liu, Guoyao Xu, Liping Zhang 0013, Yuemin Wu, Xiaopeng Yuan, Quan Chen 0002, Minyi Guo
ASPLOS (2)1
2025 TriCooling-Sim: Efficient Thermal Simulation for High-Density Micro AI Data Centers
Jinyang Guo 0001, Xinkai Wang 0003, Jing Wang 0055, Xiaofeng Hou, Chao Li 0009, Minyi Guo
NPC (2)2
2025 Power synchronization: taming massive diversified serverless functions under power constraints
Du Liu, Lu Zhang 0049, Yechen Xu, Xinkai Wang 0003, Yi-Fei Pu, Xiaofeng Hou, Chao Li 0009, Minyi Guo
Sci. China Inf. Sci.4
2024 Improving the Efficiency of Serverless Computing via Core-Level Power Management
abstract
Serverless computing has recently become a significant application paradigm in data centers. However, existing power management methods focus on optimizations at the coarse-grained server level, making them unable to handle the characteristics of these short-lived, dynamic serverless functions. In this context, the unawareness of function-level characteristics by the existing power management systems can severely degrade the energy efficiency of the data centers. To address this challenge, we design a function-level power management system. Instead of relying on server-level schedulers, we propose a novel core-level scheduling policy for serverless functions that can efficiently allocate functions to the most suitable CPU core. Additionally, we propose a power management mechanism for serverless computing that can reduce system power consumption with functions’ QoS guaranteed. Our evaluation shows that our system achieves a maximum power saving of 8.5% and an average power saving of 8% across the majority of loads without incurring any loss in tail latency, as compared to the conventional server-level scheduling system.
Du Liu, Jing Wang 0055, Xinkai Wang 0003, Chao Li 0009, Lu Zhang 0049, Xiaofeng Hou, Xiaoxiang Shi, Minyi Guo
CCGrid3
2024 SHEEO: Continuous Energy Efficiency Optimization in Autonomous Embedded Systems
abstract
The emerging trend of autonomous embedded systems minimizing human intervention has raised new questions about continuously maximizing system energy efficiency faced with stochastic runtime variance, which is costly for resource-constrained autonomous embedded systems. Considering heterogeneous hardware and variable software, we envision opportunities for vertical and horizontal shadow cycles within the AES pipeline for management facilities. This paper introduces SHEEO, a continuous energy efficiency optimizer that exploits underutilized heterogeneous computing resources to pursue variability-aware power management. To achieve this, SHEEO constantly monitors inner and outer variances and customizes reinforcement learning into two phases for stochastic runtime variance. We implement and deploy SHEEO on a commercial edge platform. The evaluation results show that SHEEO harvests up to 88% shadow cycles and improves up to 39% energy efficiency compared to state-of-the-art power management techniques with negligible overheads.
Xinkai Wang 0003, Chao Li 0009, Qizheng Lyu, Xiaofeng Hou, Jingwen Leng, Minyi Guo
ICCD1
2024 Jigsaw: Taming BEV-centric Perception on Dual-SoC for Autonomous Driving
abstract
Real-time perception is important for autonomous driving. We observe an emerging trend using one large and critical fusion-based Bird’s-Eye-View (BEV) Deep Neural Network (DNN) model to perform core perception tasks. It collaborates with a few auxiliary Perspective-View (PV) models, forming a BEV-centric paradigm. Organizing the BEV and PV models respecting their distinct real-time requirements becomes challenging, especially on the state-of-the-practice GPU-integrated dual System-on-Chip (SoC) platform. It remains unclear how to appropriately allocate the separated GPU resource to BEV and PV models, satisfying their distinct real-time requirements with latency predictability. No public solution has been proposed for this emerging software-hardware combination.This paper explores parallelism and a timeslot-filling mechanism to organize tasks. We propose Jigsaw, a specialized execution timeline management framework for BEV-centric perception on dual-SoC. First, it exploits component parallelism to carefully place BEV model components and reduce BEV model latency. Second, we recognize two types of idle GPU timeslots left by a parallelized BEV model. The stable timeslot can offer hard real-time guarantee for PV models, while the unstable timeslot could only provide soft real-time capability. Therefore, Jigsaw schedules PV models by timeslot filling to ensure latency predictability of BEV model and deadline satisfaction of PV models. The framework is implemented in compliance with the practical computing stack in modern autonomous vehicles. It is evaluated on a dual-SoC prototype connected via a PCIe bus. Results show that it achieves $1.52-1.63 \times$ speedup for the BEV model compared to no parallelism. It also ensures deadline satisfaction for PV models without interference in BEV model latency predictability.
Chao Li 0009, Xiaofeng Hou, Xinkai Wang 0003, Guangjun Bao, Bingchuan Sun, Shibo Rui, Minyi Guo
RTSS6
2024 A2: Towards Accelerator Level Parallelism for Autonomous Micromobility Systems
abstract
Autonomous micromobility systems (AMS) such as low-speed minicabs and robots are thriving. In AMS, multiple Deep Neural Networks execute in parallel on heterogeneous AI accelerators. An emerging paradigm called Accelerator Level Parallelism (ALP) suggests managing accelerators holistically. However, there lacks a specialized and practical solution populating ALP for an AMS, where the varying real-time requirements under different working scenarios bring an opportunity to dynamically tradeoff between latency and efficiency. Furthermore, accelerator heterogeneity introduces enormous configuration space, and the shared-memory architecture results in dynamic bandwidth interference. In this article, we propose A 2 , a novel AMS resource manager optimizing energy and memory space efficiency under variable latency constraints. We gain insight from prior Learn&Control scheme to design an Analyze&Adapt scheme specialized for heterogeneous AI accelerators under shared-memory architecture. It features analyzing the system thoroughly offline to support two-step adaptation online. We build a prototype of A 2 and evaluate it on a commercial edge platform. We show that A 2 achieves 32.8% improvements in power and 13.8% in memory compared with control-based methods. As for timeliness enhancement, A 2 reduces the deadline violation rate by 9.2 percentage points (12.8% → 3.6%) on average compared to directly porting Learn&Control methods.
Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Xinkai Wang 0003, Quan Chen 0002, Minyi Guo
ACM Trans. Archit. Code Optim.5
2023 Not All Resources are Visible: Exploiting Fragmented Shadow Resources in Shared-State Scheduler Architecture
abstract
With the rapid development of cloud computing, the increasing scale of clusters and task parallelism put forward higher requirements on the scheduling capability at scale. To this end, the shared-state scheduler architecture has emerged as the popular solution for large-scale scheduling due to its high scalability and utilization. In such an architecture, a central resource state view periodically updates the global cluster status to distributed schedulers for parallel scheduling. However, the schedulers obtain broader resource views at the cost of intermittently stale states, rendering resources released invisible to schedulers until the next view update. These fleeting resource fragments are referred to as shadow resources in this paper. Current shared-state solutions overlook or fail to systematically utilize the shadow resources, leaving a void in fully exploiting these invisible resources.
Xinkai Wang 0003, Yuancheng Li 0001, Chao Li 0009, Xiaofeng Hou, Jing Wang 0055, Quan Chen 0002, Jingwen Leng, Minyi Guo, Leibo Wang
SoCC1
2023 FIRST: Exploiting the Multi-Dimensional Attributes of Functions for Power-Aware Serverless Computing
abstract
Emerging cloud-native development models raise new challenges for managing server performance and power at microsecond scale. Compared with traditional cloud workloads, serverless functions exhibit unprecedented heterogeneity, variability, and dynamicity. Designing cloud-native power management schemes for serverless functions requires significant engineering effort. Current solutions remain sub-optimal since their orchestration process is often one-sided, lacking a systematic view. A key obstacle to truly efficient function deployment is the fundamental wide abstraction gap between the upper-layer request scheduling and the low-level hardware execution.In this work, we show that the optimal operating point (OOP) for energy efficiency cannot be attained without synthesizing the multi-dimensional attributes of functions. We present FIRST, a novel mechanism that enables servers to better orchestrate serverless functions. The key feature of FIRST is that it leverages a lightweight Internal Representation and meta-Scheduling (IRS) layer for collecting the maximum potential revenue from the servers. Specifically, FIRST follows a pipeline-style workflow. Its frontend components aim to analyze functions from different angles and expose their key features to the system. Meanwhile, its backend components are able to make informed function assignment decisions to avoid OOP divergence. We further demonstrate the way to create extensions based on FIRST to enable versatile cloud-native power management. In total, our design constitutes a flexible management layer that supports power-aware function deployment. We show that FIRST could allow 94% functions to be processed under the OOP, which brings up to 24% energy efficiency improvements.
Lu Zhang 0049, Chao Li 0009, Xinkai Wang 0003, Weiqi Feng, Zheng Yu 0003, Quan Chen 0002, Jingwen Leng, Minyi Guo, Shang Yue
IPDPS3
2022 Exploring Efficient Microservice Level Parallelism
abstract
The microservice architecture has recently become a driving trend in the cloud by disaggregating a monolithic application into many scenario-oriented service blocks (microservices). The decomposition process results in a highly dynamic execution scenario, in which various chained microservices contend for computing resources in different ways. While parallelism has been exploited at both the instruction/thread level and the task/request level, very limited work has been done with the grain-size of a microservice. Current parallel processing solutions are sub-optimal as they neither capture the unique characteristics of microservices nor consider the uncertainty arises in the microservice environment. In this work we introduce microservice level parallelism (MLP), a technique that aims to precisely coalesce and align parallel microservice chains for better system performance and resource utilization. We identify major issues that prevent servers from effectively exploiting MLP and we define metrics that can guide MLP optimization. We propose v-MLP, a volatility-aware MLP that is able to adapt to a highly heterogeneous and dynamic microservice environment. We show that v-MLP can reduce tail latency by up to 50% and improve resource utilization by up to 15 % under various scenarios.
Xinkai Wang 0003, Chao Li 0009, Lu Zhang 0049, Xiaofeng Hou, Quan Chen 0002, Minyi Guo
IPDPS1