VLDB 2026 Research / reviewers in the wild / expert
Yang Shi 0008
dblp:15/5233-8
· DBLP profile ↗
23ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0001-5786-3171ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 15 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DyGen: A Constant-Time Kernel Generator for Dynamic-Shape Neural NetworksabstractIn recent years, dynamic-shape neural networks have been widely adopted in intelligent applications, such as Mixture-of-Experts based large language models and computer vision tasks. However, in dynamic scenarios, operator shapes are determined at runtime. This leads to prohibitively expensive compilation times for existing static compilers, as they must search across a vast optimization space to identify the best configuration. To address the need for efficient optimization of dynamic-shape neural networks, we present DyGen (Dynamic-shape Kernel Generator)—a lightweight, two-stage compiler plug-in on GPU platforms. In the offline stage, DyGen employs deliberately crafted pruning rules to construct a compact candidate configuration set for the target hardware, then select the configuration of the high-performance kernel to train a configuration generation model. During the online stage, dynamic operator information is directly fed into the generator, which can quickly produce efficient kernel configurations without the need for costly search. Compared to state-of-the-art tensor compilers, DyGen improves inference performance by an average of 36%, while significantly reducing generation overhead from 9 seconds to 0.3 seconds. Yuhan Kang, Dong Chen 0015, Yang Shi 0008, Jianchao Yang, Zeyu Xue, Mei Wen |
DATE | 4 |
| 2026 | CVS: A Metric for Security-Aware Compilation against Side-Channel Attacks in Edge SoCs (WIP)abstractDeep learning compilers (DLCs) have become the standard approach for optimizing edge inference performance, employing techniques such as operator fusion, loop tiling, and scheduling to meet stringent resource constraints. Yet, the security implications of these optimizations remain largely unexplored. In this work, we investigate shared-memory side-channel attacks on edge SoCs and analyze how compiler optimizations reshape the leakage surface. Our study reveals that identical operators can exhibit distinct shared-resource access patterns under different compilation strategies, resulting in divergent attack outcomes. To address this, we introduce the Confusion Variance Score (CVS), a metric that quantifies compilation-induced security by measuring confusion in time-series resource traces (e.g., DRAM bandwidth). CVS integrates multidimensional dynamic time warping with statistical morphological features to ensure temporal robustness, and shows a strong negative correlation (Spearman r ≈ −0.9394) with practical attack error rates. Finally, we demonstrate the feasibility of CVS-guided compilation in TVM and TensorRT, achieving a 24 % increase in attack error rate compared to default strategies, while limiting inference latency overhead to under 5 %. Puhong Lei, Yang Shi 0008, Zhe Li 0017, Xing Mou |
LCTES | 3 |
| 2025 | WinAcc: Window-based Acceleration of Neural Networks Using Block Floating PointabstractDeep Neural Networks (DNNs) impose significant computational demands, necessitating optimizations for computational and energy efficiencies. Per-vector scaling, which applies a scaling factor to blocks of elements using narrow integer types, effectively reduces storage and computational overhead. However, the frequent occurrence of floating-point accumulations between vectors limits further improvements in energy efficiency. State-of-the-art accelerators address this challenge by grouping and summing vector products based on their exponent differences, thereby reducing the overhead associated with intra-group shifting and accumulation. Nevertheless, this approach increases the complexity of register usage and grouping logic, leading to limited energy benefits and hardware efficiency. In this context, we introduce WinAcc, a novel algorithm and architecture co-designed solution that utilizes a low-cost accumu-lator to handle the majority of data in DNNs, offering low area overhead and high energy efficiency gains. Our key insight is that the data of DNNs follows a Laplace-like distribution, which enables the use of a customized data format with a narrow dynamic range to encode most of the data. This allows for the design of a low-cost accumulator with narrow shifters and adders, significantly reducing reliance on floating-point accumulator and consequently improving energy efficiency. Compared with state-of-the-art architecture Bucket, WinAcc achieves 33.95% energy reduction across seven representative DNNs and reduces area by 9.5% while maintaining superior model performance. Xin Ju 0005, Mei Wen, Yasong Cao, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008 |
DATE | 8 |
| 2025 | SparSynergy: Unlocking Flexible and Efficient DNN Acceleration Through Multi-Level SparsityabstractTo more effectively address the computational and memory requirements of deep neural networks (DNNs), leveraging multi-level sparsity-including value-level and bit-level sparsity-has emerged as a pivotal strategy. While substantial research has been dedicated to exploring value-level and bit-level sparsity individually, the combination of both has largely been overlooked until now. In this paper, we propose SparSynergy, which-to the best of our knowledge-is the first accelerator that synergistically integrates multi-level sparsity into a unified framework, maximizing computational efficiency and minimizing memory usage. However, jointly considering multi-level sparsity is non-trivial, as it presents several challenges: (1) increased hardware overhead due to the complexity of incorporating multiple sparsity levels, (2) bandwidth-intensive data transmission during multiplexing, and (3) decreased throughput and scalability caused by bottlenecks in bit-serial computation. Our proposed SparSynergy addresses these challenges by introducing a unified sparsity format and a cooptimized hardware design. Experimental results demonstrate that SparSynergy achieves a 5.38 x geometric mean improvement in the energy-delay product (EDP) when compared with the tensor core, across workloads with varying degrees of sparsity. Furthermore, SparSynergy significantly improves accuracy retention compared to state-of-the-art accelerators for representative DNNs. Jingkui Yang, Mei Wen, Junzhong Shen, Jianchao Yang, Yasong Cao, Minjin Tang, Zhaoyun Chen, Yang Shi 0008 |
DATE | 9 |
| 2025 | SmartBlock: Adaptive Block Floating Point Quantization for Efficient DNN AccelerationabstractDeep Neural Networks (DNNs) have achieved remarkable success as model sizes continue to grow, driving the need for optimizations in both computational and energy efficiency. Block Floating Point (BFP) quantization has emerged as an effective model compression technique, offering a favorable trade-off between model accuracy and hardware cost. However, the frequent use of floating-point (FP) accumulation across BFP blocks remains a significant bottleneck, limiting further improvements in energy efficiency. State-of-the-art (SotA) accelerators mitigate this issue by introducing low-overhead accumulators with a narrower dynamic range ahead of the FP accumulator to handle a small range of values. While this approach reduces the activation of power-hungry alignment and format conversion units, it increases the complexity of the processing elements (PEs), thereby limiting the overall energy savings. Xin Ju 0005, Jingkui Yang, Mei Wen, Minjin Tang, Zhaoyun Chen, Yang Shi 0008 |
ICPP | 8 |
| 2025 | SpikeFlow: A hardware-software co-designed systolic array for spiking neural networks
Yang Shi 0008, Zhaoyun Chen, Mei Wen |
J. Syst. Archit. | 2 |
| 2025 | ESCAN: Efficient GPU sharing for cascade neural network inference
Yang Shi 0008, Zhaoyun Chen, Mei Wen |
Neural Networks | 2 |
| 2025 | MAP-SIM: A DNN-Specific Mapping Optimization Framework for Shared-Memory CPU-Systolic Array ArchitecturesabstractAs performance demands continue to rise, Shared-Memory Heterogeneous Systems (SMHSs) have been widely adopted for their ability to enable efficient communication and data sharing between different heterogeneous cores. However, existing SMHS face challenges in uneven workload distribution among heterogeneous cores and suboptimal mapping schemes, preventing them from fully leveraging their architectural advantages. To address these issues, this paper proposes a mapping-aware framework for modeling SMHSs called MAP-SIM. By performing performance modeling for CPUs and Systolic Arrays (SAs), and considering rational schemes for the partition and mapping of computational tasks, MAP-SIM aims to evaluate and optimize the computational performance of heterogeneous multicore architectures. The experimental results show that compared to previous work, MAP-SIM can increase simulation speed by 14 to 67 times and can also enhance the computational performance of SMHS by 1.4 to 4.4 times. Mei Wen, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008, Tianyu Wang 0009, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | SSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor CoresabstractSparse-Dense Matrix-Matrix Multiplication (SpMM) has emerged as a foundational primitive in HPC and AI. Recent advancements have aimed to accelerate SpMM by harnessing the powerful Tensor Cores found in modern GPUs. However, despite these efforts, existing methods frequently encounter performance degradation when ported across different Tensor Core architectures. Recognizing that scalable SpMM across multiple generations of Tensor Cores relies on the effective use of general-purpose instructions, we have meticulously developed a SpMM library named SSpMM. However, a significant conflict exists between granularity and performance in current Tensor Core instructions. To resolve this, we introduce the innovative Transpose Mapping Scheme, which elegantly implements fine-grained kernels using coarse-grained instructions. Additionally, we propose the Register Shuffle Method to further enhance performance. Finally, we introduce Sparse Vector Compression, a technique that ensures our kernels are scalable with both structured and unstructured sparsity. Our experimental results, conducted on four generations of Tensor Core GPUs using over 3,000 sparse matrices from well established matrix collections, demonstrate that SSpMM achieves an average speedup of 2.04×, 2.81×, 2.07×, and 1.87×, respectively, over the state-of-the-art SpMM solution. Furthermore, we have integrated SSpMM into PyTorch, achieving a 1.81× speedup in end-to-end Transformer inference compared to cuDNN. Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang, Zhongdi Luo, Yang Shi 0008, Zhaoyun Chen, Junzhong Shen, Johannes Langguth |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2025 | FAMS: A FrAmework of Memory-Centric Mapping for DNNs on Systolic Array AcceleratorsabstractIn recent years, deep neural networks (DNNs) have experienced rapid development. These DNNs demonstrate significant variations in architecture and scale, creating a substantial demand for domain-specific accelerators that are optimized for both high performance and low energy consumption. Systolic array accelerators, due to their efficient dataflow and parallel processing capabilities, offer significant advantages when performing computations for DNNs. Existing studies frequently overlook various hardware constraints in systolic array accelerators when representing mapping strategies. This oversight includes ignoring the differences in delays between communication and computation operations, as well as overlooking the capacities of multilevel memory hierarchies. Such omissions can lead to inaccuracies in predicting accelerator performance and inefficiencies in system design. We propose the FAMS framework, which introduces a memory-centric notation capable of fully representing the mapping of DNN operations on systolic array accelerators. Memory-centric notation moves away from the idealized assumptions of previous notations and considers various hardware constraints, thereby expanding the effective design and mapping spaces. The FAMS framework also includes a cycle-accurate simulator, which takes the hardware configurations, task descriptions, and mapping strategy represented by memory-centric notation as inputs, providing various metrics such as latency and energy consumption. The experimental results demonstrate that our proposed FAMS framework reduces latency by up to 29.7% and increases throughput by 42.4% compared to the state-of-the-art TENET framework. Additionally, under hardware configurations with a MAC delay of 2 and 3 clock cycles, the FAMS framework enhances performance by 12.0% and 25.4%, respectively. Hao Sun 0023, Junzhong Shen, Zhongyi Tang, Changwu Zhang, Yang Shi 0008, Hengzhu Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | MAP-SIM: A Performance Model for Shared-Memory Heterogeneous Systems with Mapping Awareness
Mei Wen, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008 |
ICA3PP (1) | 5 |
| 2024 | HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUsabstractThe widespread adoption of GPUs has driven the development of GPU simulators, which, in turn, lead advancements in both GPU architectures and software optimization. Trace-driven cycle-accurate Cycle-accurate simulators, which provide detailed microarchitectural models and clock-level precision, come at the cost of extended simulation times and require high computational resources. Their scalability has become a bottleneck. A growing trend is the adoption of cycle-approximate simulators, which introduce mathematical modeling of partial hardware units and utilize sampling to accelerate simulation. However, this approach faces challenges regarding the accuracy of performance predictions. To address these limitations, we introduce HyFiSS, a hybrid fidelity stall-aware GPU simulator. HyFiSS features fine-grained stall events tracking and attribution by constructing a detailed execution pipeline model for various stall events on Streaming Multiprocessors (SMs). It accurately emulates the thread block scheduler behavior using real-time scheduling logs and utilizes sampling based on thread block sets to minimize the precision loss due to fine-grained sampling points on the microarchitectural state. We achieve a balance between reliability, speed, and the level of simulation detail, especially regarding bottlenecks. By evaluating a diverse set of benchmarks, HyFiSS achieves a mean absolute percentage error in predicting active cycles that is comparable to the state-of-the-art cycle-accurate simulator Accel-Sim. Moreover, HyFiSS achieves a substantial 12.8 × speedup in the simulation efficiency compared to Accel-Sim. HyFiSS also requires at least 3.2 × less disk storage than both Accel-Sim and another state-of-the-art cycle-approximate simulator PPT-GPU due to its efficient SASS (Streaming Assembler) traces compression. With precise, per-cycle stall events statistics, HyFiSS can provide accurate GPU performance metrics and stall cause reporting. This significantly simplifies performance analysis, bottleneck identification, and performance optimization tasks for researchers, making it easier to enhance GPU performance effectively. Jianchao Yang, Mei Wen, Dong Chen 0015, Zhaoyun Chen, Zeyu Xue, Junzhong Shen, Yang Shi 0008 |
MICRO | 8 |
| 2024 | ESEN: Efficient GPU sharing of Ensemble Neural Networks
Yang Shi 0008, Zhaoyun Chen, Mei Wen |
Neurocomputing | 2 |
| 2024 | Optimizing VLIW Instruction Scheduling via a Two-Dimensional Constrained Dynamic ProgrammingabstractTypical embedded processors, such as Digital Signal Processors (DSPs), usually adopt Very Long Instruction Word (VLIW) architecture to improve computing efficiency. The performance of VLIW processors heavily relies on Instruction-Level Parallelism (ILP). Therefore, it is crucial to develop an efficient instruction scheduling algorithm to explore more ILP. While heuristic algorithms are widely used in modern compilers due to simple implementation and low computational cost, they have limitations in providing accurate solutions and are prone to local optima. On the other hand, exact algorithms can usually find the optimal solution, but their high time overhead makes them less suitable for large-scale problems. This article proposes a two-dimensional constrained dynamic programming (TDCDP) approach and a quantitative model for instruction scheduling. The TDCDP approach achieves near-optimal solutions within an acceptable time overhead. Furthermore, we integrate our TDCDP approach into mainstream compiler architecture, encompassing Pre- and Post-RA (register allocation) scheduling. We conduct a quantitative evaluation of TDCDP compared with four heuristic algorithms on a typical VLIW processor. Our approach achieves an efficiency improvement of up to 58.34% in final solutions compared with the heuristic algorithms. Additionally, the Post-RA Scheduling enhances programs with an average speedup of 14.04% than solely applying the Pre-RA Scheduling. Can Deng, Zhaoyun Chen, Yang Shi 0008, Yimin Ma, Mei Wen, Lei Luo 0002 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Automatic End-to-End Joint Optimization for Kernel Compilation on DSPsabstractDigital signal processors (DSPs) commonly adopt VLIW-SIMD architecture and are extensively applied in most compute-heavy embedded sensing applications. The performances for DSP kernels rely heavily on compilations and handwritten optimizations. Hand-crafted methods suffer from heavy burden on programmers, while state-of-the-art automatic compilation methods always focus more on a certain aspect (tiling or auto-vectorization), lacking of global and sequential vision on the intact compilation optimization process. It still requires empirical adjustments by programmers in the actual scenario.In order to release programmers from kernel tuning, we propose JOKer, an automatic end-to-end multi-level code generator for kernel joint optimization on DSPs. JOKer integrates means of optimizations in compiling process and provides an end-to-end workflow for performance tuning. It explores compilation configurations through a reinforcement learning based agent for global optimal solution and generates high performance kernel codes for DSPs automatically. Zhaoyun Chen, Yang Shi 0008, Mei Wen, Chunyun Zhang |
DAC | 3 |
| 2023 | Releasing the Potential of Tensor Core for Unstructured SpMM using Tiled-CSR FormatabstractThe GPU has become a popular platform for AI applications, thanks in part to its Tensor Cores that address performance issues. However, the Sparse Matrix Multiplication (SpMM) kernel has remained a bottleneck despite significant advances in computing power. Due to the hardware mechanism of the Tensor Core, its programming granularity does not match SpMM. In this paper, we analyze the reasons why the unstructured SpMM kernel is not suitable for the Tensor Core, and propose the Tiled Compressed Sparse Row (Tiled-CSR) compression format. To address the issue of low non-zero rates in Tiled-CSR format, we exploit the row shuffle algorithm to improve the utilization of Tensor Cores and enhance computing density. We also utilize adaptive memory access modes and 3D-Grid tiling for the SpMM kernel to reduce memory access latency. The experimental results on NVIDIA A100 GPU with matrices in the Deep Learning Matrix Collection (DLMC) demonstrate that the Tiled-CSR format improves the utilization of Tensor Cores under different sparsity, with a maximum of 3.89× at 50% sparsity and a minimum of 1.82× at 90% sparsity compared to the SR-BCRS format. Additionally, our kernel achieves an average speedup of 1.54×(up to 2.12×) over Magicube. Zeyu Xue, Mei Wen, Zhaoyun Chen, Yang Shi 0008, Minjin Tang, Jianchao Yang, Zhongdi Luo |
ICCD | 4 |
| 2022 | Exploring ILP for VLIW Architecture by Quantified Modeling and Dynamic Programming-Based Instruction SchedulingabstractExploring the instruction-level parallelism (ILP) of Very Long Instruction Word (VLIW) architecture relies on instruction scheduling. List scheduling (LS) algorithms, which are most adopted in modern compilers, have limitations in searching for optimal solutions. This paper proposes a quantifiable model for instruction scheduling and a dynamic programming-based strategy (DPS). We evaluate DPS on a specified platform and realize high efficiency. The results suggest that the DPS achieves an efficiency improvement of up to 44.72% within acceptable time overhead. Can Deng, Zhaoyun Chen, Yang Shi 0008, Xichang Kong, Mei Wen |
ASP-DAC | 3 |
| 2022 | Light: A Component Enhances Faster and More Accurate Traffic Measurement*abstractThe greatest challenge when designing an online sketch method for data flow measurement is to reduce the storage cost of sketches with little loss of accuracy and obtain a higher bandwidth. To address this issue, we proposed Light component. By storing elephant and mice flows separately, the accuracy of the Light-enhanced sketches is substantially improved, and the processing speed is significantly faster due to a significant reduction in the required computational overhead. We implement Light-enhanced sketches alongside several existing sketch methods on CPU and compare their performance. Experiments show that, under the same storage conditions, Light-enhanced sketches can greatly reduce the Average Relative Error by 1.80 to 5.07 times, as well as increase the average processing speed of each packet by 4.96 to 9.05 times compared with their original structure. Our approach also achieves a stable performance in the measurement of traffic in different traffic distributions. More importantly, Light component can be deployed to different sketch methods, which demonstrates its reusability. Jianchao Yang, Mei Wen, Yang Shi 0008 |
ICC | 4 |
| 2021 | sRouting: Towards a Better Flow Size Estimation Performance through Routing and Sketch ConfigurationabstractFlow size estimation is highly important and beneficial to various applications, including traffic engineering and anomaly detection. Considering the resource constraints and performance requirements in place, sketches are widely used to accomplish this task. However, when sketches are applied in real-life networks, the measurement performance is often decreased due to many practical reasons, including partial deployment of sketches, unique traffic characteristics, etc. In this paper, we present sRouting, a practical framework that aims to better utilize the deployed sketches. sRouting improves the measurement performance by optimizing which flows are monitored and where. We first investigate the relationship between the accuracy of a given sketch and the total number of packets, then formulate the offline problem in sRouting as an integer linear programming problem. To solve this problem efficiently, we devise a rounding-based algorithm and provide its performance guarantees. Furthermore, to handle dynamic changes in the network, we design an online adjustment algorithm capable of responding appropriately to these changes. Through experiments using real traces and typologies, we demonstrate that sRouting can significantly improve the volume of monitored traffic and reduce the measurement error without negatively impacting the network throughput. Yang Shi 0008, Mei Wen |
ICPP | 1 |
| 2021 | Automatic mapping and code optimization for OpenCL kernels on FT-matrix architecture (WIP paper)abstractFT-Matrix is a typical vector-SIMD architecture that refines the cooperation between scalar and vector units. This approach is widely used in digital signal processing, high-performance computing, and artificial intelligence, among other fields. FT-Matrix currently adopts C vector extension as the main programming model, improving the utilization efficiency of SIMD by providing explicit vector extension API. Moreover, it is difficult to efficiently transplant parallel programs (OpenCL, CUDA) adopted by users. This paper proposes an automatic mapping and code optimization method for OpenCL kernels on FT-Matrix architecture. The proposed approach solves these challenges by means of work item coalescing, slicing and rotation, and instruction-level code optimization. Preliminary results show that our method can achieve high performance and good hardware utilization for OpenCL kernels, as well as decreasing the programming difficulty on FT-Matrix. Mei Wen, Zhaoyun Chen, Yang Shi 0008, Chunyuan Zhang |
LCTES | 4 |
| 2020 | Towards High-Efficiency Data Centers via Job-Aware Network SchedulingabstractDistributed jobs typically facing competition for multiple resources in modern data centers, especially for network. Without effective network scheduling, this competition can cause low efficiency of the data center. Previous work on network scheduling has been focused on reducing flow completion time or improving per-flow fairness. Yet, its effect on improving jobs’ performance is limited by the unawareness of relationships between communication and computation. In this paper, we focus on the problem of scheduling network resources for multiple jobs, with the specific objective to reduce the job completion time (JCT), which also makes the datacenter more efficient. With an in-depth investigation of communication and computation, we identify an opportunity for accelerating jobs in a way that occupies less bandwidth for DAG-based complicated modern jobs. Accordingly, this paper proposes JIT, a job-aware network scheduler that leverages the computational graph to accelerate jobs effectively. To cater to the goal of JIT, we first develop a mathematical model and formulate the scheduling problem as an integer linear programming (ILP) problem. We further prove that it has an equivalent linear programming (LP) problem through rigorous theoretical analysis in order to solve this ILP problem efficiently. Some reasonable simplifications are also adopted to reduce the solving time of JIT to only 1 second. The proposed JIT is simulated and compared against some state-of-the-art designs, and the simulation results demonstrate that JIT can achieve an acceleration of up to 1.55 × , which successfully improves the efficiency of the data center. Yang Shi 0008, Mei Wen, Chunyuan Zhang |
ICPP | 1 |
| 2020 | Incremental Deployment of Programmable Switches for Sketch-based Network MeasurementabstractThe emergence of programmable switches has boosted lots of research around many network aspects: mea-surements, security, quality of services. To explore the ad-vantages of programmable data planes while preserving the legacy networking systems, deploying programmable switches incrementally may be a more practical solution. In this paper, we deal with the programmable switch deploy problem for sketch-based network measurement, which has been overlooked before. We first analyze the desired properties of a good deployment for sketch-based network measurement with some examples. Based on summarized lessons, we then develop two Integer Linear Programming (ILP) models, namely TraceILP and TopoILP, to solve the deployment problem. If historical traffic traces are provided, TraceILP generates better deployment with historical information. Even if no traces are provided, TopoILP can still make a reasonable strategy according to the network topology. Evaluations on real ISP and datacenter topologies show that pro-posed models guarantee a promising measurement performance with only about 40% devices upgraded to programmable ones. Yang Shi 0008, Mei Wen, Chunyuan Zhang |
ISCC | 1 |
| 2019 | KVSwitch: An In-network Load Balancer for Key-Value StoresabstractToday's cloud-based online services are underpinned by distributed key-value stores (KVSs). Keys and values are distributed across back-end servers in such scale-out systems. One primary real-life performance bottleneck occurs when storage servers suffer from load imbalance under skewed workloads. In this paper, we present KVSwitch, a centralized self-managing load balancer that leverages the power and flexibility of emerging programmable switches. The balance is achieved through dynamically predicting the hot items and creating replication strategies according to KVS loading. To overcome the challenges in realizing KVSwitch given the limitations of the switch hardware, we decompose KVSwitch's functions and carefully design them for the heterogeneous processors inside the switch. We prototype KVSwitch in a Tofino switch. Experimental results show that our solution can effectively keep the KVS servers balanced even under highly skewed workloads. Furthermore, KVSwitch only replicates 70% of hot items and consumes 9.88% of server memory rather than simply replicating all hot items to each server. Yang Shi 0008, Jiawei Fei, Mei Wen, Chunyuan Zhang |
ISCC | 1 |