EDBT 2026 Demo / reviewers in the wild / expert
Yipeng Wang 0002
dblp:36/4627-2
· DBLP profile ↗
13ranked-venue papers
6as first author
4since 2021 · last 2024
0009-0004-0215-019XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 3 first-author · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Litmus: Fair Pricing for Serverless ComputingabstractServerless computing has emerged as a market-dominant paradigm in modern cloud computing, benefiting both cloud providers and tenants. While service providers can optimize their machine utilization, tenants only need to pay for the resources they use. To maximize resource utilization, these serverless systems co-run numerous short-lived functions, bearing frequent system condition shifts. When the system gets overcrowded, a tenant's function may suffer from disturbing slowdowns. Ironically, tenants also incur higher costs during these slowdowns, as commercial serverless platforms determine costs proportional to their execution times. Qi Pei, Yipeng Wang 0002, Seunghee Shin |
ASPLOS (4) | 2 |
| 2023 | Rambda: RDMA-driven Acceleration Framework for Memory-intensive µs-scale Datacenter ApplicationsabstractResponding to the "datacenter tax" and "killer microseconds" problems for memory-intensive datacenter applications, diverse solutions including Smart NIC-based ones have been proposed. Nonetheless, they often suffer from high overhead of communications over network and/or PCIe links. To tackle the limitations of the current solutions, this paper proposes RAMBDA, a holistic network and architecture co-design solution that leverages current RDMA and emerging cache-coherent off-chip interconnect technologies. Specifically, RAMBDA consists of four hardware and software components: (1) unified abstraction of inter- and intra-machine communications synergistically managed by one-sided RDMA write and cache-coherent memory write; (2) efficient notification of requests to accelerators assisted by cache coherence; (3) cache-coherent accelerator architecture directly interacting with NIC; and (4) adaptive device-to-host data transfer for modern server memory systems comprising both DRAM and NVM exploiting state-of-the-art features in CPUs and PCIe. We prototype RAMBDA with a commercial system and evaluate three popular datacenter applications: (1) in-memory key-value store, (2) chain replication-based distributed transaction system, and (3) deep learning recommendation model inference. The evaluation shows that RAMBDA provides 30.1~69.1% lower latency, 0.2~2.5× throughput, and ~ 3× higher energy efficiency than the current state-of-the-art solutions, including Smart NIC. For those cases where Rambda performs poorly, we also envision future architecture to improve it. Jinghan Huang 0001, Jacob Nelson 0001, Dan R. K. Ports, Yipeng Wang 0002, Ren Wang 0001, Tsung-Yuan Charlie Tai, Nam Sung Kim |
HPCA | 7 |
| 2021 | QEI: Query Acceleration Can be Generic and Efficient in the CloudabstractData query operations of different data structures are ubiquitous and critical in today's data center infrastructures and applications. However, query operations are not always performance-optimal to be executed on general-purpose CPU cores. These operations exhibit insufficient memory-level parallelism and frontend bottlenecks due to unstructured control flow. Furthermore, the data access patterns are not cache- or prefetch-friendly. Based on our performance analysis on a commodity server, query operations can consume a large percentage of the CPU cycles in various modern cloud workloads. Existing accelerator solutions for query operations do not strike a balance between their generality, scalability, latency, and hardware complexity. In this paper, we propose QEI, a generic, integrated, and efficient acceleration solution for various data structure queries. We first abstract the query operations to a few regular steps and map them to a simple and hardware-friendly configurable finite automaton model. Based on this model, we develop the QEI architecture that allows multiple query operations to execute in parallel to maximize throughput. We also propose a novel way to integrate the accelerator into the CPU that balances performance, latency, and hardware cost. QEI keeps the main control logic near the L2 cache to leverage existing hardware resources in the core while distributing the data-intensive comparison logic to each last-level cache slice for higher parallelism. Our results with five representative data center workloads show that QEI can achieve 6.5× ~11.2× performance improvement in various scenarios with low overhead. Yipeng Wang 0002, Ren Wang 0001, Rangeen Basu Roy Chowdhury, Tsung-Yuan Charlie Tai, Nam Sung Kim |
HPCA | 2 |
| 2021 | Don't Forget the I/O When Allocating Your LLCabstractIn modern server CPUs, last-level cache (LLC) is a critical hardware resource that exerts significant influence on the performance of the workloads, and how to manage LLC is a key to the performance isolation and QoS in the cloud with multi-tenancy. In this paper, we argue that in addition to CPU cores, high-speed I/O is also important for LLC management. This is because of an Intel architectural innovation – Data Direct I/O (DDIO) – that directly injects the inbound I/O traffic to (part of) the LLC instead of the main memory. We summarize two problems caused by DDIO and show that (1) the default DDIO configuration may not always achieve optimal performance, (2) DDIO can decrease the performance of non-I/O workloads that share LLC with it by as high as 32%.We then present, the first LLC management mechanism that treats the I/O as the first-class citizen. Iat monitors and analyzes the performance of the core/LLC/DDIO using CPU’s hardware performance counters and adaptively adjusts the number of LLC ways for DDIO or the tenants that demand more LLC capacity. In addition, Iat dynamically chooses the tenants that share its LLC resource with DDIO to minimize the performance interference by both the tenants and the I/O. Our experiments with multiple microbenchmarks and real-world applications demonstrate that with minimal overhead, Iat can effectively and stably reduce the performance degradation caused by DDIO. Mohammad Alian, Yipeng Wang 0002, Ren Wang 0001, Ilia Kurakin, Tsung-Yuan Charlie Tai, Nam Sung Kim |
ISCA | 3 |
| 2020 | RLDRM: Closed Loop Dynamic Cache Allocation with Deep Reinforcement Learning for Network Function VirtualizationabstractNetwork function virtualization (NFV) technology attracts tremendous interests from telecommunication industry and data center operators, as it allows service providers to assign resource for Virtual Network Functions (VNFs) on demand, achieving better flexibility, programmability, and scalability. To improve server utilization, one popular practice is to deploy best effort (BE) workloads along with high priority (HP) VNFs when high priority VNF's resource usage is detected to be low. The key challenge of this deployment scheme is to dynamically balance the Service level objective (SLO) and the total cost of ownership (TCO) to optimize the data center efficiency under inherently fluctuating workloads. With the recent advancement in deep reinforcement learning, we conjecture that it has the potential to solve this challenge by adaptively adjusting resource allocation to reach the improved performance and higher server utilization. In this paper, we present a closed-loop automation system RLDRM11RLDRM: Reinforcement Learning Dynamic Resource Management to dynamically adjust Last Level Cache allocation between HP VNFs and BE workloads using deep reinforcement learning. The results demonstrate improved server utilization while maintaining required SLO for the HP VNFs. Bin Li 0018, Yipeng Wang 0002, Ren Wang 0001, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Zhu Zhou, Andrew Herdrich, Ameer Haj-Ali, Ion Stoica, Krste Asanovic |
NetSoft | 2 |
| 2019 | HALO: accelerating flow classification for scalable packet processing in NFVabstractNetwork Function Virtualization (NFV) has become the new standard in the cloud platform, as it provides the flexibility and agility for deploying various network services on general-purpose servers. However, it still suffers from sub-optimal performance in software packet processing. Our characterization study of virtual switches shows that the flow classification is the major bottleneck that limits the throughput of the packet processing in NFV, even though a large portion of the classification rules can be cached in the last level cache (LLC) in modern servers. Yipeng Wang 0002, Ren Wang 0001, Jian Huang 0006 |
ISCA | 2 |
| 2017 | Optimizing Open vSwitch to Support Millions of FlowsabstractSoftware switch has emerged as a critical component in software defined networking and network virtualization areas. Open vSwitch (OvS) is a widely used software switch which uses tuple space search algorithm for packet classification, and an exact match cache (EMC) for caching most frequently used flows. In this paper, we propose two new optimizations for OvS to further improve its performance and scalability. First aims to completely remove the sequential search overhead of the tuple space search layer of OvS, and second is a dynamic cache insertion optimization for the EMC to improve EMC effectiveness. We show that the optimizations can improve OvS's throughput by up to 3.5X for millions of active flows. Yipeng Wang 0002, Tsung-Yuan Charlie Tai, Ren Wang 0001, Sameh Gobriel, Janet Tseng, James Tsai |
GLOBECOM | 1 |
| 2017 | ObfusMem: A Low-Overhead Access Obfuscation for Trusted MemoriesabstractTrustworthy software requires strong privacy and security guarantees from a secure trust base in hardware. While chipmakers provide hardware support for basic security and privacy primitives such as enclaves and memory encryption. these primitives do not address hiding of the memory access pattern, information about which may enable attacks on the system or reveal characteristics of sensitive user data. State-of-the-art approaches to protecting the access pattern are largely based on Oblivious RAM (ORAM). Unfortunately, current ORAM implementations suffer from very significant practicality and overhead concerns, including roughly an order of magnitude slowdown, more than 100% memory capacity overheads, and the potential for system deadlock. Amro Awad, Yipeng Wang 0002, Deborah Shands, Yan Solihin |
ISCA | 2 |
| 2017 | Clone morphing: Creating new workload behavior from existing applicationsabstractComputer system designers need a deep understanding of end users' workload in order to arrive at an optimum design. However, current design practices suffer from two problems: time mismatch where designers rely on workloads available today to design systems that will be produced years into the future to run future workloads, and sparse behavior where many performance behavior is not represented by the limited set of applications available today. We propose clone morphing, a systematic method for producing new synthetic workloads (morphs) with performance behavior that does not currently exist. The morphs are generated automatically without knowing or changing the original application's source code. There are three different aspects a morph can differ from the original benchmark it is built on: temporal locality, spatial locality, and memory footprint. We showed how each of these aspects can be varied largely independently of other aspects. Furthermore, we also presented a method for merging two different applications into one that has an average behavior of both applications. We evaluated the morphs by running them on simulators and collect statistics that capture their behavior, and validated that morphs can be used for projecting future workloads and for generating new behavior that fills up the behavior map densely. Yipeng Wang 0002, Amro Awad, Yan Solihin |
ISPASS | 1 |
| 2016 | CAF: Core to Core Communication Acceleration FrameworkabstractAs the number of cores in a multicore system increases, core-to-core (C2C) communication is increasingly limiting the performance scaling of workloads that share data frequently. The traditional way cores communicate is by using shared memory space between them. However, shared memory communication fundamentally involves coherence invalidations and cache misses, which cause large performance overheads and incur a high amount of network traffic. Many important workloads incur significant C2C communication and are affected significantly by the costs, including pipelined packet processing which is widely used in software-based networking solutions. In these workloads, threads run on different cores and pass packets from one core to another for different stages of processing using software queues. Yipeng Wang 0002, Ren Wang 0001, Andrew Herdrich, James Tsai, Yan Solihin |
PACT | 1 |
| 2015 | MeToo: Stochastic Modeling of Memory Traffic Timing BehaviorabstractThe memory subsystem (memory controller, bus, andDRAM) is becoming a bottleneck in computer system performance. Optimizing the design of the multicore memory subsystem requires good understanding of the representative workload. A common practice in designing the memory subsystem is to rely on trace simulation. However, the conventional method of relying on traditional traces faces two major challenges. First, many software users are apprehensive about sharing their code (source or binaries) due to the proprietary nature of the code or secrecy of data, so representative traces are sometimes not available. Second, there is a feedback loop where memory performance affects processor performance, which in turnalters the timing of memory requests that reach the bus. Such feedback loop is difficult to capture with traces. In this paper, we present MeToo, a framework for generating synthetic memory traffic for memory subsystem design exploration. MeToo uses a small set of statistics that summarizes the performance behavior of the original applications, and generates synthetic traces or executables stochastically, allowing applications to remain proprietary. MeToo uses novel methods for mimicking the memory feedback loop. We validate MeToo clones, and show very good fit with the original applications' behavior, with an average error of only 4.2%, which is a small fraction of the errors obtained using geometric inter-arrival(commonly used in queueing models) and uniform inter-arrival. Yipeng Wang 0002, Ganesh Balakrishnan, Yan Solihin |
PACT | 1 |
| 2015 | Emulating cache organizations on real hardware using performance cloningabstractComputer system designers need a deep understanding of end users' workload in order to arrive at an optimum design. Unfortunately, many end users will not share their software to designers due to the proprietary or confidential nature of their software. Researchers have proposed workload cloning, which is a process of extracting statistics that summarize the behavior of users' workloads through profiling, followed by using them to drive the generation of a representative synthetic workload (clone). Clones can be used in place of the original workloads to evaluate computer system performance, helping designers to understand the behavior of users workload on the simulated machine models without the users having to disclose proprietary or sensitive information about the original workload. In this paper, we propose infusing environment-specific information into the clone. This Environment-Specific Clone (ESC) enables the simulation of hypothetical cache configurations directly on a machine with a different cache configuration. We validate ESC on both real systems as well as cache simulations. Furthermore, we present a case study of how page mapping affects cache performance. ESC enables such a study at native machine speed by infusing the page mapping information into clones, without needing to modify the OS or hardware. We then analyze the factors that determine how page mapping impact cache performance, and how various applications are affected differently. Yipeng Wang 0002, Yan Solihin |
ISPASS | 1 |
| 2013 | XAMP: An eXtensible Analytical Model PlatformabstractAnalytical modeling plays a unique and important role in computer architecture design, providing a first-order estimate, reducing the design search space size for simulation, or giving insights on basic relationships between various variables and parameters. However, current practices in using analytical models face considerable hurdles: models vary widely in types, predictive capability, and assumptions they rely on, creating a high barrier of entry. Consequently, the high barrier prevents wide adoption of analytical models. In this paper, we propose a tool that integrates analytical models into a single extensible platform that provides a uniform interface and model query engine. Users can perform input-based search (searching what outputs can be predicted by models given a set of available inputs), output-based search (searching what inputs are required in order for models to produce the desired outputs), or both combined. The query engine also analyzes various ways in which models can be connected so that they can be combined to extend the prediction capabilities beyond individual models' capabilities. The tool can be extended to include new models, allowing it to evolve as more models are added, and allowing collaborative model development. We refer to the tool as XAMP (eXtensible Analytical Model Platform). We believe XAMP is a useful first step toward lowering the barrier of entry for using and developing analytical models. Yipeng Wang 0002, Yan Solihin |
ISPASS | 1 |