EDBT 2026 Demo / reviewers in the wild / expert
Tsung-Yuan Charlie Tai
dblp:82/8969 · also Charlie Tai
· DBLP profile ↗
16ranked-venue papers
0as first author
7since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 since 2021Computer networks · 6 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Rambda: RDMA-driven Acceleration Framework for Memory-intensive µs-scale Datacenter ApplicationsabstractResponding to the "datacenter tax" and "killer microseconds" problems for memory-intensive datacenter applications, diverse solutions including Smart NIC-based ones have been proposed. Nonetheless, they often suffer from high overhead of communications over network and/or PCIe links. To tackle the limitations of the current solutions, this paper proposes RAMBDA, a holistic network and architecture co-design solution that leverages current RDMA and emerging cache-coherent off-chip interconnect technologies. Specifically, RAMBDA consists of four hardware and software components: (1) unified abstraction of inter- and intra-machine communications synergistically managed by one-sided RDMA write and cache-coherent memory write; (2) efficient notification of requests to accelerators assisted by cache coherence; (3) cache-coherent accelerator architecture directly interacting with NIC; and (4) adaptive device-to-host data transfer for modern server memory systems comprising both DRAM and NVM exploiting state-of-the-art features in CPUs and PCIe. We prototype RAMBDA with a commercial system and evaluate three popular datacenter applications: (1) in-memory key-value store, (2) chain replication-based distributed transaction system, and (3) deep learning recommendation model inference. The evaluation shows that RAMBDA provides 30.1~69.1% lower latency, 0.2~2.5× throughput, and ~ 3× higher energy efficiency than the current state-of-the-art solutions, including Smart NIC. For those cases where Rambda performs poorly, we also envision future architecture to improve it. Jinghan Huang 0001, Jacob Nelson 0001, Dan R. K. Ports, Yipeng Wang 0002, Ren Wang 0001, Tsung-Yuan Charlie Tai, Nam Sung Kim |
HPCA | 9 |
| 2023 | PROMPT: Learning dynamic resource allocation policies for network applications
Drew Penney, Bin Li 0018, Jaroslaw J. Sydir, Lizhong Chen, Tsung-Yuan Charlie Tai, Stefan Lee, Eoin Walsh, Thomas Long |
Future Gener. Comput. Syst. | 5 |
| 2023 | RAPID: Enabling fast online policy learning in dynamic public cloud environments
Drew Penney, Bin Li 0018, Lizhong Chen, Jaroslaw J. Sydir, Anna Drewek-Ossowicka, Ramesh Illikkal, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Andrew Herdrich |
Neurocomputing | 7 |
| 2022 | DPM-NFV: Dynamic Power Management Framework for 5G User Plane Function using Bayesian OptimizationabstractNetwork Function Virtualization (NFV), the replacement of purpose-built network appliances with software functions running on general purpose compute servers, is ubiquitous in today's telecommunication networks. The 5G User Plane Function (UPF) is an important example of an NFV workload, which enables 5G and internet communications. The UPF has strict packet drop requirements and because user traffic load can vary dramatically throughout the day, the selection of a single static configuration leads to over-provisioning of server resources. To reduce the cost of ownership, network operators can reduce power consumption during periods of low traffic load, but to do so they must ensure that packet drop requirements are met. In this paper we present DPM-NFV, a machine learning based framework that enables dynamic tuning of a real NFV system. Our methodology is composed of two phases: (1) Offline, targeted automated studies use Bayesian Optimization to infer the best configurations for various load levels; (2) Online, a run-time classifier dynamically selects the best configuration for the current load. Our results obtained on a real system demonstrate that the UPF can meet strict packet drop requirements while reducing power consumption by up to 52% with smooth traffic and up to 46% with bursty traffic. Jaroslaw J. Sydir, Bin Li 0018, Pietro Mercati, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Michael Kishinevsky, Boris Serafimov |
GLOBECOM | 4 |
| 2021 | QEI: Query Acceleration Can be Generic and Efficient in the CloudabstractData query operations of different data structures are ubiquitous and critical in today's data center infrastructures and applications. However, query operations are not always performance-optimal to be executed on general-purpose CPU cores. These operations exhibit insufficient memory-level parallelism and frontend bottlenecks due to unstructured control flow. Furthermore, the data access patterns are not cache- or prefetch-friendly. Based on our performance analysis on a commodity server, query operations can consume a large percentage of the CPU cycles in various modern cloud workloads. Existing accelerator solutions for query operations do not strike a balance between their generality, scalability, latency, and hardware complexity. In this paper, we propose QEI, a generic, integrated, and efficient acceleration solution for various data structure queries. We first abstract the query operations to a few regular steps and map them to a simple and hardware-friendly configurable finite automaton model. Based on this model, we develop the QEI architecture that allows multiple query operations to execute in parallel to maximize throughput. We also propose a novel way to integrate the accelerator into the CPU that balances performance, latency, and hardware cost. QEI keeps the main control logic near the L2 cache to leverage existing hardware resources in the core while distributing the data-intensive comparison logic to each last-level cache slice for higher parallelism. Our results with five representative data center workloads show that QEI can achieve 6.5× ~11.2× performance improvement in various scenarios with low overhead. Yipeng Wang 0002, Ren Wang 0001, Rangeen Basu Roy Chowdhury, Tsung-Yuan Charlie Tai, Nam Sung Kim |
HPCA | 5 |
| 2021 | MOBO-NFV: Automated Tuning of a Network Function Virtualization System using Multi-Objective Bayesian Optimization
Pietro Mercati, Bin Li 0018, Mesut Ali Ergin, Tsung-Yuan Charlie Tai, Michael Kishinevsky, Boris Serafimov, Subhiksha Ravisundar, Eoin Walsh, Thomas Long |
IM | 4 |
| 2021 | Don't Forget the I/O When Allocating Your LLCabstractIn modern server CPUs, last-level cache (LLC) is a critical hardware resource that exerts significant influence on the performance of the workloads, and how to manage LLC is a key to the performance isolation and QoS in the cloud with multi-tenancy. In this paper, we argue that in addition to CPU cores, high-speed I/O is also important for LLC management. This is because of an Intel architectural innovation – Data Direct I/O (DDIO) – that directly injects the inbound I/O traffic to (part of) the LLC instead of the main memory. We summarize two problems caused by DDIO and show that (1) the default DDIO configuration may not always achieve optimal performance, (2) DDIO can decrease the performance of non-I/O workloads that share LLC with it by as high as 32%.We then present, the first LLC management mechanism that treats the I/O as the first-class citizen. Iat monitors and analyzes the performance of the core/LLC/DDIO using CPU’s hardware performance counters and adaptively adjusts the number of LLC ways for DDIO or the tenants that demand more LLC capacity. In addition, Iat dynamically chooses the tenants that share its LLC resource with DDIO to minimize the performance interference by both the tenants and the I/O. Our experiments with multiple microbenchmarks and real-world applications demonstrate that with minimal overhead, Iat can effectively and stably reduce the performance degradation caused by DDIO. Mohammad Alian, Yipeng Wang 0002, Ren Wang 0001, Ilia Kurakin, Tsung-Yuan Charlie Tai, Nam Sung Kim |
ISCA | 6 |
| 2020 | RLDRM: Closed Loop Dynamic Cache Allocation with Deep Reinforcement Learning for Network Function VirtualizationabstractNetwork function virtualization (NFV) technology attracts tremendous interests from telecommunication industry and data center operators, as it allows service providers to assign resource for Virtual Network Functions (VNFs) on demand, achieving better flexibility, programmability, and scalability. To improve server utilization, one popular practice is to deploy best effort (BE) workloads along with high priority (HP) VNFs when high priority VNF's resource usage is detected to be low. The key challenge of this deployment scheme is to dynamically balance the Service level objective (SLO) and the total cost of ownership (TCO) to optimize the data center efficiency under inherently fluctuating workloads. With the recent advancement in deep reinforcement learning, we conjecture that it has the potential to solve this challenge by adaptively adjusting resource allocation to reach the improved performance and higher server utilization. In this paper, we present a closed-loop automation system RLDRM11RLDRM: Reinforcement Learning Dynamic Resource Management to dynamically adjust Last Level Cache allocation between HP VNFs and BE workloads using deep reinforcement learning. The results demonstrate improved server utilization while maintaining required SLO for the HP VNFs. Bin Li 0018, Yipeng Wang 0002, Ren Wang 0001, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Zhu Zhou, Andrew Herdrich, Ameer Haj-Ali, Ion Stoica, Krste Asanovic |
NetSoft | 4 |
| 2019 | Distributed Software Switching for Scalable Network FunctionsabstractNetwork packets have traditionally been processed by purpose built, fixed-function appliances using proprietary hardware and software designs. Today, these designs are transforming into collections of flexible software blocks running on general purpose computing hardware, with expectations of high performance and scalability. Among these pieces, software packet switching (i.e. vSwitch) plays a fundamental role, as all packets need to traverse the switch on a given platform, and quite often repeatedly. In this paper, we demonstrate that certain performance and scaling deficiencies of packet switching in software can be attributed to the centralized design and deployment of the switching software threads. While multithreading is crucial for scaling-out software performance, improper use of the multi-processor resources can lead to undesired impacts on performance, due to fixed-partitioning, sub-optimal scheduling and complex cache coherency protocol interactions involving non-negligible data-access latencies. In this study, we take advantage of symmetric multi-processor threading (SMT) and lay out principles for better software packet switching with a distributed approach. With proof of concept implementations based on the very popular Open vSwitch project and DPDK, we show distributed software switching can easily scale with the network function workload threads, and achieve more than 2× throughput by exploiting the data locality within distributed nature of the CPU cache-subsystem. Mesut Ali Ergin, James Tsai, Tsung-Yuan Charlie Tai |
NetSoft | 3 |
| 2017 | Optimizing Open vSwitch to Support Millions of FlowsabstractSoftware switch has emerged as a critical component in software defined networking and network virtualization areas. Open vSwitch (OvS) is a widely used software switch which uses tuple space search algorithm for packet classification, and an exact match cache (EMC) for caching most frequently used flows. In this paper, we propose two new optimizations for OvS to further improve its performance and scalability. First aims to completely remove the sequential search overhead of the tuple space search layer of OvS, and second is a dynamic cache insertion optimization for the EMC to improve EMC effectiveness. We show that the optimizations can improve OvS's throughput by up to 3.5X for millions of active flows. Yipeng Wang 0002, Tsung-Yuan Charlie Tai, Ren Wang 0001, Sameh Gobriel, Janet Tseng, James Tsai |
GLOBECOM | 2 |
| 2014 | Joint optimization of DVFS and low-power sleep-state selection for mobile platformsabstractTo provide the ultimate mobile user experience, extended battery life is critical to small form-factor mobile platforms such as smartphones and tablets. Dynamic voltage and frequency scaling (DVFS) and low-power CPU/platform sleep states are commonly used power management features, as they allow dynamic control of power and performance to the time-varying needs of workloads. Despite the potential power saving benefit from synergistic integration of DVFS and sleep-state selection, it is challenging to optimize them jointly for mobile workloads (e.g., video streaming), and most existing work considers them only individually. To address this problem, we study joint optimization of CPU frequency (a.k.a. CPU P-states) and CPU/platform sleep-state selections to reduce energy consumption in mobile platforms. This joint optimization becomes feasible with advanced power management techniques and power aware software development methodologies that regulate (e.g., coalesce/align) system activities, making workload characteristics and system idle duration more deterministic and predictable. We then analyze the optimal operating state that minimizes the expected platform energy consumption based on workload characteristics, and present an algorithm to adapt to it at run time. Our evaluation results on mobile workloads show that the proposed scheme can reduce system power consumption by up to 24%, compared to the conventional CPU-utilization-based approach, which seeks mainly to minimize processor energy. Alexander W. Min, Ren Wang 0001, James Tsai, Tsung-Yuan Charlie Tai |
ICC | 4 |
| 2014 | Energy-Responsive Aggregate Context for Energy Saving in a Multi-Resident EnvironmentabstractHuman activity is among the critical information for a context-aware energy saving system since knowing what activities are undertaken is important for judging if energy is well spent. Most of the prior works on energy saving do not make the best of context-awareness especially in a multiuser environment to assist the energy saving system. In addition, they often ignore whether appliances are operating implicitly or explicitly related to the context. These factors may compromise the practicality and acceptability of most of the currently available energy saving systems, thus failing to meet real user needs. Therefore, we propose Energy-Responsive Aggregate Context (ERAC) to model multi-resident activities and their associated energy consumption. Based on the relationship, implicit or explicit, between a given appliance and its associated context, an energy saving system and its users can better determine whether the power consumed by the appliance is wasted. Our experimental results demonstrate the effectiveness of the proposed approach. Ching-Hu Lu, Chao-Lin Wu, Tsung-Han Yang, Hui-Wen Yeh, Mao-Yung Weng, Li-Chen Fu, Tsung-Yuan Charlie Tai |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2013 | Energy-efficient interconnect via Router ParkingabstractThe increase in on-chip core counts in Chip Multi Processors (CMPs) has led to the adoption of interconnects such as Mesh and Torus, which consume an increasing fraction of the chip power. Moreover, as technology and voltage continue to scale down, static power consumes a larger fraction of the total power; reducing it is increasingly important for energy proportional computing. Currently, processor designers strive to send under-utilized cores into deep sleep states in order to reduce idling power and improve overall energy efficiency. However, even in state-of-the-art CMP designs, when a core goes to sleep the router attached to it remains active in order to continue packet forwarding. In this paper, we propose Router Parking - selectively power-gating routers attached to parked cores. Router Parking ensures that network connectivity is maintained, and limits the average interconnect latency impact of packet detouring around parked routers. We present two Router Parking algorithms - an aggressive approach to park as many routers as possible, and a conservative approach that parks a limited set of routers in order to keep the impact on latency increase minimal. Further, we propose an adaptive policy to choose between the two algorithms at run-time. We evaluate our algorithms using both synthetic traffic as well as real workloads taken from SPEC CPU2006 and PARSEC 2.1 benchmark suites. Our evaluation results show that Router Parking can achieve significant savings in the total interconnect energy (average of 32%, 40% and 41% for the synthetic, SPEC CPU2006, and PARSEC 2.1 workloads, respectively). Ahmad Samih, Ren Wang 0001, Anil Krishna, Christian Maciocco, Tsung-Yuan Charlie Tai, Yan Solihin |
HPCA | 5 |
| 2012 | Evaluating Dynamics and Bottlenecks of Memory Collaboration in Cluster SystemsabstractWith the fast development of highly-integrated distributed systems (cluster systems), designers face interesting memory hierarchy design choices while attempting to avoid the notorious disk swapping. Swapping to the free remote memory through Memory Collaboration has demonstrated its cost-effectiveness compared to over provisioning the cluster for peak load requirements. Recent memory collaboration studies propose several ways on accessing the under-utilized remote memory in static system configurations, without detailed exploration of the dynamic memory collaboration. Dynamic collaboration is an important aspect given the run-time memory usage fluctuations in clustered systems. Further, as the interest in memory collaboration grows, it is crucial to understand the existing performance bottlenecks, overheads, and potential optimization. In this paper we address these two issues. First, we propose an Autonomous Collaborative Memory System (ACMS) that manages memory resources dynamically at run time to optimize performance. We implement a prototype realizing the proposed ACMS, experiment with a wide range of real-world applications, and show up to 3× performance speedup compared to a non-collaborative memory system without perceivable performance impact on nodes that provide memory. Second, we analyze, in depth, the end-to-end memory collaboration overhead and pinpoint the corresponding bottlenecks. Ahmad Samih, Ren Wang 0001, Christian Maciocco, Tsung-Yuan Charlie Tai, Ronghui Duan, Jiangang Duan, Yan Solihin |
CCGRID | 4 |
| 2011 | Reducing Power Consumption for Mobile Platforms via Adaptive Traffic CoalescingabstractBattery life remains to be a critical competitive metric for today's mobile platforms that offer ubiquitous connectivity through their wireless communication interfaces. With most usage models being driven by always-on communication activities, e.g, Internet video streaming, web browsing, etc., it is imperative to understand the impact of network activities on the overall platform power, and optimize power consumption for such activities. As shown by our investigation, various real-world network-driven workloads exhibit bursty and random behavior, which motivates our work on regulating and coalescing incoming packets to reduce platform wake events. To understand the performance impact of packet coalescing, we conduct an extensive investigation to study how coalescing may affect the throughput and user experience. Armed with the deep understandings, we propose, implement and evaluate an Adaptive Traffic Coalescing (ATC) scheme that monitors the incoming traffic at the Network Interface Card (NIC), and adaptively coalesces the packets for a limited duration in the NIC buffer, thus requiring no network or eco-system support. The proposed ATC scheme effectively reduces platform wake events, and enables the platform to enter and stay in the low-power state longer for energy efficiency. We have implemented the scheme in commercial wireless NICs. Using various mobile platforms, we evaluate the power savings and performance impact of the proposed ATC scheme. Experiments show that ATC achieves significant power saving for major platform components, around 20% for real-world Internet workloads, without impacting performance and user experience. Ren Wang 0001, James Tsai, Christian Maciocco, Tsung-Yuan Charlie Tai, Jackie Wu |
IEEE J. Sel. Areas Commun. | 4 |
| 2010 | When Idle Is Not Quiet: Energy-Efficient Platform Design in Presence of Network Background and Management TrafficabstractToday's platforms offer ubiquitous network connectivity through one or more communication interfaces and while the communication devices consume a small portion of the overall platform power the impact of network connectivity and individual packet processing on the overall platform energy consumption is significant, due to the non-deterministic characteristics of network traffic. In this paper we show that significant energy is unnecessarily wasted at the platform level while it is idle and connected to one or more networks, especially unlicensed networks like WiFi or Ethernet, because the system is constantly processing what we term "background/noise" traffic. We quantify the negative impact of network connectivity on platform energy-efficiency and provide novel techniques to mitigate this impact. We implemented these mitigation techniques in our wireless network interface cards and our results indicate that when our scheme is used the total platform sleep time increases by up to 40% with no performance degradation and no user experience impact. Sameh Gobriel, Christian Maciocco, Tsung-Yuan Charlie Tai |
GLOBECOM | 3 |