Andrew Herdrich

dblp:92/7061 · also Andrew J. Herdrich · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 63% Cloud and datacenter computing · 37%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
memory bandwidth
0.512021
LIBRA: Clearing the Cloud Through Dynamic Memory Bandwidth Management · HPCA 2021
Cloud and datacenter computing
resource management
0.512021
LIBRA: Clearing the Cloud Through Dynamic Memory Bandwidth Management · HPCA 2021
Memory systems › cache management
cache allocation
0.212016
Cache QoS: From concept to reality in the Intel® Xeon® processor E5-2600 v3 product family · HPCA 2016
Memory systems
cache management
0.212016
Cache QoS: From concept to reality in the Intel® Xeon® processor E5-2600 v3 product family · HPCA 2016
Memory systems › cache management
cache monitoring
0.212016
Cache QoS: From concept to reality in the Intel® Xeon® processor E5-2600 v3 product family · HPCA 2016

Methods — techniques the papers use, named apart from their topics

hardware throttling · 0.5dynamic resource control · 0.5control policy · 0.5
YearPublicationVenuePosition
2023 RAPID: Enabling fast online policy learning in dynamic public cloud environments
Drew Penney, Bin Li 0018, Lizhong Chen, Jaroslaw J. Sydir, Anna Drewek-Ossowicka, Ramesh Illikkal, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Andrew Herdrich
Neurocomputing9
2021 LIBRA: Clearing the Cloud Through Dynamic Memory Bandwidth Management
abstract
Modern Cloud Service Providers (CSP) heavily co-schedule tasks with different priorities on the same computing node to increase server utilization. To ensure the performance of high priority jobs, CSPs usually employ Quality-of-Service (QoS) mechanisms to manage or regulate the usage of shared hardware resources. Among the critical shared hardware resources, there has been very limited analysis on effective sharing of memory bandwidth among co-scheduled jobs, mainly for two reasons: (1) The correlation between application performance and its memory bandwidth allocation is complicated. (2) An effective hardware throttling mechanism for precise memory bandwidth control is unavailable. These limitations drive CSPs to design conservative policies to ensure the performance of the high priority tasks, which significantly degrades the throughput of batch jobs and reduces the overall benefits of workload co-scheduling. This paper proposes LIBRA, a holistic framework for dynamic memory bandwidth management in production data centers. LIBRA incorporates a novel hardware throttling mechanism, Dynamic Resource Control, to support self-adaptive memory bandwidth regulation. It also employs a lightweight control policy to further enhance the bandwidth scalability for the throttled tasks. Our evaluation results on a cluster demonstrate that LIBRA is capable of increasing the performance of batch jobs by up to 52.8% compared to existing QoS schemes.
Ying Zhang 0016, Xiaowei Jiang, Ian M. Steiner, Andrew Herdrich, Kevin Shu, Ripan Das, Long Cui, Litrin Jiang
HPCA6
2020 RLDRM: Closed Loop Dynamic Cache Allocation with Deep Reinforcement Learning for Network Function Virtualization
abstract
Network function virtualization (NFV) technology attracts tremendous interests from telecommunication industry and data center operators, as it allows service providers to assign resource for Virtual Network Functions (VNFs) on demand, achieving better flexibility, programmability, and scalability. To improve server utilization, one popular practice is to deploy best effort (BE) workloads along with high priority (HP) VNFs when high priority VNF's resource usage is detected to be low. The key challenge of this deployment scheme is to dynamically balance the Service level objective (SLO) and the total cost of ownership (TCO) to optimize the data center efficiency under inherently fluctuating workloads. With the recent advancement in deep reinforcement learning, we conjecture that it has the potential to solve this challenge by adaptively adjusting resource allocation to reach the improved performance and higher server utilization. In this paper, we present a closed-loop automation system RLDRM11RLDRM: Reinforcement Learning Dynamic Resource Management to dynamically adjust Last Level Cache allocation between HP VNFs and BE workloads using deep reinforcement learning. The results demonstrate improved server utilization while maintaining required SLO for the HP VNFs.
Bin Li 0018, Yipeng Wang 0002, Ren Wang 0001, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Zhu Zhou, Andrew Herdrich, Ameer Haj-Ali, Ion Stoica, Krste Asanovic
NetSoft7
2016 CAF: Core to Core Communication Acceleration Framework
abstract
As the number of cores in a multicore system increases, core-to-core (C2C) communication is increasingly limiting the performance scaling of workloads that share data frequently. The traditional way cores communicate is by using shared memory space between them. However, shared memory communication fundamentally involves coherence invalidations and cache misses, which cause large performance overheads and incur a high amount of network traffic. Many important workloads incur significant C2C communication and are affected significantly by the costs, including pipelined packet processing which is widely used in software-based networking solutions. In these workloads, threads run on different cores and pass packets from one core to another for different stages of processing using software queues.
Yipeng Wang 0002, Ren Wang 0001, Andrew Herdrich, James Tsai, Yan Solihin
PACT3
2016 Cache QoS: From concept to reality in the Intel® Xeon® processor E5-2600 v3 product family
abstract
Over the last decade, addressing quality of service (QoS) in multi-core server platforms has been growing research topic. QoS techniques have been proposed to address the shared resource contention between co-running applications or virtual machines in servers and thereby provide better isolation, performance determinism and potentially improve overall throughput. One of the most important shared resources is cache space. Most proposals for addressing shared cache contention are based on simulations and analysis and no commercial platforms were available that integrated such techniques and provided a practical solution. In this paper, we will present the first set of shared cache QoS techniques designed and implemented in state-of-the-art commercial servers (the Intel® Xeon® processor E5-2600 v3 product family). We will describe two key technologies: (i) Cache Monitoring Technology (CMT) to enable monitoring of shared cache usage by different applications and (ii) Cache Allocation Technology (CAT) which enables redistribution of shared cache space between applications to address contention. This is the first paper to describing these techniques as they moved from concept to reality, starting from early research to product implementation. We will also present case studies highlighting the value of these techniques using example scenarios of multi-programmed workloads, virtualized platforms in datacenters and communications platforms. Finally, we will describe initial software infrastructure and enabling for industry practitioners and researchers to take advantage of these technologies for their QoS needs.
Andrew Herdrich, Edwin Verplanke, Priya Autee, Ramesh Illikkal, Chris Gianos, Ronak Singhal, Ravi R. Iyer 0001
HPCA1
2014 QoS management on heterogeneous architecture for parallel applications
abstract
Quality of service (QoS) management is widely employed to provide differentiable performance to programs with distinctive priorities on conventional chip multi-processor (CMP) platforms. Recently, heterogeneous architecture integrating diverse processor cores on the same silicon has been proposed to better serve various application domains and it is expected to be an important design paradigm of future processors. Therefore, the QoS management on emerging heterogeneous systems will be of great significance. On the other hand, parallel applications are becoming increasingly important in modern computing community in order to explore the benefit of thread-level parallelism on CMPs. However, considering the diverse characteristics of thread synchronization, data sharing, and parallelization pattern, governing the execution of multiple parallel programs with different performance requirements becomes a complicated yet significant problem. In this paper, we study QoS management for parallel applications running on heterogeneous CMP systems. We comprehensively assess a series of task-to-core mapping policies on a real heterogeneous hardware (QuickIA) by characterizing their impacts on performance of individual applications. Our evaluation results show that the proposed QoS policies are effective to improve the performance of programs with highest priority while striking good tradeoff with system fairness.
Ying Zhang 0016, Li Zhao 0002, Ramesh Illikkal, Ravi R. Iyer 0001, Andrew Herdrich, Lu Peng 0001
ICCD5
2012 Exploiting Semantics of Virtual Memory to Improve the Efficiency of the On-Chip Memory System
Bin Li 0018, Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Andrew Herdrich, Ravi R. Iyer 0001, Srihari Makineni
Euro-Par6
2009 Rate-based QoS techniques for cache/memory in CMP platforms
abstract
As we embrace the era of chip multi-processors (CMP), we are faced with two major architectural challenges: (i) QoS or performance management of disparate applications running on CPU cores contending for shared cache/memory resources and (ii) global/local power management techniques to stay within the overall platform constraints. The problem is exacerbated as the number of cores sharing the resources in a chip increase. In the past, researchers have proposed independent solutions for these two problems. In this paper, we show that rate-based techniques that are employed to address power management can be adapted to address cache/memory QoS issues. The basic approach is to throttle down the processing rate of a core if it is running a low-priority task and its execution is interfering with the performance of a high priority task due to platform resource contention (i.e. cache or memory contention). We evaluate two rate throttling mechanisms (clock modulation, and frequency scaling) for effectively managing the interference between applications running in a CMP platform and delivering QoS/performance management. We show that clock modulation is much more applicable to cache/memory QoS than frequency scaling and that resource monitoring along with rate control provides effective power-performance management in CMP platforms.
Andrew Herdrich, Ramesh Illikkal, Ravi R. Iyer 0001, Donald Newell, Vineet Chadha, Jaideep Moses
ICS1