Srihari Makineni

dblp:49/7026 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 3 first-authorSoftware engineering, systems software and programming languages · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Memory systems · 66% Electronic design automation · 7% Parallel and multicore computing · 7%
Computer networks
2 papers
Transport protocols and congestion control · 38% Internet architecture and protocols · 38% Internet of things and sensor networks · 23%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache
0.222011
ACCESS: Smart scheduling for asymmetric cache CMPs · HPCA 2011
Characterization of Direct Cache Access on multi-core systems and 10GbE · HPCA 2009
Memory systems
cache management
0.222010
CHOP: Adaptive filter-based DRAM caching for CMP server platforms · HPCA 2010
Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006
Parallel and multicore computing › task scheduling › memory-aware scheduling
cache-aware task scheduling
0.112011
ACCESS: Smart scheduling for asymmetric cache CMPs · HPCA 2011
Electronic design automation › high-level synthesis
scheduling
0.112011
ACCESS: Smart scheduling for asymmetric cache CMPs · HPCA 2011
Memory systems › cache
DRAM cache
0.112010
CHOP: Adaptive filter-based DRAM caching for CMP server platforms · HPCA 2010
Memory systems › memory hierarchy › cache hierarchy
3d stacked cache
0.112009
Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy · HPCA 2009
Memory systems › memory hierarchy
cache hierarchy
0.112009
Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy · HPCA 2009
Memory systems › cache management
direct cache access
0.112009
Characterization of Direct Cache Access on multi-core systems and 10GbE · HPCA 2009
Memory systems › cache design
reconfigurable cache
0.112009
Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy · HPCA 2009
High-performance computing › data transfer
bulk data transfer
0.112007
Hardware Support for Accelerating Data Movement in Server Platform · IEEE Trans. Computers 2007
Cloud and datacenter computing
data movement acceleration
0.112007
Hardware Support for Accelerating Data Movement in Server Platform · IEEE Trans. Computers 2007
Memory systems
cache design
0.112006
Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006
Memory systems › cache management
cache partitioning
0.112006
Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006
Memory systems › cache design
heterogeneous cache
0.112006
Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006
Internet architecture and protocols
packet processing
0.012004
Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor · HPCA 2004
Transport protocols and congestion control › transport protocols
TCP/IP
0.012004
Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor · HPCA 2004
Performance modeling and evaluation
workload characterization
0.012004
Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor · HPCA 2004
Processor architecture and microarchitecture
chip multiprocessor
0.022007
QoS policies and architecture for cache/memory in CMP platforms · SIGMETRICS 2007
Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006
Operating systems
network stack
0.012009
Characterization of Direct Cache Access on multi-core systems and 10GbE · HPCA 2009
Interconnection networks and networks-on-chip › network topology
tree networks
0.012009
Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy · HPCA 2009
Performance modeling and evaluation › simulation › architectural simulation
execution-driven simulation
0.012007
Hardware Support for Accelerating Data Movement in Server Platform · IEEE Trans. Computers 2007
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation
0.012007
QoS policies and architecture for cache/memory in CMP platforms · SIGMETRICS 2007
Cloud and datacenter computing
virtualization
0.012007
QoS policies and architecture for cache/memory in CMP platforms · SIGMETRICS 2007
Performance modeling and evaluation › workload characterization
architectural characterization
0.012004
Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor · HPCA 2004

Methods — techniques the papers use, named apart from their topics

workload characterization · 0.3simulation · 0.2cache performance prediction · 0.1OS scheduler support · 0.1filter cache · 0.1adaptive caching · 0.1page coloring · 0.1benchmarking · 0.1qos-enabled linux · 0.1execution-driven simulation · 0.1performance analysis · 0.0
YearPublicationVenuePosition
2012 Exploiting Semantics of Virtual Memory to Improve the Efficiency of the On-Chip Memory System
Bin Li 0018, Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Andrew Herdrich, Ravi R. Iyer 0001, Srihari Makineni
Euro-Par8
2011 ACCESS: Smart scheduling for asymmetric cache CMPs
abstract
In current Chip-multiprocessors (CMPs), a significant portion of the die is consumed by the last-level cache. Until recently, the balance of cache and core space has been primarily guided by the needs of single applications. However, as multiple applications or virtual machines (VMs) are consolidated on such a platform, researchers have observed that not all VMs or applications require significant amount of cache space. In order to take advantage of this phenomenon, we explore the use of asymmetric last-level caches in a CMP platform. While asymmetric cache CMPs provide the benefit of reduced power and area, it is important to build in hardware/software support to appropriately schedule applications on to cores with suitable cache capacity. In this paper, we address this problem with our ACCESS architecture comprising of: (a) asymmetric caches across a group of cores, (b) hardware support that enables prediction of cache performance on the different sized caches and (c) OS scheduler support to make use of the prediction capability and appropriately schedule applications on to core with suitable cache capacity. Measurements on a working prototype using SPEC2006 benchmarks show that our ACCESS architecture can effectively schedule jobs in an asymmetric cache CMP and provide 23% performance improvement compared to a naive scheduler, and is 97% close to an oracle scheduler in making schedules.
Xiaowei Jiang, Asit K. Mishra, Li Zhao 0002, Ravi R. Iyer 0001, Zhen Fang 0002, Sadagopan Srinivasan, Srihari Makineni, Paul Brett, Chita R. Das
HPCA7
2011 Cost-effectively offering private buffers in SoCs and CMPs
abstract
High performance SoCs and CMPs integrate multiple cores and hardware accelerators such as network interface devices and speech recognition engines. Cores make use of SRAM organized as a cache. Accelerators make use of SRAM as special-purpose storage such as FIFOs, scratchpad memory, or other forms of private buffers. Dedicated private buffers provide benefits such as deterministic access, but are highly area inefficient due to the lower average utilization of the total available storage.
Zhen Fang 0002, Li Zhao 0002, Ravi R. Iyer 0001, Carlos Flores Fajardo, German Fabila Garcia, Bin Li 0018, Steve R. King, Xiaowei Jiang, Srihari Makineni
ICS10
2010 CHOP: Adaptive filter-based DRAM caching for CMP server platforms
abstract
As manycore architectures enable a large number of cores on the die, a key challenge that emerges is the availability of memory bandwidth with conventional DRAM solutions. To address this challenge, integration of large DRAM caches that provide as much as 5× higher bandwidth and as low as 1/3rd of the latency (as compared to conventional DRAM) is very promising. However, organizing and implementing a large DRAM cache is challenging because of two primary tradeoffs: (a) DRAM caches at cache line granularity require too large an on-chip tag area that makes it undesirable and (b) DRAM caches with larger page granularity require too much bandwidth because the miss rate does not reduce enough to overcome the bandwidth increase. In this paper, we propose CHOP (Caching HOt Pages) in DRAM caches to address these challenges. We study several filter-based DRAM caching techniques: (a) a filter cache (CHOP-FC) that profiles pages and determines the hot subset of pages to allocate into the DRAM cache, (b) a memory-based filter cache (CHOP-MFC) that spills and fills filter state to improve the accuracy and reduce the size of the filter cache and (c) an adaptive DRAM caching technique (CHOP-AFC) to determine when the filter cache should be enabled and disabled for DRAM caching. We conduct detailed simulations with server workloads to show that our filter-based DRAM caching techniques achieve the following: (a) on average over 30% performance improvement over previous solutions, (b) several magnitudes lower area overhead in tag space required for cache-line based DRAM caches, (c) significantly lower memory bandwidth consumption as compared to page-granular DRAM caches.
Xiaowei Jiang, Niti Madan, Li Zhao 0002, Mike Upton, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell, Yan Solihin, Rajeev Balasubramonian
HPCA6
2009 Evaluating implications of Virtual Worlds on server architecture using Second Life
abstract
Linden Lab's Second Life is the prominent Virtual World platform in the market today. Virtual Worlds like Second Life are emerging to be a main stream server workload because of their popularity due to richness of 3D content and immersive social experience they can provide. So, it is very important for computer architects to fully understand this workload and its requirements. In this paper, our goal is to fully analyze the performance and to characterize the processing of Second Life server Simulator process. The simulator process has three key critical functions that dominate the performance characteristics of this workload. These are: 1) Physics engine that is responsible for simulating real world behaviors taking into account mass of the objects, gravity, wind force, etc., 2) Scripting engine that is responsible for executing scripts attached to the objects. Scripts is the main way of manipulating object behaviors (motion, color, etc.) in-world on the server, and 3) Simulator logic that is responsible for simulating the world which includes avatar movement, calculating visible areas and communicating with the clients. Our work includes performance scaling experiments, comparison of performance on Intel's Clovertown and Nehalem processor based server systems and collecting and analyzing architectural characterization data for this workload. Our measurements have shown that Intel's latest Xeon servers using Nehalem processors offer 20 to 50% performance improvement over previous generation processor based system, and that the physics computation is more compute and memory intensive. To get a better perspective of Second Life's requirements, we have compared this workload with three other popular commercial server workloads (TPC-E, SPECjAppServer and SPECjbb) and found out that this workload executes 2 to 10 times more floating point, multiply and divide instructions.
Srihari Makineni, Omesh Tickoo, Aaron Terrell, Jessica Young, Donald Newell
HiPC1
2009 Characterization of Direct Cache Access on multi-core systems and 10GbE
abstract
10 GbE connectivity is expected to be a standard feature of server platforms in the near future. Among the numerous methods and features proposed to improve network performance of such platforms is direct cache access (DCA) to route incoming I/O to CPU caches directly. While this feature has been shown to be promising, there can be significant challenges when dealing with high rates of traffic in a multiprocessor and multi-core environment. In this paper, we focus on two practical considerations with DCA. In the first case, we show that the performance benefit from DCA can be limited when network traffic processing rate cannot match the I/O rate. In the second case, we show that affinitizing both stack and application contexts to cores that share a cache is critical. With proper distribution and affinity, we show that a standard Linux network stack runs 32% faster for 2 KB to 64 KB I/O sizes.
Amit Kumar 0008, Ram Huggahalli, Srihari Makineni
HPCA3
2009 Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy
abstract
Cache hierarchies in future many-core processors are expected to grow in size and contribute a large fraction of overall processor power and performance. In this paper, we postulate a 3D chip design that stacks SRAM and DRAM upon processing cores and employs OS-based page coloring to minimize horizontal communication of cache data. We then propose a heterogeneous reconfigurable cache design that takes advantage of the high density of DRAM and the superior power/delay characteristics of SRAM to efficiently meet the working set demands of each individual core. Finally, we analyze the communication patterns for such a processor and show that a tree topology is an ideal fit that significantly reduces the power and latency requirements of the on-chip network. The above proposals are synergistic: each proposal is made more compelling because of its combination with the other innovations described in this paper. The proposed reconfigurable cache model improves performance by up to 19% along with 48% savings in network power.
Niti Madan, Li Zhao 0002, Naveen Muralimanohar, Aniruddha N. Udipi, Rajeev Balasubramonian, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell
HPCA7
2009 CMPSched$im: Evaluating OS/CMP interaction on shared cache management
abstract
CMPs have now become mainstream and are growing in complexity with more cores, several shared resources (cache, memory, etc) and the potential for additional heterogeneous elements. In order to manage these resources, it is becoming critical to optimize the interaction between the execution environment (operating systems, virtual machine monitors, etc) and the CMP platform. Performance analysis of such OS and CMP interactions is challenging because it requires long running full-system execution-driven simulations. In this paper, we explore an alternative approach (CMPSched$im) to evaluate the interaction of OS and CMP architectures. In particular, CMPSched$im is focused on evaluating techniques to address the shared cache management problem through better interaction between CMP hardware and operating system scheduling. CMPSched$im enables fast and flexible exploration of this interaction by combining the benefits of (a) binary instrumentation tools (Pin), (b) user-level scheduling tools (Linsched) and (c) simple core/cache simulators. In this paper, we describe CMPSched$im in detail and present case studies showing how CMPSched$im can be used to optimize OS scheduling by taking advantage of novel shared cache monitoring capabilities in the hardware. We also describe OS scheduling heuristics to improve overall system performance through resource monitoring and application classification to achieve near optimal scheduling that minimizes the effects of contention in the shared cache of a CMP platform.
Jaideep Moses, Konstantinos Aisopos, Aamer Jaleel, Ravi R. Iyer 0001, Ramesh Illikkal, Donald Newell, Srihari Makineni
ISPASS7
2008 To Snoop or Not to Snoop: Evaluation of Fine-Grain and Coarse-Grain Snoop Filtering Techniques
Jessica Young, Srihari Makineni, Ravi R. Iyer 0001, Donald Newell, Adrian Moga
Euro-Par2
2008 Achieving 10Gbps Network Processing: Are We There Yet?
Priya Govindarajan, Srihari Makineni, Donald Newell, Ravi R. Iyer 0001, Ram Huggahalli, Amit Kumar 0008
HiPC2
2008 Re-examining cache replacement policies
abstract
The replacement policies commonly used in modern processors perform an average of 57% worse than an optimal replacement policy for commercial applications using large, shared caches in a chip-multiprocessor (CMP). Recent proposals that improve the performance of smaller, uniprocessor caches with SPEC CPU workloads do not achieve similar benefits with commercial workloads and larger caches, even though these caches still perform worse than optimal. The recently proposed Shepherd Cache replacement policy reduces miss-ratios by 7.3% on average, but it relies on an impractical LRU policy and requires 5.3% overhead relative to the total cache capacity. We propose two new, practical, low-overhead replacement policies that mimic shepherd cache with significantly less meta-data overhead. First, we propose a Lightweight shepherd cache design that reduces miss-ratios by 8% on average and up to 19%, while requiring only 1.9% meta-data overhead. We also propose an extra-lightweight shepherd cache design that reduces overhead to only 0.5% when combined with a practical clock replacement policy while reducing miss-ratios by an average of 5.4% and up to 14%.
Jason Zebchuk, Srihari Makineni, Donald Newell
ICCD2
2007 CacheScouts: Fine-Grain Monitoring of Shared Caches in CMP Platforms
Li Zhao 0002, Ravi R. Iyer 0001, Ramesh Illikkal, Jaideep Moses, Srihari Makineni, Donald Newell
PACT5
2007 Constraint-Aware Large-Scale CMP Cache Design
Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Ramesh Illikkal, Jaideep Moses, Donald Newell
HiPC3
2007 QoS policies and architecture for cache/memory in CMP platforms
abstract
As we enter the era of CMP platforms with multiple threads/cores on the die, the diversity of the simultaneous workloads running on them is expected to increase. The rapid deployment of virtualization as a means to consolidate workloads on to a single platform is a prime example of this trend. In such scenarios, the quality of service (QoS) that each individual workload gets from the platform can widely vary depending on the behavior of the simultaneously running workloads. While the number of cores assigned to each workload can be controlled, there is no hardware or software support in today's platforms to control allocation of platform resources such as cache space and memory bandwidth to individual workloads. In this paper, we propose a QoS-enabled memory architecture for CMP platforms that addresses this problem. The QoS-enabled memory architecture enables more cache resources (i.e. space) and memory resources (i.e. bandwidth) for high priority applications based on guidance from the operating environment. The architecture also allows dynamic resource reassignment during run-time to further optimize the performance of the high priority application with minimal degradation to low priority. To achieve these goals, we will describe the hardware/software support required in the platform as well as the operating environment (O/S and virtual machine monitor). Our evaluation framework consists of detailed platform simulation models and a QoS-enabled version of Linux. Based on evaluation experiments, we show the effectiveness of a QoS-enabled architecture and summarize key findings/trade-offs.
Ravi R. Iyer 0001, Li Zhao 0002, Ramesh Illikkal, Srihari Makineni, Donald Newell, Yan Solihin, Lisa R. Hsu, Steven K. Reinhardt
SIGMETRICS5
2007 Hardware Support for Accelerating Data Movement in Server Platform
abstract
Data movement (memory copies) is a very common operation during network processing and application execution on servers. The performance of this operation is rather poor on today's microprocessors due to the following aspects: 1) Several long-latency memory accesses are involved because the source and/or the destination are typically in memory, 2) latency hiding techniques, such as out-of-order execution, hardware threading, and prefetching, are not very effective for bulk data movement, and 3) microprocessors move data at register (small) granularity. In this paper, we show this overhead of bulk data movement and propose the use of dedicated copy engines to minimize it. We present a detailed analysis of copy engine architectures along two dimensions: 1) on-die versus off-die and 2) synchronous versus asynchronous. These copy engine architectures are superior to traditional direct memory access (DMA) engines because they are tightly coupled to the core architecture and enable lower overhead communication and signaling. We describe the hardware support required to implement these copy engines and integrate them into server platforms. We perform a detailed case study to evaluate the performance of these copy engines. The evaluation is based on an execution-driven simulator, which was extended with detailed models of copy engines. Our simulation results show that copy engines are effective in reducing the bulk data movement overhead and, hence, hold significant promise for high-performance server platforms
Li Zhao 0002, Laxmi N. Bhuyan, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell
IEEE Trans. Computers4
2006 Communist, utilitarian, and capitalist cache policies on CMPs: caches as a shared resource
abstract
As chip multiprocessors (CMPs) become increasingly mainstream, architects have likewise become more interested in how best to share a cache hierarchy among multiple simultaneous threads of execution. The complexity of this problem is exacerbated as the number of simultaneous threads grows from two or four to the tens or hundreds. However, there is no consensus in the architectural community on what "best" means in this context. Some papers in the literature seek to equalize each thread's performance loss due to sharing, while others emphasize maximizing overall system performance. Furthermore, the specific effect of these goals varies depending on the metric used to define "performance".In this paper we label equal performance targets as Communist cache policies and overall performance targets as Utilitarian cache policies. We compare both of these models to the most common current model of a free-for-all cache (a Capitalist policy). We consider various performance metrics, including miss rates, bandwidth usage, and IPC, including both absolute and relative values of each metric. Using analytical models and behavioral cache simulation, we find that the optimal partitioning of a shared cache can vary greatly as different but reasonable definitions of optimality are applied. We also find that, although Communist and Utilitarian targets are generally compatible, each policy has workloads for which it provides poor overall performance or poor fairness, respectively. Finally, we find that simple policies like LRU replacement and static uniform partitioning are not sufficient to provide near-optimal performance under any reasonable definition, indicating that some thread-aware cache resource allocation mechanism is required.
Lisa R. Hsu, Steven K. Reinhardt, Ravi R. Iyer 0001, Srihari Makineni
PACT4
2006 Receive Side Coalescing for Accelerating TCP/IP Processing
Srihari Makineni, Ravi R. Iyer 0001, Partha Sarangam, Donald Newell, Li Zhao 0002, Ramesh Illikkal, Jaideep Moses
HiPC1
2006 Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions
abstract
CMPs enable simultaneous execution of multiple applications on the same platforms that share cache resources. Diversity in the cache access patterns of these simultaneously executing applications can potentially trigger inter-application interference, leading to cache pollution. Whereas a large cache can ameliorate this problem, the issues of larger power consumption with increasing cache size, amplified at sub-100nm technologies, makes this solution prohibitive. In this paper, in order to address the issues relating to power-aware performance of caches, we propose a caching structure that addresses the following: 1) Definition of application-specific cache partitions as an aggregation of caching units (molecules). The parameters of each molecule namely size, associativity and line size are chosen so that the power consumed by it and access time are optimal for the given technology. 2) Application-specific resizing of cache partitions with variable and adaptive associativity per cache line, way size and variable line size. 3) A replacement policy that is transparent to the partition in terms of size, heterogeneity in associativity and line size. Through simulation studies we establish the superiority of molecular cache (caches built as aggregations of molecules) that offers a 29% power advantage over that of an equivalently performing traditional cache
Keshavan Varadarajan, S. K. Nandy 0001, Vishal Sharda, Bharadwaj S. Amrutur, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell
MICRO6
2005 Hardware Support for Bulk Data Movement in Server Platforms
abstract
Bulk data movement occurs commonly in server work-loads and their performance is rather poor on today's microprocessors. We propose the use of small dedicated copy engines, and present a detailed analysis of a bulk data copy engine architecture. We describe the hardware support required to implement the copy engine and to tightly integrate it into server platforms. Our evaluation is based on an execution driven simulator that was extended with detailed models of bulk data movement engines. The simulation results show that dedicated engines are quite effective in eliminating the data movement overhead and are an attractive choice for handling bulk data in future high performance server platforms.
Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Laxmi N. Bhuyan, Donald Newell
ICCD3
2005 Anatomy and Performance of SSL Processing
abstract
A wide spectrum of e-commerce (B2B/B2C), banking, financial trading and other business applications require the exchange of data to be highly secure. The Secure Sockets Layer (SSL) protocol provides the essential ingredients of secure communications - privacy, integrity and authentication. Though it is well-understood that security always comes at the cost of performance, these costs depend on the cryptographic algorithms. In this paper, we present a detailed description of the anatomy of a secure session. We analyze the time spent on the various cryptographic operations (symmetric, asymmetric and hashing) during the session negotiation and data transfer. We then analyze the most frequently used cryptographic algorithms (RSA, AES, DES, 3DES, RC4, MD5 and SHA-1). We determine the key components of these algorithms (setting up key schedules, encryption rounds, substitutions, permutations, etc) and determine where most of the time is spent. We also provide an architectural analysis of these algorithms, show the frequently executed instructions and discuss the ISA/hardware support that may be beneficial to improving SSL performance. We believe that the performance data presented in this paper is useful to performance analysts and processor architects to help accelerate SSL performance in future processors
Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Laxmi N. Bhuyan
ISPASS3
2004 Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor
abstract
A majority of the current and next generation server applications (Web services, e-commerce, storage, etc.) employ TCP/IP as the communication protocol of choice. As a result, the performance of these applications is heavily dependent on the efficient TCP/IP packet processing within the termination nodes. This dependency becomes even greater as the bandwidth needs of these applications grow from 100 Mbps to 1 Gbps to 10 Gbps in the near future. Motivated by this, we focus on the following: (a) to understand the performance behavior of the various modes of TCP/IP processing, (b) to analyze the underlying architectural characteristics of TCP/IP packet processing and (c) to quantify the computational requirements of the TCP/IP packet processing component within realistic workloads. We achieve these goals by performing an in-depth analysis of packet processing performance on Intel's state-of-the-art low power Pentium/spl reg/ M microprocessor running the Microsoft Windows* Server 2003 operating system. Some of our key observations are - (i) that the mode of TCP/IP operation can significantly affect the performance requirements, (ii) that transmit-side processing is largely compute-intensive as compared to receive-side processing which is more memory-bound and (iii) that the computational requirements for sending/receiving packets can form a substantial component (28% to 40%) of commercial server workloads. From our analysis, we also discuss architectural as well as stack-related improvements that can help achieve higher server network throughput and result in improved application performance.
Srihari Makineni, Ravi R. Iyer 0001
HPCA1