EDBT 2026 Demo / reviewers in the wild / expert
Srihari Makineni
dblp:49/7026
· DBLP profile ↗
21ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 3 first-authorSoftware engineering, systems software and programming languages · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Memory systems · 66% Electronic design automation · 7% Parallel and multicore computing · 7% | |
| Computer networks
2 papers |
Transport protocols and congestion control · 38% Internet architecture and protocols · 38% Internet of things and sensor networks · 23% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache |
0.2 | 2 | 2011 | ACCESS: Smart scheduling for asymmetric cache CMPs · HPCA 2011 Characterization of Direct Cache Access on multi-core systems and 10GbE · HPCA 2009 |
Memory systems
cache management |
0.2 | 2 | 2010 | CHOP: Adaptive filter-based DRAM caching for CMP server platforms · HPCA 2010 Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006 |
Parallel and multicore computing › task scheduling › memory-aware scheduling
cache-aware task scheduling |
0.1 | 1 | 2011 | ACCESS: Smart scheduling for asymmetric cache CMPs · HPCA 2011 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2011 | ACCESS: Smart scheduling for asymmetric cache CMPs · HPCA 2011 |
Memory systems › cache
DRAM cache |
0.1 | 1 | 2010 | CHOP: Adaptive filter-based DRAM caching for CMP server platforms · HPCA 2010 |
Memory systems › memory hierarchy › cache hierarchy
3d stacked cache |
0.1 | 1 | 2009 | Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy · HPCA 2009 |
Memory systems › memory hierarchy
cache hierarchy |
0.1 | 1 | 2009 | Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy · HPCA 2009 |
Memory systems › cache management
direct cache access |
0.1 | 1 | 2009 | Characterization of Direct Cache Access on multi-core systems and 10GbE · HPCA 2009 |
Memory systems › cache design
reconfigurable cache |
0.1 | 1 | 2009 | Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy · HPCA 2009 |
High-performance computing › data transfer
bulk data transfer |
0.1 | 1 | 2007 | Hardware Support for Accelerating Data Movement in Server Platform · IEEE Trans. Computers 2007 |
Cloud and datacenter computing
data movement acceleration |
0.1 | 1 | 2007 | Hardware Support for Accelerating Data Movement in Server Platform · IEEE Trans. Computers 2007 |
Memory systems
cache design |
0.1 | 1 | 2006 | Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006 |
Memory systems › cache management
cache partitioning |
0.1 | 1 | 2006 | Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006 |
Memory systems › cache design
heterogeneous cache |
0.1 | 1 | 2006 | Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006 |
Internet architecture and protocols
packet processing |
0.0 | 1 | 2004 | Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor · HPCA 2004 |
Transport protocols and congestion control › transport protocols
TCP/IP |
0.0 | 1 | 2004 | Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor · HPCA 2004 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2004 | Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor · HPCA 2004 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 2 | 2007 | QoS policies and architecture for cache/memory in CMP platforms · SIGMETRICS 2007 Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regions · MICRO 2006 |
Operating systems
network stack |
0.0 | 1 | 2009 | Characterization of Direct Cache Access on multi-core systems and 10GbE · HPCA 2009 |
Interconnection networks and networks-on-chip › network topology
tree networks |
0.0 | 1 | 2009 | Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchy · HPCA 2009 |
Performance modeling and evaluation › simulation › architectural simulation
execution-driven simulation |
0.0 | 1 | 2007 | Hardware Support for Accelerating Data Movement in Server Platform · IEEE Trans. Computers 2007 |
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation |
0.0 | 1 | 2007 | QoS policies and architecture for cache/memory in CMP platforms · SIGMETRICS 2007 |
Cloud and datacenter computing
virtualization |
0.0 | 1 | 2007 | QoS policies and architecture for cache/memory in CMP platforms · SIGMETRICS 2007 |
Performance modeling and evaluation › workload characterization
architectural characterization |
0.0 | 1 | 2004 | Architectural Characterization of TCP/IP Packet Processing on the Pentium M Microprocessor · HPCA 2004 |
Methods — techniques the papers use, named apart from their topics
workload characterization · 0.3simulation · 0.2cache performance prediction · 0.1OS scheduler support · 0.1filter cache · 0.1adaptive caching · 0.1page coloring · 0.1benchmarking · 0.1qos-enabled linux · 0.1execution-driven simulation · 0.1performance analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Exploiting Semantics of Virtual Memory to Improve the Efficiency of the On-Chip Memory System
Bin Li 0018, Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Andrew Herdrich, Ravi R. Iyer 0001, Srihari Makineni |
Euro-Par | 8 |
| 2011 | ACCESS: Smart scheduling for asymmetric cache CMPsabstractIn current Chip-multiprocessors (CMPs), a significant portion of the die is consumed by the last-level cache. Until recently, the balance of cache and core space has been primarily guided by the needs of single applications. However, as multiple applications or virtual machines (VMs) are consolidated on such a platform, researchers have observed that not all VMs or applications require significant amount of cache space. In order to take advantage of this phenomenon, we explore the use of asymmetric last-level caches in a CMP platform. While asymmetric cache CMPs provide the benefit of reduced power and area, it is important to build in hardware/software support to appropriately schedule applications on to cores with suitable cache capacity. In this paper, we address this problem with our ACCESS architecture comprising of: (a) asymmetric caches across a group of cores, (b) hardware support that enables prediction of cache performance on the different sized caches and (c) OS scheduler support to make use of the prediction capability and appropriately schedule applications on to core with suitable cache capacity. Measurements on a working prototype using SPEC2006 benchmarks show that our ACCESS architecture can effectively schedule jobs in an asymmetric cache CMP and provide 23% performance improvement compared to a naive scheduler, and is 97% close to an oracle scheduler in making schedules. Xiaowei Jiang, Asit K. Mishra, Li Zhao 0002, Ravi R. Iyer 0001, Zhen Fang 0002, Sadagopan Srinivasan, Srihari Makineni, Paul Brett, Chita R. Das |
HPCA | 7 |
| 2011 | Cost-effectively offering private buffers in SoCs and CMPsabstractHigh performance SoCs and CMPs integrate multiple cores and hardware accelerators such as network interface devices and speech recognition engines. Cores make use of SRAM organized as a cache. Accelerators make use of SRAM as special-purpose storage such as FIFOs, scratchpad memory, or other forms of private buffers. Dedicated private buffers provide benefits such as deterministic access, but are highly area inefficient due to the lower average utilization of the total available storage. Zhen Fang 0002, Li Zhao 0002, Ravi R. Iyer 0001, Carlos Flores Fajardo, German Fabila Garcia, Bin Li 0018, Steve R. King, Xiaowei Jiang, Srihari Makineni |
ICS | 10 |
| 2010 | CHOP: Adaptive filter-based DRAM caching for CMP server platformsabstractAs manycore architectures enable a large number of cores on the die, a key challenge that emerges is the availability of memory bandwidth with conventional DRAM solutions. To address this challenge, integration of large DRAM caches that provide as much as 5× higher bandwidth and as low as 1/3rd of the latency (as compared to conventional DRAM) is very promising. However, organizing and implementing a large DRAM cache is challenging because of two primary tradeoffs: (a) DRAM caches at cache line granularity require too large an on-chip tag area that makes it undesirable and (b) DRAM caches with larger page granularity require too much bandwidth because the miss rate does not reduce enough to overcome the bandwidth increase. In this paper, we propose CHOP (Caching HOt Pages) in DRAM caches to address these challenges. We study several filter-based DRAM caching techniques: (a) a filter cache (CHOP-FC) that profiles pages and determines the hot subset of pages to allocate into the DRAM cache, (b) a memory-based filter cache (CHOP-MFC) that spills and fills filter state to improve the accuracy and reduce the size of the filter cache and (c) an adaptive DRAM caching technique (CHOP-AFC) to determine when the filter cache should be enabled and disabled for DRAM caching. We conduct detailed simulations with server workloads to show that our filter-based DRAM caching techniques achieve the following: (a) on average over 30% performance improvement over previous solutions, (b) several magnitudes lower area overhead in tag space required for cache-line based DRAM caches, (c) significantly lower memory bandwidth consumption as compared to page-granular DRAM caches. Xiaowei Jiang, Niti Madan, Li Zhao 0002, Mike Upton, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell, Yan Solihin, Rajeev Balasubramonian |
HPCA | 6 |
| 2009 | Evaluating implications of Virtual Worlds on server architecture using Second LifeabstractLinden Lab's Second Life is the prominent Virtual World platform in the market today. Virtual Worlds like Second Life are emerging to be a main stream server workload because of their popularity due to richness of 3D content and immersive social experience they can provide. So, it is very important for computer architects to fully understand this workload and its requirements. In this paper, our goal is to fully analyze the performance and to characterize the processing of Second Life server Simulator process. The simulator process has three key critical functions that dominate the performance characteristics of this workload. These are: 1) Physics engine that is responsible for simulating real world behaviors taking into account mass of the objects, gravity, wind force, etc., 2) Scripting engine that is responsible for executing scripts attached to the objects. Scripts is the main way of manipulating object behaviors (motion, color, etc.) in-world on the server, and 3) Simulator logic that is responsible for simulating the world which includes avatar movement, calculating visible areas and communicating with the clients. Our work includes performance scaling experiments, comparison of performance on Intel's Clovertown and Nehalem processor based server systems and collecting and analyzing architectural characterization data for this workload. Our measurements have shown that Intel's latest Xeon servers using Nehalem processors offer 20 to 50% performance improvement over previous generation processor based system, and that the physics computation is more compute and memory intensive. To get a better perspective of Second Life's requirements, we have compared this workload with three other popular commercial server workloads (TPC-E, SPECjAppServer and SPECjbb) and found out that this workload executes 2 to 10 times more floating point, multiply and divide instructions. Srihari Makineni, Omesh Tickoo, Aaron Terrell, Jessica Young, Donald Newell |
HiPC | 1 |
| 2009 | Characterization of Direct Cache Access on multi-core systems and 10GbEabstract10 GbE connectivity is expected to be a standard feature of server platforms in the near future. Among the numerous methods and features proposed to improve network performance of such platforms is direct cache access (DCA) to route incoming I/O to CPU caches directly. While this feature has been shown to be promising, there can be significant challenges when dealing with high rates of traffic in a multiprocessor and multi-core environment. In this paper, we focus on two practical considerations with DCA. In the first case, we show that the performance benefit from DCA can be limited when network traffic processing rate cannot match the I/O rate. In the second case, we show that affinitizing both stack and application contexts to cores that share a cache is critical. With proper distribution and affinity, we show that a standard Linux network stack runs 32% faster for 2 KB to 64 KB I/O sizes. Amit Kumar 0008, Ram Huggahalli, Srihari Makineni |
HPCA | 3 |
| 2009 | Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchyabstractCache hierarchies in future many-core processors are expected to grow in size and contribute a large fraction of overall processor power and performance. In this paper, we postulate a 3D chip design that stacks SRAM and DRAM upon processing cores and employs OS-based page coloring to minimize horizontal communication of cache data. We then propose a heterogeneous reconfigurable cache design that takes advantage of the high density of DRAM and the superior power/delay characteristics of SRAM to efficiently meet the working set demands of each individual core. Finally, we analyze the communication patterns for such a processor and show that a tree topology is an ideal fit that significantly reduces the power and latency requirements of the on-chip network. The above proposals are synergistic: each proposal is made more compelling because of its combination with the other innovations described in this paper. The proposed reconfigurable cache model improves performance by up to 19% along with 48% savings in network power. Niti Madan, Li Zhao 0002, Naveen Muralimanohar, Aniruddha N. Udipi, Rajeev Balasubramonian, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell |
HPCA | 7 |
| 2009 | CMPSched$im: Evaluating OS/CMP interaction on shared cache managementabstractCMPs have now become mainstream and are growing in complexity with more cores, several shared resources (cache, memory, etc) and the potential for additional heterogeneous elements. In order to manage these resources, it is becoming critical to optimize the interaction between the execution environment (operating systems, virtual machine monitors, etc) and the CMP platform. Performance analysis of such OS and CMP interactions is challenging because it requires long running full-system execution-driven simulations. In this paper, we explore an alternative approach (CMPSched$im) to evaluate the interaction of OS and CMP architectures. In particular, CMPSched$im is focused on evaluating techniques to address the shared cache management problem through better interaction between CMP hardware and operating system scheduling. CMPSched$im enables fast and flexible exploration of this interaction by combining the benefits of (a) binary instrumentation tools (Pin), (b) user-level scheduling tools (Linsched) and (c) simple core/cache simulators. In this paper, we describe CMPSched$im in detail and present case studies showing how CMPSched$im can be used to optimize OS scheduling by taking advantage of novel shared cache monitoring capabilities in the hardware. We also describe OS scheduling heuristics to improve overall system performance through resource monitoring and application classification to achieve near optimal scheduling that minimizes the effects of contention in the shared cache of a CMP platform. Jaideep Moses, Konstantinos Aisopos, Aamer Jaleel, Ravi R. Iyer 0001, Ramesh Illikkal, Donald Newell, Srihari Makineni |
ISPASS | 7 |
| 2008 | To Snoop or Not to Snoop: Evaluation of Fine-Grain and Coarse-Grain Snoop Filtering Techniques
Jessica Young, Srihari Makineni, Ravi R. Iyer 0001, Donald Newell, Adrian Moga |
Euro-Par | 2 |
| 2008 | Achieving 10Gbps Network Processing: Are We There Yet?
Priya Govindarajan, Srihari Makineni, Donald Newell, Ravi R. Iyer 0001, Ram Huggahalli, Amit Kumar 0008 |
HiPC | 2 |
| 2008 | Re-examining cache replacement policiesabstractThe replacement policies commonly used in modern processors perform an average of 57% worse than an optimal replacement policy for commercial applications using large, shared caches in a chip-multiprocessor (CMP). Recent proposals that improve the performance of smaller, uniprocessor caches with SPEC CPU workloads do not achieve similar benefits with commercial workloads and larger caches, even though these caches still perform worse than optimal. The recently proposed Shepherd Cache replacement policy reduces miss-ratios by 7.3% on average, but it relies on an impractical LRU policy and requires 5.3% overhead relative to the total cache capacity. We propose two new, practical, low-overhead replacement policies that mimic shepherd cache with significantly less meta-data overhead. First, we propose a Lightweight shepherd cache design that reduces miss-ratios by 8% on average and up to 19%, while requiring only 1.9% meta-data overhead. We also propose an extra-lightweight shepherd cache design that reduces overhead to only 0.5% when combined with a practical clock replacement policy while reducing miss-ratios by an average of 5.4% and up to 14%. Jason Zebchuk, Srihari Makineni, Donald Newell |
ICCD | 2 |
| 2007 | CacheScouts: Fine-Grain Monitoring of Shared Caches in CMP Platforms
Li Zhao 0002, Ravi R. Iyer 0001, Ramesh Illikkal, Jaideep Moses, Srihari Makineni, Donald Newell |
PACT | 5 |
| 2007 | Constraint-Aware Large-Scale CMP Cache Design
Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Ramesh Illikkal, Jaideep Moses, Donald Newell |
HiPC | 3 |
| 2007 | QoS policies and architecture for cache/memory in CMP platformsabstractAs we enter the era of CMP platforms with multiple threads/cores on the die, the diversity of the simultaneous workloads running on them is expected to increase. The rapid deployment of virtualization as a means to consolidate workloads on to a single platform is a prime example of this trend. In such scenarios, the quality of service (QoS) that each individual workload gets from the platform can widely vary depending on the behavior of the simultaneously running workloads. While the number of cores assigned to each workload can be controlled, there is no hardware or software support in today's platforms to control allocation of platform resources such as cache space and memory bandwidth to individual workloads. In this paper, we propose a QoS-enabled memory architecture for CMP platforms that addresses this problem. The QoS-enabled memory architecture enables more cache resources (i.e. space) and memory resources (i.e. bandwidth) for high priority applications based on guidance from the operating environment. The architecture also allows dynamic resource reassignment during run-time to further optimize the performance of the high priority application with minimal degradation to low priority. To achieve these goals, we will describe the hardware/software support required in the platform as well as the operating environment (O/S and virtual machine monitor). Our evaluation framework consists of detailed platform simulation models and a QoS-enabled version of Linux. Based on evaluation experiments, we show the effectiveness of a QoS-enabled architecture and summarize key findings/trade-offs. Ravi R. Iyer 0001, Li Zhao 0002, Ramesh Illikkal, Srihari Makineni, Donald Newell, Yan Solihin, Lisa R. Hsu, Steven K. Reinhardt |
SIGMETRICS | 5 |
| 2007 | Hardware Support for Accelerating Data Movement in Server PlatformabstractData movement (memory copies) is a very common operation during network processing and application execution on servers. The performance of this operation is rather poor on today's microprocessors due to the following aspects: 1) Several long-latency memory accesses are involved because the source and/or the destination are typically in memory, 2) latency hiding techniques, such as out-of-order execution, hardware threading, and prefetching, are not very effective for bulk data movement, and 3) microprocessors move data at register (small) granularity. In this paper, we show this overhead of bulk data movement and propose the use of dedicated copy engines to minimize it. We present a detailed analysis of copy engine architectures along two dimensions: 1) on-die versus off-die and 2) synchronous versus asynchronous. These copy engine architectures are superior to traditional direct memory access (DMA) engines because they are tightly coupled to the core architecture and enable lower overhead communication and signaling. We describe the hardware support required to implement these copy engines and integrate them into server platforms. We perform a detailed case study to evaluate the performance of these copy engines. The evaluation is based on an execution-driven simulator, which was extended with detailed models of copy engines. Our simulation results show that copy engines are effective in reducing the bulk data movement overhead and, hence, hold significant promise for high-performance server platforms Li Zhao 0002, Laxmi N. Bhuyan, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell |
IEEE Trans. Computers | 4 |
| 2006 | Communist, utilitarian, and capitalist cache policies on CMPs: caches as a shared resourceabstractAs chip multiprocessors (CMPs) become increasingly mainstream, architects have likewise become more interested in how best to share a cache hierarchy among multiple simultaneous threads of execution. The complexity of this problem is exacerbated as the number of simultaneous threads grows from two or four to the tens or hundreds. However, there is no consensus in the architectural community on what "best" means in this context. Some papers in the literature seek to equalize each thread's performance loss due to sharing, while others emphasize maximizing overall system performance. Furthermore, the specific effect of these goals varies depending on the metric used to define "performance".In this paper we label equal performance targets as Communist cache policies and overall performance targets as Utilitarian cache policies. We compare both of these models to the most common current model of a free-for-all cache (a Capitalist policy). We consider various performance metrics, including miss rates, bandwidth usage, and IPC, including both absolute and relative values of each metric. Using analytical models and behavioral cache simulation, we find that the optimal partitioning of a shared cache can vary greatly as different but reasonable definitions of optimality are applied. We also find that, although Communist and Utilitarian targets are generally compatible, each policy has workloads for which it provides poor overall performance or poor fairness, respectively. Finally, we find that simple policies like LRU replacement and static uniform partitioning are not sufficient to provide near-optimal performance under any reasonable definition, indicating that some thread-aware cache resource allocation mechanism is required. Lisa R. Hsu, Steven K. Reinhardt, Ravi R. Iyer 0001, Srihari Makineni |
PACT | 4 |
| 2006 | Receive Side Coalescing for Accelerating TCP/IP Processing
Srihari Makineni, Ravi R. Iyer 0001, Partha Sarangam, Donald Newell, Li Zhao 0002, Ramesh Illikkal, Jaideep Moses |
HiPC | 1 |
| 2006 | Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regionsabstractCMPs enable simultaneous execution of multiple applications on the same platforms that share cache resources. Diversity in the cache access patterns of these simultaneously executing applications can potentially trigger inter-application interference, leading to cache pollution. Whereas a large cache can ameliorate this problem, the issues of larger power consumption with increasing cache size, amplified at sub-100nm technologies, makes this solution prohibitive. In this paper, in order to address the issues relating to power-aware performance of caches, we propose a caching structure that addresses the following: 1) Definition of application-specific cache partitions as an aggregation of caching units (molecules). The parameters of each molecule namely size, associativity and line size are chosen so that the power consumed by it and access time are optimal for the given technology. 2) Application-specific resizing of cache partitions with variable and adaptive associativity per cache line, way size and variable line size. 3) A replacement policy that is transparent to the partition in terms of size, heterogeneity in associativity and line size. Through simulation studies we establish the superiority of molecular cache (caches built as aggregations of molecules) that offers a 29% power advantage over that of an equivalently performing traditional cache Keshavan Varadarajan, S. K. Nandy 0001, Vishal Sharda, Bharadwaj S. Amrutur, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell |
MICRO | 6 |
| 2005 | Hardware Support for Bulk Data Movement in Server PlatformsabstractBulk data movement occurs commonly in server work-loads and their performance is rather poor on today's microprocessors. We propose the use of small dedicated copy engines, and present a detailed analysis of a bulk data copy engine architecture. We describe the hardware support required to implement the copy engine and to tightly integrate it into server platforms. Our evaluation is based on an execution driven simulator that was extended with detailed models of bulk data movement engines. The simulation results show that dedicated engines are quite effective in eliminating the data movement overhead and are an attractive choice for handling bulk data in future high performance server platforms. Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Laxmi N. Bhuyan, Donald Newell |
ICCD | 3 |
| 2005 | Anatomy and Performance of SSL ProcessingabstractA wide spectrum of e-commerce (B2B/B2C), banking, financial trading and other business applications require the exchange of data to be highly secure. The Secure Sockets Layer (SSL) protocol provides the essential ingredients of secure communications - privacy, integrity and authentication. Though it is well-understood that security always comes at the cost of performance, these costs depend on the cryptographic algorithms. In this paper, we present a detailed description of the anatomy of a secure session. We analyze the time spent on the various cryptographic operations (symmetric, asymmetric and hashing) during the session negotiation and data transfer. We then analyze the most frequently used cryptographic algorithms (RSA, AES, DES, 3DES, RC4, MD5 and SHA-1). We determine the key components of these algorithms (setting up key schedules, encryption rounds, substitutions, permutations, etc) and determine where most of the time is spent. We also provide an architectural analysis of these algorithms, show the frequently executed instructions and discuss the ISA/hardware support that may be beneficial to improving SSL performance. We believe that the performance data presented in this paper is useful to performance analysts and processor architects to help accelerate SSL performance in future processors Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Laxmi N. Bhuyan |
ISPASS | 3 |
| 2004 | Architectural Characterization of TCP/IP Packet Processing on the Pentium M MicroprocessorabstractA majority of the current and next generation server applications (Web services, e-commerce, storage, etc.) employ TCP/IP as the communication protocol of choice. As a result, the performance of these applications is heavily dependent on the efficient TCP/IP packet processing within the termination nodes. This dependency becomes even greater as the bandwidth needs of these applications grow from 100 Mbps to 1 Gbps to 10 Gbps in the near future. Motivated by this, we focus on the following: (a) to understand the performance behavior of the various modes of TCP/IP processing, (b) to analyze the underlying architectural characteristics of TCP/IP packet processing and (c) to quantify the computational requirements of the TCP/IP packet processing component within realistic workloads. We achieve these goals by performing an in-depth analysis of packet processing performance on Intel's state-of-the-art low power Pentium/spl reg/ M microprocessor running the Microsoft Windows* Server 2003 operating system. Some of our key observations are - (i) that the mode of TCP/IP operation can significantly affect the performance requirements, (ii) that transmit-side processing is largely compute-intensive as compared to receive-side processing which is more memory-bound and (iii) that the computational requirements for sending/receiving packets can form a substantial component (28% to 40%) of commercial server workloads. From our analysis, we also discuss architectural as well as stack-related improvements that can help achieve higher server network throughput and result in improved application performance. Srihari Makineni, Ravi R. Iyer 0001 |
HPCA | 1 |