EDBT 2026 Demo / reviewers in the wild / expert
Ravi R. Iyer 0001
dblp:i/RaviRIyer · also Ravi Iyer 0001, Ravishankar Iyer 0001, Ravishankar R. Iyer 0001
· DBLP profile ↗
94ranked-venue papers
13as first author
10since 2021 · last 2025
0000-0001-5383-9561ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 76 · 10 first-author · 6 since 2021Software engineering, systems software and programming languages · 18 · 2 first-author · 2 since 2021Computer networks · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceabstractLarge Language Model (LLM) inference uses an autoregressive manner to generate one token at a time, which exhibits notably lower operational intensity compared to earlier Machine Learning (ML) models such as encoder-only transformers and Convolutional Neural Networks. At the same time, LLMs possess large parameter sizes and use key-value caches to store context information. Modern LLMs support context windows with up to 1 million tokens to generate versatile text, audio, and video content. A large key-value cache unique to each prompt requires a large memory capacity, limiting the inference batch size. Both low operational intensity and limited batch size necessitate a high memory bandwidth. However, contemporary hardware systems for ML model deployment, such as GPUs and TPUs, are primarily optimized for compute throughput. This mismatch challenges the efficient deployment of advanced LLMs and makes users to pay for expensive compute resources that are poorly utilized for the memory-bound LLM inference tasks. Yufeng Gu, Alireza Khadem, Sumanth Umesh, Xavier Servot, Onur Mutlu, Ravi R. Iyer 0001, Reetuparna Das |
ASPLOS (2) | 7 |
| 2023 | Mem-Rec: Memory Efficient Recommendation System using Alternative Representation
Gopi Krishna Jha, Anthony Thomas, Nilesh Jain, Sameh Gobriel, Tajana Rosing, Ravi R. Iyer 0001 |
ACML | 6 |
| 2023 | RAPID: Enabling fast online policy learning in dynamic public cloud environments
Drew Penney, Bin Li 0018, Lizhong Chen, Jaroslaw J. Sydir, Anna Drewek-Ossowicka, Ramesh Illikkal, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Andrew Herdrich |
Neurocomputing | 8 |
| 2023 | Eidetic: An In-Memory Matrix Multiplication Accelerator for Neural NetworksabstractThis paper presents theEideticarchitecture, which is an SRAM-based ASIC neural network accelerator that eliminates the need to continuously load weights from off-chip, while also minimizing the need to go off chip for intermediate results. Using in-situ arithmetic in the SRAM arrays, this architecture can supports a variety of precision types allowing for effective inference. We also present different data mapping policies for matrix-vector based networks (RNN and MLP) on theEideticarchitecture and describe the tradeoffs involved. With this architecture, multiple layers of a network can be concurrently mapped, storing both the layer weights and intermediate results on-chip, removing the energy and latency penalty of off-chip memory accesses. We evaluateEideticon Google's Neural Machine Translation System (GNMT) encoder and demonstrate a 17.20× increase in throughput and 7.77× reduction in average latency over a single TPUv2 chip. Charles Eckert, Arun Subramaniyan 0001, Xiaowei Wang 0005, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
IEEE Trans. Computers | 5 |
| 2022 | DPM-NFV: Dynamic Power Management Framework for 5G User Plane Function using Bayesian OptimizationabstractNetwork Function Virtualization (NFV), the replacement of purpose-built network appliances with software functions running on general purpose compute servers, is ubiquitous in today's telecommunication networks. The 5G User Plane Function (UPF) is an important example of an NFV workload, which enables 5G and internet communications. The UPF has strict packet drop requirements and because user traffic load can vary dramatically throughout the day, the selection of a single static configuration leads to over-provisioning of server resources. To reduce the cost of ownership, network operators can reduce power consumption during periods of low traffic load, but to do so they must ensure that packet drop requirements are met. In this paper we present DPM-NFV, a machine learning based framework that enables dynamic tuning of a real NFV system. Our methodology is composed of two phases: (1) Offline, targeted automated studies use Bayesian Optimization to infer the best configurations for various load levels; (2) Online, a run-time classifier dynamically selects the best configuration for the current load. Our results obtained on a real system demonstrate that the UPF can meet strict packet drop requirements while reducing power consumption by up to 52% with smooth traffic and up to 46% with bursty traffic. Jaroslaw J. Sydir, Bin Li 0018, Pietro Mercati, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Michael Kishinevsky, Boris Serafimov |
GLOBECOM | 5 |
| 2022 | EZNAS: Evolving Zero-Cost Proxies For Neural Architecture ScoringabstractNeural Architecture Search (NAS) has significantly improved productivity in the design and deployment of neural networks (NN). As NAS typically evaluates multiple models by training them partially or completely, the improved productivity comes at the cost of significant carbon footprint. To alleviate this expensive training routine, zero-shot/cost proxies analyze an NN at initialization to generate a score, which correlates highly with its true accuracy. Zero-cost proxies are currently designed by experts conducting multiple cycles of empirical testing on possible algorithms, datasets, and neural architecture design spaces. This experimentation lowers productivity and is an unsustainable approach towards zero-cost proxy design as deep learning use-cases diversify in nature. Additionally, existing zero-cost proxies fail to generalize across neural architecture design spaces. In this paper, we propose a genetic programming framework to automate the discovery of zero-cost proxies for neural architecture scoring. Our methodology efficiently discovers an interpretable and generalizable zero-cost proxy that gives state of the art score-accuracy correlation on all datasets and search spaces of NASBench-201 and Network Design Spaces (NDS). We believe that this research indicates a promising direction towards automatically discovering zero-cost proxies that can work across network architecture design spaces, datasets, and tasks. Yash Akhauri, Juan Pablo Muñoz, Nilesh Jain, Ravi R. Iyer 0001 |
NeurIPS | 4 |
| 2021 | A 93 TOPS/Watt Near-Memory Reconfigurable SAD Accelerator for HEVC/AV1/JEM EncodingabstractMotion Estimation (ME) is a major bottleneck of a Video encoding pipeline. This paper presents a low power near memory Sum of Absolute Difference (SAD) accelerator for ME. The accelerator is composed of 64 modular SAD Processing Elements (PEs) on a Reconfigurable fabric, offering maximal parallelism to support traditional and futuristic Rate-Distortion-Optimization (RDO) schemes consistent with HEVC/AV1/JEM. The accelerator offers up-to 55% speedup over State-of-art accelerators and a 7x speedup when compared to a 12 core Intel Xeon E5 processor. Our solution achieves 93 TOPS/Watt running at 500MHz frequency, capable of processing real-time 4K 30fps video. Synthesized in 22nm process, the accelerator occupies 0.08mm2 and consumes 5.46mW dynamic power. Jainaveen Sundaram, Srivatsa Rangachar Srinivasa, Dileep Kurian, Indranil Chakraborty, Sirisha Rani Kale, Nilesh Jain, Tanay Karnik, Ravi R. Iyer 0001, Anuradha Srinivasan |
DATE | 8 |
| 2021 | Compute-Capable Block RAMs for Efficient Deep Learning Acceleration on FPGAsabstractThe density of FPGA on-chip memory has been continuously increasing with modern FPGAs having thousands of block RAMs (BRAMs) distributed across their reconfigurable fabric. These distributed BRAMs can provide a tremendous amount of on-chip bandwidth for efficient acceleration of data-intensive applications. In this work, we propose enhancing the ubiquitous FPGA BRAMs with in-memory compute-capabilities. As a result, BRAMs can act as normal storage units or their bitlines can be re-purposed as SIMD lanes executing bit-serial arithmetic operations. Our proposed architectural change results in 1.6× and 2.3× increase in the peak multiply-accumulate throughput of a large Stratix 10 FPGA, at a minimal cost of only 1.8% increase in the FPGA die size and no change to the BRAM's interface to the programmable routing. Then, we present RIMA, a reconfigurable in-memory accelerator architecture for deep learning (DL) inference. RIMA exploits the proposed compute-capable BRAMs and the FPGA's reconfigurability to achieve 1.25× and 3× higher performance compared to the state-of-the-art Brainwave DL soft processor for 8-bit integer and block floating-point precisions, respectively. In addition, RIMA implemented on a Stratix 10 FPGA enhanced with compute-capable BRAMs can achieve an order of magnitude higher performance compared to a same-generation GPU. Xiaowei Wang 0005, Vidushi Goyal, Jiecao Yu, Valeria Bertacco, Andrew Boutros, Eriko Nurvitadhi, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
FCCM | 8 |
| 2021 | E2E Visual Analytics: Achieving >10X Edge/Cloud OptimizationsabstractAs visual analytics continues to rapidly grow, there is a critical need to improve the end-to-end efficiency of visual processing in edge/cloud systems. In this paper, we cover algorithms, systems and optimizations in three major areas for edge/cloud visual processing: (1) addressing storage and retrieval efficiency of visual data and meta-data by employing and optimizing visual data management systems, (2) addressing compute efficiency of visual analytics by taking advantage of co-optimization between the compression and analytics domains and (3) addressing networking (bandwidth) efficiency of visual data compression by tailoring it based on analytics tasks. We describe techniques in each of the above areas and measure its efficacy on state-of-the-art platforms (Intel Xeon), workloads and datasets. Our results show that we can achieve >10X improvements in each area based on novel algorithms, systems, and co-design optimizations. We also outline future research directions based on our findings which outline areas of further performance and efficiency advantages in end-to-end visual analytics. Chaunte W. Lacewell, Nilesh A. Ahuja, Juan Pablo Muñoz, Parual Datta, Ragaad AlTarawneh, Vui Seng Chua, Nilesh Jain, Omesh Tickoo, Ravi R. Iyer 0001 |
NAS | 9 |
| 2021 | Cache Compression with Efficient in-SRAM Data ComparisonabstractWe present a novel cache compression method that leverages the fine-grained data duplication across cache lines. We leverage the XOR operation of the in-SRAM bit-line computing peripherals, to search for compressible data over a wide range of data locations on cache, reducing the data movement requirements. To reduce the decompression latency, we design specialized compression schemes by fetching the data with the same parallelism as the original cache, according to the architecture of the last-level cache slice. The proposed compression method achieves a 2.05× compression ratio on average (up to 67×), and 4.73% of speedup on average (up to 29%), over the SPEC2006 benchmarks. Xiaowei Wang 0005, Charles Augustine, Eriko Nurvitadhi, Ravi R. Iyer 0001, Li Zhao 0002, Reetuparna Das |
NAS | 4 |
| 2020 | RLDRM: Closed Loop Dynamic Cache Allocation with Deep Reinforcement Learning for Network Function VirtualizationabstractNetwork function virtualization (NFV) technology attracts tremendous interests from telecommunication industry and data center operators, as it allows service providers to assign resource for Virtual Network Functions (VNFs) on demand, achieving better flexibility, programmability, and scalability. To improve server utilization, one popular practice is to deploy best effort (BE) workloads along with high priority (HP) VNFs when high priority VNF's resource usage is detected to be low. The key challenge of this deployment scheme is to dynamically balance the Service level objective (SLO) and the total cost of ownership (TCO) to optimize the data center efficiency under inherently fluctuating workloads. With the recent advancement in deep reinforcement learning, we conjecture that it has the potential to solve this challenge by adaptively adjusting resource allocation to reach the improved performance and higher server utilization. In this paper, we present a closed-loop automation system RLDRM11RLDRM: Reinforcement Learning Dynamic Resource Management to dynamically adjust Last Level Cache allocation between HP VNFs and BE workloads using deep reinforcement learning. The results demonstrate improved server utilization while maintaining required SLO for the HP VNFs. Bin Li 0018, Yipeng Wang 0002, Ren Wang 0001, Tsung-Yuan Charlie Tai, Ravi R. Iyer 0001, Zhu Zhou, Andrew Herdrich, Ameer Haj-Ali, Ion Stoica, Krste Asanovic |
NetSoft | 5 |
| 2019 | Bit Prudent In-Cache Acceleration of Deep Convolutional Neural NetworksabstractWe propose Bit Prudent In-Cache Acceleration of Deep Convolutional Neural Networks - an in-SRAM architecture for accelerating Convolutional Neural Network (CNN) inference by leveraging network redundancy and massive parallelism. The network redundancy is exploited in two ways. First, we prune and fine-tune the trained network model and develop two distinct methods - coalescing and overlapping - to run inferences efficiently with sparse models. Second, we propose an architecture for network models with a reduced bit width by leveraging bit-serial computation. Our proposed architecture achieves a 17.7×/3.7× speedup over server class CPU/GPU, and a 1.6× speedup compared to the relevant in-cache accelerator, with 2% area overhead each processor die, and no loss on top-1 accuracy for AlexNet. With a relaxed accuracy limit, our tunable architecture achieves higher speedups. Xiaowei Wang 0005, Jiecao Yu, Charles Augustine, Ravi R. Iyer 0001, Reetuparna Das |
HPCA | 4 |
| 2018 | Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural NetworksabstractThis paper presents the Neural Cache architecture, which re-purposes cache structures to transform them into massively parallel compute units capable of running inferences for Deep Neural Networks. Techniques to do in-situ arithmetic in SRAM arrays, create efficient data mapping and reducing data movement are proposed. The Neural Cache architecture is capable of fully executing convolutional, fully connected, and pooling layers in-cache. The proposed architecture also supports quantization in-cache. Our experimental results show that the proposed architecture can improve inference latency by 8.3× over state-of-art multi-core CPU (Xeon E5), 7.7× over server class GPU (Titan Xp), for Inception v3 model. Neural Cache improves inference throughput by 12.4× over CPU (2.2× over GPU), while reducing power consumption by 50% over CPU (53% over GPU). Charles Eckert, Xiaowei Wang 0005, Arun Subramaniyan 0001, Ravi R. Iyer 0001, Dennis Sylvester, David T. Blaauw, Reetuparna Das |
ISCA | 5 |
| 2017 | Race-to-sleep + content caching + display caching: a recipe for energy-efficient video streaming on handheldsabstractVideo streaming has become the most common application in handhelds and this trend is expected to grow in future to account for about 75% of all mobile data traffic by 2021. Thus, optimizing the performance and energy consumption of video processing in mobile devices is critical for sustaining the handheld market growth. In this paper, we propose three complementary techniques, race-to-sleep, content caching and display caching, to minimize the energy consumption of the video processing flows. Unlike the state-of-the-art frame-by-frame processing of a video decoder, the first scheme, race-to-sleep, uses two approaches, called batching of frames and frequency boosting to prolong its sleep state for saving energy, while avoiding any frame drops. The second scheme, content caching, exploits the content similarity of smaller video blocks, called macroblocks, to design a novel cache organization for reducing the memory pressure. The third scheme, in turn, takes advantage of content similarity at the display controller to facilitate display caching further improving energy efficiency. We integrate these three schemes for developing an end-to-end video processing framework and evaluate our design on a comprehensive mobile system design platform with a variety of video processing workloads. Our evaluations show that the proposed three techniques complement each other in improving performance by avoiding frame drops and reducing the energy consumption of video streaming applications by 21%, on average, compared to the current baseline design. Haibo Zhang 0005, Prasanna Venkatesh Rengasamy, Shulin Zhao 0001, Nachiappan Chidambaram Nachiappan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
MICRO | 7 |
| 2016 | Cache QoS: From concept to reality in the Intel® Xeon® processor E5-2600 v3 product familyabstractOver the last decade, addressing quality of service (QoS) in multi-core server platforms has been growing research topic. QoS techniques have been proposed to address the shared resource contention between co-running applications or virtual machines in servers and thereby provide better isolation, performance determinism and potentially improve overall throughput. One of the most important shared resources is cache space. Most proposals for addressing shared cache contention are based on simulations and analysis and no commercial platforms were available that integrated such techniques and provided a practical solution. In this paper, we will present the first set of shared cache QoS techniques designed and implemented in state-of-the-art commercial servers (the Intel® Xeon® processor E5-2600 v3 product family). We will describe two key technologies: (i) Cache Monitoring Technology (CMT) to enable monitoring of shared cache usage by different applications and (ii) Cache Allocation Technology (CAT) which enables redistribution of shared cache space between applications to address contention. This is the first paper to describing these techniques as they moved from concept to reality, starting from early research to product implementation. We will also present case studies highlighting the value of these techniques using example scenarios of multi-programmed workloads, virtualized platforms in datacenters and communications platforms. Finally, we will describe initial software infrastructure and enabling for industry practitioners and researchers to take advantage of these technologies for their QoS needs. Andrew Herdrich, Edwin Verplanke, Priya Autee, Ramesh Illikkal, Chris Gianos, Ronak Singhal, Ravi R. Iyer 0001 |
HPCA | 7 |
| 2016 | Exploiting Core Criticality for Enhanced GPU PerformanceabstractModern memory access schedulers employed in GPUs typically optimize for memory throughput. They implicitly assume that all requests from different cores are equally important. However, we show that during the execution of a subset of CUDA applications, different cores can have different amounts of tolerance to latency. In particular, cores with a larger fraction of warps waiting for data to come back from DRAM are less likely to tolerate the latency of an outstanding memory request. Requests from such cores are more critical than requests from others. Based on this observation, this paper introduces a new memory scheduler, called (C)ritica(L)ity (A)ware (M)emory (S)cheduler (CLAMS), which takes into account the latency-tolerance of the cores that generate memory requests. The key idea is to use the fraction of critical requests in the memory request buffer to switch between scheduling policies optimized for criticality and locality. If this fraction is below a threshold, CLAMS prioritizes critical requests to ensure cores that cannot tolerate latency are serviced faster. Otherwise, CLAMS optimizes for locality, anticipating that there are too many critical requests and prioritizing one over another would not significantly benefit performance. Adwait Jog, Onur Kayiran, Ashutosh Pattnaik, Mahmut T. Kandemir, Onur Mutlu, Ravi R. Iyer 0001, Chita R. Das |
SIGMETRICS | 6 |
| 2015 | Platform-aware dynamic configuration support for efficient text processing on heterogeneous system
Mi Sun Park, Omesh Tickoo, Narayanan Vijaykrishnan, Mary Jane Irwin, Ravi R. Iyer 0001 |
DATE | 5 |
| 2015 | Design of a low power SoC testchip for wearables and IoTs
May Wu, Ravi R. Iyer 0001, Yatin Hoskote, Steven Zhang, Julio Zamora-Esquivel, German Fabila Garcia, Ilya Klotchkov, Mukesh Bhartiya |
Hot Chips Symposium | 2 |
| 2015 | Domain knowledge based energy management in handheldsabstractEnergy management in handheld devices is becoming a daunting task with the growing number of accelerators, increasing memory demands and high computing capacities required to support applications with stringent QoS needs. Current DVFS techniques that modulate power states of a single hardware component, or even recent proposals that manage multiple components, can lose out opportunities for attaining high energy efficiencies that may be possible by leveraging application domain knowledge. Thus, this paper proposes a coordinated multi-component energy optimization mechanism for handheld devices, where the energy profile of different components such as CPU, memory, GPU and IP cores are considered in unison to trigger the appropriate DVFS state by exploiting the application domain knowledge. Specifically, we show that for the important class of frame-based applications, the domain knowledge - frame processing rates, component utilization and available slack - can be used to decide effective DVFS states for each component from among the numerous choices. With such knowledge, rather than a brute force search of all speed setting choices, we propose two simpler heuristics, called Greedy policy and Kaldor-Hicks compensation policy, to make the decisions at frame boundaries. Our evaluations with 7 commonly-used Android apps show that our domain-aware coordinated DVFS policies have 23% better energy efficiency than the conventionally used Android governors, and are within ~9% of an optimal policy that does not drop any frames. Nachiappan Chidambaram Nachiappan, Praveen Yedlapalli, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
HPCA | 6 |
| 2015 | Low-complexity HOG for efficient video saliencyabstractIn this paper, we propose a low-complexity histogram of oriented gradients (HOG) implementation for efficient video saliency framework. After showing how original HOG calculations present significant computation bottleneck for visual understanding pipes, we present the optimized HOG flow and algorithm for video saliency framework, which can reduce computational requirements without losing algorithmic performance. Furthermore, simplification for light-weight computations and data-reusable scanning for optimal memory usage are explained for improving system efficiency. Based on our testing and analysis, the proposed HOG implementation optimizes computational complexity and performance while maintaining the video saliency algorithm capability. Teahyung Lee, Myung Hwangbo, Tanfer Alan, Omesh Tickoo, Ravi R. Iyer 0001 |
ICIP | 5 |
| 2015 | VIP: virtualizing IP chains on handheld platformsabstractEnergy-efficient user-interactive and display-oriented applications on handhelds rely heavily on multiple accelerators (termed IP cores) to meet their periodic frame processing needs. Further, these platforms are starting to host multiple applications concurrently on the multiple CPU cores. Unfortunately, today's hardware exposes an interface that forces the host software (Android drivers) to treat each IP core as an isolated device. Consequently, the host CPU has to get involved in the (i) processing of each frame, (ii) scheduling them to ensure timely progress through the IP cores to meet their QoS needs, and (iii) explicitly having to move data from one IP core to the next, with main memory serving as the common staging area. Nachiappan Chidambaram Nachiappan, Haibo Zhang 0005, Jihyun Ryoo, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
ISCA | 7 |
| 2015 | Towards Distributed Video SummarizationabstractVideo summarization is a fertile topic in multimedia research. While the advent of modern video cameras and several social networking and video sharing websites (like YouTube, Flickr, Facebook) has led to the generation of humongous amounts of redundant video data, video summarization has emerged as an effective methodology to automatically extract a succinct and condensed representation of a given video. The unprecedented increase in the volume of video data necessitates the usage of multiple, independent computers for its storage and processing. In order to understand the overall essence of a video, it is therefore necessary to develop an algorithm which can summarize a video distributed across multiple computers. In this paper, we propose a novel algorithm for distributed video summarization. Our algorithm requires minimal communication among the computers (over which the video is stored) and also enjoys nice theoretical properties. Our empirical results on several challenging, unconstrained videos corroborate the potential of the proposed framework for real-world distributed video summarization applications. Shayok Chakraborty, Omesh Tickoo, Ravi R. Iyer 0001 |
ACM Multimedia | 3 |
| 2015 | Adaptive Keyframe Selection for Video SummarizationabstractThe explosive growth of video data in the modern era has set the stage for research in the field of video summarization, which attempts to abstract the salient frames in a video in order to provide an easily interpreted synopsis. Existing work on video summarization has primarily been static - that is, the algorithms require the summary length to be specified as an input parameter. However, video streams are inherently dynamic in nature, while some of them are relatively simple in terms of visual content, others are much more complex due to camera/object motion, changing illumination, cluttered scenes and low quality. This necessitates the development of adaptive summarization techniques, which adapt to the complexity of a video and generate a summary accordingly. In this paper, we propose a novel algorithm to address this problem. We pose the summary selection as an optimization problem and derive an efficient technique to solve the summary length and the specific frames to be selected, through a single formulation. Our extensive empirical studies on a wide range of challenging, unconstrained videos demonstrate tremendous promise in using this method for real-world video summarization applications. Shayok Chakraborty, Omesh Tickoo, Ravi R. Iyer 0001 |
WACV | 3 |
| 2014 | QoS management on heterogeneous architecture for parallel applicationsabstractQuality of service (QoS) management is widely employed to provide differentiable performance to programs with distinctive priorities on conventional chip multi-processor (CMP) platforms. Recently, heterogeneous architecture integrating diverse processor cores on the same silicon has been proposed to better serve various application domains and it is expected to be an important design paradigm of future processors. Therefore, the QoS management on emerging heterogeneous systems will be of great significance. On the other hand, parallel applications are becoming increasingly important in modern computing community in order to explore the benefit of thread-level parallelism on CMPs. However, considering the diverse characteristics of thread synchronization, data sharing, and parallelization pattern, governing the execution of multiple parallel programs with different performance requirements becomes a complicated yet significant problem. In this paper, we study QoS management for parallel applications running on heterogeneous CMP systems. We comprehensively assess a series of task-to-core mapping policies on a real heterogeneous hardware (QuickIA) by characterizing their impacts on performance of individual applications. Our evaluation results show that the proposed QoS policies are effective to improve the performance of programs with highest priority while striking good tradeoff with system fairness. Ying Zhang 0016, Li Zhao 0002, Ramesh Illikkal, Ravi R. Iyer 0001, Andrew Herdrich, Lu Peng 0001 |
ICCD | 4 |
| 2013 | OWL: cooperative thread array aware scheduling techniques for improving GPGPU performanceabstractEmerging GPGPU architectures, along with programming models like CUDA and OpenCL, offer a cost-effective platform for many applications by providing high thread level parallelism at lower energy budgets. Unfortunately, for many general-purpose applications, available hardware resources of a GPGPU are not efficiently utilized, leading to lost opportunity in improving performance. A major cause of this is the inefficiency of current warp scheduling policies in tolerating long memory latencies. Adwait Jog, Onur Kayiran, Nachiappan Chidambaram Nachiappan, Asit K. Mishra, Mahmut T. Kandemir, Onur Mutlu, Ravi R. Iyer 0001, Chita R. Das |
ASPLOS | 7 |
| 2013 | OpenCL-Based Remote Offloading Framework for Trusted Mobile Cloud ComputingabstractOpenCL has emerged as the open standard for parallel programming for heterogeneous platforms enabling a uniform framework to discover, program, and distribute parallel workloads to the diverse set of compute units in the hardware. For that reason, there have been efforts exploring the advantages of parallelism from the OpenCL framework by offloading GPGPU workloads within an HPC cluster environment. In this paper, we present an OpenCL-based remote offloading framework designed for mobile platforms by shifting the motivation and advantages of using the OpenCL framework for the HPC cluster environment into mobile cloud computing where OpenCL workloads can be exported from a mobile node to the cloud. Furthermore, our offloading framework handles service discovery, access control, and data privacy by building the framework on top of a social peer-to-peer virtual private network, Social VPN. We developed a prototype implementation and deployed it into local- and wide-area environments to evaluate the performance improvement and energy implications of the proposed offloading framework. Our results show that, depending on the complexity of the workload and the amount of data transfer, the proposed architecture can achieve more energy efficient performance by offloading than executing locally. Heungsik Eom, Pierre St. Juste, Renato J. O. Figueiredo, Omesh Tickoo, Ramesh Illikkal, Ravi R. Iyer 0001 |
ICPADS | 6 |
| 2013 | Orchestrated scheduling and prefetching for GPGPUsabstractIn this paper, we present techniques that coordinate the thread scheduling and prefetching decisions in a General Purpose Graphics Processing Unit (GPGPU) architecture to better tolerate long memory latencies. We demonstrate that existing warp scheduling policies in GPGPU architectures are unable to effectively incorporate data prefetching. The main reason is that they schedule consecutive warps, which are likely to access nearby cache blocks and thus prefetch accurately for one another, back-to-back in consecutive cycles. This either 1) causes prefetches to be generated by a warp too close to the time their corresponding addresses are actually demanded by another warp, or 2) requires sophisticated prefetcher designs to correctly predict the addresses required by a future "far-ahead" warp while executing the current warp. Adwait Jog, Onur Kayiran, Asit K. Mishra, Mahmut T. Kandemir, Onur Mutlu, Ravi R. Iyer 0001, Chita R. Das |
ISCA | 6 |
| 2013 | Reducing cache and TLB power by exploiting memory region and privilege level semantics
Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Shih-Lien Lu, Ravi R. Iyer 0001, Tong Li 0003 |
J. Syst. Archit. | 5 |
| 2012 | Optimizing datacenter power with memory system levers for guaranteed quality-of-serviceabstractCo-location of applications is a proven technique to improve hardware utilization. Recent advances in virtualization have made co-location of independent applications on shared hardware a common scenario in datacenters. Co-location, while maintaining Quality-of-Service (QoS) for each application is a complex problem that is fast gaining relevance for these datacenters. The problem is exacerbated by the need for effective resource utilization at datacenter scales. In this work, we show that the memory system is a primary bottleneck in many workloads and is a more effective focal point when enforcing QoS. We examine four different memory system levers to enforce QoS: two that have been previously proposed, and two novel levers. We compare the effectiveness of each lever in minimizing power and resource needs, while enforcing QoS guarantees. We also evaluate the effectiveness of combining various levers and show that this combined approach can yield power reductions of up to 28%. Kshitij Sudan, Sadagopan Srinivasan, Rajeev Balasubramonian, Ravi R. Iyer 0001 |
PACT | 4 |
| 2012 | Accelerator-rich architectures: Implications, opportunities and challengesabstractProviding high performance at ultra-low power for a domain of applications is possible by designing and integrating accelerators. Accelerators may be fixed-function, programmable or re-configurable in nature. Integration of many such accelerators in a system-on-chip (SoC) or chip-multiprocessor (CMP) introduces several major implications on architecture, power/performance and programmability. In this paper, we will provide an overview of the key challenges and outline research opportunities and challenges for accelerator-rich architectures and devices. We will also describe example solutions in some of these areas as a potential direction for further exploration. Ravi R. Iyer 0001 |
ASP-DAC | 1 |
| 2012 | Cache revive: architecting volatile STT-RAM caches for enhanced performance in CMPsabstractHigh density, low leakage and non-volatility are the attractive features of Spin-Transfer-Torque-RAM (STT-RAM), which has made it a strong competitor against SRAM as a universal memory replacement in multi-core systems. However, STT-RAM suffers from high write latency and energy which has impeded its widespread adoption. To this end, we look at trading-off STT-RAM's non-volatility property (data-retention-time) to overcome these problems. We formulate the relationship between retention-time and write-latency, and find optimal retention-time for architecting an efficient cache hierarchy using STT-RAM. Our results show that, compared to SRAM-based design, our proposal can improve performance and energy consumption by 18% and 60%, respectively. Adwait Jog, Asit K. Mishra, Cong Xu 0002, Yuan Xie 0001, Narayanan Vijaykrishnan, Ravi R. Iyer 0001, Chita R. Das |
DAC | 6 |
| 2012 | PCASA: Probabilistic control-adjusted Selective Allocation for shared cachesabstractChip Multi-Processors (CMPs) are designed with an increasing number of cores to enable multiple and potentially heterogeneous applications to run simultaneously on the same system. However, this results in increasing pressure on shared resources, such as shared caches. With multiple processor cores sharing the same caches, high-priority applications may end up contending with low-priority applications for cache space and suffer significant performance slow-down, hence affecting the Quality of Service (QoS). In datacenters, Service Level Agreements (SLAs) impose a reserved amount of computing resources and specific cache space per cloud customer. Thus, to meet SLAs, a deterministic capacity management solution is required to control the occupancy of all applications. In this paper, we propose a novel QoS architecture, based on Probabilistic Selective Allocation (PSA), for priority-aware caches. Further, we show that applying a control-theoretic approach (Proportional Integral controller) to dynamically adjust PSA provides accurate and fine-grained capacity management. Konstantinos Aisopos, Jaideep Moses, Ramesh Illikkal, Ravi R. Iyer 0001, Donald Newell |
DATE | 4 |
| 2012 | Exploiting Semantics of Virtual Memory to Improve the Efficiency of the On-Chip Memory System
Bin Li 0018, Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Andrew Herdrich, Ravi R. Iyer 0001, Srihari Makineni |
Euro-Par | 7 |
| 2012 | QuickIA: Exploring heterogeneous architectures on real prototypesabstractOver the last decade, homogeneous multi-core processors emerged and became the de-facto approach for offering high parallelism, high performance and scalability for a wide range of platforms. We are now at an interesting juncture where several critical factors (smaller form factor devices, power challenges, need for specialization, etc) are guiding architects to consider heterogeneous chips and platforms for the next decade and beyond. Exploring heterogeneous architectures is challenging since it involves re-evaluating architecture options, OS implications and application development. In this paper, we describe these research challenges and then introduce a heterogeneous prototype platform called QuickIA that enables rapid exploration of heterogeneous architectures employing multiple generations of Intel processors for evaluating the implications of asymmetry and FPGAs to experiment with specialized processors or accelerators. We also show example case studies using the QuickIA research prototype to highlight its value in conducting heterogeneous architecture, OS and applications research. Bhushan Chitlur, Ganapati Srinivasa, Scott Hahn, Dheeraj Reddy, David A. Koufaty, Paul Brett, Abirami Prabhakaran, Li Zhao 0002, Nelson Ijih, Suchit Subhaschandra, Sabina Grover, Xiaowei Jiang, Ravi R. Iyer 0001 |
HPCA | 14 |
| 2012 | Reducing L1 caches power by exploiting software semanticsabstractTo access a set-associative L1 cache in a high-performance processor, all ways of the selected set are searched and fetched in parallel using physical address bits. Such a cache is oblivious of memory references' software semantics such as stack-heap bifurcation of the memory space, and user-kernel ring levels. This constitutes a waste of energy since e.g., a user-mode instruction fetch will never hit a cache block that contains kernel code. Similarly, a stack access will not hit a cacheline that contains heap data. Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Shih-Lien Lu, Ravi R. Iyer 0001, Tong Li 0003 |
ISLPED | 5 |
| 2012 | Leveraging Heterogeneity in DRAM Main Memories to Accelerate Critical Word AccessabstractThe DRAM main memory system in modern servers is largely homogeneous. In recent years, DRAM manufacturers have produced chips with vastly differing latency and energy characteristics. This provides the opportunity to build a heterogeneous main memory system where different parts of the address space can yield different latencies and energy per access. The limited prior work in this area has explored smart placement of pages with high activities. In this paper, we propose a novel alternative to exploit DRAM heterogeneity. We observe that the critical word in a cache line can be easily recognized beforehand and placed in a low-latency region of the main memory. Other non-critical words of the cache line can be placed in a low-energy region. We design an architecture that has low complexity and that can accelerate the transfer of the critical word by tens of cycles. For our benchmark suite, we show an average performance improvement of 12.9% and an accompanying memory energy reduction of 15%. Niladrish Chatterjee, Manjunath Shevgoor, Rajeev Balasubramonian, Al Davis, Zhen Fang 0002, Ramesh Illikkal, Ravi R. Iyer 0001 |
MICRO | 7 |
| 2012 | Dynamic QoS management for chip multiprocessorsabstractWith the continuing scaling of semiconductor technologies, chip multiprocessor (CMP) has become the de facto design for modern high performance computer architectures. It is expected that more and more applications with diverse requirements will run simultaneously on the CMP platform. However, this will exert contention on shared resources such as the last level cache, network-on-chip bandwidth and off-chip memory bandwidth, thus affecting the performance and quality-of-service (QoS) significantly. In this environment, efficient resource sharing and a guarantee of a certain level of performance is highly desirable. Researchers have proposed different frameworks for providing QoS. Most of these frameworks focus on individual resource for QoS management. Coordinated management of multiple QoS-aware shared resources at runtime remains an open problem. Recently, there has been work that proposed a class-of-serviced based framework to jointly managing cache, NoC and memory resources simultaneously. However, the work allocates shared resources statically at the beginning of application runtime, and do not dynamically track, manage and share shared resources across applications. In this article, we address this limitation by proposing dynamic resource management policies that monitor the resource usage of applications at runtime, then steals resources from the high-priority applications for lower-priority ones. The goal is to maintain the targeted level of performance for high-priority applications while improving the performance of lower-priority applications. We use a PI (Proportional-Integral gain) feedback controller based technique to maintain stability in our framework. Our evaluation results show that our policy can improve performance for lower-priority applications significantly while maintaining the performance for high-priority application, thus demonstrating the effectiveness of our dynamic QoS resource management policy. Bin Li 0018, Li-Shiuan Peh, Li Zhao 0002, Ravi R. Iyer 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2011 | Template-based memory access engine for accelerators in SoCsabstractWith the rapid progress in semiconductor technologies, more and more accelerators can be integrated onto a single SoC chip. In SoCs, accelerators often require deterministic data access. However, as more and more applications are running simultaneous, latency can vary significantly due to contention. To address this problem, we propose a template-based memory access engine (MAE) for accelerators in SoCs. The proposed MAE can handle several common memory access patterns observed for near-future accelerators. Our evaluation results show that the proposed MAE can significantly reduce memory access latency and jitter, thus very effective for accelerators in SoCs. Bin Li 0018, Zhen Fang 0002, Ravi R. Iyer 0001 |
ASP-DAC | 3 |
| 2011 | Buffer-integrated-Cache: a cost-effective SRAM architecture for handheld and embedded platformsabstractIn an SoC, building local storage in each accelerator is area inefficient due to the low average utilization. In this paper, we present design and implementation of Buffer-integrated-Caching (BiC), which allows many buffers to be instantiated simultaneously in caches. BiC enables cores to view portions of the SRAM as cache while accelerators access other portions of the SRAM as private buffers. Carlos Flores Fajardo, Zhen Fang 0002, Ravi R. Iyer 0001, German Fabila Garcia, Li Zhao 0002 |
DAC | 3 |
| 2011 | ACCESS: Smart scheduling for asymmetric cache CMPsabstractIn current Chip-multiprocessors (CMPs), a significant portion of the die is consumed by the last-level cache. Until recently, the balance of cache and core space has been primarily guided by the needs of single applications. However, as multiple applications or virtual machines (VMs) are consolidated on such a platform, researchers have observed that not all VMs or applications require significant amount of cache space. In order to take advantage of this phenomenon, we explore the use of asymmetric last-level caches in a CMP platform. While asymmetric cache CMPs provide the benefit of reduced power and area, it is important to build in hardware/software support to appropriately schedule applications on to cores with suitable cache capacity. In this paper, we address this problem with our ACCESS architecture comprising of: (a) asymmetric caches across a group of cores, (b) hardware support that enables prediction of cache performance on the different sized caches and (c) OS scheduler support to make use of the prediction capability and appropriately schedule applications on to core with suitable cache capacity. Measurements on a working prototype using SPEC2006 benchmarks show that our ACCESS architecture can effectively schedule jobs in an asymmetric cache CMP and provide 23% performance improvement compared to a naive scheduler, and is 97% close to an oracle scheduler in making schedules. Xiaowei Jiang, Asit K. Mishra, Li Zhao 0002, Ravi R. Iyer 0001, Zhen Fang 0002, Sadagopan Srinivasan, Srihari Makineni, Paul Brett, Chita R. Das |
HPCA | 4 |
| 2011 | Cost-effectively offering private buffers in SoCs and CMPsabstractHigh performance SoCs and CMPs integrate multiple cores and hardware accelerators such as network interface devices and speech recognition engines. Cores make use of SRAM organized as a cache. Accelerators make use of SRAM as special-purpose storage such as FIFOs, scratchpad memory, or other forms of private buffers. Dedicated private buffers provide benefits such as deterministic access, but are highly area inefficient due to the lower average utilization of the total available storage. Zhen Fang 0002, Li Zhao 0002, Ravi R. Iyer 0001, Carlos Flores Fajardo, German Fabila Garcia, Bin Li 0018, Steve R. King, Xiaowei Jiang, Srihari Makineni |
ICS | 3 |
| 2011 | Shared Resource Monitoring and Throughput Optimization in Cloud-Computing DatacentersabstractMany data centers employ server consolidation to maximize the efficiency of platform resource usage. As a result, multiple virtual machines (VMs) simultaneously run on each data center platform. Contention for shared resources between these virtual machines has an undesirable and non-deterministic impact on their performance behavior in such platforms. This paper proposes the use of shared resource monitoring to (a) understand the resource usage of each virtual machine on each platform, (b) collect resource usage and performance across different platforms to correlate implications of usage to performance, and (c) migrate VMs that are resource-constrained to improve overall data center throughput and improve Quality of Service (QoS). We focus our efforts on monitoring and addressing shared cache contention and propose a new optimization metric that captures the priority of the VM and the overall weighted throughput of the data center. We conduct detailed experiments emulating data center scenarios including on-line transaction processing workloads (based on TPC-C) middle-tier workloads (based on SPECjbb and SPECjAppServer) and financial workloads (based on PARSEC). We show that monitoring shared resource contention (such as shared cache) is highly beneficial to better manage throughput and QoS in a cloud-computing data center environment. Jaideep Moses, Ravi R. Iyer 0001, Ramesh Illikkal, Sadagopan Srinivasan, Konstantinos Aisopos |
IPDPS | 2 |
| 2011 | Keynote I: The era of heterogeneity: Are we prepared?abstractUsage models and applications are rapidly changing as a new class of devices (smart phones, smart TVs, etc) and rich cloud computing services (on datacenter servers) enter the marketplace. In this talk, I will start by describing some key examples of these radical changes in usage models, applications and devices. I will then highlight why the next decade of computing (clients and servers) will be based on heterogeneous architectures consisting of asymmetric cores, accelerators and hybrid cache/memory structures. The rest of the talk will be an in-depth discussion of the power/performance analysis challenges for heterogeneous architectures, such as (i) how do we analyze applications to determine the right mix of cores and accelerators, (ii) how do we provide performance/power prediction techniques for efficient OS scheduling on heterogeneous architectures?, (iii) how do we enable runtimes and applications to achieve the required QoS on heterogeneous architectures?, (iv) how do simulation/emulation methodologies and infrastructure have to change for rapid and consistent heterogeneous architecture exploration? For each of these, I will also give examples of work that is on-going and outline potential areas for future work on performance/power analysis for heterogeneous architectures. Ravi R. Iyer 0001 |
ISPASS | 1 |
| 2011 | HeteroScouts: hardware assist for OS scheduling in heterogeneous CMPsabstractDesigning heterogeneous chip multiprocessors (CMPs) with a mix of big cores (complex superscalar out-of-order pipelines) and small cores (simple in-order pipeline) is emerging as an attractive option for future architectures. Such architectures have the potential to deliver both high performance and power efficiency but this requires operating systems (OS) or virtual machine monitors (VMMs) to efficiently schedule each software thread on the type of core that is best suited for it. In this paper, we highlight the need for architectural support for OS scheduling in a heterogeneous CMP. We propose HeteroScouts, a hardware mechanism to assist the OS to efficiently predict the performance of a task on different cores in the platform. Sadagopan Srinivasan, Ravi R. Iyer 0001, Li Zhao 0002, Ramesh Illikkal |
SIGMETRICS | 2 |
| 2011 | CoQoS: Coordinating QoS-aware shared resources in NoC-based SoCs
Bin Li 0018, Li Zhao 0002, Ravi R. Iyer 0001, Li-Shiuan Peh, Michael Leddige, Michael Espig, Donald Newell |
J. Parallel Distributed Comput. | 3 |
| 2011 | RAFT: A router architecture with frequency tuning for on-chip networks
Asit K. Mishra, Aditya Yanamandra, Reetuparna Das, Soumya Eachempati, Ravi R. Iyer 0001, Narayanan Vijaykrishnan, Chita R. Das |
J. Parallel Distributed Comput. | 5 |
| 2010 | CHOP: Adaptive filter-based DRAM caching for CMP server platformsabstractAs manycore architectures enable a large number of cores on the die, a key challenge that emerges is the availability of memory bandwidth with conventional DRAM solutions. To address this challenge, integration of large DRAM caches that provide as much as 5× higher bandwidth and as low as 1/3rd of the latency (as compared to conventional DRAM) is very promising. However, organizing and implementing a large DRAM cache is challenging because of two primary tradeoffs: (a) DRAM caches at cache line granularity require too large an on-chip tag area that makes it undesirable and (b) DRAM caches with larger page granularity require too much bandwidth because the miss rate does not reduce enough to overcome the bandwidth increase. In this paper, we propose CHOP (Caching HOt Pages) in DRAM caches to address these challenges. We study several filter-based DRAM caching techniques: (a) a filter cache (CHOP-FC) that profiles pages and determines the hot subset of pages to allocate into the DRAM cache, (b) a memory-based filter cache (CHOP-MFC) that spills and fills filter state to improve the accuracy and reduce the size of the filter cache and (c) an adaptive DRAM caching technique (CHOP-AFC) to determine when the filter cache should be enabled and disabled for DRAM caching. We conduct detailed simulations with server workloads to show that our filter-based DRAM caching techniques achieve the following: (a) on average over 30% performance improvement over previous solutions, (b) several magnitudes lower area overhead in tag space required for cache-line based DRAM caches, (c) significantly lower memory bandwidth consumption as compared to page-granular DRAM caches. Xiaowei Jiang, Niti Madan, Li Zhao 0002, Mike Upton, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell, Yan Solihin, Rajeev Balasubramonian |
HPCA | 5 |
| 2010 | Quality of service shared cache management in chip multiprocessor architectureabstractThe trends in enterprise IT toward service-oriented computing, server consolidation, and virtual computing point to a future in which workloads are becoming increasingly diverse in terms of performance, reliability, and availability requirements. It can be expected that more and more applications with diverse requirements will run on a Chip Multi-Processor (CMP) and share platform resources such as the lowest level cache and off-chip bandwidth. In this environment, it is desirable to have microarchitecture and software support that can provide a guarantee of a certain level of performance, which we refer to as performance Quality of Service . In this article, we investigated a framework would be needed to manage the shared cache resource for fully providing QoS in a CMP. We found in order to fully provide QoS, we need to specify an appropriate QoS target for each job and apply an admission control policy to accept jobs only when their QoS targets can be satisfied. We also found that providing strict QoS often leads to a significant reduction in throughput due to resource fragmentation. We proposed throughput optimization techniques that include: (1) exploiting various QoS execution modes, and (2) a microarchitecture technique, which we refer to as resource stealing, that detects and reallocates excess cache capacity from a job while preserving its QoS target. We designed and evaluated three algorithms for performing resource stealing, which differ in how aggressive they are in stealing excess cache capacity, and in the degree of confidence in meeting QoS targets. In addition, we proposed a mechanism to dynamically enable or disable resource stealing depending on whether other jobs can benefit from additional cache capacity. We evaluated our QoS framework with a full system simulation of a 4-core CMP and a recent version of the Linux Operating System. We found that compared to an unoptimized scheme, the throughput can be improved by up to 47%, making the throughput significantly closer to a non-QoS CMP. Yan Solihin, Li Zhao 0002, Ravi R. Iyer 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2009 | Architecture Support for Improving Bulk Memory Copying and Initialization PerformanceabstractBulk memory copying and initialization is one of the most ubiquitous operations performed in current computer systems by both user applications and Operating Systems. While many current systems rely on a loop of loads and stores, there are proposals to introduce a single instruction to perform bulk memory copying. While such an instruction can improve performance due to generating fewer TLB and cache accesses, and requiring fewer pipeline resources, in this paper we show that the key to significantly improving the performance is removing pipeline and cache bottlenecks of the code that follows the instructions. We show that the bottlenecks arise due to (1) the pipeline clogged by the copying instruction, (2) lengthened critical path due to dependent instructions stalling while waiting for the copying to complete, and (3) the inability to specify (separately) the cacheability of the source and destination regions. We propose FastBCI, an architecture support that achieves the granularity efficiency of a bulk copying/ initialization instruction, but without its pipeline and cache bottlenecks. When applied to OS kernel buffer management, we show that on average FastBCI achieves anywhere between 23% to 32% speedup ratios, which is roughly 3x-4x of an alternative scheme, and 1.5x-2x of a highly optimistic DMA with zero setup and interrupt overheads. Xiaowei Jiang, Yan Solihin, Li Zhao 0002, Ravi R. Iyer 0001 |
PACT | 4 |
| 2009 | HiPPAI: High Performance Portable Accelerator Interface for SoCsabstractSpecialized hardware accelerators are enabling today's System on Chip (SoC) platforms to target various applications. In this paper we show that as these SoCs evolve in complexity and usage, the programming models for such platforms need to evolve beyond the traditional driver oriented architecture. Using a test set up that employs a programmable FPGA based accelerator to implement one of the critical computation functions of a Mobile Augmented Reality based workload, we describe the performance drawbacks that a conventional programming model brings to compute environments employing hardware accelerators. We show that these performance issues become more critical as the interface latencies continue to improve over time with better hardware integration and efficient interconnect technologies. Under these usage scenarios, we show with measurements that the software overheads enforced by the current programming model, like those associated with system calls, memory copy and memory address translations account for a major part of the performance overheads. We then propose a novel High Performance Portable Accelerator Interface (HiPPAI) for SoC platforms using hardware accelerators to reduce the software overheads mentioned above. In addition, we position the new programming interface to allow for function portability between software and hardware function accelerators to reduce the application development effort. Our proposed model relies on two major building blocks for performance improvement. A uniform virtual memory addressing model based on hardware IOMMU support and direct user mode access to accelerators. We demonstrate how these enhancements reduce the overheads of system calls and address translations at the user/kernel boundary in traditional software stacks and enable function portability. Paul M. Stillwell, Vineet Chadha, Omesh Tickoo, Steven Zhang, Ramesh Illikkal, Ravi R. Iyer 0001, Donald Newell |
HiPC | 6 |
| 2009 | Optimizing communication and capacity in a 3D stacked reconfigurable cache hierarchyabstractCache hierarchies in future many-core processors are expected to grow in size and contribute a large fraction of overall processor power and performance. In this paper, we postulate a 3D chip design that stacks SRAM and DRAM upon processing cores and employs OS-based page coloring to minimize horizontal communication of cache data. We then propose a heterogeneous reconfigurable cache design that takes advantage of the high density of DRAM and the superior power/delay characteristics of SRAM to efficiently meet the working set demands of each individual core. Finally, we analyze the communication patterns for such a processor and show that a tree topology is an ideal fit that significantly reduces the power and latency requirements of the on-chip network. The above proposals are synergistic: each proposal is made more compelling because of its combination with the other innovations described in this paper. The proposed reconfigurable cache model improves performance by up to 19% along with 48% savings in network power. Niti Madan, Li Zhao 0002, Naveen Muralimanohar, Aniruddha N. Udipi, Rajeev Balasubramonian, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell |
HPCA | 6 |
| 2009 | Using checksum to reduce power consumption of display systems for low-motion contentabstractPower consumption of the display subsytem has been a relatively less explored area compared to other components of a mobile device including computing, storage, and networking units, although the former often constitutes one of the most power-hungry portions of the system. Typical applications on a mobile device such as Web browsing and text editing tend to have rather static image content; each frame hardly changes from the previous one. Efficiently detecting and handling no-motion scenarios is thus critical to extend the battery life. This paper focuses on image change detection. We propose to use checksum to detect image changes. Specifically, CRC hardware is used to optimize the power consumption of (1) refresh of a local display and (2) data compression for wireless remote display. Compared with a traditional, pixel-by-pixel comparison approach, using checksum for image change detection is not only fast, but also reduces accesses to the frame buffer, resulting in significant power savings. We have built a FPGA prototype to verify that CRC can capture image changes well enough to ensure a ¿visually lossless¿ quality. Kyungtae Han, Zhen Fang 0002, Paul Diefenbaugh, Richard Forand, Ravi R. Iyer 0001, Donald Newell |
ICCD | 5 |
| 2009 | Accelerating mobile augmented reality on a handheld platformabstractMobile Augmented Reality (MAR) is an emerging visual computing application for the mobile Internet device (MID). In one MAR usage model, the user points the handheld device to an object (like a wine bottle or a building) and the MID automatically recognizes and displays information regarding the object. Achieving this in software on the handheld requires significant compute processing for object recognition and matching. In this paper, we identify hotspot functions of the MAR workload on a low-power ×86 platform that motivates acceleration. We present the detailed design of two hardware accelerators, one for object recognition (MAR-HA) and the other for match processing (MAR-MA). We also quantify the performance and area efficiency of the hardware accelerators. Our analysis shows that hardware acceleration has the potential to improve the individual hotspot functions by as much as 20×, and overall response time by 7×. As a result, user response time can be reduced significantly. Zhen Fang 0002, Sadagopan Srinivasan, Ravi R. Iyer 0001, Donald Newell |
ICCD | 5 |
| 2009 | Rate-based QoS techniques for cache/memory in CMP platformsabstractAs we embrace the era of chip multi-processors (CMP), we are faced with two major architectural challenges: (i) QoS or performance management of disparate applications running on CPU cores contending for shared cache/memory resources and (ii) global/local power management techniques to stay within the overall platform constraints. The problem is exacerbated as the number of cores sharing the resources in a chip increase. In the past, researchers have proposed independent solutions for these two problems. In this paper, we show that rate-based techniques that are employed to address power management can be adapted to address cache/memory QoS issues. The basic approach is to throttle down the processing rate of a core if it is running a low-priority task and its execution is interfering with the performance of a high priority task due to platform resource contention (i.e. cache or memory contention). We evaluate two rate throttling mechanisms (clock modulation, and frequency scaling) for effectively managing the interference between applications running in a CMP platform and delivering QoS/performance management. We show that clock modulation is much more applicable to cache/memory QoS than frequency scaling and that resource monitoring along with rate control provides effective power-performance management in CMP platforms. Andrew Herdrich, Ramesh Illikkal, Ravi R. Iyer 0001, Donald Newell, Vineet Chadha, Jaideep Moses |
ICS | 3 |
| 2009 | CMPSched$im: Evaluating OS/CMP interaction on shared cache managementabstractCMPs have now become mainstream and are growing in complexity with more cores, several shared resources (cache, memory, etc) and the potential for additional heterogeneous elements. In order to manage these resources, it is becoming critical to optimize the interaction between the execution environment (operating systems, virtual machine monitors, etc) and the CMP platform. Performance analysis of such OS and CMP interactions is challenging because it requires long running full-system execution-driven simulations. In this paper, we explore an alternative approach (CMPSched$im) to evaluate the interaction of OS and CMP architectures. In particular, CMPSched$im is focused on evaluating techniques to address the shared cache management problem through better interaction between CMP hardware and operating system scheduling. CMPSched$im enables fast and flexible exploration of this interaction by combining the benefits of (a) binary instrumentation tools (Pin), (b) user-level scheduling tools (Linsched) and (c) simple core/cache simulators. In this paper, we describe CMPSched$im in detail and present case studies showing how CMPSched$im can be used to optimize OS scheduling by taking advantage of novel shared cache monitoring capabilities in the hardware. We also describe OS scheduling heuristics to improve overall system performance through resource monitoring and application classification to achieve near optimal scheduling that minimizes the effects of contention in the shared cache of a CMP platform. Jaideep Moses, Konstantinos Aisopos, Aamer Jaleel, Ravi R. Iyer 0001, Ramesh Illikkal, Donald Newell, Srihari Makineni |
ISPASS | 4 |
| 2009 | A case for dynamic frequency tuning in on-chip networksabstractPerformance and power are the first order design metrics for Network-on-Chips (NoCs) that have become the de-facto standard in providing scalable communication backbones for multicores/CMPs. However, NoCs can be plagued by higher power consumption and degraded throughput if the network and router are not designed properly. Towards this end, this paper proposes a novel router architecture, where we tune the frequency of a router in response to network load to manage both performance and power. We propose three dynamic frequency tuning techniques, FreqBoost, FreqThrtl and FreqTune, targeted at congestion and power management in NoCs. As enablers for these techniques, we exploit Dynamic Voltage and Frequency Scaling (DVFS) and the imbalance in a generic router pipeline through time stealing. Experiments using synthetic workloads on a 8x8 wormhole-switched mesh interconnect show that FreqBoost is a better choice for reducing average latency (maximum 40%) while, FreqThrtl provides the maximum benefits in terms of power saving and energy delay product (EDP). The FreqTune scheme is a better candidate for optimizing both performance and power, achieving on an average 36% reduction in latency, 13% savings in power (up to 24% at high load), and 40% savings (up to 70% at high load) in EDP. With application benchmarks, we observe IPC improvement up to 23% using our design. The performance and power benefits also scale for larger NoCs. Asit K. Mishra, Reetuparna Das, Soumya Eachempati, Ravi R. Iyer 0001, Narayanan Vijaykrishnan, Chita R. Das |
MICRO | 4 |
| 2009 | Hardware/Software Co-Simulation for Last Level Cache ExplorationabstractLarger last level caches are being considered for bridging the performance gap between the processors and the memory subsystem. It requires much longer simulation time to exercise the whole cache and get accurate evaluation results. In this paper, we motivate the need for a trace-driven hardware/software co-simulation approach to solve this problem. We describe the components of the hardware/software co-simulation: (a) a hardware approach for FSB (front side bus) cycle accurate long trace extraction and (b) a software simulation infrastructure to simulate arbitrary length of traces limited only by the storage system. We compare this hardware/software co-simulation approach to previous approaches (software-only and hardware FPGA-cache simulation) and articulate why our proposed approach is more flexible, more repeatable and sufficiently fast for last-level cache exploration. Evaluation results based on our hardware/software co-simulation infrastructure shows that our approach provides accurate results and shows the importance of timing information in accurate trace-driven simulations. We also demonstrate that it is not adequate to use short traces to get accurate results. Instead, the whole trace for the whole lifecycle of the workload, or at least a long trace (~5 minutes) should be used to capture the real behavior of the workloads. Tao Wang 0004, Qigang Wang, Michael Liao, Li Zhao 0002, Ravi R. Iyer 0001, Ramesh Illikkal, John Du |
NAS | 8 |
| 2009 | VM3: Measuring, modeling and managing VM shared resources
Ravi R. Iyer 0001, Ramesh Illikkal, Omesh Tickoo, Li Zhao 0002, Padma Apparao, Donald Newell |
Comput. Networks | 1 |
| 2008 | To Snoop or Not to Snoop: Evaluation of Fine-Grain and Coarse-Grain Snoop Filtering Techniques
Jessica Young, Srihari Makineni, Ravi R. Iyer 0001, Donald Newell, Adrian Moga |
Euro-Par | 3 |
| 2008 | Achieving 10Gbps Network Processing: Are We There Yet?
Priya Govindarajan, Srihari Makineni, Donald Newell, Ravi R. Iyer 0001, Ram Huggahalli, Amit Kumar 0008 |
HiPC | 4 |
| 2008 | Performance and power optimization through data compression in Network-on-Chip architecturesabstractThe trend towards integrating multiple cores on the same die has accentuated the need for larger on-chip caches. Such large caches are constructed as a multitude of smaller cache banks interconnected through a packet-based network-on-chip (NoC) communication fabric. Thus, the NoC plays a critical role in optimizing the performance and power consumption of such non-uniform cache-based multicore architectures. While almost all prior NoC studies have focused on the design of router microarchitectures for achieving this goal, in this paper, we explore the role of data compression on NoC performance and energy behavior. In this context, we examine two different configurations that explore combinations of storage and communication compression: (1) Cache compression (CC) and (2) Compression in the NIC (NC). We also address techniques to hide the decompression latency by overlapping with NoC communication latency. Our simulation results with a diverse set of scientific and commercial benchmark traces reveal that CC can provide up to 33% reduction in network latency and up to 23% power savings. Even in the case of NC - where the data is compressed only when passing through the NoC fabric of the NUCA architecture and stored uncompressed - performance and power savings of up to 32% and 21%, respectively, can be obtained. These performance benefits in the interconnect translate up to 17% reduction in CPI. These benefits are orthogonal to any router architecture and make a strong case for utilizing compression for optimizing the performance and power envelope of NoC architectures. In addition, the study demonstrates the criticality of designing faster routers in shaping the performance behavior. Reetuparna Das, Asit K. Mishra, Chrysostomos Nicopoulos, Dongkook Park, Narayanan Vijaykrishnan, Ravi R. Iyer 0001, Mazin S. Yousif, Chita R. Das |
HPCA | 6 |
| 2008 | Characterization & analysis of a server consolidation benchmarkabstractVirtualization is already becoming ubiquitous in data centers for the consolidation of multiple workloads on a single platform. However, there are very few performance studies of server consolidation workloads in the literature. In this paper, our goal is to analyze the performance characteristics of a representative server consolidation workload. To address this goal, we have carried out extensive measurement and profiling experiments of a newly proposed consolidation workload (vConsolidate). vConsolidate consists of a compute intensive workload, a web server, a mail server and a database application running simultaneously on a single platform. We start by studying the performance slowdown of each workload due to consolidation on a contemporary multi-core dual-processor Intel platform. We then look at architectural characteristics such as CPI (cycles per instruction) and L2 MP (L2 misses per instruction) I, and analyze the benefits of larger caches for such a consolidated workload. We estimate the virtualization overheads for events such as context switches, interrupts and page faults and show how these impact the performance of the workload in consolidation. Finally, we also present the execution profile of the server consolidation workload and illustrate the life of each VM in the consolidated environment. We conclude by presenting an approach to developing a preliminary performance model based on the performance. Padma Apparao, Ravi R. Iyer 0001, Donald Newell, Tom Adelmeyer |
VEE | 2 |
| 2007 | CacheScouts: Fine-Grain Monitoring of Shared Caches in CMP Platforms
Li Zhao 0002, Ravi R. Iyer 0001, Ramesh Illikkal, Jaideep Moses, Srihari Makineni, Donald Newell |
PACT | 2 |
| 2007 | qTLB: Looking Inside the Look-Aside Buffer
Omesh Tickoo, Hari Kannan, Vineet Chadha, Ramesh Illikkal, Ravi R. Iyer 0001, Donald Newell |
HiPC | 5 |
| 2007 | Constraint-Aware Large-Scale CMP Cache Design
Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Ramesh Illikkal, Jaideep Moses, Donald Newell |
HiPC | 2 |
| 2007 | Exploring DRAM cache architectures for CMP server platformsabstractAs dual-core and quad-core processors arrive in the marketplace, the momentum behind CMP architectures continues to grow strong. As more and more cores/threads are placed on-die, the pressure on the memory subsystem is rapidly increasing. To address this issue, we explore DRAM cache architectures for CMP platforms. In this paper, we investigate the impact of introducing a low latency, large capacity and high bandwidth DRAM-based cache between the last level SRAM cache and memory subsystem. We first show the potential benefits of large DRAM caches for key commercial server workloads. As the primary hurdle to achieving these benefits with DRAM caches is the tag space overheads associated with them, we identify the most efficient DRAM cache organization and investigate various options. Our results show that the combination of 8-bit partial tags and 2-way sectoring achieves the highest performance (20% to 70%) with the lowest tag space (<25%) overhead. Li Zhao 0002, Ravi R. Iyer 0001, Ramesh Illikkal, Donald Newell |
ICCD | 2 |
| 2007 | Accelerating Full-System Simulation through Characterizing and Predicting Operating System PerformanceabstractThe ongoing trend of increasing computer hardware and software complexity has resulted in the increase in complexity and overheads of cycle-accurate processor system simulation, especially in full-system simulation which not only simulates user applications, but also the operating system (OS) and system libraries. This paper seeks to address how to accelerate full-system simulation through studying, characterizing, and predicting the performance behavior of OS services. Through studying the performance behavior of OS services, we found that each OS service exhibits multiple but limited behavior points that are repeated frequently. OS services also exhibit application-specific performance behavior and largely irregular patterns of occurrences. We exploit the observation to speed up full system simulation. A simulation run is divided into two non-overlapping periods: a learning period in which performance behavior of instances of an OS service are characterized and recorded, and a prediction period in which detailed simulation is replaced with a much faster emulation mode. During a prediction period, the behavior signature of an instance of an OS service is obtained through emulation while performance of the instance is predicted based on its signature and records of the OS service's past performance behavior. Statistically-rigorous algorithms are used to determine when to switch between learning and prediction periods. We test our simulation acceleration method with a set of OS-intensive applications and a recent version of Linux OS running on top of a detailed processor and memory hierarchy model implemented on Simics, a popular full-system simulator. On average, the method needs the learning periods to cover only 11% of OS service invocations in order to produce highly accurate performance estimates. This leads to an estimated simulation speedup of 4.9times, with an average performance prediction error of only 3.2%, and a worst case error of 4.2% Seongbeom Kim, Yan Solihin, Ravi R. Iyer 0001, Li Zhao 0002, W. Cohen |
ISPASS | 4 |
| 2007 | Understanding the Memory Performance of Data-Mining Workloads on Small, Medium, and Large-Scale CMPs Using Hardware-Software Co-simulationabstractWith the amount of data continuing to grow, extracting "data of interest" is becoming popular, pervasive, and more important than ever. Data mining, as this process is known as, seeks to draw meaningful conclusions, extract knowledge, and acquire models from vast amounts of data. These compute-intensive data-mining applications, where thread-level parallelism can be effectively exploited, are the design targets of future multi-core systems. As a result, future multi-core systems will be required to process terabyte-level workloads. To understand the memory system performance of data-mining applications, this paper presents the use of hardware-software co-simulation to explore the cache design space of several multi-threaded data mining applications. Our study reveals that the workloads are memory intensive, have large working-set sizes, and exhibit good data locality. We find that large DRAM caches can be useful to address their large working-set sizes Wenlong Li 0003, Eric Q. Li, Aamer Jaleel, Jiulong Shan, Yurong Chen 0001, Qigang Wang, Ravi R. Iyer 0001, Ramesh Illikkal, Yimin Zhang 0002, Michael Liao, Jinhua Du |
ISPASS | 7 |
| 2007 | A Framework for Providing Quality of Service in Chip Multi-ProcessorsabstractThe trends in enterprise IT toward service-oriented computing, server consolidation, and virtual computing point to a future in which workloads are becoming increasingly diverse in terms of performance, reliability, and availability requirements. It can be expected that more and more applications with diverse requirements will run on a CMP and share platform resources such as the lowest level cache and off-chip bandwidth. In this environment, it is desirable to have microarchitecture and software support that can provide a guarantee of a certain level of performance, which we refer to as performance Quality of Service. In this paper, we investigate a framework that would be needed for a CMP to fully provide QoS. We found that the ability of a CMP to partition platform resources alone is not sufficient for fully providing QoS. We also need an appropriate way to specify a QoS target, and an admission control policy that accepts jobs only when their QoS targets can be satisfied. We also found that providing strict QoS often leads to a significant reduction in throughput due to resource fragmentation. We propose novel throughput optimization techniques that include: (1) exploiting various QoS execution modes, and (2) a microarchitecture technique that steals excess resources from a job while still meeting its QoS target. We evaluated our QoS framework with a full system simulation of a 4-core CMP and a recent version of the Linux Operating System. We found that compared to an unoptimized scheme, the throughput can be improved by up to 47%, making the throughput significantly closer to a non-QoS CMP. Yan Solihin, Li Zhao 0002, Ravi R. Iyer 0001 |
MICRO | 4 |
| 2007 | QoS policies and architecture for cache/memory in CMP platformsabstractAs we enter the era of CMP platforms with multiple threads/cores on the die, the diversity of the simultaneous workloads running on them is expected to increase. The rapid deployment of virtualization as a means to consolidate workloads on to a single platform is a prime example of this trend. In such scenarios, the quality of service (QoS) that each individual workload gets from the platform can widely vary depending on the behavior of the simultaneously running workloads. While the number of cores assigned to each workload can be controlled, there is no hardware or software support in today's platforms to control allocation of platform resources such as cache space and memory bandwidth to individual workloads. In this paper, we propose a QoS-enabled memory architecture for CMP platforms that addresses this problem. The QoS-enabled memory architecture enables more cache resources (i.e. space) and memory resources (i.e. bandwidth) for high priority applications based on guidance from the operating environment. The architecture also allows dynamic resource reassignment during run-time to further optimize the performance of the high priority application with minimal degradation to low priority. To achieve these goals, we will describe the hardware/software support required in the platform as well as the operating environment (O/S and virtual machine monitor). Our evaluation framework consists of detailed platform simulation models and a QoS-enabled version of Linux. Based on evaluation experiments, we show the effectiveness of a QoS-enabled architecture and summarize key findings/trade-offs. Ravi R. Iyer 0001, Li Zhao 0002, Ramesh Illikkal, Srihari Makineni, Donald Newell, Yan Solihin, Lisa R. Hsu, Steven K. Reinhardt |
SIGMETRICS | 1 |
| 2007 | I/O processing in a virtualized platform: a simulation-driven approachabstractVirtualization provides levels of execution isolation and service partition that are desirable in many usage scenarios, but its associated overheads are a major impediment for wide deployment of virtualized environments. While the virtualization cost depends heavily on workloads, it has been demonstrated that the overhead is much higher with I/O intensive workloads compared to those which are compute-intensive. Unfortunately, the architectural reasons behind the I/O performance overheads are not well understood. Early research in characterizing these penalties has shown that cache misses and TLB related overheads contribute to most of I/O virtualization cost. While most of these evaluations are done using measurements, in this paper we present an execution-driven simulation based analysis methodology with symbol annotation as a means of evaluating the performance of virtualized workloads. This methodology provides detailed information at the architectural level (with a focus on cache and TLB) and allows designers to evaluate potential hardware enhancements to reduce virtualization overhead. We apply this methodology to study the network I/O performance of Xen (as a case study) in a full system simulation environment, using detailed cache and TLB models to profile and characterize software and hardware hotspots. By applying symbol annotation to the instruction flow reported by the execution driven simulator we derive function level call flow information. We follow the anatomy of I/O processing in a virtualized platform for network transmit and receive scenarios and demonstrate the impact of cache scaling and TLB size scaling on performance. Vineet Chadha, Ramesh Illikkal, Ravi R. Iyer 0001, Jaideep Moses, Donald Newell, Renato J. O. Figueiredo |
VEE | 3 |
| 2007 | Hardware Support for Accelerating Data Movement in Server PlatformabstractData movement (memory copies) is a very common operation during network processing and application execution on servers. The performance of this operation is rather poor on today's microprocessors due to the following aspects: 1) Several long-latency memory accesses are involved because the source and/or the destination are typically in memory, 2) latency hiding techniques, such as out-of-order execution, hardware threading, and prefetching, are not very effective for bulk data movement, and 3) microprocessors move data at register (small) granularity. In this paper, we show this overhead of bulk data movement and propose the use of dedicated copy engines to minimize it. We present a detailed analysis of copy engine architectures along two dimensions: 1) on-die versus off-die and 2) synchronous versus asynchronous. These copy engine architectures are superior to traditional direct memory access (DMA) engines because they are tightly coupled to the core architecture and enable lower overhead communication and signaling. We describe the hardware support required to implement these copy engines and integrate them into server platforms. We perform a detailed case study to evaluate the performance of these copy engines. The evaluation is based on an execution-driven simulator, which was extended with detailed models of copy engines. Our simulation results show that copy engines are effective in reducing the bulk data movement overhead and, hence, hold significant promise for high-performance server platforms Li Zhao 0002, Laxmi N. Bhuyan, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell |
IEEE Trans. Computers | 3 |
| 2007 | Editorial: Special Section on CMP ArchitecturesabstractCHIP multiprocessor (CMP) architectures are formed when multiple compute cores are integrated onto the same chip, forming a single, powerful, computational entity. Nearly every major high-performance processor manufacturer has at least two cores (dual-core) on the die, and their roadmaps are increasingly multicore, signaling that the era of big, monolithic uniprocessors has ended. This results from the fact that ever-larger uniprocessors do not scale well in power/performance, area/performance, or design complexity/performance. Continued performance scaling of these processors will thus be focused primarily on increasing multithreaded throughput. The rapid adoption of small-scale CMP platforms and the quest for high performance continues to accelerate the rate at which processor manufacturers are considering adding more cores on the die. Over the last decade, there has been significant progress in research and development in both academia and industry on CMP architecture and design for client and server platforms. And, while we have successfully entered the era of CMP, there are a significant set of challenges and opportunities that are yet to be investigated deeply. Some of the broad research areas being investigated include CMP architecture alternatives (for core, cache, interconnect, and memory), CMP design and technologies (process implications, new technologies like 3D-stacking, voltage/clock domain management, etc.), CMP performance evaluation (new simulation and modeling techniques, emerging applications and execution environments like virtualization), and novel CMP architectures and use cases (asymmetric or heterogeneous architectures, accelerators, etc.). There are many questions that are still to be answered for CMP architectures. Below, we list a few of the most compelling ones. Ravi R. Iyer 0001, Dean M. Tullsen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2006 | Communist, utilitarian, and capitalist cache policies on CMPs: caches as a shared resourceabstractAs chip multiprocessors (CMPs) become increasingly mainstream, architects have likewise become more interested in how best to share a cache hierarchy among multiple simultaneous threads of execution. The complexity of this problem is exacerbated as the number of simultaneous threads grows from two or four to the tens or hundreds. However, there is no consensus in the architectural community on what "best" means in this context. Some papers in the literature seek to equalize each thread's performance loss due to sharing, while others emphasize maximizing overall system performance. Furthermore, the specific effect of these goals varies depending on the metric used to define "performance".In this paper we label equal performance targets as Communist cache policies and overall performance targets as Utilitarian cache policies. We compare both of these models to the most common current model of a free-for-all cache (a Capitalist policy). We consider various performance metrics, including miss rates, bandwidth usage, and IPC, including both absolute and relative values of each metric. Using analytical models and behavioral cache simulation, we find that the optimal partitioning of a shared cache can vary greatly as different but reasonable definitions of optimality are applied. We also find that, although Communist and Utilitarian targets are generally compatible, each policy has workloads for which it provides poor overall performance or poor fairness, respectively. Finally, we find that simple policies like LRU replacement and static uniform partitioning are not sufficient to provide near-optimal performance under any reasonable definition, indicating that some thread-aware cache resource allocation mechanism is required. Lisa R. Hsu, Steven K. Reinhardt, Ravi R. Iyer 0001, Srihari Makineni |
PACT | 3 |
| 2006 | Receive Side Coalescing for Accelerating TCP/IP Processing
Srihari Makineni, Ravi R. Iyer 0001, Partha Sarangam, Donald Newell, Li Zhao 0002, Ramesh Illikkal, Jaideep Moses |
HiPC | 2 |
| 2006 | Molecular Caches: A caching structure for dynamic creation of application-specific Heterogeneous cache regionsabstractCMPs enable simultaneous execution of multiple applications on the same platforms that share cache resources. Diversity in the cache access patterns of these simultaneously executing applications can potentially trigger inter-application interference, leading to cache pollution. Whereas a large cache can ameliorate this problem, the issues of larger power consumption with increasing cache size, amplified at sub-100nm technologies, makes this solution prohibitive. In this paper, in order to address the issues relating to power-aware performance of caches, we propose a caching structure that addresses the following: 1) Definition of application-specific cache partitions as an aggregation of caching units (molecules). The parameters of each molecule namely size, associativity and line size are chosen so that the power consumed by it and access time are optimal for the given technology. 2) Application-specific resizing of cache partitions with variable and adaptive associativity per cache line, way size and variable line size. 3) A replacement policy that is transparent to the partition in terms of size, heterogeneity in associativity and line size. Through simulation studies we establish the superiority of molecular cache (caches built as aggregations of molecules) that offers a 29% power advantage over that of an equivalently performing traditional cache Keshavan Varadarajan, S. K. Nandy 0001, Vishal Sharda, Bharadwaj S. Amrutur, Ravi R. Iyer 0001, Srihari Makineni, Donald Newell |
MICRO | 5 |
| 2005 | SpliceNP: a TCP splicer using a network processorabstractTCP Splicing can be used in content-aware switches to tremendously reduce overall request latency. In order to reduce the processing latency further, we propose to offload the protocol processing onto network processors (NPs). An NP consists of a multithreaded multiprocessor architecture that can provide high throughput for packet processing or forwarding. However, offloading any protocol software to an NP needs to be carefully designed due to its low-level programming and limited control memory size.In this paper, we first analyze the operation of TCP Splicing in detail and evaluate its performance through measurements on a Linux-based switch. Then various possibilities of workload allocation among different computation resources in an NP are presented, and the design tradeoffs are discussed. A content aware switch is implemented using IXP 2400 NP and evaluated for performance comparison. The measurement results demonstrate that our NP-based switch can reduce the http processing latency by an average of 83.3% for a 1K byte web page. The amount of reduction increases with larger file sizes. It is also shown that the packet throughput can be improved by up to 5.7x across a range of files by taking advantage of multithreading and multiprocessing, available in the NP. Li Zhao 0002, Yan Luo 0001, Laxmi N. Bhuyan, Ravi R. Iyer 0001 |
ANCS | 4 |
| 2005 | Optimal network processor topologies for efficient packet processingabstractIn this paper, we propose a novel strategy to determine the optimal network processor (NP) topology for the target application tasks. We partition network applications into different stages with the consideration of limited instruction memory of the processing elements (PEs). We develop a theoretical approach to determine an optimal topology of the PEs via multiple pipelines. The idea of multiple pipelining is to exploit the task/packet level parallelism and the pipelines are further optimized to achieve the maximum throughput and resource utilization. Simulation results verify our analytical model and demonstrate the robustness of our approach in different NP configurations. Jingnan Yao, Yan Luo 0001, Laxmi N. Bhuyan, Ravi R. Iyer 0001 |
GLOBECOM | 4 |
| 2005 | Hardware Support for Bulk Data Movement in Server PlatformsabstractBulk data movement occurs commonly in server work-loads and their performance is rather poor on today's microprocessors. We propose the use of small dedicated copy engines, and present a detailed analysis of a bulk data copy engine architecture. We describe the hardware support required to implement the copy engine and to tightly integrate it into server platforms. Our evaluation is based on an execution driven simulator that was extended with detailed models of bulk data movement engines. The simulation results show that dedicated engines are quite effective in eliminating the data movement overhead and are an attractive choice for handling bulk data in future high performance server platforms. Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Laxmi N. Bhuyan, Donald Newell |
ICCD | 2 |
| 2005 | Performance characterization of iSCSI processing in a server platformabstractThe iSCSI protocol is a key building block for enabling IP-based network storage. High performance iSCSI implementations that can support multi-gigabit storage traffic throughput at low latencies are important in facilitating the widespread deployment of this technology. Motivated by this, our work presented in this paper focuses on analyzing the underlying architectural characteristics of iSCSI packet processing and quantifying its compute/memory requirements. Our analysis and characterization methodology is based on in-depth measurement experiments of iSCSI packet processing performance on Intel/sup /spl reg// Xeon/spl trade/ processor, running the Red Hat Linux operating system. Our measurement data shows the achievable throughput and consumed CPU utilization at different disk I/O sizes. We also study the overhead of integrity checks on iSCSI performance by enabling CRC computation. To understand the source of the iSCSI processing costs, we then do a detailed analysis of the architectural characteristics in terms of path length, cycles spent per instruction, cache misses at all levels and branch mispredictions. Hormuzd M. Khosravi, Abhijeet Joglekar, Ravi R. Iyer 0001 |
IPCCC | 3 |
| 2005 | Direct Cache Access for High Bandwidth Network I/OabstractRecent I/O technologies such as PCI-Express and 10 Gb Ethernet enable unprecedented levels of I/O bandwidths in mainstream platforms. However, in traditional architectures, memory latency alone can limit processors from matching 10 Gb inbound network I/O traffic. We propose a platform-wide method called direct cache access (DCA) to deliver inbound I/O data directly into processor caches. We demonstrate that DCA provides a significant reduction in memory latency and memory bandwidth for receive intensive network I/O applications. Analysis of benchmarks such as SPECWeb9, TPC-W and TPC-C shows that overall benefit depends on the relative volume of I/O to memory traffic as well as the spatial and temporal relationship between processor and I/O memory accesses. A system level perspective for the efficient implementation of DCA is presented. Ram Huggahalli, Ravi R. Iyer 0001, Scott Tetrick |
ISCA | 2 |
| 2005 | Anatomy and Performance of SSL ProcessingabstractA wide spectrum of e-commerce (B2B/B2C), banking, financial trading and other business applications require the exchange of data to be highly secure. The Secure Sockets Layer (SSL) protocol provides the essential ingredients of secure communications - privacy, integrity and authentication. Though it is well-understood that security always comes at the cost of performance, these costs depend on the cryptographic algorithms. In this paper, we present a detailed description of the anatomy of a secure session. We analyze the time spent on the various cryptographic operations (symmetric, asymmetric and hashing) during the session negotiation and data transfer. We then analyze the most frequently used cryptographic algorithms (RSA, AES, DES, 3DES, RC4, MD5 and SHA-1). We determine the key components of these algorithms (setting up key schedules, encryption rounds, substitutions, permutations, etc) and determine where most of the time is spent. We also provide an architectural analysis of these algorithms, show the frequently executed instructions and discuss the ISA/hardware support that may be beneficial to improving SSL performance. We believe that the performance data presented in this paper is useful to performance analysts and processor architects to help accelerate SSL performance in future processors Li Zhao 0002, Ravi R. Iyer 0001, Srihari Makineni, Laxmi N. Bhuyan |
ISPASS | 2 |
| 2004 | Architectural Characterization of TCP/IP Packet Processing on the Pentium M MicroprocessorabstractA majority of the current and next generation server applications (Web services, e-commerce, storage, etc.) employ TCP/IP as the communication protocol of choice. As a result, the performance of these applications is heavily dependent on the efficient TCP/IP packet processing within the termination nodes. This dependency becomes even greater as the bandwidth needs of these applications grow from 100 Mbps to 1 Gbps to 10 Gbps in the near future. Motivated by this, we focus on the following: (a) to understand the performance behavior of the various modes of TCP/IP processing, (b) to analyze the underlying architectural characteristics of TCP/IP packet processing and (c) to quantify the computational requirements of the TCP/IP packet processing component within realistic workloads. We achieve these goals by performing an in-depth analysis of packet processing performance on Intel's state-of-the-art low power Pentium/spl reg/ M microprocessor running the Microsoft Windows* Server 2003 operating system. Some of our key observations are - (i) that the mode of TCP/IP operation can significantly affect the performance requirements, (ii) that transmit-side processing is largely compute-intensive as compared to receive-side processing which is more memory-bound and (iii) that the computational requirements for sending/receiving packets can form a substantial component (28% to 40%) of commercial server workloads. From our analysis, we also discuss architectural as well as stack-related improvements that can help achieve higher server network throughput and result in improved application performance. Srihari Makineni, Ravi R. Iyer 0001 |
HPCA | 2 |
| 2004 | Architectural Characterization of an XML-Centric Commercial Server WorkloadabstractAs XML (extensible markup language) rapidly emerges as the standard for information storage and communication, it becomes increasingly important to understand its architectural characteristics and performance implications. In This work, our goal is to characterize a representative XML-based server in a managed runtime environment such as Java. Based on detailed measurements on an Intel/spl reg/ XeonTM processor-based commercial server running a real-world XML-based server workload, we start by looking at symmetric multiprocessor (SMP) scaling characteristics and the benefits of hyper-threading technology. Using performance monitoring events provided on the processor, we present an overview of the architectural characteristics (such as clocks per instruction (CPI), cache miss rates, memory/bus utilization, branch behavior and efficiency). Using profiling tools like Intel/spl reg/ VTuneTM performance analyzer, we map these architectural/performance characteristics to the various components of application execution - helping us identify hot spots and propose potential enhancements to code generation and application software. We believe that the information presented Are useful in understanding the XML processing characteristics and may serve as a useful first step to identifying potential hardware/software optimizations for improved future performance. Padma Apparao, Ravi R. Iyer 0001, Ricardo Morin, Naren Nayak, Mahesh Bhat, David Halliwell, William Steinberg |
ICPP | 2 |
| 2004 | CQoS: a framework for enabling QoS in shared caches of CMP platformsabstractCache hierarchies have been traditionally designed for usage by a single application, thread or core. As multi-threaded (MT) and multi-core (CMP) platform architectures emerge and their workloads range from single-threaded and multithreaded applications to complex virtual machines (VMs), a shared cache resource will be consumed by these different entities generating heterogeneous memory access streams exhibiting different locality properties and varying memory sensitivity. As a result, conventional cache management approaches that treat all memory accesses equally are bound to result in inefficient space utilization and poor performance even for applications with good locality properties. To address this problem, this paper presents a new cache management framework (CQoS) that (1) recognizes the heterogeneity in memory access streams, (2) introduces the notion of QoS to handle the varying degrees of locality and latency sensitivity and (3) assigns and enforces priorities to streams based on latency sensitivity, locality degree and application performance needs. To achieve this, we propose CQoS options for priority classification, priority assignment and priority enforcement. We briefly describe CQoS priority classification and assignment options -- ranging from user-driven and developer-driven to compiler-detected and flow-based approaches. Our focus in this paper is on CQoS mechanisms for priority enforcement -- these include (1) selective cache allocation, (2) static/dynamic set partitioning and (3) heterogeneous cache regions. We discuss the architectural design and implementation complexity of these CQoS options. To evaluate the performance trade-offs for these options, we have modeled these CQoS options in a cache simulator and evaluated their performance in CMP platforms running network-intensive server workloads. Our simulation results show the effectiveness of our proposed options and make the case for CQoS in future multi-threaded/multi-core platforms since it improves shared cache efficiency and increases overall system performance as a result. Ravi R. Iyer 0001 |
ICS | 1 |
| 2004 | Characterization and Evaluation of Cache Hierarchies for Web Servers
Ravi R. Iyer 0001 |
World Wide Web | 1 |
| 2002 | Design and analysis of static memory management policies for CC-NUMA multiprocessors
Ravi R. Iyer 0001, Hu-Jun Wang, Laxmi N. Bhuyan |
J. Syst. Archit. | 1 |
| 2001 | Exploring the Cache Design Space for Web ServersabstractAs Internet usage for e-commerce, communication and entertainment expands rapidly, careful attention needs to be paid to the design of Internet servers for achieving high performance and end-user satisfaction. In this paper we explore the access characteristics of front-end web workloads using SPECweb99 as the characteristic benchmark. In particular we explore the cache design space for single and dual-processor servers running the SPECweb99 benchmark. This includes studying the performance impact of cache size, line size and associativity for various levels of the cache hierarchy. We analyze SPECweb99 cache performance in terms of the misses per instruction, miss rate, instruction fetches vs. data reads and writes, coherence due to sharing across caches the rate at which writebacks occur, spatial and temporal locality, and other such properties. The results presented in this paper can be used as a guide for architects who make cache design choices for future processors and performance analysts who build web server models to predict future performance. Ravi R. Iyer 0001 |
IPDPS | 1 |
| 2000 | Using Switch Directories to Speed Up Cache-to-Cache Transfers in CC-NUMA MultiprocessorsabstractIn this paper we propose a novel hardware caching technique, called switch directory, to reduce the communication latency in CC-NUMA multiprocessors. The main idea is to implement small fast directory caches in crossbar switches of the inter-connect medium to capture and store ownership information as the data flows from the memory module to the requesting processor. Using the stored information, the switch directory re-routes subsequent requests to dirty blocks directly to the owner cache, thus reducing the latency for home node processing such as slow DRAM directory access and coherence controller occupancies. The design and implementation details of a DiRectory Embedded Switch ARchitecture; DRESAR, are presented. We explore the performance benefits of switch directories by modeling DRESAR in a detailed execution driven simulator. Our results show that the switch directories can improve performance by up to 60% reduction in home node cache-to-cache transfers for several scientific applications and commercial workloads. Ravi R. Iyer 0001, Laxmi N. Bhuyan, Ashwini K. Nanda |
IPDPS | 1 |
| 2000 | Design and Evaluation of a Switch Cache Architecture for CC-NUMA MultiprocessorsabstractCache coherent nonuniform memory access (CC-NUMA) multiprocessors provide a scalable design for shared memory. But, they continue to suffer from large remote memory access latencies due to comparatively slow memory technology and large data transfer latencies in the interconnection network. In this paper, we propose a novel hardware caching technique, called switch cache, to improve the remote memory access performance of CC-NUMA multiprocessors. The main idea is to implement small fast caches in crossbar switches of the interconnect medium to capture and store shared data as they flow from the memory module to the requesting processor. This stored data acts as a cache for subsequent requests, thus reducing the need for remote memory accesses tremendously. The implementation of a cache in a crossbar switch needs to be efficient and robust, yet flexible for changes in the caching protocol. The design and implementation details of a CAche Embedded Switch ARchitecture, CAESAR, using wormhole routing with virtual channels is presented. We explore the design space of switch caches by modeling CAESAR in a detailed execution driven simulator and analyze the performance benefits. Our results show that the CAESAR switch cache is capable of improving the performance of CC-NUMA multiprocessors by up to 45 percent reduction in remote memory accesses for some applications. By serving remote read requests at various stages in the interconnect, we observe improvements in execution time as high as 20 percent for these applications. We conclude that switch caches provide a cost-effective solution for designing high performance CC-NUMA multiprocessors. Ravi R. Iyer 0001, Laxmi N. Bhuyan |
IEEE Trans. Computers | 1 |
| 2000 | Impact of CC-NUMA Memory Management Policies on the Application Performance of Multistage Switching NetworksabstractIn this paper, the impact of memory management policies and switch design alternatives on the application performance of cache-coherent nonuniform memory access (CC-NUMA) multiprocessors is studied in detail. Memory management plays an important role in determining the performance of NUMA multiprocessors by dictating the placement of data among the distributed memory modules. We analyze memory traces of several scientific applications for three different memory management techniques, namely buddy, round-robin, and first-touch policies, and compare their memory system performance. Interconnection network switch designs that consider virtual channels and varying number of input buffers per switch are presented. Our performance evaluation is based on execution-driven simulation methodology to capture the dynamic changes in the network traffic during execution of the applications. It is shown that the use of cut-through switching with buffers and virtual channels can Improve the average message latency tremendously. However, the choice of memory management policy affects the amount of network traffic and the network access pattern. Thus, we vary the memory management policy and confirm the performance benefits of improved switch designs. Results of sensitivity studies by varying switch design parameters, cache block size, and memory page size are also presented. We find that a combination of first-touch memory management policy and a switch design with virtual channels and increased buffer space can reduce the average message latency by as high as 70 percent. Laxmi N. Bhuyan, Ravi R. Iyer 0001, Hu-Jun Wang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1999 | Switch Cache: A Framework for Improving the Remote Memory Access Latency of CC-NUMA MultiprocessorsabstractCache coherent non-uniform memory access (CC-NUMA) multiprocessors continue to suffer from remote memory access latencies due to comparatively slow memory technology and data transfer latencies in the interconnection network. We propose a novel hardware caching technique, called switch cache. The main idea is to implement small fast caches in crossbar switches of the interconnect medium to capture and store shared data as they flow from the memory module to the requesting processor. This stored data acts as a cache for subsequent requests, thus reducing the latency of remote memory accesses tremendously. The implementation of a cache in a crossbar switch needs to be efficient and robust, yet flexible for changes in the caching protocol. The design and implementation details of a CAche Embedded Switch ARchitecture, CAESAR, using wormhole routing with virtual channels is presented. Using detailed execution-driven simulations, we find that the CAESAR switch cache is capable of improving the performance of CC-NUMA multiprocessors by reducing the number of reads served at distant remote memories by up to 45% and improving the application execution time by as high as 20%. We conclude that the switch caches provide a cost-effective solution for designing high performance CC-NUMA multiprocessors. Ravi R. Iyer 0001, Laxmi N. Bhuyan |
HPCA | 1 |
| 1999 | Comparing the memory system performance of the HP V-class and SGI Origin 2000 multiprocessors using microbenchmarks and scientific applicationsabstractAs processor technologycontinues to advance at a rapid pace, the principal performance bottleneck of shared memory systems has become the memory access latency.In order to understand the effects of cache and memory hierarchy on system latencies, performance analysts perform benchmark analysis on existing state-of-the-art multiprocessors.In this study, we present a detailed comparison of the memory system of two recent commercial ventures, the HP V-Class and the SGI Origin 2000.Our goal is to compare and contrast design techniques used in these multiprocessors to tolerate the effect of memory latency.Our experimental methodology uses microbenchmarks as well as scientific applications to characterize the user-level performance.Recent Ravi R. Iyer 0001, Nancy M. Amato, Lawrence Rauchwerger, Laxmi N. Bhuyan |
International Conference on Supercomputing | 1 |
| 1997 | Performance of Multistage Bus Networks for a Distributed Shared Memory MultiprocessorabstractA multistage bus network (MEN) is proposed to overcome some of the shortcomings of the conventional multistage interconnection networks (MINs), single bus, and hierarchical bus interconnection networks. The MBN consists of multiple stages of buses connected in a manner similar to the MINs and has the same bandwidth at each stage. A switch in an MBN is similar to that in a MIN switch except that there is a single bus connection instead of a crossbar. MBNs support bidirectional routing and there exists a number of paths between any source and destination pair. The authors develop self routing techniques for the various paths, present an algorithm to route a request along the path with minimum distance, and analyze the probabilities of a packet taking different routes. Further, they derive a performance analysis of a synchronous packet-switched MBN in a distributed shared memory environment and compare the results with those of an equivalent bidirectional MIN (BMIN). Finally, they present the execution time of various applications on the MBN and the BMIN through an execution-driven simulation. They show that the MBN provides similar performance to a BMIN while offering simplicity in hardware and more fault-tolerance than a conventional MIN. Laxmi N. Bhuyan, Ravi R. Iyer 0001, Tahsin Askar, Ashwini K. Nanda, Mohan Kumar |
IEEE Trans. Parallel Distributed Syst. | 2 |