VLDB 2026 Research / reviewers in the wild / expert
Rama Govindaraju
dblp:22/1477 · also Rama K. Govindaraju
· DBLP profile ↗
10ranked-venue papers
2as first author
3since 2021 · last 2024
0009-0008-3783-7150ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Silent Data Corruption Errors in VLSI Circuits: Implications, Challenges, and OpportunitiesabstractVLSI chips are the foundation of our computing infrastructure, and we all rely on it to function reliably. Trust of our users and the entire industry is at stake. There is increasing evidence of reliability issues with modern VLSI chips. The defect rates are orders of magnitude higher than what has traditionally been cited. Amplifying this challenge is the point that an increasing number of these chips are silently corrupting the execution context (or SDC - silent data corruption and is inconsistent with the expectation of a failstop model). We will also discuss the challenges emerging from degradation/aging. This discussion will summarize some of the experiences at Google and a sketch of what Google has been doing to address this growing challenge. We will attempt to increase awareness of the growing challenge and also the many opportunities for research to address this problem. The goal will be to make louder the call to action that Google has been championing for the last 5 years to enable an end-to-end solution that addresses this emerging and growing challenge for the entire computing industry. This is an industry wide problem and needs everyone to contribute to enable the solution space. Rama Govindaraju |
ETS | 1 |
| 2023 | Understanding and Mitigating Hardware Failures in Deep Learning Training SystemsabstractDeep neural network (DNN) training workloads are increasingly susceptible to hardware failures in datacenters. For example, Google experienced "mysterious, difficult to identify problems" in their TPU training systems due to hardware failures [7]. Although these particular problems were subsequently corrected through significant efforts, they have raised the urgency of addressing the growing challenges emerging from hardware failures impacting many DNN training workloads. Yi He 0010, Mike Hutton, Robert De Gruijl, Rama Govindaraju, Nishant Patil, Yanjing Li |
ISCA | 5 |
| 2021 | Cores that don't countabstractWe are accustomed to thinking of computers as fail-stop, especially the cores that execute instructions, and most system software implicitly relies on that assumption. During most of the VLSI era, processors that passed manufacturing tests and were operated within specifications have insulated us from this fiction. As fabrication pushes towards smaller feature sizes and more elaborate computational structures, and as increasingly specialized instruction-silicon pairings are introduced to improve performance, we have observed ephemeral computational errors that were not detected during manufacturing tests. These defects cannot always be mitigated by techniques such as microcode updates, and may be correlated to specific components within the processor, allowing small code changes to effect large shifts in reliability. Worse, these failures are often "silent" - the only symptom is an erroneous computation. Peter Hochschild, Jeffrey C. Mogul, Rama Govindaraju, Parthasarathy Ranganathan, David E. Culler, Amin Vahdat |
HotOS | 4 |
| 2019 | Kelp: QoS for Accelerated Machine Learning SystemsabstractDevelopment and deployment of machine learning (ML) accelerators in Warehouse Scale Computers (WSCs) demand significant capital investments and engineering efforts. However, even though heavy computation can be offloaded to the accelerators, applications often depend on the host system for various supporting tasks. As a result, contention on host resources, such as memory bandwidth, can significantly discount the performance and efficiency gains of accelerators. The impact of performance interference is further amplified in distributed learning, which has become increasingly common as model sizes continue to grow. In this work, we study the performance of four production machine learning workloads on three accelerator platforms. Our experiments show that these workloads are highly sensitive to host memory bandwidth contention, which can cause 40% average performance degradation when left unmanaged. To tackle this problem, we design and implement Kelp, a software runtime that isolates high priority accelerated ML tasks from memory resource interference. We evaluate Kelp with both production and artificial aggressor workloads, and compare its effectiveness with previously proposed solutions. Our evaluation shows that Kelp is effective in mitigating performance degradation of the accelerated tasks, and improves performance by 24% on average. Compared to previous work, Kelp reduces performance degradation of ML tasks by 7% and improves system efficiency by 17%. Our results further expose opportunities in future architecture designs. Haishan Zhu, David Lo 0003, Liqun Cheng, Rama Govindaraju, Parthasarathy Ranganathan, Mattan Erez |
HPCA | 4 |
| 2018 | WSMeter: A Performance Evaluation Methodology for Google's Production Warehouse-Scale ComputersabstractEvaluating the comprehensive performance of a warehouse-scale computer (WSC) has been a long-standing challenge. Traditional load-testing benchmarks become ineffective because they cannot accurately reproduce the behavior of thousands of distinct jobs co-located on a WSC. We therefore evaluate WSCs using actual job behaviors in live production environments. From our experience of developing multiple generations of WSCs, we identify two major challenges of this approach: 1) the lack of a holistic metric that incorporates thousands of jobs and summarizes the performance, and 2) the high costs and risks of conducting an evaluation in a live environment. To address these challenges, we propose WSMeter, a cost-effective methodology to accurately evaluate a WSC's performance using a live production environment. We first define a new metric which accurately represents a WSC's overall performance, taking a wide variety of unevenly distributed jobs into account. We then propose a model to statistically embrace the performance variance inherent in WSCs, to conduct an evaluation with minimal costs and risks. We present three real-world use cases to prove the effectiveness of WSMeter. In the first two cases, WSMeter accurately discerns 7% and 1% performance improvements from WSC upgrades using only 0.9% and 6.6% of the machines in the WSCs, respectively. We emphasize that naive statistical comparisons incur much higher evaluation costs (> 4 times) and sometimes even fail to distinguish subtle differences. The third case shows that a cloud customer hosting two services on our WSC quantifies the performance benefits of software optimization (+9.3%) with minimal overheads (2.3% of the service capacity). Changkyu Kim, Liqun Cheng, Rama Govindaraju, Jangwoo Kim |
ASPLOS | 5 |
| 2016 | Improving Resource Efficiency at Scale with HeraclesabstractUser-facing, latency-sensitive services, such as websearch, underutilize their computing resources during daily periods of low traffic. Reusing those resources for other tasks is rarely done in production services since the contention for shared resources can cause latency spikes that violate the service-level objectives of latency-sensitive tasks. The resulting under-utilization hurts both the affordability and energy efficiency of large-scale datacenters. With the slowdown in technology scaling caused by the sunsetting of Moore’s law, it becomes important to address this opportunity. We present Heracles, a feedback-based controller that enables the safe colocation of best-effort tasks alongside a latency-critical service. Heracles dynamically manages multiple hardware and software isolation mechanisms, such as CPU, memory, and network isolation, to ensure that the latency-sensitive job meets latency targets while maximizing the resources given to best-effort tasks. We evaluate Heracles using production latency-critical and batch workloads from Google and demonstrate average server utilizations of 90% without latency violations across all the load and colocation scenarios that we evaluated. David Lo 0003, Liqun Cheng, Rama Govindaraju, Parthasarathy Ranganathan, Christoforos E. Kozyrakis |
ACM Trans. Comput. Syst. | 3 |
| 2015 | Heracles: improving resource efficiency at scaleabstractUser-facing, latency-sensitive services, such as websearch, underutilize their computing resources during daily periods of low traffic. Reusing those resources for other tasks is rarely done in production services since the contention for shared resources can cause latency spikes that violate the service-level objectives of latency-sensitive tasks. The resulting under-utilization hurts both the affordability and energy-efficiency of large-scale datacenters. With technology scaling slowing down, it becomes important to address this opportunity. David Lo 0003, Liqun Cheng, Rama Govindaraju, Parthasarathy Ranganathan, Christoforos E. Kozyrakis |
ISCA | 3 |
| 2014 | Towards energy proportionality for large-scale latency-critical workloadsabstractReducing the energy footprint of warehouse-scale computer (WSC) systems is key to their affordability, yet difficult to achieve in practice. The lack of energy proportionality of typical WSC hardware and the fact that important workloads (such as search) require all servers to remain up regardless of traffic intensity renders existing power management techniques ineffective at reducing WSC energy use. We present PEGASUS, a feedback-based controller that significantly improves the energy proportionality of WSC systems, as demonstrated by a real implementation in a Google search cluster. PEGASUS uses request latency statistics to dynamically adjust server power management limits in a fine-grain manner, running each server just fast enough to meet global service-level latency objectives. In large cluster experiments, PEGASUS reduces power consumption by up to 20%. We also estimate that a distributed version of PEGASUS can nearly double these savings. David Lo 0003, Liqun Cheng, Rama Govindaraju, Luiz André Barroso, Christoforos E. Kozyrakis |
ISCA | 3 |
| 2004 | Architecture and Early Performance of the New IBM HPS Fabric and Adapter
Rama Govindaraju, Peter Hochschild, Don G. Grice, Kevin J. Gildea, Robert Blackmore, Carl A. Bender, Chulho Kim, Piyush Chaudhary, Jason Goscinski, Jay Herring, John Houston |
HiPC | 1 |
| 2001 | MPI-LAPI: An Efficient Implementation of MPI for IBM RS/6000 SP SystemsabstractThe IBM RS/6000 SP system is one of the most cost-effective commercially available high performance machines. IBM RS/6000 SP systems support the Message Passing Interface standard (MPI) and LAPI. LAPI is a low level, reliable and efficient one-sided communication API library implemented on IBM RS/6000 SP systems. This paper explains how the high performance of the LAPI library has been exploited in order to implement the MPI standard more efficiently than the existing MPI. It describes how to avoid unnecessary data copies at both the sending and receiving sides for such an implementation. The resolution of problems arising from the mismatches between the requirements of the MPI standard and the features of LAPI is discussed. As a result of this exercise, certain enhancements to LAPI are identified to enable an efficient implementation of MPI on LAPI. The performance of the new implementation of MPI is compared with that of the underlying LAPI itself. The latency (in polling and interrupt modes) and bandwidth of our new implementation is compared with that of the native MPI implementation on RS/6000 SP systems. The results indicate that the MPI implementation on LAPI performs comparably to or better than the original MPI implementation in most cases. Improvements of up to 17.3 percent in polling mode latency, 35.8 percent in interrupt mode latency, and 20.9 percent in bandwidth are obtained for certain message sizes. The implementation of MPI on top of LAPI also outperforms the native MPI implementation for the NAS Parallel Benchmarks. Mohammad Banikazemi, Rama Govindaraju, Robert Blackmore, Dhabaleswar K. Panda 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |