Gary Lauterbach

dblp:97/2163 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 52% Processor architecture and microarchitecture · 29% Hardware accelerators and domain-specific architectures · 8%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
clustered architecture
0.212013
A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O · HPCA 2013
Cloud and datacenter computing › virtualization
i/o virtualization
0.212013
A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O · HPCA 2013
Cloud and datacenter computing
virtualization
0.212013
A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O · HPCA 2013
Energy-efficient computing › power-performance tradeoff
performance-per-watt optimization
0.012013
A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O · HPCA 2013
Memory systems › memory access latency
cache access latency
0.011998
Low Load Latency Through Sum-Addressed Memory (SAM) · ISCA 1998

Methods — techniques the papers use, named apart from their topics

FPGA-based I/O cards · 0.2ASIC-based interconnect fabric · 0.2sum-prediction · 0.0bitwise indexing · 0.0
YearPublicationVenuePosition
2013 A novel system architecture for web scale applications using lightweight CPUs and virtualized I/O
abstract
Large web-scale applications typically use a distributed platform, like clusters of commodity servers, to achieve scalable and low-cost processing. The Map-Reduce framework and its open-source implementation, Hadoop, is commonly used to program these applications. Since these applications scale well with an increased number of servers, the cluster size is an important parameter. Cluster size however is constrained by power consumption. In this paper we present a system that uses low-power CPUs to increase the cluster size in a fixed power budget. Using low-power CPUs leads to the situation where the majority of a server's power is now consumed by the I/O sub-system. To overcome this, we develop a virtualized I/O sub-system where multiple servers share I/O resources. An ASIC based high-bandwidth interconnect fabric, and FPGA based I/O cards implement this virtualized I/O. The resulting system is the first production quality implementation of cluster-in-a-box that uses low-power CPUs. The unique design demonstrates a way to build systems using low-power CPUs, allowing a much larger number of servers in a cluster in the same power envelope. To overcome software inefficiency and increase the utilization of virtualized disk bandwidth, optimizations necessary for the operating system are also discussed. We built hardware based on these ideas and experiments on this system show a 3X average improvement in performance-per-Watt-hour compared to a commodity cluster with the same power budget.
Kshitij Sudan, Saisanthosh Balakrishnan, Sean Lie, Dhiraj Mallick, Gary Lauterbach, Rajeev Balasubramonian
HPCA6
2011 SeaMicro SM10000-64 server: Building datacenter servers using cell phone chips
Ashutosh Dhodapkar, Gary Lauterbach, Sean Lie, Dhiraj Mallick, Jim Bauman, Sundar Kanthadai, Toru Kuzuhara, Gene Shen
Hot Chips Symposium2
1998 Low Load Latency Through Sum-Addressed Memory (SAM)
abstract
Load latency contributes significantly to execution time. Because most cache accesses hit, cache-hit latency becomes an important component of expected load latency. Most modern microprocessors have base+offset addressing loads; thus effective cache-hit latency includes an addition as well as the RAM access. This paper introduces a new technique used in the UltraSPARC III microprocessor Sum-Addressed Memory (SAM), which performs true addition using the decoder of the RAM array, with very low latency. We compare SAM with other methods for reducing the add part of load latency. These methods include sum-prediction with recovery, and bitwise indexing with duplicate-tolerance. The results demonstrate the superior performance of SAM.
William L. Lynch, Gary Lauterbach, Joseph I. Chamdani
ISCA2