Robert S. Schreiber

dblp:16/2028 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0002-3057-5820ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 since 2021Software engineering, systems software and programming languages · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Memory systems · 42% Parallel and multicore computing · 19% Energy-efficient computing · 11%
Software engineering, system software, and programming languages
2 papers
Program analysis · 90% Compilers and program optimization · 10%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program analysis
data flow analysis
0.322012
On a Technique for Transparently Empowering Classical Compiler Optimizations on Multithreaded Code · ACM Trans. Program. Lang. Syst. 2012
A technique for the effective and automatic reuse of classical compiler optimizations on multithreaded code · POPL 2011
Parallel and multicore computing
parallel programming models
0.222013
On a Technique for Transparently Empowering Classical Compiler Optimizations on Multithreaded Code · ACM Trans. Program. Lang. Syst. 2012
Presto: distributed machine learning and graph processing with sparse matrices · EuroSys 2013
Parallel and multicore computing
shared-memory parallel programs
0.222012
On a Technique for Transparently Empowering Classical Compiler Optimizations on Multithreaded Code · ACM Trans. Program. Lang. Syst. 2012
A technique for the effective and automatic reuse of classical compiler optimizations on multithreaded code · POPL 2011
Distributed systems
distributed data processing
0.212013
Presto: distributed machine learning and graph processing with sparse matrices · EuroSys 2013
Memory systems › non-volatile memory
multi-level cell
0.212013
Practical nonvolatile multilevel-cell phase change memory · SC 2013
Memory systems
non-volatile memory
0.212013
Practical nonvolatile multilevel-cell phase change memory · SC 2013
Memory systems › non-volatile memory
phase change memory
0.212013
Practical nonvolatile multilevel-cell phase change memory · SC 2013
Program analysis › data flow analysis
parallel dataflow analysis
0.112012
On a Technique for Transparently Empowering Classical Compiler Optimizations on Multithreaded Code · ACM Trans. Program. Lang. Syst. 2012
Memory systems
DRAM
0.112012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Energy-efficient computing › power management
memory power management
0.112012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Memory systems › DRAM › DRAM microarchitecture
rank subsetting
0.112012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Interconnection networks and networks-on-chip › routing algorithms
adaptive routing
0.112009
HyperX: topology, routing, and packaging of efficient large-scale networks · SC 2009
Energy-efficient computing
memory energy efficiency
0.112009
Future scaling of processor-memory interfaces · SC 2009
Interconnection networks and networks-on-chip
network topology
0.112009
HyperX: topology, routing, and packaging of efficient large-scale networks · SC 2009
Memory systems › memory interface
processor-memory interface
0.112009
Future scaling of processor-memory interfaces · SC 2009
Electronic design automation › physical design
routing
0.112009
HyperX: topology, routing, and packaging of efficient large-scale networks · SC 2009
Parallel and multicore computing › data-parallel programming
data-parallel frameworks
0.012013
Presto: distributed machine learning and graph processing with sparse matrices · EuroSys 2013
Memory systems › emerging memory technologies
resistance drift
0.012013
Practical nonvolatile multilevel-cell phase change memory · SC 2013
Storage systems
storage reliability
0.012013
Practical nonvolatile multilevel-cell phase change memory · SC 2013
Compilers and program optimization
compiler optimization
0.012012
On a Technique for Transparently Empowering Classical Compiler Optimizations on Multithreaded Code · ACM Trans. Program. Lang. Syst. 2012
Hardware reliability and fault tolerance › error-correcting codes for memory
chipkill correct
0.012012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Hardware reliability and fault tolerance
memory reliability
0.012012
Improving System Energy Efficiency with Memory Rank Subsetting · ACM Trans. Archit. Code Optim. 2012
Processor architecture and microarchitecture
chip multiprocessor
0.012009
Future scaling of processor-memory interfaces · SC 2009

Methods — techniques the papers use, named apart from their topics

siloing · 0.3simulation · 0.2wearout tolerance · 0.2sparse matrix computation · 0.2encoding/decoding · 0.2data-flow transformation · 0.1data flow transformation · 0.1topology search · 0.1process technology analysis · 0.1
YearPublicationVenuePosition
2023 Wafer-Scale Fast Fourier Transforms
abstract
We have implemented fast Fourier transforms for one, two, and three-dimensional arrays on the Cerebras CS-2, a system whose memory and processing elements reside on a single silicon wafer. The wafer-scale engine (WSE) encompasses a two-dimensional mesh of roughly 850,000 processing elements (PEs) with fast local memory and equally fast nearest-neighbor interconnections.
Marcelo Orenes-Vera, Ilya Sharapov, Robert S. Schreiber, Mathias Jacquelin, Philippe Vandermersch, Sharan Chetlur
ICS3
2013 Presto: distributed machine learning and graph processing with sparse matrices
abstract
It is cumbersome to write machine learning and graph algorithms in data-parallel models such as MapReduce and Dryad. We observe that these algorithms are based on matrix computations and, hence, are inefficient to implement with the restrictive programming and communication interface of such frameworks.
Shivaram Venkataraman, Erik Bodzsar, Indrajit Roy 0001, Alvin AuYoung, Robert S. Schreiber
EuroSys5
2013 Practical nonvolatile multilevel-cell phase change memory
abstract
Multilevel-cell (MLC) phase change memory (PCM) may provide both high capacity main memory and faster-than-Flash persistent storage. But slow growth in cell resistance with time, resistance drift, can cause transient errors in MLC-PCM. Drift errors increase with time, and prior work suggests refresh before the cell loses data. The need for refresh makes MLC-PCM volatile, taking away a key advantage. Based on the observation that most drift errors occur in a particular state in four-level-cell PCM, we propose to change from four levels to three levels, eliminating the most vulnerable state. This simple change lowers cell drift error rates by many orders of magnitude: three-level-cell PCM can retain data without power for more than ten years. With optimized encoding/decoding and a wearout tolerance mechanism, we can narrow the capacity gap between three-level and four-level cells. These techniques together enable low-cost, high-performance, genuinely nonvolatile MLC-PCM.
Doe Hyun Yoon, Jichuan Chang, Robert S. Schreiber, Norman P. Jouppi
SC3
2012 Improving System Energy Efficiency with Memory Rank Subsetting
abstract
VLSI process technology scaling has enabled dramatic improvements in the capacity and peak bandwidth of DRAM devices. However, current standard DDR x DIMM memory interfaces are not well tailored to achieve high energy efficiency and performance in modern chip-multiprocessor-based computer systems. Their suboptimal performance and energy inefficiency can have a significant impact on system-wide efficiency since much of the system power dissipation is due to memory power. New memory interfaces, better suited for future many-core systems, are needed. In response, there are recent proposals to enhance the energy efficiency of main-memory systems by dividing a memory rank into subsets, and making a subset rather than a whole rank serve a memory request. We holistically assess the effectiveness of rank subsetting from system-wide performance, energy-efficiency, and reliability perspectives. We identify the impact of rank subsetting on memory power and processor performance analytically, compare two promising rank-subsetting proposals, Multicore DIMM and mini-rank, and verify our analysis by simulating a chip-multiprocessor system using multithreaded and consolidated workloads. We extend the design of Multicore DIMM for high-reliability systems and show that compared with conventional chipkill approaches, rank subsetting can lead to much higher system-level energy efficiency and performance at the cost of additional DRAM devices. This holistic assessment shows that rank subsetting offers compelling alternatives to existing processor-memory interfaces for future DDR systems.
Jung Ho Ahn, Norman P. Jouppi, Christoforos E. Kozyrakis, Jacob Leverich, Robert S. Schreiber
ACM Trans. Archit. Code Optim.5
2012 On a Technique for Transparently Empowering Classical Compiler Optimizations on Multithreaded Code
abstract
A large body of data-flow analyses exists for analyzing and optimizing sequential code. Unfortunately, much of it cannot be directly applied on parallel code, for reasons of correctness. This article presents a technique to automatically, aggressively, yet safely apply sequentially-sound data-flow transformations, without change , on shared-memory programs. The technique is founded on the notion of program references being “siloed” on certain control-flow paths. Intuitively, siloed references are free of interference from other threads within the confines of such paths. Data-flow transformations can, in general, be unblocked on siloed references. The solution has been implemented in a widely used compiler. Results on benchmarks from SPLASH-2 show that performance improvements of up to 41% are possible, with an average improvement of 6% across all the tested programs over all thread counts.
Pramod G. Joisha, Robert S. Schreiber, Prithviraj Banerjee, Hans-Juergen Boehm, Dhruva R. Chakrabarti
ACM Trans. Program. Lang. Syst.2
2011 The runtime abort graph and its application to software transactional memory optimization
abstract
Programming with atomic sections is a promising alternative to locks since it raises the abstraction and removes deadlocks at the programmer level. However, implementations of atomic sections using software transactional memory (STM) support have significant bookkeeping overheads. Additionally, because of the speculative nature of transactions, aborts can be frequent greatly lowering application performance. Thus regardless of the STM implementation, tools need to be available to programmers that provide insights into the runtime characteristics of an application as well as provide means to improve performance. This paper attempts to identify the source of an abort at the granularity of a transactional memory reference. The resulting abort patterns are captured in the form of a runtime abort graph (RAG). We show how to build this graph efficiently using compiler instrumentation. We then describe a technique that works on the RAG and automatically recommends STM policy changes to improve performance. Detailed experimental results are presented showing the tradeoffs in building the RAG and its use in reducing aborts and improving performance.
Dhruva R. Chakrabarti, Prithviraj Banerjee, Hans-Juergen Boehm, Pramod G. Joisha, Robert S. Schreiber
CGO5
2011 A technique for the effective and automatic reuse of classical compiler optimizations on multithreaded code
abstract
A large body of data-flow analyses exists for analyzing and optimizing sequential code. Unfortunately, much of it cannot be directly applied on parallel code, for reasons of correctness. This paper presents a technique to automatically, aggressively, yet safely apply sequentially-sound data-flow transformations, without change, on shared-memory programs. The technique is founded on the notion of program references being "siloed" on certain control-flow paths. Intuitively, siloed references are free of interference from other threads within the confines of such paths. Data-flow transformations can, in general, be unblocked on siloed references.
Pramod G. Joisha, Robert S. Schreiber, Prithviraj Banerjee, Hans-Juergen Boehm, Dhruva R. Chakrabarti
POPL2
2009 HyperX: topology, routing, and packaging of efficient large-scale networks
abstract
In the push to achieve exascale performance, systems will grow to over 100,000 sockets, as growing cores-per-socket and improved single-core performance provide only part of the speedup needed. These systems will need affordable interconnect structures that scale to this level. To meet the need, we consider an extension of the hypercube and flattened butterfly topologies, the HyperX, and give an adaptive routing algorithm, DAL. HyperX takes advantage of high-radix switch components that integrated photonics will make available. Our main contributions include a formal descriptive framework, enabling a search method that finds optimal HyperX configurations; DAL; and a low cost packaging strategy for an exascale HyperX. Simulations show that HyperX can provide performance as good as a folded Clos, with fewer switches. We also describe a HyperX packaging scheme that reduces system cost. Our analysis of efficiency, performance, and packaging demonstrates that the HyperX is a strong competitor for exascale networks.
Jung Ho Ahn, Nathan L. Binkert, Al Davis, Moray McLaren, Robert S. Schreiber
SC5
2009 Future scaling of processor-memory interfaces
abstract
Continuous evolution in process technology brings energy-efficiency and reliability challenges, which are harder for memory system designs since chip multiprocessors demand high bandwidth and capacity, global wires improve slowly, and more cells are susceptible to hard and soft errors. Recently, there are proposals aiming at better main-memory energy efficiency by dividing a memory rank into subsets.
Jung Ho Ahn, Norman P. Jouppi, Christoforos E. Kozyrakis, Jacob Leverich, Robert S. Schreiber
SC5