Bob Brennan

dblp:99/923 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 49% Hardware accelerators and domain-specific architectures · 21% Emerging computing paradigms · 18%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
processing-in-memory
0.622018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
DRISA: a DRAM-based reconfigurable in-situ accelerator · MICRO 2017
Memory systems › processing-in-memory
processing-using-DRAM
0.622018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
DRISA: a DRAM-based reconfigurable in-situ accelerator · MICRO 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.312018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Emerging computing paradigms › approximate and stochastic computing › stochastic computing
stochastic arithmetic
0.312018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Emerging computing paradigms › approximate and stochastic computing
stochastic computing
0.312018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Memory systems
DRAM
0.212016
DRAF: A Low-Power DRAM-Based Reconfigurable Acceleration Fabric · ISCA 2016
Memory systems › memory architecture
memory-centric architecture
0.112018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Energy-efficient computing
power management
0.112018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Memory systems
data movement
0.112017
DRISA: a DRAM-based reconfigurable in-situ accelerator · MICRO 2017
Memory systems
memory wall
0.112017
DRISA: a DRAM-based reconfigurable in-situ accelerator · MICRO 2017
Hardware accelerators and domain-specific architectures › accelerator architecture
datacenter accelerator
0.112016
DRAF: A Low-Power DRAM-Based Reconfigurable Acceleration Fabric · ISCA 2016
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.112016
DRAF: A Low-Power DRAM-Based Reconfigurable Acceleration Fabric · ISCA 2016

Methods — techniques the papers use, named apart from their topics

stochastic computing · 0.3hierarchical hybrid deterministic arithmetic · 0.3configuration context switching · 0.2bit-level reconfigurable logic · 0.2
YearPublicationVenuePosition
2018 SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator
abstract
Memory-centric architecture, which bridges the gap between compute and memory, is considered as a promising solution to tackle the memory wall and the power wall. Such architecture integrates the computing logic and the memory resources close to each other, in order to embrace large internal memory bandwidth and reduce the data movement overhead. The closer the compute and memory resources are located, the greater these benefits become. DRAM-based in-situ accelerators [1] tightly couple processing units to every memory bitline, achieving the maximum benefits among various memory-centric architectures. However, the processing units in such architectures are typically limited to simple functions like AND/OR due to strict area and power overhead constraints in DRAMs, making it difficult to accomplish complex tasks while providing high performance. In this paper, we address the challenge by applying stochastic computing arithmetic to the DRAM-based in-situ accelerator, targeting at the acceleration of error-tolerant applications such as deep learning. In stochastic computing, binary numbers are converted into stochastic bitstreams, which turns integer multiplications into simple bitwise AND operations, but at the expense of larger memory capacity/bandwidth demands. Stochastic computing is a perfect match for the DRAM-based in-situ accelerators because it addresses the in-situ accelerator's low performance problem by simplifying the operations, while leveraging the in-situ accelerator's advantage of large memory capacity/bandwidth. To further boost the performance and compensate for the numerical precision loss, we propose a novel Hierarchical and Hybrid Deterministic (H2D) stochastic computing arithmetic. Finally, we consider quantized deep neural network inference and training applications as a case study. The proposed architecture provides 2.3× improvement in performance per unit area compared with the binary arithmetic baseline, and 3.8× improvement over GPU. The proposed H2D arithmetic contributes 11× performance boost and 60% numerical precision improvement.
Shuangchen Li, Alvin Oliver Glova, Xing Hu 0001, Peng Gu 0008, Dimin Niu, Krishna T. Malladi, Hongzhong Zheng, Bob Brennan, Yuan Xie 0001
MICRO8
2017 DRISA: a DRAM-based reconfigurable in-situ accelerator
abstract
Data movement between the processing units and the memory in traditional von Neumann architecture is creating the "memory wall" problem. To bridge the gap, two approaches, the memory-rich processor (more on-chip memory) and the compute-capable memory (processing-in-memory) have been studied. However, the first one has strong computing capability but limited memory capacity/bandwidth, whereas the second one is the exact the opposite.
Shuangchen Li, Dimin Niu, Krishna T. Malladi, Hongzhong Zheng, Bob Brennan, Yuan Xie 0001
MICRO5
2016 DRAF: A Low-Power DRAM-Based Reconfigurable Acceleration Fabric
abstract
FPGAs are a popular target for application-specific accelerators because they lead to a good balance between flexibility and energy efficiency. However, FPGA lookup tables introduce significant area and power overheads, making it difficult to use FPGA devices in environments with tight cost and power constraints. This is the case for datacenter servers, where a modestly-sized FPGA cannot accommodate the large number of diverse accelerators that datacenter applications need. This paper introduces DRAF, an architecture for bit-level reconfigurable logic that uses DRAM subarrays to implement dense lookup tables. DRAF overlaps DRAM operations like bitline precharge and charge restoration with routing within the reconfigurable routing fabric to minimize the impact of DRAM latency. It also supports multiple configuration contexts that can be used to quickly switch between different accelerators with minimal latency. Overall, DRAF trades off some of the performance of FPGAs for significant gains in area and power. DRAF improves area density by 10x over FPGAs and power consumption by more than 3x, enabling DRAF to satisfy demanding applications within strict power and cost constraints. While accelerators mapped to DRAF are 2-3x slower than those in FPGAs, they still deliver a 13x speedup and an 11x reduction in power consumption over a Xeon core for a wide range of datacenter tasks, including analytics and interactive services like speech recognition.
Mingyu Gao 0001, Christina Delimitrou, Dimin Niu, Krishna T. Malladi, Hongzhong Zheng, Bob Brennan, Christoforos E. Kozyrakis
ISCA6
1986 The C-MU phonetic classification system
abstract
The Carnegie-Mellon Speech Group is working with a number of institutions in the DARPA community to develop "A New Generation English Language System" (ANGELS) to perform large vocabulary speaker-independent recognition of natural continuous speech. A major focus of this effort is the development of an acoustic-phonetic module that provides an accurate phonetic transcription of an unknown utterance. This paper describes the phonetic classification system now under development, the research approach and some preliminary results.
Ron Cole, Mike Phillips, Bob Brennan, Ben Chigier
ICASSP3