EDBT 2026 Demo / reviewers in the wild / expert
Jonathan Ross
dblp:95/1492
· DBLP profile ↗
5ranked-venue papers
0as first author
2since 2021 · last 2022
0000-0002-4193-6919ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Hardware accelerators and domain-specific architectures · 60% Interconnection networks and networks-on-chip · 24% Processor architecture and microarchitecture · 10% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.3 | 3 | 2022 | A software-defined tensor streaming multiprocessor for large-scale machine learning · ISCA 2022 Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads · ISCA 2020 In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Hardware accelerators and domain-specific architectures › tensor accelerator
tensor streaming processor |
1.0 | 2 | 2022 | A software-defined tensor streaming multiprocessor for large-scale machine learning · ISCA 2022 Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads · ISCA 2020 |
Interconnection networks and networks-on-chip › network topology › low-diameter topology
dragonfly network |
0.6 | 1 | 2022 | A software-defined tensor streaming multiprocessor for large-scale machine learning · ISCA 2022 |
Interconnection networks and networks-on-chip › network topology
large-scale interconnection networks |
0.6 | 1 | 2022 | A software-defined tensor streaming multiprocessor for large-scale machine learning · ISCA 2022 |
Processor architecture and microarchitecture
dataflow architecture |
0.4 | 1 | 2020 | Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads · ISCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › inference accelerator
neural network inference accelerator |
0.3 | 1 | 2017 | In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Hardware accelerators and domain-specific architectures › tensor accelerator
tensor processing unit |
0.3 | 1 | 2017 | In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Distributed systems
global memory |
0.2 | 1 | 2022 | A software-defined tensor streaming multiprocessor for large-scale machine learning · ISCA 2022 |
Performance modeling and evaluation › benchmarking › computer architecture benchmarking
accelerator benchmarking |
0.1 | 1 | 2017 | In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2017 | In-Datacenter Performance Analysis of a Tensor Processing Unit · ISCA 2017 |
Compilers and program optimization
compiler-hardware co-design |
0.0 | 1 | 2000 | OS and Compiler Considerations in the Design of the IA-64 Architecture · ASPLOS 2000 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 2000 | OS and Compiler Considerations in the Design of the IA-64 Architecture · ASPLOS 2000 |
Methods — techniques the papers use, named apart from their topics
producer-consumer stream programming · 0.9dataflow locality · 0.9software-defined networking · 0.6flow control · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | The Groq Software-defined Scale-out Tensor Streaming Multiprocessor : From chips-to-systems architectural overviewabstractTensor Streaming Processor (TSP) Background Dennis Abts, John Kim 0001, Garrin Kimmell, Matthew Boyd, Kris Kang, Sahil Parmar, Andrew C. Ling, Andrew Bitar, Ibrahim Ahmed 0007, Jonathan Ross |
HCS | 10 |
| 2022 | A software-defined tensor streaming multiprocessor for large-scale machine learningabstractWe describe our novel commercial software-defined approach for large-scale interconnection networks of tensor streaming processing (TSP) elements. The system architecture includes packaging, routing, and flow control of the interconnection network of TSPs. We describe the communication and synchronization primitives of a bandwidth-rich substrate for global communication. This scalable communication fabric provides the backbone for large-scale systems based on a software-defined Dragonfly topology, ultimately yielding a parallel machine learning system with elasticity to support a variety of workloads, both training and inference. We extend the TSP's producer-consumer stream programming model to include global memory which is implemented as logically shared, but physically distributed SRAM on-chip memory. Each TSP contributes 220 MiBytes to the global memory capacity, with the maximum capacity limited only by the network's scale --- the maximum number of endpoints in the system. The TSP acts as both a processing element (endpoint) and network switch for moving tensors across the communication links. We describe a novel software-controlled networking approach that avoids the latency variation introduced by dynamic contention for network links. We describe the topology, routing and flow control to characterize the performance of the network that serves as the fabric for a large-scale parallel machine learning system with up to 10,440 TSPs and more than 2 TeraBytes of global memory accessible in less than 3 microseconds of end-to-end system latency. Dennis Abts, Garrin Kimmell, Andrew C. Ling, John Kim 0001, Matthew Boyd, Andrew Bitar, Sahil Parmar, Ibrahim Ahmed 0007, Roberto DiCecco, Michael Bye, Jennifer Hwang, Jeremy Fowers, Peter Lillian, Ashwin Murthy, Elyas Mehtabuddin, Chetan Tekur, Thomas Sohmers, Kris Kang, Stephen Maresh, Jonathan Ross |
ISCA | 22 |
| 2020 | Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsabstractIn this paper, we introduce the Tensor Streaming Processor (TSP) architecture, a functionally-sliced microarchitecture with memory units interleaved with vector and matrix deep learning functional units in order to take advantage of dataflow locality of deep learning operations. The TSP is built based on two key observations: (1) machine learning workloads exhibit abundant data parallelism, which can be readily mapped to tensors in hardware, and (2) a simple and deterministic processor with producer-consumer stream programming model enables precise reasoning and control of hardware components, achieving good performance and power efficiency. The TSP is designed to exploit parallelism inherent in machine-learning workloads including instruction-level, memory concurrency, data and model parallelism, while guaranteeing determinism by eliminating all reactive elements in the hardware (e.g. arbiters, and caches). Early ResNet50 image classification results demonstrate 20.4K processed images per second (IPS) with a batch-size of one— a $4 \times$ improvement compared to other modern GPUs and accelerators [44]. Our first ASIC implementation of the TSP architecture yields a computational density of more than 1 TeraOp/s per square mm of silicon for its $25 \times 29$ mm 14nm chip operating at a nominal clock frequency of 900 MHz. The TSP demonstrates a novel hardware-software approach to achieve fast, yet predictable, performance on machine-learning workloads within a desired power envelope. Dennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren, Max Baker, Tom Hawkins, Temesghen Kahsai, Garrin Kimmell, Jennifer Hwang, Rebekah Leslie-Hurd, Michael Bye, E. R. Creswick, Matthew Boyd, Mahitha Venigalla, Evan Laforge, Jon Purdy, Purushotham Kamath, Dinesh Maheshwari, Michael Beidler, Geert Rosseel, Omar Ahmad, Gleb Gagarin, Richard Czekalski, Ashay Rane, Sahil Parmar, Jeff Werner, Jim Sproch, Adrián Macías, Brian Kurtz |
ISCA | 2 |
| 2017 | In-Datacenter Performance Analysis of a Tensor Processing UnitabstractMany architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC---called a Tensor Processing Unit (TPU) --- deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN). The heart of the TPU is a 65,536 8-bit MAC matrix multiply unit that offers a peak throughput of 92 TeraOps/second (TOPS) and a large (28 MiB) software-managed on-chip memory. The TPU's deterministic execution model is a better match to the 99th-percentile response-time requirement of our NN applications than are the time-varying optimizations of CPUs and GPUs that help average throughput more than guaranteed latency. The lack of such features helps explain why, despite having myriad MACs and a big memory, the TPU is relatively small and low power. We compare the TPU to a server-class Intel Haswell CPU and an Nvidia K80 GPU, which are contemporaries deployed in the same datacenters. Our workload, written in the high-level TensorFlow framework, uses production NN applications (MLPs, CNNs, and LSTMs) that represent 95% of our datacenters' NN inference demand. Despite low utilization for some applications, the TPU is on average about 15X -- 30X faster than its contemporary GPU or CPU, with TOPS/Watt about 30X -- 80X higher. Moreover, using the CPU's GDDR5 memory in the TPU would triple achieved TOPS and raise TOPS/Watt to nearly 70X the GPU and 200X the CPU. Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson 0001, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, Richard Ho 0001, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, Doe Hyun Yoon |
ISCA | 57 |
| 2000 | OS and Compiler Considerations in the Design of the IA-64 Architecture
Rumi Zahir, Jonathan Ross, Dale Morris, Drew Hess |
ASPLOS | 2 |