EDBT 2026 Demo / reviewers in the wild / expert
Jared Casper
dblp:23/6529
· DBLP profile ↗
12ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorArtificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Parallel and multicore computing · 62% Hardware accelerators and domain-specific architectures · 8% GPUs and heterogeneous computing · 6% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 50% Speech recognition and synthesis · 25% Deep learning architectures and training · 25% | |
| Software engineering, system software, and programming languages
3 papers |
Concurrent programming · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Database system architecture and tuning · 50% Indexing and storage engines · 50% |
Topics — the 29 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
0.5 | 1 | 2021 | Efficient large-scale language model training on GPU clusters using megatron-LM · SC 2021 |
Parallel and multicore computing › parallelization strategies
model parallelism |
0.5 | 1 | 2021 | Efficient large-scale language model training on GPU clusters using megatron-LM · SC 2021 |
Parallel and multicore computing › parallel computing › parallel machine learning
parallel training |
0.5 | 1 | 2021 | Efficient large-scale language model training on GPU clusters using megatron-LM · SC 2021 |
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Parallel and multicore computing
transactional memory |
0.2 | 3 | 2011 | Hardware acceleration of transactional memory on commodity systems · ASPLOS 2011 An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007 A practical FPGA-based framework for novel CMP research · FPGA 2007 |
Concurrent programming
concurrent data structures |
0.2 | 2 | 2010 | A practical concurrent binary search tree · PPoPP 2010 Transactional predication: high-performance concurrent sets and maps for STM · PODC 2010 |
Indexing and storage engines
columnar storage |
0.2 | 1 | 2014 | Hardware acceleration of database operations · FPGA 2014 |
Database system architecture and tuning
main-memory database |
0.2 | 1 | 2014 | Hardware acceleration of database operations · FPGA 2014 |
Hardware accelerators and domain-specific architectures › database accelerator
database operation accelerator |
0.2 | 1 | 2014 | Hardware acceleration of database operations · FPGA 2014 |
Concurrent programming › concurrency control
optimistic concurrency control |
0.2 | 2 | 2010 | A practical concurrent binary search tree · PPoPP 2010 A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007 |
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster |
0.1 | 1 | 2021 | Efficient large-scale language model training on GPU clusters using megatron-LM · SC 2021 |
Concurrent programming › concurrent data structures
concurrent binary search tree |
0.1 | 1 | 2010 | A practical concurrent binary search tree · PPoPP 2010 |
Concurrent programming › transactional memory
software transactional memory |
0.1 | 1 | 2010 | Transactional predication: high-performance concurrent sets and maps for STM · PODC 2010 |
Concurrent programming › concurrent data structures
transactional data structures |
0.1 | 1 | 2010 | Transactional predication: high-performance concurrent sets and maps for STM · PODC 2010 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Concurrent programming
transactional memory |
0.1 | 1 | 2007 | A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007 |
Memory systems › cache coherence
directory-based coherence |
0.1 | 1 | 2007 | A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007 |
Memory systems › shared memory
distributed shared memory |
0.1 | 1 | 2007 | A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007 |
Parallel and multicore computing › transactional memory
hybrid transactional memory |
0.1 | 1 | 2007 | An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007 |
Parallel and multicore computing
synchronization |
0.1 | 1 | 2007 | An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2004 | The Vector-Thread Architecture · ISCA 2004 |
Processor architecture and microarchitecture › vector processor
vector-thread architecture |
0.0 | 1 | 2004 | The Vector-Thread Architecture · ISCA 2004 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2007 | A practical FPGA-based framework for novel CMP research · FPGA 2007 |
Electronic design automation
hardware emulation |
0.0 | 1 | 2007 | A practical FPGA-based framework for novel CMP research · FPGA 2007 |
Processor architecture and microarchitecture
multicore design |
0.0 | 1 | 2007 | An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007 |
Processor architecture and microarchitecture › multithreading
multithreaded core |
0.0 | 1 | 2007 | An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007 |
Parallel and multicore computing › multiprocessor system › distributed-memory multiprocessor
NUMA systems |
0.0 | 1 | 2007 | A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007 |
Embedded and real-time systems › embedded processor
low-power embedded processor |
0.0 | 1 | 2004 | The Vector-Thread Architecture · ISCA 2004 |
Methods — techniques the papers use, named apart from their topics
tensor parallelism · 1.0pipeline parallelism · 1.0interleaved pipelining · 1.0data parallelism · 1.0hardware acceleration · 0.5batch dispatch · 0.5GPU-based inference · 0.5software transactional memory · 0.1software transactional memory techniques · 0.1optimistic synchronization · 0.1transactional memory · 0.1performance evaluation · 0.1FPGA prototyping · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Efficient large-scale language model training on GPU clusters using megatron-LMabstractLarge language models have led to state-of-the-art accuracies across several tasks. However, training these models efficiently is challenging because: a) GPU memory capacity is limited, making it impossible to fit large models on even a multi-GPU server, and b) the number of compute operations required can result in unrealistically long training times. Consequently, new methods of model parallelism such as tensor and pipeline parallelism have been proposed. Unfortunately, naive usage of these methods leads to scaling issues at thousands of GPUs. In this paper, we show how tensor, pipeline, and data parallelism can be composed to scale to thousands of GPUs. We propose a novel interleaved pipelining schedule that can improve throughput by 10+% with memory footprint comparable to existing approaches. Our approach allows us to perform training iterations on a model with 1 trillion parameters at 502 petaFLOP/s on 3072 GPUs (per-GPU throughput of 52% of theoretical peak). Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, Matei Zaharia |
SC | 3 |
| 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and MandarinabstractWe show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale. Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu |
ICML | 7 |
| 2014 | Hardware acceleration of database operationsabstractAs the amount of memory in database systems grows, entire database tables, or even databases, are able to fit in the system's memory, making in-memory database operations more prevalent. This shift from disk-based to in-memory database systems has contributed to a move from row-wise to columnar data storage. Furthermore, common database workloads have grown beyond online transaction processing (OLTP) to include online analytical processing and data mining. These workloads analyze huge datasets that are often irregular and not indexed, making traditional database operations like joins much more expensive. Jared Casper, Kunle Olukotun |
FPGA | 1 |
| 2011 | Hardware acceleration of transactional memory on commodity systemsabstractThe adoption of transactional memory is hindered by the high overhead of software transactional memory and the intrusive design changes required by previously proposed TM hardware. We propose that hardware to accelerate software transactional memory (STM) can reside outside an unmodified commodity processor core, thereby substantially reducing implementation costs. This paper introduces Transactional Memory Acceleration using Commodity Cores (TMACC), a hardware-accelerated TM system that does not modify the processor, caches, or coherence protocol. Jared Casper, Tayo Oguntebi, Sungpack Hong, Nathan Bronson, Christoforos E. Kozyrakis, Kunle Olukotun |
ASPLOS | 1 |
| 2010 | FARM: A Prototyping Environment for Tightly-Coupled, Heterogeneous ArchitecturesabstractComputer architectures are increasingly turning to parallelism and heterogeneity as solutions for boosting performance in the face of power constraints. As this trend continues, the challenges of simulating and evaluating these architectures have grown. Hardware prototypes provide deeper insight into these systems when compared to simulators, but are traditionally more difficult and costly to build. We present the Flexible Architecture Research Machine (FARM), a hardware prototyping system based on an FPGA coherently connected to a multiprocessor system. FARM substantially reduces the difficulty and cost of building hardware prototypes by providing a ready-made framework for communicating with a custom design on the FPGA. FARM ensures efficient, low-latency communication with the FPGA via a variety of mechanisms, allowing a wide range of applications to effectively utilize the system. FARM's coherent FPGA includes a cache and participates in coherence activities with the processors. This tight coupling allows for realistic, innovative architecture prototypes that would otherwise be extremely difficult to simulate. We evaluate FARM by providing the reader with a profile of the overheads introduced across the full range of communication mechanisms. This will guide the potential FARM user towards an optimal configuration when designing his prototype. Tayo Oguntebi, Sungpack Hong, Jared Casper, Nathan Bronson, Christoforos E. Kozyrakis, Kunle Olukotun |
FCCM | 3 |
| 2010 | Transactional predication: high-performance concurrent sets and maps for STMabstractConcurrent collection classes are widely used in multi-threaded programming, but they provide atomicity only for a fixed set of operations. Software transactional memory (STM) provides a convenient and powerful programming model for composing atomic operations, but concurrent collection algorithms that allow their operations to be composed using STM are significantly slower than their non-composable alternatives. Nathan Bronson, Jared Casper, Hassan Chafi, Kunle Olukotun |
PODC | 2 |
| 2010 | A practical concurrent binary search treeabstractWe propose a concurrent relaxed balance AVL tree algorithm that is fast, scales well, and tolerates contention. It is based on optimistic techniques adapted from software transactional memory, but takes advantage of specific knowledge of the the algorithm to reduce overheads and avoid unnecessary retries. We extend our algorithm with a fast linearizable clone operation, which can be used for consistent iteration of the tree. Experimental evidence shows that our algorithm outperforms a highly tuned concurrent skip list for many access patterns, with an average of 39% higher single-threaded throughput and 32% higher multi-threaded throughput over a range of contention levels and operation mixes. Nathan Bronson, Jared Casper, Hassan Chafi, Kunle Olukotun |
PPoPP | 2 |
| 2007 | ATLAS: a chip-multiprocessor with transactional memory supportabstractChip-multiprocessors are quickly becoming popular in embedded systems. However, the practical success of CMPs strongly depends on addressing the difficulty of multithreaded application development for such systems. Transactional memory (TM) promises to simplify concurrency management in multithreaded applications by allowing programmers to specify coarse-grain parallel tasks, while achieving performance comparable to fine-grain lock-based applications. This paper presents ATLAS, the first prototype of a CMP with hardware support for transactional memory. ATLAS includes 8 embedded PowerPC cores that access coherent shared memory in a transactional manner. The data cache for each core is modified to support the speculative buffering and conflict detection necessary for transactional execution. The authors have mapped ATLAS to the BEE2 multi-FPGA board to create a full-system prototype that operates at 100MHz, boots Linux, and provides significant performance and ease-of-use benefits for a range of parallel applications. Overall, the ATLAS prototype provides an excellent framework for further research on the software and hardware techniques necessary to deliver on the potential of transactional memory Njuguna Njoroge, Jared Casper, Sewook Wee, Yuriy Teslyar, Daxia Ge, Christoforos E. Kozyrakis, Kunle Olukotun |
DATE | 2 |
| 2007 | A practical FPGA-based framework for novel CMP researchabstractChip-multiprocessors are quickly gaining momentum in all segments of computing. However, the practical success of CMPs strongly depends on addressing the difficulty of multithreaded application development. To address this challenge, it is necessary to co-develop new CMP architecture with novel programming models. Currently, architecture research relies on software simulators which are too slow to facilitate interesting experiments with CMP software without using small datasets or significantly reducing the level of detail in the simulated models. An alternative to simulation is to exploit the rich capabilities of modern FPGAs to create FPGA-based platforms for novel CMP research. This paper presents ATLAS, the first prototype for CMPs with hardware support for Transactional Memory (TM), a technology aiming to simplify parallel programming. ATLAS uses the BEE2 multi-FPGA board to provide a system with 8 PowerPC cores that run at 100MHz and runs Linux. ATLAS provides significant benefits for CMP research such as 100x performance improvement over a software simulator and good visibility that helps with software tuning and architectural improvements. In addition to presenting and evaluating ATLAS, we share our observations about building a FPGA-based framework for CMP research. Specifically, we address issues such as overall performance, challenges of mapping ASIC-style CMP RTL on to FPGAs, software support, the selection criteria for the base processor, and the challenges of using pre-designed IP libraries. Sewook Wee, Jared Casper, Njuguna Njoroge, Yuriy Teslyar, Daxia Ge, Christoforos E. Kozyrakis, Kunle Olukotun |
FPGA | 2 |
| 2007 | A Scalable, Non-blocking Approach to Transactional MemoryabstractTransactional memory (TM) provides mechanisms that promise to simplify parallel programming by eliminating the need for locks and their associated problems (deadlock, livelock, priority inversion, convoying). For TM to be adopted in the long term, not only does it need to deliver on these promises, but it needs to scale to a high number of processors. To date, proposals for scalable TM have relegated livelock issues to user-level contention managers. This paper presents the first scalable TM implementation for directory-based distributed shared memory systems that is livelock free without the need for user-level intervention. The design is a scalable implementation of optimistic concurrency control that supports parallel commits with a two-phase commit protocol, uses write-back caches, and filters coherence messages. The scalable design is based on transactional coherence and consistency (TCC), which supports continuous transactions and fault isolation. A performance evaluation of the design using both scientific and enterprise benchmarks demonstrates that the directory-based TCC design scales efficiently for NUMA systems up to 64 processors Hassan Chafi, Jared Casper, Brian D. Carlstrom, Austen McDonald, Chi Cao Minh, Woongki Baek, Christoforos E. Kozyrakis, Kunle Olukotun |
HPCA | 2 |
| 2007 | An effective hybrid transactional memory system with strong isolation guaranteesabstractWe propose signature-accelerated transactional memory (SigTM), a hybrid TM system that reduces the overhead of software transactions. SigTM uses hardware signatures to track the read-set and write-set for pending transactions and perform conflict detection between concurrent threads. All other transactional functionality, including data versioning, is implemented in software. Unlike previously proposed hybrid TM systems, SigTM requires no modifications to the hardware caches, which reduces hardware cost and simplifies support for nested transactions and multithreaded processor cores. SigTM is also the first hybrid TM system to provide strong isolation guarantees between transactional blocks and nontransactional accesses without additional read and write barriers in non-transactional code. Using a set of parallel programs that make frequent use of coarsegrain transactions, we show that SigTM accelerates software transactions by 30 % to 280%. For certain workloads, SigTM can match the performance of a full-featured hardware TM system, while for workloads with large read-sets it can be up to two times slower. Overall, we show that SigTM combines the performance characteristics and strong isolation guarantees of hardware TM implementations with the low cost and flexibility of software TM systems. Chi Cao Minh, Martin Trautmann, JaeWoong Chung, Austen McDonald, Nathan Bronson, Jared Casper, Christoforos E. Kozyrakis, Kunle Olukotun |
ISCA | 6 |
| 2004 | The Vector-Thread ArchitectureabstractThe vector-thread (VT) architectural paradigm unifies the vector and multithreaded compute models. The VT abstraction provides the programmer with a control processor and a vector of virtual processors (VPs). The control processor can use vector-fetch commands to broadcast instructions to all the VPs or each VP can use thread-fetches to direct its own control flow. A seamless intermixing of the vector and threaded control mechanisms allows a VT architecture to flexibly and compactly encode application parallelism and locality, and a VT machine exploits these to improve performance and efficiency. We present SCALE, an instantiation of the VT architecture designed for low-power and high-performance embedded systems. We evaluate the SCALE prototype design using detailed simulation of a broad range of embedded applications and show that its performance is competitive with larger and more complex processors. Ronny Krashinsky, Christopher Batten, Mark Hampton, Steve Gerding, Brian Pharris, Jared Casper, Krste Asanovic |
ISCA | 6 |