Jared Casper

dblp:23/6529 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Parallel and multicore computing · 62% Hardware accelerators and domain-specific architectures · 8% GPUs and heterogeneous computing · 6%
Artificial intelligence
2 papers
Efficient and distributed learning · 50% Speech recognition and synthesis · 25% Deep learning architectures and training · 25%
Software engineering, system software, and programming languages
3 papers
Concurrent programming · 100%
Databases, data mining, and information retrieval
1 paper
Database system architecture and tuning · 50% Indexing and storage engines · 50%

Topics — the 29 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.512021
Efficient large-scale language model training on GPU clusters using megatron-LM · SC 2021
Parallel and multicore computing › parallelization strategies
model parallelism
0.512021
Efficient large-scale language model training on GPU clusters using megatron-LM · SC 2021
Parallel and multicore computing › parallel computing › parallel machine learning
parallel training
0.512021
Efficient large-scale language model training on GPU clusters using megatron-LM · SC 2021
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Parallel and multicore computing
transactional memory
0.232011
Hardware acceleration of transactional memory on commodity systems · ASPLOS 2011
An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007
A practical FPGA-based framework for novel CMP research · FPGA 2007
Concurrent programming
concurrent data structures
0.222010
A practical concurrent binary search tree · PPoPP 2010
Transactional predication: high-performance concurrent sets and maps for STM · PODC 2010
Indexing and storage engines
columnar storage
0.212014
Hardware acceleration of database operations · FPGA 2014
Database system architecture and tuning
main-memory database
0.212014
Hardware acceleration of database operations · FPGA 2014
Hardware accelerators and domain-specific architectures › database accelerator
database operation accelerator
0.212014
Hardware acceleration of database operations · FPGA 2014
Concurrent programming › concurrency control
optimistic concurrency control
0.222010
A practical concurrent binary search tree · PPoPP 2010
A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster
0.112021
Efficient large-scale language model training on GPU clusters using megatron-LM · SC 2021
Concurrent programming › concurrent data structures
concurrent binary search tree
0.112010
A practical concurrent binary search tree · PPoPP 2010
Concurrent programming › transactional memory
software transactional memory
0.112010
Transactional predication: high-performance concurrent sets and maps for STM · PODC 2010
Concurrent programming › concurrent data structures
transactional data structures
0.112010
Transactional predication: high-performance concurrent sets and maps for STM · PODC 2010
High-performance computing
performance optimization at scale
0.112016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Concurrent programming
transactional memory
0.112007
A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007
Memory systems › cache coherence
directory-based coherence
0.112007
A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007
Memory systems › shared memory
distributed shared memory
0.112007
A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007
Parallel and multicore computing › transactional memory
hybrid transactional memory
0.112007
An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007
Parallel and multicore computing
synchronization
0.112007
An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007
Parallel and multicore computing
parallel programming models
0.012004
The Vector-Thread Architecture · ISCA 2004
Processor architecture and microarchitecture › vector processor
vector-thread architecture
0.012004
The Vector-Thread Architecture · ISCA 2004
Processor architecture and microarchitecture
chip multiprocessor
0.012007
A practical FPGA-based framework for novel CMP research · FPGA 2007
Electronic design automation
hardware emulation
0.012007
A practical FPGA-based framework for novel CMP research · FPGA 2007
Processor architecture and microarchitecture
multicore design
0.012007
An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007
Processor architecture and microarchitecture › multithreading
multithreaded core
0.012007
An effective hybrid transactional memory system with strong isolation guarantees · ISCA 2007
Parallel and multicore computing › multiprocessor system › distributed-memory multiprocessor
NUMA systems
0.012007
A Scalable, Non-blocking Approach to Transactional Memory · HPCA 2007
Embedded and real-time systems › embedded processor
low-power embedded processor
0.012004
The Vector-Thread Architecture · ISCA 2004

Methods — techniques the papers use, named apart from their topics

tensor parallelism · 1.0pipeline parallelism · 1.0interleaved pipelining · 1.0data parallelism · 1.0hardware acceleration · 0.5batch dispatch · 0.5GPU-based inference · 0.5software transactional memory · 0.1software transactional memory techniques · 0.1optimistic synchronization · 0.1transactional memory · 0.1performance evaluation · 0.1FPGA prototyping · 0.1
YearPublicationVenuePosition
2021 Efficient large-scale language model training on GPU clusters using megatron-LM
abstract
Large language models have led to state-of-the-art accuracies across several tasks. However, training these models efficiently is challenging because: a) GPU memory capacity is limited, making it impossible to fit large models on even a multi-GPU server, and b) the number of compute operations required can result in unrealistically long training times. Consequently, new methods of model parallelism such as tensor and pipeline parallelism have been proposed. Unfortunately, naive usage of these methods leads to scaling issues at thousands of GPUs. In this paper, we show how tensor, pipeline, and data parallelism can be composed to scale to thousands of GPUs. We propose a novel interleaved pipelining schedule that can improve throughput by 10+% with memory footprint comparable to existing approaches. Our approach allows us to perform training iterations on a model with 1 trillion parameters at 502 petaFLOP/s on 3072 GPUs (per-GPU throughput of 52% of theoretical peak).
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, Matei Zaharia
SC3
2016 Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu
ICML7
2014 Hardware acceleration of database operations
abstract
As the amount of memory in database systems grows, entire database tables, or even databases, are able to fit in the system's memory, making in-memory database operations more prevalent. This shift from disk-based to in-memory database systems has contributed to a move from row-wise to columnar data storage. Furthermore, common database workloads have grown beyond online transaction processing (OLTP) to include online analytical processing and data mining. These workloads analyze huge datasets that are often irregular and not indexed, making traditional database operations like joins much more expensive.
Jared Casper, Kunle Olukotun
FPGA1
2011 Hardware acceleration of transactional memory on commodity systems
abstract
The adoption of transactional memory is hindered by the high overhead of software transactional memory and the intrusive design changes required by previously proposed TM hardware. We propose that hardware to accelerate software transactional memory (STM) can reside outside an unmodified commodity processor core, thereby substantially reducing implementation costs. This paper introduces Transactional Memory Acceleration using Commodity Cores (TMACC), a hardware-accelerated TM system that does not modify the processor, caches, or coherence protocol.
Jared Casper, Tayo Oguntebi, Sungpack Hong, Nathan Bronson, Christoforos E. Kozyrakis, Kunle Olukotun
ASPLOS1
2010 FARM: A Prototyping Environment for Tightly-Coupled, Heterogeneous Architectures
abstract
Computer architectures are increasingly turning to parallelism and heterogeneity as solutions for boosting performance in the face of power constraints. As this trend continues, the challenges of simulating and evaluating these architectures have grown. Hardware prototypes provide deeper insight into these systems when compared to simulators, but are traditionally more difficult and costly to build. We present the Flexible Architecture Research Machine (FARM), a hardware prototyping system based on an FPGA coherently connected to a multiprocessor system. FARM substantially reduces the difficulty and cost of building hardware prototypes by providing a ready-made framework for communicating with a custom design on the FPGA. FARM ensures efficient, low-latency communication with the FPGA via a variety of mechanisms, allowing a wide range of applications to effectively utilize the system. FARM's coherent FPGA includes a cache and participates in coherence activities with the processors. This tight coupling allows for realistic, innovative architecture prototypes that would otherwise be extremely difficult to simulate. We evaluate FARM by providing the reader with a profile of the overheads introduced across the full range of communication mechanisms. This will guide the potential FARM user towards an optimal configuration when designing his prototype.
Tayo Oguntebi, Sungpack Hong, Jared Casper, Nathan Bronson, Christoforos E. Kozyrakis, Kunle Olukotun
FCCM3
2010 Transactional predication: high-performance concurrent sets and maps for STM
abstract
Concurrent collection classes are widely used in multi-threaded programming, but they provide atomicity only for a fixed set of operations. Software transactional memory (STM) provides a convenient and powerful programming model for composing atomic operations, but concurrent collection algorithms that allow their operations to be composed using STM are significantly slower than their non-composable alternatives.
Nathan Bronson, Jared Casper, Hassan Chafi, Kunle Olukotun
PODC2
2010 A practical concurrent binary search tree
abstract
We propose a concurrent relaxed balance AVL tree algorithm that is fast, scales well, and tolerates contention. It is based on optimistic techniques adapted from software transactional memory, but takes advantage of specific knowledge of the the algorithm to reduce overheads and avoid unnecessary retries. We extend our algorithm with a fast linearizable clone operation, which can be used for consistent iteration of the tree. Experimental evidence shows that our algorithm outperforms a highly tuned concurrent skip list for many access patterns, with an average of 39% higher single-threaded throughput and 32% higher multi-threaded throughput over a range of contention levels and operation mixes.
Nathan Bronson, Jared Casper, Hassan Chafi, Kunle Olukotun
PPoPP2
2007 ATLAS: a chip-multiprocessor with transactional memory support
abstract
Chip-multiprocessors are quickly becoming popular in embedded systems. However, the practical success of CMPs strongly depends on addressing the difficulty of multithreaded application development for such systems. Transactional memory (TM) promises to simplify concurrency management in multithreaded applications by allowing programmers to specify coarse-grain parallel tasks, while achieving performance comparable to fine-grain lock-based applications. This paper presents ATLAS, the first prototype of a CMP with hardware support for transactional memory. ATLAS includes 8 embedded PowerPC cores that access coherent shared memory in a transactional manner. The data cache for each core is modified to support the speculative buffering and conflict detection necessary for transactional execution. The authors have mapped ATLAS to the BEE2 multi-FPGA board to create a full-system prototype that operates at 100MHz, boots Linux, and provides significant performance and ease-of-use benefits for a range of parallel applications. Overall, the ATLAS prototype provides an excellent framework for further research on the software and hardware techniques necessary to deliver on the potential of transactional memory
Njuguna Njoroge, Jared Casper, Sewook Wee, Yuriy Teslyar, Daxia Ge, Christoforos E. Kozyrakis, Kunle Olukotun
DATE2
2007 A practical FPGA-based framework for novel CMP research
abstract
Chip-multiprocessors are quickly gaining momentum in all segments of computing. However, the practical success of CMPs strongly depends on addressing the difficulty of multithreaded application development. To address this challenge, it is necessary to co-develop new CMP architecture with novel programming models. Currently, architecture research relies on software simulators which are too slow to facilitate interesting experiments with CMP software without using small datasets or significantly reducing the level of detail in the simulated models. An alternative to simulation is to exploit the rich capabilities of modern FPGAs to create FPGA-based platforms for novel CMP research. This paper presents ATLAS, the first prototype for CMPs with hardware support for Transactional Memory (TM), a technology aiming to simplify parallel programming. ATLAS uses the BEE2 multi-FPGA board to provide a system with 8 PowerPC cores that run at 100MHz and runs Linux. ATLAS provides significant benefits for CMP research such as 100x performance improvement over a software simulator and good visibility that helps with software tuning and architectural improvements. In addition to presenting and evaluating ATLAS, we share our observations about building a FPGA-based framework for CMP research. Specifically, we address issues such as overall performance, challenges of mapping ASIC-style CMP RTL on to FPGAs, software support, the selection criteria for the base processor, and the challenges of using pre-designed IP libraries.
Sewook Wee, Jared Casper, Njuguna Njoroge, Yuriy Teslyar, Daxia Ge, Christoforos E. Kozyrakis, Kunle Olukotun
FPGA2
2007 A Scalable, Non-blocking Approach to Transactional Memory
abstract
Transactional memory (TM) provides mechanisms that promise to simplify parallel programming by eliminating the need for locks and their associated problems (deadlock, livelock, priority inversion, convoying). For TM to be adopted in the long term, not only does it need to deliver on these promises, but it needs to scale to a high number of processors. To date, proposals for scalable TM have relegated livelock issues to user-level contention managers. This paper presents the first scalable TM implementation for directory-based distributed shared memory systems that is livelock free without the need for user-level intervention. The design is a scalable implementation of optimistic concurrency control that supports parallel commits with a two-phase commit protocol, uses write-back caches, and filters coherence messages. The scalable design is based on transactional coherence and consistency (TCC), which supports continuous transactions and fault isolation. A performance evaluation of the design using both scientific and enterprise benchmarks demonstrates that the directory-based TCC design scales efficiently for NUMA systems up to 64 processors
Hassan Chafi, Jared Casper, Brian D. Carlstrom, Austen McDonald, Chi Cao Minh, Woongki Baek, Christoforos E. Kozyrakis, Kunle Olukotun
HPCA2
2007 An effective hybrid transactional memory system with strong isolation guarantees
abstract
We propose signature-accelerated transactional memory (SigTM), a hybrid TM system that reduces the overhead of software transactions. SigTM uses hardware signatures to track the read-set and write-set for pending transactions and perform conflict detection between concurrent threads. All other transactional functionality, including data versioning, is implemented in software. Unlike previously proposed hybrid TM systems, SigTM requires no modifications to the hardware caches, which reduces hardware cost and simplifies support for nested transactions and multithreaded processor cores. SigTM is also the first hybrid TM system to provide strong isolation guarantees between transactional blocks and nontransactional accesses without additional read and write barriers in non-transactional code. Using a set of parallel programs that make frequent use of coarsegrain transactions, we show that SigTM accelerates software transactions by 30 % to 280%. For certain workloads, SigTM can match the performance of a full-featured hardware TM system, while for workloads with large read-sets it can be up to two times slower. Overall, we show that SigTM combines the performance characteristics and strong isolation guarantees of hardware TM implementations with the low cost and flexibility of software TM systems.
Chi Cao Minh, Martin Trautmann, JaeWoong Chung, Austen McDonald, Nathan Bronson, Jared Casper, Christoforos E. Kozyrakis, Kunle Olukotun
ISCA6
2004 The Vector-Thread Architecture
abstract
The vector-thread (VT) architectural paradigm unifies the vector and multithreaded compute models. The VT abstraction provides the programmer with a control processor and a vector of virtual processors (VPs). The control processor can use vector-fetch commands to broadcast instructions to all the VPs or each VP can use thread-fetches to direct its own control flow. A seamless intermixing of the vector and threaded control mechanisms allows a VT architecture to flexibly and compactly encode application parallelism and locality, and a VT machine exploits these to improve performance and efficiency. We present SCALE, an instantiation of the VT architecture designed for low-power and high-performance embedded systems. We evaluate the SCALE prototype design using detailed simulation of a broad range of embedded applications and show that its performance is competitive with larger and more complex processors.
Ronny Krashinsky, Christopher Batten, Mark Hampton, Steve Gerding, Brian Pharris, Jared Casper, Krste Asanovic
ISCA6