Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Blake Hechtman

dblp:323/5386 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 46% Parallel and multicore computing · 27% Distributed systems · 27%
Artificial intelligence
1 paper
Efficient and distributed learning · 87% Language models and text generation · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.712023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.712023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023
Distributed systems › communication optimization
communication-computation overlap
0.712023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023
Parallel and multicore computing
parallel programming models
0.712023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.612022
TPU-KNN: K Nearest Neighbor Search at Peak FLOP/s · NeurIPS 2022
Information retrieval › similarity search
nearest neighbor search
0.612022
TPU-KNN: K Nearest Neighbor Search at Peak FLOP/s · NeurIPS 2022
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.612022
TPU-KNN: K Nearest Neighbor Search at Peak FLOP/s · NeurIPS 2022
Hardware accelerators and domain-specific architectures › tensor accelerator
tensor processing unit
0.612022
TPU-KNN: K Nearest Neighbor Search at Peak FLOP/s · NeurIPS 2022
Natural language and speech › Language models and text generation
large language model
0.212023
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models · ASPLOS (1) 2023

Methods — techniques the papers use, named apart from their topics

intra-layer model parallelism · 1.3computation decomposition · 1.3recall analysis · 1.1performance modeling · 1.1
YearPublicationVenuePosition
2023 Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models
abstract
Large deep learning models have shown great potential with state-of-the-art results in many tasks. However, running these large models is quite challenging on an accelerator (GPU or TPU) because the on-device memory is too limited for the size of these models. Intra-layer model parallelism is an approach to address the issues by partitioning individual layers or operators across multiple devices in a distributed accelerator cluster. But, the data communications generated by intra-layer model parallelism can contribute to a significant proportion of the overall execution time and severely hurt the computational efficiency.
Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Tongfei Guo, Yuanzhong Xu, Zongwei Zhou
ASPLOS (1)6
2022 TPU-KNN: K Nearest Neighbor Search at Peak FLOP/s
abstract
This paper presents a novel nearest neighbor search algorithm achieving TPU (Google Tensor Processing Unit) peak performance, outperforming state-of-the-art GPU algorithms with similar level of recall. The design of the proposed algorithm is motivated by an accurate accelerator performance model that takes into account both the memory and instruction bottlenecks. Our algorithm comes with an analytical guarantee of recall in expectation and does not require maintaining sophisticated index data structure or tuning, making it suitable for applications with frequent updates. Our work is available in the open-source package of Jax and Tensorflow on TPU.
Felix Chern, Blake Hechtman, Andy Davis, David Majnemer, Sanjiv Kumar
NeurIPS2