EDBT 2026 Demo / reviewers in the wild / expert
Jack Weber
dblp:320/9848
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0002-1688-6358ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 61% High-performance computing · 30% Hardware accelerators and domain-specific architectures · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › execution model
decoupled access/execute architecture |
0.6 | 1 | 2022 | A Tensor Processing Framework for CPU-Manycore Heterogeneous Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Processor architecture and microarchitecture
many-core architecture |
0.6 | 1 | 2022 | A Tensor Processing Framework for CPU-Manycore Heterogeneous Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
High-performance computing
tensor computation |
0.6 | 1 | 2022 | A Tensor Processing Framework for CPU-Manycore Heterogeneous Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.2 | 1 | 2022 | A Tensor Processing Framework for CPU-Manycore Heterogeneous Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Methods — techniques the papers use, named apart from their topics
register transfer level modeling · 0.6pytorch · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | A Tensor Processing Framework for CPU-Manycore Heterogeneous SystemsabstractFuture CPU-manycore heterogeneous systems can provide high peak throughput by integrating thousands of simple, independent, energy-efficient cores in a single die. However, there are two key challenges to translating this high peak throughput into improved end-to-end workload performance: 1) manycore co-processors rely on simple hardware putting significant demands on the software programmer and 2) manycore co-processors use in-order cores that struggle to tolerate long memory latencies. To address the manycore programmability challenge, this article presents a dense and sparse tensor processing framework based on PyTorch that enables domain experts to easily accelerate off-the-shelf workloads on CPU-manycore heterogeneous systems. To address the manycore memory latency challenge, we use our extended PyTorch framework to explore the potential for decoupled access/execute (DAE) software and hardware mechanisms. More specifically, we propose two software-only techniques, naïve-software DAE and systolic-software DAE, along with a lightweight hardware access accelerator to further improve area-normalized throughput. We evaluate our techniques using a combination of PyTorch operator microbenchmarking and real-world PyTorch workloads running on a detailed register-transfer-level model of a 128-core manycore architecture. Our evaluation on three real-world dense and sparse tensor workloads suggests these workloads can achieve approximately 2–$6\times $performance improvement when scaled to a future 2000-core CPU-manycore heterogeneous system compared to an 18-core out-of-order CPU baseline, while potentially achieving higher area-normalized throughput and improved energy efficiency compared to general-purpose graphics processing units. Peitian Pan, Zhongyuan Zhao 0004, Krithik Ranjan, Jack Weber, Bandhav Veluri, Seyed Borna Ehsani, Max Ruttenberg, Dai Cheol Jung, Preslav Ivanov, Dustin Richmond, Michael B. Taylor, Zhiru Zhang, Christopher Batten |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |