EDBT 2026 Demo / reviewers in the wild / expert
Jeremy Intan
dblp:278/6193
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2023
0000-0003-1384-992XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 51% Parallel and multicore computing · 29% Processor architecture and microarchitecture · 20% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU computing
GPU synchronization |
0.7 | 1 | 2023 | Improving the Scalability of GPU Synchronization Primitives · IEEE Trans. Parallel Distributed Syst. 2023 |
Processor architecture and microarchitecture
atomic operations |
0.4 | 1 | 2020 | Deterministic Atomic Buffering · MICRO 2020 |
Parallel and multicore computing
deterministic execution |
0.4 | 1 | 2020 | Deterministic Atomic Buffering · MICRO 2020 |
GPUs and heterogeneous computing
GPU architecture |
0.4 | 1 | 2020 | Deterministic Atomic Buffering · MICRO 2020 |
Parallel and multicore computing › synchronization
synchronization mechanisms |
0.2 | 1 | 2023 | Improving the Scalability of GPU Synchronization Primitives · IEEE Trans. Parallel Distributed Syst. 2023 |
Methods — techniques the papers use, named apart from their topics
priority semaphore · 0.7multi-level sense-reversing barrier · 0.7deterministic scheduling · 0.4atomic buffering · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Trireme: Exploration of Hierarchical Multi-level Parallelism for Hardware AccelerationabstractThe design of heterogeneous systems that include domain specific accelerators is a challenging and time-consuming process. While taking into account area constraints, designers must decide which parts of an application to accelerate in hardware and which to leave in software. Moreover, applications in domains such as Extended Reality (XR) offer opportunities for various forms of parallel execution, including loop level, task level, and pipeline parallelism. To assist the design process and expose every possible level of parallelism, we present Trireme , a fully automated tool-chain that explores multiple levels of parallelism and produces domain-specific accelerator designs and configurations that maximize performance, given an area budget. FPGA SoCs were used as target platforms, and Catapult HLS [ 7 ] was used to synthesize RTL using a commercial 12 nm FinFET technology. Experiments on demanding benchmarks from the XR domain revealed a speedup of up to 20×, as well as a speedup of up to 37× for smaller applications, compared to software-only implementations. Georgios Zacharopoulos 0001, Adel Ejjeh, En-Yu Yang, Iulian Brumar, Jeremy Intan, Muhammad Huzaifa, Sarita V. Adve, Vikram S. Adve, Gu-Yeon Wei, David Brooks 0001 |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2023 | Improving the Scalability of GPU Synchronization PrimitivesabstractGeneral-purpose GPU applications increasingly use synchronization to enforce ordering between many threads accessing shared data. Accordingly, recently there has been a push to establish a common set of GPU synchronization primitives. However, the expressiveness of existing GPU synchronization primitives is limited. In particular the expensive GPU atomics often used to implement fine-grained synchronization make it challenging to implement efficient algorithms. Consequently, as GPU algorithms scale to millions or billions of threads, existing GPU synchronization primitives either scale poorly or suffer from livelock or deadlock issues because of heavy contention between threads accessing shared synchronization objects. We seek to overcome these inefficiencies by designing more efficient, scalable GPU barriers and semaphores. In particular, we show how multi-level sense reversing barriers and priority mechanisms for semaphores can be designed with the GPUs unique processing model in mind to improve performance and scalability of GPU synchronization primitives. Our results show that the proposed designs significantly improve performance compared to state-of-the-art solutions like CUDA Cooperative Groups and optimized CPU-style synchronization algorithms at medium and high contention levels, scale to an order of magnitude more threads, and avoid livelock in these situations unlike prior open source algorithms. Overall, across three modern GPUs the proposed barrier algorithm improves performance by an average of 33% over a GPU tree barrier algorithm and improves performance by an average of 34% over CUDA Cooperative Groups for five full-sized benchmarks at high contention levels; the new semaphore algorithm improves performance by an average of 83% compared to prior GPU semaphores. Preyesh Dalmia, Rohan Mahapatra, Jeremy Intan, Dan Negrut, Matthew D. Sinclair |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Deterministic Atomic BufferingabstractDeterministic execution for GPUs is a desirable property as it helps with debuggability and reproducibility. It is also important for safety regulations, as safety critical workloads are starting to be deployed onto GPUs. Prior deterministic architectures, such as GPUDet, attempt to provide strong determinism for all types of workloads, incurring significant performance overheads due to the many restrictions that are required to satisfy determinism. We observe that a class of reduction workloads, such as graph applications and neural architecture search for machine learning, do not require such severe restrictions to preserve determinism. This motivates the design of our system, Deterministic Atomic Buffering (DAB), which provides deterministic execution with low area and performance overheads by focusing solely on ordering atomic instructions instead of all memory instructions. By scheduling atomic instructions deterministically with atomic buffering, the results of atomic operations are isolated initially and made visible in the future in a deterministic order. This allows the GPU to execute deterministically in parallel without having to serialize its threads for atomic operations as opposed to GPUDet. Our simulation results show that, for atomic-intensive applications, DAB performs 4× better than GPUDet and incurs only a 23% slowdown on average compared to a non-deterministic GPU architecture. We also characterize the bottlenecks and provide insights for future optimizations. Yuan-Hsi Chou, Christopher Ng, Shaylin Cattell, Jeremy Intan, Matthew D. Sinclair, Joseph Devietti, Timothy G. Rogers, Tor M. Aamodt |
MICRO | 4 |