Joonho Whangbo

dblp:349/5150 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2024
0009-0008-7570-6733ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Reconfigurable computing and FPGAs · 32% Electronic design automation · 28% Hardware accelerators and domain-specific architectures · 24%
Computer graphics and multimedia
1 paper
Rendering · 100%
Artificial intelligence
1 paper
3D vision · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA-accelerated simulation
0.812024
FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024
Reconfigurable computing and FPGAs › FPGA partitioning
multi-FPGA partitioning
0.812024
FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024
Electronic design automation › hardware verification and test › functional verification
pre-silicon verification
0.812024
FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024
Electronic design automation › hardware verification and test › hardware verification
RTL simulation
0.812024
FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024
Rendering
neural rendering
0.712023
NeuRex: A Case for Neural Rendering Acceleration · ISCA 2023
Hardware accelerators and domain-specific architectures › domain-specific accelerator
compression accelerator
0.712023
CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale Systems · ISCA 2023
Cloud and datacenter computing › datacenter architecture
datacenter tax offload
0.712023
CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale Systems · ISCA 2023
Hardware accelerators and domain-specific architectures
neural rendering accelerator
0.712023
NeuRex: A Case for Neural Rendering Acceleration · ISCA 2023
Cloud and datacenter computing
autoscaling
0.212024
FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024
Reconfigurable computing and FPGAs
cloud FPGA
0.212024
FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024
Computer vision › 3D vision › 3d scene modeling › scene representation
neural scene representation
0.212023
NeuRex: A Case for Neural Rendering Acceleration · ISCA 2023

Methods — techniques the papers use, named apart from their topics

multi-resolution hash encoding · 2.0push-button user-guided partitioning · 0.8lossless compression · 0.7hardware-software co-design · 0.7
YearPublicationVenuePosition
2024 FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs
abstract
Pre-silicon validation and end-to-end system evaluation are integral parts of hardware development as they provide architects with insights about the complex interactions between various hardware components, system software, and application code. Although this process can be accelerated using FPGAs as a simulation host, existing platforms fall short when the resource requirements of a custom hardware design exceed a single FPGA. We present FireAxe, an open-source FPGA-accelerated RTL simulation platform that supports push-button user-guided partitioning across multiple FPGAs, using a compiler called FireRipper. Given a partition point, FireRipper automatically maps a monolithic RTL design onto multiple FPGAs while providing hardware designers quick feedback about the partition interface and expected simulation performance. Furthermore, FireRipper enables users to choose between an exact-mode which provides cycle-exact results with RTL-level fidelity, or a fast-mode that improves simulation rate while sacrificing fidelity only at the partition boundary. Built on FireSim, FireAxe preserves the ability to elastically scale simulations from on-premises FPGAs to cloud FPGAs. For example, pulling out a core from a systemon-chip (SoC) onto a separate FPGA, we achieve simulation rates of 1.6 MHz using on-premises FPGAs connected by direct-attach cables and 1 MHz on AWS F1 FPGAs using peer-to-peer PCIe. To show FireAxe’s ability to enable pre-silicon performance validation at unprecedented scale, we show several case studies. First, we replicate full-stack system-level effects such as latency spikes from garbage collection in a Golang application on an SoC containing 4 out-of-order (OoO) cores. We also boot Linux on, to our knowledge, the largest OoO core ever cycle-exactly simulated in academia. Lastly, we simulate a system-on-chip containing 24 OoO cores mapped onto five datacenter-class FPGAs. We discover an RTL bug when trying to run Linux user-space applications that did not appear with less substantial software stacks. This was discovered in less than 2 hours using FireAxe and would have taken weeks in a commercial software RTL simulator.
Joonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Abraham Gonzalez, Raghav Gupta 0001, Nivedha Krishnakumar, Sagar Karandikar, Borivoje Nikolic, Sophia Shao, Krste Asanovic
ISCA1
2023 CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale Systems
abstract
General-purpose lossless data compression and decompression ("(de)compression") are used widely in hyperscale systems and are key "datacenter taxes". However, designing optimal hardware compression and decompression processing units ("CDPUs") is challenging due to the variety of algorithms deployed, input data characteristics, and evolving costs of CPU cycles, network bandwidth, and memory/storage capacities.
Sagar Karandikar, Aniruddha N. Udipi, Junsun Choi, Joonho Whangbo, Jerry Zhao, Svilen Kanev, Edwin Lim, Jyrki Alakuijala, Vrishab Madduri, Sophia Shao, Borivoje Nikolic, Krste Asanovic, Parthasarathy Ranganathan
ISCA4
2023 NeuRex: A Case for Neural Rendering Acceleration
abstract
This paper presents NeuRex, an accelerator architecture that efficiently performs the modern neural rendering pipeline with an algorithmic enhancement and supporting hardware. NeuRex leverages the insights from an in-depth analysis of the state-of-the-art neural scene representation to make the multi-resolution hash encoding, which is the key operational primitive in modern neural renderings, more hardware-friendly and features a specialized hash encoding engine that enables us to effectively perform the primitive and the overall rendering pipeline. We implement and synthesize NeuRex using a commercial 28nm process technology and evaluate two versions of NeuRex (NeuRex-Edge, NeuRex-Server) on a range of scenes with different image resolutions for mobile and high-end computing platforms. Our evaluation shows that NeuRex achieves up to 9.88× and 3.11× speedups against the mobile and high-end consumer GPUs with a substantially small area overhead and lower energy consumption.
Kwanseok Choi, Joonho Whangbo, Jaewoong Sim
ISCA5