EDBT 2026 Demo / reviewers in the wild / expert
Chengyi Lux Zhang
dblp:349/5178
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2024
0009-0007-1785-2065ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Electronic design automation · 38% Reconfigurable computing and FPGAs · 31% Hardware accelerators and domain-specific architectures · 12% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA-accelerated simulation |
0.8 | 1 | 2024 | FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024 |
Reconfigurable computing and FPGAs › FPGA partitioning
multi-FPGA partitioning |
0.8 | 1 | 2024 | FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024 |
Electronic design automation › hardware verification and test › functional verification
pre-silicon verification |
0.8 | 1 | 2024 | FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024 |
Electronic design automation › hardware verification and test › hardware verification
RTL simulation |
0.8 | 1 | 2024 | FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024 |
Electronic design automation › hardware/software co-design
co-simulation |
0.7 | 1 | 2023 | RoSÉ: A Hardware-Software Co-Simulation Infrastructure Enabling Pre-Silicon Full-Stack Robotics SoC Evaluation · ISCA 2023 |
Performance modeling and evaluation
simulation |
0.7 | 1 | 2023 | RoSÉ: A Hardware-Software Co-Simulation Infrastructure Enabling Pre-Silicon Full-Stack Robotics SoC Evaluation · ISCA 2023 |
Cloud and datacenter computing
autoscaling |
0.2 | 1 | 2024 | FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024 |
Reconfigurable computing and FPGAs
cloud FPGA |
0.2 | 1 | 2024 | FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL Designs · ISCA 2024 |
Embedded and real-time systems
cyber-physical systems |
0.2 | 1 | 2023 | RoSÉ: A Hardware-Software Co-Simulation Infrastructure Enabling Pre-Silicon Full-Stack Robotics SoC Evaluation · ISCA 2023 |
Methods — techniques the papers use, named apart from their topics
push-button user-guided partitioning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL DesignsabstractPre-silicon validation and end-to-end system evaluation are integral parts of hardware development as they provide architects with insights about the complex interactions between various hardware components, system software, and application code. Although this process can be accelerated using FPGAs as a simulation host, existing platforms fall short when the resource requirements of a custom hardware design exceed a single FPGA. We present FireAxe, an open-source FPGA-accelerated RTL simulation platform that supports push-button user-guided partitioning across multiple FPGAs, using a compiler called FireRipper. Given a partition point, FireRipper automatically maps a monolithic RTL design onto multiple FPGAs while providing hardware designers quick feedback about the partition interface and expected simulation performance. Furthermore, FireRipper enables users to choose between an exact-mode which provides cycle-exact results with RTL-level fidelity, or a fast-mode that improves simulation rate while sacrificing fidelity only at the partition boundary. Built on FireSim, FireAxe preserves the ability to elastically scale simulations from on-premises FPGAs to cloud FPGAs. For example, pulling out a core from a systemon-chip (SoC) onto a separate FPGA, we achieve simulation rates of 1.6 MHz using on-premises FPGAs connected by direct-attach cables and 1 MHz on AWS F1 FPGAs using peer-to-peer PCIe. To show FireAxe’s ability to enable pre-silicon performance validation at unprecedented scale, we show several case studies. First, we replicate full-stack system-level effects such as latency spikes from garbage collection in a Golang application on an SoC containing 4 out-of-order (OoO) cores. We also boot Linux on, to our knowledge, the largest OoO core ever cycle-exactly simulated in academia. Lastly, we simulate a system-on-chip containing 24 OoO cores mapped onto five datacenter-class FPGAs. We discover an RTL bug when trying to run Linux user-space applications that did not appear with less substantial software stacks. This was discovered in less than 2 hours using FireAxe and would have taken weeks in a commercial software RTL simulator. Joonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Abraham Gonzalez, Raghav Gupta 0001, Nivedha Krishnakumar, Sagar Karandikar, Borivoje Nikolic, Sophia Shao, Krste Asanovic |
ISCA | 3 |
| 2023 | RoSÉ: A Hardware-Software Co-Simulation Infrastructure Enabling Pre-Silicon Full-Stack Robotics SoC EvaluationabstractRobotic systems, such as autonomous unmanned aerial vehicles (UAVs) and self-driving cars, have been widely deployed in many scenarios and have the potential to revolutionize the future generation of computing. To improve the performance and energy efficiency of robotic platforms, significant research efforts are being devoted to developing hardware accelerators for workloads that form bottlenecks in the robotics software pipeline. Although domain-specific accelerators can offer improved efficiency over general-purpose processors on isolated robotics benchmarks, system-level constraints such as data movement and contention over shared resources can significantly impact the achievable end-to-end acceleration. In addition, the closed-loop nature of robotic systems, where there is a tight interaction across different deployed environments, software stacks, and hardware architecture, further exacerbates the difficulties of evaluating robotics SoCs. Dima Nikiforov, Shengjun Chris Dong, Chengyi Lux Zhang, Seah Kim, Borivoje Nikolic, Sophia Shao |
ISCA | 3 |