EDBT 2026 Demo / reviewers in the wild / expert
Linus Y. Wong
dblp:339/8967 · also Yuk Wong 0001
· DBLP profile ↗
5ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0001-9171-3054ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiLFS: FPGA-Orchestrated File System for High-Level SynthesisabstractField Programmable Gate Arrays (FPGAs) deliver high performance, and High-Level Synthesis (HLS) simplifies computation description. However, modern FPGA systems cannot directly exploit the convenience and advanced features of contemporary file systems to enable performant, secure, and robust access to high-speed storage devices such as SSDs. This limitation significantly impedes the adoption of FPGAs in data- and I/O-intensive applications, including large language models (LLMs). Existing HLS storage solutions typically either rely on host CPUs to manage file systems via the operating system stack or provide only low-level block access, both of which introduce considerable performance and programmability overheads. Host-mediated access to storage incurs additional latency due to multiple round-trips through the OS kernel on the host CPU, while block-level management on the FPGA side requires substantial engineering efforts that often require recreating file system functionality, such as raw block management, security, and robustness guarantees. These challenges substantially complicate FPGA development and create a 100× gap in scalability compared to GPUs for deploying modern, large-scale machine learning models. To close this gap, we propose HiLFS, the first file system and storage stack for HLS that manages storage entirely within the FPGA. HiLFS exposes a POSIX-like file interface to HLS kernels to ease programming and maintains an on-chip cache of recently accessed file metadata to accelerate file access. It also provides rich file system features, including data integrity, crash consistency, durability guarantees, and efficient concurrent access. As such, HiLFS enables high-performance, secure, and reliable storage management, completely eliminating the need for host intervention. We prototype HiLFS on an AMD/Xilinx Alveo U200 FPGA with a Solidigm DC-P4610 SSD. On Mixtral 8×7B, HiLFS outperforms Nvidia Titan RTX by 1.1/1.3× in performance and 3.0/3.5× in energy efficiency with/without GPUDirect Storage. To the best of our knowledge, this represents the largest-scale LLM deployment on an FPGA to date. Moreover, HiLFS delivers 1.5/1.8× average latency and bandwidth improvements over state-of-the-art commercial CPU-centric storage platforms with/without PCIe P2P, while incurring 13% bandwidth and latency overhead to state-of-the-art HLS low-level block storage works. Furthermore, HiLFS reduces the LoC by 1.5× and 5.3× compared to CPU-centric and block-level storage platforms, respectively. YoungSeok Na, Linus Y. Wong, André DeHon, Jing Jane Li |
FPGA | 2 |
| 2026 | R2d2: Robotized Reconfigurable Network for Disaggregated Datacenters
Linus Y. Wong, Justin R. Yu, Zhilei Zheng, Jonathan M. Smith, André DeHon |
ISCA | 1 |
| 2024 | DONGLE 2.0: Direct FPGA-Orchestrated NVMe Storage for HLSabstractRapid growth in data size poses significant computational and memory challenges to data processing. FPGA accelerators and near-storage processing have emerged as compelling solutions for tackling the growing computational and memory requirements. Many FPGA-based accelerators have shown to be effective in processing large data sets by leveraging the storage capability of either host-attached or FPGA-attached storage devices. However, the current HLS development environment does not allow direct access to host-or FPGA-attached NVMe storage from the HLS code. As such, users must frequently hand off between HLS and host code to access data in storage, and such a process requires tedious programming to ensure functional correctness. Moreover, since the HLS code uses radically different methods to access storage compared to DRAM, the HLS codebase targeting DRAM-based platforms cannot be easily ported to NVMe-based platforms, resulting in limited code portability and reusability. Furthermore, frequent suspension of HLS kernel and synchronization between CPU and FPGA introduce significant latency overhead and require sophisticated scheduling mechanisms to hide latency. To address these challenges, we propose a new HLS storage interface named DONGLE 2.0 that enables direct FPGA-orchestrated NVMe storage access. By providing a unified interface for storage and memory access, DONGLE 2.0 allows a single-source HLS program to target multiple memory/storage devices, thus making the codebase cleaner, portable, and more efficient. DONGLE 2.0 is an extension to DONGLE 1.0 [ 1 ] but adds support for host-attached storage. While its primary focus is still on FPGA NVMe access in near-storage configurations, the added host storage support ensures its compatibility with platforms that lack native support for FPGA-attached NVMe storage. We implemented a prototype of DONGLE 2.0 using an AMD/Xilinx Alveo U200 FPGA and Solidigm DC-P4610 SSD. Our evaluation on various workloads showed a geometric mean speed-up of 2.3× and a reduction in lines of code (LoC) by 2.4× compared to the state-of-the-art commercial platform when using FPGA-attached NVMe storage. Moreover, DONGLE 2.0 demonstrated a geometric mean speed-up of 1.5× and a reduction in LoC by 2.4× compared to the state-of-the-art commercial platform when using host-attached NVMe storage. Linus Y. Wong, Jing Jane Li |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2023 | DONGLE: Direct FPGA-Orchestrated NVMe Storage for HLSabstractRapid growth in data size poses increasing computational and memory challenges to data processing. FPGA accelerators and near-storage processing are promising candidates for tackling computational and memory requirements, and many near-storage FPGA accelerators have been shown to be effective in processing large data. However, the current HLS development environment does not allow direct NVMe storage access from the HLS code. As such, users must frequently hand off between HLS and host code to access data in storage, and such a process requires tedious programming to ensure functional correctness. Moreover, since the HLS code uses radically different methods to access storage compared to DRAM, the HLS codebase targeting DRAM-based platforms cannot be easily ported to NVMe-based platforms, resulting in limited code portability and reusability. Furthermore, frequent suspension of HLS kernel and synchronization between CPU and FPGA introduce significant latency overhead and require sophisticated scheduling mechanisms to hide latency. Linus Y. Wong, Jing Jane Li |
FPGA | 1 |
| 2019 | Robust Molecular Dynamics Simulations Using Coded FFT AlgorithmabstractAs error/failure rates in supercomputers are projected to grow, computationally intensive scientific applications that lever-age large-scale parallelization will suffer from the increased error rate. In this work, we apply "coded computing" to protein folding simulations in an error-prone environment. We implemented the fast Fourier Poisson method for solving electrostatic equations at each time step of the simulation, and we utilize coded FFT algorithm to protect the compute-intensive FFT algorithm from soft errors. Through experiments on Amazon AWS, we showed that coded protein folding can be implemented with less than 10% overhead in total simulation time, and also showed that coded computing approach is faster than classical checkpointing method when the error rate is high. Linus Y. Wong, Yuqiu Zhang, Haewon Jeong, Pulkit Grover |
ICASSP | 1 |