EDBT 2026 Demo / reviewers in the wild / expert
Shih-Chieh Hsu
dblp:95/11410
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-6214-8500ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | hls4ml: A Flexible, Open Source Platform for Deep Learning Acceleration on Reconfigurable HardwareabstractWe present hls4ml , a free and open source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this article, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results. Jan-Frederik Schulte, Benjamin Ramhorst, Jovan Mitrevski, Nicolò Ghielmetti, Enrico Lupi, Dimitrios Danopoulos, Vladimir Loncar, Javier M. Duarte, David Burnette, Lauri Laatu, Stylianos Tzelepis, Konstantinos Axiotis, Quentin Berthet, Haoyan Wang, Suleyman Demirsoy, Marco Colombo, Thea Aarrestad, Sioni Summers, Maurizio Pierini, Giuseppe Di Guglielmo, Jennifer Ngadiuba, Javier Campos, Benjamin Hawks, Abhijith Gandrakota, Farah Fahim, George A. Constantinides, Zhiqiang Que, Wayne Luk, Alexander D. Tapper, Duc Hoang, Noah Paladino, Philip C. Harris, Bo-Cheng Lai, Manuel Valentin, Ryan Forelli, Seda Ogrenci Memik, Lino Gerlach, Rian Brooks Flynn, Mia Liu, Daniel Diaz 0003, Elham E Khoda, Melissa Quinnan, Russell Solares, Santosh Parajuli, Mark S. Neubauer, Christian Herwig, Ho Fung Tsoi, Dylan S. Rankin, Shih-Chieh Hsu, Scott Hauck |
ACM Trans. Reconfigurable Technol. Syst. | 52 |
| 2025 | BAQET: BRAM-aware Quantization for Efficient Transformer Inference via Stream-based Architecture on an FPGAabstractFPGAs are a compelling substrate for supporting machine learning inference. Tools such as High-Level Synthesis and hls4ml can shorten the development cycle for deploying ML algorithms on FPGAs, but can struggle to handle the large on-chip storage needed for many of these models. In particular the high BRAM usage found in many of these flows can cause Place & Route failures during synthesis. In this paper we propose using a Simulated-Annealing based flow to perform BRAM-aware quantization. This approach trades off inference accuracy with BRAM usage, to provide a high-quality inference engine that still meets on-chip resource constraints. We demonstrate this flow for Transformer-based machine learning algorithms, which include Flash Attention in a Stream-based Dataflow architecture. Our system imposes minimal accuracy drops, yet can reduce BRAM usage by 20%-50%, and improve power efficiency by 264%-812% compared to existing Transformer-based accelerators on FPGAs LingChi Yang, Chi-Jui Chen, Trung Le 0002, Bo-Cheng Lai, Scott Hauck, Shih-Chieh Hsu |
FPGA | 6 |
| 2025 | HiGTR: High-Performance FPGA Implementation of Complete GNN-based Trajectory Reconstruction for HEPabstractCharged particle trajectory reconstruction is a critical task in high-energy physics (HEP), particularly for collision analysis in the Large Hadron Collider (LHC). In the LHC, the Level-1 Trigger (L1T) system must perform trajectory reconstruction with ultralow latency and very high throughput. Graph Neural Networks (GNNs)-based trajectory reconstruction on FPGAs has shown promising performance. However, the existing FPGA-based implementations are incomplete, supporting only the GNN processing stage and lacking the graph construction and track building parts of the task flow. Yun-Chen Yang, Hsuan-Wei Yu, Bo-Cheng Lai, Shih-Chieh Hsu, Mark S. Neubauer, Santosh Parajuli |
FPGA | 4 |
| 2025 | FAIR Universe HiggsML Uncertainty Dataset and CompetitionabstractThe FAIR Universe – HiggsML Uncertainty Challenge focused on measuring the physical properties of elementary particles with imperfect simulators. Participants were required to compute and report confidence intervals for a parameter of interest regarding the Higgs boson while accounting for various systematic (epistemic) uncertainties. The dataset is a tabular dataset of 28 features and 280 million instances. Each instance represents a simulated proton-proton collision as observed at CERN’s Large Hadron Collider in Geneva, Switzerland. The features of these simulations were chosen to capture key characteristics of different types of particles. These include primary attributes, such as the energy and three-dimensional momentum of the particles, as well as derived attributes, which are calculated from the primary ones using domain-specific knowledge. Additionally, a label feature designates each instance’s type of proton-proton collision, distinguishing the Higgs boson events of interest from three background sources. As outlined in this paper, the permanent dataset release allows long-term benchmarking of new techniques. The leading submissions, including Contrastive Normalising Flows and Density Ratios estimation through classification, are described. Our challenge has brought together the physics and machine learning communities to advance our understanding and methodologies in handling systematic uncertainties within AI techniques. Wahid Bhimji, Ragansu Chakkappai, Po-Wen Chang, Yuan-Tang Chou, Sascha Diefenbacher, Jordan Dudley, Ibrahim Elsharkawy, Steven Farrell, Aishik Ghosh, Cristina Giordano, Isabelle Guyon, Christopher J. Harris 0004, Yota Hashizume, Shih-Chieh Hsu, Elham E Khoda, Claudius Krause, Benjamin Nachman, David Rousseau, Robert Schöfbeck, Maryam Shooshtari, Dennis Schwarz, Daohan Wang |
NeurIPS | 14 |
| 2023 | Low Latency Edge Classification GNN for Particle Trajectory Tracking on FPGAsabstractIn-time particle trajectory reconstruction in the Large Hadron Collider is challenging due to the high collision rate and numerous particle hits. Using GNN (Graph Neural Network) on FPGA has enabled superior accuracy with flexible trajectory classification. However, existing GNN architectures have inefficient resource usage and insufficient parallelism for edge classification. This paper introduces a resource-efficient GNN architecture on FPGAs for low latency particle tracking. The modular architecture facilitates design scalability to support large graphs. Leveraging the geometric properties of hit detectors further reduces graph complexity and resource usage. Our results on Xilinx UltraScale+ VU9P demonstrate 1625x and 1574x performance improvement over CPU and GPU respectively. Shi-Yu Huang, Yun-Chen Yang, Yu-Ru Su, Bo-Cheng Lai, Javier M. Duarte, Scott Hauck, Shih-Chieh Hsu, Jin-Xuan Hu, Mark S. Neubauer |
FPL | 7 |
| 1993 | VLSI implementation of an M-array image filter based on shift register array
Chen-Yi Lee, Jer-Min Tsai, Shih-Chieh Hsu |
Integr. | 3 |