EDBT 2026 Demo / reviewers in the wild / expert
Zhigang Wei
dblp:81/7248
· DBLP profile ↗
8ranked-venue papers
5as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XPNet: Cross-FPGA Power Prediction From High-Level Language CodeabstractMachine learning (ML) has been successfully employed to estimate power consumption for FPGAs using features derived from the results of High Level Synthesis (HLS). However, such models trained on one FPGA cannot be directly applied to another FPGA, even within the same FPGA series. Training a model for a new FPGA is time-consuming due to the significant effort required for dataset preparation. Researchers have to invest significant effort (weeks) in constructing a sufficient dataset with power value annotations, to train an accurate model for a new target FPGA. Another challenge is that existing model construction methods depend on many features extracted late in the HLS process, which are tool-specific and cannot be transferred between tools from different vendors. To address these challenges, we propose a novel cross-FPGA power modeling methodology called XPNet. With only frontend features from HLS, XPNet combines Transfer-Learning with innovative data selection techniques that enable efficient fine-tuning for a new target FPGA. With XPNet, models trained on one FPGA can be quickly adapted to a new target FPGA and used to efficiently predict the power on this new FPGA with high accuracy. Experiments with Polybench, Machsuite and CHStone demonstrate an average error of only 8.40% (10.34% if cross-vendor) when less than 1% of designs are used for the fine-tuning to the new target FPGA. In comparison to best prior model (with full training and 5.74% error), XPNet yields 232x speed up in dataset preparation and training, and 5x speed up in inference on a new FPGA. Zhigang Wei, Allison Seigler, Sean Lowe, Emily Shriver, Aman Arora 0001, Lizy Kurian John |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | ATAPP: Architecture and Technology Aware Power Predictor for Unseen FPGAS
Zhigang Wei, Aman Arora 0001, Emily Shriver, Lizy Kurian John |
FPL | 1 |
| 2024 | Cross-FPGA Power Estimation from High Level Synthesis via Transfer-LearningabstractMachine learning (ML) has been successfully employed to estimate power consumption for FPGAs using features derived from post High Level Synthesis (HLS). As a result, the power evaluation of the design bypasses time-consuming logic synthesis and implementation. However, such models have noticeable drawbacks. Firstly, the dataset preparation is time-consuming since researchers invest significant effort in constructing a sufficient dataset to train an accurate model for a target FPGA. Secondly, the model trained on one FPGA cannot be directly applied to another. Without prior knowledge about the architecture of the second FPGA, the model's power estimation on this new FPGA is of unknown confidence. To address these challenges, we propose a novel cross-FPGA power modeling methodology called XPNet that combines Transfer-Learning with an innovative data selection technique that enables efficient fine-tuning. We start by applying Transfer-Learning with our data selection methodology to adapt a GNN-based power model to a second FPGA using only 20 data samples, resulting in 6.53% error. We then explore if our approach works for lighter-weight ML-based models, such as multi-layer perception (MLP), and show less than a 1% degradation in accuracy. Additionally, we explore the impact of using Meta-Learning algorithm on our model and show that with only 40 data samples from the target FPGA, the model still manages an error of 6%. Zhigang Wei, Aman Arora 0001, Emily Shriver, Lizy Kurian John |
FPGA | 1 |
| 2023 | HLSDataset: Open-Source Dataset for ML-Assisted FPGA Design using High Level SynthesisabstractMachine Learning (ML) has been widely adopted in design exploration using high level synthesis (HLS) for faster resource, timing and power estimation at very early stages for FPGA-based design. To perform prediction accurately, high-quality and large-volume datasets are required for training ML models. However, the current datasets used in this domain are proprietary or limited in use, and practitioners have to generate their own dataset to train HLS-related ML models. This paper presents a dataset for ML-assisted FPGA design using HLS, called HLSDataset. The dataset is generated from widely used HLS C benchmarks including Polybench, Machsuite, CHStone and Rossetta. The Verilog samples are generated with a variety of directives including loop unroll, loop pipeline, and array partition to make sure optimized and realistic designs are covered. The total number of generated Verilog samples is nearly 9,000 per FPGA type. The dataset repository includes CSV (comma separated values) files containing both HLS and implementation metrics which can be easily consumed by ML model. We also include original C source code with directives, Verilog designs, post-HLS reports, post-implementation reports for each sample in the dataset, so that any metrics not present in the CSV can be easily extracted. In order to extend the dataset for future benchmarks, generation and extraction scripts are also provided. To demonstrate the effectiveness of our dataset, we undertake case studies to perform power estimation and resource usage estimation with ML models trained with our dataset. All the code and dataset are public at our github page11https://github.com/UT-LCAIML4Accel-Dataset/tree/main/fpga_ml_dataset. We believe that HLSDataset can save valuable time for researchers by avoiding the tedious process of running tools, scripting and parsing files to generate the dataset, and enable them to spend more time where it counts, that is, in training ML models. Zhigang Wei, Aman Arora 0001, Ruihao Li 0002, Lizy Kurian John |
ASAP | 1 |
| 2020 | Hamamu: Specializing FPGAs for ML Applications by Adding Hard Matrix Multiplier BlocksabstractDesigning efficient hardware for accelerating artificial intelligence (AI) and machine learning (ML) applications is a major challenge. Rapidly changing algorithms and neural network architectures make FPGA based designs an attractive solution. But the generic building blocks available in current FPGAs (Logic Blocks (LBs), multipliers, DSP blocks) limit the acceleration that can be achieved. We propose Hamamu, a modification to the current FPGA architecture that makes FPGAs specialized for ML applications. Specifically, we propose adding hard matrix multiplier blocks (matmuls) into the FPGA fabric. These matmuls are implemented using systolic arrays of MACs (Multiply-And-Accumulate) and can be connected using programmable direct interconnect between neighboring matmuls to make larger systolic matrix multipliers. We explore various matmul sizes ($2\times 2\times 2$, $4\times 4\times 4$, $8\times 8\times 8$, $16\times 16\times 16$) and various strategies to place these blocks on the FPGA (Columnar, Surround, Hybrid). We find that providing $4\times 4\times 4$ hard matrix multiplier blocks in an FPGA speeds up neural networks from MLPerf benchmarks by up to $\sim 3.9x$, compared to a Stratix-10 like FPGA with equal number of MACs, same MAC architecture and high DSP:LB ratio. Although the flexibility of the FPGA will reduce for non-ML applications, an FPGA with hard matrix multipliers is a faster, and more area efficient hardware accelerator for ML applications, compared to current FPGAs. Aman Arora 0001, Zhigang Wei, Lizy Kurian John |
ASAP | 2 |
| 2020 | Design Space Exploration for Softmax ImplementationsabstractDeep Neural Networks (DNN) are crucial components of machine learning in the big data era. Significant effort has been put into the hardware acceleration of convolution and fully-connected layers of neural networks, while not too much attention has been put on the Softmax layer. Softmax is used in terminal classification layers in networks like ResNet, and is also used in intermediate layers in networks like the Transformer. As the speed for other DNN layers keeps improving, efficient and flexible designs for Softmax are required. With the existence of several ways to implement Softmax in hardware, we evaluate various softmax hardware designs and the trade-offs between them. In order to make the design space exploration more efficient, we also develop a parameterized generator which can produce softmax designs by varying multiple aspects of a base architecture. The aspects or knobs are parallelism, accuracy, storage and precision. The goal of the generator is to enable evaluation of tradeoffs between area, delay, power and accuracy in the architecture of a softmax unit. We simulate and synthesize the generated designs and present results comparing them with the existing state-of-the-art. Our exploration reveals that the design with parallelism of 16 can provide the best area-delay product among designs with parallelism ranging from 1 to 32. It is also observed that look-up table based approximate LOG and EXP units can be used to yield almost the same accuracy as the full LOG and EXP units, while providing area and energy benefits. Additionally, providing local registers for intermediate values is seen to provide energy savings. Zhigang Wei, Aman Arora 0001, Pragenesh Patel, Lizy Kurian John |
ASAP | 1 |
| 2020 | The Case for Hard Matrix Multiplier Blocks in an FPGAabstractDesigning efficient hardware for accelerating machine learning (ML) applications is a major challenge. Rapid changing algorithms and network architectures in this field make FPGA based designs an attractive solution. But the generic building blocks available in current FPGAs (ALMs/CLBs, DSP blocks) limit the acceleration that can be achieved. We propose a modification to the current FPGA architecture that makes FPGAs specialized for ML applications. Specifically, we propose adding hard matrix multiplier blocks (matmuls) into the FPGA fabric. These matmuls are implemented using systolic arrays of MACs (Multiply-And-Accumulate) and can be connected using programmable direct interconnect between neighboring matmuls to make larger systolic matrix multipliers. We explore various matmul sizes (4x4x4, 8x8x8, 16x16x16, 32x32x32) and various strategies to place these blocks on the FPGA (clustered, surround, columnar). We recommend 4x4x4 matmul blocks with columnar placement after studying tradeoffs between area, frequency, fragmentation and channel width. Experimental results and analytical evaluation reveal that providing matmuls in an FPGA speeds up state-of-the-art neural networks (Resnet50, GNMT, Transformer, Minigo) by ~2.5x on average, compared to a DSP-heavy FPGA with equal number of MACs. Therefore, FPGAs with hard matrix multipliers can be used to design faster, more area (and hence, power) efficient hardware accelerators for ML applications, compared to current FPGAs, at the cost of reducing the flexibility of the FPGA for other applications. A matmul-heavy FPGA fabric could be a part of bigger FPGA, the rest of which can have general programmable logic, or fully ML-specific FPGAs with matmuls could be created. Aman Arora 0001, Zhigang Wei, Lizy Kurian John |
FPGA | 2 |
| 2018 | A Novel Fault-Tolerant Last-Level Cache to Improve Reliability at Near-Threshold VoltageabstractNear-threshold voltage computing (NTC) improves power and energy efficiency of cache by scaling transistor voltage. However, in large SRAM structures, such as last-level cache (LLC), a great number of bit-cell errors will occur when supply voltage scales to near-threshold voltage. In this paper, we propose a novel fault-tolerant LLC design (NFTLLC) to deal with a high failure rate which is higher than 1% at near-threshold voltage. NFTLLC corrects the single-error and compresses multi-error in Cache entry to improves the reliability of last-level cache. To validate the efficiency of NFTLLC, we implement NFTLLC and prior works in gem5, and simulate with SPEC CPU2006. The experiment shows that compared with Concertina when bit-cell failure rate is 1.1%, the performance of NFTLLC with 4-byte subblock size improves by 6.8% and the Cache capacity increases by 20.8%. Besides, miss rate decreases more than 53%, and overhead increases by 16.8% in minimum. Wei Liu 0011, Zhigang Wei, Wei Du 0001 |
ACM Great Lakes Symposium on VLSI | 2 |