VLDB 2026 Research / reviewers in the wild / expert
Bizhao Shi
dblp:258/0232
· DBLP profile ↗
16ranked-venue papers
3as first author
16since 2021 · last 2025
0000-0002-8665-1132ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 15 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TACPlace: Ultrafast Thermal-Aware Chiplet Placement with Feasibility Seeking
Xinming Wei, Bizhao Shi, Guojie Luo |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | TensTFM: Efficient Total Focusing Method for Ultrasonic Array Imaging on Dataflow AcceleratorsabstractThe Total Focusing Method (TFM) is a highresolution ultrasonic imaging technique widely adopted in nondestructive testing (NDT). However, the computational intensity and memory access demands of TFM hinder its real-time deployment, particularly on conventional computing devices such as CPUs and GPUs. This paper presents TensTFM, a novel TFM acceleration framework optimized for dataflow architectures. Firstly, we characterize the bottlenecks of existing GPU and FPGA solutions and demonstrate the potential of spatially distributed processing on dataflow accelerators. Secondly, we formulate the mapping of TFM operators onto 2D Tensix core arrays as a constrained optimization problem, and propose a simulated annealing-based strategy for efficient mapping space exploration. Thirdly, we introduce tensorized and pipelined implementations for key tasks at the operator level, including Hilbert transform and pixel-wise delay-and-sum interpolation. Finally, experiments on the Tenstorrent Wormhole architecture show that TensTFM can achieve up to$7.9 \times$throughput improvements and$32.5 \times$energy efficiency gain over the optimized GPU baselines. It can also achieve$3.1 \times$throughput improvements compared to the state-of-the-art specialized FPGA accelerators, while offering strong scalability in various imaging configurations. Jieran Zhang, Bizhao Shi, Guojie Luo |
ICCD | 2 |
| 2024 | G2PM: Performance Modeling for ACAP Architecture with Dual-Tiered Graph Representation LearningabstractPerformance estimation is a crucial component in the optimization processes of accelerator development on the Versal ACAP architecture. However, existing approaches present limitations - they are either too slow to facilitate efficient iterations, or they lack the necessary accuracy due to the specific AIE array architecture and two-level programming model of Versal ACAP. To tackle this challenge, we propose G2PM, a performance modeling technique based on a hierarchical graph representation centered on the AIE array. More specifically, we employ a hierarchical graph neural network to identify features of both kernel programs and dataflow programs, taking into account the hardware and software characteristics of the Versal ACAP architecture. In our evaluations, our method demonstrates significant improvements, achieving a mean error rate of less than 1.6% and providing a speed-up factor of 4165X compared to the simulation-based method. Tuo Dai, Bizhao Shi, Guojie Luo |
DAC | 2 |
| 2024 | PT-Map: Efficient Program Transformation Optimization for CGRA MappingabstractCoarse-Grained Reconfigurable Array (CGRA) is a parallel architecture providing high energy efficiency and spatial-temporal re-configurability. Beyond loop scheduling for throughput optimization, program transformation is also crucial in CGRA mapping to optimize overall performance and efficiency. However, existing studies on program transformation optimization face challenges in exploring the transformation space systematically and evaluating candidates efficiently, leading to sub-optimal results. To tackle these challenges, this paper introduces PT-Map, an efficient program transformation optimization framework for CGRA mapping. PT-Map defines a comprehensive transformation space and employs a CGRA-specialized top-down exploration approach. It also incorporates a bottom-up evaluation scheme using architectural parameters and a graph neural network-based predictive model. Experiments demonstrate that PT-Map achieves up to 2.95X/1.80X speedups and 59.0%/23.2% energy-delay-product (EDP) reductions over the state-of-the-art approaches MapZero and PBP, respectively. Bizhao Shi, Tuo Dai, Jiaxi Zhang 0001, Xuechao Wei, Guojie Luo |
DAC | 1 |
| 2024 | WideSA: A High Array Utilization Mapping Scheme for Uniform Recurrences on ACAPabstractThe Versal Adaptive Compute Acceleration Platform (ACAP) is a new architecture that combines AI Engines (AIEs) with reconfigurable fabric. This architecture offers significant acceleration potential for uniform recurrences in various domains, such as deep learning, high-performance computation, and signal processing. However, efficiently mapping these computations onto the Versal ACAP architecture while achieving high utilization of AIEs poses a challenge. To address this issue, we propose a mapping scheme called WideSA, which aims to accelerate uniform recurrences on the Versal ACAP architecture by leveraging the features of both the hardware and the computations. Considering the array architecture of AIEs, our approach utilizes space-time transformations based on the polyhedral model to generate legally optimized systolic array mappings. Concurrently, we have developed a routing-aware PLIO assignment algorithm tailored for communication on the AlE array, and the algorithm aims at successful compilation while maximizing array utilization. Furthermore, we introduce an automatic mapping framework. This framework is designed to generate the corresponding executable code for uniform recurrences, which encompasses the AlE kernel program, programmable logic bitstreams, and the host program. The experimental results validate the effectiveness of our mapping scheme. Specifically, when applying our scheme to matrix multiplication computations on the VCK5000 board, we achieve a throughput of 4.15TOPS on float data type, which is 1.11 x higher compared to the state-of-the-art accelerator on the Versal ACAP architecture. Tuo Dai, Bizhao Shi, Guojie Luo |
DATE | 2 |
| 2024 | BESWAC: Boosting Exact Synthesis via Wiser SAT Solver CallabstractSAT-based exact synthesis is a critical technique in logic synthesis to generate optimal circuits for given Boolean functions. The lengthy trial-and-error process limits its application in on-the-fly logic optimization and optimal netlist library construction. Previous research focuses on reducing the execution time of each trial. However, unnecessary SAT solver calls and varying execution times among encoding methods remained issues. This paper presents BESWAC to boost exact synthesis from the flow level. It leverages initial value prediction, encoding method selection, and an optional early exit to call SAT solvers efficiently and wisely. Moreover, BESWAC can seamlessly integrate existing acceleration methods focusing on individual trials. Experimental results show that BESWAC achieves a 1.79x speedup compared to state-of-the-art exact synthesis flows. Sunan Zou, Jiaxi Zhang 0001, Bizhao Shi, Guojie Luo |
DATE | 3 |
| 2024 | ImageMap: Enabling Efficient Mapping from Image Processing DSL to CGRA
Bizhao Shi, Tuo Dai, Sunan Zou, Xinming Wei, Guojie Luo |
Euro-Par (1) | 1 |
| 2024 | MuSA: Multi-Sketch Accelerator with Hybrid Parallelism and Coalesced Memory OrganizationabstractSketch algorithms are crucial for data stream analysis, offering one-pass processing, sub-linear storage, and accuracy-performance balance. FPGA-based sketch accelerator helps sketch algorithms keep up with modern network inter-connections' speed. However, deploying and optimizing multiple sketches simultaneously is not widely considered, leaving a vast optimization space untouched. This paper introduces MuSA, a multi-sketch FPGA accelerator that exploits hybrid parallelism during sketch maintenance and coalesced memory organization for merging different sketch states. MuSA supports FIFO merging and architecture-specific parameter selection for hybrid parallelism, reducing memory consumption and enabling more considerable parallelism. Evaluation results validate MuSA's effectiveness, with a 15.2 x kernel performance enhancement compared to the state-of-the-art method, enabling on-the-fly high-speed network measurement and high-velocity database analysis. Sunan Zou, Bizhao Shi, Guojie Luo |
ICCD | 2 |
| 2024 | Weave: Abstraction and Integration Flow for Accelerators of Generated ModulesabstractIn modern times, domain-specific accelerators require numerous functional components to execute complex applications in a particular domain. To ensure efficient development, the conventional approach involves decomposing, implementing, and integrating modules. Over the past decade, the generator-based method has proven to enhance the productivity of module implementation. However, current abstractions pose challenges for integrating modules implemented by generators, due to implicit interface definitions, nonunified performance modeling, and fragmented memory management. These limitations result in a lower productivity of the integration process and decreased performance of the integrated accelerators. To overcome these drawbacks, we propose Weave, an abstraction for integrating generated modules and an agile design flow for domain-specific accelerators. The Weave abstraction guides module implementation and integration with a unified performance model and memory management. Furthermore, the Weave integration flow, consisting of generation, selection, and integration phrases, enables optimization of the performance of the integrated accelerator with a design space exploration algorithm and hierarchical memory management. In the experiments, the accelerator developed by Weave achieves$1.93\times $higher performance in the deep learning domain compared to an open-source accelerator, and the integrated accelerator maintains performance for various applications with different memory access patterns. Tuo Dai, Bizhao Shi, Guojie Luo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | PowerSyn: A Logic Synthesis Framework With Early Power OptimizationabstractPower is a great concern in integrated circuits (ICs) design flow, especially in portable devices. As an early stage in electronic design automation (EDA), logic synthesis can significantly affect the quality of the design. It is essential to optimize power in logic synthesis. However, logic synthesis only has a limited concern in power due to its inaccurate estimation. This is because critical physical information is missing at this stage. Furthermore, the empirical optimization sequences need enhancement, and they are not optimal for power, while optimizing power in the early stage is effective. Technology mapping can also improve power optimization with comprehensive power metrics in this sub-15 nm era. Therefore, we propose PowerSyn, a logic synthesis framework with early power optimization. It consists of a practical power model, a power-oriented logic optimization module, and a technology mapping stage. The power model leverages probability propagation considering glitches and static power. The acrlong RL-based logic optimization generates high-quality and rapid-convergence command sequence with early power optimization. We also modify traditional technology mapping with novel power-related metrics. We evaluate PowerSyn on the EPFL benchmark suite. Experiment results show that our flow achieves an average power savings of 16.1% compared to the state-of-the-art open-source logic optimization flow. It also delivers an 8.8% and a 2.1% reduction in latency and area, respectively. The flow incurs less than 12.2% execution time overhead during inference for command generation. Sunan Zou, Jiaxi Zhang 0001, Bizhao Shi, Guojie Luo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Weave: Abstraction for Accelerator Integration of Generated ModulesabstractAs domain-specific accelerators demand multiple functional components for complex applications in a domain, the conventional wisdom for effective development involves module decomposition, module implementation, and module integration. In the recent decade, the generator-based design methodology improves the productivity of module implementation. However, with the guidance of current abstractions, it is difficult to integrate modules implemented by generators because of implicit interface definition, non-unified performance modeling, and fragmented memory management. These disadvantages cause low productivity of the integration flow and low performance of the integrated accelerators. Tuo Dai, Bizhao Shi, Guojie Luo |
FPGA | 2 |
| 2023 | RF-SIFTER: Sifting Signals at Layer-0.5 to Mitigate Wideband Cross-Technology Interference for IoTabstractIoT uplink performance is crucial for a wide variety of IoT applications such as health sensing and industrial control, which demand reliable delivery of sensor data to the cloud. However, due to the limited transmission power budget imposed on many power-constrained IoT devices, IoT uplinks are highly susceptible to cross-technology interference (CTI) caused by coexisting networks. Previous approaches to mitigating CTI have relied on MAC/PHY designs. They suffer from poor performance and limited generality in the presence of wideband CTI sources such as Wi-Fi and RF jammer, which transmit aggressively on large spectrum chunks using diverse radio technologies. Xiong Wang 0006, Jun Huang 0001, Bizhao Shi, Zhe Ou, Guojie Luo, Linghe Kong, Daqing Zhang 0001, Chenren Xu |
MobiCom | 3 |
| 2023 | Efficient Super-Resolution System With Block-Wise Hybridization and Quantized Winograd on FPGAabstractSuper-resolution (SR) techniques aim to restore a high-resolution (HR) image from low-resolution (LR) images, which are often used to assist the enhancement of image/video quality under the rapid development of HR and high-frame-rate media. Recently, neural network (NN)-based methods perform much better image reconstruction quality than classical approaches. However, the unacceptable computation complexity as well as the huge memory footprints of NNs limit the throughputs and scalability of these SR systems. In this work, we analyze several key issues in the design of NN-based SR systems first. Then, we propose a three-level systematic optimization methodology for SR systems to reduce computation overhead and keep image quality. At the algorithm level, we introduce image blocking to SR tasks and develop a block-wise SR algorithm based on the hybrid of NN and interpolation with a consistent image block evaluation metric. The configurable hybrid parameters help the SR algorithm to achieve a flexible tradeoff between the computation overhead and image quality. At the operator level, we focus on the transpose convolution operators commonly used for upsampling in SR NNs. We propose an efficient Winograd-based transposed convolution acceleration method. Through the efficient subconvolutions conversion and the Winograd specialization, this methods enables unified Winograd transformations and simplified data access patterns. At the data level, we propose a novel quantization method for Winograd-aware SR NNs to get better-quantized accuracy. Comprehensive evaluations demonstrate the effectiveness of these optimizations. Our SR system reduces a large number of multiplications with great scalability and supports 4K@120 fps and 8K@30 fps outputs with acceptable image quality degradation. Bizhao Shi, Jiaxi Zhang 0001, Zhuolun He, Xuechao Wei, Sicheng Li 0001, Guojie Luo, Hongzhong Zheng, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | EasyMAC: Design Exploration-Enabled Multiplier-Accumulator Generator Using a Canonical Architectural Representation: (Invited Paper)abstractMultiplier-accumulator (MAC) is a crucial arithmetic element widely used in digital integrated circuits. Customized MACs are necessary for different scenarios but need great effort due to the huge architecture design space. In this paper, we develop EasyMAC, a flexible Chisel-based MAC generator with a canonical architectural representation. We design a compact and canonical sequence representation to express the architecture of MACs. And the MAC generator takes the compact representation as input to gain the Verilog codes. We also give a case study on developing a heuristic design space exploration (DSE) method based on this representation. The experimental result shows the effectiveness of the representation in DSE. Using the percent relative range of the power-delay-area product as a metric to measure the optimization opportunities that this representation exposes, the relative range is 17.4% and 23.1% for 16×16 and 25×18 MACs, respectively. At last, we discuss some promising directions of EasyMAC. Jiaxi Zhang 0001, Qiuyang Gao, Yijiang Guo, Bizhao Shi, Guojie Luo |
ASP-DAC | 4 |
| 2022 | 2022 ICCAD CAD Contest Problem C: Microarchitecture Design Space ExplorationabstractIt is vital to select microarchitectures to achieve good trade-offs between performance, power, and area in the chip development cycle. Combining high-level hardware description languages and optimization of electronic design automation tools empowers microarchitecture exploration at the circuit level. Due to the extremely large design space and high runtime cost to evaluate a microarchitecture, ICCAD 2022 CAD Contest Problem C calls for an effective design space exploration algorithm to solve the problem. We formulate the research topic as a contest problem and provide benchmark suites, contest benchmark platforms, etc., for all contestants to innovate and estimate their algorithms. Sicheng Li 0001, Xuechao Wei, Bizhao Shi, Yen-Kuang Chen, Yuan Xie 0001 |
ICCAD | 4 |
| 2021 | BlockGNN: Towards Efficient GNN Acceleration Using Block-Circulant Weight MatricesabstractIn recent years, Graph Neural Networks (GNNs) appear to be state-of-the-art algorithms for analyzing non-euclidean graph data. By applying deep-learning to extract high-level representations from graph structures, GNNs achieve extraordinary accuracy and great generalization ability in various tasks. However, with the ever-increasing graph sizes, more and more complicated GNN layers, and higher feature dimensions, the computational complexity of GNNs grows exponentially. How to inference GNNs in real time has become a challenging problem, especially for some resource-limited edge-computing platforms.To tackle this challenge, we propose BlockGNN, a software-hardware co-design approach to realize efficient GNN acceleration. At the algorithm level, we propose to leverage block-circulant weight matrices to greatly reduce the complexity of various GNN models. At the hardware design level, we propose a pipelined CirCore architecture, which supports efficient block-circulant matrices computation. Basing on CirCore, we present a novel BlockGNN accelerator to compute various GNNs with low latency. Moreover, to determine the optimal configurations for diverse deployed tasks, we also introduce a performance and resource model that helps choose the optimal hardware parameters automatically. Comprehensive experiments on the ZC706 FPGA platform demonstrate that on various GNN tasks, BlockGNN achieves up to 8.3× speedup compared to the baseline HyGCN architecture and 111.9× energy reduction compared to the Intel Xeon CPU platform. Zhe Zhou 0002, Bizhao Shi, Zhe Zhang 0006, Yijin Guan, Guangyu Sun 0003, Guojie Luo |
DAC | 2 |