EDBT 2026 Demo / reviewers in the wild / expert
Puguang Liu
dblp:218/9556
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-7165-9297ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TL-Sort: A Fully Pipelined Hardware Architecture for Sorting Without Run-Drain Stalls
Hai Cao, Puguang Liu, Zhang Luo, Jihang Wang, Xingyun Qi |
APPT | 2 |
| 2026 | Blocking Is Not Stagnation: A Synchronous FPGA-CPU Architecture for Regular Expression Matching in Real-Time DPIabstractRegular expression matching is a crucial step in traffic analysis. Many hardware-based architectures are proposed to improve the matching throughput, such as FPGA. To date, however, the existing FPGA-CPU architectures are difficult to implement in DPI systems due to the following two reasons. First, existing architectures use asynchronous workflows to interact data between FPGA and CPU, making them difficult to be compatible with synchronous DPI systems. Second, asynchronous architectures require batch input, which does not meet the requirements of real-time environments. In this paper, we concentrate on the real-time deployment of a regular expression matching architecture. To improve the deployment throughput, we propose an FPGA-CPU architecture with a parallel layer between the driver and DPI systems. Then, coroutines are introduced and proved to have significant advantages. Meanwhile, some optimization methods are proposed to address idle time, memory allocation, and MMIO control. Our experiments demonstrate that directly deploying an asynchronous architecture on a synchronous DPI would result in a throughput degradation of 3 orders of magnitude. Our approach enhances throughput by 2-3 orders of magnitude. This indicates that we reach a throughput in synchronous mode that is comparable to that in asynchronous mode, and it is over 10 times faster than the software solution, making the direct deployment of asynchronous architectures on mainstream DPI systems feasible. To the best of our knowledge, this is the first attempt to improve hardware-based regular expression matching under synchronous logic, achieving both high throughput and usability. Shuhui Chen, Ziling Wei, Jincheng Zhong, Puguang Liu |
IEEE Trans. Netw. | 5 |
| 2025 | FPAMM: Fine-Grained Pipeline Architecture Accelerator for the Novel Transformer Architecture - Monarch Mixer
Hanyuan Li, Xingyun Qi, Puguang Liu, Zhenqi Li |
ICA3PP (2) | 4 |
| 2025 | PSCA: A FPGA-based Protein Structure Comparison Accelerator with Symmetric Simplified Matrix
Hui Su, Xingyun Qi, Qiang Wang 0006, Puguang Liu, Haoyu Liao |
ICA3PP (6) | 5 |
| 2025 | Enhancing Transformer Inference Efficiency on FPGA Through Fully Fusion and Integer-Only Quantization TechniquesabstractThe Transformer architecture has revolutionized the field of natural language processing (NLP) through its selfattention mechanism. However, its high computational complexity and memory requirement present significant deployment challenges on resource-constrained edge devices. While existing research predominantly focuses on accelerating linear operations via model compression and approximation techniques, the inefficiencies and high deployment costs of nonlinear operations (e.g., Softmax and LayerNorm) remain critically understudied. Although some studies have attempted to mitigate these challenges through techniques such as kernel fusion and integer-only quantization, these approaches still suffer from partial fusion and inefficient quantization with retained division operations, leaving significant efficiency gains unexploited. To bridge these gaps, we propose a fully fused Transformer accelerator that co-optimizes both linear and nonlinear operations while minimizing memory bottlenecks. For linear computations, our design incorporates a deeply optimized compute engine featuring double buffering, an output-stationary tiling strategy, and DSP-packing technology to maximize throughput. For nonlinear operations, we introduce a delayed computation strategy for vector-wise operators, effectively reducing memory bandwidth pressure and dependency stalls. Furthermore, we propose a hardware-efficient, divisionfree integer-only quantization scheme, leveraging$\log 2$quantization for Softmax and a polynomial-enhanced approximation for LayerNorm to eliminate costly floating-point units, thereby significantly reducing latency and resource overhead. Through systematic design space exploration, our solution, deployed on the Zynq Z-7100 platform, achieves 1.376 TOPS for BERT inference, demonstrating a$\mathbf{1. 6 5 - 2. 5 1} \boldsymbol{\times}$higher computational efficiency compared to prior works. Zhenqi Li, Puguang Liu, Qiang Wang 0006, Yankang Zhao, Hanyuan Li, Xingyun Qi |
ICCD | 4 |
| 2024 | Optimization of TDM Using Single-ended Transmission for Multi-FPGA PlatformsabstractIn large-scale designs, the multi-FPGA system is a popular approach for hardware acceleration and pre-silicon verification due to its scalable capabilities. Researchers aim to enhance communication bandwidth and decrease latency between FPGAs, all while working within the constraints of limited physical pins available. To address this issue, we propose an optimization for time-division multiplexing(TDM). This optimization combines the advantages of traditional logic multiplexer circuits and high-speed serial circuits by utilizing the serial-to-parallel converter (ISERDES) and parallel-to-serial converter (OSERDES). The I/OSERDES approach typically employs the low-voltage differential signaling standard (LVDS) for high-speed data transmission, necessitating two physical pins. To save one physical pin, we advocate for a single-ended transmission solution without a forward clock. Moreover, we propose a new algorithm at the receiver to align the phase of the receiver’s clock. In comparison to the LVDS-based solution, our proposed interface achieves double the communication bandwidth of inter-FPGA chips without introducing additional system latency under the same TDM rate. Haoyu Liao, Puguang Liu, Xingyun Qi |
ISCAS | 3 |
| 2022 | Reducing Network Traffic Storage Overhead: A Hardware-Accelerated Lossless Data Compression SystemabstractSkyrocketing network traffic volume brings significant overhead of traffic storage. We propose a hardware-accelerated lossless data compression system to reduce traffic storage overhead. The prototype evaluation validates the excellent resource efficiency and performance scalability of our traffic compression system. Puguang Liu, Shuhui Chen |
APNet | 1 |
| 2021 | MXQN: Mixed quantization for reducing bit-width of weights and activations in deep convolutional neural networks
Puguang Liu |
Appl. Intell. | 2 |