EDBT 2026 Demo / reviewers in the wild / expert
Hao Xiao 0001
dblp:67/7741-1
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0004-6419-9984ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A High-Performance and Configurable NTT Accelerator Based on Scalable 2-D ArchitectureabstractThe number theoretic transform (NTT), which can reduce the computational complexity of polynomial multiplication, has been widely used to accelerate cryptographic algorithms. However, due to the significant computational data volume, the limited memory bandwidth of chips limits the effectiveness of massive parallel computation in achieving high performance. In addition, the accelerator must flexibly adapt to variable parameters for diverse encryption scenarios, leading to a nonlinear growth in hardware resources and underutilization of the computation engines. Therefore, this article proposes a scalable 2-D constant-geometry (2-D CG) accelerator designed to fully utilize butterfly units (BFUs) and avoid a drastic increase in hardware resources. The proposed architecture reduces high memory bandwidth requirements by compressing BFUs from the vertical to the horizontal direction, thereby eliminating BFU idle time waiting for data transmission. We then propose a scalable 2-D CG design that performs a constant number of data access patterns within each stage to prevent the dramatic increase in complexity as the architecture scales. In addition, we propose a channel reconfigurable Barrett (CR_Barrett) architecture that computes multiple small-bitwidth modular multiplications (MMs) in parallel, ensuring that all computation engines are actively utilized. Finally, we designed and verified the proposed 2-D CG NTT accelerator on a field-programmable gate array (FPGA). Compared to state-of-the-art works, the 2-D CG achieves$1.47\times $to$7.17\times $higher throughput per slice (TPS) for scalable architectures and$2.02\times $to$6.41\times $higher TPS for runtime configurable architectures. Jianbo Guo, Jiaoyang Zhu, Hao Xiao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | High-Performance FPGA-Based LZ77 Accelerator Design With Deep-Search CapabilityabstractLZ77 is a fundamental primitive for lossless data compression in data-intensive systems. However, conventional hardware accelerators face a fundamental tradeoff between throughput and compression ratio (CR). To overcome this challenge, this brief presents a hardware accelerator tailored for deep-search-based LZ77 compression, achieving a high CR without sacrificing throughput. The proposed design introduces a multidimensional screening strategy built on a reconstructed hash table to aggressively prune invalid nodes and identify a compact set of high-quality candidates prior to high-latency string matching. By condensing the expanded search space induced by deep search, this approach significantly reduces matching complexity while preserving the compression benefits of deep exploration. Architecturally, we design a distributed hash table coupled with a cascaded matrix-simplification pipeline to alleviate memory contention and routing congestion arising from large-scale hash indexing. FPGA implementation achieves a throughput of 4.0GB/s at 250MHz and an average CR of 2.38 on the Calgary corpus. Hao Xiao 0001, Saiqin Xu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | An Efficient Sparse CNN Inference Accelerator With Balanced Intra- and Inter-PE WorkloadabstractSparse convolutional neural networks (SCNNs) which can prune trivial parameters in the network while maintaining the model accuracy has been proved to be an attractive approach to alleviate the heavy computation of convolutional neural networks (CNNs). However, the invalid data resulting from sparse patterns leads to unnecessary and irregular computation workload, which challenges the efficiency of the underlying hardware accelerators. Therefore, this article proposes an SCNN inference accelerator, which can deal with the imbalanced workload both intra- and interprocessing element (PE). A valid weight encoding (VWE) scheme is proposed to compress sparse weights into dense ones to alleviate the load imbalance intra-PE. Leveraging the VWE scheme, a randomized load rearrangement (RLR) method is proposed to dynamically schedule convolution kernels with similar sparsity into the same computation batch to alleviate the load imbalance inter-PEs. In addition, to reduce off-chip memory accesses, a recurrent weight stationary (RWS) dataflow is proposed, which adopts a small-batch and multichannel strategy to stack data from multiple channels within one off-chip access and let them compute simultaneously thereby enabling efficient reuse of on-chip data. Based on the proposed scheme, an efficient SCNN inference accelerator has been designed and verified on the field-programmable gate array (FPGA). Compared with state-of-the-art works, our design achieves$1.16\times $to$2.77\times $higher digital signal processors (DSPs) efficiency and$1.75\times $to$15\times $higher logic efficiency. Jianbo Guo, Tongqing Xu, Zhenyang Wu, Hao Xiao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Scalable and Low-Cost NTT Architecture With Conflict-Free Memory Access SchemeabstractThis brief proposes a scalable multistage and multipath architecture for variable number-theoretic transform (NTT). The proposed architecture adopts multiple parallel paths, each of which uses cascaded radix-2 butterfly units (BFUs). The radix-2 scheme simplifies the control logic and the cascaded BFU structure reduces the amount of RAM banks and the frequency of memory accesses. Moreover, a conflict-free and hardware-friendly in-place memory mapping scheme is proposed to ease the adaption to multiple paths, letting it be scalable for various throughputs. Compared with state-of-the-art works, the proposed architecture uses fewer resources and has better area-time product performance without penalty in throughput. Zhenyang Wu, Ruichen Kan, Jianbo Guo, Hao Xiao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |