VLDB 2026 Research / reviewers in the wild / expert
Qilin Hu
dblp:293/3952
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program analysis · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis
static analysis |
0.8 | 1 | 2024 | CBANA: A Lightweight, Efficient, and Flexible Cache Behavior Analysis Framework · IEEE Trans. Computers 2024 |
Memory systems
cache |
0.8 | 1 | 2024 | CBANA: A Lightweight, Efficient, and Flexible Cache Behavior Analysis Framework · IEEE Trans. Computers 2024 |
Memory systems › cache
cache behavior |
0.8 | 1 | 2024 | CBANA: A Lightweight, Efficient, and Flexible Cache Behavior Analysis Framework · IEEE Trans. Computers 2024 |
Memory systems › cache › cache behavior
cache miss analysis |
0.2 | 1 | 2024 | CBANA: A Lightweight, Efficient, and Flexible Cache Behavior Analysis Framework · IEEE Trans. Computers 2024 |
Methods — techniques the papers use, named apart from their topics
process address space analysis · 1.5loop refactoring · 1.5LLVM · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AutoNTT: Automatic Architecture Design and Exploration for Number Theoretic Transform Acceleration on FPGAsabstractFully Homomorphic Encryption (FHE), which enables homomorphic computing on encrypted data, has emerged as a promising privacy-aware computing method. However, FHE is orders-of-magnitude slower than the same computation on plain data, making it far from practical use. One of the major computation bottlenecks in FHE is the Number Theoretic Transform (NTT). While prior studies have accelerated NTT using specific architectures and FHE parameters, there still lacks a design automation tool to systematically design and explore various NTT architectures to support a diverse range of FHE parameters, such as various polynomial sizes, modulo sizes, and reduction methods. In this paper, we present AutoNTT, an open-source automatic architecture design and exploration tool to generate highly scalable NTT accelerators on FPGAs. Unlike prior studies, AutoNTT can automatically generate several optimized NTT acceleration architectures in HLS (i.e., iterative, dataflow, and hybrid architectures) with multiple common reduction methods, and support a large range of polynomial sizes (210–217) and modulo sizes ($log_{q}: 28-64$). In our auto-generated NTT architectures, we have applied many optimizations, such as polynomial and twiddle factor buffer reduction, and simplifying interconnections between different butterfly unit groups. Compared to prior studies, AutoNTT can generate NTT accelerators with 2.48× better latency and 3.61× better throughput on average, while maintaining a similar FPGA resource utilization. AutoNTT will be released soon at https://github.com/SFU-HiAccel/AutoNTT. Dilshan Kumarathunga, Qilin Hu, Zhenman Fang |
FCCM | 2 |
| 2025 | HiFA: A High-Performance and Flexible Acceleration Framework for Large-Size Number Theoretic TransformabstractZero-Knowledge Proofs (ZKP) and Homomorphic Encryption (HE) are crucial for data privacy in applications like cloud, blockchain, and analytics. However, the real-world adoption often faces performance challenges, particularly in the execution of the Number Theoretic Transform (NTT) required for polynomial multiplication involving sizes beyond \(2^{20}\) and large integer widths (e.g., 256 bits). FPGAs offer a promising platform for acceleration, but efficiently implementing large-size NTTs remains difficult due to the limited on-chip resources. The widely adopted four-step NTT method, used to relieve the need for large on-chip memory, introduces performance bottlenecks. Initially, the traditional dataflow NTT architecture may not fully exploit available compute capability, which hinders achieving peak performance. Furthermore, during the matrix transpose phase, the non-sequential access to external High-Bandwidth Memory (HBM) causes inefficiency. To address these challenges, we introduce HiFA, an FPGA-based automatic accelerator framework designed for high-performance and flexible large-size NTT computations. HiFA utilizes a stacked NTT architecture for high parallelism, maximizing HBM throughput. It supports various decomposed polynomial sizes via a novel reordering module. Additionally, a specialized cyclic shuffle module is integrated to optimize data movement during the matrix transpose step, alleviating random memory access delay. HiFA also provides an automatic Design Space Exploration (DSE) framework that identifies optimal four-step decomposition parameters and generates corresponding hardware configurations. Our experiments show that the FPGA implementation of HiFA achieves an average speedup of 2.97× and up to 7.25× improvement in latency over prior state-of-the-art FPGA solutions. Compared to prior GPU-based methods, HiFA achieves an average energy efficiency gain of 2.24×. Qilin Hu, Haotian Wang 0006, Chubo Liu, Keqin Li 0001, Kenli Li 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2024 | CBANA: A Lightweight, Efficient, and Flexible Cache Behavior Analysis FrameworkabstractCache miss analysis has become one of the most important things to improve the execution performance of a program. Generally, the approaches for analyzing cache misses can be categorized into dynamic analysis and static analysis. The former collects sampling statistics during program execution but is limited to specialized hardware support and incurs expensive execution overhead. The latter avoids the limitations but faces two challenges: inaccurate execution path prediction and inefficient analysis resulted by the explosion of the program state graph. To overcome these challenges, we propose CBANA, an LLVM- and process address space-based lightweight, efficient, and flexible cache behavior analysis framework. CBANA significantly improves the prediction accuracy of the execution path with awareness of inputs. To improve analysis efficiency and utilize the program preprocessing, CBANA refactors loop structures to reduce search space and dynamically splices intermediate results to reduce unnecessary or redundant computations. CBANA also supports configurable hardware parameter settings, and decouples the module of cache replacement policy from other modules. Thus, its flexibility is established. We evaluate CBANA by using the popular open benchmark PolyBench, graph workloads, and our synthetic workloads with good and poor data locality. Compared with the popular dynamic cache analysis tools Perf and Valgrind, the cache miss gap is less than 3.79% and 2.74% respectively with over ten thousand data accesses for the synthetic workloads, and the time reduction is up to 92.38% and 97.51% for the multiple-path workloads. Compared with the popular static cache analysis tool Heptane, CBANA achieves a time reduction of 97.71% while ensuring accuracy at the same time. Qilin Hu, Yan Ding 0004, Chubo Liu, Keqin Li 0001, Kenli Li 0001, Albert Y. Zomaya |
IEEE Trans. Computers | 1 |
| 2021 | Efficient Parallel Secure Outsourcing of Modular Exponentiation to Cloud for IoT ApplicationsabstractModular exponentiation, an operation widely utilized in cryptographic protocols to transfer text and other forms of data, can also be applied to Internet-of-Things (IoT) devices with high security requirements. However, due to the high resource consumption of modular exponentiation, IoT devices can face the problem of resource insufficient. Fortunately, the secure outsourcing scheme offers a new solution for resource-constrained devices. In this article, we apply a parallel secure outsourcing scheme to provide the possibility for modular exponentiation operation, which is used in the IoT devices. After that, the task of modular exponentiation is decomposed and we introduce the scheme in more detail. In addition, based on this scheme, we designed an extension scheme for RSA, providing enhanced security for IoT devices. Finally, the analysis of experimental results based on 512-4096 b of data indicates the superiority in scalability and time consumption over the previous schemes. Qilin Hu, Mingxing Duan, Zhibang Yang, Siyang Yu, Bin Xiao 0001 |
IEEE Internet Things J. | 1 |