EDBT 2026 Demo / reviewers in the wild / expert
Qiuping Wu
dblp:96/8979
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reconfigurable Computing Challenge: FPGA-Based WebAssembly Stack Co-ProcessorabstractLarge language models suffer from hallucinations when performing scientific computing, motivating the use of AI agents such as IronClaw that offload computation to specialized tools. IronClaw invokes tools implemented as WebAssembly (Wasm) plugins for security and extensibility, but the stack-based Wasm bytecode is mismatched with register-based processors (x86, ARM), causing runtime overhead. We propose PAWS, a native Wasm coprocessor that directly executes Wasm bytecode in hardware. PAWS features: (1) full support for all five Wasm instruction types; (2) dual digital stack circuits (operand stack and control stack) replacing register files to minimize memory access latency; (3) dedicated control logic for block-based branching; and (4) a sliding-window instruction fetch unit that decodes variable-length Wasm instructions. Evaluated on the PolyBench suite, PAWS achieves average execution latencies 28.6× lower than an Intel Xeon processor and 40.6× lower than an Nvidia Jetson TX2, making it highly suitable for IronClaw’s compute-intensive scientific applications. The design is available at https://github.com/Iris-WQP/PAWS_FPGA_softcore. Qiuping Wu, Mugeng Liu 0001, Hongxiao Zhao, Yihan Fu, Gang Huang 0001, Yun Ma 0002, Bonan Yan |
FCCM | 1 |
| 2026 | ESTroM: Element-Flow Architecture for Processing Sparse Tractable Probabilistic ModelsabstractProbabilistic Circuits (PCs) models are emerging popular tractable probabilistic models. Their internal connections are represented in the form of directed acyclic graphs (DAGs) with sum nodes and product nodes, ensuring their internal parameter efficiency and model expressiveness in terms of probabilistic inference. Despite these algorithmic advantages, executing PC still faces graph structure deployment issues. PyJuice on GPU with the block-sparse parallel computation methods causes a parallelism-sparsity gap, while DAG-style processing does not take advantage of the repetitive characteristics of PC internal nodes, resulting in low throughput. To address this challenge, this work proposes the ESTroM, an efficient architecture that provides novel graph-element (nodes/edges) parallelism with sparsity-aware compilation. Through analysis of the sum/product node computational requirements, ESTroM core uses compressed matrices for sum/product nodes DAG representations, edge-based dataflow for product node processing, and node-based dataflow for sum node processing. With intra-core rewind and intercore multicast optimizations, we develop a prototype ESTrom chip and a demonstrative system for a PC-based neural lossless compression application. Our ablation experiments show ESTrom offers a speed improvement of$2.11 \sim 3.79 \times$compared to the state-of-the-art DAG processing unit (DPU)-v2 with the same computing resources. Under various typical PC structures, ESTrom achieves a speedup of$18.7 \times$compared to DPU-v2 and$3.9 \times$compared to NVIDIA RTX 4090 GPU with PyJuice framework. In terms of neural lossless compression, ESTroM demonstrates a$1.39 \times$improvement in compression ratio compared to the industrial-standard Z-standard (Zstd) algorithms with the highest compression level, while offering$16.3 \sim 65.2 \times$improvement in compression speed compared to Zstd on Intel Xeon Gold 6230. In a nutshell, this work develops novel graph element parallelism and element-flow architecture theory with practical prototype chips and systems, revealing a new hardware-perspective path for the “scaling law” of emerging tractable probabilistic models. Anjunyi Fan, Xuejie Liu, Anji Liu, Qiuping Wu, Jiaqi Yang 0009, Yuchao Qin, Guy Van den Broeck, Yitao Liang, Bonan Yan |
HPCA | 4 |
| 2025 | PROCA: Programmable Probabilistic Processing Unit Architecture with Accept/Reject Prediction & Multicore Pipelining for Causal InferenceabstractCausal inference is an important field in data science and cognitive artificial intelligence. It requires the construction of complex probabilistic models to describe the causal relationships between random variables. Probabilistic models rely on probabilistic programming as a flexible framework. However, the computing speed of probabilistic programming is often hindered by the extensive use of Markov chain Monte Carlo (MCMC) algorithms, even though they are powerful in Bayesian inference. To accelerate MCMC, this work presents PROCA, a programmable MCMC-based probabilistic processing unit architecture. PROCA exploits processing-in-memory function units to generate new samples of Markov chains. PROCA is programmable to execute the computation for arbitrary forms of posterior distribution formulas that software probabilistic programming frameworks support. We develop a novel accept/reject prediction methodology to accelerate the sequential MCMC computation, thereby introducing efficient multi-core pipelining methods. We implement and validate the PROCA architecture with commercial process development kits. The implementation is evaluated based on 9 representative benchmarks, covering PyMC official tutorial probabilistic problems, single-variable probabilistic problems, and real-world causal inference problems. Our comprehensive experiments demonstrate that PROCA achieves a speedup of 172~4871 $\times$ compared to Intel Xeon Gold CPU, $42 \sim 1058 \times$ compared to NVIDIA A100 GPU, and $1.765 \times$ over state-of-the-art MCMC accelerators, respectively. PROCA achieves comparable statistical robustness to the software probabilistic programming frameworks. Compared with state-of-the-art MCMC domain-specific accelerators, our design boosts the energy efficiency by $9.47 \times$. Yihan Fu, Anjunyi Fan, Wenshuo Yue, Hongxiao Zhao, Daijing Shi, Qiuping Wu, Yaoyu Tao, Yuchao Yang 0001, Bonan Yan |
HPCA | 6 |
| 2025 | A 28-nm 135.19 TOPS/W Bootstrapped-SRAM Compute-in-Memory Accelerator With Layer-Wise Precision and SparsityabstractArtificial intelligence (AI) edge devices demand high energy efficiency as well as inference accuracy. SRAM-based compute-in-memory (CIM) accelerators have great potential for power reduction but still need to exploit higher throughput and better linearity performance. To meet edge-AI computing demands by CIM works, it is crucial to optimize algorithms and parameters for specific circuit systems to achieve hardware acceleration. This work firstly employs neural network search (NAS) method to find out the layer-wise optimized precisions and sparsities for convolutional neural networks (CNNs). Then, a 144-Kb charge-domain signed mixed-precision (2/4/8-bit) CIM accelerator employing bootstrapped SRAM cells with 9-transistors and 1-capacitor (9T1C) structure is proposed that incorporates a bit-level sparsity-aware analog-to-digital converter (ADC). This work not only achieves highly linear parallel accumulation operations to meet AI computing demands but also implements a hardware and software co-optimization system tailored to specific data characteristics. The design is verified on NAS-optimized networks VGG-16 and ResNet-18 using Cifar-10 dataset, which could achieve an equivalent accuracy at 4-bit of 68.68% while maintaining a high energy efficiency at 2-bit of 135.19TOPS/W by measurements. Wei Mao 0002, Dingbang Liu, Haoxiang Zhou, Fuyi Li, Kai Li 0024, Qiuping Wu, Jiaqi Yang 0009, Liuyang Zhang, Hao Yu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | Multi-bit-width CNN Accelerator with Systolic-in-Systolic Dataflow and Single DSP Multiple Multiplication SchemeabstractMulti-bit-width neural network enlightens a promising method for high performance yet energy efficient edge computing due to its balance between software algorithm accuracy and hardware efficiency. To date, FPGA has been one of the core hardware platforms for deploying various neural networks. However, it is still difficult to fully make use of the dedicated digital signal processing (DSP) blocks in FPGA for accelerating the multi-bit-width network. In this work, we develop state-of-the-art multi-bit-width convolutional neural network accelerator with novel systolic-in-systolic type of dataflow and single DSP multiple multiplication (SDMM) INT2/4/8 execution scheme. Multi-level optimizations have also been adopted to further improve the performance, including group-vector systolic array for maximizing the circuit efficiency as well as minimizing the systolic delay, and differential neural architecture search (NAS) method for the high accuracy multi-bit-width network generation. The proposed accelerator has been practically deployed on Xilinx ZCU102 with accelerating NAS optimized VGG16 and Resnet18 networks as case studies. Average performance on accelerating the convolutional layer in VGG16 and Resnet18 is 1289GOPs and 1155GOPs, respectively. Throughput for running the full multi-bit-width VGG16 network is 870.73 GOPS at 250MHz, which has exceeded all of previous CNN accelerators on the same platform. Mingqiang Huang, Yucen Liu, Sixiao Huang, Kai Li 0024, Qiuping Wu, Hao Yu 0001 |
FPGA | 5 |
| 2019 | Correlation-Averaging Methods and Kalman Filter Based Parameter Identification for a Rotational Inertial Navigation SystemabstractThe attitude accuracy of the existing rotational inertial navigation system (RINS) is affected by oscillatory attitude errors caused by the installation errors of rotation axes or inertial sensors. Additional equipment is required to estimate installation errors under dynamic conditions. Methods that use the output of a single RINS to estimate installation errors under dynamic conditions are currently lacking. To address this challenge, this study proposes an installation error estimation method that combines a correlation method, an averaging method, and the Kalman filter. The proposed method adopts a correlation method to increase the signal-to-noise ratio, an averaging method to block certain sine signals, and the Kalman filter to identify installation errors in real time. Simulation, turntable, and sea tests were conducted to verify the proposed algorithm. Results show that the estimation accuracy of installation errors is at 10 arcsec levels, which indicates that said errors are estimated accurately using the RINS output initially obtained under dynamic conditions. Peida Hu, Bingxu Chen, Qiuping Wu |
IEEE Trans. Ind. Informatics | 4 |