EDBT 2026 Demo / reviewers in the wild / expert
Zhuoyu Wu
dblp:71/6669
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-9075-7626ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLICKER: A Fine-Grained Contribution-Aware Accelerator for Real-Time 3D Gaussian SplattingabstractRecently, 3D Gaussian Splatting (3DGS) has become a mainstream rendering technique for its photorealistic quality and low latency. However, the need to process massive noncontributing Gaussian points makes it struggle on resource-limited edge computing platforms and limits its use in next-gen AR/VR devices. A contribution-based prior skipping strategy is effective in alleviating this inefficiency, but the associated contribution-testing workload becomes prohibitive when it is further applied to the edge. In this paper, we present FLICKER, a contribution-aware 3DGS accelerator that leverages a hardware–software co-design framework, including adaptive leader pixels, pixel-rectangle grouping, hierarchical Gaussian testing, and mixed-precision architecture, to achieve near-pixel-level, contribution-driven rendering with minimal overhead. Experimental results show that our design achieves up to 1.5× speedup, 2.6× energy efficiency improvement, and 14% area reduction over a state-of-the-art accelerator. Meanwhile, it also achieves 19.8× speedup and 26.7× energy efficiency compared with a common edge GPU. Wenhui Ou, Zhuoyu Wu, Yipu Zhang 0002, Dongjun Wu, Frederick Ziyang Hong, C. Patrick Yue |
DATE | 2 |
| 2026 | DepthPolyp: Pseudo-depth Guided Lightweight Segmentation for Real-Time Colonoscopy
Zhuoyu Wu, Wenhui Ou, Lexi Zhang, Pei-Sze Tan, Dongjun Wu, Junhe Zhao, Wenqi Fang, Raphael C.-W. Phan |
ICPR (7) | 1 |
| 2026 | Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoCabstractWhile recent advances in AI SoC design have focused heavily on accelerating tensor computation, the equally critical task of tensor manipulation (TM)—centered on high-volume data movement with minimal computation—remains underexplored. This work addresses that gap by introducing the TM unit (TMU): a reconfigurable, near-memory hardware block designed to execute data-movement-intensive (DMI) operators efficiently. The TMU manipulates long datastreams in a memory-to-memory fashion using a RISC-inspired execution model and a unified addressing abstraction, enabling support for both a wide range of coarse- and fine-grained tensor transformations. The proposed architecture integrates the TMU alongside a TPU within a high-throughput AI system-on-chip (SoC), leveraging double buffering and output forwarding to improve pipeline utilization. The TMU, synthesized under the SMIC 40-nm standard cell library, occupies only$0.019~\mathrm {\text {mm}^{2}}$while supporting over 10 representative DMI operators. Benchmarking shows that the TMU alone achieves up to$82.42\times $and$11.06\times $operator-level latency reduction over ARM A72 and NVIDIA Jetson TX2, respectively. When integrated with the in-house TPU, the complete system achieves a 22.89% reduction in end-to-end inference latency, demonstrating the effectiveness in reducing inference latency and the scalability of the TMU architecture across diverse tensor operators. Weiyu Zhou, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Zhuoyu Wu, Anupam Chattopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | AttenPU: An Area Efficient Attention Processor with Reconfigurable FP8 Precision and DataflowabstractEfficient numerical representation is crucial for deep learning accelerators, especially for large language models (LLMs). The 8-bit-floating-point (FP8) data representation achieves higher precision and fewer quantization efforts than integer, which has been proven inevitable in attention-based accelerators for LLMs. Therefore, area-efficient design techniques for FP8 play a central role in lowering LLMs chip’s budget. This paper presents AttenPU, which is built upon reconfigurable FP8 units and supports E4M3 for inference and E5M2 for training. Bidirectional dataflow is exploited to enable AttenPU to interact with FP32 coprocessor to reduce latency. The design achieves a low FP8-to-INT8 area ratio of 1.63×, an area efficiency of 193.5 GFLOPS/mm2, with an 87.05% reduction in the latency of RTX3090 GPU. Qiawei Zheng, Zheng Wang 0027, Zhuoyu Wu, Zhihao Du, Chao Chen 0022, Yongkui Yang, Wenqi Fang, Anupam Chattopadhyay |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | PDR-KAN: Pipeline-Driven Reconfigurable Accelerator for Kolmogorov-Arnold Networks with Cross-Mode Sparsity SupportabstractThe commonly used Multi-layer Perceptrons (MLPs) in modern AI applications significantly limit the real-time performance due to their intensive memory access operation. Recently, Kolmogorov–Arnold Networks (KANs) have gained attention for offering a similar structure to MLPs but with a more efficient parameter utilization. However, the lack of customized support in conventional hardware restricts the performance gain from its algorithmic superiority. What’s more, the early-stage development of KAN raises questions about the necessity of incurring additional hardware costs for it. In this work, we present PDR-KAN, a reconfigurable accelerator that features two distinct operating modes, one for KANs and one for MLPs. Apart from its pipeline mode and sparsity encoder (SE) for efficient KAN processing, PDR-KAN can also improve the throughput of MLPs with cross-mode sparsity support. Experiments on a real-world dataset show that PDR-KAN provides a 3.95× acceleration and 13% reduction in accuracy loss by simply replacing MLPs with KANs. For a more accurate KAN model version with 3.33× parameter scaling, the latency overhead on PDR-KAN is only 1.16× compared to the baseline KAN model. Additionally, PDR-KAN achieves a 6.73× speed-up in KAN inferences and 20.89× in energy efficiency, against Quad-core ARM Cortex-A72 CPU. Wenhui Ou, Zhuoyu Wu, Alexandra Geciova, Zheng Wang 0027, C. Patrick Yue |
ISCAS | 2 |
| 2025 | GUI-Narrator: Detecting and Captioning Computer GUI Actions
Qinchen Wu, Difei Gao, Qinghong Lin, Zhuoyu Wu, Zheng Shou 0001 |
ACM Multimedia | 4 |
| 2023 | COMPACT: Co-processor for Multi-mode Precision-adjustable Non-linear Activation FunctionsabstractNon-linear activation functions imitating neuron behaviors are ubiquitous in machine learning algorithms for time series signals while also demonstrating significant gain in precision for conventional vision-based deep learning networks. State-of-the-art implementation of such functions on GPU-like devices incurs a large physical cost, whereas edge devices adopt either linear interpolation or simplified linear functions leading to degraded precision. In this work, we design COMPACT, a co-processor with adjustable precision for multiple non-linear activation functions including but not limited to exponent, sigmoid, tangent, logarithm, and mish. Benchmarking with state-of-the-arts, COMPACT achieves a 26% reduction in the absolute error on a 1.6x widen approximation range taking advantage of the triple decomposition technique inspired by Hajduk's formula of Padé approximation. A SIMD-ISA-based vector co-processor has been implemented on FPGA which leads to a 30% reduction in execution latency but the area overhead nearly remains the same with related designs. Furthermore, COMPACT is adjustable to 46% latency improvement when the maximum absolute error is tolerant to the order of 1E-3. Wenhui Ou, Zhuoyu Wu, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang |
DATE | 2 |