Wenhui Ou

dblp:348/7372 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0009-6779-5841ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FLICKER: A Fine-Grained Contribution-Aware Accelerator for Real-Time 3D Gaussian Splatting
abstract
Recently, 3D Gaussian Splatting (3DGS) has become a mainstream rendering technique for its photorealistic quality and low latency. However, the need to process massive noncontributing Gaussian points makes it struggle on resource-limited edge computing platforms and limits its use in next-gen AR/VR devices. A contribution-based prior skipping strategy is effective in alleviating this inefficiency, but the associated contribution-testing workload becomes prohibitive when it is further applied to the edge. In this paper, we present FLICKER, a contribution-aware 3DGS accelerator that leverages a hardware–software co-design framework, including adaptive leader pixels, pixel-rectangle grouping, hierarchical Gaussian testing, and mixed-precision architecture, to achieve near-pixel-level, contribution-driven rendering with minimal overhead. Experimental results show that our design achieves up to 1.5× speedup, 2.6× energy efficiency improvement, and 14% area reduction over a state-of-the-art accelerator. Meanwhile, it also achieves 19.8× speedup and 26.7× energy efficiency compared with a common edge GPU.
Wenhui Ou, Zhuoyu Wu, Yipu Zhang 0002, Dongjun Wu, Frederick Ziyang Hong, C. Patrick Yue
DATE1
2026 DepthPolyp: Pseudo-depth Guided Lightweight Segmentation for Real-Time Colonoscopy
Zhuoyu Wu, Wenhui Ou, Lexi Zhang, Pei-Sze Tan, Dongjun Wu, Junhe Zhao, Wenqi Fang, Raphael C.-W. Phan
ICPR (7)2
2025 PDR-KAN: Pipeline-Driven Reconfigurable Accelerator for Kolmogorov-Arnold Networks with Cross-Mode Sparsity Support
abstract
The commonly used Multi-layer Perceptrons (MLPs) in modern AI applications significantly limit the real-time performance due to their intensive memory access operation. Recently, Kolmogorov–Arnold Networks (KANs) have gained attention for offering a similar structure to MLPs but with a more efficient parameter utilization. However, the lack of customized support in conventional hardware restricts the performance gain from its algorithmic superiority. What’s more, the early-stage development of KAN raises questions about the necessity of incurring additional hardware costs for it. In this work, we present PDR-KAN, a reconfigurable accelerator that features two distinct operating modes, one for KANs and one for MLPs. Apart from its pipeline mode and sparsity encoder (SE) for efficient KAN processing, PDR-KAN can also improve the throughput of MLPs with cross-mode sparsity support. Experiments on a real-world dataset show that PDR-KAN provides a 3.95× acceleration and 13% reduction in accuracy loss by simply replacing MLPs with KANs. For a more accurate KAN model version with 3.33× parameter scaling, the latency overhead on PDR-KAN is only 1.16× compared to the baseline KAN model. Additionally, PDR-KAN achieves a 6.73× speed-up in KAN inferences and 20.89× in energy efficiency, against Quad-core ARM Cortex-A72 CPU.
Wenhui Ou, Zhuoyu Wu, Alexandra Geciova, Zheng Wang 0027, C. Patrick Yue
ISCAS1
2024 Low-latency Buffering for Mixed-precision Neural Network Accelerator with MulTAP and FQPipe
abstract
Previous work has proposed precision scalable accelerators to handle mixed-precision neural network (NN) inferences on the edge, which focus on designing reconfigurable MAC arrays while leaving the issue of time-costly data buffering procedure less discussed. Besides, integer-only inference is incapable of handling emerging NN models with various non-linear activation functions. In this work, we propose a mixed-precision NN accelerator supporting int8, int16 and fp32 arithmetic with two buffering techniques namely MulTAP and FQPipe, which jointly facilitate low-latency data movement. Experiment results show that MulTAP and FQPipe boost the baseline NN accelerator with 7.7 × and 1.5 × in speed respectively, which leads to the application performance of 473.9 (int8) and 252.5 (int16) inferences per second (IPS) on YOLOv3-Tiny. Post-layout netlist with SMIC 40nm standard-cell technology demonstrates a design with an area of 26.96mm2and a power estimate of 1.83W.
Zheng Wang 0027, Wenhui Ou, Weiyu Zhou, Yongkui Yang, Chao Chen 0022
ISCAS3
2023 COMPACT: Co-processor for Multi-mode Precision-adjustable Non-linear Activation Functions
abstract
Non-linear activation functions imitating neuron behaviors are ubiquitous in machine learning algorithms for time series signals while also demonstrating significant gain in precision for conventional vision-based deep learning networks. State-of-the-art implementation of such functions on GPU-like devices incurs a large physical cost, whereas edge devices adopt either linear interpolation or simplified linear functions leading to degraded precision. In this work, we design COMPACT, a co-processor with adjustable precision for multiple non-linear activation functions including but not limited to exponent, sigmoid, tangent, logarithm, and mish. Benchmarking with state-of-the-arts, COMPACT achieves a 26% reduction in the absolute error on a 1.6x widen approximation range taking advantage of the triple decomposition technique inspired by Hajduk's formula of Padé approximation. A SIMD-ISA-based vector co-processor has been implemented on FPGA which leads to a 30% reduction in execution latency but the area overhead nearly remains the same with related designs. Furthermore, COMPACT is adjustable to 46% latency improvement when the maximum absolute error is tolerant to the order of 1E-3.
Wenhui Ou, Zhuoyu Wu, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang
DATE1