Wenyong Zhou

dblp:37/7781 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Activation-Free Implicit Neural Representation via Finite-State-Machine Based Stochastic Computing
abstract
Implicit neural representations (INRs) have revolutionized signal encoding by using neural networks to map coordinates to signal attributes. Despite their success, INRs present significant hardware implementation challenges due to complex activation functions and floating-point operations. Unlike previous efforts, such as model pruning or quantization, we address these challenges by introducing AIRFSC, a novel activationfree stochastic computing (SC) architecture that leverages finitestate machines (FSMs). AIRFSC eliminates complex activation functions and processes data efficiently through stochastic bitstreams. Our approach decomposes the input signal into a series of Fourier basis functions, enabling the FSM-based architecture to learn smooth coordinate-to-attribute mappings for accurate signal reconstruction. Extensive experiments on diverse signal types demonstrate that AIRFSC achieves reconstruction quality comparable to state-of-the-art (SOTA) INRs implemented with multi-layer perceptrons (MLPs), while significantly improving hardware efficiency. Specifically, AIRFSC reduces power and area by $\mathbf{9 6. 8 \%}$ and $\mathbf{7 2. 8 \%}$ compared to Sinusoidal Representation Networks (SIREN), and by 97.9% and 81.3% compared to Wavelet Implicit Representation (WIRE).
Xincheng Feng, Wenyong Zhou, Taiqiang Wu, Meng Li 0004, Zhengwu Liu, Ngai Wong 0001
ASP-DAC2
2026 Noise-Aware Adaptive Sampling for Robust Diffusion Models on Analog Compute-in-Memory
abstract
Diffusion models achieve state-of-the-art image generation but impose heavy computational burdens on digital computers. Compute-in-memory (CIM) architectures offer promising acceleration, but inherent noise causes severe performance degradation through weight perturbations. We find that reducing sampling steps improves robustness but limits generation versatility, and that noise at earlier steps causes more severe degradation due to error accumulation. Based on these insights, we propose EtaMix, a novel noise-aware sampling strategy that interpolates between stochastic and deterministic sampling without requiring training or hardware modifications. EtaMix applies more stochastic sampling initially to offset weight perturbations, then gradually transitions to deterministic sampling. Experimental results show EtaMix achieves up to 2.01× and 5.12× FID improvements under different noise conditions for DDPM and DDIM, respectively.
Yuannuo Feng, Wenyong Zhou, Yuexi Lv, Guangyao Wang, Zhengwu Liu, Ngai Wong 0001, Wang Kang 0001
DATE2
2026 From SMURF to HI-SMURF: Scalable Multivariate Nonlinear Function Approximation via Compact Stochastic Architectures
Xincheng Feng, Wenyong Zhou, Taiqiang Wu, Zhengwu Liu, Meng Li 0004, Ngai Wong 0001
IEEE Trans. Computers2
2026 Binary Weight Multibit Activation Quantization for Compute-in-Memory CNN Accelerators
abstract
Compute-in-memory (CIM) accelerators have emerged as a promising way for enhancing the energy efficiency of convolutional neural networks (CNNs). Deploying CNNs on CIM platforms generally requires quantization of network weights and activations to meet hardware constraints. However, existing approaches either prioritize hardware efficiency with binary weight and activation quantization at the cost of accuracy, or utilize multi-bit weights and activations for greater accuracy but limited efficiency. In this paper, we introduce a novel binary weight multi-bit activation (BWMA) method for CNNs on CIM-based accelerators. Our contributions include: deriving closed-form solutions for weight quantization in each layer, significantly improving the representational capabilities of binarized weights; and developing a differentiable function for activation quantization, approximating the ideal multi-bit function while bypassing the extensive search for optimal settings. Through comprehensive experiments on CIFAR-10 and ImageNet datasets, we show that BWMA achieves notable accuracy improvements over existing methods, registering gains of 1.44%-5.46% and 0.35%-5.37% on respective datasets. Moreover, hardware simulation results indicate that 4-bit activation quantization strikes the optimal balance between hardware cost and model performance.
Wenyong Zhou, Zhengwu Liu, Ngai Wong 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
abstract
Low-rank adaptation (LoRA) is a predominant parameter-efficient finetuning method for adapting large language models (LLMs) to downstream tasks. Meanwhile, Compute-in-Memory (CIM) architectures demonstrate superior energy efficiency due to their array-level parallel in-memory computing designs. In this article, we propose deploying the LoRA-finetuned LLMs on the hybrid CIM architecture (i.e., pretrained weights onto energy-efficient Resistive Random-Access Memory (RRAM) and LoRA branches onto noise-free Static Random-Access Memory (SRAM)), reducing the energy cost to about 3% compared with the Nvidia A100 GPU. However, the inherent noise of RRAM on the saved weights leads to performance degradation, simultaneously. To address this issue, we design a novel Hardware-aware Low-rank Adaptation (HaLoRA) method. The key insight is to train a LoRA branch that is robust toward such noise and then deploy it on noise-free SRAM, while the extra cost is negligible since the parameters of LoRAs are much fewer than pretrained weights (e.g., 0.15% for LLaMA-3.2 1B model). To improve the robustness towards the noise, we theoretically analyze the gap between the optimization trajectories of the LoRA branch under both ideal and noisy conditions and further design an extra loss to minimize the upper bound of this gap. Therefore, we can enjoy both energy efficiency and accuracy during inference. Experiments finetuning the Qwen and LLaMA series demonstrate the effectiveness of HaLoRA across multiple reasoning tasks, achieving up to 22.7 improvement in average score while maintaining robustness at various noise types and noise levels.
Taiqiang Wu, Chenchen Ding, Wenyong Zhou, Yuxin Cheng, Xincheng Feng, Wendong Xu, Chufan Shi, Zhengwu Liu, Ngai Wong 0001
ACM Trans. Design Autom. Electr. Syst.3
2025 Towards Robust RRAM-Based Vision Transformer Models with Noise-Aware Knowledge Distillation
abstract
Resistive random-access memory (RRAM)-based compute-in-memory (CIM) systems show promise in accelerating Transformer-based vision models but face challenges from inherent device non-idealities. In this work, we systematically investigate the vulnerability of Transformer-based vision models to RRAM-induced perturbations. Our analysis reveals that earlier Transformer layers are more vulnerable than later ones, and feed-forward networks (FFNs) are more susceptible to noise than multi-head self-attention (MHSA). Based on these observations, we propose a noise-aware knowledge distillation framework that enhances model robustness by aligning both intermediate features and final outputs between weight-perturbed and noise-free models. Experimental results demonstrate that our method improves accuracy by up to 1.54% and 1.49% on ViT and DeiT models under various noise conditions compared to their vanilla counterparts.
Wenyong Zhou, Zhengwu Liu, Taiqiang Wu, Chenchen Ding, Ngai Wong 0001
DATE1
2025 Enhancing Robustness of Implicit Neural Representations Against Weight Perturbations
abstract
Implicit Neural Representations (INRs) encode discrete signals in a continuous manner using neural networks, demonstrating significant value across various multimedia applications. However, the vulnerability of INRs presents a critical challenge for their real-world deployments, as the network weights might be subjected to unavoidable perturbations. In this work, we investigate the robustness of INRs for the first time and find that even minor perturbations can lead to substantial performance degradation in the quality of signal reconstruction. To mitigate this issue, we formulate the robustness problem in INRs by minimizing the difference between loss with and without weight perturbations. Furthermore, we derive a novel robust loss function to regulate the gradient of the reconstruction loss with respect to weights, thereby enhancing the robustness. Extensive experiments on reconstruction tasks across multiple modalities demonstrate that our method achieves up to a 7.5 dB improvement in peak signal-to-noise ratio (PSNR) values compared to original INRs under noisy conditions.
Wenyong Zhou, Yuxin Cheng, Zhengwu Liu, Taiqiang Wu, Ngai Wong 0001
ICASSP1
2025 MINR: Efficient Implicit Neural Representations for Multi-Image Encoding
abstract
Implicit Neural Representations (INRs) aim to parameterize discrete signals through implicit continuous functions. However, formulating each image with a separate neural network (typically, a Multi-Layer Perceptron (MLP)) leads to computational and storage inefficiencies when encoding multi-images. To address this issue, we propose MINR, sharing specific layers to encode multi-image efficiently. We first compare the layer-wise weight distributions for several trained INRs and find that corresponding intermediate layers follow highly similar distribution patterns. Motivated by this, we share these intermediate layers across multiple images while preserving the input and output layers as input-specific. In addition, we design an extra novel projection layer for each image to capture its unique features. Experimental results on image reconstruction and super-resolution tasks demonstrate that MINR can save up to 60% parameters while maintaining comparable performance. Particularly, MINR scales effectively to handle 100 images, maintaining an average peak signal-to-noise ratio (PSNR) of 34 dB. Further analysis of various backbones proves the robustness of the proposed MINR.
Wenyong Zhou, Taiqiang Wu, Zhengwu Liu, Yuxin Cheng, Ngai Wong 0001
ICASSP1
2025 Perspective-Aware 3D Gaussian Inpainting with Multi-View Consistency
abstract
3D Gaussian inpainting, a critical technique for numerous applications in virtual reality and multimedia, has made significant progress with pretrained diffusion models. However, ensuring multi-view consistency, an essential requirement for high-quality inpainting, remains a key challenge. In this work, we present PAInpainter, a novel approach designed to advance 3D Gaussian inpainting by leveraging perspective-aware content propagation and consistency verification across multi-view inpainted images. Our method iteratively refines inpainting and optimizes the 3D Gaussian representation with multiple views adaptively sampled from a perspective graph. By propagating inpainted images as prior information and verifying consistency across neighboring views, PAInpainter substantially enhances global consistency and texture fidelity in restored 3D scenes. Extensive experiments demonstrate the superiority of PAInpainter over existing methods. Our approach achieves superior 3D inpainting quality, with PSNR scores of 26.03 dB and 29.51 dB on the SPIn-NeRF and NeRFiller datasets, respectively, highlighting its effectiveness and generalization capability.
Yuxin Cheng, Binxiao Huang, Taiqiang Wu, Wenyong Zhou, Chenchen Ding, Zhengwu Liu, Graziano Chesi, Ngai Wong 0001
ICCV4
2025 Distribution-Aware Hadamard Quantization for Hardware-Efficient Implicit Neural Representations
abstract
Implicit Neural Representations (INRs) encode discrete signals using Multi-Layer Perceptrons (MLPs) with complex activation functions. While INRs achieve superior performance, they depend on full-precision number representation for accurate computation, resulting in significant hardware overhead. Previous INR quantization approaches have primarily focused on weight quantization, offering only limited hardware savings due to the lack of activation quantization. To fully exploit the hardware benefits of quantization, we propose DHQ, a novel distribution-aware Hadamard quantization scheme that targets both weights and activations in INRs. Our analysis shows that the weights in the first and last layers have distributions distinct from those in the intermediate layers, while the activations in the last layer differ significantly from those in the preceding layers. Instead of customizing quantizers individually, we utilize the Hadamard transformation to standardize these diverse distributions into a unified bell-shaped form, supported by both empirical evidence and theoretical analysis, before applying a standard quantizer. To demonstrate the practical advantages of our approach, we present an FPGA implementation of DHQ that highlights its hardware efficiency. Experiments on diverse image reconstruction tasks show that DHQ outperforms previous quantization methods, reducing latency by 32.7%, energy consumption by 40.1%, and resource utilization by up to 98.3% compared to full-precision counterparts.
Wenyong Zhou, Jiachen Ren, Taiqiang Wu, Yuxin Cheng, Zhengwu Liu, Ngai Wong 0001
ICME1
2025 Re-Activating Frozen Primitives for 3D Gaussian Splatting
Yuxin Cheng, Binxiao Huang, Wenyong Zhou, Taiqiang Wu, Zhengwu Liu, Graziano Chesi, Ngai Wong 0001
ACM Multimedia3
2024 Physics-Informed Learning for Versatile RRAM Reset and Retention Simulation
abstract
Resistive random-access memory (RRAM) constitutes an emerging and promising platform for compute-inmemory (CIM) edge AI. However, the switching mechanism and controllability of RRAM are still under debate owing to the influence of multiphysics. Although physics-informed neural networks (PINNs) are successful in achieving mesh-free multiphysics solutions in many applications, the resultant accuracy is not satisfactory in RRAM analyses. This work investigates the characteristics of RRAM devices - retention and reset transition which are described in terms of the dissolution of a conductive filament (CF) in 3-D axis-symmetric geometry. Specifically, we provide a novel neural network characterization of ion migration, Joule heating, and carrier transport, governed by the solutions of partial differential equations (PDEs). Motivated by physics-informed learning, the separation of variables (SOV) method and the neural tangent kernel (NTK) theory, we propose a customized 3-channel fully-connected network and a modified random Fourier feature (mRFF) embedding strategy to capture multiscale properties and appropriate frequency features of the self-consistent multiphysics solutions. The proposed model eliminates the need for grid meshing and temporal iterations widely used in RRAM analysis. Experiments then confirm its superior accuracy over competing physics-informed methods.
Tianshu Hou, Wenyong Zhou, Can Li 0024, Haibao Chen, Ngai Wong 0001
ASPDAC3
2024 Hybrid Module with Multiple Receptive Fields and Self-Attention Layers for Medical Image Segmentation
abstract
Recent advances in medical image segmentation models combine convolution with the attention mechanism which provides an effective approach to formulate long-term dependencies. However, many works either replaced the convolutional layers with attention layers or embedded attention layers into convolutional neural network (CNN)-based models. To explore the potential of hybrid architecture, we propose a simple cascade module that builds up multiple receptive fields using convolutional kernels with different sizes and learns global context via self-attention layers. Benefiting from the powerful representation ability of the proposed module, multilayer perceptrons (MLPs) with shift operation are adopted to bridge the encoder and decoder to reduce the model size without losing accuracy. Experiments show that our model consistently outperforms the latest 2D and 3D models by large margins on three public tasks and is more resilient to shape, size, and boundary variations. The code is available at https://github.com/cicailalala/AERFNet.
Wenbo Qi, Wenyong Zhou, Ngai Wong 0001, S. C. Chan 0001
ICASSP2
2024 Integrating FMEA and fuzzy super-efficiency SBM for risk assessment of crowdfunding project investment
Mengshan Zhu, Wenyong Zhou, Chunyan Duan
Soft Comput.2
2017 Super-Resolution Reconstruction from Single Image Based on Join Operation in Granular Computing
abstract
Improving the resolution of the image is convenient for people to study the local details of the image, and plays an important role in computer vision. The problem of generating a corresponding super-resolution (SR) image from a single low-resolution (LR) image is addressed via the join operation in the paper. Firstly, the LR image is partitioned into some patches, each patch is represented as the sphere granule set. Secondly, the join operation between two adjacent image patches is used to compensate the pixel value of SR image. Experimental results showed the feasibility and superiority via join operation by root mean square errors (RMSE) between the reconstructed SR image and the original image compared with bicubic interpolation and NNLasso.
Wenyong Zhou, Xuewen Ma, Chang-an Wu
Int. J. Pattern Recognit. Artif. Intell.2
2013 Measure Method of Fuzzy Inclusion Relation in Granular Computing
Wenyong Zhou, Chunhua Liu
ISNN (2)1