Weize Ma

dblp:335/0321 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0008-2365-5284ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HoloLUT: An Efficient LUT-Based Engine via Holistic Data Processing for Low-bit LLM Inference
abstract
Weight-only quantization enhances the efficiency of large language models (LLMs) by storing weights in low-bit integers while retaining activations at a higher precision. However, the absence of efficient mixed-precision computation support on general-purpose hardware tends to hinder the potential computational gains from low-bit LLMs. While lookup table (LUT)-based methods offer a promising alternative, conventional bit-serial architectures introduce shift-and-accumulate bottlenecks, limiting throughput and energy efficiency. To overcome these limitations, we propose HoloLUT, a novel LUT-based engine that incorporates the Unitary Data Operation paradigm. This paradigm processes weights holistically rather than bit-serially, thereby eliminating shift-and-accumulate operations. Furthermore, a precision-adaptive mapping strategy combined with a unified LUT generator allows HoloLUT to flexibly and efficiently handle various precisions with negligible hardware overhead. Implemented in 28nm CMOS technology, HoloLUT achieves 1.86 × and 2.18 × improvements in area and power efficiency, respectively, compared to state-of-the-art LUT-based accelerators, demonstrating its strong potential for deploying low-bit LLMs in resource-constrained scenarios.
Hui Wang 0083, Weize Ma, Jinming Lu, Jun Lin 0001
ACM Great Lakes Symposium on VLSI2
2026 An Efficient Hardware Accelerator for JPEG-AI Image Compression on FPGA
Weize Ma, Siyuan Leng, Tong Chen 0004, Ming Lu 0003, Zhan Ma 0001
ISCAS1
2026 WiFlow: A Precision-Scalable DNN Training Accelerator Through Winograd Algorithm and Dataflow Co-Design
abstract
To address performance degradation from the domain shift and to support user-specific services while considering privacy, security, and communication overhead, there is an urgent need for efficient on-device training accelerators for deep neural networks (DNNs). Given limited computing resources and battery capacity constraints, implementing complex DNN training on edge devices is extremely challenging. To address these issues, we introduce a Winograd-Integrated Gradient Optimization Framework (WIGOF) for cross-phase operand sharing in the Winograd domain, which significantly reduces the number of multiplications and additions. Additionally, we develop WiFlow, an efficient, precision-scalable on-device training accelerator, minimizing area and power overheads of the dedicated Winograd transformation unit. The WiFlow supports 16-bit floating point (FP16), 16-bit brain floating point (BF16), and 8 and 4-bit fixed point (INT8 and INT4), demonstrating scalable improvements in both computational throughput (TOPS) and energy efficiency (TOPS/W) at low precision. A novel data rearrangement pattern, named channel augmentation, addresses the imperfect decomposition to enhance the utilization of processing element units. Furthermore, we propose a Winograd interleaved block-execution dataflow (WInBlock), along with Hierarchical Adaptive Reuse Memory Optimization (HARM) to improve data reuse and reduce both the amount of DRAM and SRAM access. The end-to-end training of WiFlow is achieved on Xilinx XCVU440 FPGA. WiFlow is also synthesized with a 28nm CMOS technology, achieving an area efficiency of 624 GOPS/mm2and an energy efficiency of 4.4 TOPS/W at a supply voltage of 0.9V and an operating frequency of 500 MHz. WiFlow accomplishes$7.75\times $higher area efficiency and$2.91\times $higher energy efficiency in actual DNN training compared with the state-of-the-art on-device training accelerators.
Hui Wang 0083, Jinming Lu, Weize Ma, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge
abstract
Monocular Depth Estimation (MDE) has emerged as a pivotal task in computer vision, supporting numerous real-world applications. However, deploying accurate depth estimation models on resource-limited edge devices, especially Application-Specific Integrated Circuits (ASICs), is challenging due to the high computational and memory demands. Recent advancements in foundational depth estimation deliver impressive results but further amplify the difficulty of deployment on ASICs. To address this, we propose Quart-Depth which adopts post-training quantization to quantize MDE models with hardware accelerations for ASICs. Our approach involves quantizing both weights and activations to 4-bit precision, reducing the model size and computation cost. To mitigate the performance degradation, we introduce activation polishing and compensation algorithm applied before and after activation quantization, as well as a weight reconstruction method for minimizing errors in weight quantization. Furthermore, we design a flexible and programmable hardware accelerator by supporting kernel fusion and customized instruction programmability, enhancing throughput and efficiency. Experimental results demonstrate that our framework achieves competitive accuracy while enabling fast inference and higher energy efficiency on ASICs, bridging the gap between high-performance depth estimation and practical edge-device applicability. Code: https://github.com/shawnricecake/quart-depth
Xuan Shen, Weize Ma, Jing Liu 0001, Changdi Yang, Quanyi Wang, Henghui Ding, Wei Niu 0002, Yanzhi Wang 0001, Pu Zhao 0001, Jiuxiang Gu
CVPR2
2025 DiffAccel: Accelerating Diffusion Models Through Adaptive Feature Optimization and Dynamic Hardware Adaptation
Enhao Tang, Weize Ma, Yudan Jiang, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Very Large Scale Integr. Syst.2