EDBT 2026 Demo / reviewers in the wild / expert
Weize Ma
dblp:335/0321
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0008-2365-5284ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HoloLUT: An Efficient LUT-Based Engine via Holistic Data Processing for Low-bit LLM InferenceabstractWeight-only quantization enhances the efficiency of large language models (LLMs) by storing weights in low-bit integers while retaining activations at a higher precision. However, the absence of efficient mixed-precision computation support on general-purpose hardware tends to hinder the potential computational gains from low-bit LLMs. While lookup table (LUT)-based methods offer a promising alternative, conventional bit-serial architectures introduce shift-and-accumulate bottlenecks, limiting throughput and energy efficiency. To overcome these limitations, we propose HoloLUT, a novel LUT-based engine that incorporates the Unitary Data Operation paradigm. This paradigm processes weights holistically rather than bit-serially, thereby eliminating shift-and-accumulate operations. Furthermore, a precision-adaptive mapping strategy combined with a unified LUT generator allows HoloLUT to flexibly and efficiently handle various precisions with negligible hardware overhead. Implemented in 28nm CMOS technology, HoloLUT achieves 1.86 × and 2.18 × improvements in area and power efficiency, respectively, compared to state-of-the-art LUT-based accelerators, demonstrating its strong potential for deploying low-bit LLMs in resource-constrained scenarios. Hui Wang 0083, Weize Ma, Jinming Lu, Jun Lin 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | An Efficient Hardware Accelerator for JPEG-AI Image Compression on FPGA
Weize Ma, Siyuan Leng, Tong Chen 0004, Ming Lu 0003, Zhan Ma 0001 |
ISCAS | 1 |
| 2026 | WiFlow: A Precision-Scalable DNN Training Accelerator Through Winograd Algorithm and Dataflow Co-DesignabstractTo address performance degradation from the domain shift and to support user-specific services while considering privacy, security, and communication overhead, there is an urgent need for efficient on-device training accelerators for deep neural networks (DNNs). Given limited computing resources and battery capacity constraints, implementing complex DNN training on edge devices is extremely challenging. To address these issues, we introduce a Winograd-Integrated Gradient Optimization Framework (WIGOF) for cross-phase operand sharing in the Winograd domain, which significantly reduces the number of multiplications and additions. Additionally, we develop WiFlow, an efficient, precision-scalable on-device training accelerator, minimizing area and power overheads of the dedicated Winograd transformation unit. The WiFlow supports 16-bit floating point (FP16), 16-bit brain floating point (BF16), and 8 and 4-bit fixed point (INT8 and INT4), demonstrating scalable improvements in both computational throughput (TOPS) and energy efficiency (TOPS/W) at low precision. A novel data rearrangement pattern, named channel augmentation, addresses the imperfect decomposition to enhance the utilization of processing element units. Furthermore, we propose a Winograd interleaved block-execution dataflow (WInBlock), along with Hierarchical Adaptive Reuse Memory Optimization (HARM) to improve data reuse and reduce both the amount of DRAM and SRAM access. The end-to-end training of WiFlow is achieved on Xilinx XCVU440 FPGA. WiFlow is also synthesized with a 28nm CMOS technology, achieving an area efficiency of 624 GOPS/mm2and an energy efficiency of 4.4 TOPS/W at a supply voltage of 0.9V and an operating frequency of 500 MHz. WiFlow accomplishes$7.75\times $higher area efficiency and$2.91\times $higher energy efficiency in actual DNN training compared with the state-of-the-art on-device training accelerators. Hui Wang 0083, Jinming Lu, Weize Ma, Zhongfeng Wang 0001, Jun Lin 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the EdgeabstractMonocular Depth Estimation (MDE) has emerged as a pivotal task in computer vision, supporting numerous real-world applications. However, deploying accurate depth estimation models on resource-limited edge devices, especially Application-Specific Integrated Circuits (ASICs), is challenging due to the high computational and memory demands. Recent advancements in foundational depth estimation deliver impressive results but further amplify the difficulty of deployment on ASICs. To address this, we propose Quart-Depth which adopts post-training quantization to quantize MDE models with hardware accelerations for ASICs. Our approach involves quantizing both weights and activations to 4-bit precision, reducing the model size and computation cost. To mitigate the performance degradation, we introduce activation polishing and compensation algorithm applied before and after activation quantization, as well as a weight reconstruction method for minimizing errors in weight quantization. Furthermore, we design a flexible and programmable hardware accelerator by supporting kernel fusion and customized instruction programmability, enhancing throughput and efficiency. Experimental results demonstrate that our framework achieves competitive accuracy while enabling fast inference and higher energy efficiency on ASICs, bridging the gap between high-performance depth estimation and practical edge-device applicability. Code: https://github.com/shawnricecake/quart-depth Xuan Shen, Weize Ma, Jing Liu 0001, Changdi Yang, Quanyi Wang, Henghui Ding, Wei Niu 0002, Yanzhi Wang 0001, Pu Zhao 0001, Jiuxiang Gu |
CVPR | 2 |
| 2025 | DiffAccel: Accelerating Diffusion Models Through Adaptive Feature Optimization and Dynamic Hardware Adaptation
Enhao Tang, Weize Ma, Yudan Jiang, Zhongfeng Wang 0001, Jun Lin 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |