EDBT 2026 Demo / reviewers in the wild / expert
Qiwei Dang
dblp:360/1174
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0006-0326-2999ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoEA: A Mixed-Precision Edge Accelerator for CNN-MSA Models with Fine-Tuning Support
Qiwei Dang, Chengyu Ma, Zhiwang Huo, Guoming Yang, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
ASP-DAC | 1 |
| 2026 | Hierarchical-ISA Supporting Row-Wise Operands for Efficient DNN ComputationabstractDeep neural networks (DNNs) have become a cornerstone in advancing artificial intelligence, but their complexity often leads to inefficient hardware utilization due to varying structure characteristics and excessive memory accesses. Domain-specific architectures (DSAs) offer a solution by optimizing data locality through data stationary, tiling, and layer fusion, which minimize memory access and energy consumption while boosting performance. However, current approaches lack flexibility for efficient memory management at the appropriate granularity, causing misaligned accesses and decreasing reuse potential for variable-sized tiles. To this end, we propose a hierarchical Instruction Set Architecture (hierarchical-ISA) combining a RISC-V ISA and a flexible CISC-style macro-ISA (mISA). Unlike byte-level RISC-V, mISA employs row-wise tiles as the fundamental operand, enabling efficient data reuse across adjacent iterations as well as residual connections. This mISA approach simplifies DNN programming, enhances data partitioning and manipulation efficiency, and enables a hardware-software co-designed Remapping mechanism that facilitates data reuse without physical data movement. Experiments show 31.8%–72.0% reductions in off-chip memory access across MobileNet, ResNet, Swin Transformer, MobileViT, along with speedups of 2.9× to 7.4× compared to previous DNN accelerators. We also conduct comparisons under the roofline model with NVIDIA RTX A6000 and Intel Core i7-10700K. The results show that our arithmetic intensity reaches up to 26.0× that of i7-10700K and 22.6× that of A6000. Zhiwang Huo, Wenzhe Zhao 0001, Qiwei Dang, Chengyu Ma, Guoming Yang, Gelin Fu, Tian Xia 0008, Pengju Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | FP2: A 2-bit Floating-Point Format for Edge-AI Inference and Fine-TuningabstractThe increasing scale of Deep Neural Networks (DNNs) has made 2-bit quantization crucial for mitigating memory bottlenecks on edge devices. Low-bitwidth floating-point formats, offering larger dynamic ranges and avoiding quantization steps, have emerged as promising alternatives to fixed-point quantization. However, constructing viable floating-point representations with fewer than 3 bits remains challenging, as conventional formats require at least one sign bit, one exponent bit, and one mantissa bit. We address this challenge by introducing a novel data compression method that uses a 4-bit encoding space to represent two floating-point values, achieving an effective storage density of 2 bits per value. Depending on the bit width of the exponent and mantissa, we propose two different 2-bit floating-point encodings:fp2-e1m0andfp2-e0m1. Based onfp2, we introduce two computing architectures that simplify floating-point multiply-accumulate (MAC) operations into bitwise addition and logic operations, reducing floating-point computation by factors of$2\times $and$4\times $. As a result,fp2offers a practical solution for efficient inference using floating-point arithmetic on resource-constrained edge devices. Moreover, we analyze the error characteristics of thefp2data format from three perspectives. To validate the effectiveness of thefp2format, we conduct experiments on ResNet18/50 and ConvNeXt-Tiny using the CIFAR-10 and ImageNet-1K datasets. Compared tofp4, our approach reduces model size by 47%, with accuracy loss is less than 2 percentage points. Notably, on CIFAR-10, some results are close to those offp32. In contrast, when evaluated under 2-bit GPTQ,fp2demonstrates significant advantages over the baseline method on the LLAMA model. For hardware evaluation, we implement our design at the RTL level and evaluate it on both FPGA and ASIC platforms. Compared to computation architectures based onfp4, ourfp4$\times $fp2processing element (PE) array reduces area by 15% and power consumption by 8%. Furthermore, ourfp2$\times $fp2PE array achieves a remarkable 78% reduction in both area and power consumption. Qiwei Dang, Chengyu Ma, Haiduo Huang, Gelin Fu, Zhiwang Huo, Guoming Yang, Pengchen Zong, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Optimizing FPGA-Based DNN Accelerator With Shared Exponential Floating-Point FormatabstractIn recent years, low-precision fixed-point computation has become a widely used technique for neural network inference on FPGAs. However, this approach has some limitations, as certain neural networks are difficult to quantify using fixed-point arithmetic, such as those involved in super-resolution scaling, image denoising, and other scenarios that lack sufficient conditions for fine-tuning. Furthermore, deploying a floating-point precision neural network directly on an FPGA would lead to significant hardware overhead and low computational efficiency. To address this issue, this paper proposes an FPGA-friendly floating-point data format that achieves the same storage density as int8 without sacrificing inference accuracy or requiring fine-tuning. Additionally, this paper presents an FPGA-based neural network accelerator that is compatible with the proposed format, utilizing DSP resources to increase the number of DSP cascading from 7 to 16, and solving the back-to-back accumulation issue of floating-point numbers. This design achieves comparable resource consumption and execution efficiency to those of 8-bit fixed-point accelerators. Experimental results demonstrate that the accelerator proposed in this study achieves the same accuracy as the native floating point on multiple neural networks without fine-tuning, and remains high computing performance. When deployed on the Xilinx ZU9P, the performance achieves 4.072 TFlops at 250 MHz, which outperforms the previous works, including the Xilinx official DPU. Wenzhe Zhao 0001, Qiwei Dang, Tian Xia 0008, Nanning Zheng 0001, Pengju Ren |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |