EDBT 2026 Demo / reviewers in the wild / expert
Maohua Nie
dblp:355/5063
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-3515-7764ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI EnginesabstractCommercial FPGAs, such as AMD Versal devices, increasingly incorporate AI engines that exploit low-precision packed-SIMD fused multiply–accumulate (FMA) to achieve proportional throughput gains. However, trans-precision FMA (e.g., multiplying two FP16 numbers and adding their result to an FP32 accumulator), which preserves numerical stability by accumulating in higher precision, remains bottlenecked by the highest-precision, lowest-throughput operation. Dot-product accumulation (DPA) (e.g., performing a dot-product on two 4-element FP8 vectors and adding its result to an FP32 accumulator) can fully utilize the input/output bandwidth and computational resources. Existing flexible open-source FPUs, such as FPnew, do not support DPA and implement SIMD FMA on low-precision formats by replicating independent FMA lanes, which increases area, underutilizes shared arithmetic resources, and complicates the integration of DPA operations.This paper presents TransDot, a reconfigurable FPU that unifies multi-precision SIMD FMA and trans-precision DPA within a shared, reconfigurable datapath. TransDot extends the baseline design with 2-term FP16, 4-term FP8, and 8-term FP4 dot-product accumulation into FP32 using reconfigurable subcomponents. Evaluation shows that TransDot delivers 2× FP16, 4× FP8, and 8× FP4 throughput via DPA with FP32 accumulation, and 1.46× area efficiency in FP16 DPA and 2.92× area efficiency in FP8 DPA, at the cost of 37.3% larger area on average and an additional pipeline stage in dot-product mode compared to the FPnew baseline. These results demonstrate that TransDot’s area-efficient design enables scalable deployment in next-generation AMD Versal AI engines. Maohua Nie, Sin-Chen Lin, Chuanjin Richard Shi, Ang Li 0011 |
FCCM | 2 |
| 2025 | HCSAs: Hybrid Computing Systolic Arrays for Accelerating Mamba Models with Unified State Space Buffers and Energy-Efficient Dataflow
Maohua Nie, Chuanjin Richard Shi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | A Workload-Balance-Aware Accelerator Enabling Dense-to-Arbitrary-Sparse Neural NetworksabstractDeep neural networks (DNNs) have proved their great potential over various perceptual and cognitive tasks with the cost of ever-growing storage capacity and computation complexity. Sparse representations in neural networks have emerged as a compelling method to achieve substantial reductions in computational overhead and energy consumption. However, the introduction of sparsity presents challenges such as irregular memory accesses and wasted computation cycles. Traditional methods have attempted to address these challenges with varied success, unfortunately, often involving complex hardware designs or sacrificing sparsity’s benefits. In this article, we propose a software-hardware codesign solution, comprising an offline workload balancing encoding algorithm toward arbitrary sparsity in DNNs and a dedicated Processing Element array–based accelerator with a lightweight switch network. Extensive experiments are conducted to demonstrate the conclusion that our proposal is feasible to support various network structures and a wide range of sparsity ratios. With the encoding algorithm, a 1.16×–2.61× acceleration is achieved compared to the baseline. The system-on-chip measures 7.9 mm 2 , achieving an energy efficiency of 0.7 TOPS/W (dense), 2.1 TOPS/W (at 75% sparsity), and 10.3 TOPS/W (at 99.9% sparsity) at 0.9 V and 1,066-MHz clock frequency. Maohua Nie, Longfei Gou, Junmin He, Yongchuan Dong, Qiaosha Zou, Chuanjin Richard Shi |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2023 | AutoMap: Automatic Mapping of Neural Networks to Deep Learning Accelerators for Edge DevicesabstractEmerging deep neural networks (DNNs) have been emerging in applications (object detection, automatic speech recognition, etc.) deployed on edge devices. To improve the energy efficiency of edge devices, domain-specific deep learning accelerators (DLAs) are designed with limited on-chip resources. The manifold DLA designs and evolving DNN topologies bring challenges for applications mapping and scheduling on hardware resources. In this article, we propose an automatic DNN mapping framework named AutoMap, given the hardware backend information. First, a computational graph representation called extended directed weighted graph (EDWG) is proposed, which realizes unified expression for both spatial and temporal network interlayer connections. Second, an associated partitioner is implemented for splitting an EDWG into subEDWGs, which incorporates the on-chip memory constraint and facilitates weight data reuse on chip. Finally, a dynamic memory allocation strategy is utilized to alleviate the feature storing burden introduced by the multivarious network sizes and connections. Compared to the baseline mapping methods, experimental results show that our proposed automatic mapping framework can help to speedup the execution of several DNNs on state-of-the-art DLAs, ranging from$1.27\times $to$3.45\times $. The utilization of the PE array can increase from 20% to 64%. Maohua Nie, Qiaosha Zou, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |