Gelin Fu

dblp:338/8850 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-7331-0830ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical-ISA Supporting Row-Wise Operands for Efficient DNN Computation
abstract
Deep neural networks (DNNs) have become a cornerstone in advancing artificial intelligence, but their complexity often leads to inefficient hardware utilization due to varying structure characteristics and excessive memory accesses. Domain-specific architectures (DSAs) offer a solution by optimizing data locality through data stationary, tiling, and layer fusion, which minimize memory access and energy consumption while boosting performance. However, current approaches lack flexibility for efficient memory management at the appropriate granularity, causing misaligned accesses and decreasing reuse potential for variable-sized tiles. To this end, we propose a hierarchical Instruction Set Architecture (hierarchical-ISA) combining a RISC-V ISA and a flexible CISC-style macro-ISA (mISA). Unlike byte-level RISC-V, mISA employs row-wise tiles as the fundamental operand, enabling efficient data reuse across adjacent iterations as well as residual connections. This mISA approach simplifies DNN programming, enhances data partitioning and manipulation efficiency, and enables a hardware-software co-designed Remapping mechanism that facilitates data reuse without physical data movement. Experiments show 31.8%–72.0% reductions in off-chip memory access across MobileNet, ResNet, Swin Transformer, MobileViT, along with speedups of 2.9× to 7.4× compared to previous DNN accelerators. We also conduct comparisons under the roofline model with NVIDIA RTX A6000 and Intel Core i7-10700K. The results show that our arithmetic intensity reaches up to 26.0× that of i7-10700K and 22.6× that of A6000.
Zhiwang Huo, Wenzhe Zhao 0001, Qiwei Dang, Chengyu Ma, Guoming Yang, Gelin Fu, Tian Xia 0008, Pengju Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 FP2: A 2-bit Floating-Point Format for Edge-AI Inference and Fine-Tuning
abstract
The increasing scale of Deep Neural Networks (DNNs) has made 2-bit quantization crucial for mitigating memory bottlenecks on edge devices. Low-bitwidth floating-point formats, offering larger dynamic ranges and avoiding quantization steps, have emerged as promising alternatives to fixed-point quantization. However, constructing viable floating-point representations with fewer than 3 bits remains challenging, as conventional formats require at least one sign bit, one exponent bit, and one mantissa bit. We address this challenge by introducing a novel data compression method that uses a 4-bit encoding space to represent two floating-point values, achieving an effective storage density of 2 bits per value. Depending on the bit width of the exponent and mantissa, we propose two different 2-bit floating-point encodings:fp2-e1m0andfp2-e0m1. Based onfp2, we introduce two computing architectures that simplify floating-point multiply-accumulate (MAC) operations into bitwise addition and logic operations, reducing floating-point computation by factors of$2\times $and$4\times $. As a result,fp2offers a practical solution for efficient inference using floating-point arithmetic on resource-constrained edge devices. Moreover, we analyze the error characteristics of thefp2data format from three perspectives. To validate the effectiveness of thefp2format, we conduct experiments on ResNet18/50 and ConvNeXt-Tiny using the CIFAR-10 and ImageNet-1K datasets. Compared tofp4, our approach reduces model size by 47%, with accuracy loss is less than 2 percentage points. Notably, on CIFAR-10, some results are close to those offp32. In contrast, when evaluated under 2-bit GPTQ,fp2demonstrates significant advantages over the baseline method on the LLAMA model. For hardware evaluation, we implement our design at the RTL level and evaluate it on both FPGA and ASIC platforms. Compared to computation architectures based onfp4, ourfp4$\times $fp2processing element (PE) array reduces area by 15% and power consumption by 8%. Furthermore, ourfp2$\times $fp2PE array achieves a remarkable 78% reduction in both area and power consumption.
Qiwei Dang, Chengyu Ma, Haiduo Huang, Gelin Fu, Zhiwang Huo, Guoming Yang, Pengchen Zong, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Magellan: A High-Performance Loop-Guided Prefetcher for Indirect Memory Access
abstract
Graph analytics and sparse linear algebra applications heavily rely on indirect memory access (IMA).IMAs are characterized by poor temporal and spatial locality, which causes frequent high-latency DRAM accesses.While dedicated hardware prefetchers for IMA have been explored, they target narrow access patterns and tend to introduce significant hardware complexity.Software prefetching offers a promising alternative, leveraging compiler analysis to prefetch indirection patterns.However, existing software prefetchers struggle with sparse applications due to limited loop iterations and complex IMA patterns across nested loops.We propose Magellan, a novel loop-guided software prefetcher designed to detect and schedule IMA prefetches efficiently.Magellan introduces two key innovations: (1) extracting dependence graphs across loop levels to detect complex IMA patterns and (2) capturing inner-outer loop semantics to prefetch for both current and future iterations.We evaluate Magellan on 14 memory-intensive benchmarks using real-world datasets from social networks and web graphs.Compared to the best existing IMA software prefetcher, Magellan reduces cache misses by 25% and dynamic instruction counts by 14% on average.This results in a 1.14× average speedup, with performance gains of up to 1.41×.
Gelin Fu, Tian Xia 0008, Mingzhuo Yin, Prashant J. Nair, Mieszko Lis, Pengju Ren
ISCA1
2024 Differential-Matching Prefetcher for Indirect Memory Access
abstract
Indirect memory access is a critical bottleneck for modern CPUs, especially for graph analysis and sparse linear algebra applications, where the values of one data array are used to generate the fetching addresses of another array. It often causes irregular data accesses with poor temporal and spatial locality that are difficult to be captured by conventional hardware prefetchers. For many complex workloads, such indirect access patterns may have different types and are nested in a multiplelevel form. Moreover, branch mispredictions would further disturb their patterns, making them even harder to detect. As a result, existing hardware prefetchers are unable to fully prefetch complex indirect patterns. This paper proposes DMP, a low-cost hardware prefetcher to improve the memory latency in several representative irregular workloads. DMP targets four types of indirect memory access patterns including single, range, multi-level, and multi-way indirect access. DMP uses differential matching to identify an indirect access pattern in pair with its corresponding index stream. Then DMP uses a flexible prefetching mechanism to dynamically adapt the prefetching degree to maintain prefetching coverage. We evaluate the performance, energy consumption, and transistor cost of DMP among various algorithms from GAP, NAS, and HPCG benchmarks. DMP improves performance by 1.8 × (up to 5.6 ×) on average against state-of-the-art hardware prefetchers and 1.2 × (up to 2.3 ×) speedup against state-of-the-art compiler-based prefetcher Prodigy. Besides, the proposed design is optimized to take only 0.9KB of storage, making it feasible to be integrated into current CPU designs.
Gelin Fu, Tian Xia 0008, Zhongpei Luo, Wenzhe Zhao 0001, Pengju Ren
HPCA1
2023 PrSpMV: An Efficient Predictable Kernel for SpMV
abstract
Sparse Matrix-Vector Multiplication (SpMV) has been widely applied in scientific computation, industry simulation, and intelligent computation domains, which is the critical algorithm in all these applications. Due to the poor data locality, low cache usage, and extremely irregular branch patterns caused by the highly sparse and random distributions, SpMV optimization has become one of the most challenging problems for modern high-performance processors. In this paper, we study the bottlenecks of SpMV on current out-of-order CPUs and propose a novel SpMV kernel named PrSpMV to improve its performance by pursuing high predictability. Specifically, we improve the memory access regularity and locality by creating serialized access patterns so that the data prefetching efficiency and cache usage are optimized. We also improve pipeline efficiency by creating regular branch patterns to make branch prediction more accurate. Experiment results show that using the above optimization approaches, PrSpMV can eliminate nearly all branch mispredictions. Moreover, it can also significantly reduce the average L2 cache miss rate from 57% to 20% via efficiently leveraging hardware prefetchers. By using PrSpMV, stride prefetcher can be boosted with 1.31× speedup and dedicated irregular prefetcher can be improved with 1.40× speedup. Meanwhile, on commercial high-end Intel processors, it achieves 1.32× speedup against some state-of-the-art SpMV kernels.
Gelin Fu, Tian Xia 0008, Shaoru Qu, Zhongpei Luo, Pengyu Cheng, Runfan Guo, Yitong Ding, Pengju Ren
ICCD1
2023 An Energy-and-Area-Efficient CNN Accelerator for Universal Powers-of-Two Quantization
abstract
CNN model computation on edge devices is tightly restricted to the limited resource and power budgets, which motivates the low-bit quantization technology to compress CNN models into 4-bit or lower format to reduce the model size and increase hardware efficiency. Most current low-bit quantization methods use uniform quantization that maps weight and activation values onto evenly-distributed levels, which usually results in accuracy loss due to distribution mismatch. Meanwhile, some non-uniform quantization methods propose specialized representation that can better match various distribution shapes but are usually difficult to be efficiently accelerated on hardware. In order to achieve low-bit quantization with high accuracy and hardware efficiency, this paper proposes Universal Power-of-Two (UPoT), a novel low-bit quantization method that represents values as the addition of multiple power-of-two values selected from a series of subsets. By updating the subset contents, UPoT can provide adaptive quantization levels for various distributions. For each CNN model layer, UPoT automatically searches for the optimized distribution that minimizes the quantization error. Moreover, we design an efficient accelerator system with specifically optimized power-of-two multipliers and requantization units. Evaluations show that the proposed architecture can provide high-performance CNN inference with reduced circuit area and energy, and outperforms several mainstream CNN accelerators with higher ($8\times $–$65\times $) area efficiency and ($2\times $–$19\times $) energy efficiency. Further experiments of 4/3/2-bit quantization on ResNet18/50, MobileNet_V2 and EfficientNet models show that our UPoT can achieve high model accuracy which greatly outperform other state-of-the-art low-bit quantization methods by 0.3%–6%. The results indicate that our approach provides a highly-efficient accelerator for low-bit CNN model quantization with low hardware overheads and good model accuracy.
Tian Xia 0008, Boran Zhao, Gelin Fu, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 A Comprehensive Performance Model of Sparse Matrix-Vector Multiplication to Guide Kernel Optimization
abstract
Sparse Matrix-Vector Multiplication (SpMV) is important in scientific and industrial applications and remains a well-known challenge for modern CPUs due to high sparsity and irregularity. Many researchers try to improve SpMV performance by designing dedicated data formats and computation patterns. However, out-of-order superscalar CPUs have complex micro-architectures where exist complicated interactions and restrictions among software and hardware factors. It is hard to systematically study the effectiveness of optimization methods on the overall performance, as its benefits may be undermined by other factors. In this paper, we thoroughly study the execution of SpMV on modern CPUs and propose a comprehensive performance model to reveal the critical factors and their relationships. Specifically, we first study the coding characteristics of SpMV kernels to identify key factors worthy of attention. Then we model the execution of SpMV as two overlapped parts: CPU pipeline and memory latency. Both are carefully modeled with related hardware and software factors. We also model SIMD performance with the usage of specific SIMD instructions and vector registers. Experiments show that our model matches the actual execution of real-world processors. Guided by the model, we propose SpV8, a novel SpMV kernel that optimizes critical factors to improve computation efficiency and memory bandwidth. Experiments on Intel/AMD x86 and ARM AArch64 platforms show that SpV8 outperforms several state-of-the-art approaches with large margins, achieving average$3.4\times$over Intel Math Kernel Library and$1.4\times$over the best existing approach. Such results indicate that the proposed model is capable of valuable guidance for efficient SpMV optimizations.
Tian Xia 0008, Gelin Fu, Zhongpei Luo, Lucheng Zhang, Wenzhe Zhao 0001, Nanning Zheng 0001, Pengju Ren
IEEE Trans. Parallel Distributed Syst.2