Mingjia Fan

dblp:352/3277 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0002-1987-3173ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Forest to Tree: Prioritizing the Maximum Additional Delay in AQFP Circuit Design
abstract
This paper presents a fast and scalable algorithm for buffer and splitter insertion in AQFP circuits. The method maps each wire to a homeomorphic graph, constructs an additional-delay-free multi-ary forest, and merges it into an optimal tree under delay and fanout constraints. The formulation guarantees per-wire optimality in terms of maximum additional delay, total additional delay, and internal node count. A circuit-level refinement further reduces redundant insertion by identifying and adjusting critical wires. On standard AQFP benchmarks, the proposed approach achieves 2.72×, 525.70×, and 1.33× speedups over [1], [2], and [3], respectively, while maintaining comparable insertion counts and logic depths.
Yinuo Bai 0002, Mingjia Fan, Tsung-Yi Ho, Zhou Jin 0001
DATE2
2025 ReRAM-Based Process-In-Memory Accelerator for Iterative Solvers: A Systematic Survey
abstract
Iterative solvers are fundamental in scientific computing, particularly for solving large-scale linear equations, which are central to a variety of applications such as simulations and data analysis. Traditional optimization strategies for iterative solvers, however, are predominantly designed around von Neumann architectures, which suffer from significant data movement costs and the "memory wall" problem, limiting overall computational performance. In this context, processing-in-memory (PIM) architectures, especially those utilizing resistive random-access memory (ReRAM), offer a promising alternative by enabling in-situ computing, thereby reducing data movement and overcoming the storage bottleneck. These architectures have already shown substantial potential in accelerating tasks like neural network training and graph computations, and they provide new opportunities for optimizing iterative solvers. This paper systematically surveys ReRAM-based iterative solver accelerators, categorizing key contributions into four main areas: mixed-precision techniques, feedback circuit theory, floating-point computation support, and leveraging content-addressable memory (CAM) to address irregularity and sparsity. We also discuss four future research directions aimed at further improving iterative solver performance.
Boyu Geng, Mingjia Fan, Zhou Jin 0001, Weifeng Liu 0002
ISCAS2
2024 ReCG: ReRAM-Accelerated Sparse Conjugate Gradient
abstract
Solving sparse linear systems is crucial in scientific computing. Sparse Conjugate Gradient (CG) is one of the most well-known iterative solvers with high efficiency and low storage requirements. However, the performance of sparse CG solvers implemented on storage-compute separated architectures is greatly limited by the irregular memory access and the large amount of data transmission.
Mingjia Fan, Xiaoming Chen 0003, Dechuang Yang, Zhou Jin 0001, Weifeng Liu 0002
DAC1
2023 AmgR: Algebraic Multigrid Accelerated on ReRAM
abstract
Solving systems of linear equations is a fundamental problem in scientific computing, which has been extensively researched for decades. One of the most well-known solvers is Algebraic Multigrid (AMG), which is widely used in high performance computing due to its good scalability. But currently accelerating AMG relies on the traditional von Neumann architecture of storage and computation separation, which leads to a large data transmission overhead. In this work, we propose a ReRAM-based processing-in-memory (PIM) architecture named AmgR, which overcomes the limitations of the traditional von Neumann architecture for AMG acceleration.However, accelerating AMG on ReRAM is non-trivial, because (1) AMG has many computing kernels of various types; (2) there are irregular operations that cannot be directly performed using matrix-vector multiplication suitable for ReRAM, i.e., aggregation operation; (3) ReRAM has poor write endurance, and a lot of data during AMG acceleration needs to be rewritten into ReRAM, resulting in high write cost. To address these issues, firstly, we propose a flexible architecture, which can realize each kernel of AMG and is reused by many kernels to improve resource utilization. Secondly, we propose a dedicated unit to realize the aggregation operation. Finally, we present a new mapping strategy to greatly reduce the number of data handling and writes. The experimental results show that the performance of AmgR is improved by an average of one and two orders of magnitude compared to HYPRE on the CPU and AmgX on the GPU, respectively, while the energy consumption is reduced by an average of two and three orders of magnitude.
Mingjia Fan, Xiaotian Tian, Yintao He, Yiru Duan, Xiaozhe Hu, Ying Wang 0001, Zhou Jin 0001, Weifeng Liu 0002
DAC1