Siqing Fu

dblp:309/9653 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-7732-4140ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HIVE+: An Enhanced High-Priority Victim Cache to Accelerate GPU Memory Accesses
Yuhan Tang, Sheng Ma, Hanqing Li, Shengbai Luo, Jixuan Tang, Siqing Fu, Lizhou Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2025 SFC-TRNG: A Stable Fast and Compact True Random Number Generator based on Magnetic Tunnel Junction
Siqing Fu
ACM Great Lakes Symposium on VLSI4
2025 NeuroPDE: A Neuromorphic PDE Solver Based on Spintronic and Ferroelectric Devices
abstract
In recent years, new methods for solving partial differential equations (PDEs) such as Monte Carlo random walk methods have gained considerable attention. However, due to the lack of hardware-intrinsic randomness in the conventional von Neumann architecture, the performance of PDE solvers is limited. In this paper, we introduce NeuroPDE, a hardware design for neuromorphic PDE solvers that utilizes emerging spintronic and ferroelectric devices. NeuroPDE incorporates spin neurons that are capable of probabilistic transmission to emulate random walks, along with ferroelectric synapses that store continuous weights non-volatilely. The proposed NeuroPDE achieves a squared error of less than 1e-2 compared to analytical solutions when solving diffus3.48× to 315× speedup in execution time and an energy consumption advantage of 2.7× to 29.8× over advanced CMOS-based neuromorphic chips. By leveraging the inherent physical stochasticity of emerging devices, this study paves the way for future probabilistic neuromorphic computing systems.
Siqing Fu, Lizhou Wu, Chunyuan Zhang, Sheng Ma, Yuhan Tang, Jixuan Tang
ICCAD1
2025 CHQ-SC: Compact and High-Quality Stochastic Computing Framework Using Magnetic Tunnel Junction
abstract
Stochastic computing (SC) offers a compact and low-power alternative to traditional binary computing. However, CMOS-based SC using linear feedback shift registers (LFSRs) suffers from pseudo-randomness and excessive area overhead. Magnetic tunnel junction (MTJ)-based solutions leverage physical randomness but face low throughput and implementation complexity. This paper proposes CHQ-SC, a novel MTJ-based framework with key innovations: a dual-MTJ pipelined architecture achieving$312.5 M b / s$throughput; a shared circuit design enabling a compact$12.24 \mu \mathrm{m}^{2}$layout and$4.02 p J / b i t$energy; and direct MTJ true random number generator (TRNG) integration solving existing issues. In target localization tasks, the Kullback-Leibler divergence between CHQ-SC's outputs and ideal computation decreases to 0.025, approaching theoretical limits. CHQ-SC provides a practical solution for high-efficiency, high-accuracy SC systems.
Siqing Fu
ICCD4
2023 RHS-TRNG: A Resilient High-Speed True Random Number Generator Based on STT-MTJ Device
abstract
High-quality random numbers are very critical to many fields such as cryptography, finance, and scientific simulation, which calls for the design of reliable true random number generators (TRNGs). Limited by entropy source, throughput, reliability, and system integration, existing TRNG designs are difficult to be deployed in real computing systems to greatly accelerate target applications. This study proposes a TRNG circuit named resilient high-speed (RHS)-TRNG based on spin-transfer torque magnetic tunnel junction (STT-MTJ). RHS-TRNG generates resilient and high-speed random bit sequences exploiting the stochastic switching characteristics of STT-MTJ. By circuit/system codesign, we integrate RHS-TRNG into a reduced instruction set computer-V (RISC-V) processor as an acceleration component, which is driven by customized random number generation instructions. Our experimental results show that a single cell of RHS-TRNG has a random bit generation speed of up to 303 Mb/s, which is the highest among existing MTJ-based TRNGs. Higher throughput can be achieved by exploiting cell-level parallelism. RHS-TRNG also shows strong resilience against PVT variations thanks to our designs using bidirectional switching currents and dual generator units. In addition, our system evaluation results using gem5 simulator suggest that the system equipped with RHS-TRNG can achieve 3.4–$12\times $higher performance in speeding up option pricing programs than software implementations of random number generation.
Siqing Fu, Chunyuan Zhang, Hanqing Li, Sheng Ma, Lizhou Wu
IEEE Trans. Very Large Scale Integr. Syst.1
2021 A Variable-Way Address Translation Cache for the Exascale Supercomputer
Siqing Fu, Kefan Ma
ICA3PP (3)3
2021 An Incomplete Unsatisfiable Cores Extracting Algorithm to Promote Routing
abstract
Many practical applications need to localize the fundmental causes for infeasibility or failure in different applications, including formal method, electronic design automation and artificial intelligence. A paradigmatic application is finding the unroutable reason in FPGA (Field Programmable Gate Array). A small unsatisfiable subset can promote the routing tool to diagnose and localize the unroutable region and nets. For this application, an incomplete unsatisfiable cores extracting algorithm is proposed and integrated into Boolean-based routing method. The refutation-based algorithm integrates many powerful heuristics such as subsumption elimination technique. The optimal extracting minimum unsatisfiable cores algorithm called branch-and-bound algorithm is employed to compare with the incomplete algorithm. From the evaluated results on the standard FPGA routing benchmarks, it is concluded the incomplete algorithm run faster than the branch-and-bound algorithm, and two algorithms obtain the same size of unsatisfiable cores.
Siqing Fu
QRS3