Zhimin Zhang 0004

dblp:19/3058-4 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0000-8778-7149ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 JPnR: A Length-Matching Placement and Routing Framework for Single-Flux-Quantum Circuits
abstract
Superconducting rapid single-flux-quantum (RSFQ) logic is a promising candidate for advancing future computing technologies due to its low-energy consumption and high-frequency capabilities. However, precise timing alignment is crucial for its physical design, posing significant challenges in length-matching placement and routing. This paper introduces JPnR, a physical design framework tailored for RSFQ circuits, featuring a clock-aware length-matching placer and a length-matching multi-terminal router. The placer simultaneously considers both clock distribution and timing constraints, distributing clock pulses heuristically and transforming the placement problem into a single-source shortest-path problem. This allows it to minimize vertical wirelength using dynamic programming and iteratively optimize placement via a barycenter-like reordering method. The router tackles challenges related to splitter placement and length-matching multi-terminal routing using a two-layer planar Manhattan routing model. Initial routing assigns tracks based on the left-edge algorithm to minimize routing width while employing the dogleg algorithm to resolve cycles in the vertical constraint graph. Length-matching is achieved via a splitter tree-based hierarchical approach with maximum-flow-based detour insertion. Finally, a PTL region expansion strategy is employed for unsatisfied connections. Experimental results on RSFQ benchmarks demonstrate the effectiveness and efficiency of JPnR.
Rongliang Fu, Minglei Zhou, Xinda Chen, Junying Huang, Xiaochun Ye, Zhimin Zhang 0004, Tsung-Yi Ho
IEEE Trans. Computers7
2025 A GCN Accelerator with Unified Architecture
Meng Wu 0006, Mingyu Yan, Lei Deng 0003, Zhimin Zhang 0004, Xiaochun Ye, Dongrui Fan
ICA3PP (1)5
2025 JBSA: A Bit-Serial Accelerator for Deep Neural Networks Using Superconducting SFQ Logic
abstract
The potential of superconducting single flux quantum (SFQ) devices in accelerating deep neural networks (DNNs) has garnered significant attention due to their ultra-fast and lowpower switching capabilities.However, existing SFQ-based DNN accelerators face limitations in scaling up to larger-scale instances due to the stringent area constraints and complex architectures.Additionally, another challenge in SFQ-based DNN acceleration lies in bridging the gap between the ultrahigh computing speed offered by SFQ technology and the relatively low memory bandwidth.To address these challenges, we propose JBSA, an SFQ-based bit-serial accelerator for DNN inference acceleration.JBSA leverages bit-serial computing to alleviate area constraints and reduce bandwidth requirements.A bit-serial processing element is designed to implement multiply-accumulate operations using SFQ logic cells.
Huilong Jiang, Haofei Yin, Rongliang Fu, Junying Huang, Xiaochun Ye, Zhimin Zhang 0004, Tsung-Yi Ho, Dongrui Fan
ICS8
2024 JPlace: A Clock-Aware Length-Matching Placement for Rapid Single-Flux-Quantum Circuits
abstract
Superconducting rapid single-flux-quantum (RSFQ) logic has emerged as a promising candidate for future computing technology, owing to its low power consumption and high frequency characteristics. Given its ultra-high frequency operation, achieving precise timing alignment is crucial for RSFQ circuit physical design. To address the timing issue, this paper introduces JPlace, a clock-aware length-matching placement framework for RSFQ circuits. JPlace simultaneously addresses data and clock signal length matching, effectively ensuring accurate timing alignment and mitigating timing alignment challenges during the routing phase. We propose a heuristic method for constructing the clock distribution and a dynamic programming-based approach for minimizing the total vertical wirelength while maintaining fixed placement orders. Additionally, we introduce a barycenter-based reordering method to further explore the solution space and reduce wirelength. Experimental results on the RSFQ benchmark demonstrate the effectiveness and efficiency of JPlace.
Rongliang Fu, Junying Huang, Zhimin Zhang 0004, Xiaochun Ye, Tsung-Yi Ho, Dongrui Fan
DATE4
2024 Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
Meng Wu 0006, Jingkai Qiu, Mingyu Yan, Yang Zhang 0163, Zhimin Zhang 0004, Xiaochun Ye, Dongrui Fan
ICA3PP (3)6
2023 JRouter: A Multi-Terminal Hierarchical Length-Matching Router under Planar Manhattan Routing Model for RSFQ Circuits
abstract
Superconducting rapid single-flux-quantum (RSFQ) logic has shown great potential for high-energy-efficient computing systems. To ensure correct operations at ultra-high frequencies, it is necessary to incorporate length-matching constraints into the routing problem. Existing routing algorithms, however, can only address 2-pin connections or support the conventional horizontal/vertical routing model, which substantially limits the optimization space for routing solutions. This paper presents JRouter, an RSFQ router that considers the two-layer planar Manhattan routing model while simultaneously coping with splitter (SPL) placement and length-matching multi-terminal routing. JRouter contains a track-assignment-based initial routing that minimizes the initial routing width while avoiding conflicts in the horizontal constraint graph. Moreover, JRouter implements an SPL-tree-based hierarchical routing with an iterative maximum-flow-based formulation to insert the detours for multi-terminal routing. A routing region extension algorithm is also developed to insert the detours for unsatisfied connections. According to the experimental results, JRouter achieves an average routing width reduction of 35.71% and 22.46% on a 16-bit RSFQ Sklansky adder compared to Kito's and Kou's routing algorithms. For randomly generated benchmarks, JRouter reduces the routing width by an average of 38.77%, 38.20%, 21.65%, and 7.01% compared to Kito's, Kou's, and two of Yan's routing algorithms, respectively, while maintaining reasonable runtime.
Xinda Chen, Rongliang Fu, Junying Huang, Huawei Cao, Zhimin Zhang 0004, Xiaochun Ye, Tsung-Yi Ho, Dongrui Fan
ACM Great Lakes Symposium on VLSI5
2023 Design of a Compact Superconducting RSFQ Register File
abstract
In comparison to the widely-used CMOS circuits, superconducting Rapid Single Flux Quantum (RSFQ) circuits offer advantages such as fast operating frequency and low power consumption, making them a potential direction for development in the post-Moore era of digital circuits. However, designing RSFQ CPUs faces challenges, one of which is the need for a compact register file. This is because, under current RSFQ circuit process conditions, the memory module typically occupies a large chip area. This paper proposes a newly-designed Write-controllable Non-Destructive Read Out cell (WNDRO) that can limit the data writing pulse input to change the internal state of the cell. Based on the WNDRO, a compact RSFQ register file is designed that can realize random, non-destructive data reading. Additionally, a global write strategy is applied to omit the data write routing circuit and the reset control module, and the circuit design of the read and write control module is optimized. In comparison to general designs, the proposed compact register file design will save a large number of Josephson Junctions and effectively reduce the chip area. Additionally, this paper proposes several optimization logic designs for RSFQ CPU designers as a reference.
Kuozhong Zhang, Zhimin Zhang 0004, Guang-Ming Tang, Xiaochun Ye
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Equivalence Checking for Superconducting RSFQ Logic Circuits
abstract
Equivalence checking is a key component of the verification methodology for digital circuit designs. In this paper, we propose an equivalence checking framework for superconducting rapid single-flux-quantum (RSFQ) logic circuits which include acyclic circuits and bit-slice-based cyclic circuits. It consists of a structure checker and a logic checker. The structure checker is used to check whether the circuit meets the design rules of superconducting RSFQ logic circuits. The logic checker can be used to check whether two RSFQ gate-level circuits have the same logic function. For the logic checker, we propose a logic equivalence checking method based on logic cone partition. The circuit network is simplified layer by layer and iteratively partitioned into logic cones, each of which is verified by the SMT solver. The experimental results show the feasibility of our approach on superconducting RSFQ logic circuits.
Rongliang Fu, Junying Huang, Zhimin Zhang 0004
ACM Great Lakes Symposium on VLSI3
2020 Design Automation Methodology from RTL to Gate-level Netlist and Schematic for RSFQ Logic Circuits
abstract
The superconducting rapid single flux quantum (RSFQ) logic circuit has the characteristics of high speed and low power consumption, making it an attractive candidate for future supercomputers. However, computer-aided design (CAD) tools for CMOS cannot be directly applied to RSFQ logic due to their distinct properties. For instance, the RSFQ logic gate can work properly when all its fan-ins have the same logic level. This paper presents the design flow from RTL to RSFQ logic netlist and schematic. First, we implement logic synthesis for RSFQ logic circuits. It achieves path balancing while minimizing the number of DFFs. In addition, we propose an automatic schematic generator for the RSFQ logic circuits. It converts the synthesized netlist into its equivalent schematic. A layer assignment algorithm is proposed, which makes all gates layered in the order of the clock arrival time. Experimental results with ISCAS85 and EPFL benchmarks along with some Kogge-Stone adders have shown a 29.2% reduction in the number of DFFs over the breadth-first first search; moreover, 59.57% and 5.3% decrease in the number of layers of the schematic and number of edge crossings over the ELK tool.
Rongliang Fu, Zhimin Zhang 0004, Guang-Ming Tang, Junying Huang, Xiaochun Ye, Dongrui Fan, Ninghui Sun
ACM Great Lakes Symposium on VLSI2
2020 HyGCN: A GCN Accelerator with Hybrid Architecture
abstract
Inspired by the great success of neural networks, graph convolutional neural networks (GCNs) are proposed to analyze graph data. GCNs mainly include two phases with distinct execution patterns. The Aggregation phase, behaves as graph processing, showing a dynamic and irregular execution pattern. The Combination phase, acts more like the neural networks, presenting a static and regular execution pattern. The hybrid execution patterns of GCNs require a design that alleviates irregularity and exploits regularity. Moreover, to achieve higher performance and energy efficiency, the design needs to leverage the high intra-vertex parallelism in Aggregation phase, the highly reusable inter-vertex data in Combination phase, and the opportunity to fuse phase-by-phase execution introduced by the new features of GCNs. However, existing architectures fail to address these demands. In this work, we first characterize the hybrid execution patterns of GCNs on Intel Xeon CPU. Guided by the characterization, we design a GCN accelerator, HyGCN, using a hybrid architecture to efficiently perform GCNs. Specifically, first, we build a new programming model to exploit the fine-grained parallelism for our hardware design. Second, we propose a hardware design with two efficient processing engines to alleviate the irregularity of Aggregation phase and leverage the regularity of Combination phase. Besides, these engines can exploit various parallelism and reuse highly reusable data efficiently. Third, we optimize the overall system via inter-engine pipeline for inter-phase fusion and priority-based off-chip memory access coordination to improve off-chip bandwidth utilization. Compared to the state-of-the-art software framework running on Intel Xeon CPU and NVIDIA V100 GPU, our work achieves on average 1509× speedup with 2500× energy reduction and average 6.5× speedup with 10× energy reduction, respectively.
Mingyu Yan, Lei Deng 0003, Xing Hu 0001, Ling Liang 0003, Yujing Feng, Xiaochun Ye, Zhimin Zhang 0004, Dongrui Fan, Yuan Xie 0001
HPCA7
2019 iATPG: Instruction-level Automatic Test Program Generation for Vulnerabilities under DVFS attack
abstract
With the growing cost of powering and cooling, the Dynamic Voltage Frequency Scaling (DVFS) technique has been adopted in many mobiles and embedded devices nowadays. However, attackers are capable of maliciously manipulating the DVFS to threaten application programs including the security related ones. This paper proposes an instruction-level Automatic Test Program Generation (iATPG) framework, which generates test programs to test the vulnerabilities of CPU instructions under the DVFS attack. The conditions that the test program needs to meet, the testability of CPU instructions, and the iATPG algorithm are proposed. It is applied to an arm CPU in a mobile phone. Typical instructions are tested, and some are found vulnerable. The application programs using these instructions are then attacked to prove the effectiveness of the proposed framework.
Kuozhong Zhang, Junying Huang, Jing Ye 0001, Xiaochun Ye, Dongrui Fan, Huawei Li 0001, Xiaowei Li 0001, Zhimin Zhang 0004
IOLTS9
2019 Balancing Memory Accesses for Energy-Efficient Graph Analytics Accelerators
abstract
Domain-specific accelerators for graph analytics leverage a large on-chip memory in order to tackle the intensive random memory accesses, offering higher performance and energy efficiency than conventional architectures. However, limited by the inefficient usage of on-chip memory, current accelerators suffer from energy and performance bottlenecks due to the large amount of off-chip memory accesses. In this work, we introduce an online preprocessing step for the vertex-centric programming model based on our observation of imbalanced memory bandwidth utilization between two execution phases. Our scheme improves energy efficiency and performance by significantly reducing off-chip accesses in two ways. First, we sequence random off-chip memory accesses to balance memory bandwidth demands and improve the utilization of on-chip memory. Second, we prune active leaf vertices to avoid redundant memory accesses. We evaluate our method on a state-of-the-art graph analytics accelerator and achieve 1.6× speedup while reducing energy consumption by 42% on average.
Mingyu Yan, Xing Hu 0001, Shuangchen Li, Itir Akgun, Han Li 0011, Lei Deng 0003, Xiaochun Ye, Zhimin Zhang 0004, Dongrui Fan, Yuan Xie 0001
ISLPED9
2019 Instruction Vulnerability Test and Code Optimization Against DVFS Attack
abstract
With the growing cost of powering and cooling, the Dynamic Voltage Frequency Scaling (DVFS) technique has been adopted in many mobiles and embedded devices nowadays. However, attackers are capable of maliciously manipulating the DVFS to threaten application programs including the security related ones. This paper first proposes a test method to test the vulnerabilities of CPU instructions under the DVFS attack. The test program feature, the testability of CPU instructions, and the Test Program Generation Algorithm (TPGA) are proposed. It is applied to an arm CPU in a mobile phone. Typical instructions are tested, and some are found vulnerable. Then, based on the test result, a method for code optimization by instruction substitution is proposed. The application program using vulnerable instructions are then attacked and optimized to prove the effectiveness of the proposed methods.
Junying Huang, Jing Ye 0001, Xiaochun Ye, Dongrui Fan, Huawei Li 0001, Xiaowei Li 0001, Zhimin Zhang 0004
ITC-Asia8
2019 Alleviating Irregularity in Graph Analytics Acceleration: a Hardware/Software Co-Design Approach
abstract
Graph analytics is an emerging application which extracts insights by processing large volumes of highly connected data, namely graphs. The parallel processing of graphs has been exploited at the algorithm level, which in turn incurs three irregularities onto computing and memory patterns that significantly hinder an efficient architecture design. Certain irregularities can be partially tackled by the prior domain-specific accelerator designs with well-designed scheduling of data access, while others remain unsolved.
Mingyu Yan, Xing Hu 0001, Shuangchen Li, Abanti Basak, Han Li 0011, Itir Akgun, Yujing Feng, Peng Gu 0008, Lei Deng 0003, Xiaochun Ye, Zhimin Zhang 0004, Dongrui Fan, Yuan Xie 0001
MICRO12
2018 A Non-Stop Double Buffering Mechanism for Dataflow Architecture
Xu Tan 0001, Xiaochun Ye, Dongrui Fan, Lunkai Zhang, Zhimin Zhang 0004
J. Comput. Sci. Technol.8
2017 An Efficient Network-on-Chip Router for Dataflow Architecture
Xiaochun Ye, Xu Tan 0001, Lunkai Zhang, Zhimin Zhang 0004, Dongrui Fan, Ninghui Sun
J. Comput. Sci. Technol.7
2016 POSTER: An Optimization of Dataflow Architectures for Scientific Applications
abstract
Dataflow computing is proved to be promising in high-performance computing. However, traditional dataflow architectures are general-purpose and not efficient enough when dealing with typical scientific applications due to low utilization of function units. In this paper, we propose an optimization of dataflow architectures for scientific applications. The optimization introduces a request for operands mechanism and a topology-based instruction mapping algorithm to improve the efficiency of dataflow architectures. Experimental results show that the request for operands optimization achieves a 4.6% average performance improvement over the traditional dataflow architectures and the TBIM algorithm achieves a 2.28x and a 1.98x average performance improvement over SPDI and SPS algorithm respectively.
Xiaochun Ye, Xu Tan 0001, Zhimin Zhang 0004, Dongrui Fan
PACT5