Jian Zhang 0022

dblp:07/314-22 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0002-8353-6243ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Interconnection networks and networks-on-chip · 63% Integrated circuit design · 13% Emerging computing paradigms · 12%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip › deadlock handling
deadlock recovery
0.912025
Steered Bubble: An Interposer-based Deadlock Recovery Algorithm for Multi-chiplet Systems · ACM Trans. Archit. Code Optim. 2025
Interconnection networks and networks-on-chip › network-on-chip design
interposer-based noc
0.912025
Steered Bubble: An Interposer-based Deadlock Recovery Algorithm for Multi-chiplet Systems · ACM Trans. Archit. Code Optim. 2025
Hardware accelerators and domain-specific architectures › domain-specific accelerator › optimization accelerator
combinatorial optimization accelerator
0.312018
Advancing CMOS-Type Ising Arithmetic Unit into the Domain of Real-World Applications · IEEE Trans. Computers 2018
Integrated circuit design › heterogeneous integration
chiplet integration
0.312025
Steered Bubble: An Interposer-based Deadlock Recovery Algorithm for Multi-chiplet Systems · ACM Trans. Archit. Code Optim. 2025
Integrated circuit design
digital circuit design
0.112018
Advancing CMOS-Type Ising Arithmetic Unit into the Domain of Real-World Applications · IEEE Trans. Computers 2018

Methods — techniques the papers use, named apart from their topics

congestion-sense network · 0.9bubble flow control · 0.9simulated annealing · 0.3
YearPublicationVenuePosition
2025 Steered Bubble: An Interposer-based Deadlock Recovery Algorithm for Multi-chiplet Systems
abstract
Dividing a single System-on-Chip (SoC) into multiple chiplets and integrating them via an interposer can achieve an optimal balance between continuous transistor integration and monetary cost. However, potential deadlock may arise between the chiplets and the interposer. This deadlock can be avoided by applying turn restriction or injection control on the boundary routers, at the cost of additional latency and suboptimal performance. Compared to deadlock avoidance, deadlock recovery exerts less impact on network performance. Nevertheless, accurate and timely deadlock detection, along with efficient deadlock recovery, continues to pose significant challenges. Additionally, modularity is a specific concern, which involves integrating chiplets of various functions, sizes, manufacturing processes, and so on. Minimizing the negative impact of deadlock resolution while maximizing modularity is crucial for achieving the benefit of chiplets. This article proposes a modular deadlock detection strategy, Up-Down, which monitors both the upward and downward directions of vertical channels, facilitating information exchange through the congestion-sense network. When a pair of blocked upward and downward vertical channels is detected simultaneously, it is considered that an inter-chiplet deadlock has occurred. This significantly enhances the accuracy of deadlock detection by two orders of magnitude compared to time-out deadlock detection. Furthermore, this article introduces Steered Bubble, a low-cost deadlock recovery algorithm. It does so by injecting bubbles into potential deadlock cycles identified by Up-Down. These bubbles follow preset paths, ensuring efficient deadlock recovery. Experimental results indicate that the Steered Bubble results in an average performance enhancement of 1% to 10% during full-system simulations, with an area overhead of less than 2%.
Zhiqiang Chen 0006, Yongwen Wang, Jian Zhang 0022
ACM Trans. Archit. Code Optim.4
2021 Advancing DSP into HPC, AI, and beyond: challenges, mechanisms, and future directions
Chen Li 0015, Chang Liu 0019, Sheng Liu 0001, Yuanwu Lei, Jian Zhang 0022, Yang Guo 0003
CCF Trans. High Perform. Comput.6
2018 Live Demonstration: Image Segmentation on the FPGA-based Pre-calculating Ising Memory
abstract
We demonstrate image segmentation processing by using a pre-calculating Ising memory implemented on a FPGA. Results show that the FPGA-based pre-calculating Ising memory can segment a prepared image in under 100μs.
Jian Zhang 0022, Shuming Chen, Lei Wang 0011, Linghui Lv
ISCAS1
2018 Pre-Calculating Ising Memory: Low Cost Method to Enhance Traditional Memory with Ising Ability
abstract
Combinatorial optimization always contains many state search operations, which greatly reduce the efficiency of Von Neumann architecture. The Ising chip, expressing the behavior of magnetic spin systems with CMOS circuit, can efficiently support such operations. On the Ising chip, the state search can be carried out for all the spins in parallel. As the Ising chip is mainly SRAM based architecture, we propose Ising memory that enhancing the traditional memory with Ising ability, which can be easily integrated into Von Neumann architecture for both traditional data storage and efficiently solving combinatorial optimization problems. However, due to the non-memory logic for state search operations, directly integrating Ising ability into traditional memory would introduce additional 2× area overhead. To solve this problem, we propose pre-calculating structure to reduce the complexity of the state search circuit. Our proposal helps to reduce the non-memory area overhead to about 0.9× of the traditional memory. Moreover, we have physically designed an Ising memory and tested it with image segmentation problems. Our Ising memory can accelerate the segmenting processing by 26000× with only 0.0001× energy consumption. The experiment result shows that our Ising memory is a low cost method to enhance traditional memory with Ising ability for both data storage and solving combinatorial optimization problems.
Jian Zhang 0022, Shuming Chen, Lei Wang 0011, Linghui Lv
ISCAS1
2018 Advancing CMOS-Type Ising Arithmetic Unit into the Domain of Real-World Applications
abstract
Solving combinatorial optimization problems is a great challenge for Von Neumann-architecture computing. Although the Ising model could provide promising solutions for such problems, existing Ising chips, including superconductive, optical and CMOS-type circuit implementation, cannot meet the precision requirement of real-world combinatorial optimization applications. To facilitate the support for real-world applications, we propose three improvements over existing CMOS-type Ising chips: suitable narrow bit width memory cells with approximate multiply-adders, double random sources flipping method with cross random number generators and shared circuit design between adjacent spin nodes. With above improvements, we achieve high precision as well as maintaining the low cost characteristic of CMOS-type Ising chips. When searching the ground state of Ising models, our CMOS-type Ising chip can improve the precision to more than 99 percent over existing ones with about 93 percent precision. Moreover, its hardware cost is only 32 percent of the common implementation to achieve the same high precision. Specially, we have applied our Ising chip in image segmentation applications, a typical real-world application. The results show that, to find a segmentation with similar quality, our CMOS-type Ising chip can speed up the segmenting processing by 1900x with only 0.0170/00energy consumption compared with approximate algorithms operating on conventional computers.
Jian Zhang 0022, Shuming Chen
IEEE Trans. Computers1