EDBT 2026 Demo / reviewers in the wild / expert
Yang Zhang 0120
dblp:06/6785-120
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0002-3145-4625ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Parallel Tempering Processing Architecture with Multi-Spin Update for Fully-Connected Ising ModelsabstractCombinatorial optimization problems (COPs) are notoriously difficult to solve for classic Von-Neumann computers, which are ubiquitous in various domains. As a state-of-the-art hardware acceleration scheme for COPs, Ising machines are one of the promising research directions for the next generation of computing, but still suffer from the low solution accuracy and speed due to the high complexity of the fully-connected Ising model. In this work, a novel parallel tempering processing architecture (PTPA) is proposed with the modified parallel tempering algorithm, aimed at reducing search time and improving the solution quality. Several techniques are developed to further reduce hardware overhead and enhance parallelism, including the independent pipelined spin update architecture, approximated probability equations, and compact random number generators. Its prototype is implemented on FPGA with eight replicas, each replica containing 1,024 fully-connected spins and at most 64 concurrent update spins. The proposed design achieves an average cut accuracy of 99.43% within 1ms solution time on various G-set problems. Compared with the CPU-based parallel tempering implementation, it enhances the speed of solving the max-cut problems by 5,160 times. Yang Zhang 0120, Xiangrui Wang, Dong Jiang 0002, Zhanhong Huang, Gaopeng Fan, Enyi Yao |
DATE | 1 |
| 2024 | An Ising Model-Based Parallel Tempering Processing Architecture for Combinatorial OptimizationabstractCombinatorial optimization problems (COPs) are prevalent in various domains and present formidable challenges for modern computers. Searching for the ground state of the Ising model emerges as a promising approach to solve these problems. Recent studies have proposed some annealing processing architectures based on the Ising model, aimed at accelerating the solution of COPs. However, most of them suffer from low solution accuracy and inefficient parallel processing. This article presents a novel parallel tempering processing architecture (PTPA) based on the fully-connected Ising model to address these issues. The proposed modified parallel tempering algorithm supports multi-spin concurrent updates per replica and employs an efficient multi-replica swap scheme, with fast speed and high accuracy. Furthermore, an independent pipelined spin update architecture is designed for each replica, which supports replica scalability while enabling efficient parallel processing. The PTPA prototype is implemented on FPGA with 8 replicas, each with 1,024 fully-connected spins. It supports up to 64 spins for concurrent updates per replica and operates at 200 MHz. Different concurrency strategies are considered to further improve the efficiency of solving COPs. In the test of various G-set problems, PTPA achieves 3.2× faster solution speed along with 0.27% better average cut accuracy compared to a state-of-the-art FPGA-based Ising machine. Yang Zhang 0120, Xiangrui Wang, Gaopeng Fan, Yuan Cao 0003, Yiqiu Liu, Yongkui Yang, Enyi Yao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | DCAP: A Scalable Decoupled-Clustering Annealing Processor for Large-Scale Traveling Salesman ProblemsabstractThe Traveling Salesman Problem (TSP) is one of the most well-known NP-hard combinatorial optimization problems (COPs). Many social production problems can be effectively represented as instances of TSPs. However, solving large-scale TSPs remains a significant challenge for conventional Von Neumann computers. Many studies have proposed annealing processors to address large-scale COPs, but most of them focus on unconstrained problems, such as the Maxcut problem. In this paper, a scalable decoupled-clustering annealng processor (DCAP) for efficiently handling large-scale TSPs is presented. A decoupled hierarchical clustering algorithm is proposed for higher convergence speed and improved scalability. Several techniques have been developed in hardware to minimize area overhead and processing time, including a modified spin connection topology for the Ising model, an area-efficient random threshold generator, a one-step spin update scheme and a dynamic prediction method. The DCAP prototype is implemented on FPGA with an operating frequency of 125MHz. We tested our design on various TSP instances from the TSPLIB. Results show that our design outperforms the CPU- and GPU-based Neuro-Ising scheme by achieving maximum speedups of$780\times $and a 42% improvement in accuracy. With multi-chip interconnection, DCAP is able to handle problems of scale up to 85900 cities. Zhanhong Huang, Yang Zhang 0120, Xiangrui Wang, Dong Jiang 0002, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |