VLDB 2026 Research / reviewers in the wild / expert
Dong Jiang 0002
dblp:64/2623-2
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-5833-241XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multi-Precision Tensor Processing Unit for Accelerating Matrix Computations
Zongfan Wu, Dong Jiang 0002, Enyi Yao |
ISCAS | 3 |
| 2025 | An SRAM Compute-in-Memory based NTT Accelerator for CRYSTALS-KYBERabstractAmong various post-quantum cryptography (PQC) proposed by researchers, lattice-based cryptography is considered to be one of the most promising post-quantum public key cryptography systems. It has significant advantages over other PQC schemes in terms of security, computational efficiency, and versatility in designing public key encryption, digital signatures, and cryptographic negotiation protocols. Efficient implementations of the number theoretic transform (NTT) operations are crucial for many lattice-based encryption algorithms. This article presents an NTT hardware acceleration structure based on SRAM in-memory computing technology, which optimizes the key butterfly structure in the NTT algorithm by improving the 6T-SRAM array structure and incorporating near-memory computing structures. Compared with existing technologies, this accelerator can significantly reduce area and power consumption. Jinyang Hu, Xinyuan Pang, Dong Jiang 0002, Gaopeng Fan, Enyi Yao |
ISCAS | 3 |
| 2025 | A Spin Scale-Aware Self-Adaptive Ising Annealing Processing Architecture for Combinatorial Optimization ProblemsabstractThe Ising annealing processor has emerged as a promising approach to accelerate the discovery of the optimal solutions for a wide range of combinatorial optimization problems (COPs), by mapping various COPs into a unified Ising model. However, fixed computational strategies and inflexible architectures make previous designs suffer from a low hardware resource utilization rate when the numbers of the total required and real-time flipped spins vary across different COPs and iteration steps. In this paper, a novel spin scale-aware self-adaptive Ising annealing processing architecture (AIAPA) is proposed to address this problem, with an adaptive computational strategy, a custom instruction set, multi-traffic mode routers, and a fully-pipelined computing array. It can dynamically adapt to the varying scenarios during the Ising annealing process to maximize the performance of limited hardware resources. Its prototype, supporting 65k fully-connected spins, is implemented on an FPGA platform, operates at a clock frequency of 188 MHz. The AIAPA achieves up to a 24.22 times faster annealing speed compared to the state-of-the-art FPGA design on the max-cut optimization problem while maintaining a high convergence accuracy. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Longyuan Kang, Simei Yang, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | A Parallel Tempering Processing Architecture with Multi-Spin Update for Fully-Connected Ising ModelsabstractCombinatorial optimization problems (COPs) are notoriously difficult to solve for classic Von-Neumann computers, which are ubiquitous in various domains. As a state-of-the-art hardware acceleration scheme for COPs, Ising machines are one of the promising research directions for the next generation of computing, but still suffer from the low solution accuracy and speed due to the high complexity of the fully-connected Ising model. In this work, a novel parallel tempering processing architecture (PTPA) is proposed with the modified parallel tempering algorithm, aimed at reducing search time and improving the solution quality. Several techniques are developed to further reduce hardware overhead and enhance parallelism, including the independent pipelined spin update architecture, approximated probability equations, and compact random number generators. Its prototype is implemented on FPGA with eight replicas, each replica containing 1,024 fully-connected spins and at most 64 concurrent update spins. The proposed design achieves an average cut accuracy of 99.43% within 1ms solution time on various G-set problems. Compared with the CPU-based parallel tempering implementation, it enhances the speed of solving the max-cut problems by 5,160 times. Yang Zhang 0120, Xiangrui Wang, Dong Jiang 0002, Zhanhong Huang, Gaopeng Fan, Enyi Yao |
DATE | 3 |
| 2024 | Unidirectional and hierarchical on-chip interconnected architecture for large-scale hardware spiking neural networks
Junxiu Liu, Dong Jiang 0002, Qiang Fu 0019, Yuling Luo, Yaohua Deng, Sheng Qin, Shunsheng Zhang |
Neurocomputing | 2 |
| 2024 | DCAP: A Scalable Decoupled-Clustering Annealing Processor for Large-Scale Traveling Salesman ProblemsabstractThe Traveling Salesman Problem (TSP) is one of the most well-known NP-hard combinatorial optimization problems (COPs). Many social production problems can be effectively represented as instances of TSPs. However, solving large-scale TSPs remains a significant challenge for conventional Von Neumann computers. Many studies have proposed annealing processors to address large-scale COPs, but most of them focus on unconstrained problems, such as the Maxcut problem. In this paper, a scalable decoupled-clustering annealng processor (DCAP) for efficiently handling large-scale TSPs is presented. A decoupled hierarchical clustering algorithm is proposed for higher convergence speed and improved scalability. Several techniques have been developed in hardware to minimize area overhead and processing time, including a modified spin connection topology for the Ising model, an area-efficient random threshold generator, a one-step spin update scheme and a dynamic prediction method. The DCAP prototype is implemented on FPGA with an operating frequency of 125MHz. We tested our design on various TSP instances from the TSPLIB. Results show that our design outperforms the CPU- and GPU-based Neuro-Ising scheme by achieving maximum speedups of$780\times $and a 42% improvement in accuracy. With multi-chip interconnection, DCAP is able to handle problems of scale up to 85900 cities. Zhanhong Huang, Yang Zhang 0120, Xiangrui Wang, Dong Jiang 0002, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | An Annealing Processor based on 1k-Spin Fully-Connected Ising Model for Combinatorial Optimization ProblemsabstractCombinatorial optimization problems (COPs) find extensive applications in industrial and social scenarios such as transportation and communication. As the size of NP-hard COPs increases, it becomes impossible to obtain the optimal solution using an enumerative method. Recently, Ising model based annealing processors have received increasing attention due to their potential for rapidly converging to the near-optimal solutions after mapping the problem to them. This paper presents a novel annealing processor (AP) with 1024 fully-connected spins based on a modified Ising model annealing algorithm, which is more suitable for hardware implementation compared to conventional simulated annealing (SA) algorithm. The prototype is implemented using FPGA with the operation frequency up to 100MHz. We tested our design on various G-set problems with an average cut accuracy of 99.19% achieved. The proposed design outperforms the conventional CPU-based method by achieving a max speedup of 2204x for G51. Zhanhong Huang, Xiangrui Wang, Dong Jiang 0002, Yukang Huang, Enyi Yao |
ISCAS | 3 |
| 2023 | A Scalable Annealing Processing Architecture for Fully-Connected Ising ModelsabstractCombinational Optimization Problems (COPs) are prevalent in many different fields. Most of these problems are NP-hard and challenging for computers with conventional Von-Neumann architecture. Ising machines with numerous spins have the potential to solve these problems by emulating the natural annealing process of solid matter. Recent research has explored the hardware implementation of Ising machines to accelerate the convergence process of such problems at room temperature. However, most of them are suffering from low scalability and low parallel processing capability due to the huge hardware cost and high complexity. In this paper, a scalable annealing processing architecture for Ising processor is described to address these issues with a NoC computing paradigm, a distributed storage scheme, and a fully pipelined structure design. The prototype is synthesized using FPGA with the maximum operation frequency of 270MHz, achieving about 32 times faster than conventional simulated annealing method when solving the max-cut problem. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Yukang Huang, Enyi Yao |
ISCAS | 1 |
| 2023 | Hardware Spiking Neural Networks with Pair-Based STDP Using Stochastic Computing
Junxiu Liu, Yanhu Wang, Yuling Luo, Shunsheng Zhang, Dong Jiang 0002, Yifan Hua, Sheng Qin, Su Yang 0002 |
Neural Process. Lett. | 5 |
| 2023 | A Network-on-Chip-Based Annealing Processing Architecture for Large-Scale Fully Connected Ising ModelabstractCombinatorial optimization problems are prevalent in many different fields. Most of these problems are NP-hard and challenging for computers with conventional Von-Neumann architecture. Ising machines with a number of spins have the potential to solve these problems by emulating the natural annealing process of solid matter. Recent research has explored some hardware implementation methods of Ising machines to accelerate the convergence process of such problems at room temperature. However, most of them are suffering from low scalability and low parallel processing capability due to the huge hardware cost and high complexity. In this paper, a novel network-on-chip-based annealing processing architecture (NoCAPA) for a large-scale Ising processor is described to address these issues with a NoC computing paradigm, a distributed storage scheme, and a fully pipelined structure design. Several techniques are developed to further increase convergence speed and reduce hardware resource consumption, including a dynamic multithread parallel update algorithm, a router with merge and deflection abilities, and a unique multiply-accumulate operation. The prototype is implemented in FPGA with the maximum operation frequency of 200MHz, achieving up to$120.5\times $faster than conventional simulated annealing method when solving the max-cut problem while supporting high scalability. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Yongkui Yang, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Minimally buffered deflection router for spiking neural network hardware implementations
Junxiu Liu, Dong Jiang 0002, Yuling Luo, Senhui Qiu, Yongchuang Huang |
Neural Comput. Appl. | 2 |