Jinlun Ji

dblp:333/4221 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0009-5260-5726ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2026 An Adaptive Congestion-aware Approximate Communication (ACAC) scheme and implementation for Network-on-Chip systems
Shize Zhou, Wenjie Fan 0004, Yongqi Xue, Shiping Li, Songfeng Deng, Jinlun Ji, Tong Cheng, Xinyu Wang 0027, Li Li 0003
Integr.7
2025 FPGA-Par: An Efficient Algorithm for Elegant Partitioning in Multi-FPGA Systems
abstract
This work introduces FPGA-Par, an efficient graph partitioning algorithm for multi-FPGA systems. FPGA-Par utilizes an iterative balanced partitioning and supernodes transferring algorithm that transforms imbalanced partitioning problems into balanced ones, addressing the inefficiencies associated with traditional single-node movement methods, as well as eliminating the need for brute-force exploration of imbalance factors. This approach achieves partitions with flexible size ratios that satisfy resource constraints, resulting in improved partition quality. Experimental results demonstrate that FPGA-Par reduces time overhead by 55% compared to state-of-the-art imbalanced partitioning algorithms. Furthermore, it improves partition quality metrics, with Cut and Cutmaxdecreasing by 38% and 25%, respectively, and achieves a 1.25x increase in system frequency after mapping and routing on a multi-FPGA system.
Hengyue Gao, Chenyang Dai, Jinlun Ji, Qiyue Zhao, Jiangtao Yuan, Li Li 0003
ISCAS3
2025 NAME: NoC-based Accelerators Mapping Exploration for High Performance DNN Inference
abstract
The rapid advancement of deep learning, with increasingly large deep neural networks (DNNs), has led to the use of multi-core parallel processing in accelerators, utilizing Network-on-Chip (NoC) for interconnection. However, while multi-core architectures improve computational performance, they also increase data movement overhead. This paper analyzes NoC traffic patterns during DNN processing and explores mapping optimization to reduce computational and communicational overheads. We propose NoC-based Accelerators Mapping Exploration (NAME), an automated mapper for generating high-performance task mappings for NoC-based accelerators. NAME balances computational and communication overheads by dividing DNNs into scalable groups and interleaving NoC link usage, reducing traffic bottlenecks. Experiments show that NAME reduces execution time by up to 82% and 52%, respectively, compared to fixed mapping and the state-of-the-art framework AOME (Autonomous Optimal Mapping Exploration).
Jinlun Ji, Hengyue Gao, Heng Zhang 0025, Yuqi Lu, Yulong Song, Wenjie Fan 0004, Li Li 0003
ISCAS1
2024 TTNNM: Thermal- and Traffic-Aware Neural Network Mapping on 3D-NoC-based Accelerator
abstract
3D Network on Chips (3D-NoCs) have ample on-chip wiring resources and high bandwidth, yet face numerous hotspots and higher temperature gradients due to increased integration and power density. This could lead to device failure, impacting system stability. Our paper introduces a thermal- and traffic-aware mapping method for 3D-NoC-based neural network accelerators. Firstly, based on the average load of different neural network layer, we determine their mapping sequences and suitable dies. Secondly, to minimize delay and alleviate hotspot temperatures, we allocate groups to appropriate nodes. Compared with previous works, TTNNM reduces the average temperature by 3.0°C, 2.2°C, 2.4°C, temperature variance by 58.4%, 64.8%, 73.0%, maximum temperature by 9.3°C, 7.9°C, 12.0°C, and packet latency by 31.7%, 17.2%, 25.1%.
Wenjie Fan 0004, Heng Zhang 0025, Jinlun Ji, Tong Cheng, Shiping Li, Li Li 0003
ACM Great Lakes Symposium on VLSI4
2024 Automatic Generation and Optimization Framework of NoC-Based Neural Network Accelerator Through Reinforcement Learning
abstract
Choices of dataflows, which are known as intra-core neural network (NN) computation loop nest scheduling and inter-core hardware mapping strategies, play a critical role in the performance and energy efficiency of NoC-based neural network accelerators. Confronted with an enormous dataflow exploration space, this paper proposes an automatic framework for generating and optimizing the full-layer-mappings based on two reinforcement learning algorithms including A2C and PPO. Combining soft and hard constraints, this work transforms the mapping configuration into a sequential decision problem and aims to explore the performance and energy efficient hardware mapping for NoC systems. We evaluate the performance of the proposed framework on 10 experimental neural networks. The results show that compared with the direct-X mapping, the direct-Y mapping, GA-base mapping, and NN-aware mapping, our optimization framework reduces the average execution time of 10 experimental NNs by 9.09$\%$, improves the throughput by 11.27$\%$, reduces the energy by 12.62$\%$, and reduces the time-energy-product (TEP) by 14.49$\%$. The results also show that the performance enhancement is related to the coefficient of variation of the neural network to be computed.
Yongqi Xue, Jinlun Ji, Xinming Yu, Shize Zhou, Tong Cheng, Shiping Li, Kai Chen 0034, Zhonghai Lu, Li Li 0003
IEEE Trans. Computers2
2024 HAS-RL: A Hierarchical Approximate Scheme Optimized With Reinforcement Learning for NoC-Based NN Accelerators
abstract
Network-on-Chip (NoC) is a scalable on-chip communication architecture for the NN accelerator, but with the increase in the number of nodes, the communication delay becomes higher. Applications such as machine learning have a certain resilience to noisy/erroneous transmitted data. Therefore, approximate communication becomes a promising solution to improving performance by reducing traffic loads under the constraint of the acceptable maximum accuracy loss of neural networks. It is a key issue to balance the result quality and the communication delay for approximate NoC systems. The traditional approximate NoC only considers the node-to-node approximation-based dynamic traffic regulation. However, the dynamically changing traffic patterns across different nodes, different times, and different applications lead to a huge search space, which makes it hard to explore an optimal global approximation solution. In this paper, we propose a quality model for different neural networks, which presents the relationship between the quality loss and the data approximate rate. Then, a hierarchical approximate scheme optimized with reinforcement learning (HAS-RL) is proposed and we reduce the complexity of the HAS-RL by reducing the state space and action space, which will reduce the resource overhead as well. After that, we embed a global approximate controller in the NoC system, in which we deploy a policy network trained with the offline reinforcement learning algorithm to adjust the data approximate rates of each node at run time. Compared with the state-of-the-art method, the proposed scheme reduces the average network delay by 13.5% while their accuracies are similar. The proposed HAS-RL only causes an additional area overhead of 1.24% and power consumption of 0.77% compared with the traditional router design.
Shize Zhou, Yongqi Xue, Wenjie Fan 0004, Tong Cheng, Jinlun Ji, Chenyang Dai, Wenqing Song, Qinyu Chen, Chang Gao 0002, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.6
2022 Work in Progress: ACAC: An Adaptive Congestion-aware Approximate Communication Mechanism for Network-on-Chip Systems
abstract
Data-intensive applications, such as machine learning and pattern recognition, result in heavy Network-on-Chip (NoC) communication loads and a tremendous increase in the communication latency. At the same time, the error-tolerant nature of these applications makes approximate communication an effective way to relieve the sharp increase of the network latency. This paper proposes an adaptive congestion-aware approximate communication mechanism (ACAC) that can alleviate the communication congestion of NoC systems in heavy communication loads. Our cycle-accurate simulations have shown that the proposed ACAC effectively reduces the network latency similar to ABDTR under a 22% to 52% lower data approximate ratio and significantly decreases the additional compression control traffic volume under real applications.
Shize Zhou, Yongqi Xue, Jinlun Ji, Tong Cheng, Li Li 0003
CASES4
2022 AOME: Autonomous Optimal Mapping Exploration Using Reinforcement Learning for NoC-based Accelerators Running Neural Networks
abstract
Hardware mapping plays a critical role in the performance of NoC-based accelerators running large-scale neural networks (NN). Confronted with enormous mapping exploration space, traditional algorithms may find sub-optimal solutions. We conduct preliminary experiments to investigate the impact of different hardware mappings on communication latencies. Then, this paper proposes an Autonomous Optimal Mapping Exploration (AOME) architecture based on two reinforcement learning algorithms. Combining soft and hard constraints, AOME transforms the mapping process into a sequential decision problem and targets to explore the optimal mapping of the NoC system. We evaluate the performance of AOME on ten NNs. The results show that compared with the direct X mapping, the direct Y mapping, GA-base mapping, and NN-aware mapping, AOME reduces the average communication latency of ten NNs by 27.30%, 33.33%, 4.27% and 12.46% using A2C, by 27.19%, 33.21%, 4.11% and 12.31% using PPO, and improves the average communication throughput by 43.24%, 63.60%, 5.17% and 14.83% using A2C, by 43.18%, 63.68%, 5.23% and 14.87% using PPO.
Yongqi Xue, Jinlun Ji, Shize Zhou, Tong Cheng, Li Li 0003
ICCD2