Cunlu Li

dblp:171/0907 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-3724-6878ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 7 first-author · 9 since 2021Computer networks · 6 · 5 since 2021
YearPublicationVenuePosition
2026 AGILE: Achieving Max-Min Fairness and High Utilization for In-Network Bandwidth Allocation
Yani Gong, Cunlu Li, Dezun Dong
INFOCOM2
2025 BCN: Enhanced Backpressure Flow Control with Rapid Notification in Datacenter Networks
Dinghuang Hu, Dezun Dong, Cunlu Li, Zejia Zhou, Guoyuan Yuan
Comput. Networks4
2024 AQC: Achieving Precise Bandwidth Allocation with Augmented Queues for Credit-Based Proactive Congestion Control
Yani Gong, Dinghuang Hu, Cunlu Li, Guoyuan Yuan, Dezun Dong
ICA3PP (3)4
2024 DRLAR: A deep reinforcement learning-based adaptive routing framework for network-on-chips
Ke Wu 0003, Cunlu Li, Dezun Dong
Comput. Networks5
2024 A survey of machine learning for Network-on-Chips
Dezun Dong, Cunlu Li, Liquan Xiao
J. Parallel Distributed Comput.3
2023 DeTAR: A Decision Tree-Based Adaptive Routing in Networks-on-Chip
Dezun Dong, Cunlu Li, Liquan Xiao
Euro-Par4
2023 Poster Abstract: A Network-on-Chip Router Architecture for Industrial Internet-of-Thing Gateways
abstract
More processors are integrated into Industrial Internet-of-Thing gateways to perform increasing emerging applications. Network-on-chip (NoC) offers a scalable, high-throughput, and energy-efficient communicate infrastructure. However, existing NoC routers cannot guarantee differentiated quality-of-service (QoS) for diversified applications. Hence, we propose a novel NoC router architecture with the gate control mechanism to provide customized QoS.
Chenglong Li 0007, Cunlu Li, Wenwen Fu, Tao Li 0008
IPSN2
2023 LARE: A Linear Approximate Reinforcement Learning Based Adaptive Routing for Network-on-Chips
abstract
The routing algorithm is crucial for network performance in network-on-chips (NoCs). With emerging applications bring new features to NoCs with more complex and time-varying traffic, which turns the routing computation process into multi-objective optimization. However, We found that the existing routing algorithms cannot effectively achieve load-balanced between different traffic due to the static method of routing design. The routing algorithm design space will increase if all factors affecting routing are considered. Reinforcement learning (RL) methods have demonstrated promising opportunities applied to architecture design exploration. In this paper, we proposed a novel RL framework for adaptive routing design in NoCs. This method uses network information to select the best path to achieve load balance and lower communication latency at the same time. Unfortunately, with this method, the implementation overhead of RL increases rapidly as the network scale increases. Therefore, we introduce a linear function of approximate RL-based adaptive routing (LARE) to reduce implementation over-head. We conduct extensive experiments against state-of-the-art routing algorithms to evaluate our design. Simulation results demonstrate the benefits of our design under synthetic traffic workloads and real applications. In addition, LARE can achieve similar network performance with traditional RL implementation with a much lower hardware overhead.
Dezun Dong, Cunlu Li, Zicong Wang, Zongmao Zhang
ISCAS4
2023 A Deterministic Embedded End-System Tightly Coupled With TSN Schedule
abstract
Distributed real-time systems (DRTSs) composed of many embedded end-systems have been widely adopted in the industrial fields. Time-sensitive networking (TSN), as a promising communication infrastructure for DRTS, has shown great potential in industry and academia. TSN assumes that an end-system can release critical tasks to process critical packets strictly according to the prescheduled time. Unfortunately, two factors currently damage this assumption: 1) the jitter caused by system architecture during task release, task execution, and packet transmission; and 2) a TSN schedule result may exceed the execution capability of the end-system and cause conflicts. This article proposed deterministic chip (DetChip), a system-on-chip capable of deterministically implementing the TSN schedule result. DetChip supports time-triggered task release, time-predictable task execution and precise network transmission. Based on DetChip, this article first formalizes the execution capability of the end-system as end-system constraints (ECs). Existing TSN scheduling algorithms integrating ECs can solve conflicts by co-scheduling end-systems and the TSN network. Compared with previous works, DetChip only introduces few clock cycles jitter for critical task execution according to the TSN schedule result. The proposed ECs obtain$2\times$–$8\times$more conflict-free solutions for advanced scheduling algorithms with a linear increase in time overhead. Compared with the general end-system, DetChip can reduce$6\times$–$20\times$processing jitter to achieve better clock synchronization.
Chenglong Li 0007, Zonghui Li, Tao Li 0008, Cunlu Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Revisiting network congestion avoidance through adaptive packet-chaining reservation
Ke Wu 0003, Dezun Dong, Cunlu Li, Weixia Xu 0001
Comput. Networks3
2022 MUA-Router: Maximizing the Utility-of-Allocation for On-chip Pipelining Routers
abstract
As an important pipeline stage in the router of Network-on-Chips, switch allocation assigns output ports to input ports and allows flits to transit through the switch without conflicts. Previous work designed efficient switch allocation strategies by maximizing the matching efficiency in time series. However, those works neglected the interaction between different router pipeline stages. In this article, we propose the concept of Utility-of-Allocation (UoA) to indicate the quality of allocation to be practically used in on-chip routers. We demonstrate that router pipelines can interact with each other, and the UoA can be maximized if the interaction between router pipelines is taken into consideration. Based on these observations, a novel class of routers, MUA-Router, is proposed to maximize the UoA through the collaborative design (co-design) between router pipelines. MUA-Router achieves this goal in two ways and accordingly implements two novel instance router architectures. In the first, MUA-Router improves the UoA by mitigating the impact of endpoint congestion in the switch allocation, and thus Eca-Router is proposed. Eca-Router achieves an endpoint-congestion-aware switch allocation through the co-design between routing computation and switch allocation. Based on Eca-Router, CoD-Router is proposed to feed back switch allocation information to routing computation stage to provide switch allocator with more conflict-free requests. Through the co-design between pipelines, MUA-Router significantly improves the efficiency of switch allocation and the performance of the entire network. Evaluation results show that our design can achieve significant performance improvement with moderate overheads.
Cunlu Li, Dezun Dong, Xiangke Liao
ACM Trans. Archit. Code Optim.1
2022 Hybrid Memory Buffer Microarchitecture for High-Radix Routers
abstract
Hierarchical high-radix router microarchitecture consisting of small SRAM-based intermediate buffers has been used in large-scale supercomputers interconnection networks. While hierarchical organization enables efficient scaling to higher switch port count, it requires intermediate buffers which can cause performance bottleneck. Shallow intermediate buffers can cause head-of-line blocking to create backpressure towards input buffers and reduce overall performance. Increasing intermediate buffer size overcomes this problem but becomes infeasible due to the large overhead. In this work, we propose to organise decentralized intermediate buffers as a centralized buffer and leverage alternate memory technology to increase its capacity. In particular, we exploit the high-density nature of Spin-Torque Transfer Magnetic RAM (STT-MRAM) to increase intermediate buffer depth while also providing near-zero leakage power. STT-MRAM has disadvantages such as higher write latency and higher write energy. To overcome these disadvantages, we propose DeepHiR, a novel deep hybrid buffer organization (STT-MRAM and SRAM) combined with a centralized buffer organization to provide high performance with minimal cost. Although the deep intermediate buffer provided by DeepHiR can effectively improve router performance, a large amount of input buffer will still cause a lot of hardware overhead. At the same time, deeper intermediate buffers also makes it take longer for the backpressure to propagate to the source node, thereby reducing the performance of DeepHiR. Therefore, we further propose ElasHiR, which leverages elastic input buffer design in the centralized row buffer to allow a part of the centralized row buffer to act as input buffer. ElasHiR adopts reduced input buffers and automatically determines the length of input buffer in the centralized row buffer. This design minimizes the buffer resource while achieving excellent efficiency. Evaluation results show that DeepHiR can achieve 56.7 percent performance improvement in packet latency under synthetic traffic, and the cost of energy and area is moderate. ElasHiR can reduce the input buffer by 93.8 percent with performance comparable to DeepHiR.
Cunlu Li, Dezun Dong, Xiangke Liao, John Kim 0001
IEEE Trans. Computers1
2021 Evaluation of Topology-Aware All-Reduce Algorithm for Dragonfly Networks
Dezun Dong, Cunlu Li, Ke Wu 0003, Liquan Xiao
NPC3
2021 CIB-HIER: Centralized Input Buffer Design in Hierarchical High-radix Routers
abstract
Hierarchical organization is widely used in high-radix routers to enable efficient scaling to higher switch port count. A general-purpose hierarchical router must be symmetrically designed with the same input buffer depth, resulting in a large amount of unused input buffers due to the different link lengths. Sharing input buffers between different input ports can improve buffer utilization, but the implementation overhead also increases with the number of shared ports. Previous work allowed input buffers to be shared among all router ports, which maximizes the buffer utilization but also introduces higher implementation complexity. Moreover, such design can impair performance when faced with long packets, due to the head-of-line blocking in intermediate buffers. In this work, we explain that sharing unused buffers between a subset of router ports is a more efficient design. Based on this observation, we propose Centralized Input Buffer Design in Hierarchical High-radix Routers (CIB-HIER), a novel centralized input buffer design for hierarchical high-radix routers. CIB-HIER integrates multiple input ports onto a single tile and organizes all unused input buffers in the tile as a centralized input buffer. CIB-HIER only allows the centralized input buffer to be shared between ports on the same tile, without introducing additional intermediate virtual channels or global scheduling circuits. Going beyond the basic design of CIB-HIER, the centralized input buffer can be used to relieve the head-of-line blocking caused by shallow intermediate buffers, by stashing long packets in the centralized input buffer. Experimental results show that CIB-HIER is highly effective and can significantly increase the throughput of high-radix routers.
Cunlu Li, Dezun Dong, Shazhou Yang, Xiangke Liao, Guangyu Sun 0003, Yongheng Liu
ACM Trans. Archit. Code Optim.1
2019 Network Congestion Avoidance through Packet-chaining Reservation
abstract
Endpoint congestion is a bottleneck in high-performance computing (HPC) networks and severely impacts system performance, especially for latency-sensitive applications. For long messages (or flows) whose duration is far larger than the round-trip time (RTT), endpoint congestion can be effectively mitigated by proactive or reactive counter-measures such that the injection rate of each source is dynamically controlled to a proper level. However, many HPC applications produce a hybrid traffic, a mix of short and long messages, and are dominated by short messages. Existing proactive congestion avoidance methods face the great challenge of scheduling the rapidly changing traffic pattern caused by these short messages. In this paper, we leverage the advantages of proactive and reactive congestion avoidance techniques and propose the Packet-chaining Reservation Protocol (PCRP) to make a dynamic balance between flows following proactive scheduling and packets subjected to reactive network conditions. We select the chaining packets as a flexible reservation granularity between the whole flow and one packet. We allow small flows to be speculatively transmitted without being discarded and give them higher priority over the entire network. Our PCRP can respond quickly to network conditions and effectively avoid the formation of endpoint congestion and reduce the average flow delay. We conduct extensive experiments to evaluate our PCRP and compare it with the state-of-the-art proactive reservation-based protocols, Speculative Reservation Protocol (SRP) and Bilateral Flow Reservation Protocol (BFRP). The simulation results demonstrate that in our design the flow latency can be reduced by 50.2% for hotspot traffic and 28.38% for uniform traffic.
Ke Wu 0003, Dezun Dong, Cunlu Li, Shan Huang 0002
ICPP3
2019 DeepHiR: improving high-radix router throughput with deep hybrid memory buffer microarchitecture
abstract
Hierarchical high-radix router microarchitecture consisting of small SRAM-based intermediate buffers have been used in large-scale supercomputers interconnection networks. While hierarchical organization enables efficient scaling to higher switch port count, it requires intermediate buffers that can cause performance bottleneck. Shallow intermediate buffers can cause head-of-line blocking and result in backpressure towards the input buffers to reduce overall performance. Increasing intermediate buffer size overcomes this problem but is infeasible since the amount of intermediate buffer is proportional to O(p2) where p is the router radix. Adopting new memory technology with higher density can increase intermediate buffer size but is not practical in decentralized, small-size intermediate buffers.
Cunlu Li, Dezun Dong, Xiangke Liao, John Kim 0001
ICS1
2018 Eca-Router : On Achieving Endpoint Congestion Aware Switch Allocation in the On-Chip Network
abstract
As the critical pipeline stage in on-chip routers, switch allocation assigns output ports to input ports and allow flits transiting through the switch without conflicts. Previous works strive to design efficient switch allocaiton strategies by maximizing the matching at each cycle, with the information from the current cycle or multiple cycles in time series. However, those works have not taken endpoint congestion into considerations. Tree-saturation, caused by endpoint congestion, can degrade NoC performance due to the congestion fanning out from the original point to upstream routers. In this paper, a novel router design, Eca-Router, is proposed to relieve the impact of endpoint congestion by switch allocation optimization. Eca-Router detects endpoint congestion by recording the destinations of packets in switch allocation. Endpoint congestion is decided in switch allocation once there are multiple input ports competing for the same output port and the packets in these input ports contain the same destination. During switch allocation, requests that contribute to endpoint congestion will be given lower priority to be allocated, and starvation control is also introduced to ensure allocation fairness. Evaluation results show that Eca-Router is efficient in reducing packet latency.
Cunlu Li, Dezun Dong, Xiangke Liao
ICCD1
2018 BFRP: Endpoint Congestion Avoidance Through Bilateral Flow Reservation
abstract
In HPC, endpoint congestion is a bottleneck in the network and seriously affects the performance of the system. The endpoint congestion can be effectively mitigated by quickly responding to the network and reducing the injection rate of the source. However, most of the prior works do not consider the impact of flow completion time on system performance, but only focus on the packet latency and perform scheduling at packet granularity. For HPC applications, the flow completion time and throughput are the metrics that determines the speed of application execution. Although prior works reduce package latency, they do not fundamentally reduce the flow latency in flow level. In this paper, we propose the bilateral flow-based reservation protocol (BFRP). BFRP quickly responds to network conditions through light-weight bilateral reservation mechanism and can effectively avoid the formation of endpoint congestion. BFRP also schedules packets based on flows, and smallest flow is preferentially sent to decrease the average flow latency. BFRP ensures the source and destination send or receive flows according to the allocated time slices without any conflict at both ends. We evaluate our BFRP protocol against state-of-the-art reservation-based protocol, speculative reservation protocol(SRP), and the simulation results show that the flow latency can be reduced by 27.68% under hotspot traffic with fixed flow size and 24.59% under uniform traffic with fixed flow size.
Tianye Yang, Dezun Dong, Cunlu Li, Liquan Xiao
IPCCC3
2018 RoB-Router : A Reorder Buffer Enabled Low Latency Network-on-Chip Router
abstract
Traditional input-queued routers in network-on-chips (NoCs) only have a small number of virtual channels (VCs) and packets in a VC are organized in a fixed order. Such design is susceptible to head-of-line (HoL) blocking as only the packet at the head of a VC can be allocated by the switch allocator. Since switch allocation is the critical pipeline stage in on-chip routers, HoL blocking significantly degrades the performance of NoCs. In this paper, we propose to schedule packets in input buffers utilizing reorder buffer (RoB) techniques. We design VCs as RoBs to allow packets located not at the head of a VC to be allocated before the head packets. RoBs reduce the conflicts in switch allocation and mitigate the HoL blocking and thus improve the NoC performance. However, it is hard to reorder all the units in a VC due to circuit complexity and power overhead. We propose RoB-Router, which leverages elastic RoBs in VCs to only allow a part of a VC to act as RoB. RoB-Router automatically determines the length of RoB in a VC based on the number of buffered flits. This design minimizes the resource while achieving excellent efficiency. Furthermore, we propose two independent methods to improve the performance of RoB-Router. One is to optimize the packet order in input buffers by redesigning VC allocation strategy. The other combines RoB-Router with current most efficient switch allocator TS-Router. We perform evaluations and the results show that our design can achieve 46 and 15.7 percent performance improvement in packet latency under synthetic traffic and traces from PARSEC than TS-Router, and the cost of energy and area is moderate. Additionally, average packet latency reduction by our two improving methods under uniform traffic is 13 and 17 percent respectively.
Cunlu Li, Dezun Dong, Zhonghai Lu, Xiangke Liao
IEEE Trans. Parallel Distributed Syst.1
2016 MBL: A Multi-stage Bufferless High-radix Router
abstract
There is a pressing need for high-radix routers in modern HPC (High Performance Computing) interconnects and to build the exascale computers with massive clusters. In this paper, we propose MBL, a high-radix router with a multi-stage bufferless switch Clos network inside. Booksim interconnection network simulator is used to implement our arbitrating designs for the architecture and it runs well under different traffic patterns in a flattened butterfly network, with 136 ports for each router.
Wenxiang Yang, Dezun Dong, Jingyue Zhao, Cunlu Li
CLUSTER4
2016 CCAS: Contention and congestion aware switch allocation for network-on-chips
abstract
Network-on-chip system plays an important role to improve the performance of chip multiprocessor systems. As the complexity of the network increases, congestion problem has become the major performance bottleneck and seriously influence the performance of NoCs. Prior works have focused on designing effective routing algorithm based on collecting contention and congestion information to load balance the traffic. However, most prior works do not consider balancing the traffic load during switch allocation. Due to the lack of congestion information in switch allocation stage, switch allocator performs allocation only based on packet requests and thus aggravates the congestion in the ports of switch. In this paper, we propose CCAS, a new switch allocation strategy to add the contention and congestion information into the switching process to load balance the traffic and achieve efficient switch allocation. We carefully design CCAS to balance the trade-off between traffic load balance and the matching efficiency in switch allocation. We evaluate our design under synthetic traffic and traces of PARSEC benchmarks. Our evaluations show that CCAS can achieve remarkable latency reduction compared to other switch allocation strategies.
Cunlu Li, Dezun Dong, Xiangke Liao, Ji Wu 0006
ICCD1
2016 Galaxyfly: A Novel Family of Flexible-Radix Low-Diameter Topologies for Large-Scales Interconnection Networks
abstract
Interconnection network plays an essential role in the architecture of large-scale high performance computing (HPC) systems. In the paper, we construct a novel family of low-diameter topologies, Galaxyfly, using techniques of algebraic graphs over finite fields. Galaxyfly is guaranteed to retain a small constant diameter while achieving a flexible tradeoff between network scale and bisection bandwidth. Galaxyfly lowers the demands for high radix of network routers and is able to utilize routers with merely moderate radix to build exascale interconnection networks. We present effective congestion-aware routing algorithms for Galaxyfly by exploring its algebraic property. We conduct extensive simulations and analysis to evaluate the performance, cost and power consumption of Galaxyfly against state-of-the-art topologies. The results show that our design achieves better performance than most existing topologies under various routing algorithms and traffic patterns, and is cost-effective to deploy for exascale HPC systems.
Dezun Dong, Xiangke Liao, Xing Su 0004, Cunlu Li
ICS5
2015 HVCRouter: Energy Efficient Network-on-Chip Router with Heterogeneous Virtual Channels
Ji Wu 0006, Xiangke Liao, Dezun Dong, Wang Li 0003, Cunlu Li
ICA3PP (1)5