Ang Li 0045

dblp:33/2805-45 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0004-8872-7649ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2025 A Hierarchical 3-D Physical Design Method for Ultralarge-Scale Logic-on-Memory CGRA Chip
abstract
Face-to-face bonded 3-D (F2F 3D) technology, with the potential to significantly reduce chip area while enhancing performance, stands as one of the most promising ways to extend Moore’s Law. However, current 3-D physical design flows are often modifications of 2-D design flows and rely on technical personnel to manually modify technical files. Furthermore, existing research on 3-D design flow primarily focuses on module implementation, with very few studies addressing hierarchical design methods for large-scale chips. In this article, we first introduce a 3-D physical design flow which concurrently optimizes the timing of both the logic tier and the memory tier, achieving synchronized physical design for both tiers. Then, we develop a bottom-up hierarchical 3-D physical design flow to extend the 3-D design flow to large-scale chip design. Through coordinated power planning, clock tree design, and interconnect unit design, we enhance the power, performance, and area (PPA) metrics of the entire chip. Using our RTL-to-GDS physical design flow, we successfully implemented a 28-nm CMOS logic-on-memory (LoM) 3-D coarse-grained reconfigurable architecture (CGRA) chip with over 50 million gates. Experimental results demonstrate that our 3-D flow improves timing by 16.1% while reducing voltage drop by 38.6% compared to the 2-D design. In addition, the power-delay product (PDP) of the 3-D chip decreases by 10.2%, showcasing better performance.
Zizheng Dong, Shuaipeng Li, Weijia Zhu, Ang Li 0045, Qin Wang 0009, Naifeng Jing, Weiguang Sheng, Jianfei Jiang 0001, Zhigang Mao
IEEE Trans. Very Large Scale Integr. Syst.4
2024 VDA: A Simple but Efficient Virtual-Channel-Based Deadlock Avoidance Scheme for Scalable Chiplet Networks
abstract
With the escalating computation capability demands of AI and other applications, chiplet technology has emerged as a prominent force in the current market, offering scalability and cost-effectiveness. One of the most critical issues in chiplet-based systems lies in the implementation of deadlock-free routing in 2.5D architectures. However, existing routing algorithms for 2.5D chiplet-based networks typically impose turn restrictions or necessitate complex hardware modifications, posing significant obstacles to scalability and exponentially increasing design costs. To address existing issues, we propose VDA, a simple deadlock avoidance scheme with fully utilized virtual channels (VCs) and lightweight hardware overhead for scalable chiplet-based networks. By constructing a dedicated virtual network through VC assignment, we enable the existence of cyclic channel dependencies and reduce VC restrictions. Meanwhile, a loop topology at the interposer level is introduced to enhance transmission efficiency. Our evaluation demonstrates that VDA yields an average improvement of up to 32.96% in saturation throughput and reduces low-load latency by up to 13.62% under synthetic traffic patterns. Furthermore, our approach achieves an average runtime speedup of 1.7% ∼ 6.2% when executing realistic workload benchmarks compared to existing approaches, with only 0.2% area overhead.
Duo Yu, Ang Li 0045, Naifeng Jing, Jianfei Jiang 0001, Weiguang Sheng, Qin Wang 0009
ACM Great Lakes Symposium on VLSI2
2024 A Comprehensive Dataflow-Mapping Optimization for Fully Pipelined Execution in Spatial Programmable Architecture
abstract
Although spatial programmable architectures have demonstrated high-performance and programmability for a variety of applications, they suffer from the pipeline unbalancing issue which restricts resource utilization and degrades the performance. In this paper, we identify that spatial initiation interval (SpII) can quantitatively describe the impact of pipeline unbalancing on performance, so we formulate SpII for the first time in spatial architectures. To achieve an optimal SpII, we propose dataflow decomposing and integrated mapping to enable high performance dataflow-mapping on spatial architectures. Dataflow decomposing decomposes the application graph into subgraphs and runs them serially, so that it adapts the regular spatial architecture to various application dataflows, particularly for extremely unbalanced datapaths without incurring large buffering overhead. Based on the quantitative SpII, we propose integrated mapping to consider operator placing, operand routing and pipeline balancing at the same time that can find a better SpII for fully-pipelined execution on spatial architectures. The experiment results show that our proposal can gain an average of 2.1× performance speedup on a variety of application kernels over the state-of-the-art approaches.
Pengyu Liu 0004, Ang Li 0045, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 ACET: An Adaptive Clock Scheme Exploiting Comprehensive Timing Slack for Reconfigurable Processors
abstract
To ensure the correctness and reliability, digital circuits are designed with conservative timing margins to accommodate extreme variations in process, voltage, and temperature (PVT) and workload. However, worst-case scenarios rarely occur, leaving the reserved time margins unutilized, which leads to a waste of performance. This issue is particularly significant in reconfigurable processors, as they exhibit substantial workload timing slack in both spatial and temporal domains. Previous researches have mainly focused on either developing PVT slack or exploiting workload slack, but few have simultaneously considered both aspects. Additionally, directly applying existing timing enhancement techniques to reconfigurable processors is challenging due to their complex configurability and diminishing timing slack in array architectures.To address the above challenges, this paper introduces ACET, an Adaptive Clock scheme which Exploits Timing slack comprehensively through hardware-software co-optimization. On the hardware side, ACET incorporates an adaptive clock module that adjusts the clock period based on both workload and PVT conditions. The two conditions are obtained by employing a PVT delay monitor and encoding the workload-dependent delay into the configuration, respectively. Then timing information is transmitted to phase selection module for cycle-level adjustments, to leverage the temporal timing slack. On the software side, to further exploit the spatial timing slack, a scheduling algorithm is proposed, which heuristically rearranges the firing time of operations. Experiments demonstrate that ACET leads to an average performance increase of 70.1% or an equivalent energy saving of 35.6%, with the hardware overhead being only 0.56%.
Shuya Ji, Weidong Yang 0007, Jianfei Jiang 0001, Naifeng Jing, Weiguang Sheng, Ang Li 0045, Qin Wang 0009
ICCD6