Yinxiao Feng

dblp:317/0794 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-3637-1132ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Ring Road: A Scalable Polar-Coordinate-based 2D Network-on-Chip Architecture
abstract
Networks-on-chip (NoCs) are scaled out to build large-scale multi-chip networks to meet the growing demand for computing. However, the traditional router-based NoC architecture has significant limitations: 1) The overhead of routers is high, especially for long-distance traffic that traverses numerous chips and hops; 2) On/off-chip traffic is mixed up, so the local NoC performance is dragged down by the cross-chip traffic; 3) The routing design of intra/inter-chip networks is also entangled. The alternative router-less solution based on isolated multi-ring (IMR) can address some of these issues; however, the complete lack of routers leads to wiring overhead and limited scalability thus not suitable for large-scale multi-chip networks. Therefore, we are motivated to design a new network architecture that can reduce router/wiring overhead, isolate on/off-chip traffic, and decouple inter/intra-chip routing design by combining the advantages of both routers and router-less rings. In this paper, we propose Ring Road, in which high-speed isolated multi-rings are integrated with compact routers, thus achieving flexible traffic delivery with less router usage and overhead. A polar-coordinate-based description is presented to describe the topology, and corresponding deadlock-free routing algorithms are discussed. As a standalone NoC, Ring Road is low-cost, low-latency, and high-performance. It can also be scaled out into multi-chip networks without redesigning the on-chip routing regardless of inter-chip topology. With low-cost isolated ring channels, no matter how heavy the cross-chip traffic is, the local NoC performance remains consistent. Circuit implementation and post-synthesis analysis show that Ring Road achieves$1.7\sim 2\times$bandwidth-per-area and$1.6\sim 1.9\times$bandwidth-per-power compared with the traditional mesh router. Cycle-accurate simulation shows that Ring Road achieves better performance and energy efficiency at various workloads and configurations.
Yinxiao Feng, Kaisheng Ma
MICRO1
2024 Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
abstract
Existing high-performance computing (HPC) interconnection architectures are based on high-radix switches, which limits the injection/local performance and introduces latency/energy/cost overhead. The new wafer-scale packaging and high-speed wireline technologies provide high-density, low-latency, and high-bandwidth connectivity, thus promising to support direct-connected high-radix interconnection architecture.In this paper, we propose a wafer-based interconnection architecture called Switch-Less-Dragonfly-on-Wafers. By utilizing distributed high-bandwidth networks-on-chip-on-wafer, costly high-radix switches of the Dragonfly topology are eliminated while increasing the injection/local throughput and maintaining the global throughput. Based on the proposed architecture, we also introduce baseline and improved deadlock-free minimal/non-minimal routing algorithms with only one additional virtual channel. Extensive evaluations show that the Switch-Less-Dragonfly-on-Wafers outperforms the traditional switch-based Dragonfly in both cost and performance. Similar approaches can be applied to other switch-based direct topologies, thus promising to power future large-scale supercomputers.
Yinxiao Feng, Kaisheng Ma
SC1
2024 Evaluating Chiplet-based Large-Scale Interconnection Networks via Cycle-Accurate Packet-Parallel Simulation
Yinxiao Feng, Kaisheng Ma
USENIX ATC1
2023 A Scalable Methodology for Designing Efficient Interconnection Network of Chiplets
abstract
The Chiplet methodology can accelerate VLSI system development and provide better flexibility. However, it is not easy to build interconnection networks across multiple chiplets and maintain high-performance deadlock-free routing in systems of various hierarchical topologies. In particular, most on-chiplet networks are based on flat topologies such as 2D-mesh, which are inflexible and insufficient for large-scale multi-chiplet systems.To take full advantage of the multi-chiplet architecture and advanced packaging, we propose an interconnection method that can flexibly establish high-radix interconnection networks from typical 2D-mesh-NoC-based chiplets. A minus-first-based deadlock-free adaptive routing algorithm and a safe/unsafe flow control policy are introduced for these multi-chiplet interconnection networks. Additionally, a general approach network interleaving is used to balance the communication bandwidth within and between chiplets.We evaluate different architectures and traffic patterns on a cycle-accurate C++ simulator. Compared with traditional adaptive routing in 2D-mesh, our methodology can significantly improve network performance in various cases. The more chiplets there are, the more effective the method is. For 64 4×4-2D-mesh-based chiplets, The maximum injection rate increase is up to 2×, and the average latency reduction is up to 45%.
Yinxiao Feng, Kaisheng Ma
HPCA1
2023 Heterogeneous Die-to-Die Interfaces: Enabling More Flexible Chiplet Interconnection Systems
abstract
The chiplet architecture is one of the emerging methodologies and is believed to be scalable and economical. However, most current multi-chiplet systems are based on one uniform die-to-die interface, which severely limits flexibility. First, any interface has specific applicable workloads/scales/scenarios; therefore, chiplets with a uniform interface cannot be freely reused in different systems. Second, since modern computing systems must deal with complex and mixed tasks, the uniform interface does not cope well with flexible workloads, especially for large-scale systems.
Yinxiao Feng, Kaisheng Ma
MICRO1
2022 Chiplet actuary: a quantitative cost model and multi-chiplet architecture exploration
abstract
Multi-chip integration is widely recognized as the extension of Moore's Law. Cost-saving is a frequently mentioned advantage, but previous works rarely present quantitative demonstrations on the cost superiority of multi-chip integration over monolithic SoC. In this paper, we build a quantitative cost model and put forward an analytical method for multi-chip systems based on three typical multi-chip integration technologies to analyze the cost benefits from yield improvement, chiplet and package reuse, and heterogeneity. We re-examine the actual cost of multi-chip systems from various perspectives and show how to reduce the total cost of the VLSI system through appropriate multi-chiplet architecture.
Yinxiao Feng, Kaisheng Ma
DAC1