Shaolin Xiang

dblp:224/1700 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Interconnection networks and networks-on-chip · 87% Hardware accelerators and domain-specific architectures · 13%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip
die-to-die interconnect
0.612022
Application Defined On-chip Networks for Heterogeneous Chiplets: An Implementation Perspective · HPCA 2022
Interconnection networks and networks-on-chip
network-on-chip design
0.612022
Application Defined On-chip Networks for Heterogeneous Chiplets: An Implementation Perspective · HPCA 2022

Methods — techniques the papers use, named apart from their topics

application-architecture co-design · 0.6
YearPublicationVenuePosition
2022 Application Defined On-chip Networks for Heterogeneous Chiplets: An Implementation Perspective
abstract
With the help of advanced packaging technologies to integrate multiple chips (e.g., CPU, AI, IO), a chiplet-based SoC design process can enable fast system construction. However, the design of network-on-chip used within the individual chiplets and across chiplets is an extremly challenge. We introduce the design process and methodology of a bufferless multi-ring NoC for heterogeneous chiplet-based SoC. Our design is portable and can be used in diverse scenarios, like Server-CPU, AI-Processor, and Baseband-Processor.The co-design of the application, architecture, and implementation is the key to make the system power efficient and high performance. We determined many architectural design choices by reflecting an analysis of a set of target applications by application teams and several physical implementation constraints provided by development teams. In this paper, we present the pragmatic practice of our co-design effort for the NoC. As a result, the system has been proven to achieve 16TB/s bandwidth in an AI processor and low latency, in a server CPU with nearly one hundred cores.
Shaolin Xiang
HPCA3