Chaoyang Jia

dblp:294/1813 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 68% Processor architecture and microarchitecture · 18% GPUs and heterogeneous computing · 14%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache
cache optimization
0.912025
TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025
Memory systems › cache management
cache replacement
0.912025
TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025
Memory systems › data layout optimization
data alignment
0.912025
In-SRAM Parallel Data Shuffle · ACM Trans. Archit. Code Optim. 2025
Memory systems › data locality
data locality exploitation
0.912025
TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025
GPUs and heterogeneous computing › GPU memory management
GPU cache management
0.912025
TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025
Processor architecture and microarchitecture
SIMD
0.912025
In-SRAM Parallel Data Shuffle · ACM Trans. Archit. Code Optim. 2025
Memory systems › random-access memory
SRAM
0.912025
In-SRAM Parallel Data Shuffle · ACM Trans. Archit. Code Optim. 2025
Processor architecture and microarchitecture
vector processor
0.312025
In-SRAM Parallel Data Shuffle · ACM Trans. Archit. Code Optim. 2025
Graph algorithms and graph theory
graph processing
0.312025
TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025

Methods — techniques the papers use, named apart from their topics

cache management strategies · 1.7inter-bank word line data movement · 0.9
YearPublicationVenuePosition
2025 TSN Cache: Exploiting Data Localities in Graph Computing Applications
abstract
This article finds that the reusability of vertices in the same graph in graph processing differs, and the high-reuse and low-reuse vertices are stored together. These phenomena lead to the inability of existing GPU architectures to capture the reusability of graph processing. The most advanced cache optimization strategies cannot implement different management strategies for data with different reusability, which is an essential reason for graph processing’s poor performance. Therefore, we propose a TSN cache scheme for the GPU platform. This scheme employs distinct management strategies for data with varying reusability in the cache, effectively leveraging the locality of these different data types. In addition, the TSN cache scheme can also reduce the probability of cache thrashing caused by low-reuse data. This article evaluates multiple graph algorithms and datasets and shows that the TSN cache scheme achieves an average speedup of 1.38 compared with the baseline scheme.
Chaoyang Jia, Kai Lu 0001, Li Shen 0007
ACM Trans. Archit. Code Optim.1
2025 In-SRAM Parallel Data Shuffle
abstract
While Single Instruction Multiple Data (SIMD) units are widely employed in processors for neural networks, signal processing, and high-performance computing, they suffer from expensive shuffle operations dedicated to data alignment. In fact, shuffle operations only change the layout of data and ideally should be done entirely within memory. To this end, we propose Shuffle SRAM in this article, which can shuffle multiple data elements simultaneously across SRAM banks. The key idea is exploiting inter-bank word line wise data movement to shuffle data in parallel, where all data elements on the same word line of SRAM can be shuffled simultaneously, achieving a high level of parallelism. Through suitable data layout preparation and proper control, Shuffle SRAM efficiently supports a wide range of commonly used shuffle operations. Our evaluation results show that the Shuffle SRAM can reap performance benefits of 14.3× for data reorganization only applications and 1.97× for data reorganization + computation applications over conventional shuffle architecture on general-purpose processors. With Shuffle SRAM, the state-of-the-art vector processor can obtain 2.58× energy efficiency. Compared with traditional SRAM, Shuffle SRAM only increases 3.5% additional area overhead.
Chaoyang Jia, Dunbo Zhang, Qingjie Lang, Li Shen 0007
ACM Trans. Archit. Code Optim.1
2022 Compressed page walk cache
Dunbo Zhang, Chaoyang Jia, Li Shen 0007
Frontiers Comput. Sci.2
2021 Multi-level PWB and PWC for Reducing TLB Miss Overheads on GPUs
Dunbo Zhang, Chaoyang Jia, Qiong Wang 0001, Li Shen 0007
ICA3PP (2)3