VLDB 2026 Research / reviewers in the wild / expert
Chaoyang Jia
dblp:294/1813
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 68% Processor architecture and microarchitecture · 18% GPUs and heterogeneous computing · 14% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › cache
cache optimization |
0.9 | 1 | 2025 | TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025 |
Memory systems › cache management
cache replacement |
0.9 | 1 | 2025 | TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025 |
Memory systems › data layout optimization
data alignment |
0.9 | 1 | 2025 | In-SRAM Parallel Data Shuffle · ACM Trans. Archit. Code Optim. 2025 |
Memory systems › data locality
data locality exploitation |
0.9 | 1 | 2025 | TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025 |
GPUs and heterogeneous computing › GPU memory management
GPU cache management |
0.9 | 1 | 2025 | TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025 |
Processor architecture and microarchitecture
SIMD |
0.9 | 1 | 2025 | In-SRAM Parallel Data Shuffle · ACM Trans. Archit. Code Optim. 2025 |
Memory systems › random-access memory
SRAM |
0.9 | 1 | 2025 | In-SRAM Parallel Data Shuffle · ACM Trans. Archit. Code Optim. 2025 |
Processor architecture and microarchitecture
vector processor |
0.3 | 1 | 2025 | In-SRAM Parallel Data Shuffle · ACM Trans. Archit. Code Optim. 2025 |
Graph algorithms and graph theory
graph processing |
0.3 | 1 | 2025 | TSN Cache: Exploiting Data Localities in Graph Computing Applications · ACM Trans. Archit. Code Optim. 2025 |
Methods — techniques the papers use, named apart from their topics
cache management strategies · 1.7inter-bank word line data movement · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TSN Cache: Exploiting Data Localities in Graph Computing ApplicationsabstractThis article finds that the reusability of vertices in the same graph in graph processing differs, and the high-reuse and low-reuse vertices are stored together. These phenomena lead to the inability of existing GPU architectures to capture the reusability of graph processing. The most advanced cache optimization strategies cannot implement different management strategies for data with different reusability, which is an essential reason for graph processing’s poor performance. Therefore, we propose a TSN cache scheme for the GPU platform. This scheme employs distinct management strategies for data with varying reusability in the cache, effectively leveraging the locality of these different data types. In addition, the TSN cache scheme can also reduce the probability of cache thrashing caused by low-reuse data. This article evaluates multiple graph algorithms and datasets and shows that the TSN cache scheme achieves an average speedup of 1.38 compared with the baseline scheme. Chaoyang Jia, Kai Lu 0001, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | In-SRAM Parallel Data ShuffleabstractWhile Single Instruction Multiple Data (SIMD) units are widely employed in processors for neural networks, signal processing, and high-performance computing, they suffer from expensive shuffle operations dedicated to data alignment. In fact, shuffle operations only change the layout of data and ideally should be done entirely within memory. To this end, we propose Shuffle SRAM in this article, which can shuffle multiple data elements simultaneously across SRAM banks. The key idea is exploiting inter-bank word line wise data movement to shuffle data in parallel, where all data elements on the same word line of SRAM can be shuffled simultaneously, achieving a high level of parallelism. Through suitable data layout preparation and proper control, Shuffle SRAM efficiently supports a wide range of commonly used shuffle operations. Our evaluation results show that the Shuffle SRAM can reap performance benefits of 14.3× for data reorganization only applications and 1.97× for data reorganization + computation applications over conventional shuffle architecture on general-purpose processors. With Shuffle SRAM, the state-of-the-art vector processor can obtain 2.58× energy efficiency. Compared with traditional SRAM, Shuffle SRAM only increases 3.5% additional area overhead. Chaoyang Jia, Dunbo Zhang, Qingjie Lang, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 1 |
| 2022 | Compressed page walk cache
Dunbo Zhang, Chaoyang Jia, Li Shen 0007 |
Frontiers Comput. Sci. | 2 |
| 2021 | Multi-level PWB and PWC for Reducing TLB Miss Overheads on GPUs
Dunbo Zhang, Chaoyang Jia, Qiong Wang 0001, Li Shen 0007 |
ICA3PP (2) | 3 |