EDBT 2026 Demo / reviewers in the wild / expert
Gengyu Rao
dblp:281/8681
· DBLP profile ↗
2ranked-venue papers
1as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 32% Parallel and multicore computing · 21% Memory systems · 21% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
instruction set architecture |
0.6 | 1 | 2022 | SparseCore: stream ISA and processor specialization for sparse computation · ASPLOS 2022 |
Hardware accelerators and domain-specific architectures
sparse computation |
0.6 | 1 | 2022 | SparseCore: stream ISA and processor specialization for sparse computation · ASPLOS 2022 |
Hardware accelerators and domain-specific architectures › sparsity exploitation
sparse tensor computation |
0.6 | 1 | 2022 | SparseCore: stream ISA and processor specialization for sparse computation · ASPLOS 2022 |
Distributed systems
distributed graph processing |
0.4 | 1 | 2019 | Distributed Graph Processing System and Processing-in-memory Architecture with Precise Loop-carried Dependency Guarantee · ACM Trans. Comput. Syst. 2019 |
Parallel and multicore computing
graph processing |
0.4 | 1 | 2019 | Distributed Graph Processing System and Processing-in-memory Architecture with Precise Loop-carried Dependency Guarantee · ACM Trans. Comput. Syst. 2019 |
Parallel and multicore computing
parallel programming models |
0.4 | 1 | 2019 | Distributed Graph Processing System and Processing-in-memory Architecture with Precise Loop-carried Dependency Guarantee · ACM Trans. Comput. Syst. 2019 |
Memory systems › processing-in-memory
PIM-based graph processing |
0.4 | 1 | 2019 | Distributed Graph Processing System and Processing-in-memory Architecture with Precise Loop-carried Dependency Guarantee · ACM Trans. Comput. Syst. 2019 |
Memory systems
processing-in-memory |
0.4 | 1 | 2019 | Distributed Graph Processing System and Processing-in-memory Architecture with Precise Loop-carried Dependency Guarantee · ACM Trans. Comput. Syst. 2019 |
Methods — techniques the papers use, named apart from their topics
simulation · 1.1ISA extension · 1.1dependency propagation · 0.4circulant scheduling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | SparseCore: stream ISA and processor specialization for sparse computationabstractComputation on sparse data is becoming increasingly important for many applications. Recent sparse computation accelerators are designed for specific algorithm/application, making them inflexible with software optimizations. This paper proposes SparseCore, the first general-purpose processor extension for sparse computation that can flexibly accelerate complex code patterns and fast-evolving algorithms. We extend the instruction set architecture (ISA) to make stream or sparse vector first-class citizens, and develop efficient architectural components to support the stream ISA. The novel ISA extension intrinsically operates on streams, realizing both efficient data movement and computation. The simulation results show that SparseCore achieves significant speedups for sparse tensor computation and graph pattern computation. Gengyu Rao, Jingji Chen, Jason Yik, Xuehai Qian |
ASPLOS | 1 |
| 2019 | Distributed Graph Processing System and Processing-in-memory Architecture with Precise Loop-carried Dependency GuaranteeabstractTo hide the complexity of the underlying system, graph processing frameworks ask programmers to specify graph computations in user-defined functions (UDFs) of graph-oriented programming model. Due to the nature of distributed execution, current frameworks cannot precisely enforce the semantics of UDFs, leading to unnecessary computation and communication. It exemplifies a gap between programming model and runtime execution. This article proposes novel graph processing frameworks for distributed system and Processing-in-memory (PIM) architecture that precisely enforces loop-carried dependency; i.e., when a condition is satisfied by a neighbor, all following neighbors can be skipped. Our approach instruments the UDFs to express the loop-carried dependency, then the distributed execution framework enforces the precise semantics by performing dependency propagation dynamically. Enforcing loop-carried dependency requires the sequential processing of the neighbors of each vertex distributed in different nodes. We propose to circulant scheduling in the framework to allow different nodes to process disjoint sets of edges/vertices in parallel while satisfying the sequential requirement. The technique achieves an excellent trade-off between precise semantics and parallelism—the benefits of eliminating unnecessary computation and communication offset the reduced parallelism. We implement a new distributed graph processing framework SympleGraph, and two variants of runtime systems— GraphS and GraphSR —for PIM-based graph processing architecture, which significantly outperform the state-of-the-art. Youwei Zhuo, Jingji Chen, Gengyu Rao, Qinyi Luo, Yanzhi Wang 0001, Hailong Yang 0002, Depei Qian 0001, Xuehai Qian |
ACM Trans. Comput. Syst. | 3 |