EDBT 2026 Demo / reviewers in the wild / expert
Qikun Li
dblp:361/1926
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 34% Memory systems · 26% Hardware accelerators and domain-specific architectures · 24% | |
| Artificial intelligence
1 paper |
Graph learning · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
graph processing |
0.9 | 2 | 2025 | GraphTune: An Efficient Dependency-Aware Substrate to Alleviate Irregularity in Concurrent Graph Processing · ACM Trans. Archit. Code Optim. 2023 An Efficient ReRAM-based Accelerator for Asynchronous Iterative Graph Processing · ACM Trans. Archit. Code Optim. 2025 |
Hardware accelerators and domain-specific architectures
graph processing accelerator |
0.9 | 1 | 2025 | An Efficient ReRAM-based Accelerator for Asynchronous Iterative Graph Processing · ACM Trans. Archit. Code Optim. 2025 |
Memory systems
processing-in-memory |
0.9 | 1 | 2025 | An Efficient ReRAM-based Accelerator for Asynchronous Iterative Graph Processing · ACM Trans. Archit. Code Optim. 2025 |
Memory systems › processing-in-memory
ReRAM crossbar |
0.9 | 1 | 2025 | An Efficient ReRAM-based Accelerator for Asynchronous Iterative Graph Processing · ACM Trans. Archit. Code Optim. 2025 |
Machine learning › Graph learning
graph neural network |
0.8 | 1 | 2024 | CDA-GNN: A Chain-driven Accelerator for Efficient Asynchronous Graph Neural Network · DAC 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
graph neural network accelerator |
0.8 | 1 | 2024 | CDA-GNN: A Chain-driven Accelerator for Efficient Asynchronous Graph Neural Network · DAC 2024 |
Parallel and multicore computing › graph processing
concurrent graph processing |
0.7 | 1 | 2023 | GraphTune: An Efficient Dependency-Aware Substrate to Alleviate Irregularity in Concurrent Graph Processing · ACM Trans. Archit. Code Optim. 2023 |
Parallel and multicore computing › task scheduling
dependency-aware scheduling |
0.7 | 1 | 2023 | GraphTune: An Efficient Dependency-Aware Substrate to Alleviate Irregularity in Concurrent Graph Processing · ACM Trans. Archit. Code Optim. 2023 |
Distributed systems
distributed graph processing |
0.7 | 1 | 2023 | GraphTune: An Efficient Dependency-Aware Substrate to Alleviate Irregularity in Concurrent Graph Processing · ACM Trans. Archit. Code Optim. 2023 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.2 | 1 | 2024 | CDA-GNN: A Chain-driven Accelerator for Efficient Asynchronous Graph Neural Network · DAC 2024 |
Storage systems › out-of-core computation
out-of-core graph processing |
0.2 | 1 | 2023 | GraphTune: An Efficient Dependency-Aware Substrate to Alleviate Irregularity in Concurrent Graph Processing · ACM Trans. Archit. Code Optim. 2023 |
Methods — techniques the papers use, named apart from their topics
chain-driven execution · 1.5chain-aware caching · 1.5hybrid processing scheme · 0.9dependency-aware subgraph construction · 0.9synchronous execution engine · 0.7cross-iteration dependency graph · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Efficient ReRAM-based Accelerator for Asynchronous Iterative Graph ProcessingabstractGraph processing has become a central concern for many real-world applications and is well-known for its low compute-to-communication ratios and poor data locality. By integrating computing logic into memory, resistive random access memory (ReRAM) tackles the demand for high memory bandwidth in graph processing. Despite the years’ research efforts, existing ReRAM-based graph processing approaches still face the challenges of redundant computation overhead . It is because the vertices of many subgraphs are ineffectively and repeatedly processed over the ReRAM crossbars for lots of iterations so as to update their states according to the vertices of other subgraphs regardless of the dependencies among the subgraphs. In this article, we propose ASGraph , a dependency-aware ReRAM-based graph processing accelerator that overcomes the aforementioned performance bottlenecks. Specifically, ASGraph dynamically constructs the subgraph based on the dependencies between vertices’ states and then detects constructed subgraph that owns high value (it is likely that it has accumulated many state propagations from its neighbors and is able to affect more other neighbors) to be preferentially processed. In this way, it makes the vertex states propagate along the dependencies between vertices as much as possible to reduce the redundant computation. Besides, ASGraph employs a hybrid processing scheme to accelerate the state propagations of the tightly connected subgraph, thereby minimizing the redundant computations. Experimental results show that ASGraph achieves 25.5× and 4.8× speedup and 70.8× and 2.2× energy saving on average compared with the state-of-the-art ReRAM-based graph processing accelerators, that is, GraphR and GaaS-X, respectively. Jin Zhao 0003, Yu Zhang 0027, Donghao He, Qikun Li, Weihang Yin, Hao Qi 0004, Xiaofei Liao, Hai Jin 0001, Haikun Liu, Linchen Yu, Zhan Zhang 0003 |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | CDA-GNN: A Chain-driven Accelerator for Efficient Asynchronous Graph Neural NetworkabstractAsynchronous Graph Neural Network (AGNN) has attracted much research attention because it enables faster convergence speed than the synchronous GNN. However, existing software/hardware solutions suffer from redundant computation overhead and excessive off-chip communications for AGNN due to irregular state propagations along the dependency chains between vertices. This paper proposes a chain-driven asynchronous accelerator, CDA-GNN, for efficient AGNN inference. Specifically, CDA-GNN proposes a chain-driven asynchronous execution approach into novel accelerator design to regularize the vertex state propagations for fewer redundant computations and off-chip communications and also designs a chain-aware data caching method to improve data locality for AGNN. We have implemented and evaluated CDA-GNN on a Xilinx Alveo U280 FPGA card. Compared with the cutting-edge software solutions (i.e., Dorylus and AMP) and hardware solutions (i.e., BlockGNN and FlowGNN), CDA-GNN improves the performance of AGNN inference by an average of 1,173x, 182.4x, 10.2x, and 7.9x and saves energy by 2,241x, 242.2x, 12.4x, and 8.9x, respectively. Yu Zhang 0027, Ligang He, Donghao He, Qikun Li, Jin Zhao 0003, Xiaofei Liao, Hai Jin 0001, Lin Gu 0002, Haikun Liu |
DAC | 5 |
| 2023 | GraphTune: An Efficient Dependency-Aware Substrate to Alleviate Irregularity in Concurrent Graph ProcessingabstractWith the increasing need for graph analysis, massive Concurrent iterative Graph Processing (CGP) jobs are usually performed on the common large-scale real-world graph. Although several solutions have been proposed, these CGP jobs are not coordinated with the consideration of the inherent dependencies in graph data driven by graph topology. As a result, they suffer from redundant and fragmented accesses of the same underlying graph dispersed over distributed platform, because the same graph is typically irregularly traversed by these jobs along different paths at the same time. In this work, we develop GraphTune , which can be integrated into existing distributed graph processing systems, such as D-Galois, Gemini, PowerGraph, and Chaos, to efficiently perform CGP jobs and enhance system throughput. The key component of GraphTune is a dependency-aware synchronous execution engine in conjunction with several optimization strategies based on the constructed cross-iteration dependency graph of chunks. Specifically, GraphTune transparently regularizes the processing behavior of the CGP jobs in a novel synchronous way and assigns the chunks of graph data to be handled by them based on the topological order of the dependency graph so as to maximize the performance. In this way, it can transform the irregular accesses of the chunks into more regular ones so that as many CGP jobs as possible can fully share the data accesses to the common graph. Meanwhile, it also efficiently synchronizes the communications launched by different CGP jobs based on the dependency graph to minimize the communication cost. We integrate it into four cutting-edge distributed graph processing systems and a popular out-of-core graph processing system to demonstrate the efficiency of GraphTune. Experimental results show that GraphTune improves the throughput of CGP jobs by 3.1∼6.2, 3.8∼8.5, 3.5∼10.8, 4.3∼12.4, and 3.8∼6.9 times over D-Galois, Gemini, PowerGraph, Chaos, and GraphChi, respectively. Jin Zhao 0003, Yu Zhang 0027, Ligang He, Qikun Li, Xiaofei Liao, Hai Jin 0001, Lin Gu 0002, Haikun Liu, Bingsheng He, Ji Zhang 0001, Xianzheng Song, Lin Wang 0098, Jun Zhou 0011 |
ACM Trans. Archit. Code Optim. | 4 |