EDBT 2026 Demo / reviewers in the wild / expert
Yong Li 0020
dblp:93/2334-20
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-5355-6614ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
GPUs and heterogeneous computing · 45% Parallel and multicore computing · 17% Hardware accelerators and domain-specific architectures · 10% | |
| Artificial intelligence
3 papers |
Efficient and distributed learning · 55% Graph learning · 37% Language models and text generation · 8% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
inference acceleration |
0.9 | 1 | 2025 | CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration · ICML 2025 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
0.9 | 1 | 2025 | CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration · ICML 2025 |
Machine learning › Graph learning
graph neural network training |
0.7 | 1 | 2023 | Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training · USENIX ATC 2023 |
GPUs and heterogeneous computing
graph neural network training |
0.7 | 1 | 2023 | Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training · USENIX ATC 2023 |
GPUs and heterogeneous computing
multi-GPU computing |
0.7 | 1 | 2023 | Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training · USENIX ATC 2023 |
GPUs and heterogeneous computing
GPU scheduling |
0.6 | 1 | 2022 | CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUs · SC 2022 |
Parallel and multicore computing
task scheduling |
0.6 | 1 | 2022 | CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUs · SC 2022 |
Machine learning › Graph learning
graph neural network |
0.5 | 1 | 2021 | GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021 |
Distributed systems › distributed graph processing
distributed graph processing engine |
0.5 | 1 | 2021 | GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021 |
Memory systems › DRAM › DRAM architecture
high bandwidth memory |
0.5 | 1 | 2021 | FleetRec: Large-Scale Recommendation Inference on Hybrid GPU-FPGA Clusters · KDD 2021 |
High-performance computing
large-scale graph processing |
0.5 | 1 | 2021 | GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › recommendation model accelerator
recommendation inference accelerator |
0.5 | 1 | 2021 | FleetRec: Large-Scale Recommendation Inference on Hybrid GPU-FPGA Clusters · KDD 2021 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
0.3 | 1 | 2025 | CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration · ICML 2025 |
Parallel and multicore computing › parallel algorithms › parallel primitives
data-parallel primitives |
0.1 | 1 | 2021 | GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2021 | GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021 |
Methods — techniques the papers use, named apart from their topics
graph engine optimization · 1.5declarative data-parallel operators · 1.5coefficient-of-variation-based algorithm · 0.9queue-based scheduling · 0.6cost function estimation · 0.6computation-memory disaggregation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CateKV: On Sequential Consistency for Long-Context LLM Inference AccelerationabstractLarge language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. In this work, we discover that certain attention heads exhibit sequential consistency in their attention patterns, which can be persistently identified using a coefficient-of-variation-based algorithm. Inspired by this observation, we propose CateKV, a hybrid KV cache method that retains only critical token information for consistent heads, thereby reducing KV cache size and computational overhead, while preserving the majority of KV pairs in adaptive heads to ensure high accuracy. We show the unique characteristics of our algorithm and its extension with existing acceleration methods. Comprehensive evaluations on long-context benchmarks show that, while maintaining accuracy comparable to full attention, CateKV reduces memory usage by up to $2.72\times$ and accelerates decoding by $2.18\times$ in single-sample inputs, and boosts throughput by $3.96\times$ in batch scenarios. Haoyun Jiang, Haolin Li 0001, Jianwei Zhang 0012, Fei Huang 0005, Qiang Hu 0003, Minmin Sun, Shuai Xiao 0002, Yong Li 0020, Junyang Lin, Jiangchao Yao |
ICML | 8 |
| 2024 | Prediction of SiC MOSFET Power Modules Junction Temperature for Electric Vehicle Based on Electro-Thermal Coupled ModelabstractSilicon carbide (SiC) metal oxide semiconductor field-effect transistor (MOSFET) is increasingly being used in EVs inverters because of its high switching frequency, switching speed and operating temperature characteristics. However, these characteristics can lead to a rapid increase in temperature, resulting in thermal failure of the power inverter. Therefore, predicting the SiC MOSFET junction temperature is vital for effective thermal management. First, this article establishes the SiC MOSFET electrical model to calculate the loss and obtains the loss parameters through double-pulse experiments. Next, a novel coupled thermal network model is introduced, accompanied by the development of a 3D finite element model to derive the SiC MOSFET thermal parameters. Finally, a$S$iC MOSFET power module electro-thermal coupled model is developed to predict its junction temperature and the experimental outcomes demonstrated the reliability of this method. Xing Xu 0002, Yong Li 0020, Heping Ling |
INDIN | 4 |
| 2023 | Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training
Jie Sun 0017, Li Su 0005, Zuocheng Shi, Wenting Shen, Zeke Wang, Lei Wang 0004, Jie Zhang 0081, Yong Li 0020, Wenyuan Yu, Jingren Zhou 0001, Fei Wu 0001 |
USENIX ATC | 8 |
| 2022 | CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUsabstractGraph neural networks (GNNs) suffer from low GPU utilization due to frequent memory accesses. Existing concurrent training mechanisms cannot be directly adapted to GNNs because they fail to consider the impact of input irregularity. This requires pre-profiling the memory footprint of concurrent tasks based on input dimensions to ensure successful co-location on GPU. Moreover, massive training tasks generated from scenarios such as hyper-parameter tuning require flexible scheduling strategies. To address these problems, we propose CoGNN that enables efficient management of GNN training tasks on GPUs. Specifically, the CoGNN organizes the tasks in a queue and estimates the memory consumption of each task based on cost functions at operator basis. In addition, the CoGNN implements scheduling policies to generate task groups, which are iteratively submitted for execution. The experiment results show that the CoGNN can achieve shorter completion and queuing time for training tasks from diverse GNN models. Qingxiao Sun, Yi Liu 0013, Hailong Yang 0002, Ruizhe Zhang 0012, Ming Dun, Mingzhen Li 0001, Wencong Xiao, Yong Li 0020, Zhongzhi Luan, Depei Qian 0001 |
SC | 9 |
| 2021 | FleetRec: Large-Scale Recommendation Inference on Hybrid GPU-FPGA ClustersabstractWe present FleetRec, a high-performance and scalable recommendation inference system within tight latency constraints. FleetRec takes advantage of heterogeneous hardware including GPUs and the latest FPGAs equipped with high-bandwidth memory. By disaggregating computation and memory to different types of hardware and bridging their connections by high-speed network, FleetRec gains the best of both worlds, and can naturally scale out by adding nodes to the cluster. Experiments on three production models up to 114 GB show that FleetRec outperforms optimized CPU baseline by more than one order of magnitude in terms of throughput while achieving significantly lower latency. Wenqi Jiang 0001, Zhenhao He, Shuai Zhang 0007, Kai Zeng 0002, Jiansong Zhang 0001, Tongxuan Liu, Yong Li 0020, Jingren Zhou 0001, Ce Zhang 0001, Gustavo Alonso |
KDD | 8 |
| 2021 | GraphScope: A Unified Engine For Big Graph ProcessingabstractGraphScope is a system and a set of language extensions that enable a new programming interface for large-scale distributed graph computing. It generalizes previous graph processing frameworks (e.g. , Pregel, GraphX) and distributed graph databases ( e.g ., Janus-Graph, Neptune) in two important ways: by exposing a unified programming interface to a wide variety of graph computations such as graph traversal, pattern matching, iterative algorithms and graph neural networks within a high-level programming language; and by supporting the seamless integration of a highly optimized graph engine in a general purpose data-parallel computing system. A GraphScope program is a sequential program composed of declarative data-parallel operators, and can be written using standard Python development tools. The system automatically handles the parallelization and distributed execution of programs on a cluster of machines. It outperforms current state-of-the-art systems by enabling a separate optimization (or family of optimizations) for each graph operation in one carefully designed coherent framework. We describe the design and implementation of GraphScope and evaluate system performance using several real-world applications. Wenfei Fan, Tao He 0013, Longbin Lai, Xue Li 0024, Yong Li 0020, Zhao Li 0007, Zhengping Qian, Chao Tian 0001, Lei Wang 0004, Jingbo Xu 0001, Youyang Yao, Qiang Yin 0002, Wenyuan Yu, Kai Zeng 0002, Jingren Zhou 0001, Diwen Zhu |
Proc. VLDB Endow. | 5 |