EDBT 2026 Demo / reviewers in the wild / expert
Lihan Hu
dblp:277/8552
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I/O-Aware PIM Acceleration for Long-Sequence LLM Inference with Hybrid Sparse Attention
Xiaoyang Lu, Lihan Hu, Hongrui Huang, Peng Jiang 0004, Xian-He Sun |
IPDPS | 2 |
| 2025 | Improving Accuracy and Efficiency of Graph Embedding Training with Fine-Grained Parameter ManagementabstractEfficient and accurate graph embedding learning is crucial for various real-world applications. However, the large-scale nature of graph embeddings poses significant challenges, particularly in managing the massive amount of embedding parameters across CPU and GPU memories. This paper presents a fine-grained parameter management technique that significantly improves the accuracy and efficiency of graph embedding learning. Our approach leverages parameter duplication and a novel precaching strategy, which minimizes the overhead of data movement between CPU and GPU while preventing stale data usage during training. We introduce an analytical model to estimate the access frequency of embeddings, allowing for the optimal placement of embedding data between CPU and GPU. Using a zero-copy data access mechanism, our system effectively reduces training time while maintaining high accuracy. Experimental results on multiple large-scale knowledge graphs demonstrate that our approach achieves substantial performance gains compared to existing methods with improved Mean Reciprocal Rank (MRR). Lihan Hu, Peng Jiang 0004 |
IPDPS | 1 |
| 2025 | Matcha: A Language and Compiler for Backtracking-Based Subgraph MatchingabstractSubgraph matching is one of the most fundamental tasks in graph analytics. Numerous algorithms and systems have been proposed for the task. However, due to the diverse optimizations proposed in previous work and their targeting of different hardware, comparing and integrating existing techniques has become increasingly challenging. In this work, we propose Matcha, a domain-specific language for implementing subgraph matching algorithms. Compared to previous systems, Matcha provides a lower-level programming interface that allows users to express a wider variety of subgraph matching algorithms. This simplifies the comparisons of existing techniques and facilitates the development of new algorithms. We implement a compiler that translates and optimizes Matcha programs into C++/CUDA code for CPU and GPU execution. Our experiments show that Matcha can readily reproduce the performance of state-of-the-art subgraph matching systems. By incorporating additional optimizations, Matcha can achieve speedups up to 60x against the existing systems. Yihua Wei, Lihan Hu, Peng Jiang 0004 |
IPDPS | 2 |
| 2024 | cuKE: An Efficient Code Generator for Score Function Computation in Knowledge Graph EmbeddingabstractKnowledge graph embedding (KGE) plays an important role in graph mining and learning applications by converting discrete graph structures to continuous vector representations. While previous systems have focused on scaling KGE onto multiple GPUs, the score function computation on each GPU can be a performance bottleneck. Existing KGE systems implement the score functions with separate tensor operations, leading to large memory consumption and poor memory access efficiency. To overcome the issues, we propose a code generator that automatically translates Python-like definitions of KGE score functions into efficient CUDA code. Our code generator exploits the unique feature of KGE score functions and performs an aggressive fusion of tensor operations. Additionally, our generated code performs a runtime inspection to reduce redundant memory access for edges with identical indices. Experiments show that our generated code uses much less memory than previous systems and achieves an average speedup of 14.9x over TorchScript and 7.8x over TVM. Lihan Hu, Peng Jiang 0004 |
IPDPS | 1 |
| 2022 | Exposing and Exploiting Fine-Grained Block Structures for Fast and Accurate Sparse TrainingabstractSparse training is a popular technique to reduce the overhead of training large models. Although previous work has shown promising results for nonstructured sparse models, it is still unclear whether a sparse model with structural constraints can be trained from scratch to high accuracy. In this work, we study the dynamic sparse training for a class of sparse models with shuffled block structures. Compared to nonstructured models, such fine-grained structured models are more hardware-friendly and can effectively accelerate the training process. We propose an algorithm that keeps adapting the sparse model while maintaining the active parameters in shuffled blocks. We conduct experiments on a variety of networks and datasets and obtain positive results. In particular, on ImageNet, we achieve dense accuracy for ResNet50 and ResNet18 at 0.5 sparsity. On CIFAR10/100, we show that dense accuracy can be recovered at 0.6 sparsity for various models. At higher sparsity, our algorithm can still match the accuracy of nonstructured sparse training in most cases, while reducing the training time by up to 5x due to the fine-grained block structures in the models. Peng Jiang 0004, Lihan Hu, Shihui Song |
NeurIPS | 2 |
| 2020 | A disk failure prediction method based on LSTM network due to its individual specificityabstractIn current storage systems, to protect data security, disk failure prediction is required. Machine learning proved to be a method to solve the problem of disk failure prediction. However, because the disk-related values are affected by factors such as their use and usage environment, the values of different disks in the event of a failure are not the same. The normal value on one disk may be the value when another disk fails. Some studies have introduced the concept of time windows into disk failure prediction, trying to improve the ability of disk failure prediction by studying the relationship of a disk’s value change over time, and achieved good prediction results. We chose the neural network with time-series model to further validate the prediction of time series affect performance, and would like to be able to further improve the prediction performance by using a neural network. In this paper, we will introduce a disk failure prediction system based on LSTM networks. Considering the individual differences of the disks, we replace the input in the LSTM network with the continuous running records of the disks. The network will learn the disk information over a period of time and predict whether this disk will fail. With the proposed approach we are able to predict a disk will fail in next fifteen days with an average precision of 86.31. By comparing with other algorithms, our method performs well. Lihan Hu, Lixin Han, Zhenyuan Xu, Tianming Jiang, Huijun Qi |
KES | 1 |