VLDB 2026 Research / reviewers in the wild / expert
Xin He 0054
dblp:69/1798-54
· DBLP profile ↗
11ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0002-1481-3179ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InferFast: Bridging the Gap Between Unstructured LLM Sparsity and Practical GPU ThroughputabstractThe high computational and memory burden of Large Language Model (LLM) inference has spurred significant interest in model sparsification. While unstructured pruning effectively reduces parameters with minimal accuracy loss, achieving practical speedups on hardware optimized for dense matrix multiplication, such as GPUs with Tensor Cores, remains a major challenge. Traditional sparse formats incur substantial decoding overhead at practical sparsity levels (30%–90%), often causing sparse computations to underperform their dense counterparts. In this paper, we present InferFast, a high-performance inference framework that unlocks the potential of unstructured sparsity for LLMs. The core of our approach is a novel sparse encoding format called Compact Dual Position Tensor Core Bitmap Encoding (CDP-TCBE), designed for minimal decoding overhead and native compatibility with tensor core operations. Building on this format, InferFast employs a suite of system-level optimizations, including hierarchical weight reordering, efficient data movement, vectorized bitmap decoding, and a double-buffered pipeline, to maximize hardware utilization by effectively overlapping decoding, data transfer, and computation. Our evaluation demonstrates that InferFast achieves significant performance improvements over state-of-the-art dense and sparse inference engines. At the kernel level, it significantly outperforms state-of-the-art SpMM baselines and achieves an average speedup of 2.01 × compared to dense GEMM. At end-to-end framework level on OPT-30B, OPT-66B, and Llama 2-13B models, InferFast increases throughput by up to 1.82 × at medium sparsity levels (70%), effectively bridging the gap between the theoretical benefits of sparsification and practical deployment efficiency. The code of InferFast is publicly available at https://github.com/MLsys-HPC/InferFast. Weifeng Bu, Hao Chen 0002, Xin He 0054 |
ICS | 5 |
| 2025 | Cherry: Breaking the GPU Memory Wall for Large-Scale GNN Training via Micro-BatchingabstractGraph Neural Networks (GNNs) have shown remarkable performance across a variety of graph-related tasks.Recent efforts indicate that GNN performance can be enhanced through more sophisticated strategies, such as employing advanced aggregators, increasing aggregation depth, and utilizing larger sampling rates, etc.While these strategies yield promising results, it also incurs a significantly larger memory footprint that can easily surpass the GPU memory capacity.Micro-batching has emerged as a promising method to mitigate GPU memory bottleneck while preserving model accuracy.Nevertheless, integrating micro-batches into GNN Yan Wang 0022, Haoran Kong, Hao Chen 0002, Weile Jia, Dingwen Tao, Xin He 0054 |
ICS | 9 |
| 2025 | Graph Transformer-Based Dynamic Edge Interaction Encoding for Traffic PredictionabstractTraffic prediction is an essential function of intelligent transportation system for traffic control and autonomous driving. Most existing methods encode traffic spatial and temporal data separately, and then design a feature fusion module to correlate spatial and temporal features. However, spatial information is often static, and repetitive static spatial encoding leads to waste of resources, especially in large-scale traffic network prediction. In this paper, we propose a dynamic edge interaction encoding method for spatio-temporal features based on inverse Transformer (iTransformer) and Graph Transformer, named iTPGT-former. The dynamic edge interaction process is designed to embed dynamic temporal features into static edges via a convolutional embedding module. To enhance the Graph Transformer, a relative position encoding strategy based on the self-attentive score of the positive definite kernel (PDK) on graphs and a method for graph substructure encoding (GSE) via enumeration of paths are introduced. In the experimental and discussion session, the iTPGT-former is considered for accuracy, parameters, inference speed, and rich ablation experiments are provided based on six publicly available traffic datasets. The results show that iTPGT-former outperforms the baseline model in both traffic flow and traffic speed prediction. The maximum improvement is achieved in the METR-LA 60-min speed prediction task, with 15.2% reduction in Mean Absolute Percentage Error (MAPE). In addition, the inference of iTPGT-former is significantly faster than the GCN-based method. Our implementation of the iTPGT-former is available athttps://github.com/ouyangnann/iTPGTN-former. Nan Ouyang, Lei Ao, Wenkang Wan, Xiaojiang Ren, Xin He 0054 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Centimani: Enabling Fast AI Accelerator Selection for DNN Training with a Novel Performance Predictor
Murali Emani, Xiaodong Yu 0001, Dingwen Tao, Xin He 0054, Pengfei Su 0001, Keren Zhou 0001, Venkatram Vishwanath |
USENIX ATC | 5 |
| 2023 | On the Performance Intricacies of Persistent Memory Aware Storage EnginesabstractAs key components of DBMSs, various storage engines and index structures have been proposed based on incorrect assumptions before PMem hardware is publicly available. Recent studies reveal that there is a significant performance gap in evaluating index structures on real PMem platforms as compared to DRAM-based emulators. However, a comprehensive evaluation for those PMem-aware database storage engines on real PMem hardware is still missing. Meanwhile, dynamic memory management is more important on PMem systems because PMem is slower than DRAM and unfriendly to random small-writes, and ensuring crash-consistency for the metadata of PMem allocators introduces extra overhead. Therefore, it is essential to understand the performance intricacies of PMem-aware database storage engines from the perspective of PMem allocators. This paper presents a systematic evaluation of three PMem-aware database storage engines using representative workloads and a unified benchmarking framework that is integrated with four PMem allocators. Besides the commonly used metrics, the impact of different hardware configurations (such as NUMA and eADR) on performance is also considered. Through in-depth analysis, we reveal caveats and pitfalls on using or designing PMem-aware storage engines and important insights that can serve as guidelines for future development of PMem allocators and other related components. Zhiwen Chen 0006, Wenkui Che, Daokun Hu, Xin He 0054, Jianhua Sun 0002, Hao Chen 0002 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Campo: Cost-Aware Performance Optimization for Mixed-Precision Neural Network Training
Xin He 0054, Jianhua Sun 0002, Hao Chen 0002, Dong Li 0001 |
USENIX ATC | 1 |
| 2022 | CVFuzz: Detecting complexity vulnerabilities in OpenCL kernels via automated pathological input generation
Zhiwen Chen 0006, Xin He 0054, Guoyun Duan, Jianhua Sun 0002, Hao Chen 0002 |
Future Gener. Comput. Syst. | 3 |
| 2021 | Enabling energy-efficient DNN training on hybrid GPU-FPGA acceleratorsabstractDNN training consumes orders of magnitude more energy than inference and requires innovative use of accelerators to improve energy-efficiency. However, despite having complementary features, GPUs and FPGAs have been mostly used independently for the entire training process, thus neglecting the opportunity in assigning individual but distinct operations to the most suitable hardware. In this paper, we take the initiative to explore new opportunities and viable solutions in enabling energy-efficient DNN training on hybrid accelerators. To overcome fundamental challenges including avoiding training throughput loss, enabling fast design space exploration, and efficient scheduling, we propose a comprehensive framework, Hype-training, that utilizes a combination of offline characterization, performance modeling, and online scheduling of individual operations. Experimental tests using NVIDIA V100 GPUs and Intel Stratix 10 FPGAs show that, Hype-training is able to exploit a mixture of GPUs and FPGAs at a fine granularity to achieve significant energy reduction, by 44.3% on average and up to 59.7%, without any loss in training throughput. Hype-training can also enforce power caps more effectively than state-of-the-art power management mechanisms on GPUs. Xin He 0054, Hao Chen 0002, Guoyang Chen, Weifeng Zhang 0003, Dong Li 0001 |
ICS | 1 |
| 2021 | Efficient parallel A* search on multi-GPU system
Xin He 0054, Yapeng Yao, Zhiwen Chen 0006, Jianhua Sun 0002, Hao Chen 0002 |
Future Gener. Comput. Syst. | 1 |
| 2018 | Concurrent hash tables on multicore machines: Comparison, evaluation and implications
Zhiwen Chen 0006, Xin He 0054, Jianhua Sun 0002, Hao Chen 0002, Ligang He |
Future Gener. Comput. Syst. | 2 |
| 2017 | Exploring Synchronization in Cache Coherent Manycore Systems: A Case Study with Xeon PhiabstractIntel Xeon Phi is a many-core architecture, featuring more than 50 cores and 200 hardware threads. Given this scale and its other distinctive architectural features, highly-concurrent applications on Xeon Phi may behave differently than on tradi- tional multi-core systems. Yet, concurrency issues especially for synchronization intensive applications on this platform have not been thoroughly analyzed. In this paper, we conduct an extensive analysis at multiple layers, from the underlying hardware cache- coherence protocol up to the user-level applications, aiming to present the most exhaustive study of synchronization on Xeon Phi. Through a range of benchmarks, we testify the feasibility and advantage of accelerating concurrent applications with Xeon Phi. Meanwhile, we identify severe scalability issues relevant to synchronization, and solutions to these issues are discussed. We believe this work can be used as guidelines both for designing better synchronization mechanisms and in optimizing concurrent applications in order to fully exploit the capability of Xeon Phi. Xin He 0054, Zhiwen Chen 0006, Jianhua Sun 0002, Hao Chen 0002, Dong Li 0001, Zhe Quan |
ICPADS | 1 |