Xin He 0054

dblp:69/1798-54 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0002-1481-3179ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 InferFast: Bridging the Gap Between Unstructured LLM Sparsity and Practical GPU Throughput
abstract
The high computational and memory burden of Large Language Model (LLM) inference has spurred significant interest in model sparsification. While unstructured pruning effectively reduces parameters with minimal accuracy loss, achieving practical speedups on hardware optimized for dense matrix multiplication, such as GPUs with Tensor Cores, remains a major challenge. Traditional sparse formats incur substantial decoding overhead at practical sparsity levels (30%–90%), often causing sparse computations to underperform their dense counterparts. In this paper, we present InferFast, a high-performance inference framework that unlocks the potential of unstructured sparsity for LLMs. The core of our approach is a novel sparse encoding format called Compact Dual Position Tensor Core Bitmap Encoding (CDP-TCBE), designed for minimal decoding overhead and native compatibility with tensor core operations. Building on this format, InferFast employs a suite of system-level optimizations, including hierarchical weight reordering, efficient data movement, vectorized bitmap decoding, and a double-buffered pipeline, to maximize hardware utilization by effectively overlapping decoding, data transfer, and computation. Our evaluation demonstrates that InferFast achieves significant performance improvements over state-of-the-art dense and sparse inference engines. At the kernel level, it significantly outperforms state-of-the-art SpMM baselines and achieves an average speedup of 2.01 × compared to dense GEMM. At end-to-end framework level on OPT-30B, OPT-66B, and Llama 2-13B models, InferFast increases throughput by up to 1.82 × at medium sparsity levels (70%), effectively bridging the gap between the theoretical benefits of sparsification and practical deployment efficiency. The code of InferFast is publicly available at https://github.com/MLsys-HPC/InferFast.
Weifeng Bu, Hao Chen 0002, Xin He 0054
ICS5
2025 Cherry: Breaking the GPU Memory Wall for Large-Scale GNN Training via Micro-Batching
abstract
Graph Neural Networks (GNNs) have shown remarkable performance across a variety of graph-related tasks.Recent efforts indicate that GNN performance can be enhanced through more sophisticated strategies, such as employing advanced aggregators, increasing aggregation depth, and utilizing larger sampling rates, etc.While these strategies yield promising results, it also incurs a significantly larger memory footprint that can easily surpass the GPU memory capacity.Micro-batching has emerged as a promising method to mitigate GPU memory bottleneck while preserving model accuracy.Nevertheless, integrating micro-batches into GNN
Yan Wang 0022, Haoran Kong, Hao Chen 0002, Weile Jia, Dingwen Tao, Xin He 0054
ICS9
2025 Graph Transformer-Based Dynamic Edge Interaction Encoding for Traffic Prediction
abstract
Traffic prediction is an essential function of intelligent transportation system for traffic control and autonomous driving. Most existing methods encode traffic spatial and temporal data separately, and then design a feature fusion module to correlate spatial and temporal features. However, spatial information is often static, and repetitive static spatial encoding leads to waste of resources, especially in large-scale traffic network prediction. In this paper, we propose a dynamic edge interaction encoding method for spatio-temporal features based on inverse Transformer (iTransformer) and Graph Transformer, named iTPGT-former. The dynamic edge interaction process is designed to embed dynamic temporal features into static edges via a convolutional embedding module. To enhance the Graph Transformer, a relative position encoding strategy based on the self-attentive score of the positive definite kernel (PDK) on graphs and a method for graph substructure encoding (GSE) via enumeration of paths are introduced. In the experimental and discussion session, the iTPGT-former is considered for accuracy, parameters, inference speed, and rich ablation experiments are provided based on six publicly available traffic datasets. The results show that iTPGT-former outperforms the baseline model in both traffic flow and traffic speed prediction. The maximum improvement is achieved in the METR-LA 60-min speed prediction task, with 15.2% reduction in Mean Absolute Percentage Error (MAPE). In addition, the inference of iTPGT-former is significantly faster than the GCN-based method. Our implementation of the iTPGT-former is available athttps://github.com/ouyangnann/iTPGTN-former.
Nan Ouyang, Lei Ao, Wenkang Wan, Xiaojiang Ren, Xin He 0054
IEEE Trans. Intell. Transp. Syst.6
2024 Centimani: Enabling Fast AI Accelerator Selection for DNN Training with a Novel Performance Predictor
Murali Emani, Xiaodong Yu 0001, Dingwen Tao, Xin He 0054, Pengfei Su 0001, Keren Zhou 0001, Venkatram Vishwanath
USENIX ATC5
2023 On the Performance Intricacies of Persistent Memory Aware Storage Engines
abstract
As key components of DBMSs, various storage engines and index structures have been proposed based on incorrect assumptions before PMem hardware is publicly available. Recent studies reveal that there is a significant performance gap in evaluating index structures on real PMem platforms as compared to DRAM-based emulators. However, a comprehensive evaluation for those PMem-aware database storage engines on real PMem hardware is still missing. Meanwhile, dynamic memory management is more important on PMem systems because PMem is slower than DRAM and unfriendly to random small-writes, and ensuring crash-consistency for the metadata of PMem allocators introduces extra overhead. Therefore, it is essential to understand the performance intricacies of PMem-aware database storage engines from the perspective of PMem allocators. This paper presents a systematic evaluation of three PMem-aware database storage engines using representative workloads and a unified benchmarking framework that is integrated with four PMem allocators. Besides the commonly used metrics, the impact of different hardware configurations (such as NUMA and eADR) on performance is also considered. Through in-depth analysis, we reveal caveats and pitfalls on using or designing PMem-aware storage engines and important insights that can serve as guidelines for future development of PMem allocators and other related components.
Zhiwen Chen 0006, Wenkui Che, Daokun Hu, Xin He 0054, Jianhua Sun 0002, Hao Chen 0002
IEEE Trans. Knowl. Data Eng.4
2022 Campo: Cost-Aware Performance Optimization for Mixed-Precision Neural Network Training
Xin He 0054, Jianhua Sun 0002, Hao Chen 0002, Dong Li 0001
USENIX ATC1
2022 CVFuzz: Detecting complexity vulnerabilities in OpenCL kernels via automated pathological input generation
Zhiwen Chen 0006, Xin He 0054, Guoyun Duan, Jianhua Sun 0002, Hao Chen 0002
Future Gener. Comput. Syst.3
2021 Enabling energy-efficient DNN training on hybrid GPU-FPGA accelerators
abstract
DNN training consumes orders of magnitude more energy than inference and requires innovative use of accelerators to improve energy-efficiency. However, despite having complementary features, GPUs and FPGAs have been mostly used independently for the entire training process, thus neglecting the opportunity in assigning individual but distinct operations to the most suitable hardware. In this paper, we take the initiative to explore new opportunities and viable solutions in enabling energy-efficient DNN training on hybrid accelerators. To overcome fundamental challenges including avoiding training throughput loss, enabling fast design space exploration, and efficient scheduling, we propose a comprehensive framework, Hype-training, that utilizes a combination of offline characterization, performance modeling, and online scheduling of individual operations. Experimental tests using NVIDIA V100 GPUs and Intel Stratix 10 FPGAs show that, Hype-training is able to exploit a mixture of GPUs and FPGAs at a fine granularity to achieve significant energy reduction, by 44.3% on average and up to 59.7%, without any loss in training throughput. Hype-training can also enforce power caps more effectively than state-of-the-art power management mechanisms on GPUs.
Xin He 0054, Hao Chen 0002, Guoyang Chen, Weifeng Zhang 0003, Dong Li 0001
ICS1
2021 Efficient parallel A* search on multi-GPU system
Xin He 0054, Yapeng Yao, Zhiwen Chen 0006, Jianhua Sun 0002, Hao Chen 0002
Future Gener. Comput. Syst.1
2018 Concurrent hash tables on multicore machines: Comparison, evaluation and implications
Zhiwen Chen 0006, Xin He 0054, Jianhua Sun 0002, Hao Chen 0002, Ligang He
Future Gener. Comput. Syst.2
2017 Exploring Synchronization in Cache Coherent Manycore Systems: A Case Study with Xeon Phi
abstract
Intel Xeon Phi is a many-core architecture, featuring more than 50 cores and 200 hardware threads. Given this scale and its other distinctive architectural features, highly-concurrent applications on Xeon Phi may behave differently than on tradi- tional multi-core systems. Yet, concurrency issues especially for synchronization intensive applications on this platform have not been thoroughly analyzed. In this paper, we conduct an extensive analysis at multiple layers, from the underlying hardware cache- coherence protocol up to the user-level applications, aiming to present the most exhaustive study of synchronization on Xeon Phi. Through a range of benchmarks, we testify the feasibility and advantage of accelerating concurrent applications with Xeon Phi. Meanwhile, we identify severe scalability issues relevant to synchronization, and solutions to these issues are discussed. We believe this work can be used as guidelines both for designing better synchronization mechanisms and in optimizing concurrent applications in order to fully exploit the capability of Xeon Phi.
Xin He 0054, Zhiwen Chen 0006, Jianhua Sun 0002, Hao Chen 0002, Dong Li 0001, Zhe Quan
ICPADS1