EDBT 2026 Demo / reviewers in the wild / expert
Xuanteng Huang
dblp:270/8872
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0001-6647-1871ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FusedRec: Fused Embedding Communication for Distributed Recommendation Training on GPUsabstractRecent years have witnessed the wide adoption of deep learning recommendation models (DLRMs) for many online services. Unlike traditional DNN training, DLRMs leverage massive embeddings to represent sparse features, which are stored in distributed GPUs following the model parallel paradigm. Existing approaches adopt deduplication to eliminate replicated embeddings involved in AltoAll transfers to avoid unnecessary communication. In our practices, we have observed that such a deduplication design exacerbates interconnect inefficiency due to the fragmented embedding transfers with reduced message sizes, hindering the performance of distributed DLRM training. This paper introduces FusedRec, a fused embedding communication and lookup mechanism to tackle the inefficiency due to deduplication. By seeking the opportunities to fuse embeddings from multiple categories into a group, FusedRec conducts the communication in a combined shot to alleviate bandwidth under-utilization. Meanwhile, a categorical-aware hashing algorithm is integrated into FusedRec to retain the category information during lookup without extra communication. Combining with efficient unique and recovery operations, comprehensive results show FusedRec achieves a 37.8% throughput speedup in average compared to the SOTA industry implementation, without hurting the recommendation qualities of our in-house models used in online production environments. Xuanteng Huang, Riyang Hu, Jianchang Zhang, Fangying Chen |
AAAI | 1 |
| 2026 | FedCM: Fine-grained Kernel Scheduling and Management to Improve GPU SharingabstractGPU has become the de facto device to accelerate widespread machine learning and general purpose computing applications. Sharing a GPU is increasingly important to achieve higher throughput and better resource utilization. However, existing GPU sharing adopts either coarse-grained collocation approaches or interference-unaware spatial partition strategies that produce suboptimal results. In this paper, we propose FedCM, a kernel-level, collocation-based GPU sharing scheme to establish a federated use of compute and on-chip memory resources. FedCM evaluates the collocation potential of ready kernels and dispatches them in a way to maximize system throughput. During collocated execution, FedCM adopts kernel-wise management to arbitrate cache usage via customizing cache policies. The evaluation of our implementation on the off-the-shelf GPUs demonstrates that FedCM improves the overall throughput by 48.3% and 17.4%, compared to standard sharing baseline and prior state-of-the-art, respectively. Xuanteng Huang, Nong Xiao 0001 |
DATE | 2 |
| 2025 | Mpache: Interaction Aware Multi-level Cache Bypassing on GPUsabstractGraphics Processing Units (GPUs) are essential for general-purpose applications and are commonly leveraging multi-level caches to alleviate memory access pressure. However, the default cache management may lose opportunities for optimal performance in different applications. Although existing cache bypassing techniques tend to address this challenge, these methods predominantly concentrate on single-level cache, thus restricting their potential for further enhancements. To mitigate this issue, we propose Mpache, a novel software-based mechanism designed to bypass multi-level caches based on the characterization of load instructions. Mpache constructs an interaction graph and analyzes the cooperation and contention among instructions. Then, the profiling data of bypassing effectiveness guides Mpache to select the appropriate cache levels to bypass for each instruction. Finally, the design is integrated into the compiler to enable automatic bypassing for existing workloads. Evaluations on off-the-shelf GPUs show that Mpache achieves an average 1.15× speedup over the default cache policy, and effectively outperforms prior arts. Mengyue Xi, Tianyu Guo 0009, Xuanteng Huang, Zejia Lin 0001, Xianwei Zhang 0001 |
ASP-DAC | 3 |
| 2025 | PASK: Cold Start Mitigation for Inference with Proactive and Selective Kernel Loading on GPUsabstractToday, DNN inference is widely adopted, with numerous inference services being spawned from scratch across instances in scenarios such as spot serving, serverless scaling and edge computing, where frequent start-stops are required. In this work, we first delve into the inference workflow and uncover the origins of cold start when invoking a DNN model. Specifically, DNN execution is blocked by the kernel loading process to prepare the code object executing on GPU at the DL primitive library (e.g., cuDNN and MIOpen). To tackle this, we propose PASK, a kernel loading and reusing middleware to mitigate the widespread cold start issue. Unlike the reactive kernel scheduling policy used by existing frameworks, PASK adopts a proactive strategy to interleave code loading, kernel issuing and GPU computation to achieve higher hardware utilization. To further reduce the loading overhead, PASK recycles existing loaded kernels to accomplish the DNN operator, rather than introducing new kernels for every layer. Meanwhile, PASK categorically organizes the cached kernels to efficiently find the applicable kernel for reuse and thus minimize incurred runtime overhead. We implement and evaluate PASK atop of open source DNN inference engine and primitive library on off-the-shelf GPUs. Experiments demonstrate PASK is capable of alleviating the cold start overhead of popular DNN models with $5.62 \times$ speedup on average. Xuanteng Huang, Jiangsu Du, Nong Xiao 0001, Xianwei Zhang 0001 |
DAC | 1 |
| 2024 | SMILE: LLC-based Shared Memory Expansion to Improve GPU Thread Level ParallelismabstractWhile designed for massive parallelism, GPUs are frequently suffering from low thread occupancy and limited data throughput, which are typically attributed to constrained on-chip resources, such as shared memory and register file. To alleviate the pressure, last-level cache (LLC) is being substantially enlarged to support continuously growing computation and to shrink the off-chip data traffic. Nevertheless, applications can be challenging to fully utilize the excessive LLC spaces. Towards the issue, we propose to manage partial LLC in an software way instead to expand precious shared memory (SMEM), named as SMILE, helping to alleviate the low thread occupancy. SMILE splits the monolithic LLC into normal data cache and new software region, with the latter being to extend the limited SMEM. For adapting to diverse application characteristics, SMILE enables multiple splitting grades and determines the appropriate partition via online profiling. Experimental results show that SMILE achieves average performance improvements of 14.7% and 8.4% respectively, compared to the default baseline and prior state-of-the-art. Tianyu Guo 0009, Xuanteng Huang, Xianwei Zhang 0001, Nong Xiao 0001 |
DAC | 2 |
| 2024 | openLG: A Tunable and Efficient Open-source LSTM on GPUsabstractLong-Short-Term Memory (LSTM) neural networks have demonstrated exceptional proficiency in capturing both short-term and long-term dependencies within input sequences, rendering them invaluable across diverse applications, including DNA basecalling, speech recognition, reading comprehension, and scientific numerical forecasting. Despite their significance, the closed-source nature of state-of-the-art libraries, particularly the cuDNN LSTM inference kernel on NVIDIA GPUs, poses challenges in extending, optimizing, and understanding their internal workings.In this paper, we propose openLG, an open source framework to implement high performance LSTM kernel on basis of NVIDIA template kernel library CUTLASS. With flexible and transparent template configurations, developers are able to select the most suitable parameter combination for various scenarios. We formalize the LSTM computation as the time dependent and independent partitions, and then exploit efficient CUTLASS GEMM kernel for both parts. Evaluations conducted on NVIDIA GPUs demonstrate that openLG achieves up to 78% speedup on DeepBench benchmarks and an average speedup of 10% on realistic applications, effectively surpassing the performance of cuDNN. Zhaowen Shan, Xuanteng Huang, Xianwei Zhang 0001 |
IJCNN | 2 |
| 2023 | KeSCo: Compiler-based Kernel Scheduling for Multi-task GPU ApplicationsabstractNowadays, Graphics Processing Units (GPUs) dominate in a wide spectrum of computing realms and multi-task is increasingly applied in various complicated applications. To gain higher performance, multi-task programs require cumbersome programming efforts to take advantage of inter-kernel concurrency at source-code level. Although there exist works automatically scheduling kernels to enable inter-kernel concurrency, they all inevitably introduce new programming frameworks and some even bring significant performance downgrade compared to the expertise-based optimizations. To address this issue, we propose KeSCo, a compiler-based scheduler to expose kernel level concurrency in multi-task programs with trivial code modification. In compilation, KeSCo applies a strategy to schedule kernels in task queues, accounting for both load balance and synchronization cost. Also, KeSCo utilizes a customized algorithm designed for computational flow to remove redundant synchronizations. The design is further extended to support multi-process scenario, where multiple GPU processes are sharing a single context. Evaluations on representative benchmarks show that the proposed approach gains a 1.28× average speedup for multi-task scenario (1.22× for multi-process). Even with lessened programming efforts, our proposed design outperforms two state-of-the-arts GrSched and Taskflow by 1.31× and 1.16× on average, respectively. Zejia Lin 0001, Zewei Mo, Xuanteng Huang, Xianwei Zhang 0001, Yutong Lu |
ICCD | 3 |
| 2020 | MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection
Fa-Ting Hong, Xuanteng Huang, Wei-Hong Li 0001, Wei-Shi Zheng 0001 |
ECCV (13) | 2 |