EDBT 2026 Demo / reviewers in the wild / expert
Zixiao Huang 0001
dblp:254/4470-1
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0000-1273-2573ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient and Adaptable Overlapping for Computation and Communication via Signaling and ReorderingabstractGenerative models have achieved remarkable success across various applications, driving the demand for multi-GPU computing. Inter-GPU communication becomes a bottleneck in multi-GPU computing systems, particularly on consumer-grade GPUs. By exploiting concurrent hardware execution, overlapping computation and communication latency becomes an effective technique for mitigating the communication overhead. We identify that an efficient and adaptable overlapping design should satisfy (1) tile-wise overlapping to maximize the overlapping opportunity, (2) interference-free computation to maintain the original computational performance, and (3) communication agnosticism to reduce the development burden against varying communication primitives. Nevertheless, current designs fail to simultaneously optimize for all of those features. Ke Hong, Minxu Liu, Qiuli Mao, Zixiao Huang 0001, Lufang Chen, Yichong Zhang, Zhenhua Zhu 0002, Guohao Dai 0001, Yu Wang 0002 |
EuroSys | 6 |
| 2026 | STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal PlanningabstractThe rapid scaling of large language models (LLMs) has significantly increased GPU memory pressure, which is further aggravated by training optimization techniques such as virtual pipeline and recomputation that disrupt tensor lifespans and introduce considerable memory fragmentation. Such fragmentation stems from the use of online GPU memory allocators in popular deep learning frameworks like PyTorch, which disregard tensor lifespans. As a result, this inefficiency can waste as much as 43% of memory and trigger out-of-memory errors, undermining the effectiveness of optimization methods. Zixiao Huang 0001, Hao Lin 0005, Chunyang Zhu, Yueran Tang, Quanlu Zhang, Zhenhua Li 0001, Shengen Yan, Zhenhua Zhu 0002, Guohao Dai 0001, Yu Wang 0002 |
EuroSys | 1 |
| 2024 | FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAsabstractTransformer-based Large Language Models (LLMs) have made a significant impact on various domains. However, LLMs' efficiency suffers from both heavy computation and memory overheads. Compression techniques like sparsification and quantization are commonly used to mitigate the gap between LLM's computation/memory overheads and hardware capacity. However, existing GPU and transformer-based accelerators cannot efficiently process compressed LLMs, due to the following unresolved challenges: low computational efficiency, underutilized memory bandwidth, and large compilation overheads. This paper proposes FlightLLM, enabling efficient LLMs inference with a complete mapping flow on FPGAs. In FlightLLM, we highlight an innovative solution that the computation and memory overhead of LLMs can be solved by utilizing FPGA-specific resources (e.g., DSP48 and heterogeneous memory hierarchy). We propose a configurable sparse DSP chain to support different sparsity patterns with high computation efficiency. Second, we propose an always-on-chip decode scheme to boost memory bandwidth with mixed-precision support. Finally, to make FlightLLM available for real-world LLMs, we propose a length adaptive compilation method to reduce the compilation overhead. Implemented on the Xilinx Alveo U280 FPGA, FlightLLM achieves 6.0× higher energy efficiency and 1.8× better cost efficiency against commercial GPUs (e.g., NVIDIA V100S) on modern LLMs (e.g., LLaMA2-7B) using vLLM and SmoothQuant under the batch size of one. FlightLLM beats NVIDIA A100 GPU with 1.2× higher throughput using the latest Versal VHK158 FPGA. Shulin Zeng, Jun Liu 0117, Guohao Dai 0001, Tianyu Fu 0004, Wenheng Ma, Hanbo Sun, Zixiao Huang 0001, Yadong Dai, Jintao Li 0002, Kairui Wen, Xuefei Ning, Yu Wang 0002 |
FPGA | 10 |