EDBT 2026 Demo / reviewers in the wild / expert
Hongchao Du
dblp:257/1447
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-7995-1834ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breaking the I/O Bottleneck: I/O Coordination Optimization for Efficient Large-Scale LLM Fine-TuningabstractLarge Language Models (LLMs) with tens or even hundreds of billions of parameters have become the foundation of modern AI applications. However, fine-tuning such massive models is severely constrained by the limited GPU memory. Existing memory-saving systems, such as ZeRO-based offloading in DeepSpeed, reduce GPU memory usage but inevitably incur substantial I/O overhead, especially when model states reside on slow storage devices, such as NVMe SSDs. As a result, the memory bottleneck in large-scale fine-tuning is transformed into an I/O bottleneck. Although prior systems have employed strategies like parameter prefetching and partial asynchronous execution, they remain limited by synchronous I/O-communication dependencies and the lack of fine-grained read/write I/O scheduling. To address these limitations, we propose IOC, an I/O Coordination Optimization framework that maximizes pipeline parallelism across different phases of LLM fine-tuning. IOC introduces three key mechanisms: (1) An All-Gather prefetching technique based on an I/O state hash table, which completely decouples All-Gather prefetching from parameter I/O, achieving continuous overlap among I/O, communication, and computation; (2) The parameter update phase is refactored into an asynchronous pipeline with explicit I/O isolation, where the optimizer state write-back is executed in a semi-asynchronous manner, thereby mitigating read/write contention and reducing synchronization stalls; (3) Multi-disk parallelism is leveraged by introducing an additional disk to further relieve I/O contention and defer synchronization waits to the latest possible time point. Experimental results demonstrate that IOC significantly accelerates LLM finetuning while preserving low memory consumption. The end-toend fine-tuning time on the Llama-70B model is reduced by $\mathbf{2 1. 5 \%}$ and 34.3% in single-disk and multi-disk configurations compared to the baseline. Ziyang Shen, Hongchao Du, Kaihuan Lin, Yin Lin, Qiao Li 0001, Chun Jason Xue |
ISPASS | 2 |
| 2025 | ArtMem: Adaptive Migration in Reinforcement Learning-Enabled Tiered MemoryabstractWith the increasing memory demands of emerging applications, tiered memory has become a viable solution for reducing data center hardware costs.Given the low performance of the capacity tiers in tiered memory systems, optimizing memory management is crucial in improving overall system performance.This paper identifies three key limitations in existing tiered memory solutions.First, existing solutions often perform differently across different workloads, leading to suboptimal performance in some workloads.Second, they often fail to adjust migration strategies in response to low fast memory tier access rates, resulting in ineffective data placement.Third, they often miss the opportunity to dynamically tune the memory migration scope based on workload patterns, leading to unnecessary page migrations and under-utilization of tiered memory potential.This paper proposes ArtMem, a reinforcement learning (RL)-driven framework that dynamically manages tiered memory systems and adapts to workload evolution to address these limitations.ArtMem enables better placement of memory pages, enhancing system performance while reducing unnecessary migrations.Experimental evaluations show that ArtMem outperforms state-of-the-art tiering systems, achieving 35% -172% performance improvements over diverse workloads. Xinyue Yi, Hongchao Du, Yu Wang 0002, Jie Zhang 0048, Qiao Li 0001, Chun Jason Xue |
ISCA | 2 |
| 2025 | EvoP: Robust LLM Inference via Evolutionary Pruning
Shangyu Wu, Hongchao Du, Tei-Wei Kuo, Nan Guan, Chun Jason Xue |
NLPCC (1) | 2 |
| 2025 | SolFS: An Operation-Log Versioning File System for Hash-free Efficient Mobile Cloud Backup
Riwei Pan, Yu Liang 0004, Lei Li 0067, Hongchao Du, Tei-Wei Kuo, Chun Jason Xue |
USENIX ATC | 4 |
| 2023 | Multi-Granularity Shadow Paging with NVM Write Optimization for Crash-Consistent Memory-Mapped I/OabstractThe complex software stack has become the performance bottleneck of the system with high-speed Non-Volatile Memory (NVM). Memory-mapped I/O (MMIO) could avoid the long-stack overhead by bypassing the kernel, but the performance is limited by existing crash-resilient mechanisms. We propose a Multi-Granularity Shadow Paging (MGSP) strategy, which smartly utilizes the redo and undo logs as shadow logs to provide a light-weight crash-resilient mechanism for MMIO. In addition, a multi-granularity strategy is designed to provide high-performance updating and locking for reducing runtime overhead, where strong consistency is preserved with a lockfree metadata log. Experimental results show that the proposed MGSP achieves 1.1 ~ 4.21× performance improvement with write and 2.56 ~ 3.76× improvement with multi-threads write compared with the underlying file system. For SQLite, MGSP can improve the database performance by 29.4% for Mobibench and 36.5% for TPCC, on average. Hongchao Du, Qiao Li 0001, Riwei Pan, Tei-Wei Kuo, Chun Jason Xue |
HPCA | 1 |
| 2022 | A Practical Highly Paralleled ReRAM-Based DNN Accelerator by Reusing Weight Pattern RepetitionsabstractResistive random access memory (ReRAM)-based processing-in-memory (PIM) architecture has been designed to accelerate deep neural networks (DNNs) by concurring computation and memory barriers. To further improve memory and computation efficiency, the weight sparsity characteristic has been explored to optimize the ReRAM-based DNN accelerators. However, these designs only focus on compressing zero weights to eliminate ineffectual computation. In this article, we thoroughly analyze the weight distribution characteristics of several typical DNN models and observe many nonzero weight pattern repetitions (WPRs). Therefore, there is an opportunity to further improve the performance and energy efficiency by reusing these WPR. We propose a novel ReRAM-based accelerator—PattPIM, to achieve space compression and computation reuse by exploring DNN WPR based on practical ReRAM crossbars. In PattPIM, we propose a configurable WPR-aware DNN engine and a WPR-to-OU mapping scheme to save both space and computation resources. An intraprocessing engine (PE) pipeline is designed to improve the parallelism of the computation process. Furthermore, we adopt an approximate weight pattern transform algorithm to improve the DNN WPR ratio to enhance the reuse efficiency with negligible accuracy loss. Our evaluation with 6 DNN models shows that the proposed PattPIM delivers significant performance improvement, ReRAM resource efficiency and energy saving. Yuhao Zhang 0006, Zhiping Jia, Hongchao Du, Runzhen Xue, Zhaoyan Shen, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | DAP-Sketch: An accurate and effective network measurement sketch with Deterministic Admission Policy
Rui Wang 0075, Hongchao Du, Zhaoyan Shen, Zhiping Jia |
Comput. Networks | 2 |
| 2020 | PattPIM: A Practical ReRAM-Based DNN Accelerator by Reusing Weight Pattern RepetitionsabstractWeight sparsity has been explored to achieve energy efficiency for Resistive Random-access Memory (ReRAM) based DNN accelerators. However, most existing ReRAM-based DNN accelerators are based on an overidealized crossbar architecture and mainly focus on compressing zero weights. In this paper, we propose a novel ReRAM-based accelerator — PattPIM, to achieve space compression and computation reuse by studying DNN weight patterns based on practical ReRAM crossbars. We first thoroughly analyze the weight distribution characteristics of several typical DNN models and observe many non-zero weight pattern repetitions (WPRs). Thus, in PattPIM, we propose a WPR-aware DNN engine and a WPR-to-OU mapping scheme to save both space and computation resources. Furthermore, we adopt an approximate weight pattern transform algorithm to improve the DNN WPRs ratio to enhance the reuse efficiency with negligible inference accuracy loss. Our evaluation with 6 DNN models shows that the proposed PattPIM delivers significant performance improvement, ReRAM resources efficiency and energy saving. Yuhao Zhang 0006, Zhiping Jia, Yungang Pan, Hongchao Du, Zhaoyan Shen, Mengying Zhao, Zili Shao |
DAC | 4 |
| 2019 | Accurate Network Flow Measurement with Deterministic Admission Policy
Hongchao Du, Rui Wang 0075, Zhaoyan Shen, Zhiping Jia |
ICA3PP (2) | 1 |