Yida Gu

dblp:391/4056 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0007-0712-8575ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum States
abstract
AI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schrödinger equation for complex many-body systems. Among neural network quantum state (NNQS) approaches, the NNQS-SCI (Selected Configuration Interaction) method stands out as a state-of-the-art technique, recognized for its high accuracy and scalability. However, its application to larger systems is severely constrained by a hybrid CPU-GPU architecture. Specifically, centralized CPU-based global de-duplication creates a severe scalability barrier due to communication bottlenecks, while host-resident coupled-configuration generation induces prohibitive computational overheads. We introduce QiankunNet-cuSCI, a fully GPU-accelerated SCI framework designed to overcome these bottlenecks. It first integrates a distributed, load-balanced global de-duplication algorithm to minimize redundancy and communication overhead at scale. To address compute limitations, it employs specialized, fine-grained CUDA kernels for exact coupled configuration generation. Finally, to break the single-GPU memory barrier exposed by this full acceleration, it incorporates a GPU memory-centric runtime featuring GPU-side pooling, streaming mini-batches, and overlapped offloading. This design enables much larger configuration spaces and shifts the bottleneck from host-side limitations back to on-device inference. Our evaluation demonstrates that our work fundamentally expands the scale of solvable problems. On an NVIDIA A100 cluster with 64 GPUs, our work achieves up to 2.32 × end-to-end speedup over the highly-optimized NNQS-SCI baseline while preserving the same chemical accuracy. Furthermore, it demonstrates excellent distributed performance, maintaining over 90% parallel efficiency in strong scaling tests.
Daran Sun, Bowen Kan, Haoquan Long, Hairui Zhao 0002, Haoxu Li, Ankang Feng, Wenjing Huang 0002, Yida Gu, Honghui Shang, Yunquan Zhang, Dingwen Tao, Ninghui Sun, Guangming Tan
HPDC10
2026 ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Jinwu Yang, Jiaan Wu, Xinyang Ma, Hairui Zhao 0002, Yida Gu, Yuanhong Huang, Wenjing Huang 0002, Yili Ma, Zhongzhe Hu, Shaoteng Liu, Jiaxun Lu, Guangming Tan, Dingwen Tao
ISCA6
2026 PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level Interpolation
abstract
With the exponential growth of computing power, large-scale scientific simulations are producing massive volumes of data, leading to critical storage and I/O challenges. Error-bounded lossy compression has become one of the most effective solutions for reducing data size while preserving accuracy. Meanwhile, to achieve high-performance compression on such large datasets, leveraging GPUs has become increasingly essential. GPU-based lossy compressors deliver strong performance, but typically support only single-precision decompression, limiting their ability to meet the diverse accuracy requirements of scientific workflows. Progressive compressors can address this limitation by enabling on-demand precision retrieval. However, existing progressive lossy compressors on GPU still suffer from low throughput. To overcome these challenges, we present PRISM, a GPU-based progressive lossy compressor that achieves both high throughput and multi-precision retrieval, which introduces a high performance progressive framework that integrates the multiple interpolation predictors, efficient bitplane extraction, and an enhanced lossless compression that combines sign-absolute coding with zero-aware parallel algorithms. Evaluations on representative real-world datasets from five scientific domains show that PRISM significantly outperforms state-of-the-art progressive compressors on GPU, reducing retrieval data volume by over 15.6× and achieving up to 20.1× higher throughput on the NVIDIA H100 GPU under the same error bounds.
Bing Lu 0001, Hairui Zhao 0002, Dejun Luo, Wenjing Huang 0002, Yida Gu, Jinyang Liu 0003, Guangming Tan, Dingwen Tao
PPoPP6
2026 CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training
abstract
As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes—substantially outperforming existing solutions.
Yida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang 0001, Hairui Zhao 0002, Wenjing Huang 0002, Jinwu Yang, Yueyuan Zhou, Qian Zhao 0021, Haoxu Li, Zhan Wang 0003, Guangming Tan, Dingwen Tao
PPoPP1
2026 KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
Xinyang Ma, Dejun Luo, Hairui Zhao 0002, Bing Lu 0001, Wenjing Huang 0002, Yida Gu, Jinyang Liu 0003, Dingwen Tao, Guangming Tan
SIGCOMM7
2026 CML-PowF: Data Clustering Matching Based Low-overhead Multiple CPU Real-time Power Forecasting
abstract
Efficient CPU power capping is essential for energy saving and fault tolerance in parallel computing clusters, but its effectiveness depends on accurate and timely processor power forecasting with minimal sampling overhead. Existing methods often struggle to balance these factors under scalability constraints, as hardware limitations tightly bound the available sampling resources. This article focuses on the issue of high-precision real-time processor power forecasting while maintaining (or minimally increasing) the total overhead of multiprocessor power forecasting, particularly when the parallelism scale ranges from P to 2P processors or when the problem size scales from M to 2M . We propose CML-PowF , a low-overhead multiprocessor real-time power forecasting approach based on data clustering. CML-PowF integrates two key algorithms: Alg-CEF , which conducts cluster matching on the runtime characteristics of the program at the P / M scale, and models the tradeoff among forecasting error, time span, and sampling overhead. Alg-MSF , which leverages execution patterns from smaller-scale runs to determine the optimal sampling overhead and forecasting time span at the 2P / 2M scale. We evaluate CML-PowF on x86 and ARM platforms with up to 32 computing nodes (2,048 cores). Results show that it achieves 3–6% forecasting error at large scales with only 0.2–0.5% degradation compared to the P / M scale, without increasing total sampling overhead. Integrated with the PowC control system, CML-PowF effectively maintains real-time processor power below target thresholds.
Rongyu Deng, Juan Chen 0001, Yuan Yuan 0034, Yong Dong, Aolin Cao, Yida Gu, Dingwen Tao
ACM Trans. Archit. Code Optim.8
2026 TSUE+: An Efficient Update Framework With Swift Recycling Mechanism for Erasure-Coded Cluster File Systems
abstract
Compared to replication-based storage systems, erasure-coded storage incurs significantly higher overhead during data updates. To address this issue, various parity logging methods have been proposed. Nevertheless, due to the long update path and substantial amount of random I/O involved in erasure code update processes, the resulting long latency and low through put often fail to meet the requirements of high performance applications. To address this challenge, we propose TSUE+, an efficient update framework with a swift recycling mechanism. TSUE+ divides the update process into two distinct stages: in the synchronous stage, data updates are stored in the format of replica data logs, eliminating random I/O by trading space for time; in the asynchronous stage, the recorded update logs are recycled and merged into original data and parity blocks, thereby reclaiming the storage overhead incurred in the synchronization phase. By converting random I/O operations into sequential ones based on data logs, TSUE+ effectively reduces update latency; furthermore, it significantly minimizes recycling overhead using a three-layer log structure and by leveraging the spatio-temporal locality of access patterns. We evaluated TSUE+ and other state of-the-art (SOTA) update mechanisms under diverse encoding schemes, using heterogeneous storage devices—including HDDs, SATA SSDs, NVMe SSDs, and PMEM—and multiple real-world and synthetic workloads: the MSR Cambridge trace, the Alibaba Cloud trace, the Tencent Cloud trace, and multiple synthetic worst-case workloads. Among all the platforms, TSUE+ has achieved significant performance improvements compared to other update methods, it also indicates that TSUE+ can be applied to various storage devices. Additionally, we provided percentile-based tail latency tests and update tests under the worst-case environment, which further demonstrated the ro bustness of TSUE+. Moreover, by enabling prompt log recycle and avoiding unnecessary overwrites and improving update granularity through locality-aware recycling, TSUE+ not only improves update performance but also mitigates write wear on SSD devices, thereby extending their operational lifespan.
Yida Gu, Wenjing Huang 0002, Yili Ma, Dong Dai 0001, Guangming Tan, Dingwen Tao
IEEE Trans. Parallel Distributed Syst.3
2025 TSUE: A Two-Stage Data Update Method for an Erasure Coded Cluster File System
abstract
Compared to replication-based storage systems, erasure-coded storage incurs significantly higher overhead during data updates. To address this issue, various parity logging methods have been proposed. Nevertheless, due to the long update path and substantial amount of random I/O involved in erasure code update processes, the resulting long latency and low throughput often fail to meet the requirements of high performance applications. To this end, we propose a two-stage data update method called TSUE. TSUE divides the update process into a synchronous stage that records updates in a data log, and an asynchronous stage that recycles the log in real-time. TSUE effectively reduces update latency by transforming random I/O into sequential I/O, and it significantly reduces recycle overhead by utilizing a three-layer log and the spatio-temporal locality of access patterns. In SSDs cluster, TSUE significantly improves update performance, achieving improvements of 7.6× under Ali-Cloud trace, 5× under Ten-Cloud trace, while it also extends the SSD's lifespan by up to 13× through reducing the frequencies of reads/writes and of erase operations.
Yida Gu, Wenjing Huang 0002, Dong Dai 0001, Guangming Tan, Dingwen Tao
HPDC3
2025 MANS: Efficient and Portable ANS Encoding for Multi-Byte Integer Data on CPUs and GPUs
abstract
Lossless compression is a classic technique for reducing data storage and transmission requirements. Asymmetric Numeral Systems (ANS) is a high-throughput, high-ratio lossless compression algorithm, but it lacks effective support for multi-byte data and cross-platform compatibility. To address this issue, we propose an Adaptive Data Mapping (ADM) scheme, which maps multi-byte integer data into single-byte space based on the data’s characteristics, improving the compression ratio of ANS while maintaining low encoding redundancy. We also optimize the ADM algorithm and the ANS encoder for GPU and CPU architectures, respectively, and combine them to create an efficient and portable ANS encoding method for multi-byte integer data, called MANS. Experimental results show that MANS improves compression ratios by an average of 1.24 ×, achieves 870.27MB/s throughput on CPUs, and delivers up to 288.45 × and 135.86 × speedups on an NVIDIA A100 and an AMD MI210 GPU compared to the CPU version—demonstrating its efficiency and portability across platforms.
Wenjing Huang 0002, Jinwu Yang, Shengquan Yin, Haoxu Li, Yida Gu, Xing Jing, Shiyuan Fu, Hao Hu 0015, Guangming Tan, Dingwen Tao
SC5
2025 DECEPTICON: a correlation-based strategy for RNA-seq deconvolution inspired by a variation of the Anna Karenina principle
abstract
Accurately deconvoluting cellular composition from bulk RNA-seq data is pivotal for understanding the tumor microenvironment and advancing precision medicine. Existing methods often struggle to consistently and accurately quantify cell types across heterogeneous RNA-seq datasets, particularly when ground truths are unavailable. In this study, we introduce DECEPTICON, a deconvolution strategy inspired by the Anna Karenina principle, which postulates that successful outcomes share common traits, while failures are more varied. DECEPTICON selects top-performing methods by leveraging correlations between different strategies and combines them dynamically to enhance performance. Our approach demonstrates superior accuracy in predicting cell-type proportions across multiple tumor datasets, improving correlation by 23.9% and reducing root mean square error by 73.5% compared to the best of 50 analyzed strategies. Applied to The Cancer Genome Atlas (TCGA) datasets for breast carcinoma, cervical squamous cell carcinoma, and lung adenocarcinoma, DECEPTICON-based predictions showed improved differentiation between patient prognoses. This correlation-based strategy offers a reliable, flexible tool for deconvoluting complex transcriptomic data and highlights its potential in refining prognostic assessments in oncology and advancing cancer biology.
Fulan Deng, Jiawei Zou, Miaochen Wang, Yida Gu, Lianchong Gao, Henry H. Y. Tong, Wantao Chen, Lianjiang Tan, Yaoqing Chu
Briefings Bioinform.4