Song Liu 0007

dblp:80/1141-7 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
abstract
The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quantization can effectively alleviate the memory pressure caused by KV Cache. However, existing methods either rely on static one-size-fits-all precision allocation or fail to dynamically prioritize critical KV in long-context tasks, forcing memory-accuracy-throughput tradeoffs. In this work, we propose a novel mixed-precision quantization method for KV Cache named KVmix. KVmix leverages gradient-based importance analysis to evaluate how individual Key and Value projection matrices affect the model loss, enabling layer-specific bit-width allocation for mix-precision quantization. It dynamically prioritizes higher precision for important layers while aggressively quantizing less influential ones, achieving a tunable balance between accuracy and efficiency. KVmix introduces a dynamic long-context optimization strategy that adaptively keeps full-precision KV pairs for recent pivotal tokens and compresses older ones, achieving high-quality sequence generation with low memory usage. Additionally, KVmix provides efficient low-bit quantization and CUDA kernels to optimize computational overhead. On LLMs such as Llama and Mistral, KVmix achieves near-lossless inference performance with extremely low quantization configuration (Key 2.19bit Value 2.38bit), while delivering a remarkable 4.9× memory compression and a 5.3× speedup in inference throughput.
Fei Li 0042, Song Liu 0007, Weiguo Wu, Shiqiang Nie, Jinyu Wang 0002
AAAI2
2026 EADA: Efficient adaptive data augmentation
Song Liu 0007, Weiguo Wu, Jinyu Wang 0002, Shiqiang Nie
Comput. Vis. Image Underst.1
2026 BLSA: A cache-aware balanced load scheduling approach on task graphs
Song Liu 0007, Fei Li 0042, Shiqiang Nie, Jinyu Wang 0002, Weiguo Wu
Future Gener. Comput. Syst.2
2026 FDSR: Efficient Model Training via Adaptive Tensor Quantization Based on Frequency Domain Division and Similarity Data Reuse
abstract
As deep neural networks (DNNs) continue to grow in scale and complexity, GPU memory limitations have become a significant challenge for DNN model training, especially on resource-constrained commercial GPUs. While model quantization facilitates memory-efficient training, it often necessitates a tradeoff between quantization granularity and model accuracy. And quantization imposes additional computational overhead, which adversely affects the training throughput and apportions out the performance gains it brings. In this article, we propose FDSR, an adaptive tensor quantization method that leverages frequency domain division and similarity-based data reuse to break the memory bottleneck in visual model training. FDSR leverages the frequency-domain characteristics of tensors in terms of memory consumption and model accuracy, and proposes a fine-grained tensor quantization with different quantization bit-widths. It adaptively optimizes the quantization parameters according to model accuracy during training while employing sparsification according to data frequency-domain features, minimizing memory consumption and accuracy loss. To counteract the computational cost, FDSR incorporates a novel similarity-based reuse strategy that avoids redundant quantization/dequantization computations, further enhanced by a tailored Locality-Sensitive Hashing (LSH) mechanism and optimized kernels. Experimental results demonstrate that FDSR achieves an average of 10.20× activation memory compression with only 1.10% average accuracy loss across various models on the commercial GPU. Compared to the state-of-the-art quantization methods, FDSR improves memory optimization by up to 68.6% and increases throughput by up to 25.55%, with consistent performance improvements on different GPU architectures.
Song Liu 0007, Fei Li 0042, Qin Xia, Shiqiang Nie, Jinyu Wang 0002, Weiguo Wu
ACM Trans. Archit. Code Optim.1
2025 ZeroCopy: file system assisted container buffer migration in cloud computing system
abstract
Abstract In cloud computing data centers, containerized tasks are regularly scheduled from one physical host to another due to resource management requirements such as handling machine failures, rebalancing server resources, and upgrading/scaling applications. After the container running in the source host is scheduled to the target host, it suffers from I/O performance degradation until the DRAM buffer is fully rebuilt. However, migrating the DRAM buffer from the source host to the target host could also introduce intolerable downtime of containerized tasks. Especially, as the DRAM buffer capacity of the application already increases to about dozens or hundreds of GB, the cost of downtime due to container migration becomes unacceptable. Many researchers have devoted themselves to developing an effective DRAM buffer warm-up scheme to avoid the cold bootstrap issue after container migration, such as pre-copy and post-copy schemes. However, the cold bootstrap and large-capacity buffer migration issues of container scheduling are still an open research problem. In this paper, motivated by the observation that the DRAM buffer is always flushed to the storage backend before starting the container in the target host, we proposed a scheme named ZeroCopy to utilize the file system to assist the DRAM buffer migration. ZeroCopy traverses the files in the DRAM buffer and flags these files when these files are flushed into the file system, and reloads these files into DRAM after starting the container in the target host. By this scheme, the container migration procedure does not require migrating data buffers and can start within an acceptable time. We conduct a series of experiments with public cloud traces to measure several key metrics on container migration. The results show that ZeroCopy outperforms these existing schemes. The average data transmission volume is reduced by about 6.25 times compared with state-of-the-art, and the downtime of container migration is also reduced by 31.8%.
Shiqiang Nie, Tingshen Ruan, Ruijia Chen, Song Liu 0007, Weiguo Wu
CCF Trans. High Perform. Comput.5
2025 Olsync: Object-level tiering and coordination in tiered storage systems based on software-defined network
Zhike Li, Shiqiang Nie, Jinyu Wang 0002, Chi Zhang 0095, Fangxing Yu, Zhankun Zhang, Song Liu 0007, Weiguo Wu
Future Gener. Comput. Syst.8
2025 Time-constrained persistent deletion for key-value store engine on ZNS SSD
Shiqiang Nie, Jie Niu, Qihan Hu, Song Liu 0007, Weiguo Wu
Future Gener. Comput. Syst.5
2025 Scalpel: High Performance Contention-Aware Task Co-Scheduling for Shared Cache Hierarchy
abstract
For scientific computing applications that consist of many loosely coupled tasks, efficient scheduling is critical to achieve high performance and good quality of service (QoS). One of the challenges for co-running tasks is the frequent contention for shared cache hierarchy of multi-core processors. Such contention significantly increases cache miss rate and therefore, results in performance deterioration for computational tasks. This paper presents Scalpel, a contention-aware task grouping and co-scheduling approach for efficient task scheduling on shared cache hierarchy. Scalpel utilizes the shared cache access features of tasks to group them in a heuristic way, which reduces the contention within groups by achieving equal shared cache locality, while maintaining load balancing between groups. Based thereon, it proposes a two-level scheduling strategy to schedule groups to processors and assign tasks to available cores in a timely manner, while considering the impact of task scheduling on shared cache locality to minimize task execution time. Experiments show that Scalpel reduces the shared cache miss rate by up to 2.14× and optimizes the execution time by up to 1.53× for scientific computing benchmarks, compared to several baseline approaches.
Song Liu 0007, Zengyuan Zhang, Xinhe Wan, Bo Zhao 0019, Weiguo Wu
IEEE Trans. Computers1
2025 Adaptive Read Level Recording for Read Performance Improvement in 3-D NAND Flash
abstract
While low-density parity-check code has been adopted in 3-D flash for improving chip reliability, it suffers from severe read latency due to the increasing number of read retries. Recent studies propose read-level recording to mitigate the performance loss from failed read retries. However, existing schemes induce large updating overhead and achieve suboptimal results, making it critical to develop better tradeoffs among storage overhead, process variation, and performance improvement. In this article, we propose AR$^{2}$, an adaptive read-level recording scheme to improve read performance for 3-D NOT AND (NAND) flash. It consists of two designs: AR$^{2}$-win and AR$^{2}$-pre. AR$^{2}$-win records the number of read levels that fit the majority of the last$N$reads, which prevents the worst page from dominating the read level recording. AR$^{2}$-pre predicts the number of read levels for the next read based on the recorded one and a simple machine learning model, which prevents using stale recorded levels in large-capacity solid-state drives (SSDs). Our experimental results show that AR$^{2}$significantly improves the read performance for 3-D NAND flash and achieves on average 15% or more read latency reduction over the state-of-the-art.
Shiqiang Nie, Zhike Li, Fangxing Yu, Song Liu 0007, Weiguo Wu
IEEE Trans. Reliab.4
2024 pommDNN: Performance optimal GPU memory management for deep neural network training
Weiduo Chen, Xiaoshe Dong, Xinhang Chen, Song Liu 0007, Qin Xia, Qiang Wang 0062
Future Gener. Comput. Syst.4
2023 An efficient computation offloading and resource allocation algorithm in RIS empowered MEC
Xiangjun Zhang, Weiguo Wu, Song Liu 0007, Jinyu Wang 0002
Comput. Commun.3
2023 TurboStencil: You only compute once for stencil computation
Song Liu 0007, Xinhe Wan, Zengyuan Zhang, Bo Zhao 0019, Weiguo Wu
Future Gener. Comput. Syst.1
2023 DHTS: A Dynamic Hybrid Tiling Strategy for Optimizing Stencil Computation on GPUs
abstract
Stencil computation is an important class of computational modes in scientific computing applications. Loop tiling techniques have been widely studied to accelerate stencil computations on different architectures by exploiting parallelism and data locality. Recent advanced tiling methods enable the tile-wise concurrent start-up to improve the execution performance. However, such methods statically partition all dimensions of iteration space into tiles with predetermined complex shapes and sizes, and thus lead to low thread utilization and memory access efficiency on GPUs. In this paper, we present DHTS, a novel dynamic hybrid tiling strategy for stencil computations. DHTS employs static tiling on the outer dimensions to achieve concurrent start-up parallelism, while proposes a dynamic rectangular tiling method on the inner dimensions to improve thread utilization and memory access efficiency. By deriving tile size constraints, DHTS adaptively achieves equal-size workload of tiles, and therefore reducing idle threads and increasing coalesced memory accesses within tiles. We implement the proposed strategy with different complex tile shapes. Experimental results on Titan V and Tesla V100 GPUs show that DHTS effectively improves the execution performance of 2D/3D stencils compared to state-of-the-art tiling methods, and achieves the best improvement of 28×.
Song Liu 0007, Zengyuan Zhang, Weiguo Wu
IEEE Trans. Computers1
2023 MCB: a multidevice cooperative buffer management strategy for boosting the write performance of the SSD-SMR hybrid storage
Chi Zhang 0095, Shiqiang Nie, Jinyu Wang 0002, Song Liu 0007, Weiguo Wu
J. Supercomput.4
2023 BiLSTM-based Federated Learning Computation Offloading and Resource Allocation Algorithm in MEC
abstract
Mobile edge computing (MEC) driven by 5G cellular systems has recently emerged as a promising paradigm, enabling mobile devices (MDs) with limited computing resources to offload various computation-intensive tasks (such as autopilot, online game) to edge servers to enhance the data processing capabilities of MDs. However, the uncertainty of wireless channel state and data volume of offloading tasks, as well as the data security privacy of offloading tasks, bring serious challenges to computation offloading in MEC. In this article, we consider a time-varying MEC scenario and formalize the delay and energy consumption during the computation offloading process as a joint optimization problem. Then the optimization problem is decomposed into two sub-problems: intelligent task prediction and resource allocation. Different from traditional methods, we improve the federated learning (FL) algorithm and propose a thoughtful cloud-edge-client FL task prediction mechanism based on Bidirectional Long Short-Term Memory. Each participating MD trains the model locally without uploading data to the server, and periodically aggregates the model in the edge and in the cloud. The algorithm both eliminates the need to solve complex optimization problems and ensures user privacy security. Finally, experimental results show that our proposed algorithm significantly outperforms other benchmark algorithms in energy efficiency.
Xiangjun Zhang, Weiguo Wu, Jinyu Wang 0002, Song Liu 0007
ACM Trans. Sens. Networks4
2019 Accelerating Lattice Boltzmann Method by Fully Exposing Vectorizable Loops
Bin Qu, Song Liu 0007, Jiajun Yuan, Weiguo Wu
ICA3PP (1)2
2019 Revisiting the Parallel Strategy for DOACROSS Loops
Song Liu 0007, Yuanzhen Cui, Nianjun Zou, Weiguo Wu
J. Comput. Sci. Technol.1
2018 A Dynamic Parallel Strategy for DOACROSS Loops
abstract
Many parallelization methods work on exposing the pipeline/wave-front parallelism of DOACROSS loops through loop transformations. However, these methods statically assign iterations to available threads for parallel execution, and thus causing the waste of computing resources in synchronization among threads, especially in a multithreading environment. This paper proposes a brand-new parallel strategy that achieves wave-front parallelism with reduced dependences and provides dynamic tile assignment for DOACROSS loops, which has better ability to avoid threads from waiting in synchronization and utilize computing resources. The experimental results demonstrate that the proposed strategy outperforms two advanced strategies which are based on implicit barriers and POST/WAIT operations over six benchmarks on a multi-core server. The strategy also has better scalability for the increasing number of threads.
Yuanzhen Cui, Song Liu 0007, Nianjun Zou, Weiguo Wu
HPC Asia2
2018 An efficient tile size selection model based on machine learning
Song Liu 0007, Yuanzhen Cui, Weiguo Wu
J. Parallel Distributed Comput.1
2017 An Efficient Locality-Aware Task Assignment Algorithm for Minimizing Shared Cache Contention
abstract
Task scheduling can improve the performance of parallel execution through optimizing the utilization of on-chip computing resources, and thus it has been widely studied. Most of the previous work uses data access locality to predict cache behaviors for task scheduling, but usually suffering accuracy and computational time complexity issues. This paper proposes an efficient task assignment algorithm to minimize the contention for shared caches on multi-core processors among parallel independent process level tasks. The proposed algorithm leverages the property of footprint to approximately estimate the locality parameter of parallel tasks, choosing the best grouping of tasks with minimum locality value in a quick way for task assignment. The calculation time is therefore significantly reduced and the algorithm complexity is O(nlog2n). Meanwhile, the algorithm accuracy is very high. On an Intel 8 cores dual-processor system, the experimental results show that the task assignment algorithm achieves over 99% of the actual optimal performance on average and outperforms the default Linux task scheduling method by an average of over 5% for two sets of different parallel tasks.
Song Liu 0007, Xiao Xie, Yuanzhen Cui, Weiguo Wu
PDCAT1