Shiqiang Nie

dblp:180/0989 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0003-2215-7159ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 8 first-author · 17 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
abstract
The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quantization can effectively alleviate the memory pressure caused by KV Cache. However, existing methods either rely on static one-size-fits-all precision allocation or fail to dynamically prioritize critical KV in long-context tasks, forcing memory-accuracy-throughput tradeoffs. In this work, we propose a novel mixed-precision quantization method for KV Cache named KVmix. KVmix leverages gradient-based importance analysis to evaluate how individual Key and Value projection matrices affect the model loss, enabling layer-specific bit-width allocation for mix-precision quantization. It dynamically prioritizes higher precision for important layers while aggressively quantizing less influential ones, achieving a tunable balance between accuracy and efficiency. KVmix introduces a dynamic long-context optimization strategy that adaptively keeps full-precision KV pairs for recent pivotal tokens and compresses older ones, achieving high-quality sequence generation with low memory usage. Additionally, KVmix provides efficient low-bit quantization and CUDA kernels to optimize computational overhead. On LLMs such as Llama and Mistral, KVmix achieves near-lossless inference performance with extremely low quantization configuration (Key 2.19bit Value 2.38bit), while delivering a remarkable 4.9× memory compression and a 5.3× speedup in inference throughput.
Fei Li 0042, Song Liu 0007, Weiguo Wu, Shiqiang Nie, Jinyu Wang 0002
AAAI4
2026 EADA: Efficient adaptive data augmentation
Song Liu 0007, Weiguo Wu, Jinyu Wang 0002, Shiqiang Nie
Comput. Vis. Image Underst.6
2026 BLSA: A cache-aware balanced load scheduling approach on task graphs
Song Liu 0007, Fei Li 0042, Shiqiang Nie, Jinyu Wang 0002, Weiguo Wu
Future Gener. Comput. Syst.4
2026 FTL optimization for 3D NAND flash considering layer RBER variation and data precision
Shiqiang Nie, Ping She, Weiguo Wu
Future Gener. Comput. Syst.1
2026 FDSR: Efficient Model Training via Adaptive Tensor Quantization Based on Frequency Domain Division and Similarity Data Reuse
abstract
As deep neural networks (DNNs) continue to grow in scale and complexity, GPU memory limitations have become a significant challenge for DNN model training, especially on resource-constrained commercial GPUs. While model quantization facilitates memory-efficient training, it often necessitates a tradeoff between quantization granularity and model accuracy. And quantization imposes additional computational overhead, which adversely affects the training throughput and apportions out the performance gains it brings. In this article, we propose FDSR, an adaptive tensor quantization method that leverages frequency domain division and similarity-based data reuse to break the memory bottleneck in visual model training. FDSR leverages the frequency-domain characteristics of tensors in terms of memory consumption and model accuracy, and proposes a fine-grained tensor quantization with different quantization bit-widths. It adaptively optimizes the quantization parameters according to model accuracy during training while employing sparsification according to data frequency-domain features, minimizing memory consumption and accuracy loss. To counteract the computational cost, FDSR incorporates a novel similarity-based reuse strategy that avoids redundant quantization/dequantization computations, further enhanced by a tailored Locality-Sensitive Hashing (LSH) mechanism and optimized kernels. Experimental results demonstrate that FDSR achieves an average of 10.20× activation memory compression with only 1.10% average accuracy loss across various models on the commercial GPU. Compared to the state-of-the-art quantization methods, FDSR improves memory optimization by up to 68.6% and increases throughput by up to 25.55%, with consistent performance improvements on different GPU architectures.
Song Liu 0007, Fei Li 0042, Qin Xia, Shiqiang Nie, Jinyu Wang 0002, Weiguo Wu
ACM Trans. Archit. Code Optim.5
2026 WOM-FTL: An Efficient FTL for High-Density Flash Memory Through WOM-v Codes
abstract
High-density NAND flash memory, such as quadruple-level cell (QLC) flash, has has been widely adopted in emerging storage systems. However, its limited endurance and performance challenges necessitate novel solutions. Voltage-based write-once memory (WOM-v) codes have demonstrated their effectiveness in extending flash memory lifespan by reducing the erase count of flash blocks. Concurrently, secure deletion is essential to ensure data privacy in flash-based storage systems. Existing secure deletion approaches—encryption-based, erasure-based, and scrubbing-based—often face limitations such as susceptibility to deciphering or significant performance overheads. Additionally, the inherent “big block problem” in high-density flash memory complicates garbage collection (GC), further degrading system performance. To address these challenges, this paper proposes WOM-FTL, a flash translation layer (FTL) that integrates secure deletion and garbage collection (GC) with WOM-v codes to enhance both security and performance. WOM-FTL classifies request data into four categories based on access frequency and privacy requirements: hot-secure (HS), cold-secure (CS), hot-unsecure (HU) and cold-unsecure (CU). Additionally, WOM-FTL divides each block into several equal-sized sub-blocks and further classifies them into top sub-blocks and bottom sub-blocks according to their data storage characteristics. When a secure data deletion command is issued to the storage device, WOM-FTL leverages unsecure data (HU and CU) to overwrite the secure data (HS and CS). Furthermore, WOM-FTL allocates different types of request data to the corresponding sub-blocks, creating a data allocation pattern that is both scrubbing-friendly and GC-friendly. Experimental results demonstrate that WOM-FTL improves the I/O performance of storage systems by 60.91% compared to state-of-the-art solutions, providing a significant advancement in secure and efficient management of high-density flash memory.
Jinhua Cui 0001, Canghao Wen, Shiqiang Nie, Debin Liu, Yaliang Zhao, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 IFFS: An Interlaced Magnetic Recording Friendly File System
abstract
Recently, the emerging Interlaced Magnetic Recording (IMR) technology has substantially improved the areal density of disks by implementing interlaced track layout. While this track layout enhances disk storage capacity, it impairs the flexibility of write positioning. To maintain stable write performance of IMR disks, operations such as Read-Modify-Write (RMW) or Garbage Collection (GC) must be introduced, which inevitably incur extra I/Os. Especially in write-intensive workloads, the excessive additional I/Os exacerbate the write amplification effect, thereby severely degrading the overall performance of IMR disks. Although existing data management strategies strive to reduce extra I/Os via device drivers or system middleware, the semantic disparity between the disk and file system inherently limits these strategies to achieve optimal performance. To address the aforementioned challenge,this paper proposes IMR-Friendly File System (IFFS), an innovative file system tailored for IMR disks. First, we propose a semantic-aware hotness identification algorithm based on Online K-means, which redefines the data hotness metric by exploiting file system semantics to reduce data migration induced by inaccurate data classification. Second, we introduce an I/O-overhead-minimized data placement strategy that adaptively selects between in-place and out-of-place writing modes based on data hotness metrics. Furthermore, this strategy employs a log-transition mechanism to dynamically adjust write positions, effectively mitigating I/O overhead caused by RMWs and scattered read requests. Finally, we implement a multi-factor garbage collection mechanism that incorporates intra-zone data hotness, data layout, fragmentation levels, and other contextual attributes to optimize file system data management efficiency, thereby enhances the read performance on IMR disk-based storage system performance. Experimental results show that IFFS achieves an average bandwidth improvement of 12.96× over EXT4, XFS, F2FS, and the state-of-the-art work in Fio evaluation. Under YCSB workloads, IFFS improves bandwidth by an average of 56.89% while reducing latency by 34.46%.
Fangxing Yu, Chi Zhang 0095, Shiqiang Nie, Zhike Li, Weiguo Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 SecureIMR: A Plausibly Deniable Storage System for Interlaced Magnetic Recording
Jinhua Cui 0001, Canghao Wen, Shiqiang Nie
ICA3PP (1)4
2025 ZeroCopy: file system assisted container buffer migration in cloud computing system
abstract
Abstract In cloud computing data centers, containerized tasks are regularly scheduled from one physical host to another due to resource management requirements such as handling machine failures, rebalancing server resources, and upgrading/scaling applications. After the container running in the source host is scheduled to the target host, it suffers from I/O performance degradation until the DRAM buffer is fully rebuilt. However, migrating the DRAM buffer from the source host to the target host could also introduce intolerable downtime of containerized tasks. Especially, as the DRAM buffer capacity of the application already increases to about dozens or hundreds of GB, the cost of downtime due to container migration becomes unacceptable. Many researchers have devoted themselves to developing an effective DRAM buffer warm-up scheme to avoid the cold bootstrap issue after container migration, such as pre-copy and post-copy schemes. However, the cold bootstrap and large-capacity buffer migration issues of container scheduling are still an open research problem. In this paper, motivated by the observation that the DRAM buffer is always flushed to the storage backend before starting the container in the target host, we proposed a scheme named ZeroCopy to utilize the file system to assist the DRAM buffer migration. ZeroCopy traverses the files in the DRAM buffer and flags these files when these files are flushed into the file system, and reloads these files into DRAM after starting the container in the target host. By this scheme, the container migration procedure does not require migrating data buffers and can start within an acceptable time. We conduct a series of experiments with public cloud traces to measure several key metrics on container migration. The results show that ZeroCopy outperforms these existing schemes. The average data transmission volume is reduced by about 6.25 times compared with state-of-the-art, and the downtime of container migration is also reduced by 31.8%.
Shiqiang Nie, Tingshen Ruan, Ruijia Chen, Song Liu 0007, Weiguo Wu
CCF Trans. High Perform. Comput.1
2025 Olsync: Object-level tiering and coordination in tiered storage systems based on software-defined network
Zhike Li, Shiqiang Nie, Jinyu Wang 0002, Chi Zhang 0095, Fangxing Yu, Zhankun Zhang, Song Liu 0007, Weiguo Wu
Future Gener. Comput. Syst.3
2025 Time-constrained persistent deletion for key-value store engine on ZNS SSD
Shiqiang Nie, Jie Niu, Qihan Hu, Song Liu 0007, Weiguo Wu
Future Gener. Comput. Syst.1
2025 BullyDetect: Detecting School Physical Bullying With Wi-Fi and Deep Wavelet Transformer
abstract
More than 246 million children and adolescents suffer from school violence and bullying, e.g., verbal harassment, social harassment, and physical bullying, every year, according to a report from the United Nations Educational, Scientific and Cultural Organization. School violence and bullying severely harm the physical and emotional well-being of the victims, increasing the risks of depression, anxiety, sleep difficulties, lower academic achievement, dropping out of school, and even suicide attempts. Since school physical bullying always happens in the low-visibility areas spots of surveillance cameras, in this article, we propose to utilize Wi-Fi, a widely deployed infrastructure, to detect school physical bullying. We design residual wavelet transformer networks to conduct noise removal and action feature learning in an end-to-end manner. Besides, we propose two data augmentation methods in the temporal domain of Wi-Fi signals to simulate the different speeds and extents of bullying actions performed. Extensive evaluation of 20-paired volunteers demonstrates that 1) Wi-Fi can effectively detect physical school bullying; 2) the proposed approaches outperform long-short-time-memory networks, ResNet-1D, vision transformer, etc.; and 3) the proposed data augmentation methods can work as plug-and-play modules to improve the detection accuracy of all the above-mentioned approaches.
Fei Wang 0037, Lekun Xia, Fan Nai, Shiqiang Nie, Han Ding 0002, Jinsong Han
IEEE Internet Things J.5
2025 ZoomDB: Building cost-effective key-value store engine on ZNS SSD and SMR HDD
Shiqiang Nie, Chi Zhang 0095, Fangxing Yu, Yaming Li, Weiguo Wu
J. Syst. Archit.1
2025 DTB+: An enhanced data management strategy for efficient RMW reduction in IMR drives
abstract
The emerging Interlaced Magnetic Recording (IMR) technology not only achieves higher storage density than SMR, but also significantly reduces rewrite overhead by dividing tracks into bottom and top tracks and organizing them in an interlaced fashion. However, frequent updates to the bottom track can trigger a large number of Read-Modify-Write (RMW) operations during high disk space utilization, which can severely degrade the I/O performance. Addressing this issue, this paper proposes an interlaced translation layer named DTB+ to improve the write performance of IMR disks. Firstly, a workload-sensitive track heat analysis mechanism is introduced to intelligently place data to reduce track rewrite probability. Simultaneously, the zero-incremental cost region is selectively used to construct a twin-buffer architecture to reduce RMW operations. In addition, an adaptive space allocation engine based on reinforcement learning was developed to flexibly allocate and reclaim space within the twin-buffer, improving disk resource utilization . Finally, establish a flexible evicted-data transfer zone to delay the writeback operations of interference data, further reducing the additional overhead. Experimental results indicate that compared with the state-of-the-art studies, DTB+ can reduce RMWs by 63.00% and additional I/O operations by 57.41%, decrease the average write latency by 37.77%, and lower the tail latency by 53.95%.
Fangxing Yu, Chi Zhang 0095, Zhike Li, Shiqiang Nie, Weiguo Wu
J. Syst. Archit.5
2025 DPUSwap: building an infinite swap with DPU for cloud computing system
Shiqiang Nie, Jianqiang Ma, Jiaxin Shi, Weiguo Wu
J. Supercomput.1
2025 Constructing a scalable key-value store engine on multidisk system
Shiqiang Nie, Jie Niu, Fangxing Yu, Jianqiang Ma, Xingxing Zhu, Weiguo Wu
J. Supercomput.1
2025 Adaptive Read Level Recording for Read Performance Improvement in 3-D NAND Flash
abstract
While low-density parity-check code has been adopted in 3-D flash for improving chip reliability, it suffers from severe read latency due to the increasing number of read retries. Recent studies propose read-level recording to mitigate the performance loss from failed read retries. However, existing schemes induce large updating overhead and achieve suboptimal results, making it critical to develop better tradeoffs among storage overhead, process variation, and performance improvement. In this article, we propose AR$^{2}$, an adaptive read-level recording scheme to improve read performance for 3-D NOT AND (NAND) flash. It consists of two designs: AR$^{2}$-win and AR$^{2}$-pre. AR$^{2}$-win records the number of read levels that fit the majority of the last$N$reads, which prevents the worst page from dominating the read level recording. AR$^{2}$-pre predicts the number of read levels for the next read based on the recorded one and a simple machine learning model, which prevents using stale recorded levels in large-capacity solid-state drives (SSDs). Our experimental results show that AR$^{2}$significantly improves the read performance for 3-D NAND flash and achieves on average 15% or more read latency reduction over the state-of-the-art.
Shiqiang Nie, Zhike Li, Fangxing Yu, Song Liu 0007, Weiguo Wu
IEEE Trans. Reliab.1
2024 DTB: A Novel Reinforcement Learning-Assisted Data Management Strategy in Interlaced Magnetic Recording
abstract
Shingled Magnetic Recording (SMR) technology, employing a shingled track layout, has significantly enhanced areal density capability. However, this layout imposes severe write penalties when dealing with non-sequential writes. The emerging Interlaced Magnetic Recording (IMR) technology not only achieves higher storage density than SMR, but also significantly reduces rewrite overhead by dividing all tracks into bottom and top tracks and organizing them in an interlaced fashion. However, frequent updates to the bottom track can trigger a large number of Read-Modify-Write (RMW) operations during high disk space utilization, which can severely affect the I/O performance of the disk. Addressing this issue, this paper proposes an interlaced translation layer named DTB to improve the write performance of IMR disks. Firstly, a workload-sensitive track heat analysis mechanism is introduced to intelligently place data to reduce track rewrite probability. Simultaneously, the zero-incremental cost region is selectively used to construct a Twin-Buffer architecture to reduce RMW operations triggered by frequent writeback of hot data, thereby effectively curtailing the rewriting overhead. In addition, an adaptive space allocation engine was developed by analyzing the data characteristics in the buffer, and we designed a dynamic configuration model based on reinforcement learning to flexibly allocate and reclaim space within the Twin-Buffer, improving disk resource utilization and I/O performance. Experimental results indicate that DTB can reduce the number of RMWs by 63.45%, decrease the average write latency by 44.37%, and lower the tail latency by 56.84% compared with state-of-the-art studies.
Fangxing Yu, Chi Zhang 0095, Zhike Li, Shiqiang Nie, Weiguo Wu
HPCC5
2024 SmartNetSSD: Exploiting Path Resources for Read Performance Improvement in Network-Based SSDs
abstract
With the bit density improvement and the three-dimensional NAND flash techniques, solid-state drives (SSDs) dramatically increase the storage capacity and performance. However, the incorporation of multiple flash chips within a single flash channel structure in SSDs introduces the access path conflicts when servicing multiple I/O requests accessing flash chips on the same channel. To meet the increasing performance demands of modern applications, network SSDs that employ the interconnection network of flash chips, has a high potential to address the access path conflicts by fundamentally increasing the number of the available access paths. In this paper, we propose SmartNetSSD, a new scheme that utilizes the path diversity in network SSD to dramatically mitigate access path conflicts. SmartNetSSD employs the following key techniques: 1) A multi-path routing algorithm, MProuting, identifies the multiple access paths for read operations. 2) A network-based read-retry technique, NetRR, pipelines the con-secutive read-retry steps for a read operation across the multiple access paths. The experimental results show that SmartNetSSD improves I/O performance by an average of 46.29% over the state-of-the-art approach.
Jinhua Cui 0001, Shiqiang Nie, Laurence T. Yang
ICCD4
2024 DIR: Dynamic Request Interleaving for Improving the Read Performance of Aged Solid-State Drives
Shiqiang Nie, Weiguo Wu
J. Comput. Sci. Technol.1
2023 MCB: a multidevice cooperative buffer management strategy for boosting the write performance of the SSD-SMR hybrid storage
Chi Zhang 0095, Shiqiang Nie, Jinyu Wang 0002, Song Liu 0007, Weiguo Wu
J. Supercomput.2
2021 Data Pattern Aware Reliability Enhancement Scheme for 3D Solid-State Drives
abstract
3D charge-trap (CT) NAND flash-based SSD has been used widely for its large capacity, low cost per bit, and high endurance. One-shot program (OSP) scheme, as a variation of incremental step pulse programming (ISPP) scheme, has been employed to program data for CT flash, whose program unit is the Word-Line (WL) instead of the page. The existing program optimization schemes either make trade-offs among program latency and reliability by adjusting the program step voltage on demand; or remap the most error-prone cell states to others by re-encoding programmed data. However, the data pattern, which represents the ratio of 1s in data values, has not been thoroughly studied. In this paper, we observe that most small files do not contain uniform 1s and 0s among these common file types (i.e., image, audio, text, executable file), leading to programming WL cells in different states unevenly. Some cell states dominate over the WL, while others are not. Based on this observation, we propose a flexible reliability enhancement scheme based on the OSP scheme. This scheme programs the cells into different states with varied , i.e., these cells in one state, whose number is the largest in one WL, are programmed with a fine-grained (namely slow write). In contrast, the minority are programmed with a coarse-grained (namely fast write). So the reliability is improved due to averaging the major enhanced cells with the minor degraded cells without program latency overhead. A series of experiments have been conducted, and the results indicate that the proposed scheme achieves 34% read performance improvement and 16% lifetime elongation on average.
Shiqiang Nie, Weiguo Wu, Chi Zhang 0095
ACM Trans. Embed. Comput. Syst.1
2020 Layer RBER Variation Aware Read Performance Optimization for 3D Flash Memories
abstract
3D NAND flash enables the construction of large capacity Solid-State Drives (SSDs) for modern computer systems. While effectively reducing per bit cost, 3D NAND flash exhibits non-negligible process variations and thus RBER (raw bit error rate) difference across layers, which leads to sub-optimal read performance for applications with either small or large I/O requests. In this paper, we propose LRR, Layer RBER variation aware Read optimization schemes, to address the challenge. LRR consists of two schemes - LRR subpage read scheduling (SRS) and LRR fullpage allocation (FPA). SRS groups small read requests from the layers with similar RBERs to reduce the average read latency of subpage sized read requests. FPA distributes the data of a large write to multiple layers, which improves the read latency when reading from layers with large RBERs. Our experimental results show that our proposed scheme LRR reduces 46% read latency on average over the state-of-the-art.
Shiqiang Nie, Youtao Zhang, Weiguo Wu, Jun Yang 0002
DAC1
2016 VIOS: A Variation-Aware I/O Scheduler for Flash-Based Storage Systems
Jinhua Cui 0001, Weiguo Wu, Shiqiang Nie, Jianhang Huang, Zhuang Hu, Nianjun Zou, Yinfeng Wang
NPC3
2016 Exploiting Cross-Layer Hotness Identification to Improve Flash Memory System Performance
Jinhua Cui 0001, Weiguo Wu, Shiqiang Nie, Jianhang Huang, Zhuang Hu, Nianjun Zou, Yinfeng Wang
NPC3