VLDB 2026 Research / reviewers in the wild / expert
Cheng Ji 0002
dblp:32/598-2
· DBLP profile ↗
39ranked-venue papers
9as first author
20since 2021 · last 2025
0000-0002-2525-8070ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MedFS: Pursuing Low Update Overhead via Metadata-Enabled Delta Compression for Log-structured File System on Mobile Device
Chao Wu 0006, Cheng Ji 0002, Li-Pin Chang, Zongwei Zhu, Congming Gao, Weichao Guo, Yanzhi Wang 0001 |
FAST | 2 |
| 2025 | PQC-LLM: Post-Quantization Delta Compression for LLMsabstractThe massive parameter scale and substantial storage requirements of large language models (LLMs) hinder their deployment on edge devices and in real-world applications, even after quantization. To address this challenge, we propose PQC-LLM, a Post-Quantization delta Compression scheme for LLMs. PQC-LLM is motivated by two key observations: (1) quantized LLMs still exhibit redundancy on the least significant bits (LSBs); (2) State-of-the-art (SOTA) quantization techniques could only converge to a local optimum. To exploit the potential, PQC-LLM performs a tile-wise deterministic bit replacement by uniformly substituting the least significant$N$bits of all weights within the same tile with the constant bit value. We employ a network architecture search (NAS) mechanism to determine the optimal bit length and value of the uniform replacement. Subsequently, we perform delta compression on each tile using the first weight as the base value. Extensive experiments show that, PQC-LLM effectively reduces the storage requirements of quantized LLMs with negligible lossy or improved task accuracy. Yujin Zhong, Chao Wu 0006, Cheng Ji 0002 |
ICPADS | 3 |
| 2025 | Magnifier: A Chiplet Feature-Aware Test Case Generation Method for Deep Learning AcceleratorsabstractThe development of deep learning has led to increasing demands for computation and memory, making multi-chiplet accelerators a powerful solution. Multi-chiplet accelerators require more precise consideration of hardware configurations and mapping schemes in terms of computation, memory, and communication patterns compared to monolithic designs, in order to avoid underutilization of performance. However, there is currently a lack of performance testing methods specifically tailored for multi-chiplet accelerators. Existing testing methods primarily focus on correctness testing and do not address potential performance issues from a hardware perspective. To address these issues, this paper proposes Magnifier: a test case generation method for performance testing of multi-chiplet accelerators. Firstly, we analyze typical multi-chiplet accelerator prototype from the perspectives of computation, memory, and communication patterns, and summarize a chiplet feature-aware operator task set. Next, we define the test evaluation metric IPPstd and use a candidate operator set to construct a sampling space for model-level test cases. Finally, we build a GAN to learn the distribution of high-diversity test cases, enabling the rapid generation of high-quality test cases. We validate the proposed method on both simulated and real multi-chiplet accelerators. Experiments show that Magnifier can improve the metric of test cases by up to 3.42 times and significantly reduce generation time, providing valuable insights for optimizing the hardware and software of multi-chiplet accelerators. Boyu Li 0006, Zongwei Zhu, Weihong Liu, Qianyue Cao, Changlong Li 0006, Cheng Ji 0002, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | HaloFL: Efficient Heterogeneity-Aware Federated Learning Through Optimal Submodel Extraction and Dynamic Sparse AdjustmentabstractFederated learning (FL) is an advanced framework that enables collaborative training of machine learning models across edge devices. An effective strategy to enhance training efficiency is to allocate the optimal submodel based on each device’s resource capabilities. However, system heterogeneity significantly increases the difficulty of allocating submodel parameter budgets appropriately for each device, leading to the straggler problem. Meanwhile, data heterogeneity complicates the selection of the optimal submodel structure for specific devices, thereby impacting training performance. Furthermore, the dynamic nature of edge environments, such as fluctuations in network communication and computational resources, exacerbates these challenges, making it even more difficult to precisely extract appropriately sized and structured submodels from the global model. To address the challenges in heterogeneous training environments, we propose an efficient FL framework, namely, HaloFL. The framework dynamically adjusts the structure and parameter budget of submodels during training by evaluating three dimensions: 1) model-wise performance; 2) layer-wise performance; and 3) unit-wise performance. First, we design a data-aware model unit importance evaluation method to determine the optimal submodel structure for different data distributions. Next, using this evaluation method, we analyze the importance of model layers and reallocate parameters from noncritical layers to critical layers within a fixed parameter budget, further optimizing the submodel structure. Finally, we introduce a resource-aware dual-UCB multiarmed bandit agent, which dynamically adjusts the total parameter budget of submodels according to changes in the training environment, allowing the framework to better adapt to the performance differences of heterogeneous devices. Experimental results demonstrate that HaloFL exhibits outstanding efficiency in various dynamic and heterogeneous scenarios, achieving up to a 14.80% improvement in accuracy and a$3.06\times $speedup compared to existing FL frameworks. Zirui Lian, Qianyue Cao, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | Extending SSD Lifetime via Balancing Layer Endurance in 3D NAND Flash MemoryabstractBy stacking layers vertically, 3D flash memory enables continuous growth in capacity. In this paper, we study the layer variation in 3D flash blocks and find that bottom layer pages exhibit the lowest endurance, whereas middle layer pages demonstrate the highest endurance. The imbalanced endurance across different layers will diminish the overall SSD lifetime. To address this issue, we introduce a novel layer-aware write strategy, named LA-Write. It performs write-skip operations with layer-specific probabilities. The endurance of bottom layer pages with the highest probability would be noticeably improved, which could balance the layer endurance. Experiment results show that LA-Write can improve SSD lifetime by 29%. Yajuan Du, Cheng Ji 0002 |
DATE | 4 |
| 2024 | In-place Switch: Reprogramming based SLC Cache Design for Hybrid 3D SSDsabstractTo increase SSD capacity, high bit-density cells, such as Triple-Level Cell (TLC), are utilized within 3D SSDs. However, due to the inferior performance of TLC, SLC/TLC hybrid 3D SSD is designed to use a portion of the TLC space as an SLC cache to achieve high SSD performance by writing host data at the SLC speed. However, our preliminary studies indicate that the SLC cache can lead to a performance cliff if filled rapidly and cause significant write amplification when data migration occurs during idle times. In this work, we propose leveraging a reprogram operation to address these challenges. Specifically, when the SLC cache is full or during idle periods, a reprogram operation is performed to switch used SLC pages to TLC pages in place (termed In-place Switch, IPS). Subsequently, other free TLC space is allocated as the new SLC cache. IPS can continuously provide sufficient SLC cache within SSDs, significantly improving write performance and reducing write amplification. Experimental results demonstrate that IPS can reduce write latency and write amplification by up to 0.75 times and 0.53 times, respectively, compared to state-of-the-art SLC cache technologies. Xufeng Yang, Jiancong Zheng, Cheng Ji 0002, Congming Gao |
NAS | 3 |
| 2024 | Transformer with a Parallel Decoder for Image CaptioningabstractIn this paper, a parallel decoder and a word group prediction module are proposed to speed up decoding and improve the effect of captions. The features of the image extracted by the encoder are linearly projected to different word groups, and then a unique relaxed mask matrix is designed to improve the decoding speed and the caption effect. First, since image captioning is composed of many words, sentences can also be broken down into word groups or words according to their syntactic structure, and we achieve this function through constituency parsing. Second, we make full use of the extracted features to predict the size of word groups. Then, a new embedding representing the information of the word is proposed based on word embedding. Finally, with the help of word groups, we design a mask matrix to modify the decoding process so that each step of the model can produce one or more words in parallel. Experiments on public datasets demonstrate that our method can reduce the time complexity while maintaining competitive performance. Peilang Wei, Xu Liu 0006, Jun Luo 0006, Huayan Pu, Xiaoxu Huang, Shilong Wang 0001, Huajun Cao, Shouhong Yang, Xu Zhuang, Hong Yue, Cheng Ji 0002, Mingliang Zhou 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 12 |
| 2024 | FedStar: Efficient Federated Learning on Heterogeneous Communication NetworksabstractThe proliferation of multi-media applications and increased computing power of mobile devices have led to the development of personalized artificial intelligent (AI) applications that utilize the massive user-information residing on them. However, the traditional centralized training paradigm is not applicable in this scenario due to potential privacy risks and high communication overhead. Federated learning (FL) provides an option to these applications. Nevertheless, the heterogeneity of computing and communication latency among devices have posed great challenges to building efficient learning frameworks. Existing optimizations on FL either fail to speed up training on heterogeneous devices or suffer from poor communication efficiency. In this paper, we propose FedStar, an efficient FL framework that supports decentralized asynchronous training on heterogeneous communication networks. Considering the heterogeneous computing power in the network, FedStar supports running heterogeneity-aware local steps on each device. What’s more, considering the heterogeneous communication latency and possibly unreachable communication path between some devices, FedStar generates a decentralized communication topology that can achieve maximal training throughput. Finally, it adopts weighted aggregation to guarantee high convergence accuracy of global model. Theoretical analysis results show the convergence behaviour of FedStar under non-convex settings. Experimental results show that FedStar can achieve a speedup of 4.81× than the state-of-the-art FL schemes with high convergence accuracy. Qianyue Cao, Yongchun Zheng, Zongwei Zhu, Cheng Ji 0002, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Ace-Sniper: Cloud-Edge Collaborative Scheduling Framework With DNN Inference Latency Modeling on Heterogeneous DevicesabstractThe cloud–edge collaborative inference requires efficient scheduling of artificial intelligence (AI) tasks to the appropriate edge intelligence devices. Gls DNN inference latency has become a vital basis for improving scheduling efficiency. However, edge devices exhibit highly heterogeneous due to the differences in hardware architectures, computing power, etc. Meanwhile, the diverse deep neural networks (DNNs) are continuing to iterate over time. The diversity of devices and DNNs introduces high computational costs for measurement methods, while invasive prediction methods face significant development efforts and application limitations. In this article, we propose and develop Ace-Sniper, a scheduling framework with DNN inference latency modeling on heterogeneous devices. First, to address the device heterogeneity, a unified hardware resource modeling (HRM) is designed by considering the platforms as black-box functions that output feature vectors. Second, neural network similarity (NNS) is introduced for feature extraction of diverse and frequently iterated DNNs. Finally, with the results of HRM and NNS as input, the performance characterization network is designed to predict the latencies of the given unseen DNNs on heterogeneous devices, which can be combined into most time-based scheduling algorithms. Experimental results show that the average relative error of DNN inference latency prediction is 11.11%, and the prediction accuracy reaches 93.2%. Compared with the nontime-aware scheduling methods, the average waiting time for tasks is reduced by 82.95%, and the platform throughput is improved by 63% on average. Weihong Liu, Jiawei Geng, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Zirui Lian, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Heter-Train: A Distributed Training Framework Based on Semi-Asynchronous Parallel Mechanism for Heterogeneous Intelligent Transportation SystemsabstractTransportation big data (TBD) are increasingly combined with artificial intelligence to mine novel patterns and information due to the powerful representational capabilities of deep neural networks (DNNs), especially for anti-COVID19 applications. The distributed cloud-edge-vehicle training architecture has been applied to accelerate DNNs training while ensuring low latency and high privacy for TBD processing. However, multiple intelligent devices (e.g., intelligent vehicles, edge computing chips at base stations) and different networks in intelligent transportation systems lead to computing power and communication heterogeneity among distributed nodes. Existing parallel training mechanisms perform poorly on heterogeneous cloud-edge-vehicle clusters. The synchronous parallel mechanism may force fast workers to wait for the slowest worker for synchronization, thus wasting their computing power. The asynchronous mechanism has communication bottlenecks and can exacerbate the straggler problem, causing increased training iterations and even incorrect convergence. In this paper, we introduce a distributed training framework, Heter-Train. First, a communication-efficient semi-asynchronous parallel mechanism (SAP-SGD) is proposed, which can take full advantage of acceleration effect of asynchronous strategy on heterogeneous training and constrain the straggler problem by using global interval synchronization. Second, Considering the difference in node bandwidth, we design a solution for heterogeneous communication. Moreover, a novel weighted aggregation strategy is proposed to aggregate the model parameters with different versions. Finally, experimental results show that our proposed strategy can achieve up to$6.74 \times $speedups on training time, with almost no accuracy decrease. Jiawei Geng, Haipeng Jia, Zongwei Zhu, Hai Fang, Chengxi Gao, Cheng Ji 0002, Gangyong Jia, Guangjie Han, Xuehai Zhou |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | RLAlloc: A Deep Reinforcement Learning-Assisted Resource Allocation Framework for Enhanced Both I/O Throughput and QoS Performance of Multi-Streamed SSDsabstractMulti-streamed Solid-State Disks (SSDs) have attracted increasing adoption in modern flash storage devices. Despite their excellent promise, effective flash resource allocation is still limiting both their achievable I/O performance and practical implementation. To this end, we develop the first-of-its-kind framework dubbed RLAlloc, which for the first time demonstrates deep Reinforcement Learning-assisted resource Allocation for boosting both I/O throughput and QoS performance of multi-streamed SSDs. Extensive experiments consistently validate the effectiveness of RLAlloc, improving up to 39.9% on I/O throughput and 44.0% on QoS performance over the state-of-the-art competitors. Mengquan Li, Chao Wu 0006, Congming Gao, Cheng Ji 0002, Kenli Li 0001 |
DAC | 4 |
| 2023 | Transparent File Deduplication with Reduced Update Cost on Encryption Enabled Mobile DevicesabstractData deduplication has been long studied to achieve data reduction. However, deploying deduplication on encryption enabled mobile systems might consume much memory footprint and computation time for hash calculations. Moreover, frequent file updates on deduplicated files could badly degrade the deduplication efficacy due to the increased file-system metadata penalty. Considering the characteristics of mobile devices, an efficient data deduplication method is proposed in this paper. First, it separates the hash calculation into foreground and background stages. The background stage calculates the hash values of potentially duplicate files while the foreground stage quickly hashes the file which is being written using a lightweight hash algorithm. Second, a dual-level node structure is proposed to improve the file update efficacy for deduplicated files, saving more storage space against the file re-splitting. Besides, we implement a superlink call to make deduplication process compatible with file-based encryption. These methods are combined to realize a transparent file deduplication (TFDedup) approach, which eliminates redundant data and reduces the associated cost of file update. Experimental results show that TFDedup succeeds to lower the space consumption when serving file updates by 55.6% and accelerate the deduplication process by 50.3%. Junbin Ren, Cheng Ji 0002, Weiwei Jin, Weichao Guo, Yajuan Du, Zongwei Zhu |
ICPADS | 2 |
| 2023 | Ability-aware knowledge distillation for resource-constrained embedded devicesabstractDeep Neural Network (DNN) models have notably improved the efficiency of machine learning tasks. However, their high storage and computational costs restrict their deployment on resource-limited embedded devices. Knowledge distillation (KD) has emerged as a promising approach for compressing DNN models. However, two challenges in KD, namely the capacity gap problem and the time-consuming redundancy problem, have hindered its performance and efficiency in compression. To alleviate these challenges, this paper proposes a novel framework, called Ability-Aware Knowledge Distillation (AAKD). AAKD introduces a knowledge sample selection strategy and an adaptive teacher switching strategy based on the dynamic awareness of the student’s ability. This enables the framework to automatically select suitable knowledge samples and teacher networks according to the increasing representation ability of students. Extensive experiments on different datasets and models have demonstrated that AAKD can enhance the performance of compact student models, significantly improve the efficiency of distillation, and lead to higher compression rates. Yi Xiong 0003, Wenjie Zhai, Xueyong Xu, Jinchen Wang, Zongwei Zhu, Cheng Ji 0002 |
J. Syst. Archit. | 6 |
| 2023 | iAware: Interaction Aware Task Scheduling for Reducing Resource Contention in Mobile SystemsabstractTo ensure the user experience of mobile systems, the foreground application can be differentiated to minimize the impact of background applications. However, this article observes that system services in the kernel and framework layer, instead of background applications, are now the major resource competitors. Specifically, these service tasks tend to be quiet when people rarely interact with the foreground application and active when interactions become frequent, and this high overlap of busy times leads to contention for resources. This article proposes iAware, an interaction-aware task scheduling framework in mobile systems. The key insight is to make use of the previously ignored idle period and schedule service tasks to run at that period. iAware quantify the interaction characteristic based on the screen touch event, and successfully stagger the periods of frequent user interactions. With iAware, service tasks tend to run when few interactions occur, for example, when the device’s screen is turned off, instead of when the user is frequently interacting with it. iAware is implemented on real smartphones. Experimental results show that the user experience is significantly improved with iAware. Compared to the state-of-the-art, the application launching speed and frame rate are enhanced by 38.89% and 7.97% separately, with no more than 1% additional battery consumption. Yongchun Zheng, Changlong Li 0006, Yi Xiong 0003, Weihong Liu, Cheng Ji 0002, Zongwei Zhu, Lichen Yu |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2022 | Task-aware swapping for efficient DNN inference on DRAM-constrained edge systemsabstractObject detection at the edge side is a common task in various environments. The deployment of convolutional neural networks in intelligent edge systems is very challenging because of the highly constrained main-memory space. This study aims at operating neural networks with a reduced memory requirement. The basic idea is that tasks of the same type would involve the same critical subnetwork. We propose identifying the critical network connections by considering the importance of channels. During runtime, the proposed method detects the task types and timely swaps the model parameters of the critical subnetworks from the external storage into dynamic random access memory (DRAM). Compared with conventional network pruning, the proposed approach further reduced the DRAM requirement by 34.6% while maintaining a high inference accuracy. Cheng Ji 0002, Zongwei Zhu, Xianmin Wang, Wenjie Zhai, Xuemei Zong, Mingliang Zhou 0001 |
Int. J. Intell. Syst. | 1 |
| 2022 | Juggler-ResNet: A Flexible and High-Speed ResNet Optimization Method for Intrusion Detection System in Software-Defined Industrial NetworksabstractResNetsare widely used in the intrusion detection system (IDS) of software-defined industrial network to construct accurate intelligence detection of network attacks. However, the IDS based on ResNets has a long detecting interval because of the fine-grained operator and intermediate outcomes of the multi-branch architecture of ResNets. To address this problem, in this article, we propose Juggler-ResNet with a fusible residual structure that preserves the feature extraction ability of the residual structure and enables equivalent transformation to linear topology to support low latency inference service in the industrial application (e.g., malicious network behavior detection, fault diagnosis, etc.). First, we propose a fusible multibranch residual structure to avoid gradient vanishing problems in the training phase. Second, we convert it to linear-topology by using a set of equivalent fusion operators. Finally, the linear-topology model is deployed to accelerate inference speed. Our experimental results on CIFAR-10 and CIFAR-100 show that fusible residual structure can achieve 2.08-4.3x acceleration with state-of-the-art level accuracy performance. Zongwei Zhu, Wenjie Zhai, Huanghe Liu, Jiawei Geng, Mingliang Zhou 0001, Cheng Ji 0002, Gangyong Jia |
IEEE Trans. Ind. Informatics | 6 |
| 2021 | HADFL: Heterogeneity-aware Decentralized Federated Learning FrameworkabstractFederated learning (FL) supports training models on geographically distributed devices. However, traditional FL systems adopt a centralized synchronous strategy, putting high communication pressure and model generalization challenge. Existing optimizations on FL either fail to speedup training on heterogeneous devices or suffer from poor communication efficiency. In this paper, we propose HADFL, a framework that supports decentralized asynchronous training on heterogeneous devices. The devices train model locally with heterogeneity-aware local steps using local data. In each aggregation cycle, they are selected based on probability to perform model synchronization and aggregation. Compared with the traditional FL system, HADFL can relieve the central server’s communication pressure, efficiently utilize heterogeneous computing power, and can achieve a maximum speedup of 3.15x than decentralized-FedAvg and 4.68x than Pytorch distributed training scheme, respectively, with almost no loss of convergence accuracy. Zirui Lian, Weihong Liu, Zongwei Zhu, Cheng Ji 0002 |
DAC | 5 |
| 2021 | Pattern-Guided File Compression with User-Experience Enhancement for Log-Structured File System on Mobile Devices
Cheng Ji 0002, Li-Pin Chang, Riwei Pan, Chao Wu 0006, Congming Gao, Liang Shi 0001, Tei-Wei Kuo, Chun Jason Xue |
FAST | 1 |
| 2021 | Memory-efficient deep learning inference with incremental weight loading and data layout reorganization on edge systems
Cheng Ji 0002, Zongwei Zhu, Li-Pin Chang, Huanghe Liu, Wenjie Zhai |
J. Syst. Archit. | 1 |
| 2021 | iTRIM: I/O-Aware TRIM for Improving User Experience on Mobile DevicesabstractTRIM is a recommended command to deliver data invalidation information of the file system to flash storage. It is issued on both system level and device level. Since it can reduce the number of data copies during device-level garbage collection (DGC), TRIM has been widely used to improve the endurance and performance of mobile devices. Contrary to the common belief, this work identifies that the default TRIM scheme has both merit and drawback to the performance of mobile devices, especially in flash-friendly file system (F2FS), which is a commonly used file system in mobile devices. On one hand, TRIM can reduce garbage collection migration to prolong the flash lifetime as well as improving I/O throughput; On the other hand, TRIM may induce I/O contentions. This article proposes a new TRIM scheme, iTRIM, to distribute the timing overheads to system idle time. To further reduce I/O contention and improve I/O performance, the design of iTRIM considers the TRIM size, and the logical addresses' pattern of victim invalidated data. Experimental results show that iTRIM can minimize I/O contentions while retaining the benefits of the default TRIM scheme for endurance and performance. Yu Liang 0004, Cheng Ji 0002, Chenchen Fu, Rachata Ausavarungnirun, Qiao Li 0001, Riwei Pan, Liang Shi 0001, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Machine learning assisted OSP approach for improved QoS performance on 3D charge-trap based SSDsabstractThree-dimensional (3D) charge-trap based solid-state-drivers (SSDs) have become an emerging storage solution in recent years. One-shot-programming in 3D charge-trap based SSDs could deliver a maximized system input/output (I/O) throughput at the cost of degraded Quality-of-Service (QoS) performance. This paper proposes reinforcement-learning based one-shot-programming (RLOSP), a reinforcement learning based approach to improve the QoS performance for 3D charge-trap based SSDs. By learning the I/O patterns of the workload environments as well as the device internal status, the proposed approach could properly choose requests in the device queue, and allocate physical addresses for these requests during one-shot-programming. In this manner, the storage device could deliver an improved QoS performance. Experimental results reveal that the proposed approach could reduce the worst-case latency at the 99.9th percentile by 37.5%–59.2%, with an optimal system I/O throughput. Zongwei Zhu, Chao Wu 0006, Cheng Ji 0002, Xianmin Wang |
Int. J. Intell. Syst. | 3 |
| 2020 | PHDFS: Optimizing I/O performance of HDFS in deep learning cloud computing platform
Zongwei Zhu, Luchao Tan, Yinzhen Li, Cheng Ji 0002 |
J. Syst. Archit. | 4 |
| 2020 | Maximizing I/O Throughput and Minimizing Performance Variation via Reinforcement Learning Based I/O Merging for SSDsabstractMerging technique is widely adopted by I/O schedulers to maximize system I/O throughput. However, I/O merging could increase the latency of individual I/O, thus incurring prolonged I/O latencies and enlarged performance variations. Even with better system throughput, higher worst-case latency experienced by some requests could block the SSD storage system, which violates the QoS (Quality of Service) requirement. In order to improve QoS performance while providing higher I/O throughput, this paper proposes a reinforcement learning based I/O merging approach. Through learning the characteristic of various I/O patterns, the proposed approach makes merging decisions adaptively based on different I/O workloads. Evaluation results show that the proposed scheme is capable of reducing the standard deviation of I/O latency by 19.1 percent on average, worst-case latency by 7.3-60.9 percent at the 99.9th percentile compared with the latest I/O merging scheme, while maximizing system throughput. Chao Wu 0006, Cheng Ji 0002, Qiao Li 0001, Congming Gao, Riwei Pan, Chenchen Fu, Liang Shi 0001, Chun Jason Xue |
IEEE Trans. Computers | 2 |
| 2020 | Pruning Deep Reinforcement Learning for Dual User Experience and Storage Lifetime Improvement on Mobile DevicesabstractBackground segment cleaning in log-structured file system has a significant impact on mobile devices. A low triggering frequency of the cleaning activity cannot reclaim enough free space for subsequent I/O, thus incurring foreground segment cleaning and impacting the user experience. In contrast, a high triggering frequency could generate excessive block migrations (BMs) and impair the storage lifetime. Prior works address this issue either by performance-biased solutions or incurring excessive memory overhead. In this article, a pruned reinforcement learning-based approach, MOBC, is proposed. Through learning the behaviors of I/O workloads and the statuses of logical address space, MOBC adaptively reduces the number of BMs and the number of triggered foreground segment cleanings. In order to integrate MOBC to resource-constraint mobile devices, a structured pruning method is proposed to reduce the time and space cost. The experimental results show that the pruned MOBC can reduce the worst case latency by 32.5%-68.6% at the 99.9th percentile, and improve the storage endurance by 24.3% over existing approaches, with significantly reduced overheads. Chao Wu 0006, Yufei Cui, Cheng Ji 0002, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Boosting User Experience via Foreground-Aware Cache Management in UFS Mobile DevicesabstractMobile devices today often have multiple applications running simultaneously in the background. These background applications could rapidly consume storage cache resources, thus degrading the performance of foreground applications as well as the user experience. This issue could get worse as modern mobile devices are employing universal flash storage (UFS), which supports faster transmission speed and full-duplex transmission. In this article, a foreground application-aware cache management approach, FOAM, is proposed to address this issue. Through adaptive management of storage cache resources with the awareness of I/O workload patterns, UFS device features, and foreground/background information, I/O performance of foreground application is significantly improved. Experimental results show that the proposed approach could boost the performance of foreground read I/O by 45.9%, foreground write I/O by 18.4% on average compared with the existing approach. Chao Wu 0006, Qiao Li 0001, Cheng Ji 0002, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Inspection and Characterization of App File Usage in Mobile DevicesabstractWhile the computing power of mobile devices has been quickly evolving in recent years, the growth of mobile storage capacity is, however, relatively slower. A common problem shared by budget-phone users is that they frequently run out of storage space. This article conducts a deep inspection of file usage of mobile applications and their potential implications on user experience. Our major findings are as follows: First, mobile applications could rapidly consume storage space by creating temporary cache files, but these cache files quickly become obsolete after being re-used for a short period of time. Second, file access patterns of large files, especially executable files, appear highly sparse and random, and therefore large portions of file space are never visited. Third, file prefetching brings an excessive amount of file data into page cache but only a few prefetched data are actually used. The unnecessary memory pressure causes premature memory reclamation and prolongs application launching time. Through the feasibility study of two preliminary optimizations, we demonstrated a high potential to eliminate unnecessary storage and memory space consumption with a minimal impact on user experience. Cheng Ji 0002, Riwei Pan, Li-Pin Chang, Liang Shi 0001, Zongwei Zhu, Yu Liang 0004, Tei-Wei Kuo, Chun Jason Xue |
ACM Trans. Storage | 1 |
| 2020 | Process Variation Aware Read Performance Improvement for LDPC-Based nand Flash MemoryabstractWith the rapid development of technology scaling and cell density improvement for capacity increase and cost reduction, nand flash memory is confronted with degraded reliability. On one hand, while low-density parity-check (LDPC) codes have been deployed in today's nand flash memories to enhance reliability, flash read latency has still been a performance bottleneck with the increased raw bit error rates (RBER). On the other hand, significant process variations (PV) have been found on existing nand flash memories, which introduce great reliability variations among different flash blocks. Recent studies have proposed to exploit PV to improve endurance by better wear leveling or to improve write performance. These approaches are prone to allocate read data to blocks with low reliability, which further degrades read performance. This paper proposes to enhance read performance of LDPC-equipped nand flash memory by exploiting the reliability variations from PV. The paper consists of three parts. First, a block grouping approach is presented to categorize flash blocks according to their reliability. Second, according to the grouping scheme, a data placement scheme is proposed, which allocates read-hot data to flash blocks with high reliability. At the same time, the read-cold data is moved to blocks with low reliability. As a result, the read performance is enhanced. However, allocating high reliable blocks for read-hot data collides with previous PV-based wear leveling methods. To address the issue, the third part is a grouping partition scheme which limits the amount of high reliable blocks occupied by read-hot data. Therefore, read performance enhancement can be achieved and the wear leveling schemes will be impacted slightly. Experiment results present that, the proposed approach can provide significant read performance improvement on LDPC-equipped nand flash memory and is compatible with the previous PV-based wear leveling. Qiao Li 0001, Liang Shi 0001, Yejia Di, Congming Gao, Cheng Ji 0002, Yu Liang 0004, Chun Jason Xue |
IEEE Trans. Reliab. | 5 |
| 2019 | Parallel all the time: Plane Level Parallelism Exploration for High Performance SSDsabstractSolid state drives (SSDs) are constructed with multiple level parallel organization, including channels, chips, dies and planes. Among these parallel levels, plane level parallelism, which is the last level parallelism of SSDs, has the most strict restrictions. Only the same type of operations which access the same address in different planes can be processed in parallel. In order to maximize the access performance, several previous works have been proposed to exploit the plane level parallelism for host accesses and internal operations of SSDs. However, our preliminary studies show that the plane level parallelism is far from well utilized and should be further improved. The reason is that the strict restrictions of plane level parallelism are hard to be satisfied. In this work, a from plane to die parallel optimization framework is proposed to exploit the plane level parallelism through smartly satisfying the strict restrictions all the time. In order to achieve the objective, there are at least two challenges. First, due to that host access patterns are always complex, receiving multiple same-type requests to different planes at the same time is uncommon. Second, there are many internal activities, such as garbage collection (GC), which may destroy the restrictions. In order to solve above challenges, two schemes are proposed in the SSD controller: First, a die level write construction scheme is designed to make sure there are always N pages of data written by each write operation. Second, in a further step, a die level GC scheme is proposed to activate GC in the unit of all planes in the same die. Combing the die level write and die level GC, write accesses from both host write operations and GC induced valid page movements can be processed in parallel at all time. As a result, the GC cost and average write latency can be significantly reduced. Experiment results show that the proposed framework is able to significantly improve the write performance without read performance impact. Congming Gao, Liang Shi 0001, Chun Jason Xue, Cheng Ji 0002, Jun Yang 0002, Youtao Zhang |
MSST | 4 |
| 2019 | File Fragmentation in Mobile Devices: Measurement, Evaluation, and TreatmentabstractMobile devices, such as smartphones, have become a necessity in our daily life. However, users may notice that after being used for a longtime, mobile devices begin to exhibit a sluggish response. Based on an empirical study on a collection of aged smartphones, this work identified that file fragmentation is among the key factors that contribute to the progressive degradation of response time. This study takes a three-step approach: First, this study designed a set of reproducible file-system aging processes based on User-Interface (UI) script replay. Through the aging processes, it confirmed that file fragmentation quickly emerged, and SQLite files were among the most severely fragmented files. Second, based on the workloads of a selection of popular mobile applications, this study observed that file fragmentation did have an impact on user-perceived latencies. Specifically, the launching time of Chrome on an aged file system was 79 percent slower than it was on a pristine file system. Third, this study evaluated existing treatments of file fragmentation, including space preallocation, persistent journal, and file defragmentation to understand their efficacies and limitations. This study also evaluated a state-of-the-art copyless defragmenter, janusd, to show its advantage over the existing methods. Cheng Ji 0002, Li-Pin Chang, Sangwook Shane Hahn, Sungjin Lee 0001, Riwei Pan, Liang Shi 0001, Jihong Kim 0001, Chun Jason Xue |
IEEE Trans. Mob. Comput. | 1 |
| 2018 | Maximizing I/O throughput and minimizing performance variation via reinforcement learning based I/O merging for SSDs: work-in-progress
Chao Wu 0006, Cheng Ji 0002, Qiao Li 0001, Chenchen Fu, Chun Jason Xue |
CASES | 2 |
| 2018 | Selective Compression Scheme for Read Performance Improvement on Flash DevicesabstractThe increasing density and capacity of NAND flash memory leads to degraded reliability. To address the reliability issue, low-density parity-check code (LDPC) has been deployed in NAND flash memories due to its strong error correction capability. The drawback of LDPC is that, to correct data with high raw bit error rate (RBER), read latency will be amplified. To improve read performance, this paper proposes to apply lossless compression to reduce RBER on data pages. However, compression and decompression incur time overheads. Compressing all the data pages for RBER reduction will degrade write performance. In addition, the variation of compression ratio leads to variation of RBER reduction, thus varied read latency reduction. In this work, a selective data compression scheme is proposed for read performance improvement. Both read frequency and compression ratio of data are taken into consideration. Data in a flash page with high read frequency and good compressibility are prioritized for compression. Experimental results show that the proposed scheme can improve read performance by 42% on average, without impacting write performance. Qiao Li 0001, Liang Shi 0001, Riwei Pan, Cheng Ji 0002, Chun Jason Xue |
ICCD | 4 |
| 2018 | Exploiting Parallelism for Access Conflict Minimization in Flash-Based Solid State DrivesabstractSolid state drives (SSDs) have been widely deployed in personal computers, data centers, and cloud storages. In order to improve performance, SSDs are usually constructed with a number of channels with each channel connecting to a number of nand flash chips, each flash chip consisting of multiple dies and each die containing multiple planes. Based on this parallel architecture, I/O requests are potentially able to access parallel units simultaneously. Despite the rich parallelism offered by the parallel architecture, recent studies show that the utilization of flash parallel units is seriously low. This paper shows that the low parallel unit utilization is highly caused by the access conflict among I/O requests. In this paper, we propose parallel issue queueing (PIQ), a novel I/O scheduler at the host systems. PIQ groups I/O requests without conflicts into the same batch and I/O requests with conflicts into different batches. Hence, the multiple I/O requests in one batch can be fulfilled simultaneously by exploiting the rich parallelism of SSDs. Extensive experimental results show that PIQ delivers significant performance improvement especially for the applications which have heavy access conflicts. Congming Gao, Liang Shi 0001, Cheng Ji 0002, Yejia Di, Kaijie Wu 0001, Chun Jason Xue, Edwin H.-M. Sha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | An I/O Scheduling Strategy for Embedded Flash Storage Devices With Mapping CacheabstractNAND flash memory has been the default storage component in embedded systems. One of the key technologies for flash management is the address mapping scheme between logical addresses and physical addresses, which deals with the inability of in-place-updating in flash memory. Demand-based page-level mapping cache is often applied to match the cache size constraint and performance requirement of embedded storage systems. However, recent studies showed that the management overhead of mapping cache schemes is sensitive to the host I/O patterns, especially when the mapping cache is small. This paper presents a novel I/O scheduling scheme, called MAP+, to alleviate this problem. The proposed scheduling approach reorders I/O requests for performance improvement from two angles. Prioritizing the requests that will hit in the mapping cache, and grouping requests with related logical addresses into large batches. Batches of requests are reordered to further optimize request waiting time. Experimental results show that MAP+ improved upon traditional I/O schedulers by 48% and 18% in terms of read and write latencies, respectively. Cheng Ji 0002, Li-Pin Chang, Chao Wu 0006, Liang Shi 0001, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | Improving File System Performance of Mobile Storage Systems Using a Decoupled Defragmenter
Sangwook Shane Hahn, Sungjin Lee 0001, Cheng Ji 0002, Li-Pin Chang, Inhyuk Yee, Liang Shi 0001, Chun Jason Xue, Jihong Kim 0001 |
USENIX ATC | 3 |
| 2017 | Lightweight Data Compression for Mobile Flash StorageabstractData compression is beneficial to flash storage lifespan. However, because the design of mobile flash storage is highly cost-sensitive, hardware compression becomes a less attractive option. This study investigates the feasibility of data compression on mobile flash storage. It first characterizes data compressibility based on mobile apps, and the analysis shows that write traffic bound for mobile storage volumes is highly compressible. Based on this finding, a lightweight approach is introduced for firmware-based data compression in mobile flash storage. The controller and flash module work in a pipelined fashion to hide the data compression overhead. Together with this pipelined design, the proposed approach selectively compresses incoming data of high compressibility, while leaving data of low compressibility to a compression-aware garbage collector. Experimental results show that our approach greatly reduced the frequency of block erase by 50.5% compared to uncompressed flash storage. Compared to unconditional data compression, our approach improved the write latency by 10.4% at a marginal cost of 4% more block erase operations. Cheng Ji 0002, Li-Pin Chang, Liang Shi 0001, Congming Gao, Chao Wu 0006, Yuangang Wang, Chun Jason Xue |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | I/O scheduling with mapping cache awareness for flash based storage systemsabstractNAND flash memory has been the default storage component in mobile systems. One of the key technologies for flash management is the address mapping scheme between logical addresses and physical addresses, which deals with the inability of in-place-updating in flash memory. Demand-based page-level mapping cache is often applied to match the cache size constraint and performance requirement of mobile storage systems. However, recent studies showed that the management overhead of mapping cache schemes is sensitive to the host I/O patterns, especially when the mapping cache is small. This paper presents a novel I/O scheduling scheme, called MAP, to alleviate this problem. The proposed scheduling approach reorders I/O requests for performance improvement from two angles: Prioritizing the requests that will hit in the mapping cache, and grouping requests with related logical addresses into large batches. Experimental results show that MAP improved upon traditional I/O schedulers by 30% and 8% in terms of read and write latencies, respectively. Cheng Ji 0002, Chao Wu 0006, Li-Pin Chang, Liang Shi 0001, Chun Jason Xue |
EMSOFT | 1 |
| 2016 | Access Characteristic Guided Read and Write Cost Regulation for Performance Improvement on Flash Memory
Qiao Li 0001, Liang Shi 0001, Chun Jason Xue, Kaijie Wu 0001, Cheng Ji 0002, Qingfeng Zhuge, Edwin H.-M. Sha |
FAST | 5 |
| 2016 | An Empirical Study of File-System Fragmentation in Mobile Storage Systems
Cheng Ji 0002, Li-Pin Chang, Liang Shi 0001, Chao Wu 0006, Qiao Li 0001, Chun Jason Xue |
HotStorage | 1 |
| 2014 | A Thread Behavior-Based Memory Management Framework on Multi-core SmartphoneabstractMemory management systems have significantly affected the overall performance of modern multi-core smartphone systems. Android, as one of the most popular smartphone operating systems, adopts a global buddy system with the FCFS (first come, first served) principle for memory allocation, and releases requests to manage external fragmentations and maintain the memory allocation efficiency. However, extensive experimental study on thread behaviors indicates that memory external fragmentation is no longer the crucial bottleneck in most Android applications. Specifically, a thread usually allocates or releases memory in bursts, resulting in serious memory locks and inefficient memory allocation. Furthermore, the pattern of such bursting behaviors varies throughout the life cycle of a thread. The conventional FCFS policy of Android buddy system fails to adapt to such variations and thus suffers from performance degradation. In this paper, we propose a novel memory management framework, called Memory Management Based on Thread Behaviors (MMBTB), for multi-core smartphone systems. It adapts to various thread behaviors through targeted optimizations to provide efficient memory allocation. The efficiency and effectiveness of this new memory management scheme on multicore architecture is proved by a theoretical emulation model. Our experimental studies on the real Android system show that MMBTB can improve the efficiency of memory allocation by 12%-20%, confirming the theoretical analysis results. Zongwei Zhu, Xi Li 0003, Hengchang Liu, Cheng Ji 0002, Xuehai Zhou, Beilei Sun |
ICECCS | 4 |