Chao Wu 0006

dblp:45/3158-6 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0002-7829-1413ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 MedFS: Pursuing Low Update Overhead via Metadata-Enabled Delta Compression for Log-structured File System on Mobile Device
Chao Wu 0006, Cheng Ji 0002, Li-Pin Chang, Zongwei Zhu, Congming Gao, Weichao Guo, Yanzhi Wang 0001
FAST1
2025 Squat: Quant Small Language Models on the Edge
abstract
A growing trend has emerged in designing high-quality Small Language Models (SLMs) with a few million parameters. This trend is driven by the increasing concerns over cloud costs, privacy, and latency. Considering that full parameter training is feasible for SLMs on mobile devices, Quantization-Aware Training (QAT) is employed to improve efficiency by reducing computational overhead and memory footprint. However, previous QAT works adopt fine-grained quantization methods to compress models with billions of parameters on GPUs, incompatible with current commodity hardware, such as mobile and edge devices, which relies on Single Instruction Multiple Data (SIMD) instructions. Thus, the generalization of these methods to SLMs on mobile devices is limited. In this paper, we propose Squat method, an effective QAT framework with deployable quantization for SLMs on mobile devices. Specifically, we propose entropy-guided and distribution-aligned distillation to mitigate the distortion of attention information from quantization. Besides, we employ sub-8-bit token adaptive quantization, assigning varying bit widths to different tokens based on their importance. Furthermore, we develop a SIMD-based Multi-Kernel Mixed-Precision (MKMP) multiplier to support sub-8-bit mixed-precision MAC on mobile devices. Our extensive experiments verify the substantial improvements of our method compared to other QAT methods across various datasets. Furthermore, we achieve an on-device speedup of up to 2.37× compared with its FP16 counterparts, signaling a great advancement. Code: https://github.com/shawnricecake/squant
Xuan Shen, Peiyan Dong, Zhenglun Kong, Yifan Gong 0004, Changdi Yang, Yanyue Xie, Chao Wu 0006, Yanzhi Wang 0001, Pu Zhao 0001
ICCAD10
2025 PQC-LLM: Post-Quantization Delta Compression for LLMs
abstract
The massive parameter scale and substantial storage requirements of large language models (LLMs) hinder their deployment on edge devices and in real-world applications, even after quantization. To address this challenge, we propose PQC-LLM, a Post-Quantization delta Compression scheme for LLMs. PQC-LLM is motivated by two key observations: (1) quantized LLMs still exhibit redundancy on the least significant bits (LSBs); (2) State-of-the-art (SOTA) quantization techniques could only converge to a local optimum. To exploit the potential, PQC-LLM performs a tile-wise deterministic bit replacement by uniformly substituting the least significant$N$bits of all weights within the same tile with the constant bit value. We employ a network architecture search (NAS) mechanism to determine the optimal bit length and value of the uniform replacement. Subsequently, we perform delta compression on each tile using the first weight as the base value. Extensive experiments show that, PQC-LLM effectively reduces the storage requirements of quantized LLMs with negligible lossy or improved task accuracy.
Yujin Zhong, Chao Wu 0006, Cheng Ji 0002
ICPADS2
2024 Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
abstract
Large Language Models (LLMs) stand out for their impressive performance in intricate language modeling tasks. However, their demanding computational and memory needs pose obstacles for broad use on edge devices. Quantization is then introduced to boost LLMs' on-device efficiency. Recent works show that 8-bit or lower weight quantization is feasible with minimal impact on end-to-end task performance, while the activation is still not quantized. On the other hand, mainstream commodity edge devices still struggle to execute these sub-8-bit quantized networks effectively. In this paper, we propose Agile-Quant, an Activation-Guided quantization framework for faster Inference of popular Large Language Models (LLMs) on the Edge. Considering the hardware profiling and activation analysis, we first introduce a basic activation quantization strategy to balance the trade-off of task performance and real inference speed. Then we leverage the activation-aware token pruning technique to reduce the outliers and the adverse impact on attentivity. Ultimately, we utilize the SIMD-based 4-bit multiplier and our efficient TRIP matrix multiplication to implement the accelerator for LLMs on the edge. We apply our framework on different scales of LLMs including LLaMA, OPT, and BLOOM with 4-bit or 8-bit for the activation and 4-bit for the weight quantization. Experiments show that Agile-Quant achieves simultaneous quantization of model weights and activations while maintaining task performance comparable to existing weight-only quantization methods. Moreover, in the 8- and 4-bit scenario, Agile-Quant achieves an on-device speedup of up to 2.55x compared to its FP16 counterparts across multiple edge devices, marking a pioneering advancement in this domain.
Xuan Shen, Peiyan Dong, Zhenglun Kong, Zhengang Li 0001, Ming Lin 0002, Chao Wu 0006, Yanzhi Wang 0001
AAAI7
2024 LOTUS: learning-based online thermal and latency variation management for two-stage detectors on edge devices
abstract
Two-stage object detectors exhibit high accuracy and precise localization, especially for identifying small objects that are favorable for various edge applications. However, the high computation costs associated with two-stage detection methods cause more severe thermal issues on edge devices, incurring dynamic runtime frequency change and thus large inference latency variations. Furthermore, the dynamic number of proposals in different frames leads to various computations over time, resulting in further latency variations. The significant latency variations of detectors on edge devices can harm user experience and waste hardware resources. To avoid thermal throttling and provide stable inference speed, we propose Lotus, a novel framework that is tailored for two-stage detectors to dynamically scale CPU and GPU frequencies jointly in an online manner based on deep reinforcement learning (DRL). To demonstrate the effectiveness of Lotus, we implement it on NVIDIA Jetson Orin Nano and Mi 11 Lite mobile platforms. The results indicate that Lotus can consistently and significantly reduce latency variation, achieve faster inference, and maintain lower CPU and GPU temperatures under various settings. Our code is available at [link].
Yifan Gong 0004, Yushu Wu, Zheng Zhan 0001, Pu Zhao 0001, Liangkai Liu, Chao Wu 0006, Xulong Tang, Yanzhi Wang 0001
DAC6
2024 DACO: Pursuing Ultra-low Power Consumption via DNN-Adaptive CPU-GPU CO-optimization on Mobile Devices
abstract
As Deep Neural Networks (DNNs) become popular in mobile systems, their high computational and memory demands make them major power consumers, especially in limited-budget scenarios. In this paper, we propose DACO, a DNN-Adaptive CPU-GPU CO-optimization technique, to reduce the power consumption of DNNs. First, a resource-oriented classifier is proposed to quantify the computation/memory intensity of DNN models and classify them accordingly. Second, a set of rule-based policies is deduced for achieving the best-suited CPU-GPU system configuration in a coarse-grained manner. Combined with all the rules, a coarse-to-fine CPU-GPU auto-tuning approach is proposed to reach the Pareto-optimal speed and power consumption in DNN inference. Experimental results demonstrate that, compared with the existing approach, DACO could reduce power consumption by up to 71.9% while keeping an excellent DNN inference speed.
Yushu Wu, Chao Wu 0006, Geng Yuan, Yanyu Li, Weichao Guo, Jing Rao, Xipeng Shen, Bin Ren 0002, Yanzhi Wang 0001
DATE2
2024 SuperFlow: A Fully-Customized RTL-to-GDS Design Automation Flow for Adiabatic Quantum- Flux - Parametron Superconducting Circuits
abstract
Superconducting circuits, like Adiabatic Quantum- Flux-Parametron (AQFP), offer exceptional energy efficiency but face challenges in physical design due to sophisticated spacing and timing constraints. Current design tools often neglect the importance of constraint adherence throughout the entire design flow. In this paper, we propose SuperFlow, a fully-customized RTL-to-GDS design flow tailored for AQFP devices. SuperFlow leverages a synthesis tool based on CMOS technology to transform any input RTL netlist to an AQFP-based netlist. Subsequently, we devise a novel place-and-route procedure that simultaneously con-siders wirelength, timing, and routability for AQFP circuits. The process culminates in the generation of the AQFP circuit layout, followed by a Design Rule Check (DR C) to identify and rectify any layout violations. Our experimental results demonstrate that SuperFlow achieves 12.8% wirelength improvement on average and 12.1 % better timing quality compared with previous state- of-the-art placers for AQFP circuits.
Yanyue Xie, Peiyan Dong, Geng Yuan, Zhengang Li 0001, Masoud Zabihi, Chao Wu 0006, Sung-En Chang, Xue Lin 0001, Caiwen Ding, Nobuyuki Yoshikawa, Olivia Chen, Yanzhi Wang 0001
DATE6
2024 AyE-Edge: Automated Deployment Space Search Empowering Accuracy yet Efficient Real-Time Object Detection on the Edge
Chao Wu 0006, Yifan Gong 0004, Liangkai Liu, Mengquan Li, Yushu Wu, Xuan Shen, Geng Yuan, Weisong Shi, Yanzhi Wang 0001
ICCAD1
2024 Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised Learning
abstract
Deep Neural Networks (DNNs), essential for diverse applications such as visual recognition and eldercare, often require a large amount of labeled data for training, making widespread deployment of DNNs a challenging task. Self-supervised learning (SSL) emerges as a promising approach, which leverages inherent patterns within data through diverse augmentations to train models without explicit labels. However, while SSL has shown notable advancements in accuracy, its high computation costs remain a daunting impediment, particularly for resource-constrained platforms. To address this problem, we introduce SimWnW, a similarity-based efficient self-supervised learning framework. By strategically removing less important regions in augmented images and feature maps, SimWnW not only reduces computation costs but also eliminates irrelevant features that might slow down the learning process, thereby accelerating model convergence. The experimental results show that SimWnW effectively reduces the amount of computation costs in self-supervised model training without compromising accuracy. Specifically, SimWnW yields up to 54\% and 51\% computation savings in training from scratch and transfer learning tasks, respectively.
Sheng Li 0019, Chao Wu 0006, Ao Li 0004, Yanzhi Wang 0001, Xulong Tang, Geng Yuan
ICLR2
2024 Search for Efficient Large Language Models
abstract
Large Language Models (LLMs) have long held sway in the realms of artificial intelligence research. Numerous efficient techniques, including weight pruning, quantization, and distillation, have been embraced to compress LLMs, targeting memory reduction and inference acceleration, which underscore the redundancy in LLMs. However, most model compression techniques concentrate on weight optimization, overlooking the exploration of optimal architectures. Besides, traditional architecture search methods, limited by the elevated complexity with extensive parameters, struggle to demonstrate their effectiveness on LLMs. In this paper, we propose a training-free architecture search framework to identify optimal subnets that preserve the fundamental strengths of the original LLMs while achieving inference acceleration. Furthermore, after generating subnets that inherit specific weights from the original LLMs, we introduce a reformation algorithm that utilizes the omitted weights to rectify the inherited weights with a small amount of calibration data. Compared with SOTA training-free structured pruning works that can generate smaller networks, our method demonstrates superior performance across standard benchmarks. Furthermore, our generated subnets can directly reduce the usage of GPU memory and achieve inference acceleration.
Xuan Shen, Pu Zhao 0001, Yifan Gong 0004, Zhenglun Kong, Zheng Zhan 0001, Yushu Wu, Ming Lin 0002, Chao Wu 0006, Xue Lin 0001, Yanzhi Wang 0001
NeurIPS8
2024 Automated Optical Accelerator Search Toward Superior Acceleration Efficiency, Inference Robustness, and Development Speed
abstract
Remarkable breakthroughs but daunting complexities of deep learning have aroused widespread interest in dedicated deep neural network (DNN) acceleration hardware, among which optical accelerators (OAs) are particularly promising thanks to their unprecedentedly high-performance-per-watt. However, the development of OAs is much slower than that of electrical accelerators due to threefold challenges. First, the OA design space is ample and discrete, making it tough for OA optimization; Second, the ecosystem that facilitates OA development is still in its infancy. Techniques to support OA design remain less explored, limiting both the achievable performance and the innovative development of OAs; and Third, OAs are highly sensitive to fabrication-induced process variations and thermal fluctuations (i.e., PTVs), which degrades OAs’ inference robustness and even renders them unusable in practice. In this article, we develop AutOAS, the first-of-its-kind framework for Automated Optical Accelerator Search, in order to jointly boost acceleration efficiency, inference robustness, and development speed. Our AutOAS comprises four enabling components: 1) a holistic OA search space, which takes full consideration of OAs’ micro-architectures (e.g., the type, shape and size of core functional units for data computation and data access), dataflow choices, DNN-to-accelerator mapping methods, memory hierarchy and PTV mitigation techniques; 2) a PTV Regulator, which can emulate the impact of PTVs on OAs’ inference accuracy based on given PTV profiles, and enables energy-efficient PTV mitigation on OAs; 3) an O-Performance Predictor, which enables accurate yet efficient predictions of an OA’s energy, throughput (latency) and chip area according to the DNN model and OA architecture parameters; and 4) two O-Search Engines (i.e., a differentiable search engine and an evolutionary search engine), which can automatically explore the large design space of OAs and identify the optimal accelerators to maximize the acceleration targets. Based on 10 DNN models widely applied in both computer vision and sequence modeling tasks, extensive experiments and ablation studies validate the effectiveness of our PTV Regulator, O-Performance Predictor, and O-Search Engines, as well as the superior performance of AutOAS-generated OAs.
Mengquan Li, Kenli Li 0001, Chao Wu 0006, Gang Liu 0038, Mingfeng Lan, Yunchuan Qin, Zhuo Tang, Weichen Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Condense: A Framework for Device and Frequency Adaptive Neural Network Models on the Edge
abstract
With the popularity of battery-powered edge computing, an important yet under-explored problem is the supporting of DNNs for diverse edge devices. On the one hand, different edge platforms have various runtime requirements and computation/memory capabilities. Deploying the same DNN model is unsatisfiable, while designing a specialized DNN for each platform is prohibitively expensive. On the other hand, for a single edge device, DVFS is leveraged to prolong the battery, incurring significant inference speed variation for the same DNN and consequently poor user experience. To tackle this, we propose Condense, a framework providing a single adaptive model that can be reconfigured (switch to various sub-networks with different computations/parameters) instantly for diverse devices and execution frequencies without any retraining. Experiments demonstrate that Condense can simultaneously provide vast high-accuracy sub-networks with different computations and parameters corresponding to various sparsity ratios to support diverse edge devices with different runtime requirements, and reduce the speed variation under varying frequencies on each device, with a memory cost of only one set of weights.
Yifan Gong 0004, Pu Zhao 0001, Zheng Zhan 0001, Yushu Wu, Chao Wu 0006, Zhenglun Kong, Minghai Qin, Caiwen Ding, Yanzhi Wang 0001
DAC5
2023 RLAlloc: A Deep Reinforcement Learning-Assisted Resource Allocation Framework for Enhanced Both I/O Throughput and QoS Performance of Multi-Streamed SSDs
abstract
Multi-streamed Solid-State Disks (SSDs) have attracted increasing adoption in modern flash storage devices. Despite their excellent promise, effective flash resource allocation is still limiting both their achievable I/O performance and practical implementation. To this end, we develop the first-of-its-kind framework dubbed RLAlloc, which for the first time demonstrates deep Reinforcement Learning-assisted resource Allocation for boosting both I/O throughput and QoS performance of multi-streamed SSDs. Extensive experiments consistently validate the effectiveness of RLAlloc, improving up to 39.9% on I/O throughput and 44.0% on QoS performance over the state-of-the-art competitors.
Mengquan Li, Chao Wu 0006, Congming Gao, Cheng Ji 0002, Kenli Li 0001
DAC2
2023 MOC: Multi-Objective Mobile CPU-GPU Co-Optimization for Power-Efficient DNN Inference
abstract
With the emergence of DNN applications on mobile devices, plenty of attention has been attracted to their optimization. However, the impact of DNN inference tasks on device power consumption is still a lack of comprehensive study. In this work, we propose MOC, a Multi-Objective deep reinforcement learning-assisted DNN inference stage-adaptive CPU-GPU Co-optimization approach. We find through experiments that CPU-GPU parameters, including CPU core, CPU, and GPU frequency, could significantly impact the speed and power consumption of DNN inference. We empirically analyze various stages of DNN inference, including pre/post-processing and feed-forward calculating stages. Based on the analysis, a DNN demand-resource matching model is proposed to classify the DNNs into various categories. Next, a multi-objective deep reinforcement learning (MODRL)-assisted framework is proposed, which considers both the DNN type and hardware environment, to make decisions on DNN inference stage-adaptive CPU/GPU parameter tuning. Finally, a rule-based action refinement technique is introduced to tailor the search space of MOC. Extensive experiments show that, compared with existing works, MOC could substantially reduce the power consumption of DNN inference tasks by up to 74.4%, meanwhile delivering an excellent speed on mobile devices.
Yushu Wu, Yifan Gong 0004, Zheng Zhan 0001, Geng Yuan, Yanyu Li, Chao Wu 0006, Yanzhi Wang 0001
ICCAD7
2023 PackQViT: Faster Sub-8-bit Vision Transformers via Full and Packed Quantization on the Mobile
abstract
While Vision Transformers (ViTs) have undoubtedly made impressive strides in computer vision (CV), their intricate network structures necessitate substantial computation and memory resources. A decision-making process for CV tasks typically entails performing computations with low latency, which is a tricky problem for ViT models. Model quantization is a widely-used technique to optimize the hardware efficiency of deep neural networks. Full quantization under Sub-8-bit precision, in particular, is a promising solution to reduce inference latency significantly. Unfortunately, current commodity hardware, such as CPUs and GPUs, still struggles to efficiently execute these sub-8-bit quantized networks, as their SIMD instructions only support a granularity of 8 bits or wider. Also, there is a scarcity of literature that presents a full quantization paradigm for ViTs. In this paper, we propose an activation-aware fully sub-8-bit quantization-aware training (QAT) framework called PackQViT for efficient yet accurate ViT acceleration on mobile devices to facilitate real-time AI-powered decision-making. Specifically, in revisiting data activation within the ViT dataflow, two characteristics are relevant to quantization strategy and precision: the long-tailed distribution and systematic channel-wise outliers. In response, we employ either log2 quantization or clipping to address the long-tailed distribution and incorporate outlier-aware training for residual link quantization to regulate the various channel-wise outliers more consistently. Notably, due to the systematic fixed pattern, outlier-aware training approach can predict the channel indices and regularized scales of outliers in advance, thus avoiding the runtime data-adaptive selection during inference. Furthermore, we employ Int-$2^{n}$-Softmax, Int-LayerNorm, and Integer GELU to enable integer-only computation flow. Finally, we develop a SIMD-based 4-bit packed multiplier to achieve end-to-end ViT acceleration on mobile phones. Compared to prior studies on ViT quantization using 8-bit precision, PackQViT surpasses other works by an improved accuracy ranging from 0.4\% to 17.9\% for various widely used ViTs on ImageNet dataset; under 4-bit precision, PackQViT demonstrates 0.4%$\sim$2.8% higher accuracy. Compared to the baseline multiplier, our implementations on the Realme GT Android smartphone with Snapdragon 870 SoC CPU achieve 2.6x$\sim$3.7x speedup under 8-bit scenario and 3.8x$\sim$5.9x speedup under 4-bit which ensures practical real-time performance.
Peiyan Dong, Chao Wu 0006, Geng Yuan, Hao Tang 0005, Yanzhi Wang 0001
NeurIPS3
2023 Full-Reference Image Quality Assessment via Low-Level and High-Level Feature Fusion
abstract
We propose a full-reference image quality assessment (FR-IQA) method by incorporating low-level and high-level image features. First, in contrast to the preexisting deep IQA methods, which only use the features extracted by the deep network, we not only use the image gradient to replace the low-level features in the first two stages of the deep network, but also combine them with the middle-stage features of the deep network to construct the new low-level features. The deep features of shallow layers contain some unwanted noise and further result in a decline in IQA performance. Second, we combine the global features extracted by the self-attention-based model with the semantic features extracted by the convolutional neural network to form the high-level features. Instead of directly using the self-attention-based model trained on the classification task, we first train a no-reference (NR) IQA regression model on a larger dataset and then use the global features of this NR-IQA model. The self-attention-based model can capture the internal connections of the image and is more effective in extracting global information due to its larger perceptual field. In the final pooling stage, we combine the average pooling and the standard deviation pooling to obtain the dispersion and concentration of the similarity maps for a more comprehensive description of quality. Experiments show that our FR-IQA method is able to obtain competitive results on three standard IQA datasets.
Chao Wu 0006, Xiaofeng Liao 0001, Hong Yue, Xueyong Xu, Xuekai Wei, Dingcheng Wu, Mingliang Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.1
2022 All-in-One: A Highly Representative DNN Pruning Framework for Edge Devices with Dynamic Power Management
abstract
During the deployment of deep neural networks (DNNs) on edge devices, many research efforts are devoted to the limited hardware resource. However, little attention is paid to the influence of dynamic power management. As edge devices typically only have a budget of energy with batteries (rather than almost unlimited energy support on servers or workstations), their dynamic power management often changes the execution frequency as in the widely-used dynamic voltage and frequency scaling (DVFS) technique. This leads to highly unstable inference speed performance, especially for computation-intensive DNN models, which can harm user experience and waste hardware resources. We firstly identify this problem and then propose All-in-One, a highly representative pruning framework to work with dynamic power management using DVFS. The framework can use only one set of model weights and soft masks (together with other auxiliary parameters of negligible storage) to represent multiple models of various pruning ratios. By re-configuring the model to the corresponding pruning ratio for a specific execution frequency (and voltage), we are able to achieve stable inference speed, i.e., keeping the difference in speed performance under various execution frequencies as small as possible. Our experiments demonstrate that our method not only achieves high accuracy for multiple models of different pruning ratios, but also reduces their variance of inference latency for various frequencies, with minimal memory consumption of only one model and one soft mask.
Yifan Gong 0004, Zheng Zhan 0001, Pu Zhao 0001, Yushu Wu, Chao Wu 0006, Caiwen Ding, Weiwen Jiang, Minghai Qin, Yanzhi Wang 0001
ICCAD5
2021 Pattern-Guided File Compression with User-Experience Enhancement for Log-Structured File System on Mobile Devices
Cheng Ji 0002, Li-Pin Chang, Riwei Pan, Chao Wu 0006, Congming Gao, Liang Shi 0001, Tei-Wei Kuo, Chun Jason Xue
FAST4
2020 Machine learning assisted OSP approach for improved QoS performance on 3D charge-trap based SSDs
abstract
Three-dimensional (3D) charge-trap based solid-state-drivers (SSDs) have become an emerging storage solution in recent years. One-shot-programming in 3D charge-trap based SSDs could deliver a maximized system input/output (I/O) throughput at the cost of degraded Quality-of-Service (QoS) performance. This paper proposes reinforcement-learning based one-shot-programming (RLOSP), a reinforcement learning based approach to improve the QoS performance for 3D charge-trap based SSDs. By learning the I/O patterns of the workload environments as well as the device internal status, the proposed approach could properly choose requests in the device queue, and allocate physical addresses for these requests during one-shot-programming. In this manner, the storage device could deliver an improved QoS performance. Experimental results reveal that the proposed approach could reduce the worst-case latency at the 99.9th percentile by 37.5%–59.2%, with an optimal system I/O throughput.
Zongwei Zhu, Chao Wu 0006, Cheng Ji 0002, Xianmin Wang
Int. J. Intell. Syst.2
2020 Maximizing I/O Throughput and Minimizing Performance Variation via Reinforcement Learning Based I/O Merging for SSDs
abstract
Merging technique is widely adopted by I/O schedulers to maximize system I/O throughput. However, I/O merging could increase the latency of individual I/O, thus incurring prolonged I/O latencies and enlarged performance variations. Even with better system throughput, higher worst-case latency experienced by some requests could block the SSD storage system, which violates the QoS (Quality of Service) requirement. In order to improve QoS performance while providing higher I/O throughput, this paper proposes a reinforcement learning based I/O merging approach. Through learning the characteristic of various I/O patterns, the proposed approach makes merging decisions adaptively based on different I/O workloads. Evaluation results show that the proposed scheme is capable of reducing the standard deviation of I/O latency by 19.1 percent on average, worst-case latency by 7.3-60.9 percent at the 99.9th percentile compared with the latest I/O merging scheme, while maximizing system throughput.
Chao Wu 0006, Cheng Ji 0002, Qiao Li 0001, Congming Gao, Riwei Pan, Chenchen Fu, Liang Shi 0001, Chun Jason Xue
IEEE Trans. Computers1
2020 Pruning Deep Reinforcement Learning for Dual User Experience and Storage Lifetime Improvement on Mobile Devices
abstract
Background segment cleaning in log-structured file system has a significant impact on mobile devices. A low triggering frequency of the cleaning activity cannot reclaim enough free space for subsequent I/O, thus incurring foreground segment cleaning and impacting the user experience. In contrast, a high triggering frequency could generate excessive block migrations (BMs) and impair the storage lifetime. Prior works address this issue either by performance-biased solutions or incurring excessive memory overhead. In this article, a pruned reinforcement learning-based approach, MOBC, is proposed. Through learning the behaviors of I/O workloads and the statuses of logical address space, MOBC adaptively reduces the number of BMs and the number of triggered foreground segment cleanings. In order to integrate MOBC to resource-constraint mobile devices, a structured pruning method is proposed to reduce the time and space cost. The experimental results show that the pruned MOBC can reduce the worst case latency by 32.5%-68.6% at the 99.9th percentile, and improve the storage endurance by 24.3% over existing approaches, with significantly reduced overheads.
Chao Wu 0006, Yufei Cui, Cheng Ji 0002, Tei-Wei Kuo, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Boosting User Experience via Foreground-Aware Cache Management in UFS Mobile Devices
abstract
Mobile devices today often have multiple applications running simultaneously in the background. These background applications could rapidly consume storage cache resources, thus degrading the performance of foreground applications as well as the user experience. This issue could get worse as modern mobile devices are employing universal flash storage (UFS), which supports faster transmission speed and full-duplex transmission. In this article, a foreground application-aware cache management approach, FOAM, is proposed to address this issue. Through adaptive management of storage cache resources with the awareness of I/O workload patterns, UFS device features, and foreground/background information, I/O performance of foreground application is significantly improved. Experimental results show that the proposed approach could boost the performance of foreground read I/O by 45.9%, foreground write I/O by 18.4% on average compared with the existing approach.
Chao Wu 0006, Qiao Li 0001, Cheng Ji 0002, Tei-Wei Kuo, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Maximizing I/O throughput and minimizing performance variation via reinforcement learning based I/O merging for SSDs: work-in-progress
Chao Wu 0006, Cheng Ji 0002, Qiao Li 0001, Chenchen Fu, Chun Jason Xue
CASES1
2018 An I/O Scheduling Strategy for Embedded Flash Storage Devices With Mapping Cache
abstract
NAND flash memory has been the default storage component in embedded systems. One of the key technologies for flash management is the address mapping scheme between logical addresses and physical addresses, which deals with the inability of in-place-updating in flash memory. Demand-based page-level mapping cache is often applied to match the cache size constraint and performance requirement of embedded storage systems. However, recent studies showed that the management overhead of mapping cache schemes is sensitive to the host I/O patterns, especially when the mapping cache is small. This paper presents a novel I/O scheduling scheme, called MAP+, to alleviate this problem. The proposed scheduling approach reorders I/O requests for performance improvement from two angles. Prioritizing the requests that will hit in the mapping cache, and grouping requests with related logical addresses into large batches. Batches of requests are reordered to further optimize request waiting time. Experimental results show that MAP+ improved upon traditional I/O schedulers by 48% and 18% in terms of read and write latencies, respectively.
Cheng Ji 0002, Li-Pin Chang, Chao Wu 0006, Liang Shi 0001, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Lightweight Data Compression for Mobile Flash Storage
abstract
Data compression is beneficial to flash storage lifespan. However, because the design of mobile flash storage is highly cost-sensitive, hardware compression becomes a less attractive option. This study investigates the feasibility of data compression on mobile flash storage. It first characterizes data compressibility based on mobile apps, and the analysis shows that write traffic bound for mobile storage volumes is highly compressible. Based on this finding, a lightweight approach is introduced for firmware-based data compression in mobile flash storage. The controller and flash module work in a pipelined fashion to hide the data compression overhead. Together with this pipelined design, the proposed approach selectively compresses incoming data of high compressibility, while leaving data of low compressibility to a compression-aware garbage collector. Experimental results show that our approach greatly reduced the frequency of block erase by 50.5% compared to uncompressed flash storage. Compared to unconditional data compression, our approach improved the write latency by 10.4% at a marginal cost of 4% more block erase operations.
Cheng Ji 0002, Li-Pin Chang, Liang Shi 0001, Congming Gao, Chao Wu 0006, Yuangang Wang, Chun Jason Xue
ACM Trans. Embed. Comput. Syst.5
2016 I/O scheduling with mapping cache awareness for flash based storage systems
abstract
NAND flash memory has been the default storage component in mobile systems. One of the key technologies for flash management is the address mapping scheme between logical addresses and physical addresses, which deals with the inability of in-place-updating in flash memory. Demand-based page-level mapping cache is often applied to match the cache size constraint and performance requirement of mobile storage systems. However, recent studies showed that the management overhead of mapping cache schemes is sensitive to the host I/O patterns, especially when the mapping cache is small. This paper presents a novel I/O scheduling scheme, called MAP, to alleviate this problem. The proposed scheduling approach reorders I/O requests for performance improvement from two angles: Prioritizing the requests that will hit in the mapping cache, and grouping requests with related logical addresses into large batches. Experimental results show that MAP improved upon traditional I/O schedulers by 30% and 8% in terms of read and write latencies, respectively.
Cheng Ji 0002, Chao Wu 0006, Li-Pin Chang, Liang Shi 0001, Chun Jason Xue
EMSOFT2
2016 An Empirical Study of File-System Fragmentation in Mobile Storage Systems
Cheng Ji 0002, Li-Pin Chang, Liang Shi 0001, Chao Wu 0006, Qiao Li 0001, Chun Jason Xue
HotStorage4