Changlong Li 0006

dblp:148/6622-6 · DBLP profile ↗
← Back
35ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-4042-6538ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 7 first-author · 22 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Secure-collaborative autonomous driving systems for multi-mobile intelligent agents
abstract
Abstract Research on cooperative autonomous driving systems remains nascent. Within mobile intelligent agents, such as unmanned convoys, low-altitude aircraft, and unmanned ships, data security mechanisms for collaboration rely on hardware and their own adaptive control, but these introduce high latency. This paper investigates data security mechanisms for collaborative encryption in conventional computing control systems. Furthermore, considering the ineffectiveness of classical ciphertext-policy attribute-based encryption and Rivest-Shamir-Adleman algorithm (RSA) encryptions against quantum attacks within emerging mobile agent frameworks, this study explores data security issues in cooperative autonomous driving systems involving next-generation mobile agents. Specifically, for external systems requiring resistance against quantum attacks, this article explores the use of AES-256 and lattice-based cryptography. For internal systems, AES-128 encryption balances security and low latency for system data. Based on a collaborative scenario constructed for intelligent mobile agents, this study analyzes security requirements and reviews relevant defensive mechanisms. The findings indicate that the proposed autonomous driving system ensures both quantum security and low latency in the collaborative applications of mobile agents. Compared to the RSA algorithm, AES-256 reduces the inner loop latency by 93$\%$ during the secure encryption and decryption of a 16-bit control text.
Wen Ran, Changlong Li 0006, H.-M. Sha Edwin
Comput. J.2
2025 Archer: Adaptive Memory Compression with Page-Association-Rule Awareness for High-Speed Response of Mobile Devices
Changlong Li 0006, Zongwei Zhu, Chao Wang 0003, Fangming Liu, Edwin H.-M. Sha, Xuehai Zhou
FAST1
2025 Magnifier: A Chiplet Feature-Aware Test Case Generation Method for Deep Learning Accelerators
abstract
The development of deep learning has led to increasing demands for computation and memory, making multi-chiplet accelerators a powerful solution. Multi-chiplet accelerators require more precise consideration of hardware configurations and mapping schemes in terms of computation, memory, and communication patterns compared to monolithic designs, in order to avoid underutilization of performance. However, there is currently a lack of performance testing methods specifically tailored for multi-chiplet accelerators. Existing testing methods primarily focus on correctness testing and do not address potential performance issues from a hardware perspective. To address these issues, this paper proposes Magnifier: a test case generation method for performance testing of multi-chiplet accelerators. Firstly, we analyze typical multi-chiplet accelerator prototype from the perspectives of computation, memory, and communication patterns, and summarize a chiplet feature-aware operator task set. Next, we define the test evaluation metric IPPstd and use a candidate operator set to construct a sampling space for model-level test cases. Finally, we build a GAN to learn the distribution of high-diversity test cases, enabling the rapid generation of high-quality test cases. We validate the proposed method on both simulated and real multi-chiplet accelerators. Experiments show that Magnifier can improve the metric of test cases by up to 3.42 times and significantly reduce generation time, providing valuable insights for optimizing the hardware and software of multi-chiplet accelerators.
Boyu Li 0006, Zongwei Zhu, Weihong Liu, Qianyue Cao, Changlong Li 0006, Cheng Ji 0002, Xi Li 0003, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 HaloFL: Efficient Heterogeneity-Aware Federated Learning Through Optimal Submodel Extraction and Dynamic Sparse Adjustment
abstract
Federated learning (FL) is an advanced framework that enables collaborative training of machine learning models across edge devices. An effective strategy to enhance training efficiency is to allocate the optimal submodel based on each device’s resource capabilities. However, system heterogeneity significantly increases the difficulty of allocating submodel parameter budgets appropriately for each device, leading to the straggler problem. Meanwhile, data heterogeneity complicates the selection of the optimal submodel structure for specific devices, thereby impacting training performance. Furthermore, the dynamic nature of edge environments, such as fluctuations in network communication and computational resources, exacerbates these challenges, making it even more difficult to precisely extract appropriately sized and structured submodels from the global model. To address the challenges in heterogeneous training environments, we propose an efficient FL framework, namely, HaloFL. The framework dynamically adjusts the structure and parameter budget of submodels during training by evaluating three dimensions: 1) model-wise performance; 2) layer-wise performance; and 3) unit-wise performance. First, we design a data-aware model unit importance evaluation method to determine the optimal submodel structure for different data distributions. Next, using this evaluation method, we analyze the importance of model layers and reallocate parameters from noncritical layers to critical layers within a fixed parameter budget, further optimizing the submodel structure. Finally, we introduce a resource-aware dual-UCB multiarmed bandit agent, which dynamically adjusts the total parameter budget of submodels according to changes in the training environment, allowing the framework to better adapt to the performance differences of heterogeneous devices. Experimental results demonstrate that HaloFL exhibits outstanding efficiency in various dynamic and heterogeneous scenarios, achieving up to a 14.80% improvement in accuracy and a$3.06\times $speedup compared to existing FL frameworks.
Zirui Lian, Qianyue Cao, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2025 Freezing-based Memory and Process Co-design for User Experience on Resource-limited Mobile Devices
abstract
Mobile devices with limited resources are prevalent, as they have a relatively low price. Providing a good user experience with limited resources has been a big challenge. This work finds that foreground applications are often unexpectedly interfered by background applications’ memory activities. Improving user experience on resource-limited mobile devices calls for a strong collaboration between memory and process management. This article proposes Ice , a framework to optimize the user experience on resource-limited mobile devices. With Ice, processes that will cause frequent refaults in the background are identified and frozen accordingly. The frozen application will be thawed when memory condition allows. Based on the proposed Ice, this work shows that the refault can be further reduced by revisiting the LRU lists in the original kernel with app-freezing awareness (called Ice + ). Evaluation of resource-limited mobile devices demonstrates that the user experience is effectively improved with Ice. Specifically, Ice boosts the frame rate by 1.57x on average over the state of the art. The frame rate is further enhanced by 5.14% on average with Ice + .
Changlong Li 0006, Zongwei Zhu, Chun Jason Xue, Yu Liang 0004, Rachata Ausavarungnirun, Liang Shi 0001, Xuehai Zhou
ACM Trans. Comput. Syst.1
2025 A Lightweight I/O Throttling Service to Improve the User Experience of Mobile Devices
abstract
As one of the most frequently occurring operations, I/Os significantly affect the application launching time and frame rate of mobile devices, hence influencing the user experience. However, the response speed of I/O requests is still the bottleneck in practice. This paper shows that high I/O latency is usually due to the congestion inside Flash instead of the system layer. Unfortunately, Flash is treated as a black box device and cannot be modified after delivery. In this paper, we propose a novel service to address this issue without an intra-Flash modification. Specifically, this paper proposes a lightweight I/O throttling framework in mobile systems named FlashDAM. This service throttles the I/O flow to make way for I/Os that may block the foreground application. FlashDAM is the first work that proves that proper I/O throttling positively affects the user experience, contrary to the common belief. Furthermore, this paper proposes FlashDAM$^+$, an enhanced version of FlashDAM. By coordinating I/O throttling and compression, FlashDAM's effect in the system layer is minimized. We have implemented FlashDAM on real mobile devices. Experimental results illustrate that the app launching speed and frame rate are enhanced by 72% and 45% separately compared to the state-of-the-art. When enabling the compression feature of FlashDAM, that is, FlashDAM$^+$, screen jank and application launch latency are further reduced by 9.5% and 11.4%, respectively, under heavy background I/O load.
Changlong Li 0006, Zongwei Zhu, Yuyangjun Lu, Chao Wang 0003, Xuehai Zhou, Edwin H.-M. Sha
IEEE Trans. Serv. Comput.1
2024 Sparrow: Flexible Memory Deduplication in Android Systems with Similar-Page Awareness
abstract
Mobile devices have become ubiquitous in daily life. In contrast to traditional servers, mobile devices suffer from limited memory resources, leading to a significant degradation in the user experience. This paper demonstrates that the primary cause of memory consumption lies in anonymous pages associated with application heaps. Existing schemes are ineffective in deduplicating these pages due to the limited occurrence of the same anonymous pages. This paper presents Sparrow, a similar-page aware deduplication solution for mobile systems. Sparrow shows that memory pages still have the potential to deduplicate, even though the same pages are rare. An interesting observation inspires this, that is, a high number of pages having the partially-same contents. We have implemented Sparrow on real-life smartphones. Experimental results indicate that 30.45% more space can be saved with Sparrow.
Guangyu Wei, Changlong Li 0006, Rui Xu 0013, Qingfeng Zhuge, Edwin H.-M. Sha
DATE2
2024 Argus: Real-Time HQ Video Decoding with CPU Coordinating on Consumer Devices
abstract
Real-time high-quality (HQ) videos are increasingly popular in daily life (e.g., 4K video, AR/VR). However, due to the ultra-high definition and high frame rate, existing decoders cannot deal with the video frames on time. Video decoding has become a bottleneck in modern computers, especially for consumer devices. To ensure low latency, the system drops frames when the decoder is under pressure, which sacrifices the video quality. This paper shows that it is possible to realize both low latency and high quality, as we observe that computers use customized hardware to decode video frames but idle CPU resources. In this paper, we propose a new real-time HQ video decoding solution called Argus. The key insight is to make use of the wasted CPU resources. However, we will show that scheduling improper frames to the CPU can degrade, instead of improve the performance, which is out of the expect. To tackle the fundamental challenges, this paper further proposes two novel schemes: (1) a light neural network model to estimate the decoder pressure; and (2) a scheduler with frame-characteristics awareness. We have implemented Argus on both simulators and real-life consumer devices. Experimental results illustrate that Argus can reduce the tail queuing latency by ${4 3. 8 \%}$ on average. More importantly, with the coordination of CPUs, the smooth experience of video playback is effectively improved (${2. 2 \%}$ frame loss is avoided on average), compared to the state-of-the-art.
Changlong Li 0006
RTSS2
2024 Ace-Sniper: Cloud-Edge Collaborative Scheduling Framework With DNN Inference Latency Modeling on Heterogeneous Devices
abstract
The cloud–edge collaborative inference requires efficient scheduling of artificial intelligence (AI) tasks to the appropriate edge intelligence devices. Gls DNN inference latency has become a vital basis for improving scheduling efficiency. However, edge devices exhibit highly heterogeneous due to the differences in hardware architectures, computing power, etc. Meanwhile, the diverse deep neural networks (DNNs) are continuing to iterate over time. The diversity of devices and DNNs introduces high computational costs for measurement methods, while invasive prediction methods face significant development efforts and application limitations. In this article, we propose and develop Ace-Sniper, a scheduling framework with DNN inference latency modeling on heterogeneous devices. First, to address the device heterogeneity, a unified hardware resource modeling (HRM) is designed by considering the platforms as black-box functions that output feature vectors. Second, neural network similarity (NNS) is introduced for feature extraction of diverse and frequently iterated DNNs. Finally, with the results of HRM and NNS as input, the performance characterization network is designed to predict the latencies of the given unseen DNNs on heterogeneous devices, which can be combined into most time-based scheduling algorithms. Experimental results show that the average relative error of DNN inference latency prediction is 11.11%, and the prediction accuracy reaches 93.2%. Compared with the nontime-aware scheduling methods, the average waiting time for tasks is reduced by 82.95%, and the platform throughput is improved by 63% on average.
Weihong Liu, Jiawei Geng, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Zirui Lian, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 Unleashing Network/Accelerator Co-Exploration Potential on FPGAs: A Deeper Joint Search
abstract
Recently, algorithm-hardware co-exploration for neural networks (NNs) has become the key to obtaining high-quality solutions. However, previous efforts for FPGAs focus on neural architecture search (NAS) while lacking hardware architecture search (HAS), thus limiting the full potential of co-design. Although expanding the scope of HAS offers performance potential, the exponentially increased joint search space presents a formidable challenge. To address this, we propose a deep and efficient framework, which jointly searches for Networks and Accelerators for FPGAs in a balanced co-search space. First, we adjust the NAS space and then introduce a block-level bitwidth search on the software side. Meanwhile, we design a hardware-friendly quantization algorithm to facilitate hardware efficiency and accuracy. Second, we design a dataflow-configurable hardware unit with computation and memory access optimizations for quantized multiplication. Based on this, we incorporate critical heterogeneous multicore architecture exploration on the hardware side. Third, to enable rapid hardware feedback in the enlarged HAS space, we perform resource and performance modeling and design a fast hardware generation algorithm based on the genetic algorithm. Specifically, we apply optimization techniques, like mapping space pruning, greedy bandwidth allocation, and coarse-grained search, to speed up this process. We validate in edge and cloud scenarios. Experimental results show that efficiently explores a significantly larger joint space and provides high-quality solutions. Compared with previous state-of-the-art co-design works, the searched CNN-accelerator pairs improve the throughput by 2.07× ~ 7.10× and energy efficiency by 1.41× ~ 2.27× under similar accuracy on the ImageNet dataset.
Wenqi Lou, Lei Gong 0003, Chao Wang 0003, Jiaming Qian, Xuan Wang 0020, Changlong Li 0006, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 Flexible and Efficient Memory Swapping Across Mobile Devices With LegoSwap
abstract
This article presents LegoSwap, a cross-device memory swapping mechanism for mobile devices. It exploits the unbalanced utilization of memory resources across devices. With LegoSwap, remote memory is utilized in a seamless plug-and-play manner. It achieves comparable-to-local swapping performance based on existing network infrastructure. In addition, LegoSwap frees from the effect of remote I/O disconnection and minimizes the effect on remote devices. This is realized by three novel approaches: resource-dedicated swapping for fast swapping among devices, app-aware swapping for network connectivity considerations, and elastic swap area management for inter-device interference relieving. LegoSwap is implemented on real-life mobile devices. Experimental results show that LegoSwap can enhance app caching capability by 2x compared with no swapping, and improve performance by 2.3x compared with state-of-the-art remote swapping. More importantly, local swapping induced read-write conflicts are largely removed.
Changlong Li 0006, Yu Liang 0004, Liang Shi 0001, Chao Wang 0003, Chun Jason Xue, Xuehai Zhou
IEEE Trans. Parallel Distributed Syst.1
2023 ICE: Collaborating Memory and Process Management for User Experience on Resource-limited Mobile Devices
abstract
Mobile devices with limited resources are prevalent as they have a relatively low price. Providing a good user experience with limited resources has been a big challenge. This paper found that foreground applications are often unexpectedly interfered by background applications' memory activities. Improving user experience on resource-limited mobile devices calls for a strong collaboration between memory and process management. This paper proposes a framework, Ice, to optimize the user experience on resource-limited mobile devices. With Ice, processes that will cause frequent refaults in the background are identified and frozen accordingly. The frozen application will be thawed when memory condition allows. Evaluation of resource-limited mobile devices demonstrates that the user experience is effectively improved with Ice. Specifically, Ice boosts the frame rate by 1.57x on average over the state-of-the-art.
Changlong Li 0006, Yu Liang 0004, Rachata Ausavarungnirun, Zongwei Zhu, Liang Shi 0001, Chun Jason Xue
EuroSys1
2023 FlashDAM: Flexible I/O Throttling for the User Experience of Mobile Systems
abstract
I/O plays an important role in the user experience. However, their quick response cannot be ensured in mobile systems, which is always blamed by users. Our study indicates that poor response is always caused by Flash-device side congestion, instead of I/O scheduling in the system layer. Unfortunately, the Flash device is treated as a black box and is not allowed to be modified after delivery. This paper explores a new approach to address the in-device problem without any invasive modification in the Flash. In this paper, we propose FlashDAM, a flexible I/O throttling framework in mobile systems. Contrary to the common belief, FlashDAM shows that proper I/O throttling, rather than straightforward boosting, has a positive effect on the user experience. We have implemented FlashDAM on off-the-shelf smartphones. Experimental results show that the application launch speed and frame rate stability can be enhanced by 72% and 45% separately, compared to the state-of-the-art.
Changlong Li 0006, Chao Wang 0003, Xuehai Zhou, Edwin H.-M. Sha
ICCD1
2023 IOSR: Improving I/O Efficiency for Memory Swapping on Mobile Devices Via Scheduling and Reshaping
abstract
Mobile systems and applications are becoming increasingly feature-rich and powerful, which constantly suffer from memory pressure, especially for devices equipped with limited DRAM. Swapping inactive DRAM pages to the storage device is a promising solution to extend the physical memory. However, existing mobile devices usually adopt flash memory as the storage device, where swapping DRAM pages to flash memory may introduce significant performance overhead. In this paper, we first conduct an in-depth analysis of the I/O characteristics of the flash-based memory swapping, including the I/O interference and swap I/O randomness in swap subsystem. Then an I/O efficiency optimization framework for memory swapping (IOSR) is proposed to enhance the performance of flash-based memory swapping for mobile devices. IOSR consists of two methods: swap I/O scheduling (SIOS) and swap I/O pattern reshaping (SIOR). SIOS is designed to schedule the swap I/O to reduce interference with other processes I/Os. SIOR is designed to reshape the swap I/O pattern with process-oriented swap slot allocation and adaptive granularity swap read-ahead. IOSR is implemented on Google Pixel 4. Experimental results show that IOSR reduces the application switching time by 31.7% and improves the swap-in bandwidth by 35.5% on average compared to the state-of-the-art.
Wentong Li 0002, Liang Shi 0001, Changlong Li 0006, Edwin H.-M. Sha
ACM Trans. Embed. Comput. Syst.4
2023 iAware: Interaction Aware Task Scheduling for Reducing Resource Contention in Mobile Systems
abstract
To ensure the user experience of mobile systems, the foreground application can be differentiated to minimize the impact of background applications. However, this article observes that system services in the kernel and framework layer, instead of background applications, are now the major resource competitors. Specifically, these service tasks tend to be quiet when people rarely interact with the foreground application and active when interactions become frequent, and this high overlap of busy times leads to contention for resources. This article proposes iAware, an interaction-aware task scheduling framework in mobile systems. The key insight is to make use of the previously ignored idle period and schedule service tasks to run at that period. iAware quantify the interaction characteristic based on the screen touch event, and successfully stagger the periods of frequent user interactions. With iAware, service tasks tend to run when few interactions occur, for example, when the device’s screen is turned off, instead of when the user is frequently interacting with it. iAware is implemented on real smartphones. Experimental results show that the user experience is significantly improved with iAware. Compared to the state-of-the-art, the application launching speed and frame rate are enhanced by 38.89% and 7.97% separately, with no more than 1% additional battery consumption.
Yongchun Zheng, Changlong Li 0006, Yi Xiong 0003, Weihong Liu, Cheng Ji 0002, Zongwei Zhu, Lichen Yu
ACM Trans. Embed. Comput. Syst.2
2022 CDB: critical data backup design for consumer devices with high-density flash based hybrid storage
abstract
Hybrid flash based storage constructed with high-density and low-cost flash memory are becoming increasingly popular in consumer devices during the last decade. However, to protect critical data, existing methods are designed for improving reliability of consumer devices with non-hybrid flash storage. Based on evaluations and analysis, these methods will result in significant performance and lifetime degradation in consumer devices with hybrid storage. The reason is that different kinds of memory in hybrid storage have different characteristics, such as performance and access granularity. To address the above problems, a critical data backup (CDB) method is proposed to backup designated critical data with making full use of different kinds of memory in hybrid storage. Experiment results show that compared with the state-of-the-arts, CDB achieves encouraging performance and lifetime improvement.
Longfei Luo, Dingcui Yu, Liang Shi 0001, Chuanming Ding, Changlong Li 0006, Edwin H.-M. Sha
DAC5
2022 DWR: Differential Wearing for Read Performance Optimization on High-Density NAND Flash Memory
abstract
With the cost reduction and density optimization, the read performance and lifetime of high-density NAND flash memory have been significantly degraded during the last decade. Previous works proposed to optimize lifetime with wear leveling and optimize read performance with reliability improvement. However, with wearing, the reliability and read performance will be degraded along with the life of the device. To solve this problem, a differential wearing scheme (DWR) is proposed to optimize the read performance. The basic idea of DWR is to partition the flash memory into two areas and wear them at different speeds. For the area with low wearing speed, read operations are scheduled for read performance optimization. For the area with high wearing speed, write operations are scheduled but designed to avoid generating bad blocks early. Through careful design and real workloads evaluation on 3D TLC NAND flash, DWR achieves encouraging read performance optimization with negligible impacts to the lifetime.
Yunpeng Song, Qiao Li 0001, Yina Lv, Changlong Li 0006, Liang Shi 0001
DATE4
2022 CacheSifter: Sifting Cache Files for Boosted Mobile Performance and Lifetime
Yu Liang 0004, Riwei Pan, Yufei Cui, Rachata Ausavarungnirun, Xianzhang Chen, Changlong Li 0006, Tei-Wei Kuo, Chun Jason Xue
FAST7
2022 Read latency variation aware performance optimization on high-density NAND flash based storage systems
Liang Shi 0001, Yina Lv, Longfei Luo, Changlong Li 0006, Chun Jason Xue, Edwin H.-M. Sha
CCF Trans. High Perform. Comput.4
2022 Practical optimizations for lightweight distributed file system on consumer devices
Yuze Xu, Han Wang 0051, Ben Gu, Yina Lv, Longfei Luo, Changlong Li 0006, Liang Shi 0001
CCF Trans. High Perform. Comput.7
2022 Tail Latency Optimization for LDPC-Based High-Density and Low-Cost Flash Memory Devices
abstract
Flash memory has been developed with bit density improvement, technology scaling, and 3-D stacking. With this trend, its reliability has been significantly degraded. Error correction code (ECC), such as low-density parity code (LDPC), which has strong error correction capability, has been deployed to solve this problem. However, one of the critical issues of LDPC is that it would introduce a long decoding latency on devices with low reliability. In this case, tail latency would happen, which will significantly impact the quality of service. In this work, a set of smart refresh schemes is proposed to optimize the tail latency. The basic idea of the work is to refresh data when the accessed data have a long decoding latency. Two smart refresh schemes are proposed for this work. The first refresh scheme is designed to refresh data with a long access latency when they are accessed several times. The second refresh scheme is designed to periodically check data with an extremely long access latency and refresh them. To further optimize the refresh overhead caused by the above refresh schemes, a dual-ECC-based refresh scheme is proposed. Besides, a mathematical model for all proposed schemes is constructed to clarify the benefit of each scheme. The experimental results show that the proposed schemes can significantly improve the tail latency with acceptable overhead. What is more, the access performance is well maintained compared with the state-of-the-art work.
Yina Lv, Liang Shi 0001, Longfei Luo, Changlong Li 0006, Chun Jason Xue, Edwin H.-M. Sha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 MobileSwap: Cross-Device Memory Swapping for Mobile Devices
abstract
This paper presents MobileSwap, a cross-device memory swapping scheme for mobile devices. It exploits the unbalanced utilization of memory resources across devices. MobileSwap achieves comparable-to-local swapping performance based on existing network infrastructure. This is realized by two novel approaches: resource dedicated swapping for fast swapping among devices and app aware swapping for network connectivity considerations. MobileSwap is implemented and deployed on real mobile devices. Experimental results show that MobileSwap can enhance app caching capability by 2x compared with no swapping, and improve performance by 2.3x compared with state-of-the-art remote swapping. More importantly, local swapping induced read-write conflicts are largely removed.
Changlong Li 0006, Liang Shi 0001, Chun Jason Xue
DAC1
2021 SFP: Smart File-Aware Prefetching for Flash based Storage Systems
abstract
Currently, most of the Flash-based storage systems reduce the performance gap between the main memory and storage by data prefetching. However, conventional prefetching techniques perform well on hard disk drives but have limited effectiveness and efficiency on Flash. It is because the complicate data access patterns in modern systems have not been well considered. In this paper, we propose SFP, a smart file-aware prefetching scheme for Flash-based storage systems. SFP demonstrates that prefetching accuracy and efficiency can be improved comprehensively in a file-aware approach. Furthermore, three schemes are proposed: file access pattern learning, dynamic window-based file prefetching, and learning model size optimization. Experiments on the real server show that SFP reduces the access latency by up to 40% compared with the state-of-the-art with low memory and computation cost.
Han Wang 0051, Longfei Luo, Liang Shi 0001, Changlong Li 0006, Chun Jason Xue, Qingfeng Zhuge, Edwin H.-M. Sha
ACM Great Lakes Symposium on VLSI4
2021 Dynamic File Cache Optimization for Hybrid SSDs with High-Density and Low-Cost Flash Memory
abstract
Over the last few years, hybrid solid-state drives (SSDs) have been widely adopted due to their high performance and high capacity. Devices equipped with hybrid SSDs can be utilized to cache files from the network for performance improvement. However, this paper finds an interesting observation, that is, the efficiency of hybrid SSDs is significantly degraded instead of improved when too much data is cached. This is because the internal mode switching between different types of flash memory is affected by the device utilization. This paper proposes a dynamic file cache optimization scheme for hybrid SSDs, DFCache, which optimizes the device’s efficiency and limits unreasonable space consumption. DFCache includes two key ideas, dynamic cache space management, and intelligent cache file sifting. DFCache is implemented in Linux kernel and tested under real hybrid SSDs. Experimental results show that the I/O performance outperforms the state-of-the-art by up to 3.7x.
Ben Gu, Longfei Luo, Yina Lv, Changlong Li 0006, Liang Shi 0001
ICCD4
2021 Understanding and Optimizing Hybrid SSD with High-Density and Low-Cost Flash Memory
abstract
With the development of NAND flash technology, hybrid SSDs with high-density and low-cost flash memory have become the mainstream of the existing SSD architecture. In this architecture, two flash modes can be dynamically switched, such as single-level cell (SLC) mode and quad-level cell (QLC) mode. Based on evaluations and analysis of multiple real devices, this paper presents two interesting findings. They demonstrate that the coordination between the two flash-modes is not well-designed in existing architectures. This paper proposes HyFlex, which redesigns the strategies of data placement and flash-mode management of hybrid SSDs in a flexible approach. Specifically, two novel optimization strategies are proposed: velocity-based I/O scheduling (VIS) and garbage collection (GC)-aware capacity tuning (GCT). Experimental results show that HyFlex achieves encouraging performance and endurance improvement.
Liang Shi 0001, Longfei Luo, Yina Lv, Changlong Li 0006, Edwin H.-M. Sha
ICCD5
2021 LKSM: Light Weight Key-Value Store for Efficient Application Services on Local Distributed Mobile Devices
abstract
With the development of mobile network and corresponding techniques, more and more works focus on providing efficient services based on mobile devices. Furthermore, motivated by IoT, studies of local distributed mobile devices attract attentions of both industry and academia in recent years. However, existing storage systems cannot manage data and support the QoS of mobile services well. This paper presents LKSM, a light weight key-value storage system, which can be deployed on either one node or multiple nodes. To the best of our knowledge, it is the first attempt to propose key-value store in this scenario. We carefully analyze the challenges when designing the system on mobile clusters, and further propose RDS for addressing. With the help of RDS, LKSM achieves the goal of lower latency, better scalability, and higher availability. Furthermore, based on RDS, a novel data management strategy is presented, which successfully avoid energy holes of mobile clusters and achieves the tradeoff between performance and energy. We organize LKSM using a log-structured merge-tree and implement it based on LevelDB, an open source key-value storage system proposed by Google. Experiments on physical smartphones demonstrate that LKSM presents much higher performance compared with the ported LevelDB on mobile devices.
Changlong Li 0006, Hang Zhuang, Qingfeng Wang 0004, Chao Wang 0003, Xuehai Zhou
IEEE Trans. Serv. Comput.1
2020 An Empirical Study of Hybrid SSD with Optane and QLC Flash
abstract
Emerging non-volatile memory (NVM) technologies provide a new way to solve the I/O bottleneck problem. As one of the widely respected solutions, hybrid storage device performance in the real environment is worth studying. Previously, due to the delayed progress of NVM, most of the studies are proceeded on simulated devices. In this paper, an empirical study is presented on the state-of-the-art hybrid storage device - Intel Optane H10, which is designed with Optane Memory and Quad-Level Cell (QLC) NAND flash. Several interesting findings are concluded with the study, which should be well considered during the employment.
Yina Lv, Changlong Li 0006, Shouzhen Gu, Liang Shi 0001
ICCD3
2020 SEAL: User Experience-Aware Two-Level Swap for Mobile Devices
abstract
App caching is important for mobile devices, which enables fast switching and state restoration of apps by caching all the pages in memory. Memory swapping can improve app caching capability by evicting pages to the secondary storage. However, enabling memory swapping could induce jitters in interactions, which significantly degrades the user experience. As a result, storage-based swapping is disabled by default in most mobile devices. This article proposes a novel swap framework, SEAL, a user experience-aware two-level swapping, which maximizes the benefits of memory swapping and minimizes the negative impact on user experience in interactions. Inspired by a study on the access characteristics of a set of popular apps on mobile devices, the framework adopts compressed memory as the first swap level (SL1) and secondary storage as the second swap level (SL2). To optimize user experience comprehensively, three schemes are proposed. First, a novel page identification scheme is proposed to guide the page placement between these two levels. Second, a hidden page loading (HPL) scheme is proposed to load pages from SL2 to SL1 for optimized user experience during app execution. Finally, an app-granularity swapping scheme is proposed to swap data in the unit of apps. Experiments on real devices show that app caching capability is improved by 2.43× on average when enabling SEAL while minimizing the negative impact on user experience.
Changlong Li 0006, Liang Shi 0001, Yu Liang 0004, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Low-Shot Multi-label Incremental Learning for Thoracic Diseases Diagnosis
Qingfeng Wang 0004, Jie-Zhi Cheng, Hang Zhuang, Changlong Li 0006, Zhiqin Liu, Jun Huang 0005, Chao Wang 0003, Xuehai Zhou
ICONIP (7)5
2017 Efficient Distributed Smith-Waterman Algorithm Based on Apache Spark
abstract
The Smith-Waterman algorithm, which produces the optimal local alignment between pairwise sequences, is universally used as a key component in bioinformatics fields. It is more sensitive than heuristic approaches, but also more time-consuming. To speed up the algorithm, Single-Instruction Multiple-Data (SIMD) instructions have been used to parallelize the algorithm by leveraging data parallel strategy. However, SIMD-based Smith-Waterman (SW) algorithms show limited scalability. Moreover, the recent next-generation sequencing machines generate sequences at an unprecedented rate, so faster implementations of the sequence alignment algorithms are needed to keep pace. In this paper, we present CloudSW, an efficient distributed Smith-Waterman algorithm which leverages Apache Spark and SIMD instructions to accelerate the algorithm. To facilitate easy integration of distributed Smith-Waterman algorithm into third-party software, we provide application programming interfaces (APIs) service in cloud. The experimental results demonstrate that 1) CloudSW has outstanding performance and achieves up to 3.29 times speedup over DSW and 621 times speedup over SparkSW. 2) CloudSW has excellent scalability and achieves up to 529 giga cell updates per second (GCUPS) in protein database search with 50 nodes in Aliyun Cloud, which is the highest performance that has been reported as far as we know.
Changlong Li 0006, Hang Zhuang, Jiali Wang 0003, Qingfeng Wang 0004, Xuehai Zhou
CLOUD2
2017 Distributed gene clinical decision support system based on cloud computing
abstract
The clinical decision support system can effectively solve the limitations of doctors' knowledge, reduce misdiagnosis and help enhance health. The traditional genetic data storage and analysis technology based on the stand-alone environment have limited scalability, which has been difficult to meet the computational requirements of rapid genetic data growth. In this paper, we propose a distributed gene clinical decision support system, which is named as GCDSS. We implemented a prototype based on cloud computing. To speed up the data processing of GCDSS, we present a novel distributed read mapping algorithm CloudBWA that leverages batch processing strategy to map reads on Apache Spark. Evaluations show that GCDSS and its component CloudBWA achieve outstanding performance and excellent scalability. Compared with distributed algorithms, CloudBWA achieves up to 2.63 times speedup over SparkBWA.
Changlong Li 0006, Hang Zhuang, Jiali Wang 0003, Qingfeng Wang 0004, Chao Wang 0003, Xuehai Zhou
BIBM2
2017 DSA: Scalable Distributed Sequence Alignment System Using SIMD Instructions
abstract
Sequence alignment algorithms are a basic and critical component of many bioinformatics fields. With rapid development of sequencing technology, the fast growing reference database volumes and longer length of query sequence become new challenges for sequence alignment. However, the algorithms have prohibitively high time and space complexity. In this paper, we present DSA, a scalable distributed sequence alignment system that employs Apache Spark to process sequences data in a horizontally scalable distributed environment, and leverages data parallel strategy based on Single Instruction Multiple Data (SIMD) instruction to parallelize the algorithms in each core of worker node. The experimental results demonstrate that 1) DSA has outstanding performance and achieves up to 201x speedup over SparkSW. 2) DSA has excellent scalability and achieves near linear speedup when increasing the number of nodes in cluster.
Changlong Li 0006, Hang Zhuang, Jiali Wang 0003, Qingfeng Wang 0004, Jinhong Zhou, Xuehai Zhou
CCGrid2
2017 Light Weight Key-Value Store for Efficient Services on Local Distributed Mobile Devices
abstract
With the development of mobile network and corresponding techniques, more and more works focus on providing efficient services based on mobile devices. Furthermore, motivated by IoT, studies of local distributed mobile devices attract attentions of both industry and academia in recent years. However, existing storage systems cannot manage data and support the QoS of mobile services well. This paper presents LKSM, a light weight key-value storage system, which can be deployed on either one node or multiple nodes. To the best of our knowledge, it is the first attempt to propose key-value store in this scenario. We carefully analyze the challenges when designing the system on mobile cluster, and further propose RDS for addressing. With the help of RDS, LKSM achieves the goal of lower latency, better scalability, and higher availability. We organize LKSM using a log-structured merge-tree, and implement it based on LevelDB, an open source key-value storage system proposed by Google. Experiments on physical smartphones demonstrate that LKSM presents much higher performance compared with the ported LevelDB on mobile devices.
Changlong Li 0006, Hang Zhuang, Jiali Wang 0003, Chao Wang 0003, Xuehai Zhou
ICWS1
2017 Natural Language Processing Service Based on Stroke-Level Convolutional Networks for Chinese Text Classification
abstract
With the development of deep learning and artificial intelligence, more and more research apply neural networks to natural language processing tasks. However, while the majority of these research take English corpus as the dataset, few studies have been done using Chinese corpus. Meanwhile, Existing Chinese processing algorithms typically regard Chinese word or Chinese character as the basic unit but ignore the deeper information into the Chinese character. In Chinese linguistic, strokes are the basic unit of Chinese character who are similar to letters of the English word. Inspired by the recent success of deep learning at character-level, we delve deeper to Chinese stroke level for Chinese language processing and developed it into service for Chinese text classification. In this paper, we dig the basic feature of the strokes considering the similar Chinese character components and propose a new method to leverage Chinese stroke for learning the continuous representation of Chinese character and develop it into a service for Chinese text classification. We develop a dedicated neural architecture based on the convolutional neural network to effectively learn character embedding and apply it to Chinese word similarity judgment and Chinese text classification. Both experiments results show that the stroke level method is effective for Chinese language processing.
Hang Zhuang, Chao Wang 0003, Changlong Li 0006, Qingfeng Wang 0004, Xuehai Zhou
ICWS3
2016 A Fast and Better Hybrid Recommender System Based on Spark
Jiali Wang 0003, Hang Zhuang, Changlong Li 0006, Zhuocheng He, Xuehai Zhou
NPC3