Kan Zhong

dblp:149/3972 · DBLP profile ↗
← Back
46ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0001-5278-0406ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 9 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 D ${ }^{2}$ Write: Accelerating Erasure-Coded Writes With Distributed Encoding and Decoupled Transmission
Canghai Yang, Wanyi Guo, Zhiwang Yu, Chaoxia Qin, Kan Zhong, Duo Liu 0002
ICDCS5
2026 Zero-Cost Merging Transitioning for Large-Scale Erasure-Coded Storage Systems
Canghai Yang, Kan Zhong, Zhiwang Yu, Wanyi Guo, Chaoxia Qin, Duo Liu 0002
IWQoS2
2026 RefineDedup: efficient deduplication for mobile systems via application-wise learning
Wei Li 0322, Xianzhang Chen, Xingjie Zhou, Duo Liu 0002, Yujuan Tan, Ao Ren, Kan Zhong, Lei Qiao 0002
Sci. China Inf. Sci.7
2026 Optimizing F2FS performance with the inter-zone parallelism in small-zone ZNS SSDs
Linbo Long, Xinrui Dong, Ting Wu 0012, Jingcheng Shen, Kan Zhong
Future Gener. Comput. Syst.5
2026 CD-ANN: Scalable Approximate Nearest Neighbor search on client-side devices
Chaoxia Qin, Yixiong Tang, Bing Guo 0003, Kan Zhong, Duo Liu 0002
J. Syst. Archit.4
2026 SAHChain: A Hybrid Storage Blockchain System Supporting Semantic Expressiveness and Retrieval
Chaoxia Qin, Duo Liu 0002, Bing Guo 0003, Yujuan Tan, Ao Ren, Kan Zhong, Liang Liang 0002
IEEE Trans. Computers6
2025 RAN: Accelerating Data Repair with Available Nodes in Erasure-Coded Storage
abstract
Distributed storage systems ensure data availability through fault-tolerant mechanisms, with erasure coding widely adopted for its low storage overhead. However, erasure coding generates significant repair traffic during data recovery, severely degrading performance. Recent repair algorithms aim to alleviate network bottlenecks at congested nodes, but they primarily address downlink bottlenecks while neglecting uplink constraints, which fundamentally limit repair efficiency. Furthermore, these algorithms lack a systematic approach for handling diverse failure scenarios, complicating recover implementation. In this paper, we propose RAN, an aggregation-based repair algorithm that alleviates both uplink and downlink bottlenecks by optimizing bandwidth utilization across all available nodes and aggregating network transfers via programmable network devices. Additionally, RAN systematically maximizes repair performance across diverse failure scenarios through a unified procedure. Experiments on Amazon EC2 show that RAN improves repair throughput by up to$\mathbf{6 8. 9 \%}$for degraded read and$\mathbf{2 6 6. 6 \%}$for full-node recovery compared to state-of-the-art algorithms.
Canghai Yang, Kan Zhong, Yujuan Tan, Ao Ren, Duo Liu 0002
CLUSTER2
2025 LIO-DPC: Accurate and Fast LiDAR-Inertial Odometry with Dynamic Pose Chain
abstract
LiDAR-inertial odometry is widely used in robotics navigation, autonomous driving, and drone operation to provide precise, low-latency motion estimation. Filter-based methods are fast but suffer from significant cumulative errors. Graph optimization methods reduce cumulative errors through loop closure detection but are computationally expensive. In this work, we propose LIO-DPC, a framework that combines the benefits of the filter-based approach and graph-based approach. First, we propose a dynamic pose chain optimization method. It generates an initial pose chain using the fast filter. This is followed by applying computationally efficient local graph optimization to a set of local pose chains to generate refined relative poses, which are then used to update the motion estimation. Second, we propose a loop sparsification approach to select representative loops that are both temporally and spatially proximate, to reduce the computational complexity in graph optimization and minimize loop errors. Extensive experiments demonstrate that LIO-DPC achieves real-time performance and outperforms state-of-the-art methods in accuracy.
Yuexin Mu, Ao Ren, Duo Liu 0002, Zihao Zhang 0002, Haojie Lu, Longyi Zhou, Huachen Tan, Kan Zhong, Yujuan Tan, Chaoxia Qin
DAC8
2025 CoSF: A Co-Optimization Framework for Operator Splitting and Fusion
Wei Li 0322, Ao Ren, Qingqiu Lan, Haining Fang, Zhenyu Wang 0002, Yujuan Tan, Kan Zhong, Duo Liu 0002
Euro-Par (1)7
2025 Cocache: An Accurate and Low-Overhead Dynamic Caching Method for GNNs
Zhaoyang Zeng, Yujuan Tan, Zhuoxin Bai, Kan Zhong, Duo Liu 0002, Ao Ren
Euro-Par (2)6
2025 MPNAS: Multimodal Sentiment Analysis Pruning via Neural Architecture Search
abstract
With the rapid development of social media, sentiment analysis from multimodal posts has garnered significant attention in recent years. However, the substantial size of these models impedes their deployment on resource-constrained embedded devices. Although pruning has been extensively studied to reduce the size of unimodal models, specific challenges remain for Multimodal Sentiment Analysis (MSA) models. First, existing techniques prune fixed original models into sparse models, while our findings indicate that different model architectures of identical size yield varying performance outcomes. Second, prior studies fail to explore the unique characteristics of MSA models, resulting in suboptimal pruning performance. To address these challenges, we propose MPNAS, a unified pruning framework via Neural Architecture Search (NAS) for MSA models. Specifically, we formulate pruning as a NAS problem and analyze MSA model characteristics to guide the subnet search. We conduct an initial coarse-grained NAS on the original model, expanding the search space slightly to identify suitable subnets that enhance pruning rates and accuracy. Subsequently, we refine coarse-grained subnets in a fine-grained NAS stage, where MSA model characteristics guide the search process. Extensive experiments on three representative datasets demonstrate the superiority of our approach over existing methods.
Binyan Zhang, Ao Ren, Zihao Zhang 0002, Moming Duan, Duo Liu 0002, Yujuan Tan, Kan Zhong
ICASSP7
2025 CAST: An Efficient Framework for Schedules Performance Prediction Based on Compact ASTs
abstract
With the advances of deep learning, efficient model inference is crucial. Deep learning compilers optimize inference by decomposing models into subgraphs and searching schedules for them, whose evaluation relies on accurate cost models. Existing methods suffer from high transformation overheads or limited prediction accuracy caused by insufficient structural representation of subgraphs and schedules. To address these limitations, we propose CAST, a framework that predicts schedule performance based on Abstract Syntax Trees (ASTs). CAST proposes AST classification based on structural similarity and class-specific cost models. Experiments show CAST achieves significantly reduced prediction errors and up to$13 \times$higher efficiency than prior methods.
Qingqiu Lan, Ao Ren, Zhenyu Wang 0002, Wei Li 0322, Hongbin Zhu, Yujuan Tan, Duo Liu 0002, Kan Zhong, Chaoxia Qin
ICCD8
2025 DualSpar: A Dual-Granularity Memory Framework with Adaptive Sparsity for Efficient LLM Inference
abstract
The block-based inference engine, powered by noncontiguous key-value (KV) cache management, has emerged as a new paradigm for large language model (LLM) inference due to its efficient memory utilization. However, in large-batch, longcontext workloads, the substantial demand for KV cache remains a major bottleneck in inference performance. Existing research leverages the sparsity of the attention mechanisms by removing non-critical tokens to limit the KV cache size. However, we observe that in block-based inference engines, current sparse methods require defragmentation after token removal to maintain tensor continuity, incurring significant overhead. Additionally, we find a correlation between request length and sparsity potential, yet existing methods apply a uniform sparsity strategy at batch level without dynamic adjustment. To address these, we propose DualSpar, a novel sparse KV cache framework with dual-granularity memory and adaptive sparsity strategy. First, it binds token importance to KV cache block granularity, achieving low-overhead pre-consolidation. Second, it incorporates system load and request length into sparsity decision, fully exploiting the sparse potential of different requests while reducing the queuing latency in large-batch processing. Evaluations show that DualSpar achieves up to a 3.16× throughput improvement, a 3.75× faster time-to-first-token (TTFT), and an 87.2% reduction in defragmentation overhead while maintaining high accuracy.
Yujuan Tan, Zhuoxin Bai, Sanle Zhao, Yujiao Wang, Zongjie Wang, Ao Ren, Kan Zhong
ICCD8
2025 Co-GNN: A Co-optimization Framework for Memory and Computation in Sampling-Based GNN Training
Yan Gan, Yujuan Tan, Yujiao Wang, Zongjie Wang, Duo Liu 0002, Ao Ren, Kan Zhong, Chaoxia Qin, Mingrui Qiang
ICIC (21)8
2025 RobTrack: A Robust 3D Multi-object Tracking Method for Edge Devices
Mingrui Qiang, Ao Ren, Yujuan Tan, Jing Yu 0026, Zhuoxin Bai, Duo Liu 0002, Kan Zhong, Chaoxia Qin
ICIC (5)7
2025 FASP: A Fast and Accurate Framework for Schedule Performance Evaluation
abstract
With the widespread application of deep neural networks, improving inference efficiency has become increasingly critical. To speed up the inference, deep learning compilers search for high-performance schedules for the DNN tensor programs. During the process, cost models have been extensively studied to evaluate the performance of the schedules, such that high-performance ones can be efficiently obtained. However, existing methods suffer from either high overhead or low accuracy of performance evaluation, both of which limit the efficiency of the final schedule. To address these issues, we propose FASP, a fast and accurate framework for schedule performance evaluation, based on Abstract Syntax Trees (ASTs). First, we propose a redundancy-aware ASTs reduction method to generate compact ASTs for more accurate feature extraction. Second, we propose a feature extraction method based on compact ASTs, which extracts features by accounting for computation nodes, loop nodes, and their structural relationships. Third, we propose a composition-similarity-driven ASTs classification method and a class-specific cost model architecture for more accurate performance evaluation. FASP overcomes the limitations of prior methods by significantly reducing evaluation errors. Experiments show its excellent performance in both single-model and cross-model evaluation, with errors ranging from 6 % to$\mathbf{1 3 \%}$. Moreover, FASP can obtain high-performance schedules with$13 \times$lower latency.
Qingqiu Lan, Ao Ren, Zhenyu Wang 0002, Wei Li 0322, Hongbin Zhu, Yujuan Tan, Duo Liu 0002, Kan Zhong, Chaoxia Qin
ICPADS8
2025 PIM-IoT: Enabling hierarchical, heterogeneous, and agile Processing-in-Memory in IoT systems
Kan Zhong, Qiao Li 0001, Ao Ren, Yujuan Tan, Xianzhang Chen, Linbo Long, Duo Liu 0002
Future Gener. Comput. Syst.1
2025 GNNBoost: Accelerating sampling-based GNN training on large scale graph by optimizing data preparation
Yujuan Tan, Yan Gan, Zhaoyang Zeng, Zhuoxin Bai, Lei Qiao 0002, Duo Liu 0002, Kan Zhong, Ao Ren
J. Syst. Archit.7
2025 Overlapping Aware Data Placement Optimizations for LSM Tree-Based Store on ZNS SSDs
abstract
Solid State Drives (SSDs) based on the NVMe Zoned Namespaces (ZNS) interface can notably reduce the costs of address mapping, garbage collection, and over-provisioning by dividing the storage space into multiple zones for sequential writes and random reads. The Log-Structured Merge (LSM) tree, which is extensively used in key-value storage systems, converts random writes to sequential writes, hence a suitable scenario to utilize ZNS SSDs. However, LSM tree associated data significantly varies in lifetime due to the levels and merging mechanisms of the LSM tree. Therefore, without an accurate method to estimate data lifetime, data with disparate lifetimes may be placed in the same zone, thus causing low space utilization and high write amplification within the SSD. To address these issues, the article proposes two data overlapping aware optimizations to realize intelligent data placement: a zone allocation scheme and a garbage collection scheme. The key technique of these optimizations is an accurate data-lifetime estimation by considering both the associated tree level of the data and the data overlapping ratio between the data and those in the neighboring level. Using the estimation technique, the zone allocation optimization can place data with similar lifetimes in the same zone. Besides, the garbage collection optimization can reclaim zones in an adaptive manner based on overlapping ratios to reduce the amount of data migration. Experimental results demonstrate that the optimization schemes effectively reduce garbage collection-incurred data copy by average factors of 2.11× and 1.50× in comparison to a conventional work and a state-of-the-art work, respectively. Consequently, the proposed work successfully alleviates the write amplification effect by 18% and 6%, compared to the conventional work and the state-of-the-art work, respectively.
Jingcheng Shen, Linbo Long, Zhenhua Tan, Congming Gao, Kan Zhong, Masao Okita, Fumihiko Ino
ACM Trans. Archit. Code Optim.6
2024 Rethinking Literary Plagiarism in LLMs through the Lens of Copyright Laws
Huachen Tan, Moming Duan, Duo Liu 0002, Haojie Lu, Yuexin Mu, Longyi Zhou, Ao Ren, Yujuan Tan, Kan Zhong
ACML9
2024 DPC: DPU-accelerated High-Performance File System Client
abstract
To achieve efficient file access to the file system backend, file system clients employ various intricate optimization techniques, such as local data/metadata caching and direct data access. However, these techniques impose a significant load on the host CPU, posing substantial challenges to the valuable CPU resources.
Kan Zhong, Zhiwang Yu, Qiao Li 0001, Xianqiang Luo, Linbo Long, Yujuan Tan, Ao Ren, Duo Liu 0002
ICPP1
2024 BGS: Accelerate GNN training on multiple GPUs
Yujuan Tan, Zhuoxin Bai, Duo Liu 0002, Zhaoyang Zeng, Yan Gan, Ao Ren, Xianzhang Chen, Kan Zhong
J. Syst. Archit.8
2024 WA-Zone: Wear-Aware Zone Management Optimization for LSM-Tree on ZNS SSDs
abstract
ZNS SSDs divide the storage space into sequential-write zones, reducing costs of DRAM utilization, garbage collection, and over-provisioning. The sequential-write feature of zones is well-suited for LSM-based databases, where random writes are organized into sequential writes to improve performance. However, the current compaction mechanism of LSM-tree results in widely varying access frequencies (i.e., hotness) of data and thus incurs an extreme imbalance in the distribution of erasure counts across zones. The imbalance significantly limits the lifetime of SSDs. Moreover, the current zone-reset method involves a large number of unnecessary erase operations on unused blocks, further shortening the SSD lifetime. Considering the access pattern of LSM-tree, this article proposes a wear-aware zone-management technique, termed WA-Zone , to effectively balance inter- and intra-zone wear in ZNS SSDs. In WA-Zone, a wear-aware zone allocator is first proposed to dynamically allocate data with different hotness to zones with corresponding lifetimes, enabling an even distribution of the erasure counts across zones. Then, a partial-erase-based zone-reset method is presented to avoid unnecessary erase operations. Furthermore, because the novel zone-reset method might lead to an unbalanced distribution of erasure counts across blocks in a zone, a wear-aware block allocator is proposed. Experimental results based on the FEMU emulator demonstrate the proposed WA-Zone enhances the ZNS-SSD lifetime by 5.23×, compared with the baseline scheme.
Linbo Long, Shuiyong He, Jingcheng Shen, Renping Liu 0002, Zhenhua Tan, Congming Gao, Duo Liu 0002, Kan Zhong
ACM Trans. Archit. Code Optim.8
2024 Optimizing Garbage Collection for ZNS SSDs via In-storage Data Migration and Address Remapping
abstract
The NVMe Zoned Namespace (ZNS) is a high-performance interface for flash-based solid-state drives (SSDs), which divides the logical address space into fixed-size and sequential-write zones. Meanwhile, ZNS SSDs eliminate in-device garbage collection (GC) by shifting the responsibility of GC to the host. However, the host-side GC of ZNS SSDs is not efficient. On the one hand, data migration during GC first moves data to the host buffer and then writes back the transferred data to the new location in the SSD, resulting in an unnecessary end-to-end transfer overhead. On the other hand, due to the pre-configured mapping between zones and blocks, GC incurs a large block-to-block rewrite overhead, i.e., even if most of the data in a block of the victim zone is valid, the valid data will still be rewritten to another block in the target zone. To address these issues, this article proposes a novel ZNS SSD design that features dynamic zone mapping, termed Brick-ZNS . Brick-ZNS implements two key functionalities: in-storage data migration and address remapping. New ZNS commands are first designed to realize in-storage data migration to avoid the end-to-end transfer overhead of GC while ensuring performance predictability. Then, a remapping strategy exploiting parallel physical blocks is proposed to reduce the large block-to-block rewrite overhead while ensuring zone-level access parallelism. The basic idea of the strategy is to directly remap the parallel physical blocks with a sufficient amount of valid data in the victim zone to the target zone, hence avoiding the large block-to-block rewrite overhead. Based on a full-stack SSD emulator, the evaluation results show that Brick-ZNS improves write throughput by 25% and SSD lifetime by 1.41×.
Zhenhua Tan, Linbo Long, Jingcheng Shen, Renping Liu 0002, Congming Gao, Kan Zhong
ACM Trans. Archit. Code Optim.6
2024 LightFS: A Lightweight Host-CSD Coordinated File System Optimizing for Heavy Small File Accesses
abstract
Computational storage drive (CSD) improves the data processing efficiency by processing the data within the storage. However, existing CSDs rely on the host-centric file systems to manage the data, where the layouts of files are retrieved by the host and sent to the CSD, resulting in additional I/O overhead and reduced processing efficiency, especially in heavy small file accesses. Moreover, the lack of consistency mechanisms poses potential consistency issues. To address these challenges, we propose LightFS, a lightweight host-CSD coordinated file system for the CSD file management. To reduce task offloading overhead, LightFS builds an index file$.ndpmeta$which summarizes the files’ metadata and shares between the host and CSD to enable CSD to retrieve the file layout in storage directly. To ensure consistency, LightFS employs a metadata locker and an update synchronizer. The metadata locker leverages the out-of-place update feature of the flash to capture a snapshot of the file to be written without any data copy, while the update synchronizer triggers metadata updates by monitoring the addresses of written blocks to ensure that the modified file is successfully written to the CSD. We implement and evaluate LightFS on a real testbed, and the results demonstrate that LightFS achieves$3.66\times $performance improvement on the average in real-world operations.
Zhaoyan Shen, Duo Liu 0002, Xianzhang Chen, Kan Zhong, Zhaoyang Zeng, Yujuan Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 FreePrune: An Automatic Pruning Framework Across Various Granularities Based on Training-Free Evaluation
abstract
Network pruning is an effective technique that reduces the computational costs of networks while maintaining accuracy. However, pruning requires expert knowledge and hyperparameter tuning, such as determining the pruning rate for each layer. Automatic pruning methods address this challenge by proposing an effective training-free metric to quickly evaluate the pruned network without fine-tuning. However, most existing automatic pruning methods only investigate a certain pruning granularity, and it remains unclear whether metrics benefit automatic pruning at different granularities. Neural architecture search also studies training-free metrics to accelerate network generation. Nevertheless, whether they apply to pruning needs further investigation. In this study, we first systematically analyze various advanced training-free metrics for various granularities in pruning, and then we investigate the correlation between the training-free metric score and the after-fine-tuned model accuracy. Based on the analysis, we proposed FreePrune score, a more general metric compatible with all pruning granularities. Aiming at generating high-quality pruned networks and unleashing the power of FreePrune score, we further propose FreePrune, an automatic framework that can rapidly generate and evaluate the candidate networks, leading to a final pruned network with both high accuracy and pruning rate. Experiments show that our method achieves high correlation on various pruning granularities and comprehensively improves the accuracy.
Ning Liu 0007, Haining Fang, Qiu Lin, Yujuan Tan, Xianzhang Chen, Duo Liu 0002, Kan Zhong, Ao Ren
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2023 Revisiting Swapping in User-Space With Lightweight Threading
abstract
Memory-intensive applications, such as in-memory databases, caching systems, and key-value stores, are increasingly demanding larger main memory to fit their working sets. Conventional swapping can enlarge the memory capacity by paging out inactive pages to backend stores. However, existing swapping solutions suffer several performance and compatibility issues, making them unsuitable for high-concurrency and memory-intensive applications. In this article, we redesign the swapping system and propose Lightswap, a high-performance user-space swapping solution that supports paging with both local SSDs and remote memories. First, to avoid kernel involvement, we propose to leverage the extended Berkeley packet filter (eBPF) for handling page faults (PFs) in user space and further eliminate the heavy I/O stack with the help of user-space I/O drivers. Then, we co-design the PF handling with lightweight thread (LWT) scheduling to improve system throughput and reduce the end-to-end PF latency. Finally, we propose a try-catch framework in Lightswap to deal with swap-in errors which have been exacerbated by the scaling in process technology. We implement Lightswap in our production-level system and evaluate it with various benchmarks. Results show that Lightswap achieves scalable PF notification latency ($4 \mu \text{s}$under 128 LWTs), reduces the PF handling latency by 3–5 times, and improves the throughput of memcached by more than 40% compared with the state-of-art swapping systems.
Kan Zhong, Wenlin Cui, Qiao Li 0001, Zhe Yang 0012, Youyou Lu, Xiaodan Yan, Siwei Luo, Qizhao Yuan, Keji Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Downsizing Without Downgrading: Approximated Dynamic Time Warping on Nonvolatile Memories
abstract
In recent years, time-series data have emerged in a variety of application domains, such as wireless sensor networks and surveillance systems. To identify the similarity between time-series data, the Euclidean distance and its variations are common metrics that quantify the differences between time-series data. However, the Euclidean distance is limited by its inability to elastically shift with the time axis, which motivates the development of dynamic time warping (DTW) algorithms. While DTW algorithms have been proven very useful in diversified applications like speech recognition, their efficacy might be seriously affected by the resolution of the time-series data. However, high-resolution time-series data might take up a gigantic amount of main memory and storage space, which will slow down the DTW analysis procedure. This makes the upscaling of DTW analysis more challenging, especially for in-memory data analytics platforms with limited nonvolatile memory space. In this paper, we propose a strategy to downsample time-series data to significantly reduce their size without seriously affecting the precision of the results obtained by DTW algorithms (downsizing without downgrading). In other words, this paper proposes a technique to remove the unimportant details that are largely ignored by DTW algorithms. The efficacy of the proposed technique is verified by a series of experimental studies, where the results are quite encouraging.
Duo Liu 0002, Xingni Li, Po-Chun Huang, Yingjian Ling, Kan Zhong, Renping Liu 0002, Xianzhang Chen, Liang Liang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2019 Towards Fast and Lightweight Checkpointing for Mobile Virtualization Using NVRAM
abstract
Checkpointing is a key enabler of hibernation, live migration and fault-tolerance for virtual machines (VMs) in mobile devices. However, checkpointing a VM is usually heavyweight: the VM's entire memory needs to be dumped to storage, which induces a significant amount of (slow) I/O operations, degrading system performance and user experience. In this paper, we propose FLIC, a fast and lightweight checkpointing machinery for virtualized mobile devices by taking advantages of recent byte-addressable, non-volatile memory (NVRAM). Instead of saving the VM's entire memory to storage, we store its working set pages in NVRAM, avoiding accessing slow flash memory (compared to server-grade SSDs). To further reduce the write activities to flash memory, we propose an energy-efficient data deduplication to eliminate redundant data in VM snapshot and save storage space. Experimental results based on an Exynos 5250 SoC show that our approach can effectively improve the performance of checkpointing in mobile virutalization and save energy.
Kan Zhong, Duo Liu 0002, Yunsong Wu, Linbo Long, Weichen Liu 0001, Jinting Ren, Renping Liu 0002, Liang Liang 0002, Zili Shao, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.1
2018 In-Situ AI: Towards Autonomous and Incremental Deep Learning for IoT Systems
abstract
Recent years have seen an exploration of data volumes from a myriad of IoT devices, such as various sensors and ubiquitous cameras. The deluge of IoT data creates enormous opportunities for us to explore the physical world, especially with the help of deep learning techniques. Traditionally, the Cloud is the option for deploying deep learning based applications. However, the challenges of Cloud-centric IoT systems are increasing due to significant data movement overhead, escalating energy needs, and privacy issues. Rather than constantly moving a tremendous amount of raw data to the Cloud, it would be beneficial to leverage the emerging powerful IoT devices to perform the inference task. Nevertheless, the statically trained model could not efficiently handle the dynamic data in the real in-situ environments, which leads to low accuracy. Moreover, the big raw IoT data challenges the traditional supervised training method in the Cloud. To tackle the above challenges, we propose In-situ AI, the first Autonomous and Incremental computing framework and architecture for deep learning based IoT applications. We equip deep learning based IoT system with autonomous IoT data diagnosis (minimize data movement), and incremental and unsupervised training method (tackle the big raw IoT data generated in ever-changing in-situ environments). To provide efficient architectural support for this new computing paradigm, we first characterize the two In-situ AI tasks (i.e. inference and diagnosis tasks) on two popular IoT devices (i.e. mobile GPU and FPGA) and explore the design space and tradeoffs. Based on the characterization results, we propose two working modes for the In-situ AI tasks, including Single-running and Co-running modes. Moreover, we craft analytical models for these two modes to guide the best configuration selection. We also develop a novel two-level weight shared In-situ AI architecture to efficiently deploy In-situ tasks to IoT node. Compared with traditional IoT systems, our In-situ AI can reduce data movement by 28-71%, which further yields 1.4X-3.3X speedup on model update and contributes to 30-70% energy saving.
Mingcong Song, Kan Zhong, Jiaqi Zhang 0002, Yang Hu 0001, Duo Liu 0002, Weigong Zhang, Jing Wang 0055, Tao Li 0006
HPCA2
2018 Towards Efficient Microarchitecture Design of Simultaneous Localization and Mapping in Augmented Reality Era
abstract
Recently, augmented reality technologies are debuting to the mainstream markets. Simultaneous Localization and Mapping (SLAM), which serves as the core to drive augmented reality, enables mobile devices to understand "the reality" by recognizing and understanding the surrounding space. However, enabling SLAM on mobile devices still faces challenges due to computation and power limitations. Thus, this paper explores the microarchitecture design space of visual SLAM. First, we conduct a characterization to make the image signal processing stage in the front-end image acquisition stage configurable and explore SLAM's sensitivity to individual stages. Our characterization shows that only demosaicing, gamma compression, and denoising in the image signal processing process have significant influences on SLAM's accuracy. Thus, the image signal processor in SoC in traditional mobile devices can be replaced by an approximation logic to save hardware overhead and energy. Second, we propose a new CGRA architecture, called SL-CGRA, specially tailored to SLAM's workload characteristics. It features a two-level memory design, which includes data-redirection layer support for on-chip memory, and an efficient CGRA memory controller stacked through Through-Via Silicon (TSV) for highly efficient off-chip memory access. Besides, PEs in SL-CGRA support different execution modes to exploit the data-level parallelism and task-level parallelism of SLAM. Our evaluation results show that SL-CGRA achieves good performance and energy efficiency.
Huixiang Chen 0001, Yuting Dai, Kan Zhong, Tao Li 0006
ICCD4
2017 SmartSwap: High-Performance and User Experience Friendly Swapping in Mobile Systems
abstract
With high-performance mobile processors and large main memory, smartphones are now integrated with more applications and richer functionality than ever. This poses larger memory and storage space demands, however, most mobile systems have limited memory space, which in turn affects user satisfaction. For example, application response time could become longer due to limited memory capacity. Swapping is an effective way to extend memory capacity, but often lead to poor performance in smartphones.
Duo Liu 0002, Kan Zhong, Jinting Ren, Tao Li 0006
DAC3
2017 System reliability evaluation considering parameter variations of a single-phase inverter with integrated active power decoupling
abstract
Electrolytic Capacitor (E-cap) is one of the lifetime bottlenecks in power electronic converters. In the last two decades, various active power decoupling circuits have been proposed to improve the reliability of the DC link by eliminating the DC-link E-caps. However, additional active devices could change the stresses of the existing converters, whether the system reliability is improving or not is still an open question. This paper investigates the reliability of the single-phase H-bridge inverter with active power decoupling circuit. The parameter variations in IGBT lifetime model are considered. The Weibull distribution of the key components is obtained from Monte Carlo analysis, and the reliability of the whole system is estimated by the system Reliability Block Diagram (RBD) method. As a case study, 2 kW single-phase H-bridge inverter with passive E-caps and active power decoupling circuits are presented. It is shown that the active power decoupling method is applied to H bridge inverter, and the lifetime of decoupling capacitor can be improved significantly, but it has different effects on system reliability in different applications. In addition, the difference on system reliability of fixed parameter and parameter variations is shown in conclusions.
Erjie Qi, Mingxuan Qi, Dingjun Zeng, Guorong Zhu, Kan Zhong
IECON7
2017 Revisiting swapping in mobile systems with SwapBench
Duo Liu 0002, Liang Liang 0002, Kan Zhong, Linbo Long, Meikang Qiu, Zili Shao, Edwin H.-M. Sha
Future Gener. Comput. Syst.4
2017 Non-Volatile Memory Based Page Swapping for Building High-Performance Mobile Devices
abstract
Smartphones are getting increasingly high-performance with advances in mobile processors and larger main memories to support feature-rich applications. However, the storage subsystem has always been a prohibitive factor that slows down the pace of reaching even higher performance while maintaining good user experience. Despite today's smartphones are equipped with larger-than-ever main memories, they consume more energy and still run out of memory. But the slow NAND flash based storage vetoes the possibility of swapping-an important technique to extend main memory-and leaves a system that constantly terminates user applications under memory pressure. In this paper, we propose NVM-Swap by revisiting swapping for smartphones with fast, byte-addressable, non-volatile memory (NVM) technologies. Instead of using flash, we build the swap area with NVM, to allow high performance without sacrificing user experience. NVM-Swap supports Lazy Swap-in, which can reduce memory copy operations by giving the swapped out pages a second chance to stay in byte-addressable NVM backed swap area. To avoid fast worn-out of certain NVM, we also propose Heap-Wear, a wear leveling algorithm that distributes writes in NVM more evenly. Evaluation results based on the Google Nexus 5 smartphone show that our solution can effectively enhance smartphone performance and achieve better wear-leveling of NVM.
Duo Liu 0002, Kan Zhong, Lingbo Long, Zili Shao
IEEE Trans. Computers2
2017 Durable Address Translation in PCM-Based Flash Storage Systems
abstract
Phase change memory (PCM) is a promising DRAM alternative because of its non-volatility, high density, low standby power and close-to-DRAM performance. These features make PCM an attractive solution to optimize the management of NAND flash memory in embedded systems. However, PCM's limited write endurance hinders its application in embedded systems. Therefore, how to manage flash memory with PCM-particularly guarantee PCM a reasonable lifetime-becomes a challenging issue. In this paper, we propose to partially replace DRAM using PCM to optimize the management of flash memory metadata for better system reliability in the presence of power failure and system crash. To prolong PCM's lifetime, we present a write-activity-aware PCM-assisted flash memory management scheme, called PCM-FTL. By differentiating sequential and random I/O behaviors, a novel two-level mapping mechanism and a customized wear-leveling scheme are developed to reduce writes to PCM and extend its lifetime. We evaluate PCM-FTL with a variety of general-purpose and mobile I/O workloads. Experimental results show that PCM-FTL can significantly reduce write activities and achieve an even distribution of writes in PCM with very low overhead.
Duo Liu 0002, Kan Zhong, Tianzheng Wang 0001, Yi Wang 0003, Zili Shao, Edwin H.-M. Sha, Jingling Xue
IEEE Trans. Parallel Distributed Syst.2
2017 Building NVRAM-Aware Swapping Through Code Migration in Mobile Devices
abstract
Mobile applications are becoming increasingly feature-rich and powerful, but also dependent on large main memories, which consume a large portion of system energy, especially for devices equipped with 4/6 GB DRAM. Swapping inactive DRAM pages to byte-addressable, non-volatile memory (NVRAM) is a promising solution to this problem. However, most NVRAMs have limited write endurance and the current victim pages selecting algorithm does not aware it. Therefore, to make it practical, the design of an NVRAM based swapping system must also consider endurance. In this paper, we target at prolonging the lifetime of NVRAM based swap area in mobile devices by reducing the write activities to NVRAM based swap area. Different from traditional wisdom, such as wear leveling and hot/cold data identification, we propose to build a system called nCode, which exploits the fact that code pages are easy to identify, read-only, and therefore a perfect candidate for swapping. Utilizing NVRAM's byte-addressability, we support execute-in-place (XIP) of the code pages in the swap area, without copying them back to DRAM based main memory. Experimental results based on the Google Nexus 5 smartphone show that nCode can effectively prolong the lifetime of NVRAM under various workloads.
Kan Zhong, Duo Liu 0002, Lingbo Long, Jinting Ren, Edwin H.-M. Sha
IEEE Trans. Parallel Distributed Syst.1
2016 FLIC: Fast, lightweight checkpointing for mobile virtualization using NVRAM
Kan Zhong, Duo Liu 0002, Liang Liang 0002, Linbo Long, Zili Shao
DATE1
2016 A compiler assisted wear leveling for morphable PCM in embedded systems
Linbo Long, Edwin H.-M. Sha, Duo Liu 0002, Liang Liang 0002, Kan Zhong
J. Syst. Archit.5
2016 Morphable Resistive Memory Optimization for Mobile Virtualization
abstract
Virtualization offers significant benefits, such as better isolation and security for mobile systems. However, the limited amount of memory and virtualization's memory-demanding nature make it challenging to virtualize mobile systems efficiently. In this paper, we utilize morphable resistive memories to design a high-performance mobile system with an extensible memory space. With morphable resistive memories, a simple and effective page management technique, Balloonfish, is proposed to convert the memory cell state between multilevel and single-level for achieving a balance between performance and memory space. First, an application-specific page allocation is proposed for managing morphable resistive memories in virtualized mobile systems. Besides, we use a balloon-style algorithm to balance memory allocation among multiple virtual machines. Our evaluation based on the Samsung Exynos 5250 system-on-chip with various real Android applications shows that our system achieves 28.63% performance improvement compared with the baseline scheme.
Linbo Long, Duo Liu 0002, Liang Liang 0002, Kan Zhong, Zili Shao, Edwin H.-M. Sha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2016 Energy-Efficient In-Memory Paging for Smartphones
abstract
Smartphones are becoming increasingly energy-hungry to support feature-rich applications, posing a lot of pressure on battery lifetime and making energy consumption a non-negligible issue. In particular, dynamic random access memory (DRAM)-based main memory subsystem is a major contributor to the energy consumption of mobile devices. In this paper, we propose direct read (DR). Swap, an energy-efficient in-memory paging design to reduce energy consumption in smartphones. In DR. Swap, we adopt emerging energy-efficient nonvolatile memory (NVM) and use it as the swap area. Utilizing NVMs byte-addressability, we propose DR which guarantees zero memory copy for read-only requests when accessing a page in swap area. To better understand the energy consumption of swapping, we build an energy model to analyze the energy consumption of different paging architectures. We evaluate DR. Swap based on the Google Nexus 5 smartphone, experimental results show that our technique can reduce more than 50% energy consumption compared to DRAM backed swapping.
Kan Zhong, Duo Liu 0002, Liang Liang 0002, Linbo Long, Yi Wang 0003, Edwin H.-M. Sha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2015 Balloonfish: Utilizing morphable resistive memory in mobile virtualization
abstract
Virtualization offers significant benefits such as better isolation and security for mobile systems. However, the limited amount of memory and virtualization's memory-demanding nature makes it challenging to virtualize mobile systems efficiently. In this paper, we utilize morphable resistive memories to design a high-performance mobile system with extensible memory space. With morphable resistive memory, we convert the memory cell state between multi-level and single-level to achieve a balance between performance and memory space. Our evaluation based on the Samsung Exynos 5250 SoC with real Android applications shows that our system achieve 27% performance improvement compared with the baseline scheme.
Linbo Long, Duo Liu 0002, Kan Zhong, Zili Shao, Edwin H.-M. Sha
ASP-DAC4
2015 nCode: limiting harmful writes to emerging mobile NVRAM through code swapping
Kan Zhong, Duo Liu 0002, Linbo Long, Weichen Liu 0001, Qingfeng Zhuge, Edwin H.-M. Sha
DATE1
2014 Building high-performance smartphones via non-volatile memory: The swap approach
abstract
Smartphones are getting increasingly high-performance with advances in mobile processors and larger main memories to support feature-rich applications. However, the storage subsystem has always been a prohibitive factor that slows down the pace of reaching even higher performance while maintaining good user experience. Despite today's smartphones are equipped with larger-than-ever main memories, they consume more energy and still run out of memory. But the slow NAND flash based storage vetoes the possibility of swapping---an important technique to extend main memory---and leaves a system that constantly terminates user applications under memory pressure.
Kan Zhong, Tianzheng Wang 0001, Linbo Long, Duo Liu 0002, Weichen Liu 0001, Zili Shao, Edwin H.-M. Sha
EMSOFT1
2014 DR. Swap: energy-efficient paging for smartphones
abstract
Smartphones are becoming increasingly energy-hungry to support feature-rich applications, posing a lot of pressure on battery lifetime and making energy consumption a non-negligible issue. In particular, DRAM is among the most demanding components in energy consumption. In this paper, we propose DR. Swap, an energy-efficient paging design to reduce energy consumption in smartphones. We adopt emerging energy-efficient non-volatile memory (NVM) and use it as the swap area. Utilizing NVM's byte-addressability, we propose direct read which guarantees zero-copy for read-only pages in the swap area. Experimental results based on the Google Nexus 5 smartphone show that our technique can effectively reduce energy consumption.
Kan Zhong, Tianzheng Wang 0001, Dan Zhang 0011, Xianlu Luo, Duo Liu 0002, Weichen Liu 0001, Edwin H.-M. Sha
ISLPED1
2014 Enhancing lifetime of NVM-based main memory with bit shifting and flipping
abstract
Non-volatile memory (NVM) is considered as the most promising candidate of main memory due to many attractive properties, such as shock-resistivity, non-volatility, high density and near zero leakage power. However, the write endurance and high write energy consumption greatly limit its adoption in modern memory systems. In this paper, we propose a write reduction technique, called Min-Shift, to reduce the total number of writes to NVM. The basic idea is to re-encode the data to be written via bit shifting and flipping. This is motivated by the fact that NVM write operation takes more latency than read, and the energy cost of the value to be written varies a lot. The effectiveness of Min-Shift is verified by mathematical analysis and experiment. Experimental results show that the proposed technique can reduce the number of writes by 57.3% on average. The lifetime of NVM is 2.3× longer than before.
Xianlu Luo, Duo Liu 0002, Kan Zhong, Dan Zhang 0011, Weichen Liu 0001
RTCSA3