VLDB 2026 Research / reviewers in the wild / expert
Qing'an Li
dblp:63/8378 · also Qingan Li
· DBLP profile ↗
47ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0003-0110-5405ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 7 first-author · 10 since 2021Software engineering, systems software and programming languages · 15 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Security and privacy · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pyls: Enabling Python Hardware Synthesis with Dynamic Polymorphism via LCRS EncodingabstractPython dominates AI development and is the most widely used dynamic programming language, but synthesizing its polymorphic functions into hardware remains challenging. Existing HLS solutions support only static subsets of Python, forcing CPU offload with costly communication overhead. We present Pyls, the first framework that synthesizes dynamically polymorphic Python into monolithic hardware via Left-Child Right-Sibling (LCRS) encoding. Key to our approach is representing all Python objects as LCRS trees, enabling uniform hardware handling of dynamic types. Pyls automatically converts objects to fixed-width formats, generates XLS IR designs, and implements a tree memory architecture for efficient runtime type resolution. On FPGA platforms, Pyls demonstrates speedups of 5.19× and 3.98× over two ASIC CPUs, 303.29× over a soft-core processor, and 282.66× over a heterogeneous SoC design. Bolei Tong, Yongyan Fang, Chaorui Wang, Qing'an Li, Jingling Xue, Mengting Yuan 0001 |
CGO | 4 |
| 2026 | Accelerating App Recompilation across Android System Updates by Code ReusingabstractAndroid utilizes Ahead-of-Time (AOT) compilation technology to precompile applications and stores the compiled code in OAT files, thereby improving app performance. When the Android system is updated, the old OAT files become invalidated. Applications will fall back to interpreted execution, resulting in degraded performance. To accommodate the frequent updates to the Android system that commonly occur on a monthly basis for most smartphone manufacturers, apps must be frequently recompiled into OAT files to promptly restore optimal app performance. However, recompiling applications is a resource-consuming process that cannot be completed quickly. Users have to endure issues such as device overheating and lag, which are caused by performance degradation after system updates.This paper evaluated popular Android apps across different system updates and made an important observation: up to 99% of the compiled code can be reused across different system updates, rendering most existing recompilation efforts unnecessary. Based on this observation, this paper proposes a method to accelerate app recompilation across Android system updates by reusing the old OAT files. We evaluated the proposed method with eight popular apps, on ten open-source Android system pairs and one closed-source Android system pair provided by a smartphone manufacturer. These Android system pairs have the same Android Runtime (ART) version and execute AOT compilation in both speed and speed-profile modes. Experimental results show that the proposed method reuses approximately 95% of compiled methods, achieving average speedups of 2.12× in CPU time and 1.39× in wall-clock time in speed-profile mode. In speed mode, the proposed method reuses about 99% of compiled methods, achieving average speedups of 5.15× in CPU time and 2.80× in wall-clock time, respectively. The proposed method not only accelerates app recompilation but also generates OAT files identical to those generated by native AOT compilation, without introducing security issues. Therefore, it holds significant promise for real-world deployment and has the potential to enhance user experience by speeding up the generation of new OAT files for applications. Mengfei Xie, Futeng Yang, Jiang Ma, Jianming Fu, Chun Jason Xue, Qing'an Li |
CGO | 9 |
| 2026 | CMakeSonar: A Static Approach to Detecting CMake Bugs with a Fine-Grained Type SystemabstractAs build systems and their scripts grow in size and complexity, detecting bugs in build configurations becomes increasingly challenging due to the rich functionality and weak typing of build scripting languages. This paper introduces CM ake S onar , the first static approach to precisely identifying semantic bugs in CMake scripts. CM ake S onar addresses this challenge by (1) designing a fine-grained type system that captures the runtime semantics of CMake values, and (2) performing a flow-sensitive analysis that detects inconsistent and ill-typed value usages by solving type constraints. Our approach identifies configuration and usage errors that can silently affect build correctness, portability, and deployment safety. In our evaluation, CM ake S onar identifies 155 bugs across 36 real-world CMake projects on GitHub, of which 23 have been accepted and fixed by developers. With a false positive rate of 4.32 % and a recall of 97.48 % , CM ake S onar demonstrates that precise static analysis can effectively uncover high-impact bugs in untyped build systems. Haotian Han, Zihang Zhong, Qing'an Li, Jingling Xue, Mengting Yuan 0001 |
Proc. ACM Program. Lang. | 3 |
| 2025 | MTE4JNI: A Memory Tagging Method to Protect Java Heap Memory from Illicit Native Code AccessabstractWith the proliferation of mobile devices in daily life, ensuring the security and performance of these devices has become crucial. On Android, the Java Native Interface (JNI) acts as a bridge, allowing native libraries to directly access Java heap memory via raw pointers, bypassing Java's built-in safety checks. While this offers powerful functionality and performance, it also threatens the memory safety of the Java heap. Recently, Memory Tagging Extension (MTE) is introduced into the ARM architectures to enhance memory safety, reducing software vulnerabilities caused by illegal memory operations. This paper proposes MTE4JNI, an MTE-based JNI checking method, to protect Java heap memory from illicit native code access. Experimental results on real Android devices demonstrate that, compared to the currently employed guarded copy method, the proposed MTE4JNI method provides superior memory safety protection, while significantly reducing the runtime overhead on average by 11x and 27x for single-threaded and multi-threaded environments, respectively. Huinan Chen, Jiang Ma, Chun Jason Xue, Qing'an Li |
CGO | 4 |
| 2025 | Calibro: Compilation-Assisted Linking-Time Binary Code Outlining for Code Size Reduction in Android ApplicationsabstractRecent Android systems have employed pre-compilation technology to boost app launch speed and runtime performance. However, this generates large OAT files that over-consume scarce memory and storage resources in mobile devices. This paper conducts an evaluation of code redundancy in popular production android applications and observes that the code redundancy is up to 25%. To reduce the code size via redundancy elimination, this paper proposes Calibro, a compilation-assisted linking-time binary code outlining method. Calibro consists of two parts, the Compilation-Time code Outlining (CTO) and the Linking-Time Binary code Outlining (LTBO) with information collected at compilation- time. Additionally, a paralleled suffix tree method is proposed to reduce the building time overhead, and a hot function filtering method is proposed to effectively mitigate run-time performance degradation caused by code outlining. Experimental results show that compared to the baseline (the original AOSP version with all available code size optimization enabled), the proposed approach reduces code size in Android applications by more than 15.19% on average, with negligible runtime performance degradation and tolerable building time overhead. Hence the proposed code outlining approach is promising for production deployment. Hanming Sun, Wenhan Shang, Mengting Yuan 0001, Jingqin Fu, Jiang Ma, Chun Jason Xue, Qing'an Li |
CGO | 8 |
| 2025 | MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight QuantizationabstractSmall language models (SLMs) are gaining attention for their lower computational and memory needs while maintaining strong performance.However, efficiently deploying SLMs on resource-constrained devices remains a significant challenge.Post-training quantization (PTQ) is a widely used compression technique that reduces memory usage and inference computation, yet existing methods face challenges in inefficient bit-width allocation and insufficient fine-grained quantization adjustments, leading to suboptimal performance, particularly at lower bit-widths.To address these challenges, we propose multi-level weight quantization (MLWQ), which facilitates the efficient deployment of SLMs.Our method enables more effective bit-width allocation by jointly considering inter-layer loss and intralayer salience.Furthermore, we propose a finegrained partitioning of intra-layer salience to support the tweaking of quantization parameters within each group.Experimental results indicate that MLWQ achieves competitive performance compared to state-of-the-art methods, providing an effective approach for the efficient deployment of SLMs while maintaining model accuracy. Chun Hu, Shangyu Wu, Chun Jason Xue, Qing'an Li |
EMNLP | 6 |
| 2025 | Hardware Implementation of Modified Noisy Gradient Descent Bit-Flipping DecodersabstractRecently, modified noisy gradient descent bit flipping (MNGDBF) algorithms have been proposed to eliminate the Gaussian random generators required in the original noisy gradient descent bit flipping (NGDBF) algorithm for the decoding of low-density parity-check (LDPC) codes. In this paper, a platform has been established for the hardware implementation of the LDPC encoder, the Gaussian noise channel generator and the decoder using MNGDBF decoding algorithms based on the use of field programmable gate array (FPGA). With 6-bit quantization, results show that the error performance achieved by hardware experiments is similar to that obtained by floating-point computer simulations. Moreover, we propose a new MNGDBF algorithm, in which the flipping condition is determined with the help of the summation of all inversion functions. Experimental results show that the proposed algorithm has superior performance than the original MNGDBF algorithm. Qing'an Li, Wai Man Tam, Francis C. M. Lau 0002 |
ISCAS | 1 |
| 2025 | Accelerating graph substitutions in DNN optimization by heuristic algorithmsabstractAbstract Graph substitution is a key optimization technique used in deep learning frameworks. Traditional search-based methods are one way to address the problem of graph substitution. However, with the ongoing expansion of deep neural networks (DNNs), the exploration of their vast equivalent graph search space becomes increasingly time-consuming. In this paper, we propose two heuristic methods to accelerate the search process in graph substitution, offering a relatively novel direction compared to existing methods. The first method employs a Memory-Augmented heuristic to optimize computation graphs. To further enhance the efficiency of computation graph optimization, the second method uses the simulated annealing method. This method adds computation graphs with degraded performance into the candidate set with a certain probability. The experimental results show that without significant compromise of inference performance, these two methods can find graph substitutions delivering similar DNN computing performance compared to existing searching methods, while the overall searching time can be reduced from hours to seconds. The source code is available at https://github.com/hudevictor/MAS-SAS . Chun Hu, Yufan Huang, Mengting Yuan 0001, Qing'an Li |
Neural Process. Lett. | 6 |
| 2024 | More Apps, Faster Hot-Launch on Mobile Devices via Fore/Background-aware GC-Swap Co-designabstractFaster app launching is crucial for the user experience on mobile devices. Apps launched from a background cached state, called hot-launching, have much better performance than apps launched from scratch. To increase the number of hot-launches, leading mobile vendors now cache more apps in the background by enabling swap. Recent work also proposed reducing the Java heap to increase the number of cached apps. However, this paper found that existing methods deteriorate app hot-launch performance while increasing the number of cached apps. To simultaneously improve the number of cached apps and hot-launch performance, this paper proposes Fleet, a foreground/background-aware GC-swap co-design framework. To enhance app-caching capacity, Fleet limits the tracing range of GC to background objects only, avoiding touching long-lifetime foreground objects. To improve hot-launch performance, Fleet identifies objects that will be accessed during the next hot-launch and uses runtime information to guide the swap scheme in the OS. In addition, Fleet aggregates small objects with similar access patterns into the same pages to improve swap efficiency. We implemented Fleet in AOSP and evaluated its performance with different types of apps. Experimental results show that Fleet achieves a 1.59× faster hot-launch time and caches 1.21× more apps than Android. Jiacheng Huang 0002, Yunmo Zhang, Junqiao Qiu, Yu Liang 0004, Rachata Ausavarungnirun, Qing'an Li, Chun Jason Xue |
ASPLOS (3) | 6 |
| 2024 | CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective SparsificationabstractDeploying large language models (LLMs) on edge devices presents significant challenges due to the substantial computational overhead and memory requirements.Activation sparsification can mitigate these resource challenges by reducing the number of activated neurons during inference.Existing methods typically employ thresholding-based sparsification based on the statistics of activation tensors.However, they do not model the impact of activation sparsification on performance, resulting in suboptimal performance degradation.To address the limitations, this paper reformulates the activation sparsification problem to explicitly capture the relationship between activation sparsity and model performance.Then, this paper proposes CHESS , a general activation sparsification approach via CHannel-wise thrEsholding and Selective Sparsification.First, channel-wise thresholding assigns a unique threshold to each activation channel in the feed-forward network (FFN) layers.Then, selective sparsification involves applying thresholding-based activation sparsification to specific layers within the attention modules.Finally, we detail the implementation of sparse kernels to accelerate LLM inference.Experimental results demonstrate that the proposed CHESS achieves lower performance degradation over eight downstream tasks while activating fewer parameters than existing methods, thus speeding up the LLM inference by up to 1.27x. Shangyu Wu, Weidong Wen, Chun Jason Xue, Qing'an Li |
EMNLP | 5 |
| 2024 | Rolling bearing fault diagnosis method using time-frequency information integration and multi-scale TransFusion network
Zifei Xu, Kezhong Shi, Xiaohui Zhong, Zhiqiang Liao, Qing'an Li |
Knowl. Based Syst. | 9 |
| 2023 | Shared Dictionary Compression for Efficient Mobile Software DistributionabstractThe distribution of software to small devices, such as smartphones, requires significant resources in terms of network bandwidth, server storage, and energy consumption during the upload and download process. In addition, software distribution platforms like Apple's App Store and Google's Play Store impose restrictions on maximum software size, emphasizing the importance of code size reduction. While traditional compression tools can help minimize redundancies within files, this paper demonstrates the presence of considerable redundancies across different files that are not addressed by existing methods. To remedy this, we propose a novel approach to reduce code size by compressing instructions at the intermediate representation (IR) layer, which is supported as a distribution format by Apple. Our method identifies and extracts common sub-strings among instructions across all IR files, creating a shared dictionary. When combined with conventional compression tools for packaging, this technique achieves an average mobile software size reduction of 24.49% compared to using zip compression alone, thereby alleviating software distribution costs for small devices. Jinheng Li, Qiao Li 0001, Qing'an Li, Chun Jason Xue |
RTCSA | 3 |
| 2023 | Statistical Type Inference for Incomplete ProgramsabstractWe propose a novel two-stage approach, Stir, for inferring types in incomplete programs that may be ill-formed, where whole-program syntactic analysis often fails. In the first stage, Stir predicts a type tag for each token by using neural networks, and consequently, infers all the simple types in the program. In the second stage, Stir refines the complex types for the tokens with predicted complex type tags. Unlike existing machine-learning-based approaches, which solve type inference as a classification problem, Stir reduces it to a sequence-to-graph parsing problem. According to our experimental results, Stir achieves an accuracy of 97.37 % for simple types. By representing complex types as directed graphs (type graphs), Stir achieves a type similarity score of 77.36 % and 59.61 % for complex types and zero-shot complex types, respectively. Yaohui Peng, Qiongling Yang, Hanwen Guo, Qing'an Li, Jingling Xue, Mengting Yuan 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2023 | Effective Stack Wear Leveling for NVMabstractWith the rapid growth of data processed by computer systems, nonvolatile memory (NVM), represented by phase change memory (PCM), is regarded as a promising next-generation storage technology as it offers superior advantages over DRAM. However, PCM suffers from a severe write durability problem, leading to an extremely short lifespan under the uneven write patterns of real-world programs. We observe that loops are one of the primary causes of uneven writes on the stack. To alleviate this problem, we present Loop2Recursion, a compiler-assisted stack wear leveling technique that automatically transforms loops into recursive functions. In addition, we propose several optimizations to reduce the stack sizes and instruction counts of the generated recursive functions, two schemes to limit recursion depth, and selective loop transformation for cache-enabled architectures. Experimental results demonstrate that Loop2Recursion outperforms state-of-the-art methods by significantly improving stack wear leveling with a greatly reduced performance overhead. Jifeng Wu, Wei Li 0241, Mengting Yuan 0001, Chun Jason Xue, Jingling Xue, Qing'an Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Lamina: Low Overhead Wear Leveling for NVM with Bounded TailabstractEmerging non-volatile memory (NVM) has been considered as a promising candidate for the next generation memory architecture because of its excellent characteristics. However, the endurance of NVM is much lower than DRAM. Without additional wear management technology, its lifetime can be very short, which extremely limits the use of NVM. This paper observes that the tail wear with a very small percentage of extreme deviation significantly hurts the lifetime of NVM, which the existing methods do not effectively solve. We present Lamina to address the tail wear issue, in order to improve the lifetime of NVM. Lamina consists of two parts: bounded tail wear leveling (BTWL) and lightweight wear enhancement (LWE). BTWL is used to make the wear degree of all pages close to the average value and control the upper limit of tail wear. LWE improves the accuracy of BTWL by exploiting the locality to interpolate low-frequency sampling schemes in virtual memory space. Our experiments show that compared with the state-of-the-art methods, Lamina can significantly improve the lifetime of NVM with low overhead. Jiacheng Huang 0002, Min Peng 0002, Chun Jason Xue, Qing'an Li |
ASP-DAC | 5 |
| 2022 | Recovering Container Class Types in C++ BinariesabstractWe present TIARA, a novel approach to recovering container classes in c++ binaries. Given a variable address in a c++ binary, TIARA first applies a new type-relevant slicing algorithm incorporated with a decay function, TSLICE, to obtain an inter-procedural forward slice of instructions expressed as a CFG to summarize how the variable is used in the binary (as our primary contribution). TIARA then makes use of a GCN (Graph Convolutional Network) to learn and predict the container type for the variable (as our secondary contribution). According to our evaluation, TIARA can advance the state of the art in inferring commonly used container types in a set of eight large real-world COTS c++ binaries efficiently (in terms of the overall analysis time) and effectively (in terms of precision, recall and F1 score). Xuezheng Xu, Qing'an Li, Mengting Yuan 0001, Jingling Xue |
CGO | 3 |
| 2022 | GLite: a fast and efficient automatic graph-level optimizer for large-scale DNNsabstractWe propose a scalable graph-level optimizer named GLite to speed up search-based optimizations on large neural networks. GLite leverages a potential-based partitioning strategy to partition large computation graphs into small subgraphs without losing profitable substitution patterns. To avoid redundant subgraph matching, we propose a dynamic programming algorithm to reuse explored matching patterns. The experimental results show that GLite reduces the running time of search-based optimizations from hours to milliseconds, without compromising in inference performance. Jiaqi Li 0006, Min Peng 0002, Qing'an Li, Meizheng Peng, Mengting Yuan 0001 |
DAC | 3 |
| 2022 | An empirical study of the effectiveness of IR-based bug localization for large-scale industrial projects
Wei Li 0241, Qing'an Li, Yunlong Ming, Weijiao Dai, Shi Ying 0001, Mengting Yuan 0001 |
Empir. Softw. Eng. | 2 |
| 2022 | A mobile edge computing-based applications execution framework for Internet of Vehicles
Rui Zhang 0083, Qing'an Li, Chao Ma 0008, Xiaochuan Shi |
Frontiers Comput. Sci. | 3 |
| 2022 | A spatio-temporal sequence-to-sequence network for traffic flow prediction
Shuqin Cao, Jia Wu 0001, Dan Wu 0006, Qing'an Li |
Inf. Sci. | 5 |
| 2022 | A Fully Authenticated Diffie-Hellman Protocol and Its Application in WSNsabstractThe secure authenticated key establishment between nodes in Wireless Sensor Networks (WSNs) has not been fully solved in the existing schemes. It’s a good idea to apply the Diffie-Hellman protocol to address it perfectly, but the existing authenticated Diffie-Hellman (ADH) protocols are not perfect because their authentication are partial or delayed. In this paper, we first present a concept of full authentication and propose a new fully authenticated Diffie-Hellman (FADH) prototype with light-certificate-based authentication. And then based on the theory of elliptic curve cryptography, we construct the TinyADH (Tiny Authenticated Diffie-Hellman) protocol with applying the FADH in WSNs. Compared with the existing similar solutions, TinyADH has lower communication overload, is easier to implement into existing standards, and more secure under equivalent computational complexity. The experimental results show that using this scheme for a successful key agreement between two nodes averagely takes about 54 seconds on TelosB. Moreover, the simulation results indicate that repeated key agreement can improve the secure connectivity rate. However, considering the cost performance ratio, it is advisable to take 2 runs of the negotiation. Fajun Sun, Selena He, Jun Zhang 0058, Qing'an Li, Yanxiang He |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Task Offloading with Task Classification and Offloading Nodes Selection for MEC-Enabled IoVabstractThe Mobile Edge Computing (MEC)-based task offloading in the Internet of Vehicles (IoV) scenario, which transfers computational tasks to mobile edge nodes and fixed edge nodes with available computing resources, has attracted interest in recent years. The MEC-based task offloading can achieve low latency and low operational cost under the tasks delay constraints. However, most existing research generally focuses on how to divide and migrate these tasks to the other devices. This research ignores delay constraints and offloading node selection for different tasks. In this article, we design the MEC-enabled IoV architecture, in which all vehicles and MEC servers act as offloading nodes. Mobile offloading nodes (i.e., vehicles) and fixed offloading nodes (i.e., MEC servers) provide low latency offloading services cooperatively through roadside units. Then we propose the task offloading scheme that considers task classification and offloading nodes selection (TO-TCONS). Our goal is to minimize the total execution time of tasks. In TO-TCONS Scheme, we divide the task offloading into the same region offloading mode and cross-region offloading mode, which is based on the delay constraints of tasks and the travel time of the target vehicle. Moreover, we propose the mobile offloading nodes selection strategy to select offloading nodes for each task, which evaluates offloading candidates for each task based on computing resources and transmission rates. Simulation results demonstrate that TO-TCONS Scheme is indeed capable of reducing total latency of tasks execution under the delay constraints in MEC-enabled IoV. Rui Zhang 0083, Shuqin Cao, Xinrong Hu, Shan Xue 0001, Dan Wu 0006, Qing'an Li |
ACM Trans. Internet Techn. | 7 |
| 2021 | A Spatial-Temporal Graph Attention Network for Multi-intersection Traffic Light ControlabstractTraffic light control is an extremely challenging problem in transportation. Recently, an increasing number of studies have employed reinforcement learning approaches to deal with traffic light control problems. However, these studies often observe the target intersection independently and ignore the dynamic effects from the surrounding intersections. Moreover, the dependence of historical information on current traffic conditions is also not fully exploited in multi-intersection traffic light control. In this work, we propose a spatial-temporal graph attention network based multi-agent deep Q-learning model. Our model can capture not only spatial features but also temporal features. Specifically, we first utilize the Graph Attention Network (GAT) to obtain the spatial relationship between the central intersection and surrounding intersections. Then Long Short-Term Memory (LSTM) network and the self-attention mechanism are used to get the time dimension information. Finally, we adopt a deep Q-learning network (DQN) to predict the action-value of each feasible action and choose the signal phase according to the size of the action-value. Experimental results on synthetic and realworld datasets confirm that our method can lead to both the least travel time and the maximum throughput, compared with some existing baselines. Qing'an Li, Min Wang 0017, Jianxin Li 0001, Dan Wu 0006 |
IJCNN | 2 |
| 2020 | Loop2Recursion: Compiler-Assisted Wear Leveling for Non-Volatile MemoryabstractNon-Volatile Memory (NVM) technologies, such as Phase Change Memory (PCM), herald the next generation of main memory as they offer superior features compared with DRAM. Unfortunately, NVM's limited write endurance hinders its adoption as its lifetime can be extremely short under skew writes. This paper observes that the loops in programs are one of the primary causes of uneven writes as they introduce the hot data and cause a large number of stack frames to be allocated to the same locations. To alleviate this problem, we present Loop2Recursion, a compile-time wear leveling technique for transforming loops into recursions automatically. Our approach is flexible as it can avoid a substantial memory overhead by limiting the depth of recursion. Experimental results demonstrate that Loop2Recursion can significantly improve the wear leveling over stack area compared to the state-of-the-art methods, while incurring only negligible performance overhead. Wei Li 0241, Mengting Yuan 0001, Chun Jason Xue, Jingling Xue, Qing'an Li |
ICCD | 6 |
| 2019 | A Wear Leveling Aware Memory Allocator for Both Stack and Heap Management in PCM-based Main Memory SystemsabstractPhase change memory (PCM) has been considered as a replacement of DRAM, due to its potentials in high storage density and low leakage power. However, the limited write endurance presents critical challenges. Various wear leveling techniques have been proposed to mitigate this issue from different perspectives, including both hardware and software levels. This paper proposes a wear leveling aware memory allocator, which (1) always prefers allocating memory blocks with less writes upon memory requests, and (2) leaves blocks allocated more than a threshold value unallocable temporarily. Furthermore, for the first time, this allocator provides a uniform management scheme for both stack and heap areas, thus could better balance writes in stack and heap areas. Experimental evaluations show that, compared to state-of-the-art memory allocators (i.e., glibc malloc, NVMalloc and Walloc), the proposed memory allocator improves the PCM wear leveling, in terms of CoV (a wear leveling indicator) by 41.9%, 30.3%, and 35.8%, respectively. Wei Li 0241, Ziqi Shuai, Chun Jason Xue, Mengting Yuan 0001, Qing'an Li |
DATE | 5 |
| 2017 | Stack-Size Sensitive On-Chip Memory Backup for Self-Powered Nonvolatile ProcessorsabstractWearable devices gain increasing popularity since they can collect important information for healthcare and well-being purposes. Compared with battery, energy harvesting is a better power source for these wearable devices due to many advantages. However, harvested energy is naturally unstable and program execution will be interrupted frequently. Nonvolatile processors demonstrate promising advantages to back up volatile state before the system energy is depleted. However, it also introduces non-negligible energy and area overhead. In this paper, we aim to reduce the amount of data that need to be backed up during a power failure. Based on the observation that stack size varies along program execution, we propose to analyze the application program and identify efficient backup positions, by which the stack content to back up can be significantly reduced. The evaluation results show an average of 45.7% reduction on nonvolatile stack size for stack backup, with 0.58% storage overhead. In the mean time, with the proposed schemes, the energy utilization and program forward progress can be greatly improved compared with instant backup. Mengying Zhao, Chenchen Fu, Qing'an Li, Mimi Xie, Yongpan Liu, Jingtong Hu, Zhiping Jia, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Write Mode Aware Loop Tiling for High Performance Low Power Volatile PCM in Embedded SystemsabstractArchitecting PCM, especially MLC PCM, as main memory for MCUs is a promising technique to replace conventional DRAM deployment. However, PCM/MLC PCM suffers from long write latency and large write energy. Recent work has proposed a compiler directed dual-write (CDDW) scheme to combat the drawbacks of PCM by adopting fast or slow mode for different write operations. For large-scale loops, we observe that write instances' lifetime is very long and can only be written by the expensive slow mode. This paper proposes a write mode aware loop tiling approach to effectively reduce the lifetime of write instances and maximize the number of efficient fast writes in loops. The experimental results show that the proposed approach improves performance by 50.8 percent and reduces dynamic energy by 32.0 percent across a set of benchmarks compared to the CDDW approach on average. Keni Qiu, Qing'an Li, Jingtong Hu, Weigong Zhang, Chun Jason Xue |
IEEE Trans. Computers | 2 |
| 2015 | Compiler directed automatic stack trimming for efficient non-volatile processorsabstractWearable devices are becoming increasingly important in our daily lives. Energy harvesting instead of battery is a better power source for these wearable devices due to many advantages. However, harvested energy is often unstable and program execution will be frequently interrupted. Non-volatile processors demonstrate promising advantages to back up volatile state before the system energy is depleted. But Non-volatile processors require additional memory for backing up, thus introducing non-negligible overhead in terms of energy, runtime as well as chip area. In this work, we target at non-volatile register reduction for energy harvesting based wearable devices. This paper proposes to stack trimming the memory footprint via a novel compiler directed method. The evaluation results deliver on average 28.6% reduction of non-volatile register files for backing up stack area, with ultra low runtime overhead. Qing'an Li, Mengying Zhao, Jingtong Hu, Yongpan Liu, Yanxiang He, Chun Jason Xue |
DAC | 1 |
| 2015 | Software assisted non-volatile register reduction for energy harvesting based cyber-physical system
Mengying Zhao, Qing'an Li, Mimi Xie, Yongpan Liu, Jingtong Hu, Chun Jason Xue |
DATE | 2 |
| 2015 | Compiler-Assisted Refresh Minimization for Volatile STT-RAM CacheabstractSpin-transfer torque RAM (STT-RAM) has been proposed to build on-chip caches because of its attractive features such as high storage density and ultra low leakage power. However, long write latency and high write energy are the two challenges for STT-RAM. Recently, researchers propose to improve the write performance of STT-RAM by relaxing its non-volatility property. To avoid data losses resulting from volatility, refresh schemes have been proposed. However, refresh operations consume additional overhead. In this paper, we propose to significantly reduce the number of refresh operations through re-arranging program data layout at compilation time. An N-refresh scheme is also proposed to further reduce the number of refreshes. Experimental results show that, on average, the proposed methods can reduce the number of refresh operations by 84.2 percent, and reduce the dynamic energy consumption by 38.0 percent for volatile STT-RAM caches while incurring only 4.1 percent performance degradation. Qing'an Li, Yanxiang He, Jianhua Li 0003, Liang Shi 0001, Yiran Chen 0001, Chun Jason Xue |
IEEE Trans. Computers | 1 |
| 2014 | Write Mode Aware Loop Tiling for High Performance Low Power Volatile PCMabstractArchitecting PCM, especially MLC PCM, as main memory for MCUs is a promising technique to replace conventional DRAM deployment. However, PCM/MLC PCM suffers from long write latency and large write energy. Recent work has proposed a compiler directed dual-write (CDDW) scheme to combat the drawbacks of PCM by adopting fast or slow write mode for different write operations. We observe that write instances' lifetime is very long and can only be written by the expensive slow mode for large-scale loops. This paper proposes a write mode aware loop tiling approach to effectively reduce the lifetime of write instances and maximize the number of efficient fast writes in loops. The experimental results show that the proposed approach improves performance by 50.8% and reduces dynamic energy by 32.0% across a set of benchmarks compared to the CDDW approach on average. Keni Qiu, Qing'an Li, Chun Jason Xue |
DAC | 2 |
| 2014 | A wear-leveling-aware dynamic stack for PCM memory in embedded systemsabstractPhase Change Memory (PCM) is a promising DRAM replacement in embedded systems due to its attractive characteristics such as extremely low leakage power, high storage density and good scalability. However, PCM's low endurance constrains its practical applications. In this paper, we propose a wear leveling aware dynamic stack to extend PCM's lifetime when it is adopted in embedded systems as main memory. Through a dynamic stack, the memory space is circularly allocated to stack frames, and thus an even usage of PCM memory is achieved. The experimental results show that the proposed method can significantly reduce the write variation on PCM cells and enhance the lifetime of PCM memory. Qing'an Li, Yanxiang He, Chun Jason Xue |
DATE | 1 |
| 2014 | Register allocation for hybrid register architecture in nonvolatile processorsabstractNonvolatile processors (NVP) have been an emerging topic in recent years due to its zero standby power, data retention and instant-on features. The conventional full replacement architecture in NVP has drawbacks of large area overhead and high backup energy. This paper provides a partial replacement based hybrid register architecture to significantly abate above problems. However, the hybrid register architecture can induce potential critical data loss and backup errors. In this paper, we propose a critical-data overflow aware register allocation (CORA). Different from other register allocation methods, CORA efficiently reduces the possibility of critical data spilling and backup errors. The experiment results show that CORA reduces the critical data overflow rate by up to 52%. The hybrid register architecture reduces the chip area by 45.1% and backup energy by 82.8% when using CORA. Hongyang Jia, Yongpan Liu, Qing'an Li, Chun Jason Xue, Huazhong Yang |
ISCAS | 4 |
| 2014 | Migration-Aware Loop Retiming for STT-RAM-Based Hybrid Cache in Embedded SystemsabstractRecently hybrid cache architecture consisting of both spin-transfer torque RAM (STT-RAM) and SRAM has been proposed for energy efficiency. In hybrid caches, migration-based techniques have been proposed. A migration technique dynamically moves write-intensive and read-intensive data between STT-RAM and SRAM to explore the advantages of hybrid cache. Meanwhile, migrations also introduce extra reads and writes during data movements. For stencil loops with read and write data dependencies, we observe that migration overhead is significant, and migrations closely correlate to the interleaved read and write memory access pattern in a memory block. This paper proposes a loop retiming framework during compilation to reduce the migration overhead by changing the interleaved memory access pattern. With the proposed loop retiming technique, the interleaved memory accesses can be significantly reduced so that migration overhead is mitigated, and energy efficiency of hybrid cache is significantly improved. The experimental results have shown that, with the proposed methods, on average, the migration number is reduced up to 27.1% and the cache dynamic energy is reduced up to 14.0%. Keni Qiu, Mengying Zhao, Qing'an Li, Chenchen Fu, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | Thread Progress Aware Coherence Adaption for Hybrid Cache Coherence ProtocolsabstractFor chip multiprocessor systems (CMPs), the interference on shared resources such as on-chip caches typically leads to unbalanced progress among threads. Because of the inherent synchronization primitives, such as barriers and locks, cores running fast threads have to waste precious cycles to wait for cores with slow progress, which leads to performance and energy inefficiency. For the purpose of improving performance and reducing energy consumption, this paper proposes to adapt the cache coherence policy for threads according to their delay-tolerant levels. Specifically, this paper proposes Thread progrEss Aware Coherence Adaption (TEACA) which utilizes the thread progress information as hints for coherence adaption. TEACA dynamically utilize the memory system statistics to estimate the progress of threads. Based on the estimated thread progress information, TEACA categorizes threads into leader threads and laggard threads. The thread categorization decisions are then leveraged for efficient coherence adaption on CMP systems supporting hybrid coherence protocols. Experimental results show that, on a 64-core CMP system, TEACA outperforms directory protocol in application execution time and a recently proposed hybrid protocol in both application execution time and energy dissipation. Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yinlong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | WCET-Aware Re-Scheduling Register Allocation for Real-Time Embedded Systems With Clustered VLIW ArchitectureabstractWorst-case execution time (WCET) is one of the most important metric in real-time embedded system design. For embedded systems with clustered very long instruction word (VLIW) architecture, register allocation, instruction scheduling, and cluster assignment are three key activities for code optimization, which have profound impact on WCET. At the same time, these three activities exhibit a phase ordering problem, i.e., independently performing register allocation, scheduling, and cluster assignment could have a negative effect on the other phases, thereby generating sub-optimal compiled code. In this paper, a compiler level optimization, namely WCET-aware re-scheduling register allocation, is proposed to achieve WCET minimization for real-time embedded systems with clustered VLIW architecture. The novelty of the proposed approach is that the effects of register allocation, instruction scheduling, and cluster assignment on the quality of generated code are taken into account for WCET minimization. These three compilation processes are integrated into a single phase to obtain a balanced result. The proposed technique is implemented in Trimaran 4.0. The experimental results show that the proposed technique can reduce WCET effectively, by 34% on average. Yazhi Huang, Liang Shi 0001, Jianhua Li 0003, Qing'an Li, Chun Jason Xue |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | Compiler-Assisted STT-RAM-Based Hybrid Cache for Energy Efficient Embedded SystemsabstractHybrid caches consisting of static RAM (SRAM) and spin-torque transfer (STT)-RAM have been proposed recently for energy efficiency. To explore the advantages of hybrid cache, most of the management strategies for hybrid caches employ migration-based techniques to dynamically move write-intensive data from STT-RAM to SRAM. These techniques involve additional access operations, and thus lead to extra overheads. In this paper, we propose two compilation-based approaches to improve the energy efficiency and performance of STT-RAM-based hybrid cache by reducing the migration overheads. The first approach, migration-aware data layout, is proposed to reduce the migrations by rearranging the data layout. The second approach, migration-aware cache locking, is proposed to reduce the migrations by locking migration-intensive memory blocks into SRAM part of hybrid cache. Furthermore, experiments show that these two methods can be combined to reduce more migrations. The reduction of migration overheads can improve the energy efficiency and performance of STT-RAM-based hybrid cache. Experimental results show that, combining these two methods, on average, the number of write operations on STT-RAM is reduced by 17.6%, the number of migrations is reduced by 38.9%, the total dynamic energy is reduced by 15.6%, and the total access latency is reduced by 13.8%. Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Mengying Zhao, Chun Jason Xue, Yanxiang He |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | A Unified Write Buffer Cache Management Scheme for Flash MemoryabstractNAND flash memory has been widely adopted in embedded systems as secondary storage. However, the further development of flash memory strongly hinges on the tackling of its inherent implausible characteristics, including read-and-write speed asymmetry, inability of in-place updates, and performance-harmful erase operations. While write buffer cache (WBC) has been proposed to enhance the performance of write operations, the development of a unified WBC management scheme that is effective for diverse types of access patterns is still a challenging task. In this paper, a novel WBC management scheme named expectation-based least recently used (ExLRU) is proposed to improve the performance of flash memory through effectively reducing the number of erase operations and write activities. Different from the previous works, ExLRU accurately maintains access history information in the WBC, based on which a novel cost model is constructed to select data with the minimum write cost to write to flash memory. An efficient ExLRU implementation with negligible overhead is developed. Simulation results show that ExLRU outperforms state-of-the-art WBC management schemes under various workloads. Liang Shi 0001, Jianhua Li 0003, Qing'an Li, Chun Jason Xue, Chengmo Yang, Xuehai Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Compiler-assisted refresh minimization for volatile STT-RAM cacheabstractSpin-Transfer Torque RAM (STT-RAM) has been proposed to build on-chip caches because of its attractive features: high storage density and negligible leakage power. Recently, researchers propose to improve the write performance of STT-RAM by relaxing its non-volatility property. To avoid data loss resulting from volatility, refresh schemes are proposed. However, refresh operations consume additional energy. In this paper, we propose to reduce the number of refresh operations through re-arranging program data layout at compilation time. An N-refresh scheme is also proposed. Experimental results show that, on average, the proposedmethods can reduce the number of refresh operations by 73.3%, and reduce the dynamic energy consumption by 27.6%. Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Chun Jason Xue, Yiran Chen 0001, Yanxiang He |
ASP-DAC | 1 |
| 2013 | Minimizing code size via page selection optimization on partitioned memory architecturesabstractFor 8-bit microcontrollers, bank-switching is commonly used to increase memory capacity. The disadvantage of this technique is that bank (page) selection instructions are introduced when switching active data (program) bank. The page selection problem is to minimize the number of page selection instructions inserted. While previous efforts work on optimizing bank selection instructions for the data segment, our work focuses on minimizing page selection instructions for the program segment. Minimizing page selection instructions is a more challenging problem as the size of each procedure being allocated is affected by the number of inserted page selection instructions. In this paper, we first give a formal definition of the page selection problem, and then we formulate the problem as an Integer Linear Programming (ILP) to find the optimal solution. We introduce a tabu search heuristic algorithm, TMSEARCH, to solve the problem efficiently. The experimental results show that ILP can find optimal solutions for small-scale problems, and TMSEARCH is able to find good solutions for all benchmarks within reasonable time. Com-pared to a commercial compiler, TMSEARCH reduces total code size between 0.04% and 19.3%, and reduces page selection instructions between 24.3% and 78.7%. Mengting Yuan 0001, Chun Jason Xue, Qing'an Li, Yingchao Zhao 0001 |
CASES | 4 |
| 2013 | Cache coherence enabled adaptive refresh for volatile STT-RAM
Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yiran Chen 0001, Yinlong Xu 0001 |
DATE | 3 |
| 2013 | Compiler directed write-mode selection for high performance low power volatile PCMabstractMicro-Controller Units (MCUs) are widely adopted ubiquitous computing devices. Due to tight cost and energy constraints, MCUs often integrate very limited internal RAM memory on top of Flash storage, which exposes Flash to heavy write traffic and results in short system lifetime. Architecting emerging Phase Change Memory (PCM) is a promising approach for MCUs due to its fast read speed and long write endurance. Qing'an Li, Lei Jiang 0001, Youtao Zhang, Yanxiang He, Chun Jason Xue |
LCTES | 1 |
| 2013 | Low-energy volatile STT-RAM cache design using cache-coherence-enabled adaptive refreshabstractSpin-Torque Transfer RAM (STT-RAM) is a promising candidate for SRAM replacement because of its excellent features, such as fast read access, high density, low leakage power, and CMOS technology compatibility. However, wide adoption of STT-RAM as cache memories is impeded by its long write latency and high write power. Recent work proposed improving the write performance through relaxing the retention time of STT-RAM cells. The resultant volatile STT-RAM needs to be periodically refreshed to prevent data loss. When volatile STT-RAM is applied as the last-level cache (LLC) in chip multiprocessor (CMP) systems, frequent refresh operations could dissipate significant extra energy. In addition, refresh operations could severely conflict with normal read/write operations to degrade overall system performance. Therefore, minimizing the performance impact caused by refresh operations is crucial for the adoption of volatile STT-RAM. In this article, we propose Cache-Coherence-Enabled Adaptive Refresh (CCear) to minimize the number of refresh operations for volatile STT-RAM, adopted as the LLC for CMP systems. Specifically, CCear interacts with cache coherence protocol and cache management policy to minimize the number of refresh operations on volatile STT-RAM caches. Full-system simulation results show that CCear performs close to an ideal refresh policy with low overhead. Compared with state-of-the-art refresh policies, CCear simultaneously improves the system performance and reduces the energy consumption. Moreover, the performance of CCear could be further enhanced using small filter caches to accommodate the not-refreshed private STT-RAM blocks. Jianhua Li 0003, Liang Shi 0001, Qing'an Li, Chun Jason Xue, Yiran Chen 0001, Yinlong Xu 0001, Wei Wang 0237 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2013 | Task Allocation on Nonvolatile-Memory-Based Hybrid Main MemoryabstractIn this paper, we consider the task allocation problem on a hybrid main memory composed of nonvolatile memory (NVM) and dynamic random access memory (DRAM). Compared to the conventional memory technology DRAM, the emerging NVM has excellent energy performance since it consumes orders of magnitude less leakage power. On the other hand, most types of NVMs come with the disadvantages of much shorter write endurance and longer write latency as opposed to DRAM. By leveraging the energy efficiency of NVM and long write endurance of DRAM, this paper explores task allocation techniques on hybrid memory for multiple objectives such as minimizing the energy consumption, extending the lifetime, and minimizing the memory size. The contributions of this paper are twofold. First, we design the integer linear programming (ILP) formulations that can solve different objectives optimally. Then, we propose two sets of heuristic algorithms including three polynomial time offline heuristics and three online heuristics. Experiments show that compared to the optimal solutions generated by the ILP formulations, the offline heuristics can produce near-optimal results. Wanyong Tian, Yingchao Zhao 0001, Liang Shi 0001, Qing'an Li, Jianhua Li 0003, Chun Jason Xue, Minming Li, Enhong Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | MAC: migration-aware compilation for STT-RAM based hybrid cache in embedded systemsabstractHybrid caches consisting of both STT-RAM and SRAM have been proposed recently for energy efficiency. To explore the advantages of hybrid cache, most work on hybrid caches employs migration based strategies to dynamically move write-intensive data from STT-RAM to SRAM. Migrations require additional read and write operations for data movement and may lead to significant overheads. To address this issue, this paper proposes a Migration-Aware Compilation (MAC) approach to improve the energy efficiency and performance of STT-RAM based hybrid cache. By re-arranging data layout, the data access pattern in memory blocks is changed such that the number of migrations is reduced without any hardware modification. The reduction of migration overheads in turn improves energy efficiency and performance. The experimental results show that with the proposed approach, on average, the number of write operations on STT-RAM is reduced by 13.4%, the number of migrations is reduced by 16.1%, the total dynamic energy is reduced by 8.5%, and the total latency is reduced by 12.1%. Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Chun Jason Xue, Yanxiang He |
ISLPED | 1 |
| 2012 | Compiler-assisted preferred caching for embedded systems with STT-RAM based hybrid cacheabstractAs technology scales down, energy consumption is becoming a big problem for traditional SRAM-based cache hierarchies. The emerging Spin-Torque Transfer RAM (STT-RAM) is a promising replacement for large on-chip cache due to its ultra low leakage power and high storage density. However, write operations on STT-RAM suffer from considerably higher energy consumption and longer latency than SRAM. Hybrid cache consisting of both SRAM and STT-RAM has been proposed recently for both performance and energy efficiency. Most management strategies for hybrid caches employ migration-based techniques to dynamically move write-intensive data from STT-RAM to SRAM. These techniques lead to extra overheads. In this paper, we propose a compiler-assisted approach, preferred caching, to significantly reduce the migration overhead by giving migration-intensive memory blocks the preference for the SRAM part of the hybrid cache. Furthermore, a data assignment technique is proposed to improve the efficiency of preferred caching. The reduction of migration overhead can in turn improve the performance and energy efficiency of STT-RAM based hybrid cache. The experimental results show that, with the proposed techniques, on average, the number of migrations is reduced by 21.3%, the total latency is reduced by 8.0% and the total dynamic energy is reduced by 10.8%. Qing'an Li, Mengying Zhao, Chun Jason Xue, Yanxiang He |
LCTES | 1 |
| 2011 | Minimizing Schedule Length via Cooperative Register Allocation and Loop Scheduling for Embedded SystemsabstractLoops are typically the most computation intensive sections for embedded applications. Therefore, it is important to minimize the overall schedule length for loops during the compilation process. Register allocation and instruction scheduling are two key activities during a compilation process. These two activities exhibit a phase ordering problem: Instruction scheduling before register allocation could lengthen live ranges of variables which will create more conflicts and more costly spills; Register allocation before scheduling may result in the same register assignment to two independent variables which will limit the choices available for scheduling. This paper proposes a cooperative re-scheduling register allocation technique for loops that combines these two critical stages together to minimize the schedule length. The novelty of the proposed approach is that the responsibility for balancing the phase ordering problem lies within the register allocator, which can re-schedule the aggressive initial scheduling to minimize the schedule length. Experimental results show that the proposed approach can reduce overall schedule length by 12% on average compared to previous techniques. Yazhi Huang, Qing'an Li, Chun Jason Xue |
TrustCom | 2 |