VLDB 2026 Research / reviewers in the wild / expert
Keni Qiu
dblp:132/8163
· DBLP profile ↗
39ranked-venue papers
15as first author
11since 2021 · last 2026
0000-0002-5851-777XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 14 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Parallel Serving for Vision-Language Models via Modal Decoupling and SchedulingabstractVision-Language Models (VLMs) have demonstrated strong performance in tasks such as image captioning and visual question answering. Under mixed workloads, however, the differing inference pipelines for text-only and multimodal requests create heterogeneity that existing serving systems fail to optimize—leading to high latency and poor fairness. We propose Duet-Infer, a modality-aware serving framework that enhances single-GPU serving efficiency for VLMs through three key contributions: (i) parallel computation enabled by preprocessing parallelism and decoupled vision-language execution, (ii) a shared memory manager that eliminates weight redundancy and supports efficient encoder cache sharing, and (iii) a fairness-aware scheduler that reduces delays for multimodal requests without penalizing text-only ones. Implemented within vLLM and evaluated on realistic workloads, DuetInfer reduces P99 TTFT by up to 33.7% and end-to-end latency by up to 20%. Yijia Yang, Yubo Deng, Yuanchao Xu 0002, Keni Qiu |
DATE | 5 |
| 2025 | MIVAS: Adaptive Residual Value Mining for Task Scheduling in Self-Powered SystemsabstractExisting schedulers for self-powered real-time systems either discard deadline-missing tasks, wasting their residual value, or harness it inefficiently. Furthermore, effective priority strategies for mixed (periodic/sporadic) tasks remain lacking. To address these issues, we propose a value model framework that captures value decay under energy fluctuations. Building on this framework, we present MIVAS, an energy-aware scheduler that integrates task urgency and value efficiency in priority assignment and strategically mines residual value from delayed tasks. Experiments show that MIVAS consistently achieves higher total value than the EA-EDF scheduler under different utilization levels. Xuejin Li, Keni Qiu |
CODES+ISSS | 2 |
| 2024 | Toward Energy-efficient STT-MRAM-based Near Memory Computing Architecture for Embedded SystemsabstractConvolutional Neural Networks (CNNs) have significantly impacted embedded system applications across various domains. However, this exacerbates the real-time processing and hardware resource-constrained challenges of embedded systems. To tackle these issues, we propose spin-transfer torque magnetic random-access memory (STT-MRAM)-based near memory computing (NMC) design for embedded systems. We optimize this design from three aspects: Fast-pipelined STT-MRAM readout scheme provides higher memory bandwidth for NMC design, enhancing real-time processing capability with a non-trivial area overhead. Direct index compression format in conjunction with digital sparse matrix-vector multiplication (SpMV) accelerator supports various matrices of practical applications that alleviate computing resource requirements. Custom NMC instructions and stream converter for NMC systems dynamically adjust available hardware resources for better utilization. Experimental results demonstrate that the memory bandwidth of STT-MRAM achieves 26.7 GB/s. Energy consumption and latency improvement of digital SpMV accelerator are up to 64× and 1,120× across sparsity matrices spanning from 10% to 99.8%. Single-precision and double-precision elements transmission increased up to 8× and 9.6×, respectively. Furthermore, our design achieves a throughput of up to 15.9× over state-of-the-art designs. Yueting Li 0001, He Zhang 0011, Biao Pan, Keni Qiu, Wang Kang 0001, Jun Wang 0041, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2024 | REC: REtime Convolutional Layers to Fully Exploit Harvested Energy for ReRAM-based CNN AcceleratorsabstractAs the Internet of Things (IoTs) increasingly combines AI technology, it is a trend to deploy neural network algorithms at edges and make IoT devices more intelligent than ever. Moreover, energy-harvesting technology-based IoT devices have shown the advantages of green and low-carbon economy, convenient maintenance, and theoretically infinite lifetime, and so on. However, the harvested energy is often unstable, resulting in low performance due to the fact that a fixed load cannot sufficiently utilize the harvested energy. To address this problem, recent works focusing on ReRAM-based convolutional neural networks (CNN) accelerators under harvested energy have proposed hardware/software optimizations. However, those works have overlooked the mismatch between the power requirement of different CNN layers and the variation of harvested power. Motivated by the above observation, this article proposes a novel strategy, called REC , that retimes convolutional layers of CNN inferences to improve the performance and energy efficiency of energy harvesting ReRAM-based accelerators. Specifically, at the offline stage, REC defines different power levels to fit the power requirements of different convolutional layers. At runtime, instead of sequentially executing the convolutional layers of an inference one by one, REC retimes the execution timeframe of different convolutional layers so as to accommodate different CNN layers to the changing power inputs. What is more, REC provides a parallel strategy to fully utilize very high power inputs. Moreover, a case study is presented to show that REC is effective to improve the real-time accomplishment of periodical critical inferences because REC provides an opportunity for critical inferences to preempt the process window with a high power supply. Our experimental results show that the proposed REC scheme achieves an average performance improvement of 6.1× (up to 16.5×) compared to the traditional strategy without the REC idea. The case study results show that the REC scheme can significantly improve the success rate of periodical critical inferences’ real-time accomplishment. Kunyu Zhou, Keni Qiu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | ResCheck: Resilient Checkpointing for Energy Harvesting SystemsabstractCheckpointing is a key technique to guarantee execution correctness and ensure progress forwarding in energy harvesting systems. However, checkpointing itself introduces system overhead due to extra operations of data movements between volatile memory and nonvolatile memory. Moreover, execution rollback to the latest checkpoint under a power failure can cancel some obtained progress and this waste is highly correlated to the latest checkpoint interval. These two kinds of overhead can be quite high if the checkpoint interval setting mismatches the power input characteristic. Unlike previous checkpointing schemes emphasizing on optimizing data copy overhead, this paper further takes into account the characteristic of input power sources and proposes a resilient checkpointing scheme, ResCheck, which can capture the power input changes and thus accordingly mitigate checkpoint overhead and rollback cost. The proposed ResCheck, directed by a lightweight neural network-based power level predictor, is capable of adjusting the checkpoint intervals to fit different power levels at runtime. In this way, an energy harvesting system equipped with ResCheck can achieve both fewer checkpoint number and lower execution rollback punishment. Our experimental results show that ResCheck can reduce an average checkpoint number of 24.4% and 14.6% over the conventional periodic checkpointing scheme and the state-of-the-art iCheck scheme respectively. Meanwhile, ResCheck improves average performance as well as energy efficiency by more than three times compared to iCheck. Keni Qiu, Chuting Xu, Kunyu Zhou, Dehui Qiu |
ICCD | 1 |
| 2023 | EagerReuse: An Efficient Memory Reuse Approach for Complex Computational GraphabstractMemory reuse is a promising approach for deep neural network (DNN) to reduce memory consumption because it does not introduce any additional runtime overhead. We observe that existing memory reuse algorithms consider only the effect of an individual data feature (either tensor size or tensor lifetime) on memory reuse and ignore the relative position relationship (RPR) among tensors. As computational graphs grow slightly more complex, the mining of memory reuse becomes insufficient. To address this issue, we propose a new memory reuse algorithm—EagerReuse, which can exploit more memory reuse opportunities by analyzing RPR among tensors and reusing them as quickly as possible. We evaluated the algorithms with inference models in TensorFlow Model Garden, and the results show that the EagerReuse outperforms the state-of-the-art algorithms in three out of seven cases. For more complex computational graphs, EagerReuse can achieve better memory usage with slightly higher but acceptable overhead. Ruyi Qian, Bojun Cao, Mengjuan Gao, Qinwen Shi, Yuanchao Xu 0002, Qirun Huo, Keni Qiu |
ICPADS | 8 |
| 2023 | Experimental Demonstration of STT-MRAM-based Nonvolatile Instantly On/Off System for IoT Applications: Case StudiesabstractEnergy consumption has been a big challenge for electronic devices, particularly for battery-powered Internet of Things (IoT) equipment. To address such a challenge, on the one hand, low-power electronic design methodologies and novel power management techniques have been proposed, such as nonvolatile memories and instantly on/off systems; on the other hand, the energy harvesting technology by collecting signals from human activity or the environment has attracted widespread attention in the IoT area. However, the system with self-powered energy harvesting may suffer frequent energy failures or fluctuating energy conditions, which degrade system reliability and user experience. Therefore, how to make the system under unreliable power inputs operate correctly and efficiently is one of the most critical issues for energy harvesting technology. In this article, we built an instantly on/off system based on nonvolatile STT-MRAM for IoT applications, which can instantly power on/off under different conditions of the harvested energy. The system powers on and operates normally when the harvested energy is enough (over the preset threshold); otherwise, the system powers off and stores the operational data back to the nonvolatile STT-MRAM. We described implementations of the hardware/software co-designed architecture (with image acquisition as an example) based on the commercialized 32 MB STT-MRAM, and we experimentally demonstrated the system functionality and efficiency under five typical energy harvesting scenarios, including radio frequency, thermal, solar, piezoelectric, and WIFI. Our experimental results show that the power consumption and data restore time were reduced by 15.1% and 714 times, respectively, in comparison with the DRAM-based counterpart. Yueting Li 0001, Wang Kang 0001, Kunyu Zhou, Keni Qiu, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | REC: REtime convolutional layers in energy harvesting ReRAM-based CNN acceleratorsabstractAs the Internet of Things (IoTs) increasingly combines AI technology, it is a trend to deploy neural network algorithms at edges and make IoT devices more intelligent than ever. Moreover, the energy harvesting technology-based IoT devices have shown the advantages of green economy, convenient maintenance, and theoretically infinite lifetime, etc. However, the harvested energy is often unstable, resulting in low performance due to the fact that a fixed load can't sufficiently utilize the harvested energy. To address this problem, recent works focusing on ReRAM-based convolutional neural networks (CNN) accelerators under harvested energy have proposed hardware/software optimizations. However, those works have overlooked the mismatch between the power requirement of different CNN layers and the variation of harvested power. Kunyu Zhou, Keni Qiu |
CF | 2 |
| 2022 | Smart scheduler: an adaptive NVM-aware thread scheduling approach on NUMA systems
Yuetao Chen, Keni Qiu, Haipeng Jia, Yunquan Zhang, Limin Xiao 0001, Lei Liu 0037 |
CCF Trans. High Perform. Comput. | 2 |
| 2022 | Publisher Correction: Smart scheduler: an adaptive NVM-aware thread scheduling approach on NUMA systems
Yuetao Chen, Keni Qiu, Haipeng Jia, Yunquan Zhang, Limin Xiao 0001, Lei Liu 0037 |
CCF Trans. High Perform. Comput. | 2 |
| 2021 | MaxTracker: Continuously Tracking the Maximum Computation Progress for Energy Harvesting ReRAM-based CNN AcceleratorsabstractThere is an ongoing trend to increasingly offload inference tasks, such as CNNs, to edge devices in many IoT scenarios. As energy harvesting is an attractive IoT power source, recent ReRAM-based CNN accelerators have been designed for operation on harvested energy. When addressing the instability problems of harvested energy, prior optimization techniques often assume that the load is fixed, overlooking the close interactions among input power, computational load, and circuit efficiency, or adapt the dynamic load to match the just-in-time incoming power under a simple harvesting architecture with no intermediate energy storage. Targeting a more efficient harvesting architecture equipped with both energy storage and energy delivery modules, this paper is the first effort to target whole system, end-to-end efficiency for an energy harvesting ReRAM-based accelerator. First, we model the relationships among ReRAM load power, DC-DC converter efficiency, and power failure overhead. Then, a maximum computation progress tracking scheme ( MaxTracker ) is proposed to achieve a joint optimization of the whole system by tuning the load power of the ReRAM-based accelerator. Specifically, MaxTracker accommodates both continuous and intermittent computing schemes and provides dynamic ReRAM load according to harvesting scenarios. We evaluate MaxTracker over four input power scenarios, and the experimental results show average speedups of 38.4%/40.3% (up to 51.3%/84.4%), over a full activation scheme (with energy storage) and order-of-magnitude speedups over the recently proposed (energy storage-less) ResiRCA technique. Furthermore, we also explore MaxTracker in combination with the Capybara reconfigurable capacitor approach to offer more flexible tuners and thus further boost the system performance. Keni Qiu, Nicholas Jao, Kunyu Zhou, Yongpan Liu, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2020 | Insights and Optimizations on IR-drop Induced Sneak-Path for RRAM Crossbar-based ConvolutionsabstractRRAM crossbar structure has been proposed to accelerate the convolution computation neural networks because its current-mode weighted summation operation intrinsically matches the dominant multiplication-and-accumulation (MAC) operations. However, there is an inevitable IR-drop problem with the RRAM crossbar, which may introduce sneak-path and thus reduce the accuracy of neural network algorithms and the system reliability. This work addresses the sneak-path problem caused by the IR-drop in a RRAM crossbar. We first present the characteristics of variation distribution of the sneak-path through numerous experiments, taking into account RRAM cell resistance, input voltage, and cell location in a crossbar. Then we propose optimization strategies from the hardware and software perspectives respectively to mitigate the variations resulting from sneak-path. The experimental results show that the proposed methods can compensate the accuracy of algorithms. Keni Qiu |
ASP-DAC | 3 |
| 2020 | Design Insights of Non-volatile Processors and Accelerators in Energy Harvesting SystemsabstractThere is growing interest in deploying energy harvesting processors and accelerators in Internet of Things (IoT). Energy harvesting harnesses the energy scavenged from the environment to power a system. Although it has many advantages over battery-operated systems such as lightweight, compact size, and no necessity of recharging and maintenance, it may suffer frequently power-down and a fluctuating power supply even with power on. Non-volatile processor (NVP) is a promising architecture for effective computing in energy harvesting scenarios. Recently, non-volatile accelerators (NVA) have been proposed to perform computations of deep learning algorithms. In this paper, we overview the recent studies of NVP and NVA across the layers of hardware, architecture, software and their co-design. Especially, we present the design insights of how the state-of-the-art works adapt their specific designs to the intermittent and fluctuating power conditions with the energy harvesting technology. Finally, we discuss recent trends using NVP and NVA in energy harvesting scenarios. Keni Qiu, Mengying Zhao, Zhenge Jia, Jingtong Hu, Chun Jason Xue, Kaisheng Ma, Xueqing Li 0002, Yongpan Liu, Narayanan Vijaykrishnan |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | ResiRCA: A Resilient Energy Harvesting ReRAM Crossbar-Based Accelerator for Intelligent Embedded ProcessorsabstractMany recent works have shown substantial efficiency boosts from performing inference tasks on Internet of Things (IoT) nodes rather than merely transmitting raw sensor data. However, such tasks, e.g., convolutional neural networks (CNNs), are very compute intensive. They are therefore challenging to complete at sensing-matched latencies in ultra-low-power and energy-harvesting IoT nodes. ReRAM crossbar-based accelerators (RCAs) are an ideal candidate to perform the dominant multiplication-and-accumulation (MAC) operations in CNNs efficiently, but conventional, performance-oriented RCAs, while energy-efficient, are power hungry and ill-optimized for the intermittent and unstable power supply of energy-harvesting IoT nodes. This paper presents the ResiRCA architecture that integrates a new, lightweight, and configurable RCA suitable for energy harvesting environments as an opportunistically executing augmentation to a baseline sense-and-transmit battery-powered IoT node. To maximize ResiRCA throughput under different power levels, we develop the ResiSchedule approach for dynamic RCA reconfiguration. The proposed approach uses loop tiling-based computation decomposition, model duplication within the RCA, and inter-layer pipelining to reduce RCA activation thresholds and more closely track execution costs with dynamic power income. Experimental results show that ResiRCA together with ResiSchedule achieve average speedups and energy efficiency improvements of 8× and 14× respectively compared to a baseline RCA with intermittency-unaware scheduling. Keni Qiu, Nicholas Jao, Mengying Zhao, Cyan Subhra Mishra, Gulsum Gudukbay Akbulut, Sethu Jose, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan |
HPCA | 1 |
| 2019 | Leveraging Energy Cycle Regularity to Predict Adaptive Mode for Non-volatile ProcessorsabstractAmbient energy harvesting technique is currently an ideal alternative to the state-of-the-art batteries for the power supply of IoT edge devices. Due to the intermittent power supply of the ambient energy, the systems suffer data loss and procedure rollbacks. NVPs have been proposed to ameliorate this problem through storing the volatile data into NVM when power fails and coping back them when power resumes. Recent studies have shown that NVPs can enter the retention mode when a power failure occurs so as to further mitigate the backup and recovery overheads through waiting for power resumption instead of immediate backup. However, the effectiveness of retention-based mechanism highly depends on energy prediction, which usually results in a complicated and slow mode decision process. If we use simple mode decision mechanism, the system may often enter inappropriate mode. In this work, we observe an interesting phenomenon that quite a few ambient energy waveforms exhibit regularity that the duration of the power outage in one energy cycle is quite akin to the adjacent ones. Addressing the mode decision issue upon power failures and exploiting the power regularity property, we build a fast history adaptive mechanism to accurately determine the backup or retention modes for a NVP system upon power dropping to a threshold. The metrics of energy cycle length, historical mode ratio and resumption time are defined to direct the proposed two-phase mode decision process. Experimental evaluations demonstrate the proposed prediction mechanism achieves up to 1.38X execution progress and up to 41.9% improvement on energy utilization over the conventional scheme. Zejun Shi, Dongqin Zhou, Keni Qiu, Jiwu Shu |
ASAP | 3 |
| 2019 | Checkpointing-Aware Loop Tiling for Energy Harvesting Powered Nonvolatile ProcessorsabstractAs power failures often occur in energy harvesting powered nonvolatile processors (NVPs), checkpointing is needed during program execution. It is observed that checkpointing is implemented with high overhead in applications with loops, because a large amount of data needs backup during loop execution. As such, we are motivated to reduce the amount of checkpointing data by analyzing data locality and shortening data lifetime in loops. This paper proposes a checkpointing-aware loop tiling technique which targets to reduce the checkpointing and recovering overheads for loops. Specifically, we first derive the optimal tile size for nested loops considering checkpointing distance and data dependencies. Then, the implementations of checkpointing and recovering for tiled loops are presented. Finally, the experiments are conducted to evaluate the effectiveness of the proposed method. The experimental results show that compared to the no-tiling method, the checkpointing-aware loop tiling method reduces the checkpointing and recovering data by 36.2% on average and reduces the total execution time and dynamic energy for checkpointing and recovering by 27.2% and 22.9% on average, respectively. Keni Qiu, Mengying Zhao, Jingtong Hu, Yongpan Liu, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Dual-threshold directed execution progress maximization for nonvolatile processorsabstractTo meet the needs of the Internet of Things (IoTs) devices, energy harvesting systems are proposed to power the systems instead of battery. Addressing the problem that harvested energy is unstable, nonvolatile processors (NVPs) have been proposed to hold intermediate data and avoid frequent program restarting from the beginning. However, NVPs often suffer a lot of waste on energy and system sources that can not be used for program execution owing to the frequent backup and recovery operations. To further improve the performance of NVPs, the paper proposes a dual-threshold method to maximize execution progress by enabling a system to hibernate to wait for power resumption instead of backing up data directly upon power interruptions. In particular, the optimal high and low thresholds, and the switches of system hibernation and backup, are discussed in details in order to achieve the goal of maximizing computation progress. The evaluation results show an average of up to 82.3% reduction on power failures and 1.5x speedup for forwarding progress by the proposed dual-threshold method compared to the conventional single threshold scheme. Dongqin Zhou, Keni Qiu, Yongpan Liu |
CF | 2 |
| 2018 | Power optimization through peripheral circuit reusing integrated with loop tiling for RRAM crossbar-based CNNabstractConvolutional neural networks (CNNs) have been proposed to be widely adopted to make predictions on a large amount of data in modern embedded systems. Prior studies have shown that convolutional computations which consist of numbers of multiply and accumulate (MAC) operations, serve as the most computationally expensive portion in CNN. Compared to the manner of executing MAC operations in GPU and FPGA, CNN implementation in the RRAM crossbar-based computing system (RCS) demonstrates the outstanding advantages of high performance and low power. However, the current design is energy-unbalanced among the three parts of RRAM crossbar computation, peripheral circuits and memory accesses, the latter two factors can significantly limit the potential gains of RCS. Addressing the problem of high power overhead of peripheral circuits in RCS, the Peripheral Circuit Unit (PeriCU)-Reuse scheme has been proposed to meet given power budget. In this paper, it is further observed that memory accesses can be bypassed if two adjacent layers are assigned in different PeriCUs. In this way, memory accesses can be reduced and thus the performance and power can be improved. A loop tiling technique is proposed to save memory accesses. The experiments of two convolutional applications validate that the proposed loop tiling technique can reduce energy consumption by 61.7%. Yuanhui Ni, Weiwen Chen, Wenjuan Cui, Yuanchun Zhou, Keni Qiu |
DATE | 5 |
| 2018 | A peripheral circuit reuse structure integrated with a retimed data flow for low power RRAM crossbar-based CNNabstractConvolutional computations implemented in RRAM crossbar-based Computing System (RCS) demonstrate the outstanding advantages of high performance and low power. However, current designs are energy-unbalanced among the three parts of RRAM crossbar computation, peripheral circuits and memory accesses, and the latter two factors can significantly limit the potential gains of RCS. Addressing the problem of high power overhead of peripheral circuits in RCS, this paper proposes a Peripheral Circuit Unit (PeriCU)-Reuse scheme to meet power budgets in energy constrained embedded systems. The underlying idea is to put the expensive ADCs/DACs onto spotlight and arrange multiple convolution layers to be sequentially served by the same PeriCU. In the solution, the first step is to determine the number of PeriCUs which are organized by cycle frames. Inside a cycle frame, the layers are computed in parallel inter-PeriCUs while sequentially intra-PeriCU. Furthermore, a layer retiming technique is exploited to further improve the energy of RCS by assigning two adjacent layers within the same PeriCU so as to bypass the energy consuming memory accesses. The experiments of five convolutional applications validate that the PeriCU-Reuse scheme integrated with the retiming technique can efficiently meet variable power budgets, and further reduce energy consumption efficiently. Keni Qiu, Weiwen Chen, Yuanchao Xu 0002, Lixue Xia, Yu Wang 0002, Zili Shao |
DATE | 1 |
| 2018 | Live Demonstration: A self-powered ultraviolet radiation monitoring platform based on nonvolatile processorabstractThis live demonstration shows a self-powered hardware platform for healthcare application on accurately monitoring ultraviolet(UV) radiation. Nonvolatile processor(NVP)[1] based sensor nodes harvest energy from solar panels and are connected to a Rohm ZigBee chip as the gateway node to support the network level communication. The gateway node finally uploads UV radiation data to a PC or a workstation for data analysis and display. Our platform provides two sensing modes (Data-First and Delay-First) under different UV radiation patterns to achieve high performance. Yongpan Liu, Yixiong Yang, Keni Qiu |
ISCAS | 5 |
| 2018 | Exporting Transactional Interface to Applications in Log-Structured File SystemsabstractThe following topics are dealt with: storage management; cloud computing; mobile computing; learning (artificial intelligence); optimisation; graph theory; cache storage; resource allocation; flash memories; telecommunication network routing. Youyou Lu, Keni Qiu, Zejun Shi, Hongsuk Choi, Jiwu Shu |
NAS | 3 |
| 2018 | Efficient energy management by exploiting retention state for self-powered nonvolatile processors
Keni Qiu, Zhiyao Gong, Dongqin Zhou, Weiwen Chen, Yuanchao Xu 0002, Yongpan Liu |
J. Syst. Archit. | 1 |
| 2017 | Expected Completion Time Aware Message Scheduling for UM-BUS Interconnected SystemabstractIn real-time embedded systems, periodic messages need to be transmitted at the expected time because of timing sensitive requirements. In this paper, we take advantage of the characteristics of UM-BUS, a novel serial bus with the capability of multi-lane concurrent transmissions, and investigate the scheduling problem to reduce the deviation to the expected completion time of messages. By configuring different lanes to change the bus utilization, two sets of experiments were implemented to evaluate the effectiveness of the proposed algorithm. The results show that the heuristic algorithm works effectively and can achieve a deviation within 1.52% which is significantly smaller comparing to the existing scheduling algorithms. Jiqin Zhou, Weigong Zhang, Keni Qiu, Ruiying Bai |
ISORC | 3 |
| 2017 | Retention state-enabled and progress-driven energy management for self-powered nonvolatile processorsabstractEnergy harvesting instead of battery is a better power source for wearable devices due to many advantages such as long operation time without maintenance and comfort to users. However, harvested energy is naturally unstable and program execution will be interrupted frequently. To solve this problem, nonvolatile processor (NVP) has been proposed because it can back up volatile state before the system energy is depleted. However, this backup process also introduces non-negligible energy and area overhead. To improve the performance of NVP, retention state has been proposed recently which can enable a system to retain the volatile data to wait for power resumption instead of saving data immediately. The goal of this paper is to forward program execution as much as possible by exploiting retention state. Specifically, two objectives are achieved. The first objective is to minimize power failures of the system if there is a great probability to get power resumption during retention state. The second objective of this paper is to achieve maximum computation efficiency if it is unlikely to avoid power failure. Compared to the instant backup scheme, evaluation results report that power failure can be reduced by 81.6% and computation efficiency can be increased by 2.5x by the proposed retention state-aware energy management strategy. Zhiyao Gong, Keni Qiu, Dongqin Zhou, Weiwen Chen, Yuanchao Xu 0002, Yongpan Liu |
RTCSA | 2 |
| 2017 | Data re-allocation enabled cache locking for embedded systems
Chun Jason Xue, Keni Qiu, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Mengying Zhao |
J. Syst. Archit. | 2 |
| 2017 | On the Implication of NTC versus Dark Silicon on Emerging Scale-Out Workloads: The Multi-Core Architecture PerspectiveabstractThe end of Dennard's scaling poses computer systems, especially the datacenters, in front of both power and utilization walls. One possible solution to combat the power and utilization walls is dark silicon where transistors are under-utilized in the chip, but this will result in a diminishing performance. Another solution is Near-Threshold Voltage Computing (NTC) which operates transistors in the near-threshold region and provides much more flexible tradeoffs between power and performance. However, prior efforts largely focus on a specific design option based on the legacy desktop applications, therefore, lacking comprehensive analysis of emerging scale-out applications with multiple design options when dark silicon and/or NTC are/is applied. In this paper, we characterize different perspectives including performance, energy efficiency and reliability in the context of NTC/dark silicon cloud processors running emerging scale-out workloads on various architecture designs. We find NTC is generally an effective way to alleviate the power challenge over scale-out applications compared with dark silicon, it can improve performance by 1.6X, energy efficiency by 50 percent and the reliability problem can be relieved by ECC. Meanwhile, we also observe tiled-OoO architecture improves the performance by 20~370 percent and energy efficiency by 40~600 percent over alternative architecture designs, making it a preferable design paradigm for scale-out workloads. We believe that our observations will provide insights for the design of cloud processors under dark silicon and/or NTC. Jing Wang 0055, Xin Fu 0001, Weigong Zhang, Keni Qiu, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | Refresh-aware loop scheduling for high performance low power volatile STT-RAMabstractThe highlighted advantages of low leakage power, high storage density and immunity to electronic magnetic radiation make STT-RAM a promising candidate to build cache, SPM or main memory in embedded systems. However, write operations on STT-RAM have considerably longer latency and higher energy consumption than conventional SRAM. To solve this problem, researchers have proposed to relax STT-RAM's non-volatility and to have it work in a fast and low power mode. Under this volatile mode, refresh operations are needed to guarantee data correctness if their lifespan is larger than the retention time. It is observed that this refresh overhead is significant for data in stencil loops with the characteristic of constant read and write dependencies. This paper proposes a loop scheduling technique which can traverse loops in a new direction such that data lifespan can be greatly shortened. Therefore, overall refresh overhead can be efficiently mitigated so as to improve performance and reduce power consumption. The experimental results indicate that access latency and dynamic energy can be improved by 21.4~96.0% and 22.0~95.5% respectively by the proposed scheduling scheme. Keni Qiu, Junpeng Luo, Zhiyao Gong, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Tao Li 0006, Chun Jason Xue |
ICCD | 1 |
| 2016 | An adaptive Non-Uniform Loop Tiling for DMA-based bulk data transfers on many-core processorabstractMesh Network-on-Chip (NoC) is a key fabric to interconnect many cores with desirable scalability, reliability and interoperability. We observe that DMA-based bulk data block transfer exhibits non-negligible NoC latency due to heavy congestions. Loop tiling is an effective way to partition data space for SPM+DMA-based data block transfer. Nevertheless, we observe that the unbalanced NoC latency can degrade the effectiveness of loop tiling in a uniform fashion. In this paper, we propose a NoC-aware Non-Uniform Loop Tiling (NULT) scheme to improve DMA performance. A NULT framework is built on the proposed model to adaptively hide DMA latency into computation time and reduce the overall execution time. The framework first groups cores into different families taking into account their distance-to-data in NoC. Then a heuristic method is presented to solve the near optimal tiling factors for each core family. In this way, different core families are assigned non-uniform tiling sizes. We evaluate the NULT scheme on the NIRGAM platform. Compared to the traditional uniform tiling approach, the proposed NULT technique shows more benefit to overlap memory access time and computation time and thus reduce the overall execution time of a loop nest. Keni Qiu, Yuanhui Ni, Weigong Zhang, Jing Wang 0055, Chun Jason Xue, Tao Li 0006 |
ICCD | 1 |
| 2016 | Exploring Variation-Aware Fault-Tolerant Cache under Near-Threshold ComputingabstractNear threshold voltage computing enables transistor voltage scaling to continue with Moore's Law projection and dramatically improves power and energy efficiency. However, reducing the supply voltage to near-threshold level significantly increases the susceptibility of on-chip caches to process variations, leading to the high error rate. Most existing fault-tolerant schemes significantly sacrifice cache capacity and performance. In this paper, we propose a novel fault-tolerant cache architecture at near-threshold computing, which is suitable for high error rate memories. We first propose a variation-aware skewed-associative cache, and then redirect the faulty blocks to the error-free blocks based on it to explore the fault-tolerance cache design. Unlike previous cache reconfiguration schemes for the fault tolerance, our cache design does not need to sacrifice or disable any fault-free blocks to form a completely functional set. We use all error-free blocks and have the least cache capacity waste. More importantly, since the aging impact could also cause cell failures, our skewed cache takes the aggregated process variation and aging impact into the consideration. Last but not least, our skewed cache design avoids the complex remapping from faulty blocks to the error-free blocks and minimizes the hardware overheads. Our evaluation results show that our variation-aware fault-tolerant cache design exhibits strong capability to tolerate the high error rate, and more excitingly, its effectiveness on reducing the cache miss rate and improving the performance is even more obvious as the supply voltage scales down to the near-threshold region. Jing Wang 0055, Yanjun Liu 0005, Weigong Zhang, Kezhong Lu, Keni Qiu, Xin Fu 0001, Tao Li 0006 |
ICPP | 5 |
| 2016 | Redesigning software and systems for non-volatile processors on self-powered devicesabstractWearable devices gain increasing popularity since they can collect important information for healthcare and well-being purposes. Compared with battery, energy harvesting is a better power source for these wearable devices due to many advantages. However, harvested energy is naturally unstable and program execution will be interrupted frequently. Nonvolatile processor (NVP) demonstrates promising advantages to back up volatile state before the system energy is depleted. Due to the backup and resumption procedures resulted from frequent power failures, non-volatile processor exhibits different characteristics from traditional processors, necessitating a set of adaptive design and optimization strategies. Recently, there have been both hardware and software researches aiming to develop correct and efficient non-volatile processors. In this paper, we summarize the software-level techniques for NVP, covering error-correctness schemes, backup timing determination, backup content optimization, adaptive software modifications and NVP simulators and tools, to provide an overview of state-of-the-art NVP research from the software and system level. Mengying Zhao, Keni Qiu, Yuan Xie 0001, Jingtong Hu, Chun Jason Xue |
VLSI-SoC | 2 |
| 2016 | Reducing Synchronization Cost for Single-Level Store in Mobile Systems
Yuanchao Xu 0002, Hu Wan 0001, Keni Qiu, Tao Li 0006, Weigong Zhang |
J. Comput. Sci. Technol. | 3 |
| 2016 | Write Mode Aware Loop Tiling for High Performance Low Power Volatile PCM in Embedded SystemsabstractArchitecting PCM, especially MLC PCM, as main memory for MCUs is a promising technique to replace conventional DRAM deployment. However, PCM/MLC PCM suffers from long write latency and large write energy. Recent work has proposed a compiler directed dual-write (CDDW) scheme to combat the drawbacks of PCM by adopting fast or slow mode for different write operations. For large-scale loops, we observe that write instances' lifetime is very long and can only be written by the expensive slow mode. This paper proposes a write mode aware loop tiling approach to effectively reduce the lifetime of write instances and maximize the number of efficient fast writes in loops. The experimental results show that the proposed approach improves performance by 50.8 percent and reduces dynamic energy by 32.0 percent across a set of benchmarks compared to the CDDW approach on average. Keni Qiu, Qing'an Li, Jingtong Hu, Weigong Zhang, Chun Jason Xue |
IEEE Trans. Computers | 1 |
| 2014 | Write Mode Aware Loop Tiling for High Performance Low Power Volatile PCMabstractArchitecting PCM, especially MLC PCM, as main memory for MCUs is a promising technique to replace conventional DRAM deployment. However, PCM/MLC PCM suffers from long write latency and large write energy. Recent work has proposed a compiler directed dual-write (CDDW) scheme to combat the drawbacks of PCM by adopting fast or slow write mode for different write operations. We observe that write instances' lifetime is very long and can only be written by the expensive slow mode for large-scale loops. This paper proposes a write mode aware loop tiling approach to effectively reduce the lifetime of write instances and maximize the number of efficient fast writes in loops. The experimental results show that the proposed approach improves performance by 50.8% and reduces dynamic energy by 32.0% across a set of benchmarks compared to the CDDW approach on average. Keni Qiu, Qing'an Li, Chun Jason Xue |
DAC | 1 |
| 2014 | Migration-Aware Loop Retiming for STT-RAM-Based Hybrid Cache in Embedded SystemsabstractRecently hybrid cache architecture consisting of both spin-transfer torque RAM (STT-RAM) and SRAM has been proposed for energy efficiency. In hybrid caches, migration-based techniques have been proposed. A migration technique dynamically moves write-intensive and read-intensive data between STT-RAM and SRAM to explore the advantages of hybrid cache. Meanwhile, migrations also introduce extra reads and writes during data movements. For stencil loops with read and write data dependencies, we observe that migration overhead is significant, and migrations closely correlate to the interleaved read and write memory access pattern in a memory block. This paper proposes a loop retiming framework during compilation to reduce the migration overhead by changing the interleaved memory access pattern. With the proposed loop retiming technique, the interleaved memory accesses can be significantly reduced so that migration overhead is mitigated, and energy efficiency of hybrid cache is significantly improved. The experimental results have shown that, with the proposed methods, on average, the migration number is reduced up to 27.1% and the cache dynamic energy is reduced up to 14.0%. Keni Qiu, Mengying Zhao, Qing'an Li, Chenchen Fu, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | Error Model Guided Joint Performance and Endurance Optimization for Flash MemoryabstractAs flash memory has better performance than hard disks, it has been widely applied in embedded systems, personal computers, and data centers as storage components. However, endurance and write performance are the two key challenges in the deployment of flash memory. In this paper, with the awareness of errors induced from write operations, endurance, and retention time, a stage-based optimization approach is proposed to improve the write performance and endurance at different usage stages of flash memory. A series of trace-driven simulations show that the proposed approach outperforms a set of state-of-the-art approaches in terms of write performance and lifetime. Liang Shi 0001, Keni Qiu, Mengying Zhao, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Branch Prediction-Directed Dynamic Instruction Cache Locking for Embedded SystemsabstractCache locking is a cache management technique to preclude the replacement of locked cache contents. Cache locking is often adopted to improve cache access predictability in Worst-Case Execution Time (WCET) analysis. Static cache locking methods have been proposed recently to improve Average-Case Execution Time (ACET) performance. This article presents an approach, Branch Prediction-directed Dynamic Cache Locking (BPDCL), to improve system performance through cache conflict miss reduction. In the proposed approach, the control flow graph of a program is first partitioned into disjoint execution regions, then memory blocks worth locking are determined by calculating the locking profit for each region. These two steps are conducted during compilation time. At runtime, directed by branch predictions, locking routines are prefetched into a small high-speed buffer. The predetermined cache locking contents are loaded and locked at specific execution points during program execution. Experimental results show that the proposed BPDCL method exhibits an average improvement of 25.9%, 13.8%, and 8.0% on cache miss rate reduction in comparison to cases with no cache locking, the static locking method, and the dynamic locking method, respectively. Keni Qiu, Mengying Zhao, Chun Jason Xue, Alex Orailoglu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Migration-aware loop retiming for STT-RAM based hybrid cache for embedded systemsabstractIn hybrid cache architecture consisting of both STT-RAM and SRAM, migration based techniques have been proposed. The migration technique dynamically moves write-intensive and read-intensive data between STT-RAM and SRAM to explore the advantage of hybrid cache. Meanwhile, migrations induce extra read and write overhead during data movements. For loops with intensive data array operations, we observe that migration overhead is significant and migrations closely correlate to the interleaved read and write access pattern in a memory block. This paper proposes a loop retiming framework to reduce the migration overhead by changing the interleaved memory access pattern. The experimental results show that with the proposed method, migrations are significantly reduced without any hardware modification. As a result, energy efficiency and performance of hybrid cache can be improved. Keni Qiu, Mengying Zhao, Chenchen Fu, Liang Shi 0001, Chun Jason Xue |
ASAP | 1 |
| 2013 | Branch Prediction directed Dynamic instruction Cache Locking for embedded systemsabstractCache locking is a cache management technique to preclude the replacement of locked cache contents. Cache locking is often used to improve cache access predictability in Worst-Case Execution Time (WCET) analysis. Static cache locking methods have been proposed recently to improve average system performance. This paper presents an approach, Branch Prediction directed Dynamic Cache Locking (BPDCL), to improve average system performance through effective cache conflict miss reduction in different execution regions. In this proposed approach, the control flow graph of a program is partitioned into regions and memory blocks worth locking for each region are calculated during compilation time. At runtime, directed by branch predictions, locking routines are prefetched into a high-speed buffer. The pre-determined cache locking contents are loaded and locked at specific execution points during program execution. Experimental results show that the proposed BPDCL method exhibits an average improvement of 21.8% and 10.3% on cache miss rate reduction in comparison to the case with no cache locking and the static locking method respectively. Keni Qiu, Mengying Zhao, Chun Jason Xue, Alex Orailoglu |
RTCSA | 1 |
| 2013 | Data re-allocation enabled cache locking for embedded systemsabstractCache locking is a cache management technique to preclude the replacement of locked contents. Recently, instruction cache locking has been applied to improve average-case execution time (ACET). However, we observe that the prior instruction cache locking method shows very limited performance improve-ment for data cache. The main reason lies in that, data access similarity in data memory blocks is weaker than that in code memory blocks. This paper proposes a data re-allocation enabled cache locking approach which can significantly enhance locking efficiency for data cache and thus improve system performance. The experimental results show that with the proposed approach, on average, the miss rate is reduced by 9.1% and execution cycles are reduced by 9.4% across a suite of benchmarks. Keni Qiu, Mengying Zhao, Chenchen Fu, Chun Jason Xue |
VLSI-SoC | 1 |