EDBT 2026 Demo / reviewers in the wild / expert
Mimi Xie
dblp:151/6207
· DBLP profile ↗
42ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0003-1973-2909ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 5 first-author · 18 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AERO: Adaptive and Efficient Runtime-Aware OTA Updates for Energy-Harvesting IoTabstractEnergy-harvesting (EH) Internet of Things (IoT) devices operate under intermittent energy availability, which disrupts task execution and makes energy-intensive over-the-air (OTA) updates particularly challenging. Conventional OTA update mechanisms rely on reboots and incur significant overhead, rendering them unsuitable for intermittently powered systems. Recent live OTA update techniques reduce reboot overhead but still lack mechanisms to ensure consistency when updates interact with runtime execution. This paper presents AERO, an Adaptive and Efficient Runtime-Aware OTA update mechanism that integrates update tasks into the device’s Directed Acyclic Graph (DAG) and schedules them alongside routine tasks under energy and timing constraints. By identifying update-affected execution regions and dynamically adjusting dependencies, AERO ensures consistent update integration while adapting to intermittent energy availability. Experiments on representative workloads demonstrate improved update reliability and efficiency compared to existing live update approaches. Wei Wei 0060, Jingye Xu, Sahidul Islam, Dakai Zhu 0001, Mimi Xie |
DATE | 6 |
| 2026 | Intelligent UAV Coordination for Data Freshness in Energy-Harvesting IoT NetworksabstractUAV-assisted energy harvesting (EH) IoT networks enable data collection in resource-limited areas but face critical challenges from obstructed communication, intermittent harvested energy, and limited UAV flight time, which degrade data freshness. This work proposes a three-tier Coordinated UAV Scheduling and Charging (CUSC) framework: (1) a deep reinforcement learning (DRL) scheduler that selects the next cluster head (these also serve as charging stations) based on Age of Information (AoI) and energy constraints to optimize inter-cluster flight scheduling; (2) a Christofides-based procedure that computes an intra-cluster visitation order for unresponsive sensors; and (3) an AoI-aware dynamic UAV-charging scheme at each cluster head that sets the charging order and duration in real time based on current AoI and the local queue state. These cluster heads also aggregate nearby sensor data, which reduces the need for time-consuming UAV-to-sensor direct communication. Our framework achieves a 51.9% reduction in peak AoI for 25 km2 monitoring areas while maintaining a 2:1 flight-to-charging ratio through optimized station selection. The combined approach improves data collection by 1.7× with only 70% inter-cluster deviation, demonstrating effective coordination between mobility planning, intra-cluster routing, and adaptive charging in EH-IoT. Mason Conkel, Mimi Xie, Yufang Jin, Wenlu Wang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | NetVault: A Lightweight IP Protection Framework for Inference on Embedded IoT DevicesabstractDeploying deep neural network (DNN) models at the edge enables real-time inference while enhancing data privacy. However, the intellectual property (IP) of these devices faces significant risks from malicious end-users. Traditional IP protection techniques often incur high costs and rely on specialized hardware. In this paper, we present NetVault, a novel and lightweight framework designed to proactively secure the IP of DNNs. Specifically, NetVault sabotages the inference accuracy by selectively altering the model’s weights in critical layers before the device enters an unprotected offline stage and encodes these weights’ locations. We leverage an embedded AES accelerator to fast encrypt the weight locations and original values of the weights and use the RSA mechanism to secure the encryption key to ensure that the compromised model is rendered ineffective for unauthorized users. Performance evaluations on typical embedded devices demonstrate that NetVault introduces negligible energy and time overhead. Extensive tests across various neural network architectures and datasets further validate the efficacy and practicality of NetVault as a robust IP protection solution. Nathan Wiatrek, Mimi Xie |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | SCREME: A Scalable Framework for Resilient Memory DesignabstractThe continuing advancement of memory technology has not only fueled a surge in performance, but also substantially exacerbated reliability challenges. Traditional solutions have primarily focused on improving the efficiency of protection schemes, i.e., Error Correction Codes (ECC), under the assumption that allocating additional memory space for ECC parity is always costly and therefore unsustainable as parity scales. We break the stereotype by proposing an orthogonal approach that provides additional, cost-effective memory space for resilient memory design. In particular, we recognize that ECC chips (used for parity storage) do not necessarily require the same performance level as regular data chips. This offers two-fold benefits: First, the bandwidth originally provisioned for a regularperformance ECC chip can instead be used to accommodate multiple low-performance chips. Second, the cost of ECC chips can be effectively reduced, as lower performance often correlates with lower expense. In addition, we observe that server-class memory chips are often provisioned with ample, yet underutilized I/O resources. This suggests an opportunity to repurpose these resources for flexible on-DIMM interconnections. Based on the above two insights, we finally propose SCREME, a scalable memory framework that leverages cost-effective yet slower chips - byproducts of rapid technology evolution - to meet the growing reliability demands driven by this evolution. Mimi Xie, Yanan Guo 0002, Huize Li, Xin Xin 0008 |
PACT | 2 |
| 2025 | ReLAD: Retrieval-Augmented LAtent Diffusion for Complex Aerial Image Generation
Douglas J. Townsell, Lingwei Chen, Mimi Xie |
IEEE Big Data | 3 |
| 2025 | Autonomous UAV-Assisted IoT Systems with Deep Reinforcement Learning Based Data FerryabstractEmerging unmanned aerial vehicle (UAV) technology offers reliable, flexible, and controllable techniques for transfer-ring data collected by wireless internet of things (IoT) devices located in remote areas. However, deploying UAVs faces limitations in mission distance to recharging, especially when recharge occurs far from the monitoring. To address these challenges, we propose smart charging stations installed within the monitoring area equipped with energy-harvest features and communication modules. These stations can replenish the UAV's energy and act as cluster heads by collecting information from IoT devices within their jurisdiction. This allows a UAV to operate continuously by downloading while charging and forwarding the data to the remote server during flight. Despite these improvements, the unpredictable nature of energy-harvest devices and charging needs can lead to stale or obsolete information at cluster heads. The limited communication range may prevent the cluster heads from establishing connections with all nodes in their jurisdiction. To overcome these issues, we proposed an age-of-information-aware data ferry algorithm using deep reinforcement learning to determine the UAV's flight path. The deep reinforcement learning agent, running on cluster heads, utilizes a global state gathered by the UAV to output the location of the next stop, which can be a cluster head or an IoT device. The experiments show that the algorithm can minimize the age of information without diminishing data collection. Mason Conkel, Mimi Xie, Yufang Jin |
DATE | 3 |
| 2025 | AeroDiffusion: Complex Aerial Image Synthesis with Keypoint-Aware Text Descriptions and Feature-Augmented Diffusion ModelsabstractAerial imagery provides crucial insights for various fields, including remote monitoring, environmental assessment, and autonomous navigation. However, the availability of aerial image datasets is limited due to privacy concerns and imbalanced data distribution, impeding the development of robust deep learning models. Recent advancements in text-guided image synthesis offer a promising approach to enrich and diversify these datasets. Despite progress, existing generative models face challenges in synthesizing realistic aerial images due to the lack of paired text-aerial datasets, the complexity of densely packed objects, and the limitations of modeling object relationships. In this paper, we introduce AeroDiffusion, a novel framework designed to overcome these challenges by leveraging large language models (LLMs) for keypoint-aware text description generation and a feature-augmented diffusion process for realistic image synthesis. Our approach integrates region-level feature extraction to preserve small objects and multimodal feature alignment to improve textual descriptions of complex aerial scenes. AeroDiffusion is the first to extend deep generative models for high-resolution, text-guided aerial image generation, including the creation of images from novel viewpoints. We contribute a new paired text-aerial image dataset and demonstrate the effectiveness of our model, achieving an FID score of 78.15 across five benchmarks, significantly outperforming state-of-the-art models such as DDPM (217.95), Stable Diffusion (119.13), and ARLDM (111.59). Douglas J. Townsell, Mimi Xie, Fathi H. Amsaad 0001, Varshitha Reddy Thanam |
DATE | 2 |
| 2025 | CipherPrune: Efficient and Scalable Private Transformer InferenceabstractPrivate Transformer inference using cryptographic protocols offers promising solutions for privacy-preserving machine learning; however, it still faces significant runtime overhead (efficiency issues) and challenges in handling long-token inputs (scalability issues). We observe that the Transformer's operational complexity scales quadratically with the number of input tokens, making it essential to reduce the input token length. Notably, each token varies in importance, and many inputs contain redundant tokens. Additionally, prior private inference methods that rely on high-degree polynomial approximations for non-linear activations are computationally expensive. Therefore, reducing the polynomial degree for less important tokens can significantly accelerate private inference. Building on these observations, we propose \textit{CipherPrune}, an efficient and scalable private inference framework that includes a secure encrypted token pruning protocol, a polynomial reduction protocol, and corresponding Transformer network optimizations. At the protocol level, encrypted token pruning adaptively removes unimportant tokens from encrypted inputs in a progressive, layer-wise manner. Additionally, encrypted polynomial reduction assigns lower-degree polynomials to less important tokens after pruning, enhancing efficiency without decryption. At the network level, we introduce protocol-aware network optimization via a gradient-based search to maximize pruning thresholds and polynomial reduction conditions while maintaining the desired accuracy. Our experiments demonstrate that CipherPrune reduces the execution overhead of private Transformer inference by approximately $6.1\times$ for 128-token inputs and $10.6\times$ for 512-token inputs, compared to previous methods, with only a marginal drop in accuracy. The code is publicly available at https://github.com/UCF-Lou-Lab-PET/cipher-prune-inference. Yancheng Zhang, Mengxin Zheng, Mimi Xie, Mingzhe Zhang 0005, Lei Jiang 0001, Qian Lou |
ICLR | 4 |
| 2025 | InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer InteractionabstractThis paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video.
Unlike existing approaches that either build intricate workflows around a single large model or only provide workflow modularity, our agent integrates tool-based and pure vision agents within a highly modular architecture, enabling different models to collaboratively solve decoupled tasks in a step-by-step manner.
Our generality is demonstrated by our ability to evaluate not only pure vision-based real-world benchmarks (i.e., OSWorld), but also more general or tool-intensive benchmarks (e.g., GAIA and SWE-Bench).
Specifically,
we
achieve a $\mathbf{7.27\\%}$ accuracy gain over Claude-Computer-Use on OSWorld.
Codes and evaluation scripts are included in the supplementary material and will be released as open-source. Weitai Kang, Winson Chen, Shan Zuo, Mimi Xie, Ali Payani, Mingyi Hong 0001, Caiwen Ding |
NeurIPS | 7 |
| 2025 | STARS: Semantics-Aware Text-guided Aerial Image Refinement and Synthesis
Douglas J. Townsell, Lingwei Chen, Mimi Xie |
Comput. Vis. Image Underst. | 3 |
| 2025 | Intermittent OTA Code Update Framework for Tiny Energy Harvesting DevicesabstractThe widespread deployment of various tiny energy harvesting devices has facilitated the expansion of Internet of Things (IoT) applications, notably in remote and hard-to-reach areas. Once deployed, a critical limitation of these devices is their inability to adapt code to evolving environmental conditions or user requirements. This challenge primarily arises from frequent power interruptions during code updates in energy harvesting devices, unlike their battery-powered counterparts, which can lead to significant errors or system failures. In response, we have designed an innovative framework for facilitating intermittent over-the-air (OTA) code updates in tiny energy harvesting devices. Our approach incorporates Intermittent-aware Update Operations, including insert, modify, delete, and copy, tailored for a variety of update scenarios while accommodating intermittent power and resource constraints. Furthermore, We have designed a Fault-tolerant bootloader that enables the intermittent update capability. This advanced bootloader enables code updates without system reboots and ensures correct task resumption of both routine and update tasks. This not only conserves energy by reducing the need for repetitive reboots but also ensures consistent code updates despite frequent power failures. Additionally, our framework integrates an update-aware checkpointing mechanism to provide reliable backups for both routine tasks and update tasks. This proposed framework presents a general solution for enabling intermittent code updates in tiny energy harvesting devices. Our experimental results demonstrate that the proposed approach outperforms existing approaches under conditions of insufficient harvested energy. Wei Wei 0060, Sahidul Islam, Jishnu Banerjee, Shyamala Palanisamy, Mimi Xie |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Autotile: Autonomous Task-tiling for Deep Inference on Battery-less Embedded SystemabstractDeep Neural Networks (DNNs) are increasingly applied in various intelligent applications for enhanced accuracy for in-situ decision-making. Considering the cost and longevity, those intelligent applications usually employ energy harvesting (EH) for power supply. Nevertheless, due to inherent intermittency, EH power can frequently disrupt the runtime operation, resulting in subsequent forward progress loss when executing long computations of DNN inference. To address this issue, DNN tiling has been employed where the input data is partitioned into multiple smaller tiles for efficient runtime operation. However, under the energy harvesting scenarios, the size of the tiles can influence runtime energy efficiency significantly under different EH conditions. Therefore, we proposed environmentally adaptive dynamic DNN tiling methods to optimize energy efficiency and runtime reliability. The experimental results on a real testbed show that the proposed technique can outperform the state-of-the-art methods by 19.24% on average. Jishnu Banerjee, Sahidul Islam, Wei Wei 0060, Mimi Xie |
ACM Great Lakes Symposium on VLSI | 5 |
| 2024 | Real-time intelligent on-device monitoring of heart rate variability with PPG sensors
Jingye Xu, Yuntong Zhang 0001, Mimi Xie, Wei Wang 0054, Dakai Zhu 0001 |
J. Syst. Archit. | 3 |
| 2024 | Intelligent Networking for Energy Harvesting Powered IoT SystemsabstractAs the next-generation battery substitute for IoT system, energy harvesting (EH) technology revolutionizes the IoT industry with environmental friendliness, ubiquitous accessibility, and sustainability, which enables various self-sustaining IoT applications. However, due to the weak and intermittent nature of EH power, the performance of EH-powered IoT systems as well as its collaborative routing mechanism can severely deteriorate, rendering unpleasant data package loss during each power failure. Such a phenomenon makes conventional routing policies and energy allocation strategies impractical. Given the complexity of the problem, reinforcement learning (RL) appears to be one of the most promising and applicable methods to address this challenge. Nevertheless, although the energy allocation and routing policy are jointly optimized by the RL method, due to the energy restriction of EH devices, the inappropriate configuration of multi-hop network topology severely degrades the data collection performance. Therefore, this article first conducts a thorough mathematical discussion and develops the topology design and validation algorithm under energy harvesting scenarios. Then, this article develops DeepIoTRouting , a distributed and scalable deep reinforcement learning (DRL)-based approach, to address the routing and energy allocation jointly for the energy harvesting powered distributed IoT system. The experimental results show that with topology optimization, DeepIoTRouting achieves at least 38.71% improvement on the amount of data delivery to sink in a 20-device IoT network, which significantly outperforms state-of-the-art methods. Tao Liu 0023, Jeff Zhang 0001, Mehdi Sookhak, Mimi Xie |
ACM Trans. Sens. Networks | 6 |
| 2023 | Dynamic Sparse Training via Balancing the Exploration-Exploitation Trade-offabstractOver-parameterization of deep neural networks (DNNs) has shown high prediction accuracy for many applications. Although effective, the large number of parameters hinders its popularity on resource-limited devices and has an outsize environmental impact. Sparse training (using a fixed number of nonzero weights in each iteration) could significantly mitigate the training costs by reducing the model size. However, existing sparse training methods mainly use either random-based or greedy-based drop-and-grow strategies, resulting in local minimal and low accuracy. In this work, to assist explainable sparse training, we propose important weights Exploitation and coverage Exploration to characterize Dynamic Sparse Training (DST-EE), and provide quantitative analysis of these two metrics. We further design an acquisition function and provide the theoretical guarantees for the proposed method and clarify its convergence property. Experimental results show that sparse models (up to 98% sparsity) obtained by our proposed method outperform the SOTA sparse training methods on a wide variety of deep learning tasks. On VGG-19 / CIFAR-100, ResNet-50 / CIFAR-10, ResNet-50 / CIFAR-100, our method has even higher accuracy than dense models. On ResNet-50 / ImageNet, the proposed method has up to 8.2% accuracy improvement compared to SOTA sparse training methods. Shaoyi Huang, Bowen Lei, Dongkuan Xu, Hongwu Peng, Mimi Xie, Caiwen Ding |
DAC | 6 |
| 2023 | Virtual Summer Camp for High School Students with Disabilities - An Experience ReportabstractIn the past years, the authors held computer programming and machine learning summer camp for high-school students with disabilities. Due to the pandemic, the summer camp was offered virtually in 2020 and 2021. This paper reports our experience of teaching this summer camp. The main goal of the summer camp was to let students with disabilities get first-hand experience of working in STEM fields to encourage them to pursue STEM careers. The curriculum was primarily composed of hands-on activities for Python programming and Computer Vision. Besides lectures and programming tasks, there were also guest speakers and an external panel to offer their personal experiences of working in STEM fields. Wei Wang 0054, Kathy B. Ewoldt, Mimi Xie, Alberto M. Mestas-Nuñez, Sean Soderman, Jeffrey Wang |
SIGCSE (1) | 3 |
| 2023 | ELIXIR: An Expedient Connection Paradigm for Self-Powered IoT DevicesabstractIoT devices usually work under power-constrained scenarios like outdoor environmental monitoring. Considering the cost and sustainability, in the long run, energy-harvesting technology is preferable for powering IoT devices. Since harvesting power is intrinsically weak and transient, the connection between IoT devices not only cannot be maintained constantly but is also extremely difficult to be established. In order to communicate, those devices should have a synchronized timeline so that both transmitter and receiver can start connection simultaneously. Yet due to the transient nature of ambient energy, volatile time data can be easily tampered with by the unstable power supply or corrupted completely by frequent power outages. As a result, unsynchronized IoT devices require significant efforts in time and energy to reconnect and communicate, which further escalates the performance degradation of the edge network. To adapt to ubiquitous self-powered scenarios on the IoT edge, we propose ELIXIR, an expedient connection paradigm, to synchronize swiftly and autonomously in a joint effort for self-powered IoT devices. The lightweight and highly efficient natures enable ELIXIR to be integrated into low-power IoT devices easily and efficiently. The experimental results show that the proposed ELIXIR can avoid a connection loop and help speed up the connection 2.83X and 1.84X faster than baseline methods. Yanzhi Wang 0001, Mimi Xie |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune ParadigmabstractShaoyi Huang, Dongkuan Xu, Ian Yen, Yijue Wang, Sung-En Chang, Bingbing Li, Shiyang Chen, Mimi Xie, Sanguthevar Rajasekaran, Hang Liu, Caiwen Ding. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shaoyi Huang, Dongkuan Xu, Ian En-Hsu Yen, Yijue Wang, Sung-En Chang, Shiyang Chen 0004, Mimi Xie, Sanguthevar Rajasekaran, Hang Liu 0001, Caiwen Ding |
ACL (1) | 8 |
| 2022 | Energy Harvesting Aware Multi-Hop Routing Policy in Distributed IoT System Based on Multi-Agent Reinforcement LearningabstractEnergy harvesting technologies offer a promising solution to sustainably power an ever-growing number of Internet of Things (IoT) devices. However, due to the weak and transient natures of energy harvesting, IoT devices have to work intermittently rendering conventional routing policies and energy allocation strategies impractical. To this end, this paper, for the very first time, developed a distributed multi-agent reinforcement algorithm known as global actor-critic policy (GAP) to address the problem of routing policy and energy allocation together for the energy harvesting powered IoT system. At the training stage, each IoT device is treated as an agent and one universal model is trained for all agents to save computing resources. At the inference stage, packet delivery rate can be maximized. The experimental results show that the proposed GAP algorithm achieves ~ 1.28× and ~ 1.24× data transmission rate than that of the Q-table and ESDSRAA algorithm, respectively. Tao Liu 0023, Mimi Xie, Longzhuang Li, Dulal C. Kar |
ASP-DAC | 3 |
| 2022 | Enabling Fast Deep Learning on Tiny Energy-Harvesting IoT DevicesabstractEnergy harvesting (EH) IoT devices that operate intermittently without batteries, coupled with advances in deep neural networks (DNNs), have opened up new opportunities for en-abling sustainable smart applications. Nevertheless, implementing those computation and memory-intensive intelligent algorithms on EH devices is extremely difficult due to the challenges of limited resources and intermittent power supply that causes frequent failures. To address those challenges, this paper proposes a methodology that enables fast deep learning with low-energy accelerators for tiny energy harvesting devices. We first propose RAD, a resource-aware structured DNN training framework, which employs block circulant matrix and structured pruning to achieve high compression for leveraging the advantage of various vector operation accelerators. A DNN implementation method, ACE, is then proposed that employs low-energy accelerators to profit maximum performance with small energy consumption. Finally, we further design FLEX, the system support for inter-mittent computation in energy harvesting situations. Experimental results from three different DNN models demonstrate that RAD, ACE, and FLEX can enable fast and correct inference on energy harvesting devices with up to 4.26X runtime reduction, up to 7. 7X energy reduction with higher accuracy over the state-of-the-art. Sahidul Islam, Jieren Deng, Shanglin Zhou, Caiwen Ding, Mimi Xie |
DATE | 6 |
| 2022 | M2M-Routing: Environmental Adaptive Multi-agent Reinforcement Learning based Multi-hop Routing Policy for Self-Powered IoT SystemsabstractEnergy harvesting (EH) technologies facilitate the trending proliferation of IoT devices with sustainable power supplies. However, the intrinsic weak and unstable nature of EH results in frequent and unpredictable power interruptions in EH IoT devices, which further causes unpleasant packet loss or reconnection failures in IoT network. Therefore, conventional routing and energy allocation methods are inefficient in the EH environments. The complexity of the EH environment caused a stumbling block to an intelligent routing policy and energy allocation. To address the problems, this work proposes an environment adaptive Deep Reinforcement Learning (DRL)-based multi-hop routing policy, M2M-Routing, to jointly optimize energy allocation and routing policy and mitigate these challenges through leveraging the offline computation resources. We prepare multi-models for the complex energy harvesting environment offline. By searching a historically similar power trace to identify the model ID, the prepared DRL model is selected to manage energy allocation and routing policy on the query power traces. Simulation results indicate that M2M-Routing improves the amount of data delivery by ~ 3 × to ~ 4 × compared with baselines. Jeff Zhang 0001, Mimi Xie, Tao Liu 0023, Wenlu Wang |
DATE | 3 |
| 2022 | EVE: Environmental Adaptive Neural Network Models for Low-Power Energy Harvesting SystemabstractIoT devices are increasingly being implemented with neural network models to enable smart applications. Energy harvesting (EH) technology that harvests energy from ambient environment is a promising alternative to batteries for powering those devices due to the low maintenance cost and wide availability of the energy sources. However, the power provided by the energy harvester is low and has an intrinsic drawback of instability since it varies with the ambient environment. This paper proposes EVE, an automated machine learning (autoML) co-exploration framework to search for desired multi-models with shared weights for energy harvesting IoT devices. Those shared models incur significantly reduced memory footprint with different levels of model sparsity, latency, and accuracy to adapt to the environmental changes. An efficient on-device implementation architecture is further developed to efficiently execute each model on device. A run-time model extraction algorithm is proposed that retrieves individual model with negligible overhead when a specific model mode is triggered. Experimental results show that the neural networks models generated by EVE is on average 2.5× times faster than the baseline models without pruning and shared weights. Sahidul Islam, Shanglin Zhou, Yufang Jin, Wujie Wen, Caiwen Ding, Mimi Xie |
ICCAD | 7 |
| 2022 | Sparsity-Aware Intelligent Spatiotemporal Data Sensing for Energy Harvesting IoT SystemabstractIn this era of the Internet of Things (IoT), the increasing number of IoT devices benefit from energy harvesting (EH) technology which enables a sustainable data acquisition process, including data sensing, communication, and storing for promoting the well beings of the society. However, intermittent and low EH power confines the data acquisition process. Specifically, due to frequent harvesting power outages and depletion of energy for expensive data transmission to the IoT edge server, insufficient energy is allocated for data sensing resulting in the missing of key information. To address this issue, this article proposes a sparsity-aware spatiotemporal data sensing framework for EH IoT devices to minimize the data sensing rate/energy while acquiring comprehensive information and reserving sufficient energy. In this framework, the IoT devices sample critical sparse spatiotemporal data, and then the sparse data are sent to the edge server for reconstruction. To maximize the reconstruction accuracy subject to the limited power supply and intermittent work patterns of EH devices, we first propose the QR-based algorithm QR-ST to initiate a sensing scheduling for each EH device. Due to the unstable and intermittent work pattern, the schedule needs to be dynamically fine-tuned based on environmental inputs. Therefore, we further propose a multiagent deep reinforcement learning-based method named S-Agents for the IoT edge server to globally select the sensing devices at each time slot, where the spatial and temporal features of reconstructed data are guaranteed. Experimental results show that the proposed framework reduced the reconstruction error by 66.30% compared with baselines. Mimi Xie, Caleb Scott |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Memory-aware Efficient Deep Learning Mechanism for IoT DevicesabstractDeep learning neural networks are of critical importance to enable next-generation IoT devices. However, due to the limited computation power, memory space, and energy, it remains a grand challenge to deploy those algorithms on IoT devices efficiently as they demand high computation, energy, and memory footprint. Numerous pruning methods of deep learning algorithms have been proposed to minimize the latency, energy, and weights. However, few consider the running time memory footprint and the overhead caused by data movement between the volatile memory and non-volatile memory. This paper proposes four novel memory-ware mechanisms for implementing CNN models on self-restrained embedded systems. The proposed techniques maximize the use of high-speed volatile memory and provide three implementation choices to achieve the minimum energy cost, SRAM space usage, and inference latency, as well as a hybrid trade-off choice of the three features. The experimental evaluation compares their energy cost, time latency, and required run-time memory footprint and demonstrates high implementation efficiency. Jishnu Banerjee, Sahidul Islam, Wei Wei 0060, Dakai Zhu 0001, Mimi Xie |
ASAP | 6 |
| 2021 | Binary Complex Neural Network Acceleration on FPGA : (Invited Paper)abstractBeing able to learn from complex data with phase information is imperative for many signal processing applications. Today’s real-valued deep neural networks (DNNs) have shown efficiency in latent information analysis but fall short when applied to the complex domain. Deep complex networks (DCN), in contrast, can learn from complex data, but have high computational costs; therefore, they cannot satisfy the instant decision-making requirements of many deployable systems dealing with short observations or short signal bursts. Recent, Binarized Complex Neural Network (BCNN), which integrates DCNs with binarized neural networks (BNN), shows great potential in classifying complex data in real-time. In this paper, we propose a structural pruning based accelerator of BCNN, which is able to provide more than 5000 frames/s inference throughput on edge devices. The high performance comes from both the algorithm and hardware sides. On the algorithm side, we conduct structural pruning to the original BCNN models and obtain 20 × pruning rates with negligible accuracy loss; on the hardware side, we propose a novel 2D convolution operation accelerator for the binary complex neural network. Experimental results show that the proposed design works with over 90% utilization and is able to achieve the inference throughput of 5882 frames/s and 4938 frames/s for complex NIN-Net and ResNet-18 using CIFAR-10 dataset and Alveo U280 Board. Hongwu Peng, Shanglin Zhou, Scott Weitze, Sahidul Islam, Tong Geng, Ang Li 0006, Wei Zhang 0052, Minghu Song, Mimi Xie, Hang Liu 0001, Caiwen Ding |
ASAP | 10 |
| 2020 | FTRANS: energy-efficient acceleration of transformers using FPGAabstractIn natural language processing (NLP), the "Transformer" architecture was proposed as the first transduction model replying entirely on self-attention mechanisms without using sequence-aligned recurrent neural networks (RNNs) or convolution, and it achieved significant improvements for sequence to sequence tasks. The introduced intensive computation and storage of these pre-trained language representations has impeded their popularity into computation and memory constrained devices. The field-programmable gate array (FPGA) is widely used to accelerate deep learning algorithms for its high parallelism and low latency. However, the trained models are still too large to accommodate to an FPGA fabric. In this paper, we propose an efficient acceleration framework, Ftrans, for transformer-based large scale language representations. Our framework includes enhanced block-circulant matrix (BCM)-based weight representation to enable model compression on large-scale language representations at the algorithm level with few accuracy degradation, and an acceleration design at the architecture level. Experimental results show that our proposed framework significantly reduce the model size of NLP models by up to 16 times. Our FPGA design achieves 27.07× and 81 × improvement in performance and energy efficiency compared to CPU, and up to 8.80× improvement in energy efficiency compared to GPU. Santosh Pandey 0001, Haowen Fang, Yanjun Lyv, Ji Li 0006, Jieyang Chen, Mimi Xie, Lipeng Wan 0001, Hang Liu 0001, Caiwen Ding |
ISLPED | 7 |
| 2019 | Modeling and Optimization for Self-powered Non-volatile IoT Edge Devices with Ultra-low Harvesting PowerabstractEnergy harvesters are becoming increasingly popular as power sources for IoT edge devices. However, one of the intrinsic problems of energy harvester is that harvesting power is often weak and frequently interrupted. Therefore, energy harvesting powered edge devices have to work intermittently. To maintain execution progress, execution states need to be checkpointed into the non-volatile memory before each power failure. In this way, previous execution states can be resumed after power comes back again. Nevertheless, frequent checkpointing and low charging efficiency generate significant energy overhead. To alleviate these problems, this article conducts a thorough energy efficiency analysis and proposes three algorithms to maximize the energy efficiency of program execution. First, a non-volatile processor-aware task scheduling algorithm is proposed to reduce the size of checkpointing data. Second, a tentative checkpointing avoidance technique is proposed to avoid checkpointing for further reduction of checkpointing overhead. Finally, a dynamic wake-up strategy is proposed to wake up the edge device at proper voltages where the total hardware and software overhead is minimized for further energy efficiency maximization. The experiments on a real testbed demonstrate that, with the proposed algorithms, an edge device is resilient to the extremely weak and intermittent power supply and the energy efficiency can be achieved more than 2× higher than the fundamental baseline and 1.5× higher than the state-of-the-art technique. Mimi Xie, Song Han 0003, Zhi-Hong Mao, Jingtong Hu |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2018 | AIM: Fast and energy-efficient AES in-memory implementation for emerging non-volatile main memoryabstractNon-volatile main memory-based systems pose an opportunity for an attacker to readily access sensitive information on the memory because of its long retention time. While real-time memory encryption with dedicated AES engine can address this vulnerability, it incurs extra performance and energy overheads. As an alternative, we propose an AES in-memory implementation, AIM, to encrypt the whole/part of the memory only when it is necessary. We leverage the benefits offered by the inmemory computing architecture to address the challenges of the bandwidth intensive encryption application. We take advantage of NVM's intrinsic logic operation capability to implement the AES task. Embracing the massive parallelism inside the memory, AIM outperforms existing mechanisms with higher throughput yet lower energy consumption. Compared with state-of-the-art AES engine running at 2.1GHz, AIM can speed up the encryption process by 80 χ for a 1GB NVM. Mimi Xie, Shuangchen Li, Alvin Oliver Glova, Jingtong Hu, Yuangang Wang, Yuan Xie 0001 |
DATE | 1 |
| 2018 | ENZYME: An Energy-Efficient Transient Computing Paradigm for Ultralow Self-Powered IoT Edge DevicesabstractInternet of Things (IoT) edge devices usually work under power-constrained scenarios like outdoor environmental monitoring. Considering the cost and sustainability in a long run, energy harvesting technology is preferable for edge devices. Nevertheless, the harvesting power is generally weak and unstable, making it difficult for edge devices to maintain the normal functionality. Hence, it is crucial to improve the energy efficiency of edge devices. This paper proposes a software paradigm, ENZYME, to improve the energy efficiency of edge devices for transient computing with ultralow energy harvesting power supplies. ENZYME consists of two lightweight yet highly efficient software modules including Routine Handler and frequency modulator (FM). Routine Handler assists power regulator to maximize power extraction from energy harvesters with proper operation routines. Further, FM maximizes the utility of the extracted energy for program execution via efficient runtime clock frequency modulation. The lightweight and highly efficient natures enable ENZYME to be integrated into low-power IoT edge devices easily and efficiently. Experimental results demonstrate that ENZYME achieves more than 8.8% energy efficiency over state-of-the-art techniques with Routine Handler, and 35.71% extra energy utility with FM upon applying Routine Handler. Mimi Xie, Jingtong Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Avoiding Data Inconsistency in Energy Harvesting Powered Embedded SystemsabstractEnergy harvesting is becoming a favorable alternative to power future generation embedded systems, as it is more environmentally and user friendly. However, energy harvesting powered embedded systems suffer from frequent execution interruption due to unstable energy supply. To tackle this problem, nonvolatile memory has been deployed to save the whole volatile state for computation. When power resumes, the processor can restore the state back to volatile memories and continue execution. However, without careful consideration, the process of checkpointing and resuming could cause inconsistency between volatile and nonvolatile memories, which leads to irreversible errors. In this article, we propose a consistency-aware adaptive checkpointing scheme that ensures correctness for all checkpoints. The proposed technique efficiently identifies all possible inconsistency positions in programs and inserts auxiliary code to ensure correctness by offline analysis. In addition, adaptive checkpointing assisted register file profiling and online tracking techniques further reduce the overhead of each checkpoint. Evaluation results show that the proposed checkpointing strategy can successfully eliminate inconsistency errors and greatly reduce the checkpointing overhead. Mimi Xie, Mengying Zhao, Yongpan Liu, Chun Jason Xue, Jingtong Hu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2018 | Securing Emerging Nonvolatile Main Memory With Fast and Energy-Efficient AES In-Memory Implementation
Mimi Xie, Shuangchen Li, Alvin Oliver Glova, Jingtong Hu, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | A lightweight progress maximization scheduler for non-volatile processor under unstable energy harvestingabstractEnergy harvesting techniques become increasingly popular as power supplies for embedded systems. However, the harvested energy is intrinsically unstable. Thus, the program execution may be interrupted frequently. Although the development of non-volatile processors (NVP) can save and restore execution states, both hardware and software challenges exist for energy harvesting powered embedded systems. On the hardware side, existing power detector only signals the ``poor'' quality of the harvested power based on a preset threshold voltage. The inappropriate setting of this threshold will make the NVP based embedded system suffer from either unnecessary checkpointing or checkpointing failures. On the software side, not all tasks can be checkpointed. Once the power is off, these tasks will have to restart from the beginning. In this paper, a task scheduler is proposed to maximize task progress by prioritizing tasks which cannot be checkpointed when power is weak so that they can finish before the power outage. To assist task scheduling, three additional modules including voltage monitor, checkpointing handler, and routine handler, are proposed. Experimental results show increased overall task progress and reduced energy consumption. Mimi Xie, Yongpan Liu, Yanzhi Wang 0001, Chun Jason Xue, Yuangang Wang, Yiran Chen 0001, Jingtong Hu |
LCTES | 2 |
| 2017 | Maximize energy utilization for ultra-low energy harvesting powered embedded systemsabstractEnergy harvesting systems become increasingly popular as power sources for many embedded systems. However, the harvesting power is often weak and the execution is frequently interrupted. Therefore, embedded systems have to work intermittently. To maintain the execution progress for better energy utilization, embedded systems need to save all execution states and program stacks into the non-volatile memory before each power failure, which is known as checkpointing. Then, embedded system waits until power comes back on again. By then the system can resume previous execution state. Nevertheless, frequent checkpointing incurs extra energy overhead. Besides, the charging efficiency of the storage capacitor reduces as the capacitor is charging up. These problems reduce the energy efficiency, resulting in less program execution progress. To alleviate these problems, this paper proposes three algorithms. First, a priority-based task scheduling (PTS) is proposed to prioritize the execution of tasks which have less checkpointing contents for a lower software overhead. Then, a tentative checkpointing avoidance (TCA) technique is proposed to avoid unnecessary checkpointing for further reduction of software overhead. Finally, a dynamic wake-up strategy (DWS) is proposed to wake up the system at proper voltages where the total hardware and software overheads are minimized. In this way, the system can achieve further energy efficiency improvements. The experiments on a real testbed show that the proposed prioritized algorithms enable the targeted embedded system to be resilient to extremely weak and intermittent power while achieving better energy utilization. Mimi Xie, Jingtong Hu |
RTCSA | 2 |
| 2017 | Stack-Size Sensitive On-Chip Memory Backup for Self-Powered Nonvolatile ProcessorsabstractWearable devices gain increasing popularity since they can collect important information for healthcare and well-being purposes. Compared with battery, energy harvesting is a better power source for these wearable devices due to many advantages. However, harvested energy is naturally unstable and program execution will be interrupted frequently. Nonvolatile processors demonstrate promising advantages to back up volatile state before the system energy is depleted. However, it also introduces non-negligible energy and area overhead. In this paper, we aim to reduce the amount of data that need to be backed up during a power failure. Based on the observation that stack size varies along program execution, we propose to analyze the application program and identify efficient backup positions, by which the stack content to back up can be significantly reduced. The evaluation results show an average of 45.7% reduction on nonvolatile stack size for stack backup, with 0.58% storage overhead. In the mean time, with the proposed schemes, the energy utilization and program forward progress can be greatly improved compared with instant backup. Mengying Zhao, Chenchen Fu, Qing'an Li, Mimi Xie, Yongpan Liu, Jingtong Hu, Zhiping Jia, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Exploiting Multiple Write Modes of Nonvolatile Main Memory in Embedded SystemsabstractExisting Nonvolatile Memories (NVMs) have many attractive features to be the main memory of embedded systems. These features include low power, high density, and better scalability. Recently, Multilevel Cell (MLC) NVM has gained more and more popularity as it can provide a higher density than the traditional Single-Level Cell (SLC) NVM. However, there are also drawbacks in MLC NVM, namely, limited write endurance and expensive write operation. These two drawbacks have to be overcome before MLC NVM can be practically adopted as the main memory. In MLC Nonvolatile Main Memory (NVMM), two different types of write operations with very diverse data retention times are allowed. The first type maintains data for years but takes a longer time to write and is detrimental to the endurance. The second type maintains data for a short period but takes a shorter time to write. By observing that much of the data written to main memory is temporary and does not need to last long during the execution of a program, in this article, we propose novel task scheduling and write operation selection algorithms to improve MLC NVMM endurance and program efficiency. An Integer Linear Programming (ILP) formulation is first proposed to obtain optimal results. Since ILP takes exponential time to solve, we also propose the Multiwrite Mode-Aware Scheduling (MMAS) algorithm to achieve a near-optimal solution in polynomial time. Additionally, the Dynamical Memory Block Screening (DMS) algorithm is proposed to achieve wear leveling. The experimental results demonstrate that the proposed techniques can greatly improve the lifetime of the MLC NVMM as well as the efficiency of the program. Mimi Xie, Chengmo Yang, Yiran Chen 0001, Jingtong Hu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Checkpoint-aware instruction scheduling for nonvolatile processor with multiple functional unitsabstractEmbedded systems powered with harvested energy experience frequent execution interruption due to unstable energy source. Nonvolatile (NV) register based processor is proposed to realize fast resume after power failure. The states in the volatile registers are checkpointed to NV registers. However, frequent checkpointing causes performance degradation and consumes excessive power. In this paper, we propose the checkpoint aware instruction scheduling (CAIS) algorithm to reduce the writes to NV registers. Experiments show that CAIS can improve performance and reduce power consumption. Mimi Xie, Jingtong Hu, Chengmo Yang, Yiran Chen 0001 |
ASP-DAC | 1 |
| 2015 | Fixing the broken time machine: consistency-aware checkpointing for energy harvesting powered non-volatile processorabstractEnergy harvesting has become a favorable alternative to batteries for wearable embedded systems since it is more environmental and user friendly. However, harvested energy is intrinsically unstable, which could frequently interrupt a processor's execution. To tackle this problem, non-volatile processors have been proposed to checkpoint the whole volatile processor state into attached non-volatile memories periodically. When power resumes, the processor can copy the checkpointed state back to volatile memories and continue execution. However, without careful consideration, the process of checkpointing and resuming could cause inconsistency among different memory addresses and lead to irreversible errors. In this paper, we present a consistency aware checkpointing scheme that ensures correctness for all checkpoints. The proposed technique efficiently identifies all possible inconsistency positions in programs and inserts auxiliary code to ensure correctness. Evaluation results show that the proposed checkpointing technique can successfully eliminate inconsistency errors and greatly reduce the checkpointing overhead. Mimi Xie, Mengying Zhao, Jingtong Hu, Yongpan Liu, Chun Jason Xue |
DAC | 1 |
| 2015 | Software assisted non-volatile register reduction for energy harvesting based cyber-physical system
Mengying Zhao, Qing'an Li, Mimi Xie, Yongpan Liu, Jingtong Hu, Chun Jason Xue |
DATE | 3 |
| 2015 | Nonvolatile main memory aware garbage collection in high-level language virtual machineabstractNon-volatile memories (NVMs) such as Phase Change Memory (PCM) have been considered as promising candidates of next generation main memory for embedded systems due to their attractive features. These features include low power, high density, and better scalability. However, most existing NVMs suffer from two drawbacks, namely, limited write endurance and expensive write operation in terms of both time and energy. These problems are worsen when modern high-level languages employ virtual machine with garbage collector that generates a large amount of extra writes on non-volatile main memory. To tackle this challenge, this paper proposes three techniques: Living Objects Remapping (LORE), Dead Object Stamping (DOS), and Smart Wiping with Maximum Likelihood Estimation (SMILE) to reduce the unnecessary writes when garbage collector handles objects. The experimental results show that the proposed techniques not only significantly reduce the writes during each garbage collection cycle but also greatly improve the performance of virtual machine. Mimi Xie, Chengmo Yang, Zili Shao, Jingtong Hu |
EMSOFT | 2 |
| 2015 | Low Overhead Software Wear Leveling for Hybrid PCM + DRAM Main Memory on Embedded SystemsabstractPhase change memory (PCM) is a promising DRAM replacement in embedded systems due to its attractive characteristics, such as low-cost, shock-resistivity, nonvolatility, high density, and low leakage power. However, relatively low endurance has limited its practical applications. In this paper, in addition to existing hardware level optimizations, we propose software enabled wear-leveling techniques to further extend PCMs lifetime when it is adopted in embedded systems. Most existing software optimization techniques focus on reducing the total number of writes to PCM, but none of them consider wear leveling, in which the writes are distributed more evenly over the PCM. An integer linear programming formulation and a polynomial-time algorithm, the software wear-leveling algorithm, are proposed in this paper to achieve wear leveling without hardware overhead. According to the experimental results, the proposed techniques can reduce the number of writes on the most-written addresses by more than 80% when compared with a greedy algorithm, and by more than 60% when compared with the existing optimal data allocation algorithm with under 6% memory access overhead. Jingtong Hu, Mimi Xie, Chun Jason Xue, Qingfeng Zhuge, Edwin H.-M. Sha |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Wear-leveling for PCM main memory on embedded system via page management and process schedulingabstractPhase Change Memory (PCM) has been considered as a leading candidate to replace the traditional DRAM in embedded systems due to its promising characteristics such as low leakage power, low cost, non-volatility, and high scalability. One of the constraints that undermine the credential of PCM as main memory is its limited write endurance. In this paper, we develop wear-leveling techniques purely on the Operating System (OS) level to extend lifetime of PCM. Without extra hardware support, OS management is more flexible to be integrated into existing embedded systems. To achieve wear-leveling, the Periodical Page Swapping (PPS), Rearrangement Inequality Based Page Allocation (RIPA), and Write Intensity Based Process Scheduling (WIPS) algorithms are proposed in this paper on OS level. The experimental results show that the proposed techniques can significantly extend the lifetime of PCM main memory. Mimi Xie, Jingtong Hu, Meikang Qiu, Qingfeng Zhuge |
RTCSA | 2 |
| 2014 | Non-volatile registers aware instruction selection for embedded systemsabstractIt is common that embedded systems are powered by limited and unstable power supply. In order to improve the reliability of embedded systems against unstable power supply, non-volatile memory (e.g. FRAM) based registers are proposed for embedded processors. FRAM-based registers have many advantages over traditional CMOS-based volatile registers such as non-volatility and power-economy. However, similar to other non-volatile memories (NVM), write operations to FRAM consume more time and power compared with read operations and limit the lifetime of the registers. Existing compiler optimization techniques never take the writes to registers into consideration. Therefore, code generated by a traditional compiler has an adverse effect on processors with non-volatile registers. This paper aims at improving the lifetime and efficiency of non-volatile registers based embedded processors by generating NV register friendly code. To achieve the goal, in this paper, we investigate the usage of memory access instructions and propose the NV Register Aware Instruction Selection (NAIS) algorithm to reduce the write operations on non-volatile registers. According to the experimental results, the proposed algorithm can reduce the writes on NV registers by 66.89% on average when compared with GCC [1]. Thus the lifetime of NV registers is extended to 2 times as long as before on average. The time cost is reduced by 56.68% and the energy consumption is reduced by 59.76% on average. Mimi Xie, Jingtong Hu, Chun Jason Xue, Qingfeng Zhuge |
RTCSA | 1 |