Qinru Qiu

dblp:64/1120 · DBLP profile ↗
← Back
118ranked-venue papers
16as first author
17since 2021 · last 2026
0000-0003-2546-0655ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 78 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 28 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 12 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 since 2021Computer networks · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning to Sense Without ADCs: Exploiting Phasic Responses from Diffusive Memristors
abstract
This paper presents a neuromorphic system that integrates a low-cost, energy-efficient, and bio-realistic spike encoder with a spiking neural network (SNN). The two are jointly optimized via an online learning algorithm to enable temporal pattern detection across multi-channel sensor inputs. At the core of the system is a novel analog-to-spike converter based on diffusive memristors, which replaces conventional ADCs and provide a fundamentally different encoding scheme with substantially lower power consumption and reduced device footprint. However, the inherent variability of diffusive memristor devices introduces significant challenges for both frontend and backend design. To address this, we propose an Adaptive Gain Unit (AGU) and a frontend-backend co-adaptation strategy that supports real-time online learning updates of both the AGU and the classifier network. Experimental results on nine publicly available time-series datasets show that this adaptation improves accuracy by 4.19% on average. Furthermore, compared to conventional 8-bit ADCs and state-of-the-art level-crossing ADCs, the diffusive memristor–based system achieves comparable classification accuracy while offering orders of magnitude lower power consumption and improved area efficiency.
Zhenhang Zhang, Jingang Jin, Rui Zuo, J. Joshua Yang, Qinru Qiu
DATE8
2026 SAFER-AiD: Saccade-Assisted Foveal-peripheral vision Enhanced Reconstruction for Adversarial Defense
abstract
Adversarial attacks significantly challenge the safe deployment of deep learning models, particularly in real-world applications. Traditional defenses often rely on computationally intensive optimization (e.g., adversarial training or data augmentation) to improve robustness, whereas the human visual system achieves inherent robustness to adversarial perturbations through evolved biological mechanisms. We hypothesize that attention guided non-homogeneous sparse sampling and predictive coding plays a key role in this robustness. To test this hypothesis, we propose a novel defense framework incorporating three key biological mechanisms: foveal-peripheral processing, saccadic eye movements, and cortical filling-in. Our approach employs reinforcement learning-guided saccades to selectively capture multiple foveal-peripheral glimpses, which are integrated into a reconstructed image before classification. This biologically inspired preprocessing effectively mitigates adversarial noise, preserves semantic integrity, and notably requires no retraining or fine-tuning of downstream classifiers, enabling seamless integration with existing systems. Experiments on the ImageNet dataset demonstrate that our method improves system robustness across diverse classifiers and attack types, while significantly reducing training overhead compared to both biologically and non-biologically inspired defense techniques. Code available at https://github.com/jliu206/SAFER-AiD.
Daniel Tso, Yiming Bu, Qinru Qiu
WACV4
2025 Why the Agent Made that Decision: Contrastive Explanation Learning for Reinforcement Learning
abstract
Reinforcement learning (RL) has demonstrated remarkable success in solving complex decision-making problems, yet its adoption in critical domains is hindered by the lack of interpretability in its decision-making processes. Existing explainable AI (xAI) approaches often fail to provide meaningful explanations for RL agents, particularly because they overlook the contrastive nature of human reasoning—answering "why this action instead of that one?" To address this gap, we propose a novel framework of contrastive learning to explain RL selected actions, named VisionMask. VisionMask is trained to generate explanations by explicitly contrasting the agent's chosen action with alternative actions in a given state using a self-supervised manner. We demonstrate the efficacy of our method through experiments across diverse RL environments, evaluating it in terms of faithfulness, robustness and complexity. Our results show that VisionMask significantly improves human understanding of agent behavior while maintaining accuracy and fidelity. Furthermore, we present examples illustrating how VisionMask can be used for counterfactual analysis. This work bridges the gap between RL and xAI, paving the way for safer and more interpretable RL systems.
Rui Zuo, Simon Khan, Zifan Wang 0004, Garrett E. Katz, Qinru Qiu
IJCAI5
2025 Event Prediction and Quality-adaptive Attention for Energy and Data Efficient Dynamic Visual Sensing
abstract
Most machine learning models require a significant amount of sensor data to perceive and interpret situations accurately. This data often includes spatial and temporal redundancy, leading to unnecessary consumption of sensor and processor resources. Dynamic Vision Sensors (DVS) partially address this by detecting only changes within a scene; however, this approach can be further optimized. Inspired by the predictive sensing capabilities of the human brain, this work leverages spike-driven transformer and predictive temporal attention to suppress sensor inputs when they are predictable. A quality-adaptive temporal attention mechanism is presented to monitor the prediction quality and guide sensor gating. The system not only saves energy by gating the physical sensors but also leverages predicted sensor data, which contains more (predictable) salient features and less (unpredictable) noise, for energy reduction in the downstream tasks. Experiment results show that the proposed framework reduces sensor workload by approximately 66.4%. It also reduces spike activities by 39.4%, which leads to a 31.9% energy reduction during inference, with marginal performance loss.
Yiming Bu, Simon Khan, Qinru Qiu
IJCNN4
2025 Linearithmic Clean-up for Vector-Symbolic Key-Value Memory with Kroneker Rotation Products
abstract
A computational bottleneck in current Vector-Symbolic Architectures (VSAs) is the "clean-up" step, which decodes the noisy vectors retrieved from the architecture. Clean-up typically compares noisy vectors against a "codebook" of prototype vectors, incurring computational complexity that is quadratic or similar. We present a new codebook representation that supports efficient clean-up, based on Kroneker products of rotation-like matrices. The resulting clean-up time complexity is linearithmic, i.e. $\mathcal{O}(N\text{log}N)$, where $N$ is the vector dimension and also the number of vectors in the codebook. Clean-up space complexity is $\mathcal{O}(N)$. Furthermore, the codebook is not stored explicitly in computer memory: It can be represented in $\mathcal{O}(\text{log}N)$ space, and individual vectors in the codebook can be materialized in $\mathcal{O}(N)$ time. At the same time, asymptotic memory capacity remains comparable to standard approaches. Computer experiments confirm these results, demonstrating several orders of magnitude more scalability than baseline VSA techniques.
Ruipeng Liu, Qinru Qiu, Simon Khan, Garrett E. Katz
NeSy2
2024 SOLSA: Neuromorphic Spatiotemporal Online Learning for Synaptic Adaptation
abstract
Spiking neural networks (SNNs) are bio-plausible computing models with high energy efficiency. The temporal dynamics of neurons and synapses enable them to detect temporal patterns and generate sequences. While Backpropagation Through Time (BPTT) is traditionally used to train SNNs, it is not suitable for online learning of embedded applications due to its high computation and memory cost as well as extended latency. In this work, we present Spatiotemporal Online Learning for Synaptic Adaptation (SOLSA), which is specifically designed for online learning of SNNs composed of Leaky Integrate and Fire (LIF) neurons with exponentially decayed synapses and soft reset. The algorithm not only learns the synaptic weight but also adapts the temporal filters associated to the synapses. Compared to the BPTT algorithm, SOLSA has much lower memory requirement and achieves a more balanced temporal workload distribution. Moreover, SOLSA incorporates enhancement techniques such as scheduled weight update, early stop training and adaptive synapse filter, which speed up the convergence and enhance the learning performance. When compared to other non-BPTT based SNN learning, SOLSA demonstrates an average learning accuracy improvement of 14.2%. Furthermore, compared to BPTT, SOLSA achieves a 5% higher average learning accuracy with a 72% reduction in memory cost.
Zhenhang Zhang, Jingang Jin, Haowen Fang, Qinru Qiu
ASPDAC4
2024 Prompt-Based Domain Incremental Learning with Modular Classification Layer
abstract
Continual learning is referred to as machine learning model’s ability to learn a sequence of tasks or data over time without forgetting previously learned knowledge. In particular, we focus on domain incremental learning where the model trained on one domain, e.g., real-life photos, has to be augmented to accommodate data from new domains, e.g., cartoon images. Typical incremental learning relies on rehearsal-based methods, which store trained samples in a buffer and replay them during training alongside new data. This results in significant memory overhead and raise concerns about data privacy. Recently, prompt-based methods address these challenges and outperform them by utilizing the pre-trained Vision Transformer (ViT) and replacing the replay buffer with a prompt pool. However, existing prompt-based models fail to capture domain-specific knowledge and perform poorly in domain incremental learning. In this paper, we propose “domain prompt incremental learning via dynamic neural network”, which combines the advantages of architecture-based and prompt-based methods. Specifically, our framework maintains two types of prompt: a instance-level prompt that improves the model’s generalization ability is shared across all input samples; and a domain prompt that encodes domain-specific knowledge is assigned for each task. Furthermore, a separated classification head is trained for each domain so that the model has a pre-trained ViT and an ensemble of classification layers, one for each domain. The experimental results shows that our approach outperforms state-of-the-art methods by 2.3% in average on two domain incremental learning (DIL) benchmarks.
Boyu Wang 0005, Qinru Qiu
ECAI3
2024 Improved Efficiency Based on Learned Saccade and Continuous Scene Reconstruction From Foveated Visual Sampling
abstract
High accuracy, low latency and high energy efficiency represent a set of contradictory goals when searching for system solutions for image classification and detection. While high-quality images naturally result in more precise detection and classification, they also result in a heavier computational workload for imaging and processing, reduce camera refresh rates, and increase the volume of data communication between the camera and processor. Taking inspiration from the foveal-peripheral sampling mechanism, saccade mechanism observed in the human visual system and the filling-in phenomena of brain, we have developed an active scene reconstruction architecture based on multiple foveal views. This model stitches together information from foveal and peripheral vision, which are sampled from multiple glances. Assisted by a reinforcement learning-based saccade mechanism, our model reduces the required input pixels by over 90\% per frame while maintaining the same level of performance in image recognition as with the original images. We evaluated the effectiveness of our model using the GTSRB dataset and the ImageNet dataset. Using an equal number of input pixels, our study demonstrates a 5\% higher image recognition accuracy compared to state-of-the-art foveal-peripheral vision systems. Furthermore, we demonstrate that our foveal sampling/saccadic scene reconstruction model exhibits significantly lower complexity and higher data efficiency during the training phase compared to existing approaches.
Yiming Bu, Daniel Tso, Qinru Qiu
ICLR4
2023 Multi-Agent Cooperative Games Using Belief Map Assisted Training
abstract
In a multi-agent system, agents share their local observations to gain global situational awareness for decision making and collaboration using a message passing system. When to send a message, how to encode a message, and how to leverage the received messages directly affect the effectiveness of the collaboration among agents. When training a multi-agent cooperative game using reinforcement learning (RL), the message passing system needs to be optimized together with the agent policies. This consequently increases the model’s complexity and poses significant challenges to the convergence and performance of learning. To address this issue, we propose the Belief-map Assisted Multi-agent System (BAMS), which leverages a neuro-symbolic belief map to enhance training. The belief map decodes the agent’s hidden state to provide a symbolic representation of the agent’s understanding of the environment and other agents’ status. The simplicity of symbolic representation allows the gathering and comparison of the ground truth information with the belief, which provides an additional channel of feedback for the learning. Compared to the sporadic and delayed feedback coming from the reward in RL, the feedback from the belief map is more consistent and reliable. Agents using BAMS can learn a more effective message passing network to better understand each other, resulting in better performance in the game. We evaluate BAMS’s performance in a cooperative predator and prey game with varying levels of map complexity and compare it to previous multi-agent message passing models. The simulation results showed that BAMS reduced training epochs by 66%, and agents who apply the BAMS model completed the game with 34.62% fewer steps on average.
Qinwei Huang, Alex B. Wu, Simon Khan, Hai Li 0001, Qinru Qiu
ECAI6
2023 Catch You if Pay Attention: Temporal Sensor Attack Diagnosis Using Attention Mechanisms for Cyber-Physical Systems
abstract
In Cyber-Physical Systems (CPS), sensor data integrity is crucial since acting on malicious sensor data can cause serious consequences, given the tight coupling between cyber components and physical systems. While extensive works focus on sensor attack detection, attack diagnosis that aims to find out when the attack starts has not been well studied yet. This temporal sensor attack diagnosis problem is equally important because many recovery methods rely on the accurate determination of trustworthy historical data. To address this problem, we propose a lightweight data-driven solution to achieve real-time sensor attack diagnosis. Our novel solution consists of five modules, with the attention and diagnosis ones as the core. The attention module not only helps accurately predict future sensor measurements but also computes statistical attention scores for the diagnosis module. Based on our unique observation that the score fluctuates sharply once an attack launches, the diagnosis module determines the onset of an attack through monitoring the fluctuation. Evaluated on high-dimensional high-fidelity simulators and a testbed, our solution demonstrates robust and accurate temporal diagnosis results while incurring millisecond-level computational overhead on Raspberry Pi.
Zifan Wang 0004, Lin Zhang 0039, Qinru Qiu, Fanxin Kong
RTSS3
2022 Neural Network Pruning and Fast Training for DRL-based UAV Trajectory Planning
abstract
Deep reinforcement learning (DRL) has been applied for optimal control of autonomous UAV trajectory generation. The energy and payload capacity of small UAVs impose constraints on the complexity and size of the neural network. While Model compression has the potential to optimize the trained neural network model for efficient deployment on em-bedded platforms, pruning a neural network for DRL is more difficult due to the slow convergence in the training before and after pruning. In this work, we focus on improving the speed of DRL training and pruning. New reward function and action exploration are first introduced, resulting in convergence speedup by 34.14%. The framework that integrates pruning and DRL training is then presented with an emphasize on how to reduce the training cost. The pruning does not only improve computational performance of inference, but also reduces the training effort with-out compromising the quality of the trajectory. Finally, experimental results are presented. We show that the integrated training and pruning framework reduces 67.16% of the weight and improves trajectory success rate by 1.7%. It achieves a 4.43x reduction of the floating-point operations for the inference, resulting a measured 41.85% run time reduction.
Yilan Li, Haowen Fang, Qinru Qiu
ASP-DAC5
2022 Guest Editorial: IEEE TC Special Issue On Software, Hardware and Applications for Neuromorphic Computing
abstract
The papers in this special section focus on software, hardware, and computer applications for neuromorphic computing. Inspired by biological neural systems, neuromorphic computing has drawn much attention for its great potential of achieving machine intelligence at extremely low energy dissipation. Bio-inspired computing models have been investigated for information encoding, sparse representation, event driven communication/computation, and online learning. This new computing paradigm triggered a recent wave of innovations in software and hardware architecture and emerging device technology, which consequently enabled many novel applications. This is an exemplar of research area where the application, computing model, architecture and circuit level design are tightly coupled to deliver unprecedented functionality and energy efficiency.
Yiran Chen 0001, Qinru Qiu
IEEE Trans. Computers2
2021 Neuromorphic Algorithm-hardware Codesign for Temporal Pattern Learning
abstract
Neuromorphic computing and spiking neural networks (SNN) mimic the behavior of biological systems and have drawn interest for their potential to perform cognitive tasks with high energy efficiency. However, some factors such as temporal dynamics and spike timings prove critical for information processing but are often ignored by existing works, limiting the performance and applications of neuromorphic computing. On one hand, due to the lack of effective SNN training algorithms, it is difficult to utilize the temporal neural dynamics. Many existing algorithms still treat neuron activation statistically. On the other hand, utilizing temporal neural dynamics also poses challenges to hardware design. Synapses exhibit temporal dynamics, serving as memory units that hold historical information, but are often simplified as a connection with weight. Most current models integrate synaptic activations in some storage medium to represent membrane potential and institute a hard reset of membrane potential after the neuron emits a spike. This is done for its simplicity in hardware, requiring only a “clear” signal to wipe the storage medium, but destroys temporal information stored in the neuron.In this work, we derive an efficient training algorithm for Leaky Integrate and Fire neurons, which is capable of training a SNN to learn complex spatial temporal patterns. We achieved competitive accuracy on two complex datasets. We also demonstrate the advantage of our model by a novel temporal pattern association task. Codesigned with this algorithm, we have developed a CMOS circuit implementation for a memristor-based network of neuron and synapses which retains critical neural dynamics with reduced complexity. This circuit implementation of the neuron model is simulated to demonstrate its ability to react to temporal spiking patterns with an adaptive threshold.
Haowen Fang, Brady Taylor, Ziru Li, Zaidao Mei, Hai Li 0001, Qinru Qiu
DAC6
2021 In-Hardware Learning of Multilayer Spiking Neural Networks on a Neuromorphic Processor
abstract
Although widely used in machine learning, backpropagation cannot directly be applied to SNN training and is not feasible on a neuromorphic processor that emulates biological neuron and synapses. This work presents a spike-based backpropagation algorithm with biological plausible local update rules and adapts it to fit the constraint in a neuromorphic hardware. The algorithm is implemented on Intel’s Loihi chip enabling low power in-hardware supervised online learning of multilayered SNNs for mobile applications. We test this implementation on MNIST, Fashion-MNIST, CIFAR-10 and MSTAR datasets with promising performance and energy-efficiency, and demonstrate a possibility of incremental online learning with the implementation.
Amar Shrestha, Haowen Fang, Daniel Patrick Rider, Zaidao Mei, Qinru Qiu
DAC5
2021 1S1R-Based Stable Learning through Single-Spike-Encoded Spike-Timing-Dependent Plasticity
abstract
Spike-timing-dependent plasticity (STDP) is emerging as a simple and biologically-plausible approach to learning, and specialized digital implementations are readily available. Memristor technology has been embraced as a much denser solution than digital static random-access memory (SRAM) implementations of STDP synapses, with plasticity capabilities built into the physics of these devices. One-selector-one-memristor (1S1R) arrays using volatile memristor devices as selectors are capable of the desired synaptic behavior using efficient spike-events, but previous literature has only explored the dynamics of single 1S1R synapses, or groups of synapses for single neurons. When placed in the context of an SNN, unintentional synapse disturbances are revealed that must be addressed. We present1a technique for STDP-based learning, enabled for dense 1S1R technology and utilizing efficient single-spike encoding. This technique leverages the array's dynamics to produce models that are stable, resilient to noise, and power-efficient.
Brady Taylor, Amar Shrestha, Qinru Qiu, Hai Li 0001
ISCAS3
2021 Introduction of Special Issue on Hardware and Algorithms for Efficient Machine Learning-Part 1
abstract
No abstract available.
Yiran Chen 0001, Qinru Qiu, Yingyan (Celine) Lin
ACM J. Emerg. Technol. Comput. Syst.2
2021 Introduction to the Special Issue on Hardware and Algorithms for Efficient Machine Learning - Part 2
abstract
No abstract available.
Yiran Chen 0001, Qinru Qiu, Yingyan (Celine) Lin
ACM J. Emerg. Technol. Comput. Syst.2
2020 Embedding Compression with Isotropic Iterative Quantization
abstract
Continuous representation of words is a standard component in deep learning-based NLP models. However, representing a large vocabulary requires significant memory, which can cause problems, particularly on resource-constrained platforms. Therefore, in this paper we propose an isotropic iterative quantization (IIQ) approach for compressing embedding vectors into binary ones, leveraging the iterative quantization technique well established for image retrieval, while satisfying the desired isotropic property of PMI based models. Experiments with pre-trained embeddings (i.e., GloVe and HDC) demonstrate a more than thirty-fold compression ratio with comparable and sometimes even improved performance over the original real-valued embedding vectors.
Siyu Liao, Jie Chen 0007, Yanzhi Wang 0001, Qinru Qiu, Bo Yuan 0001
AAAI4
2020 Database and Benchmark for Early-stage Malicious Activity Detection in 3D Printing
abstract
Increasing malicious users have sought practices to leverage 3D printing technology to produce unlawful tools in criminal activities. It is of vital importance to enable 3D printers to identify the objects to be printed and terminate at early stage if illegal objects are identified. Deep learning yields significant rises in performance in the object recognition tasks. However, the lack of large-scale databases in 3D printing domain stalls the advancement of automatic illegal weapon recognition. This paper presents a new 3D printing image database, namely C3PO, which compromises two subsets for the different system working scenarios. We extract images from the numerical control programming code files of 22 3D models, and then categorize the images into 10 distinct labels. These two sets are designed for identifying: (i). printing knowledge source (G-code) at beginning of manufacturing, (ii). printing procedure during manufacturing. Importantly, we demonstrate that the weapons can be recognized in either scenario using deep learning based approaches using our proposed database. The quantitative results are promising, and the future exploration of the database and the crime prevention in 3D printing are demanding tasks.
Zhe Li 0001, Hongjia Li 0003, Qiyuan An, Qinru Qiu, Wenyao Xu, Yanzhi Wang 0001
ASP-DAC5
2020 Encoding, Model, and Architecture: Systematic Optimization for Spiking Neural Network in FPGAs
abstract
Spiking neural network (SNN) has drawn research interests as it mimics dynamic activities of human brain and has the potential to perform real-time cognitive tasks. However, latency, throughput and flexibility of existing hardware implemented SNNs are limited. The conventional rate coding is inefficient in terms of accuracy and latency. Oversimplified SNN models adopted by neuromorphic hardware discard characteristics such as neuron dynamics and filter effects etc., which are critical for neural information processing. Recent research advancements show that the potential of SNN can be better utilized by moving beyond rate-based model and considering temporal information embedded in the spike sequences. However, these works employ complex biologically realistic SNN models, posing challenges to hardware complexity. Furthermore, most existing neuromorphic hardware are developed for specific SNN models, or aiming at replicating biological behaviors. There is a lack of general methodology for SNN design optimization. Novel hardware architecture and systematic optimization techniques are required for efficient FPGA implementation and support flexible SNN models.
Haowen Fang, Zaidao Mei, Amar Shrestha, Yilan Li, Qinru Qiu
ICCAD6
2020 MAGNet: Multi-Region Attention-Assisted Grounding of Natural Language Queries at Phrase Level
abstract
Grounding free-form textual queries necessitates an understanding of these textual phrases and its relation to the visual cues to reliably reason about the described locations. Spatial attention networks are known to learn this relationship and focus its gaze on salient objects in the image. Thus, we propose to utilize spatial attention networks for image-level visual-textual fusion preserving local (word) and global (phrase) information to refine region proposals with an in-network Region Proposal Network (RPN) and detect single or multiple regions for a phrase query. We focus only on the phrase query - ground truth pair (referring expression) for a model independent of the constraints of the datasets i.e. additional attributes, context etc. For such referring expression dataset ReferIt game, our Multi-region Attention-assisted Grounding network (MAGNet) achieves over 12% improvement over the state-of-the-art. Without the context from image captions and attribute information in Flickr30k Entities, we still achieve competitive results compared to the state-of-the-art.
Amar Shrestha, Krittaphat Pugdeethosapol, Haowen Fang, Qinru Qiu
ICPR4
2020 Exploiting Neuron and Synapse Filter Dynamics in Spatial Temporal Learning of Deep Spiking Neural Network
abstract
The recently discovered spatial-temporal information processing capability of bio-inspired Spiking neural networks (SNN) has enabled some interesting models and applications. However designing large-scale and high-performance model is yet a challenge due to the lack of robust training algorithms. A bio-plausible SNN model with spatial-temporal property is a complex dynamic system. Synapses and neurons behave as filters capable of preserving temporal information. As such neuron dynamics and filter effects are ignored in existing training algorithms, the SNN downgrades into a memoryless system and loses the ability of temporal signal processing. Furthermore, spike timing plays an important role in information representation, but conventional rate-based spike coding models only consider spike trains statistically, and discard information carried by its temporal structures. To address the above issues, and exploit the temporal dynamics of SNNs, we formulate SNN as a network of infinite impulse response (IIR) filters with neuron nonlinearity. We proposed a training algorithm that is capable to learn spatial-temporal patterns by searching for the optimal synapse filter kernels and weights. The proposed model and training algorithm are applied to construct associative memories and classifiers for synthetic and public datasets including MNIST, NMNIST, DVS 128 etc. Their accuracy outperforms state-of-the-art approaches.
Haowen Fang, Amar Shrestha, Qinru Qiu
IJCAI4
2020 Multivariate Time Series Classification Using Spiking Neural Networks
abstract
There is an increasing demand to process streams of temporal data in energy-limited scenarios such as embedded devices, driven by the advancement and expansion of Internet of Things (IoT) and Cyber-Physical Systems (CPS). Spiking neural network has drawn attention as it enables low power consumption by encoding and processing information as sparse spike events, which can be exploited for event-driven computation. Recent works also show SNNs' capability to process spatial temporal information. Such advantages can be exploited by power-limited devices to process real-time sensor data. However, most existing SNN training algorithms focus on vision tasks and temporal credit assignment is not addressed. Furthermore, widely adopted rate encoding ignores temporal information, hence it's not suitable for representing time series. In this work, we present an encoding scheme to convert time series into sparse spatial temporal spike patterns. A training algorithm to classify spatial temporal patterns is also proposed. Proposed approach is evaluated on multiple time series datasets in the UCR repository and achieved performance comparable to deep neural networks.
Haowen Fang, Amar Shrestha, Qinru Qiu
IJCNN3
2020 Automatic Image Labeling with Click Supervision on Aerial Images
abstract
Manually generating annotated bounding boxes for object detection is time consuming. Although human-annotation is the most accurate approach, machine learning models can provide additional assistance. In this paper, we propose a human in a loop automatic image labeling framework focusing on aerial images with less features for detection. The proposed model consists of two main parts, prediction model and adjustment model. The user first provides click location to prediction model to generate a bounding box of a specific object. The bounding box is then fine-tuned by the adjustment model for more accurate size and location. A feedback and retrain mechanism is implemented that allows the users to manually adjust the generated bounding box and provide feedback to incrementally train the adjustment network during runtime. This unique online learning feature enables user to generalize existing model to target classes not initially presented in the training set, and gradually improves the specificity of the model to those new targets online. We demonstrate promising results on Neovision 2 Heli dataset. Compared to the state-of-the-art method, our prediction model achieves a higher detection rate, and our adjustment model improves the IOU by up to 45%.
Krittaphat Pugdeethosapol, Morgan Bishop, Dennis Bowen, Qinru Qiu
IJCNN4
2020 GISNet: Graph-Based Information Sharing Network For Vehicle Trajectory Prediction
abstract
The trajectory prediction is a critical and challenging problem in the design of an autonomous driving system. Many AI-oriented companies, such as Google Waymo, Uber and DiDi, are investigating more accurate vehicle trajectory prediction algorithms. However, the prediction performance is governed by lots of entangled factors, such as the stochastic behaviors of surrounding vehicles, historical information of self-trajectory, and relative positions of neighbors, etc. In this paper, we propose a novel graph-based information sharing network (GISNet) that allows the information sharing between the target vehicle and its surrounding vehicles. Meanwhile, the model encodes the historical trajectory information of all the vehicles in the scene. Experiments are carried out on the public NGSIM US-101 and I-80 Dataset and the prediction performance is measured by the Root Mean Square Error (RMSE). The quantitative and qualitative experimental results show that our model significantly improves the trajectory prediction accuracy, by up to 50.00%, compared to existing models.
Haowen Fang, Qinru Qiu
IJCNN4
2020 Mission-Aware Spatio-Temporal Deep Learning Model for UAS Instantaneous Density Prediction
abstract
The number of daily sUAS operations in uncontrolled low altitude airspace is expected to reach into the millions in a few years. Therefore, UAS density prediction has become an emerging and challenging problem. In this paper, a deep learning-based UAS instantaneous density prediction model is presented. The model takes two types of data as input: 1) the historical density generated from the historical data, and 2) the future sUAS mission information. The architecture of our model contains four components: Historical Density Formulation module, UAS Mission Translation module, Mission Feature Extraction module, and Density Map Projection module. The training and testing data are generated by a python based simulator which is inspired by the multi-agent air traffic resource usage simulator (MATRUS) framework. The quality of prediction is measured by the correlation score and the Area Under the Receiver Operating Characteristics (AUROC) between the predicted value and simulated value. The experimental results demonstrate outstanding performance of the deep learning-based UAS density predictor. Compared to the baseline models, for simplified traffic scenario where no-fly zones and safe distance among sUASs are not considered, our model improves the prediction accuracy by up to 15.2% and its correlation score reaches 0.947. In a more realistic scenario, where the no-fly zone avoidance and the safe distance among sUASs are maintained using A* routing algorithm, our model can still achieve 0.822 correlation score. Meanwhile, the AUROC can reach 0.951 for the hot spot prediction.
Wentian Bai, Wentan Bai, Carlos E. Caicedo Bastidas, Mustafa Cenk Gursoy, Qinru Qiu
IJCNN7
2019 E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs
abstract
Recurrent Neural Networks (RNNs) are becoming increasingly important for time series-related applications which require efficient and real-time implementations. The two major types are Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) networks. It is a challenging task to have real-time, efficient, and accurate hardware RNN implementations because of the high sensitivity to imprecision accumulation and the requirement of special activation function implementations. Recently two works have focused on FPGA implementation of inference phase of LSTM RNNs with model compression. First, ESE uses a weight pruning based compressed RNN model but suffers from irregular network structure after pruning. The second work C-LSTM mitigates the irregular network limitation by incorporating block-circulant matrices for weight matrix representation in RNNs, thereby achieving simultaneous model compression and acceleration. A key limitation of the prior works is the lack of a systematic design optimization framework of RNN model and hardware implementations, especially when the block size (or compression ratio) should be jointly optimized with RNN type, layer size, etc. In this paper, we adopt the block-circulant matrixbased framework, and present the Efficient RNN (E-RNN) framework for FPGA implementations of the Automatic Speech Recognition (ASR) application. The overall goal is to improve performance/energy efficiency under accuracy requirement. We use the alternating direction method of multipliers (ADMM) technique for more accurate block-circulant training, and present two design explorations providing guidance on block size and reducing RNN training trials. Based on the two observations, we decompose E-RNN in two phases: Phase I on determining RNN model to reduce computation and storage subject to accuracy requirement, and Phase II on hardware implementations given RNN model, including processing element design/optimization, quantization, activation implementation, etc. 1 Experimental results on actual FPGA deployments show that E-RNN achieves a maximum energy efficiency improvement of 37.4× compared with ESE, and more than 2× compared with C-LSTM, under the same accuracy.
Zhe Li 0001, Caiwen Ding, Siyue Wang, Wujie Wen, Youwei Zhuo, Qinru Qiu, Wenyao Xu, Xue Lin 0001, Xuehai Qian, Yanzhi Wang 0001
HPCA7
2019 An Event-driven Neuromorphic System with Biologically Plausible Temporal Dynamics
abstract
Driven by the expanse of Internet of Things (IoT) and Cyber-Physical Systems (CPS), there is an increasing demand to process streams of temporal data on embedded devices with limited energy and power resources. Among all potential solutions, neuromorphic computing with spiking neural networks (SNN) that mimic the behavior of brain, have recently been placed at the forefront. Encoding information into sparse and distributed spike events enables low-power implementations, and the complex spatial temporal dynamics of synapses and neurons enable SNNs to detect temporal pattern. However, most existing hardware SNN implementations use simplified neuron and synapse models ignoring synapse dynamic, which is critical for temporal pattern detection and other applications that require temporal dynamics. To adopt a more realistic synapse model in neuromorphic platform its significant computation overhead must be addressed. In this work, we propose an FPGA-based SNN with biologically realistic neuron and synapse for temporal information processing. An encoding scheme to convert continuous real-valued information into sparse spike events is presented. The event-driven implementation of synapse dynamic model and its hardware design that is optimized to exploit the sparsity are also presented. Finally, we train the SNN on various temporal pattern-learning tasks and evaluate its performance and efficiency as compared to rate-based models and artificial neural networks on different embedded platforms. Experiments show that our work can achieve 10X speed up and 196X gains in energy efficiency compared with GPU.
Haowen Fang, Amar Shrestha, Yilan Li, Qinru Qiu
ICCAD5
2019 Fast and Accurate Trajectory Tracking for Unmanned Aerial Vehicles based on Deep Reinforcement Learning
abstract
Continuous trajectory control of fixed-wing unmanned aerial vehicles (UAVs) is complicated when considering hidden dynamics. Due to UAV multi degrees of freedom, tracking methodologies based on conventional control theory, such as Proportional-Integral-Derivative (PID) has limitations in response time and adjustment robustness, while a model based approach that calculates the force and torques based on UAV's current status is complicated and rigid. We present an actor-critic reinforcement learning framework that controls UAV trajectory through a set of desired waypoints. A deep neural network is constructed to learn the optimal tracking policy and reinforcement learning is developed to optimize the resulting tracking scheme. The experimental results show that our proposed approach can achieve 58.14% less position error, 21.77% less system power consumption and 9.23% faster attainment than the baseline. The actor network consists of only linear operations, hence Field Programmable Gate Arrays (FPGA) based hardware acceleration can easily be designed for energy efficient real-time control.
Yilan Li, Hongjia Li 0003, Zhe Li 0001, Haowen Fang, Amit K. Sanyal, Yanzhi Wang 0001, Qinru Qiu
RTCSA7
2019 Temporal and Spatial Routing for Large Scale Safe and Connected UAS Traffic Management in Urban Areas
abstract
Small Unmanned Aircraft Systems (sUAS) will be an important component of the smart city and intelligent transportation environments of the near future. The demand for sUAS related applications, such as commercial delivery and land surveying, is expected to grow rapidly in next few years. In general, sUAS traffic scheduling and management functions are needed to coordinate the launching of sUAS from different launch sites and plan their trajectories to avoid conflict while considering several other constraints such as expected arrival time, minimum flight energy, and availability of communication resources. However, as the airbone sUAS density grows in a certain area, it is difficult to foresee the potential airspace and communications resource conflicts and make immediate decisions to avoid them. To address this challenge, we present a temporal and spatial routing algorithm for sUAS trajectory management in a high density urban area. It plans sUAS movements in a spatial and temporal maze with the consideration of obstacles that are either static or dynamic in time. The routing allows the sUAS to avoid static no-fly areas (i.e. static obstacles) or other in-flight sUAS and areas that have congested communication resources (i.e. dynamic obstacles). The algorithm is evaluated using an agent-based simulation platform. The simulation results show that the proposed algorithm outperforms reference route management algorithms in many areas, especially in processing speed and memory efficiency. Detailed comparisons are provided for the sUAS flight time, the overall throughput, the conflict rate and communication resource utilization. The results demonstrate that our proposed algorithm can be used as a solution to improve the efficiency of airspace and communication resource utilization for next generation smart city and smart transportation.
Haowen Fang, Franco Basti, Mustafa Cenk Gursoy, Carlos E. Caicedo Bastidas, Qinru Qiu
RTCSA8
2019 Normalization and dropout for stochastic computing-based deep convolutional neural networks
Ji Li 0006, Zhe Li 0001, Ao Ren, Caiwen Ding, Jeffrey T. Draper, Shahin Nazarian, Qinru Qiu, Bo Yuan 0001, Yanzhi Wang 0001
Integr.8
2019 HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks
abstract
Deep convolutional neural networks (DCNNs) are one of the most promising deep learning techniques and have been recognized as the dominant approach for almost all recognition and detection tasks. The computation of DCNNs is memory intensive due to large feature maps and neuron connections, and the performance highly depends on the capability of hardware resources. With the recent trend of wearable devices and Internet of Things, it becomes desirable to integrate the DCNNs onto embedded and portable devices that require low power and energy consumptions and small hardware footprints. Recently stochastic computing (SC)-DCNN demonstrated that SC as a low-cost substitute to binary-based computing radically simplifies the hardware implementation of arithmetic units and has the potential to satisfy the stringent power requirements in embedded devices. In SC, many arithmetic operations that are resource-consuming in binary designs can be implemented with very simple hardware logic, alleviating the extensive computational complexity. It offers a colossal design space for integration and optimization due to its reduced area and soft error resiliency. In this paper, we present HEIF, a highly efficient SC-based inference framework of the large-scale DCNNs, with broad applications including (but not limited to) LeNet-5 and AlexNet, that achieves high energy efficiency and low area/hardware cost. Compared to SC-DCNN, HEIF features: 1) the first (to the best of our knowledge) SC-based rectified linear unit activation function to catch up with the recent advances in software models and mitigate degradation in application-level accuracy; 2) the redesigned approximate parallel counter and optimized stochastic multiplication using transmission gates and inverse mirror adders; and 3) the new optimization of weight storage using clustering. Most importantly, to achieve maximum energy efficiency while maintaining acceptable accuracy, HEIF considers holistic optimizations on cascade connection of function blocks in DCNN, pipelining technique, and bit-stream length reduction. Experimental results show that in large-scale applications HEIF outperforms previous SC-DCNN by the throughput of 4.1×, by area efficiency of up to 6.5×, and achieves up to 5.6× energy improvement.
Zhe Li 0001, Ji Li 0006, Ao Ren, Ruizhe Cai, Caiwen Ding, Xuehai Qian, Jeffrey T. Draper, Bo Yuan 0001, Jian Tang 0008, Qinru Qiu, Yanzhi Wang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.10
2018 Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework
abstract
Hardware accelerations of deep learning systems have been extensively investigated in industry and academia. The aim of this paper is to achieve ultra-high energy efficiency and performance for hardware implementations of deep neural networks (DNNs). An algorithm-hardware co-optimization framework is developed, which is applicable to different DNN types, sizes, and application scenarios. The algorithm part adopts the general block-circulant matrices to achieve a fine-grained tradeoff of accuracy and compression ratio. It applies to both fully-connected and convolutional layers and contains a mathematically rigorous proof of the effectiveness of the method. The proposed algorithm reduces computational complexity per layer from O(n2) to O(n log n) and storage complexity from O(n2) to O(n), both for training and inference. The hardware part consists of highly efficient Field Programmable Gate Array (FPGA)-based implementations using effective reconfiguration, batch processing, deep pipelining, resource re-using, and hierarchical control. Experimental results demonstrate that the proposed framework achieves at least 152X speedup and 71X energy efficiency gain compared with IBM TrueNorth processor under the same test accuracy. It achieves at least 31X energy efficiency gain compared with the reference FPGA-based work.
Yanzhi Wang 0001, Caiwen Ding, Zhe Li 0001, Geng Yuan, Siyu Liao, Bo Yuan 0001, Xuehai Qian, Jian Tang 0008, Qinru Qiu, Xue Lin 0001
AAAI10
2018 Optimizing Data Transfers for Improved Performance on Shared GPUs Using Reinforcement Learning
abstract
Optimizing resource utilization is a critical issue in cloud and cluster-based computing systems. In such systems, computing resources often consist of one or more GPU devices, and much research has already been conducted on means for maximizing compute resources through shared execution strategies. However, one of the most severe resource constraints in these scenarios is the data transfer channel between the host (i.e., CPU) and the device (i.e., GPU). Data transfer contention has been shown to have a significant impact on performance, yet methods for optimizing such contention have not been thoroughly studied. Techniques that have been examined make certain assumptions which limit effectiveness in the general case. In this paper, we introduce a heuristic which selectively aggregates transfers in order to maximize system performance by optimizing the transfer channel bandwidth. We compare this heuristic to traditional first-come-first-served approach, and apply Monte Carlo reinforcement learning to find an optimal policy for message aggregation. Finally, we evaluate the performance of Monte Carlo reinforcement learning with an arbitrarily-initialized policy. We demonstrate its effectiveness in learning optimal data transfer policy without detailed system characterization, which will enable a general adaptable solution for resource management of future systems.
Ryan S. Luley, Qinru Qiu
CCGrid2
2018 C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs
abstract
Recently, significant accuracy improvement has been achieved for acoustic recognition systems by increasing the model size of Long Short-Term Memory (LSTM) networks. Unfortunately, the ever-increasing size of LSTM model leads to inefficient designs on FPGAs due to the limited on-chip resources. The previous work proposes to use a pruning based compression technique to reduce the model size and thus speedups the inference on FPGAs. However, the random nature of the pruning technique transforms the dense matrices of the model to highly unstructured sparse ones, which leads to unbalanced computation and irregular memory accesses and thus hurts the overall performance and energy efficiency.
Shuo Wang 0009, Zhe Li 0001, Caiwen Ding, Bo Yuan 0001, Qinru Qiu, Yanzhi Wang 0001, Yun Liang 0001
FPGA5
2018 Learning Topics Using Semantic Locality
abstract
The topic modeling discovers the latent topic probability of the given text documents. To generate the more meaningful topic that better represents the given document, we proposed a new feature extraction technique which can be used in the data preprocessing stage. The method consists of three steps. First, it generates the word/word-pair from every single document. Second, it applies a two-way TF-IDF algorithm to word/word-pair for semantic filtering. Third, it uses the K-means algorithm to merge the word pairs that have the similar semantic meaning. Experiments are carried out on the Open Movie Database (OMDb), Reuters Dataset and 20NewsGroup Dataset. The mean Average Precision score is used as the evaluation metric. Comparing our results with other state-of-the-art topic models, such as Latent Dirichlet allocation and traditional Restricted Boltzmann Machines. Our proposed data preprocessing can improve the generated topic accuracy by up to 12.99 %.
Krittaphat Pugdeethosapol, Sheng Lin 0001, Zhe Li 0001, Caiwen Ding, Yanzhi Wang 0001, Qinru Qiu
ICPR7
2018 Scalable NoC-based Neuromorphic Hardware Learning and Inference
abstract
Bio-inspired neuromorphic hardware is a research direction to approach brain's computational power and energy efficiency. Spiking neural networks (SNN) encode information as sparsely distributed spike trains and employ spike-timingdependent plasticity (STDP) mechanism for learning. Existing hardware implementations of SNN are limited in scale or do not have in-hardware learning capability. In this work, we propose a low-cost scalable Network-on-Chip (NoC) based SNN hardware architecture with fully distributed in-hardware STDP learning capability. All hardware neurons work in parallel and communicate through the NoC. This enables chip-level interconnection, scalability and reconfigurability necessary for deploying different applications. The hardware is applied to learn MNIST digits as an evaluation of its learning capability. We explore the design space to study the trade-offs between speed, area and energy. How to use this procedure to find optimal architecture configuration is also discussed.
Haowen Fang, Amar Shrestha, De Ma, Qinru Qiu
IJCNN4
2018 AnRAD: A Neuromorphic Anomaly Detection Framework for Massive Concurrent Data Streams
abstract
The evolution of high performance computing technologies has enabled the large-scale implementation of neuromorphic models and pushed the research in computational intelligence into a new era. Among the machine learning applications, unsupervised detection of anomalous streams is especially challenging due to the requirements of detection accuracy and real-time performance. Designing a computing framework that harnesses the growing computing power of the multicore systems while maintaining high sensitivity and specificity to the anomalies is an urgent research topic. In this paper, we propose anomaly recognition and detection (AnRAD), a bioinspired detection framework that performs probabilistic inferences. We analyze the feature dependency and develop a self-structuring method that learns an efficient confabulation network using unlabeled data. This network is capable of fast incremental learning, which continuously refines the knowledge base using streaming data. Compared with several existing anomaly detection approaches, our method provides competitive detection quality. Furthermore, we exploit the massive parallel structure of the AnRAD framework. Our implementations of the detection algorithm on the graphic processing unit and the Xeon Phi coprocessor both obtain substantial speedups over the sequential implementation on general-purpose microprocessor. The framework provides real-time service to concurrent data streams within diversified knowledge contexts, and can be applied to large problems with multiple local patterns. Experimental results demonstrate high computing performance and memory efficiency. For vehicle behavior detection, the framework is able to monitor up to 16000 vehicles (data streams) and their interactions in real time with a single commodity coprocessor, and uses less than 0.2 ms for one testing subject. Finally, the detection network is ported to our spiking neural network simulator to show the potential of adapting to the emerging neuromorphic architectures.
Qiuwen Chen, Ryan S. Luley, Qing Wu 0002, Morgan Bishop, Richard W. Linderman, Qinru Qiu
IEEE Trans. Neural Networks Learn. Syst.6
2017 Towards acceleration of deep convolutional neural networks using stochastic computing
abstract
In recent years, Deep Convolutional Neural Network (DCNN) has become the dominant approach for almost all recognition and detection tasks and outperformed humans on certain tasks. Nevertheless, the high power consumptions and complex topologies have hindered the widespread deployment of DCNNs, particularly in wearable devices and embedded systems with limited area and power budget. This paper presents a fully parallel and scalable hardware-based DCNN design using Stochastic Computing (SC), which leverages the energy-accuracy trade-off through optimizing SC components in different layers. We first conduct a detailed investigation of the Approximate Parallel Counter (APC) based neuron and multiplexer-based neuron using SC, and analyze the impacts of various design parameters, such as bit stream length and input number, on the energy/power/area/accuracy of the neuron cell. Then, from an architecture perspective, the influence of inaccuracy of neurons in different layers on the overall DCNN accuracy (i.e., software accuracy of the entire DCNN) is studied. Accordingly, a structure optimization method is proposed for a general DCNN architecture, in which neurons in different layers are implemented with optimized SC components, so as to reduce the area, power, and energy of the DCNN while maintaining the overall network performance in terms of accuracy. Experimental results show that the proposed approach can find a satisfactory DCNN configuration, which achieves 55X, 151X, and 2X improvement in terms of area, power and energy, respectively, while the error is increased by 2.86%, compared with the conventional binary ASIC implementation.
Ji Li 0006, Ao Ren, Zhe Li 0001, Caiwen Ding, Bo Yuan 0001, Qinru Qiu, Yanzhi Wang 0001
ASP-DAC6
2017 SC-DCNN: Highly-Scalable Deep Convolutional Neural Network using Stochastic Computing
abstract
With the recent advance of wearable devices and Internet of Things (IoTs), it becomes attractive to implement the Deep Convolutional Neural Networks (DCNNs) in embedded and portable systems. Currently, executing the software-based DCNNs requires high-performance servers, restricting the widespread deployment on embedded and mobile IoT devices. To overcome this obstacle, considerable research efforts have been made to develop highly-parallel and specialized DCNN accelerators using GPGPUs, FPGAs or ASICs.
Ao Ren, Zhe Li 0001, Caiwen Ding, Qinru Qiu, Yanzhi Wang 0001, Ji Li 0006, Xuehai Qian, Bo Yuan 0001
ASPLOS4
2017 Real-time anomaly detection for streaming data using burst code on a neurosynaptic processor
abstract
Real-time anomaly detection for streaming data is a desirable feature for mobile devices or unmanned systems. The key challenge is how to deliver required performance under the stringent power constraint. To address the paradox between performance and power consumption, brain-inspired hardware, such as the IBM Neurosynaptic System, has been developed to enable low power implementation of large-scale neural models. Meanwhile, inspired by the operation and the massive parallel structure of human brain, carefully structured inference model has been demonstrated to give superior detection quality than many traditional models while facilitates neuromorphic implementation. Implementing inference based anomaly detection on the neurosynaptic processor is not straightforward due to hardware limitations. This work presents a design flow and component library that flexibly maps learned detection network to the TrueNorth architecture. Instead of traditional rate code, burst code is adopted in the design, which represents numerical value using the phase of a burst of spike trains. This does not only reduce the hardware complexity, but also increases the results accuracy. A Corelet library, NeoInfer-TN, is developed for basic operations in burst code and two-phase pipelines are constructed based on the library components. The design can be configured for different tradeoffs between detection accuracy and throughput/energy. We evaluate the system using intrusion detection data streams. The results show higher detection rate than some conventional approaches and real-time performance, with only 50mW power consumption. Overall, it achieves 108operations per watt-second.
Qiuwen Chen, Qinru Qiu
DATE2
2017 Structural design optimization for deep convolutional neural networks using stochastic computing
abstract
Deep Convolutional Neural Networks (DCNNs) have been demonstrated as effective models for understanding image content. The computation behind DCNNs highly relies on the capability of hardware resources due to the deep structure. DCNNs have been implemented on different large-scale computing platforms. However, there is a trend that DCNNs have been embedded into light-weight local systems, which requires low power/energy consumptions and small hardware footprints. Stochastic Computing (SC) radically simplifies the hardware implementation of arithmetic units and has the potential to satisfy the small low-power needs of DCNNs. Local connectivities and down-sampling operations have made DCNNs more complex to be implemented using SC. In this paper, eight feature extraction designs for DCNNs using SC in two groups are explored and optimized in detail from the perspective of calculation precision, where we permute two SC implementations for inner-product calculation, two down-sampling schemes, and two structures of DCNN neurons. We evaluate the network in aspects of network accuracy and hardware performance for each DCNN using one feature extraction design out of eight. Through exploration and optimization, the accuracies of SC-based DCNNs are guaranteed compared with software implementations on CPU/GPU/binary-based ASIC synthesis, while area, power, and energy are significantly reduced by up to 776x, 190x, and 32835x.
Zhe Li 0001, Ao Ren, Ji Li 0006, Qinru Qiu, Bo Yuan 0001, Jeffrey T. Draper, Yanzhi Wang 0001
DATE4
2017 Softmax Regression Design for Stochastic Computing Based Deep Convolutional Neural Networks
abstract
Recently, Deep Convolutional Neural Networks (DCNNs) have made tremendous advances, achieving close to or even better accuracy than human-level perception in various tasks. Stochastic Computing (SC), as an alternate to the conventional binary computing paradigm, has the potential to enable massively parallel and highly scalable hardware implementations of DCNNs. In this paper, we design and optimize the SC based Softmax Regression function. Experiment results show that compared with a binary SR, the proposed SC-SR under longer bit stream can reach the same level of accuracy with the improvement of 295X, 62X, 2617X in terms of power, area and energy, respectively. Binary SR is suggested for future DCNNs with short bit stream length input whereas SC-SR is recommended for longer bit stream.
Ji Li 0006, Zhe Li 0001, Caiwen Ding, Ao Ren, Bo Yuan 0001, Qinru Qiu, Jeffrey T. Draper, Yanzhi Wang 0001
ACM Great Lakes Symposium on VLSI7
2017 Energy-efficient, high-performance, highly-compressed deep neural network design using block-circulant matrices
abstract
Deep neural networks (DNNs) have emerged as the most powerful machine learning technique in numerous artificial intelligent applications. However, the large sizes of DNNs make themselves both computation and memory intensive, thereby limiting the hardware performance of dedicated DNN accelerators. In this paper, we propose a holistic framework for energy-efficient high-performance highly-compressed DNN hardware design. First, we propose block-circulant matrix-based DNN training and inference schemes, which theoretically guarantee Big-O complexity reduction in both computational cost (from O(n2) to O(n log n)) and storage requirement (from O(n2) to O(n)) of DNNs. Second, we dedicatedly optimize the hardware architecture, especially on the key fast Fourier transform (FFT) module, to improve the overall performance in terms of energy efficiency, computation performance and resource cost. Third, we propose a design flow to perform hardware-software co-optimization with the purpose of achieving good balance between test accuracy and hardware performance of DNNs. Based on the proposed design flow, two block-circulant matrix-based DNNs on two different datasets are implemented and evaluated on FPGA. The fixed-point quantization and the proposed block-circulant matrix-based inference scheme enables the network to achieve as high as 3.5 TOPS computation performance and 3.69 TOPS/W energy efficiency while the memory is saved by 108X ~ 116X with negligible accuracy degradation.
Siyu Liao, Zhe Li 0001, Xue Lin 0001, Qinru Qiu, Yanzhi Wang 0001, Bo Yuan 0001
ICCAD4
2017 A spike-based long short-term memory on a neurosynaptic processor
abstract
Low-power brain-inspired hardware systems have gained significant traction in recent years. They offer high energy efficiency and massive parallelism due to the distributed and asynchronous nature of neural computation through low-energy spikes. One such platform is the IBM TrueNorth Neurosynaptic System. Recently TrueNorth compatible representation learning algorithms have emerged, achieving close to state-of-the-art performance in various datasets. An exception is its application in temporal sequence processing models such as recurrent neural networks (RNNs), which is still at the proof of concept level. This is partly due to the hardware constraints in connectivity and syn-aptic weight resolution, and the inherent difficulty in capturing temporal dynamics of an RNN using spiking neurons. This work presents a design flow that overcomes the aforementioned difficulties and maps a special case of recurrent networks called Long Short-Term Memory (LSTM) onto a spike-based platform. The framework is built on top of various approximation techniques, weight and activation discretization, spiking neuron sub-circuits that implements the complex gating mechanisms and a store-and-release technique to enable neuron synchronization and faithful storage. While many of the techniques can be applied to map LSTM to any SNN simulator/emulator, here we demonstrate this approach on the TrueNorth chip adhering to its constraints. Two benchmark LSTM applications, parity check and Extended Reber Grammar, are evaluated and their accuracy, energy and speed tradeoffs are analyzed.
Amar Shrestha, Khadeer Ahmed, Yanzhi Wang 0001, David P. Widemann, Adam Moody, Brian Van Essen, Qinru Qiu
ICCAD7
2017 A Hierarchical Framework of Cloud Resource Allocation and Power Management Using Deep Reinforcement Learning
abstract
Automatic decision-making approaches, such as reinforcement learning (RL), have been applied to (partially) solve the resource allocation problem adaptively in the cloud computing system. However, a complete cloud resource allocation framework exhibits high dimensions in state and action spaces, which prohibit the usefulness of traditional RL techniques. In addition, high power consumption has become one of the critical concerns in design and control of cloud computing systems, which degrades system reliability and increases cooling cost. An effective dynamic power management (DPM) policy should minimize power consumption while maintaining performance degradation within an acceptable level. Thus, a joint virtual machine (VM) resource allocation and power management framework is critical to the overall cloud computing system. Moreover, novel solution framework is necessary to address the even higher dimensions in state and action spaces. In this paper, we propose a novel hierarchical framework for solving the overall resource allocation and power management problem in cloud computing systems. The proposed hierarchical framework comprises a global tier for VM resource allocation to the servers and a local tier for distributed power management of local servers. The emerging deep reinforcement learning (DRL) technique, which can deal with complicated control problems with large state space, is adopted to solve the global tier problem. Furthermore, an autoencoder and a novel weight sharing structure are adopted to handle the high-dimensional state space and accelerate the convergence speed. On the other hand, the local tier of distributed server power managements comprises an LSTM based workload predictor and a model-free RL based power manager, operating in a distributed manner. Experiment results using actual Google cluster traces show that our proposed hierarchical framework significantly saves the power consumption and energy usage than the baseline while achieving no severe latency degradation. Meanwhile, the proposed framework can achieve the best trade-off between latency and power/energy consumption in a server cluster.
Ning Liu 0007, Zhe Li 0001, Jielong Xu, Sheng Lin 0001, Qinru Qiu, Jian Tang 0008, Yanzhi Wang 0001
ICDCS6
2017 Hardware-driven nonlinear activation for stochastic computing based deep convolutional neural networks
abstract
Recently, Deep Convolutional Neural Networks (DCNNs) have made unprecedented progress, achieving the accuracy close to, or even better than human-level perception in various tasks. There is a timely need to map the latest software DCNNs to application-specific hardware, in order to achieve orders of magnitude improvement in performance, energy efficiency and compactness. Stochastic Computing (SC), as a low-cost alternative to the conventional binary computing paradigm, has the potential to enable massively parallel and highly scalable hardware implementation of DCNNs. One major challenge in SC based DCNNs is designing accurate nonlinear activation functions, which have a significant impact on the network-level accuracy but cannot be implemented accurately by existing SC computing blocks. In this paper, we design and optimize SC based neurons, and we propose highly accurate activation designs for the three most frequently used activation functions in software DCNNs, i.e, hyperbolic tangent, logistic, and rectified linear units. Experimental results on LeNet-5 using MNIST dataset demonstrate that compared with a binary ASIC hardware DCNN, the DCNN with the proposed SC neurons can achieve up to 61X, 151X, and 2X improvement in terms of area, power, and energy, respectively, at the cost of small precision degradation. In addition, the SC approach achieves up to 21X and 41X of the area, 41X and 72X of the power, and 198200X and 96443X of the energy, compared with CPU and GPU approaches, respectively, while the error is increased by less than 3.07%. ReLU activation is suggested for future SC based DCNNs considering its superior performance under a small bit stream length.
Ji Li 0006, Zhe Li 0001, Caiwen Ding, Ao Ren, Qinru Qiu, Jeffrey T. Draper, Yanzhi Wang 0001
IJCNN6
2017 Stable spike-timing dependent plasticity rule for multilayer unsupervised and supervised learning
abstract
Spike-Timing Dependent Plasticity (STDP), the canonical learning rule for spiking neural networks (SNN), is gaining tremendous interest because of its simplicity, efficiency and biological plausibility. However, to date, multilayer feed-forward networks of spiking neurons are either only partially trained using STDP or pre-trained using traditional deep neural networks which are converted to deep spiking neural networks or a two-layer network where STDP learnt features are manually labelled. In this work, we present a low-cost, simplified, yet stable STDP rule for layer-wise unsupervised and supervised training of a multilayer feed-forward SNN. We propose to approximate Bayesian neuron using Stochastic Integrate and Fire (SIF) neuron model and introduce a supervised learning approach using teacher neurons to train the classification layer with one neuron per class. A SNN is trained for classification of handwritten digits with multiple layers of spiking neurons, including both the feature extraction and classification layer, using the proposed STDP rule. Our method achieves comparable to better accuracy on MNIST dataset than manually labelled two layer networks for the same sized hidden layer. We also analyze the parameter space to provide rationales for parameter fine-tuning and provide additional methods to improve noise resilience and input intensity variations. We further propose a Quantized 2-Power Shift (Q2PS) STDP rule, which reduces the implementation cost of digital hardware while achieves comparable performance.
Amar Shrestha, Khadeer Ahmed, Yanzhi Wang 0001, Qinru Qiu
IJCNN4
2017 CirCNN: accelerating and compressing deep neural networks using block-circulant weight matrices
abstract
Large-scale deep neural networks (DNNs) are both compute and memory intensive. As the size of DNNs continues to grow, it is critical to improve the energy efficiency and performance while maintaining accuracy. For DNNs, the model size is an important factor affecting performance, scalability and energy efficiency. Weight pruning achieves good compression ratios but suffers from three drawbacks: 1) the irregular network structure after pruning, which affects performance and throughput; 2) the increased training complexity; and 3) the lack of rigirous guarantee of compression ratio and inference accuracy.
Caiwen Ding, Siyu Liao, Yanzhi Wang 0001, Zhe Li 0001, Ning Liu 0007, Youwei Zhuo, Chao Wang 0051, Xuehai Qian, Yu Bai 0004, Geng Yuan, Jian Tang 0008, Qinru Qiu, Xue Lin 0001, Bo Yuan 0001
MICRO14
2016 DSCNN: Hardware-oriented optimization for Stochastic Computing based Deep Convolutional Neural Networks
abstract
Deep Convolutional Neural Networks (DCNN), a branch of Deep Neural Networks which use the deep graph with multiple processing layers, enables the convolutional model to finely abstract the high-level features behind an image. Large-scale applications using DCNN mainly operate in high-performance server clusters, GPUs or FPGA clusters; it is restricted to extend the applications onto mobile/wearable devices and Internet-of-Things (IoT) entities due to high power/energy consumption. Stochastic Computing is a promising method to overcome this shortcoming used in specific hardware-based systems. Many complex arithmetic operations can be implemented with very simple hardware logic in the SC framework, which alleviates the extensive computation complexity. The exploration of network-wise optimization and the revision of network structure with respect to stochastic computing based hardware design have not been discussed in previous work. In this paper, we investigate Deep Stochastic Convolutional Neural Network (DSCNN) for DCNN using stochastic computing. The essential calculation components using SC are designed and evaluated. We propose a joint optimization method to collaborate components guaranteeing a high calculation accuracy in each stage of the network. The structure of original DSCNN is revised to accommodate SC hardware design's simplicity. Experimental Results show that as opposed to software inspired feature extraction block in DSCNN, an optimized hardware oriented feature extraction block achieves as higher as 59.27% calculation precision. And the optimized DSCNN can achieve only 3.48% network test error rate compared to 27.83% for baseline DSCNN using software inspired feature extraction block.
Zhe Li 0001, Ao Ren, Ji Li 0006, Qinru Qiu, Yanzhi Wang 0001, Bo Yuan 0001
ICCD4
2016 Simulation of bayesian learning and inference on distributed stochastic spiking neural networks
abstract
The ability of neural networks to perform pattern recognition, classification and associative memory, is essential to applications such as image and speech recognition, natural language understanding, decision making etc. In spiking neural networks (SNNs), information is encoded as sparsely distributed train of spikes, which allows learning through the spike-timing dependent plasticity (STDP) property. SNNs can potentially achieve very large scale implementation and distributed learning due to the inherent asynchronous and sparse inter-neuron communications. In this work, we develop an efficient, scalable and flexible SNN simulator, which supports learning through STDP. The simulator is ideal for biologically inspired neuron models for computation but not for biologically realistic models. Bayesian neuron model for SNNs that is capable of online and fully-distributed STDP learning is introduced. The function of the simulator is validated using two networks representing two different applications from unsupervised feature extraction to inference based sentence construction.
Khadeer Ahmed, Amar Shrestha, Qinru Qiu
IJCNN3
2016 Probabilistic inference using stochastic spiking neural networks on a neurosynaptic processor
abstract
Spiking neural networks are rapidly gaining popularity for their ability to perform efficient computation akin to the way a brain processes information. It has the potential to achieve low cost and high energy efficiency due to the distributed nature of neural computation and the use of low energy spikes for information exchange. A stochastic spiking neural network naturally can be used to realize Bayesian inference. IBM's TrueNorth is a neurosynaptic processor that has more than 1 million digital spiking neurons and 268 million digital synapses with less than 200 mW peak power. In this paper we propose the first work that converts an inference network to a spiking neural network that runs on the TrueNorth processor. Using inference-based sentence construction as a case study, we discuss algorithms that transform an inference network to a spiking neural network, and a spiking neural network to TrueNorth corelet designs. In our experiments, the TrueNorth spiking neural network constructed sentences have a matching accuracy of 88% while consuming an average power of 0.205 mW.
Khadeer Ahmed, Amar Shrestha, Qinru Qiu, Qing Wu 0002
IJCNN3
2016 Enhancing bidirectional association between deep image representations and loosely correlated texts
abstract
The problem of bridging the gap between image and natural language has gained more and more attention in recent years. This paper continues to push the study and improves the bidirectional retrieval performance across the modalities. Unlike previous works that target at single sentence densely describing the image objects, we extend the focus to associating deep image representations with noisy texts that are only loosely correlated. Based on text-image fragment embedding, our model employs a sequential configuration, connects two embedding stages together. The first stage learns the relevancy of the text fragments, and the second stage uses the filtered output from the first one to improve the matching results. The model also integrates multiple convolutional neural networks (CNN) to construct the image fragments, in which rich context information such as human faces can be extracted to increase the alignment accuracy. The proposed method is evaluated with both synthetic dataset and real-world dataset collected from picture news website. The results show up to 50% ranking performance improvement over the comparison models.
Qiuwen Chen, Qinru Qiu
IJCNN2
2016 Towards memristor based accelerator for sparse matrix vector multiplication
abstract
In the last few years, memristor crossbar array is drawing increasing attention from the research community as a promising neuromorphic computing accelerator. In this work, we investigate the hardware acceleration of a sparse matrix vector (SpMV) multiplication engine based on memristor crossbar array. We demonstrate that naive matrix coefficient mapping is infeasible and unpractical if the matrix has large dimensions. To combat this problem, we extend the traditional Cuthill-McKee algorithm used for matrix restructuring, and propose a generalized sparse matrix reordering (GSMR) technique, which leverages linear transformation to effectively break down any rectangular unsymmetrical matrices into minimum number of sub-blocks that fit into the reasonably sized crossbar array. Simulated results show that our proposed design achieves appealing performances in terms of speed and energy efficiency compared to both CPU and GPU platforms. In addition, a memristor crossbar array utilizing GSMR outperforms its counterpart with no-GSMR by 90% performance improvements and 44% energy reduction.
Qinru Qiu
ISCAS2
2015 Cloning your mind: security challenges in cognitive system designs and their solutions
abstract
With the booming of big-data applications, cognitive information processing systems that leverage advanced data processing technologies, e.g., machine learning and data mining, are widely used in many industry fields. Although these technologies demonstrate great processing capability and accuracy in the relevant applications, several security and safety challenges are also emerging against these learning based technologies. In this paper, we will first introduce several security concerns in cognitive system designs. Some real examples are then used to demonstrate how the attackers can potentially access the confidential user data, replicate a sensitive data processing model without being granted the access to the details of the model, and obtain some key features of the training data by using the services publically accessible to a normal user. Based on the analysis of these security challenges, we also discuss several possible solutions that can protect the information privacy and security of cognitive systems during different stages of the usage.
Beiye Liu, Chunpeng Wu, Hai Li 0001, Yiran Chen 0001, Qing Wu 0002, Mark Barnell, Qinru Qiu
DAC7
2015 FPGA Acceleration of Recurrent Neural Network Based Language Model
abstract
Recurrent neural network (RNN) based language model (RNNLM) is a biologically inspired model for natural language processing. It records the historical information through additional recurrent connections and therefore is very effective in capturing semantics of sentences. However, the use of RNNLM has been greatly hindered for the high computation cost in training. This work presents an FPGA implementation framework for RNNLM training acceleration. At architectural level, we improve the parallelism of RNN training scheme and reduce the computing resource requirement for computation efficiency enhancement. The hardware implementation primarily targets at reducing data communication load. A multi-thread based computation engine is utilized which can successfully mask the long memory latency and reuse frequent accessed data. The evaluation based on the Microsoft Research Sentence Completion Challenge shows that the proposed FPGA implementation outperforms traditional class-based modest-size recurrent networks and obtains 46.2% in training accuracy. Moreover, experiments at different network sizes demonstrate a great scalability of the proposed framework.
Sicheng Li 0001, Chunpeng Wu, Hai Li 0001, Boxun Li, Yu Wang 0002, Qinru Qiu
FCCM6
2015 Self-structured confabulation network for fast anomaly detection and reasoning
abstract
Inference models such as the confabulation network are particularly useful in anomaly detection applications because they allow introspection to the decision process. However, building such network model always requires expert knowledge. In this paper, we present a self-structuring technique that learns the structure of a confabulation network from unlabeled data. Without any assumption of the distribution of data, we leverage the mutual information between features to learn a succinct network configuration, and enable fast incremental learning to refine the knowledge bases from continuous data streams. Compared to several existing anomaly detection methods, the proposed approach provides higher detection performance and excellent reasoning capability. We also exploit the massive parallelism that is inherent to the inference model and accelerate the detection process using GPUs. Experimental results show significant speedups and the potential to be applied to real-time applications with high-volume data streams.
Qiuwen Chen, Qing Wu 0002, Morgan Bishop, Richard W. Linderman, Qinru Qiu
IJCNN5
2015 The applications of memristor devices in next-generation cortical processor designs
abstract
Discovery of memristor opened a new era of the research on universal memory thanks to many attractive properties demonstrated by this emerging device. In this paper, we switch our research focus to neuromorphic computing, which, same as memory technology, significantly benefits from the technical advances of memristor. Particularly, we present the implementation of cortical processor augmented with neuromorphic computing accelerators (NCAs) for cognitive applications, including: 1) the design details and basic operations of the NCA based on memristor crossbars; and 2) the integration between conventional pipeline and NCAs. At the end, we also discuss the scalability of our proposed NCA designs.
Hai Li 0001, Beiye Liu, Xiaoxiao Liu 0001, Mengjie Mao, Yiran Chen 0001, Qing Wu 0002, Qinru Qiu
ISCAS7
2014 The stochastic modeling of TiO2 memristor and its usage in neuromorphic system design
abstract
Memristor, the fourth basic circuit element, has shown great potential in neuromorphic circuit design for its unique synapse-like feature. However, though the continuous resistance state of memristor has been expected, obtaining and maintaining an arbitrary intermediate state cannot be well controlled in nowadays memristive system. In addition, the stochastic switching behaviors have been widely observed. To facilitate the investigation on memristor-based hardware implementation, we built a stochastic behavior model of TiO2memristive devices based on the real experimental results. By leveraging the stochastic behavior of memristors, a macro cell design composed of multiple parallel connecting memristors can be successfully used in implementing the weight storage unit and the stochastic neuron — the two fundamental components in neural network (NN)s, providing a feasible solution in memristor-based hardware implementation.
Miao Hu 0002, Yu Wang 0002, Qinru Qiu, Yiran Chen 0001, Hai Li 0001
ASP-DAC3
2014 Battery aware stochastic QoS boosting in mobile computing devices
abstract
Mobile computing has been weaved into everyday lives to a great extend. Their usage is clearly imprinted with user's personal signature. The ability to learn such signature enables immense potential in workload prediction and resource management. In this work, we investigate the user behavior modeling and apply the model for energy management. Our goal is to maximize the quality of service (QoS) provided by the mobile device (i.e., smartphone), while keep the risk of battery depletion below a given threshold. A Markov Decision Process (MDP) is constructed from history user behavior. The optimal management policy is solved using linear programing. Simulations based on real user traces validate that, compared to existing battery energy management techniques, the stochastic control performs better in boosting the mobile devices' QoS without significantly increasing the chance of battery depletion.
Hao Shen 0007, Qiuwen Chen, Qinru Qiu
DATE3
2014 Contention aware frequency scaling on CMPs with guaranteed quality of service
abstract
Workload consolidation is usually performed in datacenters to improve server utilization for higher energy efficiency. One of the key issues related to workload consolidation is contention for shared resources such as last level cache, main memory, memory controller, etc. Dynamic voltage and frequency scaling (DVFS) of CPU is another effective technique that has widely been used to trade the performance for power reduction. We have found that the degree of resource contention of a system affects its performance sensitivity to CPU frequency. In this paper, we apply machine learning techniques to construct a model that quantifies runtime performance degradation caused by resource contention and frequency scaling. The inputs of our model are readings from Performance Monitoring Units (PMU) screened using standard feature selection technique. The model is tested on an SMT-enabled chip multi-processor and it reaches up to 90% accuracy. Experimental results show that, guided by the performance model, runtime power management techniques such as DVFS can achieve more accurate power and performance tradeoff without violating the quality of service (QoS) agreement. The QoS violation of the proposed system is significantly lower than systems that have no performance degradation information.
Hao Shen 0007, Qinru Qiu
DATE2
2014 Accelerating pattern matching in neuromorphic text recognition system using Intel Xeon Phi coprocessor
abstract
Neuromorphic computing systems refer to the computing architecture inspired by the working mechanism of human brains. The rapidly reducing cost and increasing performance of state-of-the-art computing hardware allows large-scale implementation of machine intelligence models with neuromorphic architectures and opens the opportunity for new applications. One such computing hardware is Intel Xeon Phi coprocessor, which delivers over a TeraFLOP of computing power with 61 integrated processing cores. How to efficiently harness such computing power to achieve real time decision and cognition is one of the key design considerations. This paper presents an optimized implementation of Brain-State-in-a-Box (BSB) neural network model on the Xeon Phi coprocessor for pattern matching in the context of intelligent text recognition of noisy document images. From a scalability standpoint on a High Performance Computing (HPC) platform we show that efficient workload partitioning and resource management can double the performance of this many-core architecture for neuromorphic applications.
Khadeer Ahmed, Qinru Qiu, Parth Malani, Mangesh Tamhankar
IJCNN2
2014 Bio-inspired computing with resistive memories - models, architectures and applications
abstract
The traditional Von Neumann architecture has constrained the potential for applying massively parallel architecture to embedded high performance computing where we must optimize the size, weight and power of the system. Inspired by highly parallel biological systems, such as the human brain, the neuromorphic architecture offers a promising novel computing paradigm for compact and energy efficient platforms. The discovery of memristor devices provided the element we need with unprecedented efficiency in realizing such a computing architecture. There are still many challenges left to meet our goal of a fully functional bio-inspired computer. Here we will discuss our research in memristor crossbar based architecture, adaptation of this architecture for cogent confabulation models, and potential applications of the bio-inspired computer.
Qing Wu 0002, Beiye Liu, Yiran Chen 0001, Hai Li 0001, Qiuwen Chen, Qinru Qiu
ISCAS6
2013 Improving energy efficiency for energy harvesting embedded systems
abstract
While the energy harvesting system (EHS) supplies green energy to the embedded system, it also suffers from uncertainty and large variation in harvesting rate. This constraint can be remedied by using efficient energy storage. Hybrid Electrical Energy Storage (HEES) system is proposed recently as a cost effective approach with high power conversion efficiency and low self-discharge. In this paper, we propose a fast heuristic algorithm to improve the efficiency of charge allocation and replacement in an EHS/HEES equipped embedded system. The goal of our algorithm is to minimize the energy overhead on the DC-DC converter while satisfying the task deadline constraints of the embedded workload and maximizing the energy stored in the HEES system. We first provide an approximated but accurate power consumption model of the DC-DC converter. Based on this model, the optimal operating point of the system can be analytically solved. Integrated with the dynamic reconfiguration of the HEES bank, our algorithm provides energy efficiency improvement and runtime overhead reduction compared to previous approaches.
Yukan Zhang, Qinru Qiu
ASP-DAC3
2013 Improving charging efficiency with workload scheduling in energy harvesting embedded systems
abstract
In energy harvesting embedded systems, if the harvested power is sufficient for the workload, extra power will be stored in the electrical energy storage (EES) bank. How much energy can be stored is affected by many factors including the efficiency of the energy harvesting module, the input/output voltage of the DC-DC converters, the status of the EES elements, and the characteristics of the workload. This paper investigates the impact of workload scheduling of the embedded system on the storage efficiency of the EES bank. We first provide an approximated but accurate power consumption model of the DC-DC converter. Based on this model, we analytically prove that an optimal workload schedule is to always execute high power task first. Experimental results confirm that proposed scheduling strategy outperforms all other possible scheduling and increases the amount of stored energy by up to 10.41% in average.
Yukan Zhang, Qinru Qiu
DAC3
2013 User-aware energy efficient streaming strategy for smartphone based video playback applications
abstract
We propose a methodology to design user-aware streaming strategies for energy efficient smartphone video playback applications (e.g. YouTube). Our goal is to manage the streaming process to minimize the sleep and wake penalty of cellular module and at the same time avoid the energy waste from excessive downloading. The problem is modeled as a stochastic inventory system, where the real length of video playback requested by the smartphone user is considered as demand that follows a stochastic process. Through user behavior analysis, a Gaussian Mixture Model (GMM) is constructed to predict the user demand in video playback, and then an energy efficient video downloading strategy will be determined progressively during the playback process. Experimental results show that compared to a static downloading strategy that is optimized by exhaustive trail, our method can reduce the wasted energy by 10 percent in average.
Hao Shen 0007, Qinru Qiu
DATE2
2013 A neuromorphic architecture for anomaly detection in autonomous large-area traffic monitoring
abstract
The advanced sensing and imaging capability of today's sensor networks enables real time monitoring in a large area. In order to provide continuous monitoring and prompt situational awareness, an abstract-level autonomous information processing framework is developed that is able to detect various categories of abnormal traffic events with unsupervised learning. The framework is based on cogent confabulation model, which performs statistical inference in a manner inspired by human neocortex system. It enables detection and recognition of abnormal target vehicles within the context of surrounding traffic activities and previous events using likelihood-ratio test. A neuromorphic architecture is proposed which accelerates the computation for real-time detection by leveraging memristor crossbar arrays.
Qiuwen Chen, Qinru Qiu, Hai Li 0001, Qing Wu 0002
ICCAD2
2013 A Parallel Neuromorphic Text Recognition System and Its Implementation on a Heterogeneous High-Performance Computing Cluster
abstract
Given the recent progress in the evolution of high-performance computing (HPC) technologies, the research in computational intelligence has entered a new era. In this paper, we present an HPC-based context-aware intelligent text recognition system (ITRS) that serves as the physical layer of machine reading. A parallel computing architecture is adopted that incorporates the HPC technologies with advances in neuromorphic computing models. The algorithm learns from what has been read and, based on the obtained knowledge, it forms anticipations of the word and sentence level context. The information processing flow of the ITRS imitates the function of the neocortex system. It incorporates large number of simple pattern detection modules with advanced information association layer to achieve perception and recognition. Such architecture provides robust performance to images with large noise. The implemented ITRS software is able to process about 16 to 20 scanned pages per second on the 500 trillion floating point operations per second (TFLOPS) Air Force Research Laboratory (AFRL)/Information Directorate (RI) Condor HPC after performance optimization.
Qinru Qiu, Qing Wu 0002, Morgan Bishop, Robinson E. Pino, Richard W. Linderman
IEEE Trans. Computers1
2013 Achieving autonomous power management using reinforcement learning
abstract
System level power management must consider the uncertainty and variability that come from the environment, the application and the hardware. A robust power management technique must be able to learn the optimal decision from past events and improve itself as the environment changes. This article presents a novel on-line power management technique based on model-free constrained reinforcement learning (Q-learning). The proposed learning algorithm requires no prior information of the workload and dynamically adapts to the environment to achieve autonomous power management. We focus on the power management of the peripheral device and the microprocessor, two of the basic components of a computer. Due to their different operating behaviors and performance considerations, these two types of devices require different designs of Q-learning agent. The article discusses system modeling and cost function construction for both types of Q-learning agent. Enhancement techniques are also proposed to speed up the convergence and better maintain the required performance (or power) constraint in a dynamic system with large variations. Compared with the existing machine learning based power management techniques, the Q-learning based power management is more flexible in adapting to different workload and hardware and provides a wider range of power-performance tradeoff.
Hao Shen 0007, Qing Wu 0002, Qinru Qiu
ACM Trans. Design Autom. Electr. Syst.5
2012 Tag-assisted sentence confabulation for intelligent text recognition
abstract
Autonomous and intelligent recognition of printed or handwritten text image is one of the key features to achieve situational awareness. A neuromorphic model based intelligent text recognition (ITR) system has been developed in our previous work, which recognizes texts based on word level and sentence level context represented by statistical information of characters and words. While quite effective, sometimes the existing ITR system still generates results that are grammatically incorrect because it ignores semantic and syntactic properties of sentences. In this work, we improve the accuracy of the existing ITR system by incorporating parts-of-speech tagging into the text recognition procedure. Our experimental results show that the tag-assisted text recognition improves sentence level success rate by 33% in average.
Qinru Qiu, Morgan Bishop, Qing Wu 0002
CISDA2
2012 A game theoretic resource allocation for overall energy minimization in mobile cloud computing system
abstract
Cloud computing and virtualization techniques provide mobile devices with battery energy saving opportunities by allowing them to offload computation and execute code remotely. When the cloud infrastructure consists of heterogeneous servers, the mapping between mobile devices and servers plays an important role in determining the energy dissipation on both sides. From an environmental impact perspective, any energy dissipation related to computation should be counted. To achieve energy sustainability, it is important reducing the overall energy consumption of the mobile systems and the cloud infrastructure. Furthermore, reducing cloud energy consumption can potentially reduce the cost of mobile cloud users because the pricing model of cloud services is pay-by-usage. In this paper, we propose a game-theoretic approach to optimize the overall energy in a mobile cloud computing system. We formulate the energy minimization problem as a congestion game, where each mobile device is a player and his strategy is to select one of the servers to offload the computation while minimizing the overall energy consumption. We prove that the Nash equilibrium always exists in this game and propose an efficient algorithm that could achieve the Nash equilibrium in polynomial time. Experimental results show that our approach is able to reduce the total energy of mobile devices and servers compared to a random approach and an approach which only tries to reduce mobile devices alone.
Yukan Zhang, Qinru Qiu, Yung-Hsiang Lu
ISLPED3
2012 Introduction to the special section on adaptive power management for energy and temperature-aware computing systems
abstract
No abstract available.
Ayse K. Coskun, Yung-Hsiang Lu, Qinru Qiu
ACM Trans. Design Autom. Electr. Syst.3
2012 A Multi-Agent Framework for Thermal Aware Task Migration in Many-Core Systems
abstract
In deep submicrometer era, thermal hot spots, and large temperature gradients significantly impact system reliability, performance, cost, and leakage power. As the system complexity increases, it is more and more difficult to perform thermal management in a centralized manner because of state explosion and the overhead of monitoring the entire chip. In this paper, we propose a framework for distributed thermal management in many-core systems where balanced thermal profile can be achieved by proactive task migration among neighboring cores. The framework has a low cost agent residing in each core that observes the local workload and temperature and communicates with its nearest neighbor for task migration and exchange. By choosing only those migration requests that will result in balanced workload without generating thermal emergency, the proposed framework maintains workload balance across the system and avoids unnecessary migration. Experimental results show that, our distributed management policy achieves almost the same performance as a global management policy when the tasks are initially randomly distributed. Compared with existing proactive task migration technique, our approach generates less hotspot, less migration overhead with negligible performance overhead.
Qinru Qiu, Qing Wu 0002
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Harvesting-Aware Power Management for Real-Time Systems With Renewable Energy
abstract
In this paper, we propose a harvesting-aware power management algorithm that targets at achieving good energy efficiency and system performance in energy harvesting real-time systems. The proposed algorithm utilizes static and adaptive scheduling techniques combined with dynamic voltage and frequency selection to achieve good system performance under timing and energy constraints. In our approach, we simplify the scheduling and optimization problem by separating constraints in timing and energy domains. The proposed algorithm achieves improved system performance by exploiting task slack with dynamic voltage and frequency selection and minimizing the waste on harvested energy. Experimental results show that the proposed algorithm improves the system performance in deadline miss rate and the minimum storage capacity requirement for zero deadline miss rate. Comparing to the existing algorithms, the proposed algorithm achieves better performance in terms of the deadline miss rate and the minimum storage capacity under various settings of workloads and harvested energy profiles.
Qing Wu 0002, Qinru Qiu
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Dynamic thermal management for multimedia applications using machine learning
abstract
Multimedia applications are expected to form the largest portion of workload in general purpose PC and portable devices. The ever-increasing computation intensity of multimedia applications elevates the processor temperature and consequently impairs the reliability and performance of the system. In this paper, we propose to perform dynamic thermal management using reinforcement learning algorithm for multimedia applications. The proposed learning model does not need any prior knowledge of the workload information or the system thermal and power characteristics. It learns the temperature change and workload switching patterns by observing the temperature sensor and event counters on the processor, and finds the management policy that provides good performance-thermal tradeoff during the runtime. We validated our model on a Dell personal computer with Intel Core 2 processor. Experimental results show that our approach provides considerable performance improvements with marginal increase in the percentage of thermal hotspot comparing to existing workload phase detection approach.
Qinru Qiu
DAC2
2011 An FPGA-Based Distributed Computing System with Power and Thermal Management Capabilities
abstract
Runtime power and thermal management has attracted substantial interests in multi-core distributed embedded systems. Fast performance evaluation is an essential step in the research of distributed power and thermal management. Compared to software simulation, an FPGA-based evaluation platform provides fast emulation speed which enables us to test the performance of power/thermal management policies with real-life applications and OS. Compared to computer clusters, an FPGA-based platform has the flexibility to be configured into any network topology and hardware sniffer for performance monitoring can be added easily. This paper presents an FPGA based emulator of multi-core distributed embedded system designed to support the research in runtime power/thermal management. The system consists of multiple FPGAs connecting through Ethernet with each FPGA configured as a multi-core system. Hardware and software supports are provided to carry out basic power/thermal management actions including inter-core or inter-FPGA communications, runtime temperature monitoring and dynamic frequency scaling.
Hao Shen 0007, Qinru Qiu
ICCCN2
2011 Unified perception-prediction model for context aware text recognition on a heterogeneous many-core platform
abstract
Existing optical character recognition (OCR) software tools can perform text image detection and pattern recognition with fairly high accuracy, however their performance will be significantly impaired when the image of the character is partially blocked or smudged. Such missing information does not hinder the human perception because we predict the missing part based on the word level and sentence level context of the character. In order to mimic the human cognitive behavior, we developed a hybrid cognitive architecture combining two neuromorphic computing models, i.e. brain-state-in-a-box (BSB) and cogent confabulation, to achieve context-aware text recognition. The BSB model performs the character recognition from input image while the confabulation models perform the context-aware prediction based on the word and sentence knowledge bases. The software tool is implemented on an 1824-core computing cluster. Its accuracy and performance are analyzed in the paper.
Qinru Qiu, Qing Wu 0002, Richard W. Linderman
IJCNN1
2010 Distributed task migration for thermal management in many-core systems
abstract
In the deep submicron era, thermal hot spots and large temperature gradients significantly impact system reliability, performance, cost and leakage power. As the system complexity increases, it is more and more difficult to perform thermal management in a centralized manner because of state explosion and the overhead of monitoring the entire chip. In this paper, we propose a framework for distributed thermal management for many-core systems where balanced thermal profile can be achieved by proactive task migration among neighboring cores. The framework has a low cost agent residing in each core that observes the local workload and temperature and communicates with its nearest neighbor for task migration/exchange. By choosing only those migration requests that will result balanced workload without generating thermal emergency, the proposed framework maintains workload balance across the system and avoids unnecessary migration. Experimental results show that, compared with existing proactive task migration technique, our approach generates less hotspots and smoother thermal gradient with less migration overhead and higher processing throughput.
Parth Malani, Qinru Qiu
DAC3
2010 Enhanced Q-learning algorithm for dynamic power management with performance constraint
abstract
This paper presents a novel power management techniques based on enhanced Q-learning algorithms. By exploiting the sub modularity and monotonic structure in the cost function of a power management system, the enhanced Q-learning algorithm is capable of exploring ideal trade-offs in the power-performance design space and converging to a better power management policy. We further propose a linear adaption algorithm that adapts the Lagrangian multiplier ¿ to search for the power management policy that minimizes the power consumption while delivering the exact required performance. Experimental results show that, comparing to the existing expert-based power management, the proposed Q-learning based power management achieves up to 30% and 60% reduction in power saving for synthetic workload and real workload, respectively while in average maintain a performance within 7% variation of the given constraint.
Qinru Qiu
DATE3
2010 Load-matching adaptive task scheduling for energy efficiency in energy harvesting real-time embedded systems
abstract
In this paper we present a load matching task scheduling algorithm for energy harvesting real-time embedded systems using a realistic model for the battery charging and discharging processes. The proposed approach addresses two important issues that have not been considered by previous work: load matching and battery charge/discharge overhead. The new algorithm increases available energy by managing the system load through task scheduling so that the energy harvesting module delivers maximum power output. It further improves the system wide energy efficiency by considering the charging and discharging overhead when deciding if the harvested energy should be used to charge the battery or directly on the circuits. Experimental results show that, comparing to the best of the existing techniques the proposed algorithm improves the system wide energy efficiency by 8.0% to 56.3% and reduces deadline misses by 13.3% to 81.8% under different workload conditions.
Qing Wu 0002, Qinru Qiu
ISLPED4
2009 An adaptive scheduling and voltage/frequency selection algorithm for real-time energy harvesting systems
abstract
In this paper we propose an adaptive scheduling and voltage/frequency selection algorithm which targets at energy harvesting systems. The proposed algorithm adjusts the processor operating frequency under the timing and energy constraints based on workload information so that the system-wide energy efficiency is achieved. In this approach, we decouple the timing and energy constraints and simplify the original scheduling problem by separating constraints in timing and energy domains. The proposed algorithm utilizes maximum task slack for energy saving. Experimental results show that the proposed method improves the system performance in remaining energy, deadline miss rate and the minimum storage capacity requirement for zero deadline miss rate. Comparing to the existing algorithms, the new algorithm decreases the deadline miss rate by at least 23%, and the minimum storage capacity by at least 20% under various processor utilizations.
Qing Wu 0002, Qinru Qiu
DAC3
2009 Adaptive power management using reinforcement learning
abstract
System level power management must consider the uncertainty and variability that comes from the environment, the application and the hardware. A robust power management technique must be able to learn the optimal decision from past history and improve itself as the environment changes. This paper presents a novel online power management technique based on model-free constrained reinforcement learning (RL). It learns the best power management policy that gives the minimum power consumption for a given performance constraint without any prior information of workload. Compared with existing machine learning based power management techniques, the RL based learning is capable of exploring the trade-off in the power-performance design space and converging to a better power management policy. Experimental results show that the proposed RL based power management achieves 24% and 3% reduction in power and latency respectively comparing to the existing expert based power management.
Qinru Qiu
ICCAD3
2008 Energy Aware Dynamic Voltage and Frequency Selection for Real-Time Systems with Energy Harvesting
abstract
In this paper, an energy aware dynamic voltage and frequency selection (EA-DVFS) algorithm is proposed. The EA-DVFS algorithm adjusts the processor's behavior depending on the summation of the stored energy and the harvested energy in a future duration. Specifically, if the system has sufficient energy, tasks are executed at full speed; otherwise, the processor slows down task execution to save energy. Simulation results show that when the utilization is low, the EA-DVFS algorithm gives a deadline miss rate that is at least 50% lower than the one given by the lazy scheduling policy. Similarly, when the workload is low, the minimum storage size is reduced by at least 25%.
Qinru Qiu, Qing Wu 0002
DATE2
2008 Adaptive Scheduling and Voltage Scaling for Multiprocessor Real-time Applications with Non-deterministic Workload
abstract
The computational workload of some real-time applications varies significantly during runtime, which makes the task scheduling and power management a challenge. One of the major influences to the workload of an application is the selection of conditional branches which may activate or deactivate a large set of operations. Focusing on real-time applications with variable workload which is due to random branch selection, this paper presents a framework of task mapping, scheduling and dynamic voltage and frequency scaling (DVFS) for a multiprocessor system. The proposed framework maintains workload awareness using dynamic profiling of branch probability. The profiled information is utilized by the scheduling and DVFS algorithm that are adopted in this framework to generate statistically optimal solution.
Parth Malani, Prakash Mukre, Qinru Qiu, Qing Wu 0002
DATE3
2008 A Framework of Stochastic Power Management Using Hidden Markov Model
abstract
The effectiveness of stochastic power management relies on the accurate system and workload model and effective policy optimization. Workload modeling is a machine learning procedure that finds the intrinsic pattern of the incoming tasks based on the observed workload attributes. Markov Decision Process (MDP) based model has been widely adopted for stochastic power management because it delivers provable optimal policy. Given a sequence of observed workload attributes, the hidden Markov model (HMM) of the workload is trained. If the observed workload attributes and states in the workload model do not have one-to-one correspondence, the MDP becomes a Partially Observable Markov Decision Process (POMDP). This paper presents a framework of modeling and optimization for stochastic power management using HMM and POMDP. The proposed technique discovers the HMM of the workload by maximizing the likelihood of the observed attribute sequence. The POMDP optimization is formulated and solved as a quadraticly constrained linear programming (QCLP). Compared with traditional optimization technique, which is based on value iteration, the QCLP based optimization provides superior policy by enabling stochastic control.
Qinru Qiu
DATE2
2008 Full-chip leakage current estimation based on statistical sampling techniques
abstract
In this paper, we propose statistical sampling techniques in estimating the mean and distribution of full-chip leakage current under process variations. The stratified random sampling procedures are used to estimate the mean and variance of the full-chip leakage, under intra-die and inter-die process variations. Statistical quantile estimation method is then applied to estimate the cumulative distribution function. Experimental results show that, comparing to simple random sampling, the proposed approaches improve the estimation speed by 2.7X, on average.
Qinru Qiu, Qing Wu 0002
ACM Great Lakes Symposium on VLSI2
2008 Performance optimization for pattern recognition using associative neural memory
abstract
In this paper, we present our work in the implementation and performance optimization of the recall operation of the Brain-State-in-a-Box (BSB) model on the Cell Broadband Engine processor. We have applied optimization techniques on different parts of the algorithm to improve the overall computing and communication performance of the BSB recall algorithm. Runtime measurements show that, we have been able to achieve about 70% of the theoretical peak performance of the processor.
Qing Wu 0002, Prakash Mukre, Richard W. Linderman, Thomas Renz, Daniel J. Burns, Michael J. Moore, Qinru Qiu
ICME7
2008 Accelerating cogent confabulation: An exploration in the architecture design space
abstract
Cogent confabulation is a computation model that mimics the Hebbian learning, information storage, inter-relation of symbolic concepts, and the recall operations of the brain. The model has been applied to cognitive processing of language, audio and visual signals. In this project, we focus on how to accelerate the computation which underlie confabulation based sentence completion through software and hardware optimization. On the software implementation side, appropriate data structures can improve the performance of the software by more than 5,000X. On the hardware implementation side, the cogent confabulation algorithm is an ideal candidate for parallel processing and its performance can be significantly improved with the help of application specific, massively parallel computing platforms. However, as the complexity and parallelism of the hardware increases, cost also increases. Architectures with different performance-cost tradeoffs are analyzed and compared. Our analysis shows that although increasing the number of processors or the size of memories per processor can increase performance, the hardware cost and performance improvements do not always exhibit a linear relation. Hardware configuration options must be carefully evaluated in order to achieve good cost performance tradeoffs.
Qinru Qiu, Daniel J. Burns, Michael J. Moore, Richard W. Linderman, Thomas Renz, Qing Wu 0002
IJCNN1
2008 A probabilistic technique for full-chip leakage estimation
abstract
In this paper, we propose a probability-based algorithm to estimate full-chip leakage without knowing layout information, under intra-die and inter-die process variations. Through modeling process variations into a random vector, we show that the standard cell leakage can be modeled as an inverse Gaussian random variable and further demonstrate that full-chip leakage can also be approximated to be an inverse Gaussian random variable. Hence, the leakage estimation problem is reduced to the estimation of the mean value and variance of the full-chip leakage. Experimental results show that the proposed algorithm is over 1000X faster than Monte Carlo simulation while the maximum estimation error is less than 6%.
Qinru Qiu, Qing Wu 0002
ISLPED2
2008 Bus encoding for simultaneous delay and energy optimization
abstract
In this paper we propose two bus encoding algorithms that optimize both bus delay and energy dissipation based on the probabilistic characteristics of data on data buses. The first algorithm minimizes the crosstalk transitions by inserting temporal redundancy and achieves optimal energy. The second algorithm reduces crosstalk more aggressively to achieve optimal bus delay by mapping the original data to low-energy opposite-transition-forbidden codes. Experimental results show that they outperform the existing heuristic bus encoding algorithms by 15.7% to 58.8% in average energy dissipation and 11.4% to 58.4% in average delay.
Qing Wu 0002, Qinru Qiu
ISLPED3
2007 Hybrid Architecture for Accelerating DNA Codeword Library Searching
abstract
A large and reliable DNA codeword library is the key to the success of DNA based computing. Searching for the set of reliable DNA codewords is an NP-hard problem, which can take days on the state-of-art high performance cluster computers. This work presents a hybrid architecture that consists of a general purpose microprocessor and a hardware accelerator for accelerating the discovery of DNA reverse complement, edit distance codes. Two applications of this architecture were implemented and evaluated, including a code generator that uses a genetic algorithm (GA) to produce nearly locally optimal codes in a few minutes, and a code extender that uses exhaustive search to produce locally optimum codes in about 1.5 hours for the case of length 16 codes. The experimental results demonstrate that the GA can find ~99% of the words in locally optimum libraries, and that the hybrid architecture provides more than 1000X speed-up compared to a software only implementation
Qinru Qiu, Daniel J. Burns, Qing Wu 0002, Prakash Mukre
CIBCB1
2007 Architectural Design and Complexity Analysis of Large-Scale Cortical Simulation on a Hybrid Computing Platform
abstract
Research and development in modeling and simulation of human cognizance functions requires a high-performance computing platform for manipulating large-scale mathematical models. Traditional computing architectures cannot fulfill the attendant needs in terms of arithmetic computation and communication bandwidth. In this work, we propose a novel hybrid computing architecture for the simulation and evaluation of large-scale associative neural memory models. The proposed architecture achieves very high computing and communication performances by combining the technologies of hardware-accelerated computing, parallel distributed data operation and the publish/subscribe protocol. Analysis has been done on the computation and data bandwidth demands for implementing a large-scale brain-state-in-a-box (BSB) model. Compared to the traditional computing architecture, the proposed architecture can achieve at least 100X speedup.
Qing Wu 0002, Qinru Qiu, Richard W. Linderman, Daniel J. Burns, Michael J. Moore, Dennis Fitzgerald
CISDA2
2007 Stochastic modeling and optimization for robust power management in a partially observable system
abstract
As the hardware and software complexity grows, it is unlikely for the power management hardware/software to have a full observation of the entire system status. In this paper, we propose a new modeling and optimization technique based on partially observable Markov decision process (POMDP) for robust power management, which can achieve near-optimal power savings, even when only partial system information is available. Three scenarios of partial observations that may occur in an embedded system are discussed and their modeling techniques are presented. The experimental results show that, compared with power management policy derived from traditional Markov decision process model that assumes the system is fully observable, the new power management technique gives significantly better performance and energy tradeoff
Qinru Qiu, Qing Wu 0002
DATE1
2007 Hardware Acceleration for Thermodynamic Constrained DNA Code Generation
Qinru Qiu, Prakash Mukre, Morgan Bishop, Daniel J. Burns, Qing Wu 0002
DNA1
2007 Hardware acceleration of multi-deme genetic algorithm for the application of DNA codeword searching
abstract
A large and reliable DNA codeword library is key to the success of DNA based computing. Searching for sets of reliable DNA codewords is an NP-hard problem, which can take days on state-of-art high performance cluster computers. This work presents a hybrid architecture that consists of a general purpose microprocessor and a hardware accelerator for accelerating the multi-deme genetic algorithm (GA) for the application of DNA codeword searching. The presented architecture provides more than 1000X speed-up compared to a software only implementation. A code extender that uses exhaustive search to produce locally optimum codes in about 1.5 hours for the case of length 16 codes is also described. The experimental results demonstrate that the GA can find ~99% of the words in locally optimum libraries. Finally, we investigate the performance impact of migration, mating and mutation functions in the hardware accelerator. The analysis shows that a modified GA without mating is the most effective for DNA codeword searching.
Qinru Qiu, Daniel J. Burns, Prakash Mukre, Qing Wu 0002
GECCO1
2007 Resource-aware High Performance Scheduling for Embedded MPSoCs With the Application of MPEG Decoding
abstract
In this paper, we propose a scheduling algorithm to minimize the resource contentions and the processing latency for applications running on a multiprocessor system-on-chip (MPSoC) platform. The scheduling algorithm is applied on an MPSoC MPEG decoder to improve the system performance. Application specific task partition and mapping techniques are further investigated. The experimental results show an average improvement of 17% in total latency when comparing to the ad-hoc scheduled method.
Parth Malani, Qinru Qiu
ICME3
2007 Profile-Based Low Power Scheduling for Conditional Task Graph: A Communication Aware Approach
abstract
This work focuses on power optimization of realtime applications with conditional execution running on a dynamic voltage scaling (DVS) enabled multiprocessor system. A novel algorithm is proposed that performs simultaneous task mapping and ordering followed by task stretching of a conditional task graph (CTG). The algorithm minimizes the mathematical expectation of energy dissipation of non-deterministic applications with random branch selection by utilizing the task execution profile. Compared with existing scheduling algorithm, the experimental results show that our algorithm has 32% energy reduction in average.
Parth Malani, Prakash Mukre, Qinru Qiu
ISCAS3
2007 Power optimization for conditional task graphs in DVS enabled multiprocessor systems
abstract
In this paper, we focus on power optimization of real-time applications with conditional execution running on a dynamic voltage scaling (DVS) enabled multiprocessor system. The targeted system consists of heterogeneous processing elements with non-negligible inter-processor communication delay and energy. Given a conditional task graph (CTG), we have developed novel online and offline algorithms that perform simultaneous task mapping and ordering followed by task stretching. Both algorithms minimize the mathematical expectation of energy dissipation of non-deterministic applications by considering the probabilistic distribution of branch selection. Compared with existing CTG scheduling algorithms, our online and offline scheduling algorithms reduce energy by 28% and 39% in average, respectively.
Parth Malani, Prakash Mukre, Qinru Qiu
VLSI-SoC3
2006 Workload prediction and dynamic voltage scaling for MPEG decoding
abstract
In this paper we present three efficient DVS techniques for an MPEG decoder. Their energy reduction is comparable to that of the optimal solution. A workload prediction model is also developed based on the block level statistics of each MPEG frame. Compared with previous works, the new model exhibits a remarkable improvement in accuracy of the prediction. The experimental results show that, with the new prediction model, the presented DVS techniques achieve more energy reduction than previous works while delivering the same Quality of Service (QoS)
Parth Malani, Qinru Qiu, Qing Wu 0002
ASP-DAC3
2006 Low-power bus encoding using an adaptive hybrid algorithm
abstract
In this paper, we propose an adaptive low-power bus encoding algorithm based on weighted code mapping (WCM) and the delayed bus technique. The WCM algorithm transforms an original bus data vector to a low-energy code through one-to-one mapping. The code mapping is determined by the data probabilistic distribution in the original sequence. The WCM algorithm considers both the self and coupling capacitance of the bus wires. In addition, we found that applying the delayed-bus technique can further reduce the bus energy. A window-based adaptive encoding algorithm is proposed to improve the energy saving by adaptively changing the code mapping when significant data changes are detected. Experimental results show that the proposed algorithm outperforms the existing heuristic bus encoding algorithms by 20~60% in energy dissipation.
Avnish R. Brahmbhatt, Qing Wu 0002, Qinru Qiu
DAC4
2006 Distributed genetic algorithm for energy-efficient resource management in sensor networks
abstract
In this work we consider energy-efficient resource management in an environment monitoring and hazard detection sensor network. Our goal is to allocate different detection methods to different sensor nodes in the way such that the required detection probability can be achieved while the network lifetime is maximized. The optimization algorithm is designed based on the Island multi-deme genetic algorithm (GA). The experimental results show that our algorithm increases the network lifetime by approximately 14.4% in average compared with the heuristic approaches. We also investigate the effect of the configuration parameters on the searching quality of the proposed distributed GA. A regression model is derived empirically that estimates the runtime of the distributed GA given the configuration parameters such as the sub-population size, parallelism, and migration rate. Once the model has been fit to a group of data, it can be utilized to find the efficient configurations of the proposed algorithm.
Qinru Qiu, Qing Wu 0002, Daniel J. Burns, Douglas Holzhauer
GECCO1
2006 Task Merging for Dynamic Power Management of Cyclic Applications in Real-Time Multi-Processor Systems
abstract
In this paper we propose the method of task merging and idle period clustering for dynamic power management (DPM) in a real-time system with multiple processing elements. We show that with good task scheduling, the energy and delay overheads due to power mode switching can be reduced significantly, while the opportunity for the system to switch to low power modes can be further improved. New on-line and off-line task scheduling algorithms are proposed that minimize the number of idle time intervals under the deadline and precedence constraints. A simple DPM policy is then used to save the energy dissipation during the idle time intervals. Experimental results show that, comparing to the DPM schemes without proper task scheduling, the proposed method reduces the number of power mode switching by 56% in average.
Qinru Qiu, Qing Wu 0002
ICCD2
2006 Adaptive low-power bus encoding based on weighted code mapping
abstract
In this paper, we propose an adaptive low-power bus encoding algorithm based on weighted code mapping (WCM). The WCM algorithm transforms an original bus data vector to a low-energy code through one-to-one mapping. The code mapping is determined by the data probability distribution in the original sequence. The WCM algorithm considers both the self and coupling capacitance of the bus wires. A window-based adaptive encoding algorithm is proposed to improve the energy saving by adaptively changing the code mapping for different data probability characteristics. Experimental results show that the proposed algorithm outperforms the existing coding algorithms by either significantly lower computation/hardware complexity or higher energy savings.
Avnish R. Brahmbhatt, Qinru Qiu, Qing Wu 0002
ISCAS3
2006 Design considerations for digital circuits using organic thin film transistors on a flexible substrate
abstract
Organic thin film transistor (OTFT) is the basic device for building analog and digital circuits and systems on a flexible substrate. Our studies show that, complementary design logic (CMOS), which is most common in MOSFET circuit design, is not suitable for OTFT circuits. In this work, we propose a new design logic that fits the characteristics of OTFT devices. Considering the fact that the mobility of the p-type OTFT is much better than the n-type OTFT, the proposed design logic uses mostly p-type OTFTs to implement logic functions. Circuit simulations have been done to compare the proposed design logic with other existing ones. The results show that the proposed design logic can improve the performance of the circuit by 2/spl sim/4/spl times/, lower the energy dissipation by 2/spl sim/6/spl times/, and reduce the circuit area by 2/spl sim/10/spl times/.
Qing Wu 0002, Qinru Qiu
ISCAS3
2006 Lifetime aware resource management for sensor network using distributed genetic algorithm
abstract
In this work we consider lifetime-aware resource management for sensor network using distributed genetic algorithm (GA). Our goal is to allocate different detection methods to different sensor nodes in the way such that the required detection probability can be achieved while the network lifetime is maximized. The contribution of this paper is twofold. Firstly, the resource management problem is formulated as a constraint optimization problem and is solved using a distributed GA. Secondly, empirical analysis results are provided that reveals the relationship between the configuration parameters and the quality of the search. A regression model is designed to estimate the runtime of the distributed GA given the configuration parameters. The model is utilized to find energy efficient configurations of the algorithm.
Qinru Qiu, Qing Wu 0002, Daniel J. Burns, Douglas Holzhauer
ISLPED1
2006 Low-Density Parity-Check Coded Distributed Space-Time Cooperative System
abstract
The concatenation scheme of low-density parity-check (LDPC) codes and space-time block coding (LDPC-STBC) has been proved to be effective in the multiple-input multiple-output (MIMO) system. In this paper, we extend the LDPC-STBC to the distributed communication system with non-regenerative relays. Two LDPC coded distributed space-time cooperative (LDPC-DSTC) schemes are proposed, which utilize different signal combination methods at the destination. The decoding algorithms are discussed. The performance of these proposed schemes is examined in wireless communication systems with one and two relays. Simulation results show that both of the proposed LDPC-DSTC schemes reduce the transmission error and offer the diversity gain, and that the location of the relays has profound impact on the system performance
Peiliang Qiu, Qinru Qiu
VTC Spring4
2005 Partitioned bus coding for energy reduction
abstract
For VLSI design in deep submicron technology, the bus energy reduction has become more and more important. This paper studies the bus partition scheme for the Transition Pattern Coding (TPC). The genetic algorithm based approach is used. A closed-form expression is derived to calculate the energy dissipation for the partitioned bus with TPC coding. A general bus model with coupling capacitance is considered during the energy estimation and optimization. The resulted partitioned bus coding reduces the encoding and decoding complexity of the original TPC. The experimental results show that the TPC with careful bus partitioned saves up to 16.9% the energy of the TPC with random bus partition.
Peiliang Qiu, Qinru Qiu
ASP-DAC3
2004 ESACW: an adaptive algorithm for transmission power reduction in wireless networks
abstract
In this paper we propose a new algorithm for reducing the energy dissipation of a wireless ad-hoc network. We first show that the performance and energy dissipation is a function of the probability of packet collision, which can be varied by changing the minimum contention window (CWmin) parameter. Then we propose an algorithm, based on the IEEE 802.11 protocol, which can dynamically adjust CWmin for better performance and power. Experimental results show that, comparing to the original protocol, the proposed method can save 30% to 60% energy dissipation, and achieve similar or better performance.
Peiliang Qiu, Qinru Qiu
ISLPED3
2001 Dynamic Power Management in a Mobile Multimedia System with Guaranteed Quality-of-Service
abstract
In this paper we address the problem of dynamic power management in a distributed multimedia system with a required quality of service (QoS). Using a generalized stochastic Petri net model where the non-exponential inter-arrival time distribution of the incoming requests is captured by a stage method, we provide a detailed model of the power-managed multimedia system under general QoS constraints. Based on this mathematical model, the power-optimal policy is obtained by solving a linear programming problem. We compare the new problem formulation and solution technique to previous dynamic power management techniques that can only optimize power under delay constraints and demonstrate that these other techniques yield policies with higher power dissipation by over-constraining the delay target in an attempt to indirectly satisfy the QoS constraints. In contrast, our new method correctly formulates the power management problem under QoS constraints and obtains the optimal solution.
Qinru Qiu, Qing Wu 0002, Massoud Pedram
DAC1
2001 Stochastic modeling of a power-managed system-construction andoptimization
abstract
The goal of a dynamic power management policy is to reduce the power consumption of an electronic system by putting system components into different states, each representing a certain performance and power consumption level. The policy determines the type and timing of these transitions based on the system history, workload, and performance constraints. In this paper we propose a new abstract model of a power-managed electronic system. We formulate the problem of system-level power management as a controlled optimization problem based on the theories of continuous-time Markov derision processes and stochastic networks. This problem is solved exactly using linear programming or heuristically using "policy iteration." Our method is compared with existing heuristic methods for different workload statistics. Experimental results show that the power management method based on a Markov decision process outperforms heuristic methods by as much as 44% in terms of power dissipation savings for a given level of system performance.
Qinru Qiu, Massoud Pedram
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2001 Estimation of peak power dissipation in VLSI circuits using thelimiting distributions of extreme order statistics
abstract
In this paper, we present a statistical method for estimating the peak power dissipation in very large scale integrated (VLSI) circuits. The method is based on the theory of extreme order statistics and its application to the probabilistic distributions of the cycle-by-cycle power consumption, the maximum-likelihood estimation, and the Monte-Carlo simulation. It enables us to predict the maximum power of a VLSI circuit in the set of constrained input vector pairs as well as the complete set of all possible input vector pairs. The simulation-based nature of the proposed method allows us to avoid the limitations of a gate-level delay model and a gate-level circuit structure. Most significantly, the proposed method produces maximum power estimates to satisfy user-specified error and confidence levels. Experimental results show that this method typically produces maximum power estimates within 5% of the actual value and with a 90% confidence level by only simulating less than 2500 input vectors.
Qing Wu 0002, Qinru Qiu, Massoud Pedram
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 An interleaved dual-battery power supply for battery-operated electronics
abstract
After a detailed analysis and discussion of two important characteristics of today's battery cells (i.e., their current-capacity and current-voltage curves), this paper describes the design principles and architecture of a dual-battery power supply system for portable electronics. The key idea is to integrate two battery types with different energy capacity and current rate curves into the power supply system, and then use them in an interleaved manner in response to varying current requirement of the VLSI circuit that is powered by this dual-battery system. Analytical and empirical results demonstrate the effectiveness of the new battery architecture in maximizing the service life of a battery system with fixed volume (or weight).
Qing Wu 0002, Qinru Qiu, Massoud Pedram
ASP-DAC2
2000 Dynamic power management of complex systems using generalized stochastic Petri nets
abstract
In this paper, we introduce a new technique for modeling and solving the dynamic power management (DPM) problem for systems with complex behavioral characteristics such as concurrency, synchronization, mutual exclusion and conflict. We model a power-managed distributed computing system as a controllable Generalized Stochastic Petri Net (GSPN) with cost. The obtained GSPN model is automatically converted to an equivalent continuous-time Markov decision process. Given the delay constraints, the optimal power management policy for system components as well as the optimal dispatch policy for requests are calculated by solving a linear programming problem based on the Markov decision process. Experimental results show that the proposed technique can achieve more than 20% power saving compared to other existing DPM techniques.
Qinru Qiu, Qing Wu 0002, Massoud Pedram
DAC1
1999 Dynamic Power Management Based on Continuous-Time Markov Decision Processes
abstract
AbstractThis paper introduces a continuous-time, controllable Markov process model of a power-managed system.The system model is composed of the corresponding stochastic models of the service queue and the service provider.The system environment is modeled by a stochastic service request process.The problem of dynamic power management in such a system is formulated as a policy optimization problem and solved using an efficient "policy iteration" algorithm.Compared to previous work on dynamic power management, our formulation allows better modeling of the various system components, the power-managed system as a whole, and its environment.In addition it captures dependencies between the service queue and service provider status.Finally, the resulting power management policy is asynchronous, hence it is more power-efficient and more useful in practice.Experimental results demonstrate the effectiveness of our policy optimization algorithm compared to a number of heuristic (time-out and Npolicy) algorithms.I.
Qinru Qiu, Massoud Pedram
DAC1
1999 Stochastic modeling of a power-managed system: construction and optimization
abstract
Article Stochastic modeling of a power-managed system: construction and optimization Share on Authors: Qinru Qiu Department of Electrical Engineering-Systems, University of Southern California, Los Angeles, CA Department of Electrical Engineering-Systems, University of Southern California, Los Angeles, CAView Profile , Qing Wu Department of Electrical Engineering-Systems, University of Southern California, Los Angeles, CA Department of Electrical Engineering-Systems, University of Southern California, Los Angeles, CAView Profile , Massoud Pedram Department of Electrical Engineering-Systems, University of Southern California, Los Angeles, CA Department of Electrical Engineering-Systems, University of Southern California, Los Angeles, CAView Profile Authors Info & Claims ISLPED '99: Proceedings of the 1999 international symposium on Low power electronics and designAugust 1999 Pages 194–199https://doi.org/10.1145/313817.313923Online:17 August 1999Publication History 38citation199DownloadsMetricsTotal Citations38Total Downloads199Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Qinru Qiu, Qing Wu 0002, Massoud Pedram
ISLPED1
1998 Maximum Power Estimation Using the Limiting Distributions of Extreme Order Statistics
abstract
In this paper we present a statistical method for estimating the maximum power consumption in VLSI circuits. The method is based on the theory of extreme order statistics applied to the probabilistic distribution of the cycle-based power consumption, maximum likelihood estimation, and Monte-Carlo simulation. The method can predict the maximum power in the constrained space of given input vector pairs as well as the complete space of all possible input vector pairs. The simulation-based nature of the proposed method allows one to avoid the limitations imposed by simple gate-level delay models and handle arbitrary circuit structures. The proposed method can produce maximum power estimates to satisfy user-specified error and confidence levels. Experimental results show that this method provides maximum power estimates within 5% of the actual value and with a 90% confidence level by simulating, on average, about 2500 vector pairs.
Qinru Qiu, Qing Wu 0002, Massoud Pedram
DAC1
1998 Cycle-accurate macro-models for RT-level power analysis
abstract
In this paper, we present a methodology and techniques for generating cycle-accurate macro-models for register transfer (RT)-level power analysis. The proposed macro-model predicts not only the cycle-by-cycle power consumption of a module, but also the moving average of power consumption and the power profile of the module over time. We propose an exact power function and approximation steps to generate our power macro-model. First-order temporal correlations and spatial correlations of up to order three are considered in order to improve the estimation accuracy. A variable reduction algorithm is designed to eliminate the "insignificant" variables using a statistical sensitivity test. Population stratification is employed to increase the model fidelity. Experimental results show our macro-models with 15 or fewer variables, exhibit <5% error for average power and <20% errors for cycle-by-cycle power estimation compared to circuit simulation results using Powermill.
Qing Wu 0002, Qinru Qiu, Massoud Pedram, Chih-Shun Ding
IEEE Trans. Very Large Scale Integr. Syst.2
1997 Cycle-accurate macro-models for RT-level power analysis
abstract
In this paper we present a methodology and techniques for generating cycle-accurate macro-models for RT-level power analysis. The proposed macro-model predicts not only the cycle-by-cycle power consumption of a module, but the power profile of the module over time.The proposed methodology consists of three steps: module equation form generation and variable selection, variable reduction, and population stratification.First order temporal correlations and spatial correlations of up to order 3 are considered to improve the estimation accuracy.Experimental results show that, the macro-models have 15 or less variables and exhibit 4% error in average power, and 45% errors in cycle-by-cycle power compared to circuit simulation results using Powermill.I.
Qinru Qiu, Qing Wu 0002, Massoud Pedram, Chih-Shun Ding
ISLPED1