EDBT 2026 Demo / reviewers in the wild / expert
Robert P. Dick
dblp:84/523 · also Robert Paul Dick
· DBLP profile ↗
123ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0001-5428-9530ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 71 · 11 first-author · 2 since 2021Software engineering, systems software and programming languages · 21 · 1 first-authorArtificial intelligence and machine learning · 15 · 14 since 2021Computer networks · 15 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Human-computer interaction and ubiquitous computing · 9Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The GOODPUT System: A Machine Learning-Driven Optimization Framework for Dynamic Spectrum Control in Heterogeneous WLANs
Robert P. Dick |
NSDI | 2 |
| 2025 | A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal LanguageabstractIncrease in data, size, or compute can lead to sudden learning of specific capabilities by a neural network---a phenomenon often called "emergence". Beyond scientific understanding, establishing the causal factors underlying such emergent capabilities is crucial to enable risk regulation frameworks for AI. In this work, we seek inspiration from study of emergent properties in other fields and propose a phenomenological definition for the concept in the context of neural networks. Our definition implicates the acquisition of general regularities underlying the data-generating process as a cause of sudden performance growth for specific, narrower tasks. We empirically investigate this definition by proposing an experimental system grounded in a context-sensitive formal language, and find that Transformers trained to perform tasks on top of strings from this language indeed exhibit emergent capabilities. Specifically, we show that once the language's underlying grammar and context-sensitivity inducing regularities are learned by the model, performance on narrower tasks suddenly begins to improve. We then analogize our network's learning dynamics with the process of percolation on a bipartite graph, establishing a formal phase transition model that predicts the shift in the point of emergence observed in our experiments when intervening on the data regularities. Overall, our experimental and theoretical frameworks yield a step towards better defining, characterizing, and predicting emergence in neural networks. Ekdeep Singh Lubana, Kyogo Kawaguchi, Robert P. Dick, Hidenori Tanaka |
ICLR | 3 |
| 2025 | Denoising Reuse: Exploiting Inter-Frame Motion Consistency for Efficient Video GenerationabstractDenoising-based diffusion models have attained impressive image synthesis; however, their applications on videos can lead to unaffordable computational costs due to the per-frame denoising operations. In pursuit of efficient video generation, we present a Diffusion Reuse MOtion (Dr. Mo) network to accelerate the video-based denoising process. Our crucial observation is that the latent representations in early denoising steps between adjacent video frames exhibit high consistencies with motion clues. Inspired by the discovery, we propose to accelerate the video denoising process by incorporating lightweight, learnable motion features. Specifically, Dr. Mo will only compute all denoising steps for base frames. For a non-based frame, Dr. Mo will propagate the pre-computed based latents of a particular step with inter-frame motions to obtain a fast estimation of its coarse-grained latent representation, from which the denoising will continue to obtain more sensitive and fine-grained representations. On top of this, Dr. Mo employs a meta-network named Denoising Step Selector (DSS) to dynamically determine the step to perform motion-based propagations for each frame, ensuring the correct transformation of multi-granularity visual features. Extensive evaluations on video generation and editing tasks indicate that Dr. Mo delivers widely applicable acceleration for diffusion-based video generations while effectively retaining the visual quality and style. Video generation and visualization results can be found athttps://drmo-denoising-reuse.github.io. Yixuan Chen 0003, Yujiang Wang 0001, Mingzhi Dong, Dongsheng Li 0002, Rui Zhu 0006, David A. Clifton, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 11 |
| 2025 | Efficient Video Redaction at the Edge: Human Motion Tracking for Privacy ProtectionabstractComputationally efficient, camera-based, real-time human position tracking on low-end, edge devices would enable numerous applications, including privacy-preserving video redaction and analysis. Unfortunately, running most deep neural network based models in real time requires expensive hardware, making widespread deployment difficult, particularly on edge devices. Shifting inference to the cloud increases the attack surface, generally requiring that users trust cloud servers, and increases demands on wireless networks in deployment venues. Our goal is to determine the extreme to which edge video redaction efficiency can be taken, with a particular interest in enabling, for the first time, low-cost, real-time deployments with inexpensive commodity hardware. We present an efficient solution to the human detection (and redaction) problem based on singular value decomposition (SVD) background removal and describe a novel time-efficient and energy-efficient sensor-fusion algorithm that leverages human position information in real-world coordinates to enable real-time visual human detection and tracking at the edge. These ideas are evaluated using a prototype built from (resource-constrained) commodity hardware representative of commonly used low-cost IoT edge devices. The speed and accuracy of the system are evaluated via a deployment study, and it is compared with the most advanced relevant alternatives. The multi-modal system operates at a frame rate ranging from 20 FPS to 60 FPS, achieves a wIoU 0.3 score (see Section 5.4 ) ranging from 0.71 to 0.79, and successfully performs complete redaction of privacy-sensitive pixels with a success rate of 91%–99% in human head regions and 77%–91% in upper body regions, depending on the number of individuals present in the field of view. These results demonstrate that it is possible to achieve adequate efficiency to enable real-time redaction on inexpensive, commodity edge hardware. Haotian Qiao, Vidya Srinivas, Peter A. Dinda, Robert P. Dick |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | Introduction to Special Issue on Large Language Models for Electronic System Design AutomationabstractLarge Language Models are having a substantial impact on electronic design automation in areas ranging from hardware architecture to verification and optimization. The special issue provides a snapshot of work on this topic. This introduction describes and provides context for the research area, describes the organization of the special issue, and provides terse summaries of each of its papers. Robert P. Dick, Hammond A. Pearce, Li Shang 0002, Fan Yang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2024 | A Context-Oriented Multi-Scale Neural Network for Fire SegmentationabstractExisting image-based fire segmentation techniques use convolutional neural networks to handle complicated scenes. Such approaches perform poorly when flame sizes vary greatly and when backgrounds are complex. In this paper, we describe a novel Context-Oriented Multi-Scale Network for fire segmentation. We construct a multi-scale aggregation module that combines semantic information at different levels in the neural network in order to recognize fires with different shapes and sizes. We also describe a Context-Oriented Module, which increases the receptive field of the network by utilizing relationships of all pixels in the feature map in order to obtain features that more effectively discriminate between fire and non-fire pixels. Experimental results demonstrate that our proposed model has a $2.7 \%$ higher mean Intersection over Union (mIoU) accuracy than previous fire detection methods. Tony Zhang, Robert P. Dick |
ICIP | 2 |
| 2024 | In-Context Learning Dynamics with Random Binary SequencesabstractLarge language models (LLMs) trained on huge text datasets demonstrate intriguing capabilities, achieving state-of-the-art performance on tasks they were not explicitly trained for. The precise nature of LLM capabilities is often mysterious, and different prompts can elicit different capabilities through in-context learning. We propose a framework that enables us to analyze in-context learning dynamics to understand latent concepts underlying LLMs’ behavioral patterns. This provides a more nuanced understanding than success-or-failure evaluation benchmarks, but does not require observing internal activations as a mechanistic interpretation of circuits would. Inspired by the cognitive science of human randomness perception, we use random binary sequences as context and study dynamics of in-context learning by manipulating properties of context data, such as sequence length. In the latest GPT-3.5+ models, we find emergent abilities to generate seemingly random numbers and learn basic formal languages, with striking in-context learning dynamics where model outputs transition sharply from seemingly random behaviors to deterministic repetition. Eric J. Bigelow, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka, Tomer D. Ullman |
ICLR | 3 |
| 2024 | Mechanistically analyzing the effects of fine-tuning on procedurally defined tasksabstractFine-tuning large pre-trained models has become the de facto strategy for developing both task-specific and general-purpose machine learning systems, including developing models that are safe to deploy. Despite its clear importance, there has been minimal work that explains how fine-tuning alters the underlying capabilities learned by a model during pretraining: does fine-tuning yield entirely novel capabilities or does it just modulate existing ones? We address this question empirically in synthetic, controlled settings where we can use mechanistic interpretability tools (e.g., network pruning and probing) to understand how the model's underlying capabilities are changing. We perform an extensive analysis of the effects of fine-tuning in these settings, and show that: (i) fine-tuning rarely alters the underlying model capabilities; (ii) a minimal transformation, which we call a `wrapper', is typically learned on top of the underlying model capabilities, creating the illusion that they have been modified; and (iii) further fine-tuning on a task where such ``wrapped capabilities'' are relevant leads to sample-efficient revival of the capability, i.e., the model begins reusing these capabilities after only a few gradient steps. This indicates that practitioners can unintentionally remove a model's safety wrapper merely by fine-tuning it on a, e.g., superficially unrelated, downstream task. We additionally perform analysis on language models trained on the TinyStories dataset to support our claims in a more realistic setup. Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka, Tim Rocktäschel, Edward Grefenstette, David Krueger 0001 |
ICLR | 4 |
| 2024 | Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation ModelabstractStepwise inference protocols, such as scratchpads and chain-of-thought, help language models solve complex problems by decomposing them into a sequence of simpler subproblems. To unravel the underlying mechanisms of stepwise inference we propose to study autoregressive Transformer models on a synthetic task that embodies the multi-step nature of problems where stepwise inference is generally most useful. Specifically, we define a graph navigation problem wherein a model is tasked with traversing a path from a start to a goal node on the graph. We find we can empirically reproduce and analyze several phenomena observed at scale: (i) the stepwise inference reasoning gap, the cause of which we find in the structure of the training data; (ii) a diversity-accuracy trade-off in model generations as sampling temperature varies; (iii) a simplicity bias in the model’s output; and (iv) compositional generalization and a primacy bias with in-context exemplars. Overall, our work introduces a grounded, synthetic framework for studying stepwise inference and offers mechanistic hypotheses that can lay the foundation for a deeper understanding of this phenomenon. Mikail Khona, Maya Okawa, Jan Hula, Rahul Ramesh, Kento Nishi, Robert P. Dick, Ekdeep Singh Lubana, Hidenori Tanaka |
ICML | 6 |
| 2024 | Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable TasksabstractTransformers trained on huge text corpora exhibit a remarkable set of capabilities, e.g., performing simple logical operations. Given the inherent compositional nature of language, one can expect the model to learn to compose these capabilities, potentially yielding a combinatorial explosion of what operations it can perform on an input. Motivated by the above, we aim to assess in this paper “how capable can a transformer become?”. Specifically, we train autoregressive Transformer models on a data-generating process that involves compositions of a set of well-defined monolithic capabilities. Through a series of extensive and systematic experiments on this data-generating process, we show that: (1) autoregressive Transformers can learn compositional structures from small amounts of training data and generalize to exponentially or even combinatorially many functions; (2) composing functions by generating intermediate outputs is more effective at generalizing to unseen compositions, compared to generating no intermediate outputs; (3) biases in the order of the compositions in the training data, results in Transformers that fail to compose some combinations of functions; and (4) the attention layers seem to select the capability to apply while the feed-forward layers execute the capability. Rahul Ramesh, Ekdeep Singh Lubana, Mikail Khona, Robert P. Dick, Hidenori Tanaka |
ICML | 4 |
| 2024 | Once Read is Enough: Domain-specific Pretraining-free Language Models with Cluster-guided Sparse Experts for Long-tail Domain KnowledgeabstractLanguage models (LMs) only pretrained on a general and massive corpus usually cannot attain satisfying performance on domain-specific downstream tasks, and hence, applying domain-specific pretraining to LMs is a common and indispensable practice.
However, domain-specific pretraining can be costly and time-consuming, hindering LMs' deployment in real-world applications.
In this work, we consider the incapability to memorize domain-specific knowledge embedded in the general corpus with rare occurrences and long-tail distributions as the leading cause for pretrained LMs' inferior downstream performance.
Analysis of Neural Tangent Kernels (NTKs) reveals that those long-tail data are commonly overlooked in the model's gradient updates and, consequently, are not effectively memorized, leading to poor domain-specific downstream performance.
Based on the intuition that data with similar semantic meaning are closer in the embedding space, we devise a Cluster-guided Sparse Expert (CSE) layer to actively learn long-tail domain knowledge typically neglected in previous pretrained LMs.
During pretraining, a CSE layer efficiently clusters domain knowledge together and assigns long-tail knowledge to designate extra experts. CSE is also a lightweight structure that only needs to be incorporated in several deep layers.
With our training strategy, we found that during pretraining, data of long-tail knowledge gradually formulate isolated, outlier clusters in an LM's representation spaces, especially in deeper layers. Our experimental results show that only pretraining CSE-based LMs is enough to achieve superior performance than regularly pretrained-finetuned LMs on various downstream tasks, implying the prospects of domain-specific-pretraining-free language models. Mengyi Chen, Jixian Zhou, Yubin Shi, Yixuan Chen 0003, Mingzhi Dong, Yujiang Wang 0001, Dongsheng Li 0002, Rui Zhu 0006, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
NeurIPS | 11 |
| 2023 | LoRa Synchronized Energy-Efficient LPWANabstractThis work addresses the need for low-power wide-area networks (LPWANs) that remain energy efficient even when scaled to cover large geographic areas without power delivery infrastructure. LoRaWAN is a widely used LPWAN that requires a star (or star-of-stars) topology in which only leaf nodes are energy efficient enough to be powered by compact and inexpensive batteries. We describe the design of LoRa Synchronized Energy-Efficient LPWAN (SEEL), a novel protocol for multi-hop LoRa networks that allows battery-powered nodes to forward messages to difficult-to-access locations. SEEL performs link-quality-aware, dynamic formation of a tree topology network enabling nodes to deliver their packets to a central gateway while minimizing node energy consumption. We evaluate SEEL via a month-long deployment in an outdoor, real-world environment covering 9.0 km2and analyze its delivery rates and energy efficiency; we find SEEL nodes, on average, function 6.6× as long as always-on nodes would in the same deployment setting. We conduct a follow-up deployment of SEEL to compare its dynamic network formation protocol with a static network formation protocol; the SEEL network drops 16% fewer node-to-node packets than would a static topology network, even in near-ideal circumstances for a static topology network. Vidya Srinivas, Robert P. Dick |
ICC | 4 |
| 2023 | Spatial-Frequency Network for Segmentation of Remote Sensing ImagesabstractWe describe a deep learning system for satellite image segmentation. Our CNN model embeds contextual feature dependencies in both spatial and frequency domains. Its Spatial Weighting Module uses a multi-scale pooling layer to represent correlations at longer length scales in the spatial domain. Its Frequency Weighting Module uses frequency-domain information to better discriminate between object classes. Experimental results on the Potsdam dataset demonstrate that our model has a 1.9% higher average F1 accuracy than previous methods. Tony Zhang, Robert P. Dick |
ICIP | 2 |
| 2023 | Over-parameterized Model Optimization with Polyak-Łojasiewicz Condition
Yixuan Chen 0003, Yubin Shi, Mingzhi Dong, Dongsheng Li 0002, Yujiang Wang 0001, Robert P. Dick, Qin Lv, Fan Yang 0001, Ning Gu 0001, Li Shang 0002 |
ICLR | 7 |
| 2023 | Mechanistic Mode ConnectivityabstractWe study neural network loss landscapes through the lens of mode connectivity, the observation that minimizers of neural networks retrieved via training on a dataset are connected via simple paths of low loss. Specifically, we ask the following question: are minimizers that rely on different mechanisms for making their predictions connected via simple paths of low loss? We provide a definition of mechanistic similarity as shared invariances to input transformations and demonstrate that lack of linear connectivity between two models implies they use dissimilar mechanisms for making their predictions. Relevant to practice, this result helps us demonstrate that naive fine-tuning on a downstream dataset can fail to alter a model’s mechanisms, e.g., fine-tuning can fail to eliminate a model’s reliance on spurious attributes. Our analysis also motivates a method for targeted alteration of a model’s mechanisms, named connectivity-based fine-tuning (CBFT), which we analyze using several synthetic datasets for the task of reducing a model’s reliance on spurious attributes. Ekdeep Singh Lubana, Eric J. Bigelow, Robert P. Dick, David Krueger 0001, Hidenori Tanaka |
ICML | 3 |
| 2023 | Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskabstractModern generative models exhibit unprecedented capabilities to generate extremely realistic data. However, given the inherent compositionality of the real world, reliable use of these models in practical applications requires that they exhibit the capability to compose a novel set of concepts to generate outputs not seen in the training data set. Prior work demonstrates that recent diffusion models do exhibit intriguing compositional generalization abilities, but also fail unpredictably. Motivated by this, we perform a controlled study for understanding compositional generalization in conditional diffusion models in a synthetic setting, varying different attributes of the training data and measuring the model's ability to generate samples out-of-distribution. Our results show: (i) the order in which the ability to generate samples from a concept and compose them emerges is governed by the structure of the underlying data-generating process; (ii) performance on compositional tasks exhibits a sudden "emergence" due to multiplicative reliance on the performance of constituent tasks, partially explaining emergent phenomena seen in generative models; and (iii) composing concepts with lower frequency in the training data to generate out-of-distribution samples requires considerably more optimization steps compared to generating in-distribution samples. Overall, our study lays a foundation for understanding emergent capabilities and compositionality in generative models from a data-centric perspective. Maya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka |
NeurIPS | 3 |
| 2023 | Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized ModelsabstractDespite their prevalence in deep-learning communities, over-parameterized models convey high demands of computational costs for proper training. This work studies the fine-grained, modular-level learning dynamics of over-parameterized models to attain a more efficient and fruitful training strategy. Empirical evidence reveals that when scaling down into network modules, such as heads in self-attention models, we can observe varying learning patterns implicitly associated with each module's trainability. To describe such modular-level learning capabilities, we introduce a novel concept dubbed modular neural tangent kernel (mNTK), and we demonstrate that the quality of a module's learning is tightly associated with its mNTK's principal eigenvalue $\lambda_{\max}$. A large $\lambda_{\max}$ indicates that the module learns features with better convergence, while those miniature ones may impact generalization negatively. Inspired by the discovery, we propose a novel training strategy termed Modular Adaptive Training (MAT) to update those modules with their $\lambda_{\max}$ exceeding a dynamic threshold selectively, concentrating the model on learning common features and ignoring those inconsistent ones. Unlike most existing training schemes with a complete BP cycle across all network modules, MAT can significantly save computations by its partially-updating strategy and can further improve performance. Experiments show that MAT nearly halves the computational cost of model training and outperforms the accuracy of baselines. Yubin Shi, Yixuan Chen 0003, Mingzhi Dong, Dongsheng Li 0002, Yujiang Wang 0001, Robert P. Dick, Qin Lv, Fan Yang 0001, Tun Lu, Ning Gu 0001, Li Shang 0002 |
NeurIPS | 7 |
| 2022 | Image-Based Air Quality Forecasting Through Multi-Level AttentionabstractThe problem of air quality forecasting is important but also challenging because air quality is affected by a diverse set of complex factors. This paper describes the first image-based air quality forecasting model. It fuses a history of PM2.5measurements with colocated images. We construct a multi-level attention-based recurrent network that uses images and PM2.5data to represent variation over space and time. Experiments on Shanghai data show that our model improves PM2.5RMSE prediction accuracy by 15.8% and MAE by 10.9% compared to previous forecasting methods. In addition, we evaluate the impact of each model component via ablation studies. Tony Zhang, Robert P. Dick |
ICIP | 2 |
| 2022 | Recursive Disentanglement Network
Yixuan Chen 0003, Yubin Shi, Dongsheng Li 0002, Yujiang Wang 0001, Mingzhi Dong, Robert P. Dick, Qin Lv, Fan Yang 0001, Li Shang 0002 |
ICLR | 7 |
| 2022 | Orchestra: Unsupervised Federated Learning via Globally Consistent ClusteringabstractFederated learning is generally used in tasks where labels are readily available (e.g., next word prediction). Relaxing this constraint requires design of unsupervised learning techniques that can support desirable properties for federated training: robustness to statistical/systems heterogeneity, scalability with number of participants, and communication efficiency. Prior work on this topic has focused on directly extending centralized self-supervised learning techniques, which are not designed to have the properties listed above. To address this situation, we propose Orchestra, a novel unsupervised federated learning technique that exploits the federation’s hierarchy to orchestrate a distributed clustering task and enforce a globally consistent partitioning of clients’ data into discriminable clusters. We show the algorithmic pipeline in Orchestra guarantees good generalization performance under a linear probe, allowing it to outperform alternative techniques in a broad range of conditions, including variation in heterogeneity, number of clients, participation ratio, and local epochs. Ekdeep Singh Lubana, Chi Ian Tang, Fahim Kawsar, Robert P. Dick, Akhil Mathur |
ICML | 4 |
| 2022 | A Reinforcement-Learning-Based Energy-Efficient Framework for Multi-Task Video Analytics PipelineabstractDeep-learning-based video processing has yielded transformative results in recent years. However, the video analytics pipeline is energy-intensive due to high data rates and reliance on complex inference algorithms, which limits its adoption in energy-constrained applications. Motivated by the observation of high and variable spatial redundancy and temporal dynamics in video data streams, we design and evaluate an adaptive-resolution optimization framework to minimize the energy use of multi-task video analytics pipelines. Instead of heuristically tuning the input data resolution of individual tasks, our framework utilizes deep reinforcement learning to dynamically govern the input resolution and computation of the entire video analytics pipeline. By monitoring the impact of varying resolution on the quality of high-dimensional video analytics features, hence the accuracy of video analytics results, the proposed end-to-end optimization framework learns the best non-myopic policy for dynamically controlling the resolution of input video streams to globally optimize energy efficiency. Governed by reinforcement learning, optical flow is incorporated into the framework to minimize unnecessary spatio-temporal redundancy that leads to re-computation, while preserving accuracy. The proposed framework is applied to video instance segmentation which is one of the most challenging computer vision tasks, and achieves better energy efficiency than all baseline methods of similar accuracy on the YouTube-VIS dataset. Mingzhi Dong, Yujiang Wang 0001, Da Feng, Qin Lv, Robert P. Dick, Dongsheng Li 0002, Tun Lu, Ning Gu 0001 |
IEEE Trans. Multim. | 6 |
| 2021 | A Gradient Flow Framework For Analyzing Network Pruning
Ekdeep Singh Lubana, Robert P. Dick |
ICLR | 2 |
| 2021 | Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep LearningabstractInspired by BatchNorm, there has been an explosion of normalization layers in deep learning. Recent works have identified a multitude of beneficial properties in BatchNorm to explain its success. However, given the pursuit of alternative normalization layers, these properties need to be generalized so that any given layer's success/failure can be accurately predicted. In this work, we take a first step towards this goal by extending known properties of BatchNorm in randomly initialized deep neural networks (DNNs) to several recently proposed normalization layers. Our primary findings follow: (i) similar to BatchNorm, activations-based normalization layers can prevent exponential growth of activations in ResNets, but parametric techniques require explicit remedies; (ii) use of GroupNorm can ensure an informative forward propagation, with different samples being assigned dissimilar activations, but increasing group size results in increasingly indistinguishable activations for different samples, explaining slow convergence speed in models with LayerNorm; and (iii) small group sizes result in large gradient norm in earlier layers, hence explaining training instability issues in Instance Normalization and illustrating a speed-stability tradeoff in GroupNorm. Overall, our analysis reveals a unified set of mechanisms that underpin the success of normalization methods in deep learning, providing us with a compass to systematically explore the vast design space of DNN normalization layers. Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka |
NeurIPS | 2 |
| 2020 | Online Resource Management for Improving Reliability of Real-Time Systems on "Big-Little" Type MPSoCsabstractHeterogeneous multiprocessor systems on a chips (MPSoCs) consisting of cores with different performance/power characteristics are widely used in many real-time embedded systems, where both soft-error reliability and lifetime reliability are key concerns. Although existing efforts have investigated related problems, they either focus on one of the two reliability concerns or propose time-consuming scheduling algorithms that cannot adequately address runtime workload and environmental variations. This paper introduces an online framework which is adaptive to runtime variations and maximizes soft-error reliability while satisfying the lifetime reliability constraint for soft real-time systems executing on MPSoCs that are composed of high-performance cores and low-power (LP) cores. Based on each core's executing frequency and utilization, the framework performs workload migration between high-performance cores and LP cores to reduce power consumption and improve soft-error reliability. Experimental results based on different hardware platforms show that the proposed approach reduces the probability of failures due to soft errors by at least 17% and 50% on average compared to a number of representative existing approaches that satisfy the same lifetime reliability constraints. Yue Ma 0001, Junlong Zhou, Thidapat Chantem, Robert P. Dick, Shige Wang, Xiaobo Sharon Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Improving Reliability of Soft Real-Time Embedded Systems on Integrated CPU and GPU PlatformsabstractMultiprocessor systems on a chip consisting of integrated CPUs and GPUs are suitable platforms for real-time embedded applications requiring massively parallel processing. For such applications, lifetime reliability due to permanent faults and soft-error reliability due to transient faults are major concerns. Detailed execution profiling has revealed that a CUDA task's CPU execution time significantly increases if the task executes on a different core than the operating system (OS). Based on this observation, an extended task model is introduced to consider the execution time dependencies among tasks and the OS. A hybrid framework is proposed to improve soft-error reliability while satisfying a lifetime reliability constraint for soft real-time systems executing on integrated CPU and GPU platforms. This framework: 1) reduces the total utilization of cores and improves soft-error reliability via off-line task mapping; 2) achieves a higher lifetime reliability through task migration at run time; and 3) improves soft-error reliability by dynamically scaling frequencies of CPU and GPU cores. The experimental results show that the proposed framework leads to a system that can execute without soft errors for at least 4 days (4 times) and 6 days (6 times) longer, on average, than existing approaches. Yue Ma 0001, Junlong Zhou, Thidapat Chantem, Robert P. Dick, Shige Wang, Xiaobo Sharon Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Machine Foveation: An Application-Aware Compressive Sensing FrameworkabstractEmbedded vision applications generally face tight resource constraints. Biological vision systems are optimized to operate under similar conditions; they use highly heterogeneous sensing patterns to capture only the most valuable information within scenes. Our exploration of similar approaches in embedded systems has led to the design of Machine Foveation-a general-purpose, application-aware compressive sensing related framework that uses a cascaded network architecture integrating an autoencoder and application network to determine the importance of each pixel. The cascaded structure results in inherent regulation of the autoencoder network, forcing it to learn a representation that retains a given feature only if it is crucial to the overall application. The framework further uses scene awareness for reducing the number of bits necessary to represent the image data. This reduces sensed data at minimal or no decrease in task accuracy and reduces signal communication latency and corresponding energy consumption in embedded systems. For example, when evaluated on the Fashion-MNIST data set, channel bandwidth requirements are reduced by 77.37% and signal communication latency is reduced by 64.6%, with an accuracy loss of only 0.32%. Ekdeep Singh Lubana, Vinayak Aggarwal, Robert P. Dick |
DCC | 3 |
| 2019 | Minimalistic Image Signal Processing for Deep Learning ApplicationsabstractIn-sensor energy-efficient deep learning accelerators have the potential to enable the use of deep neural networks in embedded vision applications. However, their negative impact on accuracy has been severely underestimated. The inference pipeline used in prior in-sensor deep learning accelerators bypasses the image signal processor (ISP), thereby disrupting the conventional vision pipeline and undermining accuracy of machine learning algorithms trained on conventional, post-ISP datasets. For example, the detection accuracy of an off-the-shelf Faster RCNN algorithm in a vehicle detection scenario reduces by 60%. To make in-sensor accelerators practical, we describe energy-efficient operations that yield most of the benefits of an ISP and reduce covariate shift between the training (ISP processed images) and target (RAW images) distributions. For the vehicle detection problem, our approach improves accuracy by 25-60%. Relative to the conventional ISP pipeline, energy consumption and response time improve by 30% and 34%, respectively. Ekdeep Singh Lubana, Robert P. Dick, Vinayak Aggarwal, Pyari Mohan Pradhan |
ICIP | 2 |
| 2019 | Estimation of Multiple Atmospheric Pollutants Through Image AnalysisabstractMultiple atmospheric pollutants, such as PM2.5, PM10, and NO2, degrades air quality in many parts of the world. Fine-grained air pollution data can help combat the problem, but conventional monitoring stations are too expensive to support high spatial resolution; image-based estimates have the potential to improve spatial coverage. We estimate pollutant concentrations from images using the position-and color-dependent properties of scattering and absorption. We are the first to use images to estimate pollutant concentrations in systems with multiple pollutants. We achieve this by considering the differences in scattering and absorption spectra between different pollutants. Our system improves the accuracy of PM2.5, PM10, and NO2estimation by 22% for single-scene images in Beijing and Shanghai compared to the best existing image-based techniques. Tony Zhang, Robert P. Dick |
ICIP | 2 |
| 2019 | Multi-Group Encoder-Decoder Networks to Fuse Heterogeneous Data for Next-Day Air Quality PredictionabstractAccurate next-day air quality prediction is essential to enable warning and prevention measures for cities and individuals to cope with potential air pollution, such as vehicle restriction, factory shutdown, and limiting outdoor activities. The problem is challenging because air quality is affected by a diverse set of complex factors. There has been prior work on short-term (e.g., next 6 hours) prediction, however, there is limited research on modeling local weather influences or fusing heterogeneous data for next-day air quality prediction. This paper tackles this problem through three key contributions: (1) we leverage multi-source data, especially high-frequency grid-based weather data, to model air pollutant dynamics at station-level; (2) we add convolution operators on grid weather data to capture the impacts of various weather parameters on air pollutant variations; and (3) we automatically group (cross-domain) features based on their correlations, and propose multi-group Encoder-Decoder networks (MGED-Net) to effectively fuse multiple feature groups for next-day air quality prediction. The experiments with real-world data demonstrate the improved prediction performance of MGED-Net over state-of-the-art solutions (4.2% to 9.6% improvement in MAE and 9.2% to 16.4% improvement in RMSE). Qin Lv, Duanfeng Gao, Si Shen, Robert P. Dick, Michael Hannigan, Qi Liu 0052 |
IJCAI | 5 |
| 2018 | Improving reliability for real-time systems through dynamic recoveryabstractTechnology scaling has increased concerns about transient faults due to soft errors and permanent faults due to lifetime wear processes. Although researchers have investigated related problems, they have either considered only one of the two reliability concerns or presented simple recovery allocation algorithms that cannot effectively use available time slack to improve soft-error reliability. This paper introduces a framework for improving soft-error reliability while satisfying lifetime reliability and real-time constraints. We present a dynamic recovery allocation technique that guarantees to recover any failed task if the remaining slack is adequate. Based on this technique, we propose two scheduling algorithms for task sets with different characteristics to improve system-level soft-error reliability. Lifetime reliability requirements are satisfied by reducing core frequencies for appropriate tasks, thereby reducing wear due to temperature and thermal cycling. Simulation results show that the proposed framework reduces the probability of failure by at least 8% and 73% on average compared to existing approaches. Yue Ma 0001, Thidapat Chantem, Robert P. Dick, Xiaobo Sharon Hu |
DATE | 3 |
| 2018 | Digital Foveation: An Energy-Aware Machine Vision FrameworkabstractIn machine vision applications, imaging systems and analysis algorithms are generally interdependent and energy intensive. We describe a machine vision energy minimization framework in which imaging hardware and vision algorithms are co-designed and tightly integrated. Digital foveation is inspired by the human vision system, which uses a spatially varying sensing architecture to generate oculomotory feedback and capture a series of high-resolution images using the densely sampling fovea. A multiround process with bidirectional information flow between camera hardware and analysis software optimizes energy consumption while preserving accuracy. By using existing hardware mechanisms, namely, row / column skipping, random access via readout circuitry, and frame preservation, digital foveation adapts to the chosen analysis algorithm. It aims to transmit and process only the necessary parts of the scene under consideration. This framework is general across a wide range of embedded machine vision applications and enables large improvements in energy efficiency. When evaluated for an embedded license plate recognition vision application, it reduces system energy consumption by 81.3% with at most 0.65% reduction in accuracy. Ekdeep Singh Lubana, Robert P. Dick |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | An on-line framework for improving reliability of real-time systems on "big-little" type MPSoCsabstractHeterogeneous MPSoCs consisting of cores with different performance/power behaviors are widely used in many power-constrained real-time systems. Both soft-error reliability and lifetime reliability are key concerns in such systems. Although existing work have investigated related problems, they either focus on one of the two reliability concerns or propose complicated scheduling algorithms that cannot adequately address run-time workload and environment variations. This paper introduces an on-line heuristic to maximize soft-error reliability while satisfying a lifetime reliability constraint for soft real-time systems executed on MPSoCs composed of high-performance cores and low-power cores. Based on the run-time cores' frequencies and utilizations, the heuristic performs workload migration between the high-performance cores and low-power cores to achieve improved soft-error reliability. Experimental results from both a hardware platform and a simulator show that the proposed algorithm reduces the probability of faults by at least 30% compared to a number of representative existing approaches while satisfying the same lifetime reliability constraints. Yue Ma 0001, Thidapat Chantem, Robert P. Dick, Shige Wang, Xiaobo Sharon Hu |
DATE | 3 |
| 2017 | Gazelle: Energy-Efficient Wearable Analysis for RunningabstractRunning is one of the most popular sports with hundreds of millions of participants worldwide. Good running form is the key to fast, efficient, and injury-free running. Existing kinematic analysis technologies, such as high-speed camera systems, are expensive, difficult to operate, and exclusive to sports physiology laboratories and elite athletes. Miniature MEMS-based motion sensors enable portable high-precision kinematic analysis, but suffer from high energy consumption hence short battery lifetime, especially for continued online analysis for running. This paper presents Gazelle, a wearable online analysis system for running that is compact, lightweight, accurate, and highly energy efficient; intended for runners of all levels. To enable long-term maintenance-free mobile analysis for running, Sparse Adaptive Sensing (SAS) is proposed, which selectively identifies the best sampling points to maintain high accuracy while greatly reducing sensing and analysis energy overheads. Experimental results demonstrate 97.7 percent accuracy with 76.9 to 99 percent reduced energy consumption (83.6 percent average reduction under real-world testing)-a one-order-of-magnitude improvement over existing solutions. SAS enables > 200 days of continuous high-precision operation using only a coin-cell battery. Since 2014, Gazelle has been used by over 100 elite and recreational runners during daily training and at top-level races like the Kona Ironman World Championships and New York Marathon. Qi Liu 0052, James Williamson, Wyatt Mohrman, Qin Lv, Robert P. Dick |
IEEE Trans. Mob. Comput. | 6 |
| 2017 | Improving System-Level Lifetime Reliability of Multicore Soft Real-Time SystemsabstractThis paper studies the problem of maximizing multicore system lifetime reliability, an important design consideration for many real-time embedded systems. Existing work has investigated the problem, but has neglected important failure mechanisms. Furthermore, most existing algorithms are too slow for online use, and thus cannot address runtime workload and environment variations. This paper presents an online framework that maximizes system lifetime reliability through reliability-aware utilization control. It focuses on homogeneous multicore soft real-time systems. It selectively employs a comprehensive reliability estimation tool to deal with a variety of failure mechanisms at the system level. A model-predictive controller adjusts utilization by manipulating core frequencies, thereby reducing temperature, and an online heuristic adjusts the controller sampling window length to decrease the reliability effects of thermal cycling. Experiments with a real quad-core ARM processor and a simulator demonstrate that the proposed approach improves system mean time to failure by 50% on average and 141% in the best case, compared with existing techniques. Yue Ma 0001, Thidapat Chantem, Robert P. Dick, Xiaobo Sharon Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Data sensing and analysis: Challenges for wearablesabstractWearables are a leading category in the Internet of Things. Compared with mainstream mobile phones, wearables target one order of magnitude form factor reduction, and offer the potential of providing ubiquitous, personalized services to end users. Aggressive reduction in size imposes serious limits on battery capacity. Wearables are equipped with a range of sensors, such as accelerometers and gyroscopes. Most economical sensors were developed for mobile phones, with energy consumptions more appropriate for phones than for ultra-compact wearables. This article describes the energy challenges for wearable sensing technologies, with a primary focus on the most widely used wearable sensors: MEMS-based inertial measurement units. Using sports and fitness wearables as the pilot application, we analyze the energy characteristics of MEMS IMU data sensing, analysis, and wireless communication. We then discuss the technologies needed to solve the power and energy consumptions challenges for wearables. James Williamson, Qi Liu 0052, Fenglong Lu, Wyatt Mohrman, Robert P. Dick |
ASP-DAC | 6 |
| 2015 | Embedded system and application aware design of deregulated energy delivery systemsabstractVoltage regulation circuity can be removed from embedded systems, commonly saving 30% of the printed circuit board area. Most unregulated systems suffer from performance degradation or premature failure as battery voltage decreases. Researchers have described a technique called power deregulation, in which performance is preserved by activating processor cores as battery voltage decreases. Unfortunately, the operating lifespans of such systems may be limited due to mismatches between the battery discharge profiles and the voltage-performance characteristics of processors. This paper provides a technique for codesign of power deregulated systems in which battery, processor, and workload characteristics are jointly considered. This technique extends lifespans of deregulated embedded systems. We also describe new battery and system performance and power models, some of which are based on a recently commercialized silver-zinc battery technology that is well suited to power deregulated systems. These models support lifetime simulation of regulated and deregulated embedded systems. If it is possible to find an appropriate match between embedded system and battery characteristics, the proposed design technique results in lifetimes similar to those of regulated systems and eliminates the need for bulky power regulation components. Xuejing He, Robert P. Dick, Russ Joseph |
CASES | 2 |
| 2015 | Improving Lifetime of Multicore Soft Real-Time Systems through Global Utilization ControlabstractSystem lifetime reliability is an important design consideration for many real-time embedded systems. Increasing integrated circuit power density and the subsequent rise in chip temperature negatively impact the lifetime reliability of such systems. Although existing thermal-aware methods are effective in reducing temperature, they cannot increase, and may even hamper, the system lifetime reliability. The complicated relationship between temperature and system lifetime requires that reliability be considered explicitly during system design. This paper presents a reliability-aware utilization control framework for homogeneous multicore soft real-time systems. The framework employs a model predictive controller to increase the system lifetime by manipulating the utilization of real-time tasks. An online heuristic algorithm is introduced to adjust the controller's sampling window in order to reduce the effects of thermal cycling on reliability. Simulation results show that the proposed approach can improve the system mean time to failure by at least 43% and as much as 369% compared to existing techniques. Yue Ma 0001, Thidapat Chantem, Xiaobo Sharon Hu, Robert P. Dick |
ACM Great Lakes Symposium on VLSI | 4 |
| 2015 | Performance and Energy Consumption Analysis of a Delay-Tolerant Network for Censorship-Resistant CommunicationabstractDelay Tolerant Networks (DTNs) composed of commodity mobile devices have the potential to support communication applications resistant to blocking and censorship, as well as certain types of surveillance. We analyze the performance and energy consumption of such a network, and consider the impact of random and targeted denial-of-service and censorship attacks. To gather wireless connectivity traces for a DTN composed of human-carried commodity smartphones, we implemented and deployed a prototype DTN-based micro-blogging application, called 1am, in a college town. We analyzed the system during a time period with 111 users. Although the study provided detailed enough connectivity traces to enable analysis, message posting was too infrequent to draw strong conclusions based on user-initiated messages, alone. We therefore simulated more frequent message initiations and used measured connectivity traces to analyze message propagation. Using a flooding protocol, we found that with an adoption rate of 0.2% of a college town's student and faculty population, the median one-week delivery rate is 85% and the median delivery delay is 13 hours. We also found that the network delivery rate and delay are robust to denial-of service and censorship attacks eliminating more than half of the participants. Using a measurement-based energy model, we also found that the DTN system would use less than 10.0% of a typical smartphone's battery energy per day in a network of 2,500 users. David R. Bild, David Adrian, Gulshan Singh, Robert P. Dick, Dan S. Wallach, Z. Morley Mao |
MobiHoc | 5 |
| 2015 | The Mason Test: A Defense Against Sybil Attacks in Wireless Networks Without Trusted AuthoritiesabstractWireless networks are vulnerable to Sybil attacks, in which a malicious node poses as many identities in order to gain disproportionate influence. Many defenses based on spatial variability of wireless channels exist, but depend either on detailed, multi-tap channel estimation-something not exposed on commodity 802.11 devices-or valid RSSI observations from multiple trusted sources, e.g., corporate access points-something not directly available in ad hoc and delay-tolerant networks with potentially malicious neighbors. We extend these techniques to be practical for wireless ad hoc networks of commodity 802.11 devices. Specifically, we propose two efficient methods for separating the valid RSSI observations of behaving nodes from those falsified by malicious participants. Further, we note that prior signalprint methods are easily defeated by mobile attackers and develop an appropriate challenge-response defense. Finally, we present the Mason test, the first implementation of these techniques for ad hoc and delay-tolerant networks of commodity 802.11 devices. We illustrate its performance in several real-world scenarios. David R. Bild, Robert P. Dick, Z. Morley Mao, Dan S. Wallach |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | Aggregate Characterization of User Behavior in Twitter and Analysis of the Retweet GraphabstractMost previous analysis of Twitter user behavior has focused on individual information cascades and the social followers graph, in which the nodes for two users are connected if one follows the other. We instead study aggregate user behavior and the retweet graph with a focus on quantitative descriptions. We find that the lifetime tweet distribution is a type-II discrete Weibull stemming from a power law hazard function, that the tweet rate distribution, although asymptotically power law, exhibits a lognormal cutoff over finite sample intervals, and that the inter-tweet interval distribution is a power law with exponential cutoff. The retweet graph is small-world and scale-free, like the social graph, but less disassortative and has much stronger clustering. These differences are consistent with it better capturing the real-world social relationships of and trust between users than the social graph. Beyond just understanding and modeling human communication patterns and social networks, applications for alternative, decentralized microblogging systems---both predicting real-word performance and detecting spam---are discussed. David R. Bild, Robert P. Dick, Z. Morley Mao, Dan S. Wallach |
ACM Trans. Internet Techn. | 3 |
| 2014 | Efficient location aware intrusion detection to protect mobile devicesabstractThis paper addresses the problem of efficient intrusion detection for mobile devices via correlating the user’s location and time data. We developed two statistical profiling approaches for modeling the normal spatio–temporal behavior of the users: one based on an empirical cumulative probability measure and the other based on the Markov properties of trajectories. An anomaly is detected when the probability of a particular (location, time) evolution matching the normal behavior of a given user becomes lower than a certain threshold, determined by controlling the recall rate of the model of the normal user’s behavior. We used compression techniques to reduce processing overhead while maintaining high accuracy. Our evaluation based on the Reality Mining and Geolife data sets shows that the proposed system is capable of detecting a potential intrusion within 15 min and with 94 % accuracy. Sausan Yazji, Peter Scheuermann, Robert P. Dick, Goce Trajcevski, Ruoming Jin |
Pers. Ubiquitous Comput. | 3 |
| 2013 | Enhancing multicore reliability through wear compensation in online assignment and schedulingabstractSystem reliability is a crucial concern especially in multicore systems which tend to have high power density and hence temperature. Existing reliability-aware methods are either slow and non-adaptive (offline techniques) or do not use task assignment and scheduling to compensate for uneven core wear states (online techniques). In this article, we present a dynamically-activated task assignment and scheduling algorithm based on theoretical results that explicitly optimizes system life-time. We also propose a data distillation method that dramatically reduces the size of the thermal profiles to make full system reliability analysis viable online. Simulation results show that our algorithm results in between 27–291% improvement to system lifetime compared to existing techniques for four-core systems. Thidapat Chantem, Xiang Yun, Xiaobo Sharon Hu, Robert P. Dick |
DATE | 4 |
| 2013 | A Hybrid Sensor System for Indoor Air Quality MonitoringabstractIndoor air quality is important. It influences human productivity and health. Personal pollution exposure can be measured using stationary or mobile sensor networks, but each of these approaches has drawbacks. Stationary sensor network accuracy suffers because it is difficult to place a sensor in every location people might visit. In mobile sensor networks, accuracy and drift resistance are generally sacrificed for the sake of mobility and economy. We propose a hybrid sensor network architecture, which contains both stationary sensors (for accurate readings and calibration) and mobile sensors (for coverage). Our technique uses indoor pollutant concentration prediction models to determine the structure of the hybrid sensor network. In this work, we have (1) developed a predictive model for pollutant concentration that minimizes prediction error; (2) developed algorithms for hybrid sensor network construction; and (3) deployed a sensor network to gather data on the airflow in a building, which are later used to evaluate the prediction model and hybrid sensor network synthesis algorithm. Our modeling technique reduces sensor network error by 40.4% on average relative to a technique that does not explicitly consider the inaccuracies of individual sensors. Our hybrid sensor network synthesis technique improves personal exposure measurement accuracy by 35.8% on average compared with a stationary sensor network architecture. Xiang Yun, Ricardo Piedrahita, Robert P. Dick, Michael Hannigan, Qin Lv |
DCOSS | 3 |
| 2013 | Hallway based automatic indoor floorplan construction using room fingerprintsabstractPeople spend approximately 70% of their time indoors. Understanding the indoor environments is therefore important for a wide range of emerging mobile personal and social applications. Knowledge of indoor floorplans is often required by these applications. However, indoor floorplans are either unavailable or obtaining them requires slow, tedious, and error-prone manual labor. Yifei Jiang, Xiang Yun, Qin Lv, Robert P. Dick, Michael Hannigan |
UbiComp | 6 |
| 2013 | Personalized multi-modality image management and search for mobile devices
Changyun Zhu, Qin Lv, Robert P. Dick |
Pers. Ubiquitous Comput. | 5 |
| 2013 | HAPPE: Human and Application-Driven Frequency Scaling for Processor Power EfficiencyabstractConventional dynamic voltage and frequency scaling techniques use high CPU utilization as a predictor for user dissatisfaction, to which they react by increasing CPU frequency. In this paper, we demonstrate that for many interactive applications, perceived performance is highly dependent upon the particular user and application, and is not linearly related to CPU utilization. This observation reveals an opportunity for reducing power consumption. We propose Human and Application driven frequency scaling for Processor Power Efficiency (HAPPE), an adaptive user-and-application-aware dynamic CPU frequency scaling technique. HAPPE continuously adapts processor frequency and voltage to the learned performance requirement of the current user and application. Adaptation to user requirements is quick and requires minimal effort from the user (typically a handful of key strokes). Once the system has adapted to the user's performance requirements, the user is not required to provide continued feedback but is permitted to provide additional feedback to adjust the control policy to changes in preferences. HAPPE was implemented on a Linux-based laptop and evaluated in 22 hours of controlled user studies. Compared to the default Linux CPU frequency controller, HAPPE reduces the measured system-wide power consumption of CPU-intensive interactive applications by 25 percent on average while maintaining user satisfaction. Lei Yang 0017, Robert P. Dick, Gokhan Memik, Peter A. Dinda |
IEEE Trans. Mob. Comput. | 2 |
| 2012 | ARIEL: automatic wi-fi based room fingerprinting for indoor localizationabstractPeople spend the majority of their time indoors, and human indoor activities are strongly correlated with the rooms they are in. Room localization, which identifies the room a person or mobile phone is in, provides a powerful tool for characterizing human indoor activities and helping address challenges in public health, productivity, building management, etc. Existing room localization methods, however, require labor-intensive manual annotation of individual rooms. Yifei Jiang, Qin Lv, Robert P. Dick, Michael Hannigan |
UbiComp | 5 |
| 2012 | Collaborative calibration and sensor placement for mobile sensor networksabstractMobile sensing systems carried by individuals or machines make it possible to measure position- and time-dependent environmental conditions, such as air quality and radiation. The low-cost, miniature sensors commonly used in these systems are prone to measurement drift, requiring occasional re-calibration to provide accurate data. Requiring end users to periodically do manual calibration work would make many mobile sensing systems impractical. We therefore argue for the use of collaborative, automatic calibration among nearby mobile sensors, and provide solutions to the drift estimation and placement problems posed by such a system. Xiang Yun, Lan S. Bai, Ricardo Piedrahita, Robert P. Dick, Qin Lv, Michael Hannigan |
IPSN | 4 |
| 2012 | Understanding the impact of laptop power saving options on user satisfaction using physiological sensorsabstractSeveral techniques are available to save power consumption in laptop computers. However, their effect on user satisfaction has not been well studied. We analyze how user satisfaction is affected by these techniques and show that, within a fixed power budget, some techniques cause more dissatisfaction than others. Second, we study the use of physiological sensors and show that the sensor readings are stable across times when no technique is applied, whereas they show statistically significant changes when power-saving techniques are employed. Finally, we demonstrate a prediction mechanism using these sensors that predicts user satisfaction with over 80% accuracy. Matthew Schuchhardt, Benjamin Scholbrock, Utku Pamuksuz, Gokhan Memik, Peter A. Dinda, Robert P. Dick |
ISLPED | 6 |
| 2012 | Introduction to special section SCPS'09abstractNo abstract available. Robert P. Dick, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2012 | Static NBTI Reduction Using Internal Node ControlabstractNegative Bias Temperature Instability (NBTI) is a significant reliability concern for nanoscale CMOS circuits. Its effects on circuit timing can be especially pronounced for circuits with standby-mode equipped functional units, because these units can be subjected to static NBTI stress for extended periods of time. This article describes Internal Node Control (INC), in which the inputs to some individual gates are directly manipulated to prevent this static NBTI fatigue. We prove that the INC selection problem is NP -complete and present a linear-time heuristic that can quickly determine near-optimal placements. This near-optimality is confirmed by comparing results for small benchmarks against optimal solutions from a mixed integer linear programming formulation of our problem. We evaluate the heuristic on the ISCAS85 benchmarks and the Synopsys DesignWare Library. Our heuristic reduces static NBTI-induced delay over a ten year period by 30--60% and can reduce total path delay by an average 9.4% when NBTI degradation is severe. The INC placements and sleep signal routing require only a 1.6% increase in area. David R. Bild, Robert P. Dick, Gregory E. Bok |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2011 | Automated construction of fast and accurate system-level models for wireless sensor networksabstractRapidly and accurately estimating the impact of design decisions on performance metrics is critical to both the manual and automated design of wireless sensor networks. Estimating system-level performance metrics such as lifetime, data loss rate, and network connectivity is particularly challenging because they depend on many factors, including network design and structure, hardware characteristics, communication protocols, and node reliability. This paper describes a new method for automatically building efficient and accurate predictive models for a wide range of system-level performance metrics. These models can be used to eliminate or reduce the need for simulation during design space exploration. We evaluate our method by building a model for the lifetime of networks containing up to 120 nodes, considering both fault processes and battery energy depletion. With our adaptive sampling technique, only 0.27% of the potential solutions are evaluated via simulation. Notably, one such automatically produced model outperforms the most advanced manually designed analytical model, reducing error by 13% while maintaining very low model evaluation overhead. We also propose a new, more general definition of system lifetime that accurately captures application requirements and decouples the specification of requirements from implementation decisions. Lan S. Bai, Robert P. Dick, Pai H. Chou, Peter A. Dinda |
DATE | 2 |
| 2011 | Simplified programming of faulty sensor networks via code transformation and run-time interval computationabstractDetecting and reacting to faults is an indispensable capability for many wireless sensor network applications. Unfortunately, implementing fault detection and error correction algorithms is challenging. Programming languages and fault tolerance mechanisms for sensor networks have historically been designed in isolation. This is the first work to combine them. Our goal is to simplify the design of fault-tolerant sensor networks. We describe a system that makes it unnecessary for sensor network application developers and users to understand the intricate implementation details of fault detection and tolerance techniques, while still using their domain knowledge to support fault detection, error correction, and error estimation mechanisms. Our FACTS system translates low-level faults into their consequences for application-level data quality, i.e., consequences domain experts can appreciate and understand. FACTS is an extension of an existing sensor network programming language; its compiler and runtime libraries have been modified to support automatic generation of code for on-line fault detection and tolerance. This code determines the impacts of faults on the accuracies of the results of potentially complex data aggregation and analysis expressions. We evaluate the overhead of the proposed system on code size, memory use, and the accuracy improvements for data analysis expressions using a small experimental testbed and simulations of large-scale networks. Lan S. Bai, Robert P. Dick, Peter A. Dinda, Pai H. Chou |
DATE | 2 |
| 2011 | Integrated circuit white space redistribution for temperature optimizationabstractThermal problems are important for integrated circuits with high power densities. Three-dimensional stacked-wafer integrated circuit technology reduces interconnect lengths and improves performance compared to two-dimensional integration. However, it intensifies thermal problems. One remedy is to redistribute white space during floorplanning. In this paper, we propose a two-phase algorithm to redistribute white space. In the first phase, the lateral heat flow white space redistribution problem is formulated as a minimum cycle ratio problem, in which the maximum power density is minimized. Since this phase only considers lateral heat flow, it also works for traditional two-dimensional integrated circuits. In the second phase, to consider inter-layer heat flow in three-dimensional integrated circuits, we discretize the chip into an array of tiles and use a dynamic programming algorithm to minimize the maximum stacked tile power consumption. We compared our algorithms with a previously proposed technique based on mathematical programming. Our iterative minimum cycle ratio algorithm achieves 35% more reduction in peak temperature. Our two-phase algorithm achieves 4.21× reduction in peak temperature for three-dimensional integrated circuits compared to applying the first phase, alone. Yuankai Chen, Hai Zhou 0001, Robert P. Dick |
DATE | 3 |
| 2011 | MAQS: a personalized mobile sensing system for indoor air quality monitoringabstractMost people spend more than 90% of their time indoors; indoor air quality (IAQ) influences human health, safety, productivity, and comfort. This paper describes MAQS, a personalized mobile sensing system for IAQ monitoring. In contrast with existing stationary or outdoor air quality sensing systems, MAQS users carry portable, indoor location tracking sensors that provide personalized IAQ information. To improve accuracy and energy efficiency, MAQS incorporates three novel techniques: (1) an accurate temporal n-gram augmented Bayesian room localization method that requires few Wi-Fi fingerprints; (2) an air exchange rate based IAQ sensing method, which measures general IAQ using only CO2 sensors; and (3) a zone-based proximity detection method for collaborative sensing, which saves energy and enables data sharing among users. MAQS has been deployed and evaluated via user study. Detailed evaluation results demonstrate that MAQS supports accurate personalized IAQ monitoring and quantitative analysis with high energy efficiency. Yifei Jiang, Lei Tian 0004, Ricardo Piedrahita, Xiang Yun, Omkar Mansata, Qin Lv, Robert P. Dick, Michael Hannigan |
UbiComp | 8 |
| 2011 | MAQS: a mobile sensing system for indoor air qualityabstractMost people spend more than 90% of their time indoors. Indoor air quality (IAQ) influences human health, safety, productivity, and comfort. This demo introduces MAQS, a personalized mobile sensing system for IAQ monitoring. In contrast with existing stationary or outdoor air quality sensing systems, MAQS users carry portable, indoor location tracking sensors that provide personalized IAQ information. To improve accuracy and energy efficiency, MAQS incorporates three novel techniques: (1) an accurate temporal n-gram augmented Bayesian room localization method; (2) an air exchange rate based IAQ sensing method; and (3) a zone-based proximity detection method for collaborative sensing. Yifei Jiang, Lei Tian 0004, Ricardo Piedrahita, Xiang Yun, Omkar Mansata, Qin Lv, Robert P. Dick, Michael Hannigan |
UbiComp | 8 |
| 2011 | Indoor localization without infrastructure using the acoustic background spectrumabstractWe introduce a new technique for determining a mobile phone's indoor location even when Wi-Fi infrastructure is unavailable or sparse. Our technique is based on a new ambient sound fingerprint called the Acoustic Background Spectrum (ABS). An ABS serves well as a room fingerprint because it is compact, easily computed, robust to transient sounds, and surprisingly distinctive. As with other fingerprint-based localization techniques, location is determined by measuring the current fingerprint and then choosing the "closest" fingerprint from a database. An experiment involving 33 rooms yielded 69% correct fingerprint matches meaning that, in the majority of observations, the fingerprint was closer to a previous visit's fingerprint than to any fingerprints from the other 32 rooms. An implementation of ABS-localization called Batphone is publicly available for Apple iPhones. We used Batphone to show the benefit of using ABS-localization together with a commercial Wi-Fi-based localization method. In this second experiment, adding ABS improved room-level localization accuracy from 30% (Wi-Fi only) to 69% (Wi-Fi and ABS). While Wi-Fi-based localization has difficulty distinguishing nearby rooms, Batphone performs just as well with nearby rooms; it can distinguish pairs of adjacent rooms with 92% accuracy. Stephen P. Tarzia, Peter A. Dinda, Robert P. Dick, Gokhan Memik |
MobiSys | 3 |
| 2011 | Demo: indoor localization without infrastructure using the acoustic background spectrumabstractWe demonstrate an indoor localization technique to be presented as a full paper at this MobiSys conference. In that paper, we introduce a new technique for determining mobile phone's indoor location even when Wi-Fi infrastructure is unavailable or sparse. Our technique is based on a new ambient sound fingerprint called the Acoustic Background Spectrum (ABS). This demonstration has two components. First, it shows attendees a live view of the ABS in the demonstration hall. This allows attendees to test the claim that the ABS is stable and robust to transient noise. Attendees can speak or make noise and observe the consequent effect on ABS. In the second demonstration, short sound recordings from various rooms are played on a set of headphones. Simultaneously, a photograph of the room and a plot of the ABS is shown. This demonstrates the ambient sound variations present in a set of sample rooms and allows attendees to test their own ability to distinguish locations based on sound. Stephen P. Tarzia, Peter A. Dinda, Robert P. Dick, Gokhan Memik |
MobiSys | 3 |
| 2011 | Temperature-Aware Scheduling and Assignment for Hard Real-Time Applications on MPSoCsabstractIncreasing integrated circuit (IC) power densities and temperatures may hamper multiprocessor system-on-chip (MPSoC) use in hard real-time systems. This paper formalizes the temperature-aware real-time MPSoC assignment and scheduling problem and presents an optimal phased steady-state mixed integer linear programming-based solution that considers the impact of scheduling and assignment decisions on MPSoC thermal profiles to directly minimize the chip peak temperature. We also introduce a flexible heuristic framework for task assignment and scheduling that permits system designers to trade off accuracy for running time when solving large problem instances. Finally, for task sets with sufficient slack, we show that inserting idle times between task executions can further reduce the peak temperature of the MPSoC quite significantly. Thidapat Chantem, Xiaobo Sharon Hu, Robert P. Dick |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Full-Spectrum Spatial-Temporal Dynamic Thermal Analysis for Nanometer-Scale Integrated CircuitsabstractThis paper presents NanoHeat, a multi-resolution full-chip dynamic integrated circuit (IC) thermal analysis solution, that is accurate down to the scale of individual gates and transistors. NanoHeat unifies nanoscale and macroscale dynamic thermal physics models, for accurate characterization of heat transport from the gate and transistor level up to the chip-package level. A non-homogeneous Arnoldi-based analysis method is proposed for accurate and fast dynamic thermal analysis through a unified adaptive spatial-temporal refinement process. NanoHeat is capable of covering the complete spatial and temporal modeling spectrum of IC thermal analysis. The accuracy and efficiency of NanoHeat are evaluated, and NanoHeat has been applied to a large industry design. The importance of considering fine-grain temperature information is illustrated by using NanoHeat to estimate temperature-dependent negative-bias-temperature-instability (NBTI) effects. NanoHeat has been implemented and publicly released for free academic and personal use. Zyad Hassan, Nicholas Allec, Fan Yang 0001, Robert P. Dick, Xuan Zeng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2010 | Performance and power modeling in a multi-programmed multi-core environmentabstractThis paper describes a fast, automated technique for accurate on-line estimation of the performance and power consumption of interacting processes in a multi-programmed, multi-core environment. The proposed technique does not require modifying hardware or applications. The performance model uses reuse distance histograms, cache access frequencies, and the relationship between the throughput and cache miss rate of each process to predict throughput. The system-level power model is derived using multi-variable linear regression, accounting for cache contention. Both models are validated on multiple real multi-core systems using SPEC CPU2000 benchmarks; their performance and power estimates are within 3.5% of measured values on average. We explain how to integrate the two models for power estimation during process assignment, helpful for power-aware assignment. Xi Chen 0068, Robert P. Dick, Z. Morley Mao |
DAC | 3 |
| 2010 | Properties of and improvements to time-domain dynamic thermal analysis algorithmsabstractTemperature has a strong influence on integrated circuit (IC) performance, power consumption, and reliability. However, accurate thermal analysis can impose high computation costs during the IC design process. We analyze the performance and accuracies of a variety of time-domain dynamic thermal analysis techniques and use our findings to propose a new analysis technique that improves performance by 38-138× relative to popular methods such as the fourth-order globally adaptive Runge-Kutta method while maintaining accuracy. More precisely, we prove that the step sizes of step doubling based globally adaptive fourth-order Runge-Kutta method and Runge-Kutta-Fehlberg methods always converge to a constant value regardless of the initial power profile, thermal profile, and error threshold during dynamic thermal analysis. Thus, these widely-used techniques are unable to adapt to the requirements of individual problems, resulting in poor performance. We also determine the effect of using a number of temperature update functions and step size adaptation methods for dynamic thermal analysis, and identify the most promising approach considered. Based on these observations, we propose FATA, a temporally-adaptive technique for fast and accurate dynamic thermal analysis. Xi Chen 0068, Robert P. Dick |
DATE | 2 |
| 2010 | Memory access aware on-line voltage control for performance and energy optimizationabstractThis paper describes an off-chip memory access-aware runtime DVFS control technique that minimizes energy consumption subject to constraints on application execution times. We consider application phases and the implications of changing cache miss rates on the ideal power control state. We first propose a two-stage DVFS algorithm based on formulating the throughput-constrained energy minimization problem as a multiple-choice knapsack problem (MCKP). This algorithm uses a power model that adapts to application phase changes by observing processor hardware performance counter values. The solutions it produces provide upper bounds on the energy savings achievable under a performance constraint. However, this algorithm assumes a priori (oracle or profiling-based) knowledge of application phase change behavior. To relax this assumption, we propose P-DVFS, an predictive DVFS algorithm for on-line minimization of energy consumption under a performance constraint without requiring a priori knowledge of an application's behavior. P-DVFS uses hardware performance counter based performance and power models. It predicts remaining execution time online in order to control voltage and frequency settings to optimize energy consumption and performance. The P-DVFS problem is formulated as a multiple-choice knapsack problem, which can be efficiently and optimally solved online. We evaluated P-DVFS using direct measurement of a real DVFS-equipped system. When bounding performance loss to at most 20% of that at the maximum frequency and voltage, P-DVFS leads to energy consumptions within 1.83% of the optimal solution for our problem instances on average with a maximum deviation of 4.83%. In addition to producing results approaching those of an oracle formulation, P-DVFS reduces power consumption for our problem instances by 9.93% on average, and up to 25.64%, compared with the most advanced related work. Xi Chen 0068, Robert P. Dick |
ICCAD | 3 |
| 2010 | Reliability, thermal, and power modeling and optimizationabstractThis tutorial provides an overview of challenges to designing and implementing reliable integrated circuits and systems, and suggests areas for future study. It illustrates some concepts in detail, explaining the challenges of appropriately considering the impact of temperature on reliability in fault-tolerant systems. Finally, it points out considerations that may influence adoption of reliability modeling and optimization techniques and stresses the importance of considering the most relevant fault processes during reliability modeling and optimization. Robert P. Dick |
ICCAD | 1 |
| 2010 | Cache contention and application performance prediction for multi-core systemsabstractThe ongoing move to chip multiprocessors (CMPs) permits greater sharing of last-level cache by processor cores but this sharing aggravates the cache contention problem, potentially undermining performance improvements. Accurately modeling the impact of inter-process cache contention on performance and power consumption is required for optimized process assignment. However, techniques based on exhaustive consideration of process-to-processor mappings and cycle-accurate simulation are inefficient or intractable for CMPs, which often permit a large number of potential assignments. This paper proposes CAMP, a fast and accurate shared cache aware performance model for multi-core processors. CAMP estimates the performance degradation due to cache contention of processes running on CMPs. It uses reuse distance histograms, cache access frequencies, and the relationship between the throughput and cache miss rate of each process to predict its effective cache size when running concurrently and sharing cache with other processes, allowing instruction throughput estimation.We also provide an automated way to obtain process-dependent characteristics, such as reuse distance histograms, without offline simulation, operating system (OS) modification, or additional hardware. We tested the accuracy of CAMP using 55 different combinations of 10 SPEC CPU2000 benchmarks on a dual-core CMP machine. The average throughput prediction error was 1.57%. Xi Chen 0068, Robert P. Dick, Z. Morley Mao |
ISPASS | 3 |
| 2010 | Efficient Intrusion Detection for Mobile Devices Using Spatio-temporal Mobility Patterns
Sausan Yazji, Robert P. Dick, Peter Scheuermann, Goce Trajcevski |
MobiQuitous | 2 |
| 2010 | Online memory compression for embedded systemsabstractMemory is a scarce resource during embedded system design. Increasing memory often increases packaging costs, cooling costs, size, and power consumption. This article presents CRAMES, a novel and efficient software-based RAM compression technique for embedded systems. The goal of CRAMES is to dramatically increase effective memory capacity without hardware or application design changes, while maintaining high performance and low energy consumption. To achieve this goal, CRAMES takes advantage of an operating system's virtual memory infrastructure by storing swapped-out pages in compressed format. It dynamically adjusts the size of the compressed RAM area, protecting applications capable of running without it from performance or energy consumption penalties. In addition to compressing working data sets, CRAMES also enables efficient in-RAM filesystem compression, thereby further increasing RAM capacity. CRAMES was implemented as a loadable module for the Linux kernel and evaluated on a battery-powered embedded system. Experimental results indicate that CRAMES is capable of doubling the amount of RAM available to applications running on the original system hardware. Execution time and energy consumption for a broad range of examples are rarely affected. When physical RAM is reduced to 62.5% of its original quantity, CRAMES enables the target embedded system to support the same applications with reasonable performance and energy consumption penalties (on average 9.5% and 10.5%), while without CRAMES those applications either may not execute or suffer from extreme performance degradation or instability. In addition to presenting a novel framework for dynamic data memory compression and in-RAM filesystem compression in embedded systems, this work identifies the software-based compression algorithms that are most appropriate for use in low-power embedded systems. Lei Yang 0017, Robert P. Dick, Haris Lekatsas, Srimat T. Chakradhar |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2010 | High-performance operating system controlled online memory compressionabstractOnline memory compression is a technology that increases the amount of memory available to applications by dynamically compressing and decompressing their working datasets on demand. It has proven extremely useful in embedded systems with tight physical RAM constraints. The technology can be used to increase functionality, reduce size, and reduce cost, without modifying applications or hardware. This article presents a new software-based online memory compression algorithm for embedded systems. In comparison with the best algorithms used in online memory compression, our new algorithm has a competitive compression ratio but is twice as fast. In addition, we describe several practical problems encountered in developing an online memory compression infrastructure and present solutions. We present a method of adaptively managing the uncompressed and compressed memory regions during application execution. This memory management scheme adapts to the predicted memory requirements of applications. It permits efficient compression for a wide range of applications. We have evaluated our techniques on a portable embedded device and have found that the memory available to applications can be increased by 2.5× with negligible performance and power consumption penalties, and with no changes to hardware or applications. Our techniques allow existing applications to execute with less physical memory. They also allow applications with larger working datasets to execute on unchanged embedded system hardware, thereby increasing functionality. Lei Yang 0017, Robert P. Dick, Haris Lekatsas, Srimat T. Chakradhar |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2010 | C-Pack: A High-Performance Microprocessor Cache Compression AlgorithmabstractMicroprocessor designers have been torn between tight constraints on the amount of on-chip cache memory and the high latency of off-chip memory, such as dynamic random access memory. Accessing off-chip memory generally takes an order of magnitude more time than accessing on-chip cache, and two orders of magnitude more time than executing an instruction. Computer systems and microarchitecture researchers have proposed using hardware data compression units within the memory hierarchies of microprocessors in order to improve performance, energy efficiency, and functionality. However, most past work, and all work on cache compression, has made unsubstantiated assumptions about the performance, power consumption, and area overheads of the proposed compression algorithms and hardware. It is not possible to determine whether compression at levels of the memory hierarchy closest to the processor is beneficial without understanding its costs. Furthermore, as we show in this paper, raw compression ratio is not always the most important metric. In this work, we present a lossless compression algorithm that has been designed for fast on-line data compression, and cache compression in particular. The algorithm has a number of novel features tailored for this application, including combining pairs of compressed lines into one cache line and allowing parallel compression of multiple words while using a single dictionary and without degradation in compression ratio. We reduced the proposed algorithm to a register transfer level hardware design, permitting performance, power consumption, and area estimation. Experiments comparing our work to previous work are described. Xi Chen 0068, Lei Yang 0017, Robert P. Dick, Haris Lekatsas |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Scheduled voltage scaling for increasing lifetime in the presence of NBTIabstractNegative Bias Temperature Instability (NBTI) is a leading reliability concern for integrated circuits (ICs). It gradually increases the threshold voltages of PMOS transistors, thereby increasing delay. We propose scheduled voltage scaling, a technique that gradually increases the operating voltage of the IC to compensate for NBTI-related performance degradation. Scheduled voltage scaling has the potential to increase IC lifetime by 46% relative to the conventional approach using guard banding for ICs fabricated using a 45 nm process. Lide Zhang, Robert P. Dick |
ASP-DAC | 2 |
| 2009 | Process variation characterization of chip-level multiprocessorsabstractWithin-die variation in leakage power consumption is substantial and increasing for chip-level multiprocessors (CMPs) and multiprocessor systems-on-chip. Dealing with this problem via conservative assumptions is sub-optimal. Instead, operating systems may adapt task assignment and power management decisions to the variable characteristics of cores, improving system-wide power consumption and performance. Researchers have proposed such adaptation techniques. However, they rely on knowledge of CMP process variation (PV) maps. These maps are not provided by processor vendors, providing them would impose additional cost during the testing process, and static maps would not permit adaptation to aging effects. Further progress on developing and validating PV aware control techniques for CMPs requires access to PV maps for real processors. We present an online technique to extract the PV maps of CMPs. Potentially automatic temperature measurements with built-in on-die sensors during the execution of characterization workloads are used to determine variation in leakage power consumption. The proposed technique is applied to real CMPs, and the resulting PV maps are used within a PV aware task assignment and scheduling algorithm. Lide Zhang, Lan S. Bai, Robert P. Dick, Russ Joseph |
DAC | 3 |
| 2009 | Minimization of NBTI performance degradation using internal node controlabstractNegative Bias Temperature Instability (NBTI) is a significant reliability concern for nanoscale CMOS circuits. Its effects on circuit timing can be especially pronounced for circuits with standby-mode equipped functional units because these units can be subjected to static NBTI stress for extended periods of time. This paper proposes internal node control, in which the inputs to individual gates are directly manipulated to prevent this static NBTI fatigue. We give a mixed integer linear program formulation for an optimal solution to this problem. The optimal placement of internal node control yields an average 26.7% reduction in NBTI-induced delay over a ten year period for the ISCAS85 benchmarks. We find that the problem is NP-complete and present a linear-time heuristic that can be used to quickly find near-optimal solutions. The heuristic solutions are, on average, within 0.17% of optimal and all were within 0.60% of optimal. David R. Bild, Gregory E. Bok, Robert P. Dick |
DATE | 3 |
| 2009 | Latency criticality aware on-chip communicationabstractPacket-switched interconnect fabric is a promising on-chip communication solution for many-core architectures. It offers high throughput and excellent scalability for on-chip data and protocol transactions. The main problem posed by this communication fabric is the potentially-high and nondeterministic network latency caused by router data buffering and resource arbitration. This paper describes a new method to minimize on-chip network latency, which is motivated by the observation that only a small percentage of on-chip data and protocol traffic is latency-critical. Existing work focusing on minimizing average network latency is thus suboptimal. Such techniques expend most of the design, area, and power overhead accelerating latency-noncritical traffic for which there is no corresponding application-level speedup. We propose run-time techniques that identify latency-critical traffic by leveraging network data-transaction and protocol information. Latency-critical traffic is permitted to bypass router pipeline stages and latency-noncritical traffic. These techniques are evaluated via a router design that has been implemented using TSMC 65nm technology. Detailed network latency simulation and hardware characterization demonstrate that, for latency-critical traffic, the proposed solution closely approximates the ideal interconnect even under heavy load while preserving throughput for both latency-critical and noncritical traffic. Robert P. Dick, Yihe Sun |
DATE | 4 |
| 2009 | Energy-efficient spatially-adaptive clustering and routing in wireless sensor networksabstractWireless sensor networks hold the potential to open new domains to distributed data acquisition. However, low-cost battery-powered nodes are often used to implement such networks, resulting in tight energy and communication bandwidth constraints. Cluster-based data compression and aggregation helps to reduce communication energy consumption. However, neglecting to adapt cluster sizes to local network conditions has limited the efficiency of previous clustering schemes. We have found that sensor node distances and densities are key factors in clustering. To the best of our knowledge, this is the first work taking these factors into consideration when adaptively forming data aggregation clusters. Compared with previous uniform-size clustering techniques, the proposed algorithm achieves up to 24% communication energy savings in uniform density networks and 36% savings in non-uniform density networks. Hengyu Long, Yongpan Liu, Xiaoguang Fan, Robert P. Dick, Huazhong Yang |
DATE | 4 |
| 2009 | Sonar-based measurement of user presence and attentionabstractWe describe a technique to detect the presence of computer users. This technique relies on sonar using hardware that already exists on commodity laptop computers and other electronic devices. It leverages the fact that human bodies have a different effect on sound waves than air and other objects. We conducted a user study in which 20 volunteers used a computer equipped with our ultrasonic sonar software. Our results show that it is possible to detect the presence or absence of users with near perfect accuracy after only ten seconds of measurement. We find that this technique can differentiate varied user positions and actions, opening the possibility of future use in estimating attention level. Stephen P. Tarzia, Robert P. Dick, Peter A. Dinda, Gokhan Memik |
UbiComp | 2 |
| 2009 | Battery allocation for wireless sensor network lifetime maximization under cost constraintsabstractWireless sensor networks hold the potential to open new domains to distributed data acquisition. However, such networks are prone to premature failure because some nodes deplete their batteries more rapidly than others due to workload variations, non-uniform communication, and heterogenous hardware. Many-to-one traffic patterns are common in sensor networks, further increasing node power consumption heterogeneity. Most previous sensor network lifetime enhancement techniques focused on balancing power distribution, based on the assumption of uniform battery capacity allocation among homogeneous nodes. Hengyu Long, Yongpan Liu, Robert P. Dick, Huazhong Yang |
ICCAD | 4 |
| 2009 | Archetype-based design: Sensor network programming for application experts, not just programming experts
Lan S. Bai, Robert P. Dick, Peter A. Dinda |
IPSN | 2 |
| 2009 | Online work maximization under a peak temperature constraintabstractIncreasing power densities and the high cost of low thermal resistance packages and cooling solutions make it impractical to design processors for worst-case temperature scenarios. As a result, packages and cooling solutions are designed for less than worst-case power densities and dynamic voltage and frequency scaling (DVFS) is used to prevent dangerous on-chip temperatures at run time. Unfortunately, DVFS can cause unpredicted drops in performance (e.g., long response times). We propose and optimally solve the problem of thermally-constrained online work maximization for general-purpose computing systems on uniprocessors with discrete speed levels and non-negligible transition overheads. Simulation results show that our approach completes 47.7% on average and up to 68.0% more cycles than a naive policy. Thidapat Chantem, Xiaobo Sharon Hu, Robert P. Dick |
ISLPED | 3 |
| 2009 | User- and process-driven dynamic voltage and frequency scalingabstractWe describe and evaluate two new, independently-applicable power reduction techniques for power management on processors that support dynamic voltage and frequency scaling (DVFS): user-driven frequency scaling (UDFS) and process-driven voltage scaling (PDVS). In PDVS, a CPU-customized profile is derived offline that encodes the minimum voltage needed to achieve stability at each combination of CPU frequency and temperature. On a typical processor, PDVS reduces the voltage below the worst-case minimum operating voltages given in datasheets. UDFS, on the other hand, dynamically adapts CPU frequency to the individual user and the workload through direct user feedback. Our UDFS algorithms dramatically reduce typical operating frequencies and voltages while maintaining performance at a satisfactory level for each user. We evaluate our techniques independently and together through user studies conducted on a Pentium M laptop running Windows applications. We measure the overall system power and temperature reduction achieved by our methods. Combining PDVS and the best UDFS scheme reduces measured system power by 49.9% (27.8% PDVS, 22.1% UDFS), averaged across all our users and applications, compared to Windows XP DVFS. The average temperature of the CPU is decreased by 13.2degC. User trace-driven simulation to evaluate the CPU only indicates average CPU dynamic power savings of 57.3% (32.4% PDVS, 24.9% UDFS), with a maximum reduction of 83.4%. In a multitasking environment, the same UDFS+PDVS technique reduces the CPU dynamic power by 75.7% on average. Bin Lin 0002, Arindam Mallik, Peter A. Dinda, Gokhan Memik, Robert P. Dick |
ISPASS | 5 |
| 2009 | iScope: personalized multi-modality image search for mobile devicesabstractMobile devices are becoming a primary medium for personal information gathering, management, and sharing. Managing personal image data on mobile platforms is a difficult problem due to large data set size, content diversity, heterogeneous individual usage patterns, and resource constraints. This article presents a user-centric system, called iScope, for personal image management and sharing on mobile devices. iScope uses multi-modality clustering of both content and context information for efficient image management and search, and online learning techniques for predicting images of interest. It also supports distributed content-based search among networked devices while maintaining the same intuitive interface, enabling efficient information sharing among people. We have implemented iScope and conducted in-field experiments using networked Nokia N810 portable Internet tablets. Energy efficiency was a primary design focus during the design and implementation of the iScope search algorithms. Experimental results indicate that iScope improves search time and search energy by 4.1X and 3.8X on average, relative to browsing. Changyun Zhu, Qin Lv, Robert P. Dick |
MobiSys | 5 |
| 2009 | Evaluating a BASIC approach to sensor network node programmingabstractSensor networks have the potential to empower domain experts from a wide range of fields. However, presently they are notoriously difficult for these domain experts to program, even though their applications are often conceptually simple. We address this problem by applying the BASIC programming language to sensor networks and evaluating its effectiveness. BASIC has proven highly successful in the past in allowing novices to write useful programs on home computers. Our contributions include a user study evaluating how well novice (no programming experience) and intermediate (some programming experience) users can accomplish simple sensor network tasks in BASIC and in TinyScript (a principally event-driven high-level language for node-oriented programming) and an evaluation of power consumption issues in BASIC. 45--55% of novice users can complete simple tasks in BASIC, while only 0--17% can do so in TinyScript. In both languages, users generally are most successful using imperative loop-oriented programming. The use of an interpreter, such as our BASIC implementation, has little impact on the power consumption of applications in which computational demands are low. Further, when in final form, BASIC can be compiled to reduce power consumption even further. J. Scott Miller, Peter A. Dinda, Robert P. Dick |
SenSys | 3 |
| 2009 | Implicit User Re-authentication for Mobile Devices
Sausan Yazji, Xi Chen 0068, Robert P. Dick, Peter Scheuermann |
UIC | 3 |
| 2009 | Multiscale Thermal Analysis for Nanometer-Scale Integrated CircuitsabstractThermal analysis has long been essential for designing reliable high-performance cost-effective integrated circuits (ICs). Increasing power densities are making this problem more important. Characterizing the thermal profile of an IC quickly enough to allow feedback on the thermal effects of tentative design changes is a daunting problem, and its complexity is increasing. The move to nanometer-scale fabrication processes is increasing the importance of thermal phenomena such as ballistic phonon transport. The accurate thermal analysis of nanometer-scale ICs containing hundreds of millions of devices requires characterization of heat transport across multiple length scales. These scales range from the nanometer scale (device-level impact) to the centimeter scale (cooling package impact). Existing chip-package thermal analysis methods based on classical Fourier heat transfer cannot capture nanometer-scale thermal effects. However, accurate device-level modeling techniques, such as molecular dynamics methods, are far too slow for use in full-chip IC thermal analysis. In this paper, we propose and develop ThermalScope, a multiscale thermal analysis method for nanometer-scale IC design. It unifies microscopic and macroscopic thermal modeling methods, i.e., the Boltzmann transport equation and Fourier modeling methods. Moreover, it supports adaptive multiresolution modeling. Together, these ideas enable the efficient and accurate characterization of nanometer-scale heat transport as well as the chip-package-level heat flow. ThermalScope is designed for full-chip thermal analysis of billion-transistor nanometer-scale IC designs, with accuracy at the scale of individual devices. ThermalScope enables the accurate characterization of various temperature-related effects, such as temperature-dependent leakage power and temperature-timing dependences. ThermalScope has been implemented in software and used for the full-chip thermal analysis and temperature-dependent leakage analysis of an IC design with more than 150 million transistors. ThermalScope will be publicly released for free academic and personal use. Zyad Hassan, Nicholas Allec, Robert P. Dick, V. Venkatraman, Ronggui Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2009 | MEMMU: Memory expansion for MMU-less embedded systemsabstractRandom access memory (RAM) is tightly constrained in the least expensive, lowest-power embedded systems such as sensor network nodes and portable consumer electronics. The most widely used sensor network nodes have only 4 to 10KB of RAM and do not contain memory management units (MMUs). It is difficult to implement complex applications under such tight memory constraints. Nonetheless, price and power-consumption constraints make it unlikely that increases in RAM in these systems will keep pace with the increasing memory requirements of applications. We propose the use of automated compile-time and runtime techniques to increase the amount of usable memory in MMU-less embedded systems. The proposed techniques do not increase hardware cost, and require few or no changes to existing applications. We have developed runtime library routines and compiler transformations to control and optimize the automatic migration of application data between compressed and uncompressed memory regions, as well as a fast compression algorithm well suited to this application. These techniques were experimentally evaluated on Crossbow TelosB sensor network nodes running a number of data-collection and signal-processing applications. Our results indicate that available memory can be increased by up to 50% with less than 10% performance degradation for most benchmarks. Lan S. Bai, Lei Yang 0017, Robert P. Dick |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2009 | Characterization of Single-Electron Tunneling Transistors for Designing Low-Power Embedded SystemsabstractMinimizing power consumption is vitally important in embedded system design; power consumption determines battery lifespan. Ultra-low-power designs may even permit embedded systems to operate without batteries by scavenging energy from the environment. Moreover, managing power dissipation is now a key factor in integrated circuit packaging and cooling. As a result, embedded system price, size, weight, and reliability are all strongly dependent on power dissipation. Recent developments in nanoscale devices open new alternatives for low-power embedded system design. Among these, single-electron tunneling transistors (SETs) hold the promise of achieving the lowest power consumption. Unfortunately, most analysis of SETs has focused on single devices instead of architectures, making it difficult to determine whether they are appropriate for low-power embedded systems. Evaluating the use of SETs in large-scale digital systems requires novel architectural and circuit design. SET-based design imposes numerous challenges resulting from low driving strength, relatively large static power consumption, and the presence of reliability problems resulting from random background charge effects. We propose a fault-tolerant, hybrid SET/CMOS, reconfigurable architecture, named IceFlex, that can be tailored to specific requirements and allows tradeoffs among power consumption, performance requirements, operation temperature, fabrication cost, and reliability. Using IceFlex as a testbed, we characterize the benefits and limitations of SETs in embedded system designs. In particular, we focus on the use of SETs in room-temperature ultra-low-power embedded systems such as wireless sensor network nodes. We also consider high-performance applications such as multimedia consumer electronics. We see this work as a first step in determining the potential of ultra-low-power embedded system design using SETs. Changyun Zhu, Zhengyu Gu, Robert P. Dick, Robert G. Knobel |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Multi-optimization power management for chip multiprocessorsabstractThe emergence of power as a first-class design constraint has fueled the proposal of a growing number of run-time power optimizations. Many of these optimizations trade-off power saving opportunity for a variable performance loss which depends on application characteristics and program phase. Furthermore, the potential benefits of these optimizations are sometimes non-additive, and it can be difficult to identify which combinations of these optimizations to apply. Trial-and-error approaches have been proposed to adaptively tune a processor. However, in a chip multiprocessor, the cost of individually configuring each core under a wide range of optimizations might be prohibitive under simple trial-and-error approaches. Russ Joseph, Robert P. Dick |
PACT | 3 |
| 2008 | PICSEL: measuring user-perceived performance to control dynamic frequency scalingabstractThe ultimate goal of a computer system is to satisfy its users. The success of architectural or system-level optimizations depends largely on having accurate metrics for user satisfaction. We propose to derive such metrics from information that is close to flesh and apparent to the user rather than from information that is close to metal and hidden from the user. We describe and evaluate PICSEL, a dynamic voltage and frequency scaling (DVFS) technique that uses measurements of variations in the rate of change of a computer's video output to estimate user-perceived performance. Our adaptive algorithms, one conservative and one aggressive, use these estimates to dramatically reduce operating frequencies and voltages for graphically-intensive applications while maintaining performance at a satisfactory level for the user. We evaluate PICSEL through user studies conducted on a Pentium M laptop running Windows XP. Experiments performed with 20 users executing three applications indicate that the measured laptop power can be reduced by up to 12.1%, averaged across all of our users and applications, compared to the default Windows XP DVFS policy. User studies revealed that the difference in overall user satisfaction between the more aggressive version of PICSEL and Windows DVFS were statistically insignificant, whereas the conservative version of PICSEL actually improved user satisfaction when compared to Windows DVFS. Arindam Mallik, Jack Cosgrove, Robert P. Dick, Gokhan Memik, Peter A. Dinda |
ASPLOS | 3 |
| 2008 | Adaptive Filesystem Compression for Embedded SystemsabstractEmbedded system secondary storage size is often constrained, yet storage demands are growing as a result of increasing application complexity and storage of personal data and multimedia flies. Filesystem compression offers a solution. This paper formalizes the problem of automatic filesystem compression using multiple compression algorithms. The average latency of on-line file accesses is optimized under a constraint on filesystem capacity. Our solution is based on predictive control. Predicted latency implications are used to solve the file compression state selection problem using a multiple choice knapsack problem formulation. This approach is evaluated on filesystem traces and compared with other efficient heuristics. Our approach results in 34.1% reduction in file access latency compared to a straight-forward heuristic that decompresses frequently-accessed files and compresses least recently used files with more aggressive compression algorithms. It reduces file access latency by 67.7% compared to uniformly compressing files to the shallowest level required to meet storage capacity constraints. Lan S. Bai, Haris Lekatsas, Robert P. Dick |
DATE | 3 |
| 2008 | Temperature-Aware Scheduling and Assignment for Hard Real-Time Applications on MPSoCsabstractThermal effects in MPSoCs may cause the violation of timing constraints in real-time systems. This paper presents a mixed integer linear programming based solution to this problem. Tasks are assigned and scheduled to an MPSoC to minimize peak temperature, subject to real-time constraints. The proposed approach outperforms existing methods, reducing peak temperature by up to 24.66degC and by an average of 8.75degC when compared to minimal-energy solutions. We also present a heuristic for use on large problem instances. Steady- state thermal analysis is used for tasks with long execution times compared to the RC thermal time constants of the cores. Transient analysis is used otherwise. The steady-state analysis based heuristic finds solutions with at most 3.40degC deviation from optimal peak temperature (0.22degC on average) while improving upon existing technique by as much as 25.71degC and 10.86degC on average. The transient analysis based heuristic further reduce peak temperature by 1degC in the best case and 0.17degC on average. Thidapat Chantem, Robert P. Dick, Xiaobo Sharon Hu |
DATE | 2 |
| 2008 | Operating System Controlled Processor-Memory Bus EncryptionabstractUnencrypted data appearing on the processor- memory bus can result in security violations, e.g., allowing attackers to gather keys to financial accounts and personal data. Although on-chip bus encryption hardware can solve this problem, it requires hardware redesign or increases processor cost. Application redesign to prevent sensitive data from appearing on the processor-memory bus is extremely difficult. We propose and evaluate a processor-memory bus encryption technique for embedded systems that requires no changes to applications or hardware. This technique exploits cache locking or scratchpad memory, features present in many embedded processors, permitting the operating system (OS) virtual memory infrastructure to automatically encrypt data belonging to protected processes as they are written to off-chip memory. Pages belonging to unprotected processes are stored unencrypted to prevent performance and energy consumption penalties. We evaluate the proposed bus encryption technique using full system simulation. Experimental results indicate that it is possible to prevent the working data sets of processes from appearing on the processor-memory bus in plaintext, without using dedicated hardware and without changing applications. The OS based technique results in 1.37times slowdown for protected processes for processors with 512 KB of L2 cache and 1.78times slowdown for processors with 256 KB of L2 cache. There are negligible performance penalties for unprotected processes. Xi Chen 0068, Robert P. Dick, Alok N. Choudhary |
DATE | 2 |
| 2008 | Design and Implementation of a High-Performance Microprocessor Cache Compression AlgorithmabstractAbstract Researchers have proposed using hardware data compression units within the memory hierarchies of microprocessors in order to improve performance, energy efficiency, and functionality. However, most past work, and in particular work on cache compression, has made unsubstantiated assumptions about the performance, power consumption, and area overheads of the required compression hardware. We present a lossless compression algorithm that has been designed for on-line memory hierarchy compression, and cache compression in particular. We reduced our algorithm to a register transfer level hardware implementation, permitting performance, power consumption, and area estimation. The results of experiments comparing our work to previous work are presented. Xi Chen 0068, Lei Yang 0017, Haris Lekatsas, Robert P. Dick |
DCC | 4 |
| 2008 | State space abstraction for parameterized self-stabilizing embedded systemsabstractSelf-stabilizing systems are systems that automatically recover from any transient fault. Proving the correctness of a parameterized self-stabilizing system, i.e., a system composed of an arbitrary number of processes, is a challenging task. For the verification of parameterized systems the method of control abstraction has been developed. However, control abstraction can only be applied to systems in which each process has a fixed number of observable variables. In this article, we propose a technique to abstract a parameterized self-stabilizing system, whose processes have a parameterized number of observable variables, to a system with fixed number of observable variables. This enables the use of control abstraction for verification. The proposed technique targets low-atomicity, shared-memory, asynchronous systems. We establish the completeness of the method under reasonable conditions and demonstrate its effectiveness by applying it on a number of self-stabilizing distributed systems. Nikolaos D. Liveris, Hai Zhou 0001, Robert P. Dick, Prithviraj Banerjee |
EMSOFT | 3 |
| 2008 | ThermalScope: multi-scale thermal analysis for nanometer-scale integrated circuitsabstractThermal analysis has long been essential for designing reliable, high-performance, cost-effective integrated circuits (ICs). Increasing power densities are making this problem more important. Characterizing the thermal profile of an IC quickly enough to allow feedback on the thermal effects of tentative design changes is a daunting problem, and its complexity is increasing. The move to nanoscale fabrication processes is increasing the importance of quantum thermal phenomena such as ballistic phonon transport. Accurate thermal analysis of nanoscale ICs containing hundreds of millions of devices requires characterization of thermal effects on length scales that vary by several orders of magnitude, from nanoscale quantum thermal effects to centimeter-scale cooling package impact. Existing chip.package thermal analysis methods based on classical Fourier heat transfer cannot capture nanoscale quantum thermal effects. However, accurate device-level modeling techniques, such as molecular dynamics methods, are far too slow for use in full-chip IC thermal analysis. In this work, we propose and develop ThermalScope, a multi-scale thermal analysis method for nanoscale IC design. It unifies microscopic and macroscopic thermal physics modeling methods, i.e., the Fourier and Boltzmann transport modeling methods. Moreover, it supports adaptive multi-resolution modeling. Together, these ideas enable efficient and accurate characterization of nanoscale quantum heat transport as well as chip.package level heat flow. ThermalScope is designed for full-chip thermal analysis of billion-transistor nanoscale IC designs, with accuracy at the scale of individual devices. ThermalScope enables accurate characterization of temperature-related effects, such as variation in leakage power and delay. ThermalScope has been implemented in software and used for full-chip thermal analysis and temperature-dependent leakage analysis of an IC design with more than 150 million transistors. It will be publicly released for free academic and personal use. Nicholas Allec, Zyad Hassan, Robert P. Dick, Ronggui Yang |
ICCAD | 4 |
| 2008 | Temperature-aware test scheduling for multiprocessor systems-on-chipabstractIncreasing power densities due to process scaling, combined with high switching activity and poor cooling environments during testing, have the potential to result in high integrated circuit (IC) temperatures. This has the potential to damage ICs and cause good ICs to be discarded due to temperature-induced timing faults. We first study the power impact of scan chain testing for the ISCAS89 benchmarks. We find that the scan-chain test power consumption is 1.6× higher for at-speed testing than normal operating power consumption. We conclude that if the testing frequency is less than half of the normal frequency, then the testing power consumption may in fact be lower. However, due to differences in the cooling environments, the peak die temperatures may still be higher. Second, we present an optimal formulation for minimal-duration temperature-constrained test scheduling. Our results improve on the test schedule time of the best existing algorithm by 10.8% on average for a packaged IC thermal environment. We also present an efficient heuristic that generally produces the same results as the optimal algorithm, while requiring little CPU time, even for large problem instances. David R. Bild, Sanchit Misra, Thidapat Chantem, Prabhat Kumar 0002, Robert P. Dick, Xiaobo Sharon Hu, Alok N. Choudhary |
ICCAD | 5 |
| 2008 | Learning and Leveraging the Relationship between Architecture-Level Measurements and Individual User SatisfactionabstractThe ultimate goal of computer design is to satisfy the end-user. In particular computing domains, such as interactive applications, there exists a variation in user expectations and user satisfaction relative to the performance of existing computer systems. In this work, we leverage this variation to develop more efficient architectures that are customized to end-users. We first investigate the relationship between microarchitectural parameters and user satisfaction. Specifically, we analyze the relationship between hardware performance counter (HPC) readings and individual satisfaction levels reported by users for representative applications. Our results show that the satisfaction of the user is strongly correlated to the performance of the underlying hardware. More importantly, the results show that user satisfaction is highly user-dependent. To take advantage of these observations, we develop a framework called Individualized Dynamic Voltage and Frequency Scaling (iDVFS). We study a group of users to characterize the relationship between the HPCs and individual user satisfaction levels. Based on this analysis, we use artificial neural networks to model the function from HPCs to user satisfaction for individual users. This model is then used online to predict user satisfaction and set the frequency level accordingly. A second set of user studies demonstrates that iDVFS reduces the CPU power consumption by over 25% in representative applications as compared to the Windows XP DVFS algorithm. Alex Shye, Berkin Özisikyilmaz, Arindam Mallik, Gokhan Memik, Peter A. Dinda, Robert P. Dick, Alok N. Choudhary |
ISCA | 6 |
| 2008 | Power to the people: Leveraging human physiological traits to control microprocessor frequencyabstractAny architectural optimization aims at satisfying the end user. However, modern architectures execute with little to no knowledge about the individual user. If architectures could determine whether their users are satisfied, they could provide higher efficiency; improved reliability, reduced power consumption, increased security, and a better user experience. A major reason for this limitation is their input devices. Specifically, the traditional input devices (e.g., the mouse and keyboard) provide limited information about the user. In this paper, we make a case for the addition of new biometric input devices for providing the computer information about the userpsilas physiological traits. We explore three biometric devices as potential sensors: an eye tracker, a galvanic skin response (GSR) sensor, and force sensors. We first present two user studies that explore the link between the sensor readings and user satisfaction when the performance of the processor is varied as a video game is being played. In the first study, we drastically drop the processor clock frequency at a set point in the game. In the second study, we set the clock frequency to randomly-selected levels during game play. Both studies show that there are significant changes in human physiological traits as performance decreases. More importantly, we show that physiological changes correlate strongly to the satisfaction levels reported by the users. Based upon these observations, we construct a Physiological Traits-based Power-management (PTP) system that can be applied to existing dynamic voltage and frequency scaling (DVFS) schemes. We apply PTP to a typical CPU-utilization-based adaptive DVFS policy and evaluate our scheme using a third user study. An aggressive version of our PTP scheme reduces the total system power consumption of a laptop by up to 33.3% for an application averaged across users (18.1% averaged across three applications), while a conservative version reduces the total system power consumption by up to 25.6% across users (11.4% averaged across three applications). Alex Shye, Yan Pan 0010, Benjamin Scholbrock, J. Scott Miller, Gokhan Memik, Peter A. Dinda, Robert P. Dick |
MICRO | 7 |
| 2008 | Three-Dimensional Chip-Multiprocessor Run-Time Thermal ManagementabstractThree-dimensional integration has the potential to improve the communication latency and integration density of chip-level multiprocessors (CMPs). However, the stacked high-power density layers of 3D CMPs increase the importance and difficulty of thermal management. In this paper, we investigate the 3D CMP run-time thermal management problem and describe efficient management techniques. This paper makes the following main contributions: 1) It identifies and describes the critical concepts required for optimal thermal management, namely the methods by which heterogeneity in both workload power characteristics and processor core thermal characteristics should be exploited; and 2) it proposes an efficient proactive continuously engaged hardware and operating system thermal management technique governed by optimal thermal management polices. The proposed technique is evaluated using multiprogrammed and multithreaded benchmarks in an integrated power, performance, and temperature full-system simulation environment. We find that proactive power-thermal budgeting allows a 30% improvement in instruction throughput compared to a proactive thermal management approach that bases decisions only upon local information. The software components of the proposed thermal management technique have been implemented in the Linux 2.6.8 kernel. This source code will be publicly released. The analysis and technique developed in this paper provide a general solution for future 3D and 2D CMPs. Changyun Zhu, Zhenyu (Peter) Gu, Robert P. Dick, Russ Joseph |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2008 | Application-Specific MPSoC Reliability OptimizationabstractThis paper presents modeling and estimation techniques permitting the temperature-aware optimization of application-specific multiprocessor system-on-chip (MPSoC) reliability. Technology scaling and increasing power densities make MPSoC lifetime reliability problems more severe. MPSoC reliability strongly depends on system-level MPSoC architecture, redundancy, and thermal profile during operation. We propose an efficient temperature-aware MPSoC reliability analysis and prediction technique that enables MPSoC reliability optimization via redundancy and temperature-aware design planning. Reliability, performance, and area are concurrently optimized. Simulation results indicate that the proposed approach has the potential to substantially improve MPSoC system mean time to failure with small area overhead. Zhenyu (Peter) Gu, Changyun Zhu, Robert P. Dick |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2007 | Towards An Ultra-Low-Power Architecture Using Single-Electron Tunneling TransistorsabstractMinimizing power consumption is vitally important in embedded system design; power consumption determines battery lifespan. Ultra-low-power designs may even permit embedded systems to operate without batteries, e.g., by scavenging energy from the environment. Moreover, managing power dissipation is now a key factor in integrated circuit packaging and cooling. As a result, embedded system price, size, weight, and reliability are all strongly dependent on power dissipation. Changyun Zhu, Zhenyu (Peter) Gu, Robert P. Dick, Robert G. Knobel |
DAC | 4 |
| 2007 | Accurate temperature-dependent integrated circuit leakage power estimation is easy
Yongpan Liu, Robert P. Dick, Huazhong Yang |
DATE | 2 |
| 2007 | 3D-STAF: scalable temperature and leakage aware floorplanning for three-dimensional integrated circuitsabstractThermal issues are a primary concern in the threedimensional (3D) integrated circuit (IC) design. Temperature, area, and wire length must be simultaneously optimized during 3D floorplanning, significantly increasing optimization complexity. Most existing floorplanners use combinatorial stochastic optimization techniques, hampering performance and scalability when used for 3D floorplanning. In this work, we propose and evaluate a scalable, temperature-aware, force-directed floorplanner called 3D-STAF. Force-directed techniques, although efficient at reacting to physical information such as temperature gradients, must eventually eliminate overlap. This can cause significant displacement when used for heterogeneous blocks. To smooth the transition from an unconstrained 3D placement to a legalized, layer-assigned floorplan, we propose a three-stage force-directed optimization flow combined with new legalization techniques that eliminate white spaces and block overlapping during multi-layer floorplanning. A temperature-dependent leakage model is used within 3D-STAF to permit optimization based on the feedback loop connecting thermal profile and leakage power consumption. 3D-STAF has good performance that scales well for large problem instances. Compared to recently published 3D floorplanning work, 3D-STAF improves the area by 6%, wire length by 16%, via count by 22%, peak temperature by 6% while running nearly 4× faster on average. Pingqiang Zhou, Yuchun Ma, Zhuoyuan Li 0003, Robert P. Dick, Hai Zhou 0001, Xianlong Hong, Qiang Zhou 0001 |
ICCAD | 4 |
| 2007 | Lucid dreaming: reliable analog event detection for energy-constrained applicationsabstractExisting sensor network architectures are based on the assumption that data will be polled. Therefore, they are not adequate for long-term battery-powered use in applications that must sense or react to events that occur at unpredictable times. In response, and motivated by a structural autonomous crack monitoring (ACM) application from civil engineering that requires bursts of high resolution sampling in response to aperiodic vibrations in buildings and bridges, we have designed, implemented, and evaluated lucid dreaming, a hardware--software technique to dramatically decrease sensor node power consumption in this and other event-driven sensing applications. Sasha Jevtic, Mathew Kotowsky, Robert P. Dick, Peter A. Dinda, Charles Dowding |
IPSN | 3 |
| 2007 | Power reduction through measurement and modeling of users and CPUs: summaryabstractDynamic Voltage and Frequency Scaling (DVFS) is one of the most commonly used power reduction techniques in high-performance processors. DVFS varies the frequency and voltage of a microprocessor in real-time according to processing needs. Although there are different versions of DVFS, at its core DVFS adapts power consumption and performance to the current workload of the CPU. Specifically, existing DVFS techniques in high-performance processors select an operating point (CPU frequency and voltage) based on the utilization of the processor. This approach integrates OS-level control, but such control is pessimistic. Existing DVFS techniques are pessimistic about the user. Indeed, they ignore the user, assuming that CPU utilization or the OS events prompting it are sufficient proxies. A high CPU utilization simply leads to a high frequency and high voltage, regardless of the user’s satisfaction or expectation of performance. Existing DVFS techniques are pessimistic about the CPU. They assume worst-case manufacturing process variation and operating temperature by basing their policies on loose worstcase bounds given by the processor manufacturer. A voltage level for a given frequency is set such that even the worst shipped processor of a given generation will be stable at the highest specified temperature. In response to these observations, we have developed, implemented, and evaluated the following two new power management techniques that can be readily employed independently or together. We elaborate on these techniques in detail elsewhere [4]. Bin Lin 0002, Arindam Mallik, Peter A. Dinda, Gokhan Memik, Robert P. Dick |
SIGMETRICS | 5 |
| 2007 | Unified Incremental Physical-Level and High-Level SynthesisabstractAchieving design closure is one of the biggest challenges for modern very large-scale integration system designers. This problem is exacerbated by the lack of high-level design-automation tools that consider the increasingly important impact of physical features, such as interconnect, on integrated circuit area, performance, and power consumption. Using physical information to guide decisions in the behavioral-level stage of system design is essential to solve this problem. In this paper, we present an incremental floorplanning high-level-synthesis system. This system integrates high-level and physical-design algorithms to concurrently improve a design's schedule, resource binding, and floorplan, thereby allowing the incremental exploration of the combined behavioral-level and physical-level design space. Compared with previous approaches that repeatedly call loosely coupled floorplanners for physical estimation, this approach has the benefits of efficiency, stability, and better quality of results. The average CPU time speedup resulting from unifying incremental physical-level and high-level synthesis is 24.72times and area improvement is 13.76%. The low power consumption of a state-of-the-art low-power interconnect-aware high-level-synthesis algorithm is maintained. The benefits of concurrent behavioral-level and physical-design optimization increased for larger problem instances. Zhenyu (Peter) Gu, Jia Wang 0003, Robert P. Dick, Hai Zhou 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | SLOPES: Hardware-Software Cosynthesis of Low-Power Real-Time Distributed Embedded Systems With Dynamically Reconfigurable FPGAsabstractIn this paper, we present a multiobjective hardware-software cosynthesis system, called SLOPES, for multirate low-power real-time distributed embedded systems consisting of dynamically reconfigurable field-programmable gate arrays (FPGAs), processors, and heterogeneous communication resources. This cosynthesis algorithm simultaneously optimizes system price and average power consumption. First, we present an evolutionary algorithm that automatically determines the quantities and types of system resources, assigns tasks to different potentially reconfigurable processing elements, and assigns communication events to communication resources. Second, we propose a dynamic priority multirate scheduling algorithm to determine the times at which all the tasks and communication events in the system occur. This two-dimensional scheduling algorithm determines task priorities based on real-time constraints and detailed frame-by-frame FPGA reconfiguration overhead information. Experimental results indicate that the proposed method reduces schedule length by an average of 34.3% and reconfiguration energy by an average of 40.4%, compared to a method that does not consider the effect of partial reconfiguration during synthesis. SLOPES yields multiple system architectures that tradeoff system price and average power consumption under real-time constraints Robert P. Dick, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | ISAC: Integrated Space-and-Time-Adaptive Chip-Package Thermal AnalysisabstractEver-increasing integrated circuit (IC) power densities and peak temperatures threaten reliability, performance, and economical cooling. To address these challenges, thermal analysis must be embedded within IC synthesis. However, this requires accurate three-dimensional chip-package heat flow analysis. This has typically been based on numerical methods that are too computationally intensive for numerous repeated applications during synthesis or design. Thermal analysis techniques must be both accurate and fast for use in IC synthesis. This paper presents a novel accurate incremental spatially and temporally adaptive chip-package thermal analysis technique called ISAC for use in IC synthesis and design. It is common for IC temperature variation to strongly depend on position and time. ISAC dynamically adapts spatial- and temporal-modeling granularities to achieve high efficiency while maintaining accuracy. Both steady-state and dynamic thermal analyses are accelerated by the proposed heterogeneous spatial-resolution adaptation and asynchronous thermal-element time-marching techniques. Each technique enables orders-of-magnitude improvement in performance while preserving accuracy when compared with other state-of-the-art adaptive steady-state and dynamic IC thermal analysis techniques. Experimental results indicate that these improvements are sufficient to make accurate dynamic and steady-state thermal analysis practical within the inner loops of IC synthesis algorithms. ISAC has been validated against reliable commercial thermal analysis tools using industrial and academic synthesis test cases and chip designs. It has been implemented as a software package suitable for integration in IC synthesis and design flows and has been publicly released Yonghong Yang, Zhenyu (Peter) Gu, Changyun Zhu, Robert P. Dick |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2006 | TAPHS: thermal-aware unified physical-level and high-level synthesisabstractThermal effects are becoming increasingly important during integrated circuit design. Thermal characteristics influence reliability, power consumption, cooling costs, and performance. It is necessary to consider thermal effects during all levels of the design process, from the architectural level to the physical level. However, design-time temperature prediction requires access to block placement, wire models, power profile, and a chip-package thermal model. Thermal-aware design and synthesis necessarily couple architectural-level design decisions (e.g., scheduling) with physical design (e.g., floorplanning) and modeling (e.g., wire and thermal modeling). This article proposes an efficient and accurate thermal-aware floor-planning high-level synthesis system that makes use of integrated high-level and physical-level thermal optimization techniques. Voltage islands are automatically generated via novel slack distribution and voltage partitioning algorithms in order to reduce the design's power consumption and peak temperature. A new thermal-aware floorplanning technique is proposed to balance chip thermal profile, thereby further reducing peak temperature. The proposed system was used to synthesize a number of benchmarks, yielding numerous designs that trade off peak temperature, integrated circuit area, and power consumption. The proposed techniques reduces peak temperature by 12.5degC on average. When used to minimize peak temperature with a fixed area, peak temperature reductions are common. Under a constraint on peak temperature, integrated circuit area is reduced by 9.9% on average Zhenyu (Peter) Gu, Yonghong Yang, Jia Wang 0003, Robert P. Dick |
ASP-DAC | 4 |
| 2006 | Automated compile-time and run-time techniques to increase usable memory in MMU-less embedded systemsabstractRandom access memory (RAM) is tightly-constrained in many embedded systems. This is especially true for the least expensive, lowest-power embedded systems, such as sensor network nodes and portable consumer electronics. The most widely-used sensor network nodes have only 4-10 KB of RAM and do not contain memory management units (MMUs). It is very difficult to implement increasingly complex applications under such tight memory constraints. Nonetheless, price and power consumption constraints make it unlikely that increases in RAM in these systems will keep pace with the requirements of applications.We propose the use of automated compile-time and run-time techniques to increase the amount of usable memory in MMU-less embedded systems. The proposed techniques do not increase hardware cost, and are designed to require few or no changes to existing applications. We have developed a fast compression algorithm well suited to this application, as well as run-time library routines and compiler transformations to control and optimize the automatic migration of application data between compressed and uncompressed memory regions. These techniques were experimentally evaluated on Crossbow TelosB sensor network nodes running a number of data collection and signal processing applications. The results indicate that available memory can be increased by up to 50% with less than 10% performance degradation for most benchmarks. Lan S. Bai, Lei Yang 0017, Robert P. Dick |
CASES | 3 |
| 2006 | High-performance operating system controlled memory compressionabstractThis article describes a new software-based on-line memory compression algorithm for embedded systems and presents a method of adaptively managing the uncompressed and compressed memory regions during application execution. The primary goal of this work is to save memory in disk-less embedded systems, resulting in greater functionality, smaller size, and lower overall cost, without modifying applications or hardware. In comparison with algorithms that are commonly used in on-line memory compression, our new algorithm has a comparable compression ratio but is twice as fast. The adaptive memory management scheme effectively responds to the predicted needs of applications and prevents on-line memory compression deadlock, permitting reliable and efficient compression for a wide range of applications. We have evaluated our technique on an embedded portable device and have found that the memory available to applications can be increased by 150%, allowing the execution of applications with larger working data sets, or allowing existing applications to run with less physical memory. Lei Yang 0017, Haris Lekatsas, Robert P. Dick |
DAC | 3 |
| 2006 | Adaptive chip-package thermal analysis for synthesis and designabstractEver-increasing integrated circuit(IC) power densities and peak temperatures threaten reliability, performance, and economical cooling. To address these challenges, thermal analysis must be embedded within IC synthesis. However, detailed thermal analysis requires accurate three-dimensional chip-package heat flow analysis. This has typically been based on numerical methods that are too computationally intensive for numerous repeated applications during synthesis or design. Thermal analysis techniques must be both accurate and fast for use in IC synthesis. This article presents a novel, accurate, incremental, self-adaptive, chip-package thermal analysis technique, called ISAC, for use in IC synthesis and design. It is common for IC temperature variation to strongly depend on position and time. ISAC dynamically adapts spatial and temporal modeling granularity to achieve high efficiency while maintaining accuracy. Both steady-state and dynamic thermal analysis are accelerated by the proposed heterogeneous spatial resolution adaptation and temporally decoupled element time marching techniques. Each technique enables orders of magnitude improvement in performance while preserving accuracy when compared with other state-of-the-art adaptive steady-state and dynamic IC thermal analysis techniques. Experimental results indicate that these improvements are sufficient to make accurate dynamic and static thermal analysis practical with in the inner loops of IC synthesis algorithms. ISAC has been validated against reliable commercial thermal analysis tools using industrial and academic synthesis test cases and chip designs. It has been implemented as a software package suitable for integration in IC synthesis and design flows and has been publicly released. Yonghong Yang, Zhenyu (Peter) Gu, Changyun Zhu, Robert P. Dick |
DATE | 5 |
| 2006 | Adaptive multi-domain thermal modeling and analysis for integrated circuit synthesis and designabstractChip-package thermal analysis is necessary for the design and synthesis of reliable, high-performance, low-power, compact integrated circuits (ICs). Many methods of IC thermal analysis suffer performance or accuracy problems that prevent use in IC synthesis and hinder use in architectural design. Yonghong Yang, Changyun Zhu, Zhenyu (Peter) Gu, Robert P. Dick |
ICCAD | 5 |
| 2005 | FD-HGAC: a hybrid heuristic/genetic algorithm hardware/software co-synthesis framework with fault detectionabstractEmbedded real-time systems are becoming increasingly complex. To combat the rising design cost of those systems, co-synthesis tools that map tasks to systems containing both software and specialized hardware have been developed. As system transient fault rates increase due to technology scaling, embedded systems must be designed in fault tolerant ways to maintain system reliability. This paper presents and analyzes FD-HGAC, a tool using a genetic algorithm and heuristics to design real-time systems with partial fault detection. Results of numerous trials of the tool are shown to produce systems with average 22% detection coverage that incurs no cost or performance penalty. John Conner, Yuan Xie 0001, Mahmut T. Kandemir, Robert P. Dick, Greg M. Link |
ASP-DAC | 4 |
| 2005 | Incremental exploration of the combined physical and behavioral design spaceabstractAchieving design closure is one of the biggest headaches for modern VLSI designers. This problem is exacerbated by high-level design automation tools that ignore increasingly important factors such as the impact of interconnect on the area and power consumption of integrated circuits. Bringing physical information up into the logic level or even behavioral-level stages of system design is essential to solve this problem. In this paper, we present an incremental floorplanning high-level synthesis system. This system integrates high-level and physical design algorithms to concurrently improve a system's schedule, resource binding, and floorplan, thereby allowing the incremental exploration of the combined behavioral-level and physical-level design space. Compared with previous approaches that repeatedly call loosely coupled floorplanners for physical estimation, this approach has the benefit of effi- ciency, stability, and better quality of results. For designs containing functional units with non-unity aspect ratios, the average CPU time improved by 369 %, the area improved by 14.24%, and power improved by 4%. Zhenyu (Peter) Gu, Jia Wang 0003, Robert P. Dick, Hai Zhou 0001 |
DAC | 3 |
| 2004 | Energy-aware deterministic fault tolerance in distributed real-time embedded systemsabstractWe investigate a unified approach for fault tolerance and dynamic power management in distributed real-time embedded systems. Coordinated checkpointing is used to achieve fault tolerance, and power management is carried out using dynamic voltage scaling. We present feasibility-of-scheduling tests for coordinated checkpointing schemes for a constant processor speed as well as for DVS-enabled processors that can operate at variable speeds. Simulation results based on the CORDS hardware/software co-synthesis system show that, compared to fault-oblivious methods, the proposed approach significantly reduces power consumption while guaranteeing timely task completion in the presence of faults. Ying Zhang 0041, Robert P. Dick, Krishnendu Chakrabarty |
DAC | 2 |
| 2004 | COWLS: hardware-software cosynthesis of wireless low-power distributed embedded client-server systemsabstractIn this paper, we present COWLS, a hardware-software cosynthesis algorithm that targets embedded systems composed of servers and low-power clients that communicate with each other through a channel of limited bandwidth, e.g., a wireless link. A novel scheduling algorithm is used to pipeline the execution of tasks that serve multiple clients associated with a given server. COWLS simultaneously optimizes the price of the client-server system, the power consumption of the clients, and the response times of tasks that have only soft deadlines, while meeting all of the hard deadlines. It produces numerous solutions that trade off different architectural features, e.g., price, power consumption, and response time, of an embedded client-server system. As far as we know, this is the first synthesis algorithm of its kind. We present the experimental results for numerous pseudorandom examples, a low-power client-server camera system, as well as the rest of the benchmarks within a publicly released embedded system synthesis benchmark suite. Robert P. Dick, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | DESP: A Distributed Economics-Based Subcontracting Protocol for Computation Distribution in Power-Aware Mobile Ad Hoc NetworksabstractIn this paper, we present a new economics-based power-aware protocol, called the distributed economic subcontracting protocol (DESP) that dynamically distributes task computation among mobile devices in an ad hoc wireless network. Mobile computation devices may be energy buyers, contractors, or subcontractors. Tasks are transferred between devices via distributed bargaining and transactions. When additional energy is required, buyers and contractors negotiate energy prices within their local markets. Contractors and subcontractors spend communication and computation energy to relay or execute buyers' tasks. Buyers pay the negotiated price for this energy. Decision-making algorithms are proposed for buyers, contractors, and subcontractors, each of which has a different optimization goal. We have built a wireless network simulator, called ESIM, to assist in the design and analysis of these algorithms. When the average communication energy required transferring a task is less than the average energy required to execute a task, our experimental results indicate that markets based on our protocol and decision-making algorithms fairly and effectively allocate energy resources among different tasks in both cooperative and competitive scenarios. Robert P. Dick, Niraj K. Jha |
IEEE Trans. Mob. Comput. | 2 |
| 2003 | Analysis of power dissipation in embedded systems using real-time operating systemsabstractThe increasing complexity and software content of embedded systems has led to the frequent use of system software to help applications access hardware resources easily and efficiently. In this paper, we present a method for detailed analysis of real-time operating system (RTOS) power consumption. RTOSs form an important component of the system software layer. Despite the widespread use of, and significant role played by, RTOSs in mobile and low-power embedded systems, little is known about their power-consumption effects. This paper presents a method of producing a hierarchical energy-consumption profile for applications as they interact with an RTOS. As a proof-of-concept, we use our infrastructure to produce the power profiles for a commercial RTOS, /spl mu/C/OS-II, running several applications on an embedded system based on the Fujitsu SPARClite processor. These examples demonstrate that an RTOS can have a significant impact on power consumption. We discuss ways in which application software can be designed to use an RTOS in a power-efficient manner. We believe that this is a first step toward establishing a systematic approach to power optimization of embedded systems containing RTOSs. Robert P. Dick, Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | Power analysis of embedded operating systemsabstractThe increasing complexity and software content of embedded systems has led to the frequent use of system software that helps applications access underlying hardware resources easily and efficiently. In this paper, we analyze the power consumption of real-time operating systems (RTOSs), which form an important component of the system software layer. Despite the widespread use of, and significant role played by, RTOSs in mobile and low-power embedded systems, little is known about their power consumption characteristics. This work presents the power profiles for a commercial RTOS, μC/OS, running several applications on an embedded system based on the Fujitsu SPARClite processor. Our work demonstrates that the RTOS can consume a significant fraction of the system power and, in addition, impact the power consumed by other software components. We illustrate the ways in which application software can be designed to use the RTOS in a power-efficient manner. We believe that this work is a first step towards establishing a systematic approach to RTOS power modeling and optimization. Robert P. Dick, Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha |
DAC | 1 |
| 1999 | MOCSYN: Multiobjective Core-Based Single-Chip System SynthesisabstractIn this paper we present a system synthesis algorithm, called MOCSYN, which partitions and schedules embedded system specifications to intellectual property cores in an integrated circuit. Given a system specification consisting of multiple periodic task graphs as well as a database of core and integrated circuit characteristics, MOCSYN synthesizes real-time heterogeneous single-chip hardware software architectures using an adaptive multiobjective genetic algorithm that is designed to escape local minima. The use of multiobjective optimization allows a single system synthesis run to produce multiple designs which trade off different architectural features. Integrated circuit price, power consumption, and area are optimized under hard real-time constraints. MOCSYN differs from previous work by considering problems unique to single-chip systems. It solves the problem of providing clock signals to cores composing a system-on-a-chip. It produces a bus structure which balances ease of layout with the reduction of bus contention. In addition, it carries out floorplan block placement within its inner loop allowing accurate estimation of global communication delays and power consumption. Robert P. Dick, Niraj K. Jha |
DATE | 1 |
| 1999 | Corrections to "mogac: a multiobjective genetic algorithm for hardware-software cosynthesis of distributed embedded systems"abstractIn the above-named article [ibid., vol. 16, pp. 920-935, Oct. 1998] there is an error in the experimental results reported for the Hou 3&4 (clustered) benchmark, due to a typographical error in its task collection input file. The last line of the The Hou 3&4 (clustered) row of Table IV on page 932 of the original article should be replaced with Table I of this short paper. For this example, average price increases when the effort level is changed from three to four. In general, solution quality increases, i.e., the price decreases, with increasing optimizer effort. This trend is not expressed in this case as a result of the limited number of samples (30). When run with 500 samples, price strictly decreases with increasing effort. The Hou 3&4 (clustered) row of Table VI on page 933 of the above paper1 should be replaced with Table II of this short paper. Robert P. Dick, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | CORDS: hardware-software co-synthesis of reconfigurable real-time distributed embedded systemsabstractArticle CORDS: hardware-software co-synthesis of reconfigurable real-time distributed embedded systems Share on Authors: Robert P. Dick Department of Electrical Engineering, Princeton University, 2 Princeton, New Jersey Department of Electrical Engineering, Princeton University, 2 Princeton, New JerseyView Profile , Niraj K. Jha Department of Electrical Engineering, Princeton University, 2 Princeton, New Jersey Department of Electrical Engineering, Princeton University, 2 Princeton, New JerseyView Profile Authors Info & Claims ICCAD '98: Proceedings of the 1998 IEEE/ACM international conference on Computer-aided designNovember 1998 Pages 62–67https://doi.org/10.1145/288548.288561Online:01 November 1998Publication History 66citation479DownloadsMetricsTotal Citations66Total Downloads479Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Robert P. Dick, Niraj K. Jha |
ICCAD | 1 |
| 1998 | MOGAC: a multiobjective genetic algorithm for hardware-software cosynthesis of distributed embedded systemsabstractIn this paper, we present a hardware-software cosynthesis system, called MOGAC, that partitions and schedules embedded system specifications consisting of multiple periodic task graphs. MOGAC synthesizes real-time heterogeneous distributed architectures using an adaptive multiobjective genetic algorithm that can escape local minima. Price and power consumption are optimized while hard real-time constraints are met. MOGAC places no limit on the number of hardware or software processing elements in the architectures it synthesizes. Our general model for bus and point-to-point communication links allows a number of link types to be used in an architecture. Application-specific integrated circuits consisting of multiple processing elements are modeled. Heuristics are used to tackle multirate systems, as well as systems containing task graphs whose hyperperiods are large relative to their periods. The application of a multiobjective optimization strategy allows a single cosynthesis run to produce multiple designs that trade off different architectural features. Experimental results indicate that MOGAC has advantages over previous work in terms of solution quality and running time. Robert P. Dick, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1997 | MOGAC: a multiobjective genetic algorithm for the co-synthesis of hardware-software embedded systemsabstractWe present a hardware-software co-synthesis system, called MOGAC, that partitions and schedules embedded system specifications consisting of multiple periodic task graphs. MOGAC synthesizes real-time heterogeneous distributed architectures using an adaptive multiobjective genetic algorithm that can escape local minima. Price and power consumption are optimized while hard real-time constraints are met. MOGAC places no limit on the number of hardware or software processing elements in the architectures it synthesizes. Our general model for bus and point-to-point communication links allows a number of link types to be used in an architecture. Application-specific integrated circuits consisting of multiple processing elements are modeled. Heuristics are used to tackle multi-rate systems, as well as systems containing task graphs whose hyperperiods are large relative to their periods. The application of a multiobjective optimization strategy allows a single co-synthesis run to produce multiple designs which trade off different architectural features. Experimental results indicate that MOGAC has advantages over previous work in terms of solution quality and running time. Robert P. Dick, Niraj K. Jha |
ICCAD | 1 |