EDBT 2026 Demo / reviewers in the wild / expert
Rajesh K. Gupta 0001
dblp:213/9138-1 · also Rajesh Gupta 0001
· DBLP profile ↗
192ranked-venue papers
16as first author
19since 2021 · last 2025
0000-0002-6489-7633ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 126 · 13 first-author · 2 since 2021Software engineering, systems software and programming languages · 42 · 3 first-authorComputer networks · 21 · 4 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 since 2021Theory of computation · 9 · 2 first-authorDatabases, data management, data science and information retrieval · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ZeroHAR: Sensor Context Augments Zero-Shot Wearable Action RecognitionabstractWearable Human Action Recognition (wHAR) uses motion sensor data to identify human movements, which is essential for mobile and wearable devices. However, traditional wHAR systems are only trained on a limited set of activities. Hence, they fail to generalize to diverse human motions, prompting Zero-Shot Learning (ZSL). Existing ZSL methods for wHAR focus solely on augmenting labels, such as representing them as attribute matrices, images, videos, or text. We propose ZeroHAR that enhances ZSL by not just focusing on activity labels, but by augmenting motion data with sensor context features. Our approach incorporates information about the sensor type, the Cartesian axis of the data, and the sensor's body position, providing the model with crucial spatial and biomechanical insights. This helps the model generalize better to new actions. First, we train the model by aligning the latent space of the motion time-series with its corresponding sensor context, while distancing it from unrelated sensor contexts. Finally, we train the model using the target activity descriptions. We tested our method against eight baselines on five benchmark HAR datasets with various sensors, placements, and activities. Our model shows exceptional generalizability across 18 motion time series classification benchmark datasets, outperforming the best baselines by 262% in the zero-shot setting. Ranak Roy Chowdhury, Ritvik Kapila, Ameya Panse, Xiyuan Zhang 0001, Diyan Teng, Rashmi Kulkarni, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
AAAI | 8 |
| 2025 | Matching Skeleton-based Activity Representations with Heterogeneous Signals for HARabstractIn human activity recognition (HAR), activity labels have typically been encoded in one-hot format, which has a recent shift towards using textual representations to provide contextual knowledge. Here, we argue that HAR should be anchored to physical motion data, as motion forms the basis of activity and applies effectively across sensing systems, whereas text is inherently limited. We propose SKELAR, a novel HAR framework that pretrains activity representations from skeleton data and matches them with heterogeneous HAR signals. Our method addresses two major challenges: (1) capturing core motion knowledge without context-specific details. We achieve this through a self-supervised coarse angle reconstruction task that recovers joint rotation angles, invariant to both users and deployments; (2) adapting the representations to downstream tasks with varying modalities and focuses. To address this, we introduce a self-attention matching module that dynamically prioritizes relevant body parts in a data-driven manner. Given the lack of corresponding labels in existing skeleton data, we establish MASD, a new HAR dataset with IMU, WiFi, and skeleton, collected from 20 subjects performing 27 activities. This is the first broadly applicable HAR dataset with time-synchronized data across three modalities. Experiments show that SKELAR achieves the state-of-the-art performance in both full-shot and few-shot settings. We also demonstrate that SKELAR can effectively leverage synthetic skeleton data to extend its use in scenarios without skeleton collections. Shuheng Li, Jiayun Zhang, Xiaohan Fu, Xiyuan Zhang 0001, Jingbo Shang, Rajesh K. Gupta 0001 |
SenSys | 6 |
| 2025 | Contextual Inference From Sparse Shopping Transactions Based on Motif PatternsabstractInferring contextual information such as demographics from historical transactions is valuable to public agencies and businesses. Existing methods are data-hungry and do not work well when the available records of transactions are sparse. We consider here specifically inference of demographic information using limited historical grocery transactions from a few random trips that a typical business or public service organization may see. We propose a novel method calledDemoMotifto build a network model from heterogeneous data and identify subgraph patterns (i.e., motifs) that enable us to infer demographic attributes. We then design a novel motif context selection algorithm to find specific node combinations significant to certain demographic groups. Finally, we learn representations of households using these selected motif instances as context, and employ a standard classifier (e.g., SVM) for inference. For evaluation purposes, we use three real-world consumer datasets, spanning different regions and time periods in the U.S. We evaluate the framework for predicting three attributes: ethnicity, seniority of household heads, and presence of children. Extensive experiments and case studies demonstrate thatDemoMotifis capable of inferring household demographics using only a small number (e.g., fewer than 10) of random grocery trips, significantly outperforming the state-of-the-art. Jiayun Zhang, Xinyang Zhang 0002, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Large Language Models for Time Series: A Survey
Xiyuan Zhang 0001, Ranak Roy Chowdhury, Rajesh K. Gupta 0001, Jingbo Shang |
IJCAI | 3 |
| 2024 | UniMTS: Unified Pre-training for Motion Time SeriesabstractMotion time series collected from low-power, always-on mobile and wearable devices such as smartphones and smartwatches offer significant insights into human behavioral patterns, with wide applications in healthcare, automation, IoT, and AR/XR. However, given security and privacy concerns, building large-scale motion time series datasets remains difficult, hindering the development of pre-trained models for human activity analysis. Typically, existing models are trained and tested on the same dataset, leading to poor generalizability across variations in device location, device mounting orientation, and human activity type. In this paper, we introduce UniMTS, the first unified pre-training procedure for motion time series that generalizes across diverse device latent factors and activities. Specifically, we employ a contrastive learning framework that aligns motion time series with text descriptions enriched by large language models. This helps the model learn the semantics of time series to generalize across activities. Given the absence of large-scale motion time series data, we derive and synthesize time series from existing motion skeleton data with all-joint coverage. We use spatio-temporal graph networks to capture the relationships across joints for generalization across different device locations. We further design rotation-invariant augmentation to make the model agnostic to changes in device mounting orientations. Our model shows exceptional generalizability across 18 motion time series classification benchmark datasets, outperforming the best baselines by 340% in the zero-shot setting, 16.3% in the few-shot setting, and 9.2% in the full-shot setting. Xiyuan Zhang 0001, Diyan Teng, Ranak Roy Chowdhury, Shuheng Li, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
NeurIPS | 6 |
| 2024 | How Few Davids Improve One Goliath: Federated Learning in Resource-Skewed Edge Computing EnvironmentsabstractReal-world deployment of federated learning requires orchestrating clients with widely varied compute resources, from strong enterprise-grade devices in data centers to weak mobile and Web-of-Things devices. Prior works have attempted to downscale large models for weak devices and aggregate shared parts among heterogeneous models. A typical architectural assumption is that there are equally many strong and weak devices. In reality, however, we often encounter resource skew where a few (1 or 2) strong devices hold substantial data resources, alongside many weak devices. This poses challenges-the unshared portion of the large model rarely receives updates or gains benefits from weak collaborators. Jiayun Zhang, Shuheng Li, Haiyu Huang 0003, Zihan Wang 0001, Xiaohan Fu, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
WWW | 7 |
| 2023 | PrimeNet: Pre-training for Irregular Multivariate Time SeriesabstractReal-world applications often involve irregular time series, for which the time intervals between successive observations are non-uniform. Irregularity across multiple features in a multi-variate time series further results in a different subset of features at any given time (i.e., asynchronicity). Existing pre-training schemes for time-series, however, often assume regularity of time series and make no special treatment of irregularity. We argue that such irregularity offers insight about domain property of the data—for example, frequency of hospital visits may signal patient health condition—that can guide representation learning. In this work, we propose PrimeNet to learn a self-supervised representation for irregular multivariate time-series. Specifically, we design a time sensitive contrastive learning and data reconstruction task to pre-train a model. Irregular time-series exhibits considerable variations in sampling density over time. Hence, our triplet generation strategy follows the density of the original data points, preserving its native irregularity. Moreover, the sampling density variation over time makes data reconstruction difficult for different regions. Therefore, we design a data masking technique that always masks a constant time duration to accommodate reconstruction for regions of different sampling density. We learn with these tasks using unlabeled data to build a pre-trained model and fine-tune on a downstream task with limited labeled data, in contrast with existing fully supervised approach for irregular time-series, requiring large amounts of labeled data. Experiment results show that PrimeNet significantly outperforms state-of-the-art methods on naturally irregular and asynchronous data from Healthcare and IoT applications for several downstream tasks, including classification, interpolation, and regression. Ranak Roy Chowdhury, Jiacheng Li 0003, Xiyuan Zhang 0001, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
AAAI | 5 |
| 2023 | Unleashing the Power of Shared Label Structures for Human Activity RecognitionabstractCurrent human activity recognition (HAR) techniques regard activity labels as integer class IDs without explicitly modeling the semantics of class labels. We observe that different activity names often have shared structures. For example, "open door" and "open fridge" both have "open" as the action; "kicking soccer ball" and "playing tennis ball" both have "ball" as the object. Such shared structures in label names can be translated to the similarity in sensory data and modeling common structures would help uncover knowledge across different activities, especially for activities with limited samples. In this paper, we propose SHARE, a HAR framework that takes into account shared structures of label names for different activities. To exploit the shared structures, SHARE comprises an encoder for extracting features from input sensory time series and a decoder for generating label names as a token sequence. We also propose three label augmentation techniques to help the model more effectively capture semantic structures across activities, including a basic token-level augmentation, and two enhanced embedding-level and sequence-level augmentations utilizing the capabilities of pre-trained models. SHARE outperforms state-of-the-art HAR models in extensive experiments on seven HAR benchmark datasets. We also evaluate in few-shot learning and label imbalance settings and observe even more significant performance gap. Xiyuan Zhang 0001, Ranak Roy Chowdhury, Jiayun Zhang, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
CIKM | 5 |
| 2023 | Towards Diverse and Coherent Augmentation for Time-Series ForecastingabstractTime-series data augmentation mitigates the issue of insufficient training data for deep learning models. Yet, existing augmentation methods are mainly designed for classification, where class labels can be preserved even if augmentation alters the temporal dynamics. We note that augmentation designed for forecasting requires diversity as well as coherence with the original temporal dynamics. As time-series data generated by real-life physical processes exhibit characteristics in both the time and frequency domains, we propose to combine Spectral and Time Augmentation (STAug) for generating more diverse and coherent samples. Specifically, in the frequency domain, we use the Empirical Mode Decomposition to decompose a time series and reassemble the subcomponents with random weights. This way, we generate diverse samples while being coherent with the original temporal relationships as they contain the same set of base components. In the time domain, we adapt a mix-up strategy that generates diverse as well as linearly in-between coherent samples. Experiments on five real-world time-series datasets demonstrate that STAug outperforms the base models without data augmentation as well as state-of-the-art augmentation methods. Xiyuan Zhang 0001, Ranak Roy Chowdhury, Jingbo Shang, Rajesh K. Gupta 0001, Dezhi Hong |
ICASSP | 4 |
| 2023 | Minimally Supervised Contextual Inference from Human Mobility: An Iterative Collaborative Distillation FrameworkabstractThe context about trips and users from mobility data is valuable for mobile service providers to understand their customers and improve their services. Existing inference methods require a large number of labels for training, which is hard to meet in practice. In this paper, we study a more practical yet challenging setting—contextual inference using mobility data with minimal supervision (i.e., a few labels per class and massive unlabeled data). A typical solution is to apply semi-supervised methods that follow a self-training framework to bootstrap a model based on all features. However, using a limited labeled set brings high risk of overfitting to self-training, leading to unsatisfactory performance. We propose a novel collaborative distillation framework STCOLAB. It sequentially trains spatial and temporal modules at each iteration following the supervision of ground-truth labels. In addition, it distills knowledge to the module being trained using the logits produced by the latest trained module of the other modality, thereby mutually calibrating the two modules and combining the knowledge from both modalities. Extensive experiments on two real-world datasets show STCOLAB achieves significantly more accurate contextual inference than various baselines. Jiayun Zhang, Xinyang Zhang 0002, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
IJCAI | 4 |
| 2023 | Navigating Alignment for Non-identical Client Class Sets: A Label Name-Anchored Federated Learning FrameworkabstractTraditional federated classification methods, even those designed for non-IID clients, assume that each client annotates its local data with respect to the same universal class set. In this paper, we focus on a more general yet practical setting, non-identical client class sets, where clients focus on their own (different or even non-overlapping) class sets and seek a global model that works for the union of these classes. If one views classification as finding the best match between representations produced by data/label encoder, such heterogeneity in client class sets poses a new significant challenge-local encoders at different clients may operate in different and even independent latent spaces, making it hard to aggregate at the server. We propose a novel framework, FedAlign1, to align the latent spaces across clients from both label and data perspectives. From a label perspective, we leverage the expressive natural language class names as a common ground for label encoders to anchor class representations and guide the data encoder learning across clients. From a data perspective, during local training, we regard the global class representations as anchors and leverage the data points that are close/far enough to the anchors of locally-unaware classes to align the data encoders across clients. Our theoretical analysis of the generalization performance and extensive experiments on four real-world datasets of different tasks confirm that FedAlign outperforms various state-of-the-art (non-IID) federated classification methods. Jiayun Zhang, Xiyuan Zhang 0001, Xinyang Zhang 0002, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
KDD | 5 |
| 2023 | Physics-Informed Data Denoising for Real-Life Sensing SystemsabstractSensors measuring real-life physical processes are ubiquitous in today's interconnected world. These sensors inherently bear noise that often adversely affects the performance and reliability of the systems they support. Classic filtering approaches introduce strong assumption on the time or frequency characteristics of sensory measurements, while learning-based denoising approaches typically rely on using ground truth clean data to train a denoising model, which is often challenging or prohibitive to obtain for many real-world applications. We observe that in many scenarios, the relationships between different sensor measurements (e.g., location and acceleration) are analytically described by laws of physics (e.g., second-order differential equation). By incorporating such physics constraints, we can guide the denoising process to improve performance even in the absence of ground truth data. In light of this, we design a physics-informed denoising model that leverages the inherent algebraic relationships between different measurements governed by the underlying physics. By obviating the need for ground truth clean data, our method offers a practical denoising solution for real-world applications. We conducted experiments in various domains, including inertial navigation, CO2 monitoring, and HVAC control, and achieved state-of-the-art performance compared with existing denoising methods. Our method can denoise data in real time (4ms for a sequence of 1s) for low-cost noisy sensors and produces results that closely align with those from high-precision, high-cost alternatives, leading to an efficient, cost-effective approach for more accurate sensor-based systems. Xiyuan Zhang 0001, Xiaohan Fu, Diyan Teng, Chengyu Dong, Keerthivasan Vijayakumar, Jiayun Zhang, Ranak Roy Chowdhury, Junsheng Han, Dezhi Hong, Rashmi Kulkarni, Jingbo Shang, Rajesh K. Gupta 0001 |
SenSys | 12 |
| 2022 | TARNet: Task-Aware Reconstruction for Time-Series TransformerabstractTime-series data contains temporal order information that can guide representation learning for predictive end tasks (e.g., classification, regression). Recently, there are some attempts to leverage such order information to first pre-train time-series models by reconstructing time-series values of randomly masked time segments, followed by an end-task fine-tuning on the same dataset, demonstrating improved end-task performance. However, this learning paradigm decouples data reconstruction from the end task. We argue that the representations learnt in this way are not informed by the end task and may, therefore, be sub-optimal for the end-task performance. In fact, the importance of different timestamps can vary significantly in different end tasks. We believe that representations learnt by reconstructing important timestamps would be a better strategy for improving end-task performance. In this work, we propose TARNet, Task-Aware Reconstruction Network, a new model using Transformers to learn task-aware data reconstruction that augments end-task performance. Specifically, we design a data-driven masking strategy that uses self-attention score distribution from end-task training to sample timestamps deemed important by the end task. Then, we mask out data at those timestamps and reconstruct them, thereby making the reconstruction task-aware. This reconstruction task is trained alternately with the end task at every epoch, sharing parameters in a single model, allowing the representation learnt through reconstruction to improve end-task performance. Extensive experiments on tens of classification and regression datasets show that TARNet significantly outperforms state-of-the-art baseline models across all evaluation metrics. Ranak Roy Chowdhury, Xiyuan Zhang 0001, Jingbo Shang, Rajesh K. Gupta 0001, Dezhi Hong |
KDD | 4 |
| 2022 | SQEE: A Machine Perception Approach to Sensing Quality Evaluation at the Edge by Uncertainty QuantificationabstractCyber-physical systems are starting to adopt neural network (NN) models for a variety of smart sensing applications. While several efforts seek better NN architectures for system performance improvement, few attempts have been made to study the deployment of these systems in the field. Proper deployment of these systems is critical to achieving ideal performance, but the current practice is largely empirical via trials and errors, lacking a measure of quality. Sensing quality should reflect the impact on the performance of NN models that drive machine perception tasks. However, traditional approaches either evaluate statistical difference that exists objectively, or model the quality subjectively via human perception. Shuheng Li, Jingbo Shang, Rajesh K. Gupta 0001, Dezhi Hong |
SenSys | 3 |
| 2022 | ESC-GAN: Extending Spatial Coverage of Physical SensorsabstractScientific discoveries and studies about our physical world have long benefited from large-scale and planetary sensing, from weather forecasting to wildfire monitoring. However, the limited deployment of sensors in the environment due to cost or physical access constraints has lagged behind our ever-growing need for increased data coverage and higher resolution, impeding timely and precise monitoring and understanding of the environment. Therefore, we seek to extend the spatial coverage of analysis based on existing sensory data, that is, to "generate" data for locations where no historical data exists. This problem is fundamentally different and more challenging than the traditional spatio-temporal imputation that assumes data for any particular location are only partially missing across time. Inspired by the success of Generative Adversarial Network (GAN) in imputation, we propose a novel ESC-GAN. We observe that there are local patterns in nearby locations, as well as trends in a global manner (e.g., temperature drops as altitude increases regardless of the location). As local patterns may exhibit at different scales (from meters to kilometers), we employ a multi-branch generator to aggregate information of different granularity. More specifically, each branch in the generator contains 1) randomly masked 3D partial convolutions at different resolutions to capture the local patterns and 2) global attention modules for global similarity. Next, we adversarially train a 3D convolution-based discriminator to distinguish the generator's output from the ground truth. Extensive experiments on three geo-sensor datasets demonstrate that ESC-GAN outperforms state-of-the-art methods on extending spatial coverage and also achieves the best results on a traditional spatio-temporal imputation task. Xiyuan Zhang 0001, Ranak Roy Chowdhury, Jingbo Shang, Rajesh K. Gupta 0001, Dezhi Hong |
WSDM | 4 |
| 2022 | Performance Analysis of Timing-Speculative ProcessorsabstractWe propose a framework to estimate the number of timing errors experienced by an application as it runs on a timing-speculative processor. It takes a hybrid approach combining an accurate gate-level dynamic timing analysis engine to find timing errors in the processor’s control network with a fast architecture-level execution-driven simulator based on a path activation model of the datapath. We develop an instruction-level error model that estimates the likelihood of an instruction experiencing a timing error, capturing the effects of process and data variations as well as inter-instruction correlations caused by the error recovery scheme used by the processor. Finally, we utilize two well-known laws of applied statistics, the law of small numbers and the law of large numbers, to estimate, with bounded inaccuracy, the total number of timing errors experienced by a specific application. Our experiments show that the combination of running application and its input data can change performance by as much as 25 percent, demonstrating that application-specific analysis is necessary for accurate evaluation of timing-speculative processors and should be used to inform design decisions and assess the suitability of applications for timing speculation. Omid Assare, Rajesh K. Gupta 0001 |
IEEE Trans. Computers | 2 |
| 2021 | Associative Convolutional LayersabstractWe provide a general and easy to implement method for reducing the number of parameters of Convolutional Neural Networks (CNNs) during the training and inference phases. We introduce a simple trainable auxiliary neural network which can generate approximate versions of “slices” of the sets of convolutional filters of any CNN architecture from a low dimensional “code” space. These slices are then concatenated to form the sets of filters in the CNN architecture. The auxiliary neural network, which we call “Convolutional Slice Generator” (CSG), is unique to the network and provides the association among its convolutional layers. We apply our method to various CNN architectures including ResNet, DenseNet, MobileNet and ShuffleNet. Experiments on CIFAR-10 and ImageNet-1000, without any hyper-parameter tuning, show that our approach reduces the network parameters by approximately $2\times$ while the reduction in accuracy is confined to within one percent and sometimes the accuracy even improves after compression. Interestingly, through our experiments, we show that even when the CSG takes random binary values for its weights that are not learned, still acceptable performances are achieved. To show that our approach generalizes to other tasks, we apply it to an image segmentation architecture, Deeplab V3, on the Pascal VOC 2012 dataset. Results show that without any parameter tuning, there is $\approx 2.3\times$ parameter reduction and the mean Intersection over Union (mIoU) drops by $\approx 3%$. Finally, we provide comparisons with several related methods showing the superiority of our method in terms of accuracy. Hamed Omidvar, Vahideh Akhlaghi, Massimo Franceschetti, Rajesh K. Gupta 0001 |
AISTATS | 5 |
| 2021 | UniTS: Short-Time Fourier Inspired Neural Networks for Sensory Time Series ClassificationabstractDiscovering patterns in time series data is essential to many key tasks in intelligent sensing systems, such as human activity recognition and event detection. These tasks involve the classification of sensory information from physical measurements such as inertial or temperature change measurements. Due to differences in the underlying physics, existing methods for classification use handcrafted features combined with traditional learning algorithms, or employ distinct deep neural models to directly learn from raw data. Shuheng Li, Ranak Roy Chowdhury, Jingbo Shang, Rajesh K. Gupta 0001, Dezhi Hong |
SenSys | 4 |
| 2021 | EditorialabstractAfter years of selfless service, Prof. Xin Li has stepped down from the Deputy EIC role at IEEE TCAD in order to focus on his recent appointment as Dean of Graduate Studies at his school. It is with mixed feeling I have to convey this news: Li has been the key operational leader and a direct contact to the editorial board members for more than four years. In the process, he has carried enormous institutional memory and offered guidance on measures that we have devised and implemented over the years. I have relied on this guidance as a crucial input into making substantial changes in TCAD policies and processes. Yet, it is a well-deserved elevation, and our heartiest congratulations to Prof. Li. Rajesh K. Gupta 0001, David Atienza 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Ember - energy management of batteryless event detection sensors with deep reinforcement learning: demo abstractabstractBatteryless sensors avoid battery replacement at the cost of slowing down or stopping their operations when there is not sufficient energy to harvest in the environment. While this strategy can work for some applications, event-based applications still remain a challenge as events arrive sporadically and energy availability is uncertain. One solution is to only turn On a sensor right before an event is happening to both detect the event and save as much energy as possible. Therefore, the system has to correctly predict events while managing limited resource availability. In this demo, we present Ember, an energy management system based on deep reinforcement learning to duty cycle event-driven sensors in low-energy conditions. We show how our system learns environmental patterns over time and makes decisions to maximize the event detection rate for batteryless energy-harvesting sensor nodes subject to low energy availability. Furthermore, we show a novel self-supervised data collection algorithm that helps Ember in discovering new environmental patterns over time. For more details, we refer readers to the full paper of Ember [2]. Francesco Fraternali, Bharathan Balaji, Michael Barrow, Dezhi Hong, Rajesh K. Gupta 0001 |
SenSys | 5 |
| 2020 | Ember: energy management of batteryless event detection sensors with deep reinforcement learningabstractEnergy management can extend the lifetime of batteryless, energy-harvesting systems by judiciously utilizing the energy available. Duty cycling of such systems is especially challenging for event detection, as events arrive sporadically and energy availability is uncertain. If the node sleeps too much, it may miss important events; if it depletes energy too quickly, it will stop operating in low energy conditions and miss events. Thus, accurate event prediction is important in making this tradeoff. We propose Ember, an energy management system based on deep reinforcement learning to duty cycle event-driven sensors in low energy conditions. We train a policy using historical real-world data traces of motion, temperature, humidity, pressure, and light events. The resulting policy can learn to capture up to 95% of the events without depleting the node. Without historical data for training when deploying a node at a new location, we propose a self-supervised mechanism to collect ground-truth data while learning from the data at the same time. Ember learns to capture the majority of events within a week without any historical data and matches the performance of the policies trained with historical data in a few weeks. We deployed 40 nodes running Ember for indoor sensing and demonstrate that the learned policies generalize to real-world settings as well as outperform state-of-the-art techniques. Francesco Fraternali, Bharathan Balaji, Dhiman Sengupta, Dezhi Hong, Rajesh K. Gupta 0001 |
SenSys | 5 |
| 2020 | Local Binary Pattern NetworksabstractEmerging edge devices such as sensor nodes are increasingly being tasked with non-trivial tasks related to sensor data processing and even application-level inferences from this sensor data. These devices are, however, extraordinarily resource-constrained in terms of CPU power (often Cortex M0-3 class CPUs), available memory (in few KB to MBytes), and energy. Under these constraints, we explore a novel approach to character recognition using local binary pattern networks, or LBPNet, that can learn and perform bit-wise operations in an end-to-end fashion. LBPNet has its advantage for characters whose features are composed of structured strokes and distinctive outlines. LBPNet uses local binary comparisons and random projections in place of conventional convolution (or approximation of convolution) operations, providing an important means to improve memory efficiency as well as inference speed. We evaluate LBPNet on a number of character recognition benchmark datasets as well as several object classification datasets and demonstrate its effectiveness and efficiency. Jeng-Hau Lin, Justin Lazarow, Yunfan Yang, Dezhi Hong, Rajesh K. Gupta 0001, Zhuowen Tu |
WACV | 5 |
| 2020 | ACES: Automatic Configuration of Energy Harvesting Sensors with Reinforcement LearningabstractMany modern smart building applications are supported by wireless sensors to sense physical parameters, given the flexibility they offer and the reduced cost of deployment. However, most wireless sensors are powered by batteries today, and large deployments are inhibited by the requirement of periodic battery replacement. Energy harvesting sensors provide an attractive alternative, but they need to provide adequate quality of service to applications given the uncertainty of energy availability. We propose ACES, which uses reinforcement learning to maximize sensing quality of energy harvesting sensors for periodic and event-driven indoor sensing with available energy. Our custom-built sensor platform uses a supercapacitor to store energy and Bluetooth Low Energy to relay sensors data. Using simulations and real deployments, we use the data collected to continually adapt the sensing of each node to changing environmental patterns and transfer learning to reduce the training time in real deployments. In our 60-node deployment lasting 2 weeks, nodes stop operations only 0.1% of the time, and collection of data is comparable with current battery-powered nodes. We show that ACES reduces the node duty-cycle period by an average of 33% compared to three prior reinforcement learning techniques while continuously learning environmental changes over time. Francesco Fraternali, Bharathan Balaji, Yuvraj Agarwal, Rajesh K. Gupta 0001 |
ACM Trans. Sens. Networks | 4 |
| 2019 | Accurate Estimation of Program Error Rate for Timing-Speculative ProcessorsabstractWe propose a framework that estimates the error rate experienced by an application as it runs on a timing-speculative processor. The framework uses an instruction error model that is comparable in accuracy to low-level simulations---as it considers the effects of operand values, preceding instructions, datapath configuration, and error correction scheme, as well as process variation, including its spatial correlation property---and yet efficient enough to allow its application in Monte Carlo experiments to characterize large program input datasets. We then use statistical limit theorems to estimate program error rate and quantify the effect of inter-instruction correlations. Omid Assare, Rajesh K. Gupta 0001 |
DAC | 2 |
| 2019 | Accelerating Local Binary Pattern Networks with Software-Programmable FPGAsabstractFueled by the success of mobile devices, the computational demands on these platforms have been rising faster than the computational and storage capacities or energy availability to perform tasks ranging from recognizing speech, images to automated reasoning and cognition. While the success of convolutional neural networks (CNNs) have contributed to such a vision, these algorithms stay out of the reach of limited computing and storage capabilities of mobile platforms. It is clear to most researchers that such a transition can only be achieved by using dedicated hardware accelerators on these platforms. However, CNNs with arithmetic-intensive operations remain particularly unsuitable for such acceleration both computationally as well as for the high memory bandwidth needs of highly parallel processing required. In this paper, we implement and optimize an alternative genre of networks, local binary pattern network (LBPNet) which eliminates arithmetic operations by combinatorial operations thus substantially boosting the efficiency of hardware implementation. LBPNet is built upon a radically different view of the arithmetic operations sought by conventional neural networks to overcome limitations posed by compression and quantization methods used for hardware implementation of CNNs. This paper explores in depth the design and implementation of both an architecture and critical optimizations of LBPNet for realization in accelerator hardware and provides a comparison of results with the state-of-art CNN on multiple datasets. Jeng-Hau Lin, Atieh Lotfi, Vahideh Akhlaghi, Zhuowen Tu, Rajesh K. Gupta 0001 |
DATE | 5 |
| 2019 | Towards verified programming of embedded devicesabstractWe propose a type-driven approach to building verified safe and correct IoT applications. Today's IoT applications are plagued with bugs that can cause physical damage. This is largely because developers account for physical constraints using ad-hoc techniques. Accounting for such constrains in a more principled fashion demands reasoning about the composition of all the software and hardware components of the application. Our proposed framework takes a step in this direction by (1) using refinement types to make make physical constraints explicit and (2) imposing an event-driven programing discipline to simplify the reasoning of system-wide properties to that of an event queue. In taking this approach, our framework makes it possible for developers to build verified IoT application by making it a type error for code to violate physical constraints. Jean-Pierre Talpin, Jean-Joseph Marty, Shravan Narayan, Deian Stefan, Rajesh K. Gupta 0001 |
DATE | 5 |
| 2019 | Real Time Principal Component AnalysisabstractBy processing the data in motion, real-time data processing enables us to extract instantaneous results from online input data that ensures timely responsiveness to events as well as a much enhanced capacity to process large data sets. This is especially important when decision loops include querying and processing data on the web where size and latency considerations make it impossible to process raw data in real-time. This makes dimensionality reduction techniques, like principal component analysis (PCA), an important data preprocessing tool to gain insights into data. In this paper, we propose a variant of PCA, that is suited for real-time applications. In the real-time version of the PCA problem, we maintain a window over the most recent data and project every incoming row of data into lower dimensional subspace, which we generate as the output of the model. The goal is to minimize the reconstruction error of the output from the input. We use the reconstruction error as the termination criteria to update the eigenspace as new data arrives. To verify whether our proposed model can capture the essence of the changing distribution of large datasets in real-time, we have implemented the algorithm and evaluated performance against carefully designed simulations that change distributions of data sources over time in a controllable manner. Furthermore, we have demonstrated that our algorithm can capture the changing distributions of real-life datasets by running simulations on datasets from a variety of real-time applications e.g. localization, customer expenditure, etc. We propose algorithmic enhancements that rely upon spectral analysis to improve dimensionality reduction. Results show that our method can successfully capture the changing distribution of data in a real-time scenario, thus enabling real-time PCA. Ranak Roy Chowdhury, Muhammad Abdullah Adnan, Rajesh K. Gupta 0001 |
ICDE | 3 |
| 2019 | New models and methods for programming cyber-physical systems (keynote)abstractEmerging cyber-physical systems are distributed systems in constant interaction with their physical environments through sensing and actuation at network edges. Over the past decade, the embedded and control systems community have vigorously pursued a vision of coupled feedback-controlled systems with a broad range of real-life applications from transportation, smart buildings to human health. These efforts have continued to push intelligent processing to edge and near-edge devices, provide new capabilities for improved sensing with high quality timing information, establish limits on the quality of time and its impact on the stability of control algorithms etc. Rajesh K. Gupta 0001, Jason Koh, Dezhi Hong |
LCTES | 1 |
| 2019 | Serving deep neural networks at the cloud edge for vision applications on mobile platformsabstractThe proliferation of high resolution cameras on embedded devices along with the growing maturity of deep neural networks (DNNs) has spawned powerful mobile vision applications. To enable applications on mobile devices, the offloading approach processes live video streams using DNNs on server-class GPU accelerators. However, their use in latency constrained applications is particularly challenging because of the large and unpredictable round-trip latency from mobile devices to the cloud computing resources. As a consequence, system designers routinely look for ways to offload to local servers at the cloud edge, known as the cloudlet. This paper explores the potential of serving multiple DNNs using the cloudlet model to implement complex vision applications on mobile devices. We present DeepQuery, a new mobile offloading system that is capable to serve DNNs with different structures for a wide range of tasks including object detection and tracking, scene graph detection, and video description. DeepQuery provides application programming interfaces to offload applications programed as Directed Acyclic Graphs of DNN queries, and employs data parallelization and input batching techniques to reduce processing delays. To improve GPU utilization, it co-locates real-time and delay-tolerant tasks on shared GPUs, and exploits a predictive and plan-ahead approach to alleviate resource contention caused by co-locating. We evaluate DeepQuery and demonstrate its effectiveness using several real world applications. Dezhi Hong, Rajesh K. Gupta 0001 |
MMSys | 3 |
| 2018 | LEMAX: learning-based energy consumption minimization in approximate computing with quality guaranteeabstractApproximate computing aims to trade accuracy for energy efficiency. Various approximate methods have been proposed in the literature that demonstrate the effectiveness of relaxing accuracy requirements in a specific unit. This provides a basis for exploring simultaneous use of multiple approximate units to improve efficiency under guarantees on quality of results. In this paper, we explore the effect of combining multiple approximate units on the energy consumption and identify the best setting that minimizes energy consumption under a quality constraint. Our approach also enables changes in unit configurations throughout the program. To do this effectively, we need a method to examine the combined impact of multiple approximate units on the output quality, and configure individual units accordingly. To solve this problem, we propose LEMAX that uses gradient descent approach to identify the best configuration of the individual approximate units for a given program. We evaluate the efficacy of LEMAX in minimizing the energy consumption of several machine learning applications with varying size (i.e., number of operations) under different quality constraints. Our evaluation shows that the configuration provided by LEMAX for a system with multiple approximate units improves the energy consumption by on average, 97.7%, 83.12%, and 73.95% for quality loss of 5%, 2% and 0.5%, respectively, compared to configurations obtained for a system with a single approximate resource. Vahideh Akhlaghi, Sicun Gao, Rajesh K. Gupta 0001 |
DAC | 3 |
| 2018 | Energy-efficient neural networks using approximate computation reuseabstractAs a problem-solving method, neural networks have shown broad success for medical applications, speech recognition, and natural language processing. Current hardware implementations of neural networks exhibit high energy consumption due to the intensive computing workloads. This paper proposes a methodology to design an energy-efficient neural network that effectively exploits computation reuse opportunities. To do so, we use Bloom filters (BFs) by tightly integrating them with computation units. BFs store and recall frequently occurring input patterns to reuse computations. We expand the opportunities for computation reuse by storing frequent input patterns specific to a given layer and using approximate pattern matching with hashing for limited data precision. This reconfigurable matching is key to achieving a “controllable approximation” for neural networks. To lower the energy consumption of BFs, we also use low-pow memristor arrays to implement BFs. Our experimental results show that for convolutional neural networks, the BFs enable 47.5% energy saving of multiplication operations, while incurring only 1% accuracy drop. While the actual savings will vary depending upon the extent of approximation and reuse, this paper presents a method for reducing computing workloads and improving energy efficiency. Xun Jiao 0002, Vahideh Akhlaghi, Yu Jiang 0001, Rajesh K. Gupta 0001 |
DATE | 4 |
| 2018 | Embedded software for robotics: challenges and future directions: special sessionabstractThis paper surveys recent challenges and solutions in the design, implementation, and verification of embedded software for robotics. Emphasis is placed on mobile robots, like self-driving cars. In design, it addresses programming support for robotic systems, secure state estimation, and ROS-based monitor generation. In the implementation phase, it describes the synthesis of control software using finite precision arithmetic, real-time platforms and architectures for safety-critical robotics, efficient implementation of neural network based-controllers, and standards for computer vision applications. The issues in verification include verification of neural network-based robotic controllers, and falsification of closed-loop control systems. The paper also describes notable open-source robotic platforms. Along the way, we highlight important research problems for developing the next generation of high-performance, low-resource-usage, correct embedded software. Houssam Abbas, Indranil Saha 0001, Yasser Shoukry, Rüdiger Ehlers, Georgios Fainekos, Rajesh K. Gupta 0001, Rupak Majumdar, Dogan Ulus |
EMSOFT | 6 |
| 2018 | Reliability-Aware Data Placement for Heterogeneous Memory ArchitectureabstractSystem reliability is a first-class concern as technology continues to shrink, resulting in increased vulnerability to traditional sources of errors such as single event upsets. By tracking access counts and the Architectural Vulnerability Factor (AVF), application data can be partitioned into groups based on how frequently it is accessed (its "hotness") and its likelihood to cause program execution error (its "risk"). This is particularly useful for memory systems which exhibit heterogeneity in their performance and reliability such as Heterogeneous Memory Architectures – with a typical configuration combining slow, highly reliable memory with faster, less reliable memory. This work demonstrates that current state of the art, performance-focused data placement techniques affect reliability adversely. It shows that page risk is not necessarily correlated with its hotness; this makes it possible to identify pages that are both hot and low risk, enabling page placement strategies that can find a good balance of performance and reliability. This work explores heuristics to identify and monitor both hotness and risk at run-time, and further proposes static, dynamic, and program annotation-based reliability-aware data placement techniques. This enables an architect to choose among available memories with diverse performance and reliability characteristics. The proposed heuristic-based reliability-aware data placement improves reliability by a factor of 1.6x compared to performance-focused static placement while limiting the performance degradation to 1%. A dynamic reliability-aware migration scheme, which does not require prior knowledge about the application, improves reliability by a factor of 1.5x on average while limiting the performance loss to 4.9%. Finally, program annotation-based data placement improves the reliability by 1.3x at a performance cost of 1.1%. Manish Gupta 0010, Vilas Sridharan, Andreas Prodromou, Ashish Venkat, Dean M. Tullsen, Rajesh K. Gupta 0001 |
HPCA | 7 |
| 2018 | SnaPEA: Predictive Early Activation for Reducing Computation in Deep Convolutional Neural NetworksabstractDeep Convolutional Neural Networks (CNNs) perform billions of operations for classifying a single input. To reduce these computations, this paper offers a solution that leverages a combination of runtime information and the algorithmic structure of CNNs. Specifically, in numerous modern CNNs, the outputs of compute-heavy convolution operations are fed to activation units that output zero if their input is negative. By exploiting this unique algorithmic property, we propose a predictive early activation technique, dubbed SnaPEA. This technique cuts the computation of convolution operations short if it determines that the output will be negative. SnaPEA can operate in two distinct modes, exact and predictive. In the exact mode, with no loss in classification accuracy, SnaPEA statically re-orders the weights based on their signs and periodically performs a single-bit sign check on the partial sum. Once the partial sum drops below zero, the rest of computations can simply be ignored, since the output value will be zero in any case. In the predictive mode, which trades the classification accuracy for larger savings, SnaPEA speculatively cuts the computation short even earlier than the exact mode. To control the accuracy, we develop a multi-variable optimization algorithm that thresholds the degree of speculation. As such, the proposed algorithm exposes a knob to gracefully navigate the trade-offs between the classification accuracy and computation reduction. Compared to a state-of-the-art CNN accelerator, SnaPEA in the exact mode, yields, on average, 28% speedup and 16% energy reduction in various modern CNNs without affecting their classification accuracy. With 3% loss in classification accuracy, on average, 67.8% of the convolutional layers can operate in the predictive mode. The average speedup and energy saving of these layers are 2.02x and 1.89x, respectively. The benefits grow to a maximum of 3.59x speedup and 3.14x energy reduction. Compared to static pruning approaches, which are complimentary to the dynamic approach of SnaPEA, our proposed technique offers up to 63% speedup and 49% energy reduction across the convolution layers with no loss in classification accuracy. Vahideh Akhlaghi, Amir Yazdanbakhsh, Kambiz Samadi, Rajesh K. Gupta 0001, Hadi Esmaeilzadeh |
ISCA | 4 |
| 2018 | Introducing Automatic Time Stamping (ATS) with a Reference Implementation in SwiftabstractThe need for associating a time with the arrival of data is prevalent in many applications but this is even more the case in Cyber Physical Systems (CPS) which measure quantities from the real world. One common attribute of any measured real world quantity is the time of when it was acquired. Automating the process of time stamping data upon arrival frees the programmer from having to deal with this task manually thus reducing the number of errors, shrinking the code size and making the code more readable and maintainable. This paper explores the concept of variables that are automatically time stamped by the runtime system and its impact in building real-time applications. Such time stamping support enables seamless integration of a time model that is kept updated by real-time events within well-defined synchronization time bounds. This paper discusses the design of the Automatic Time Stamping (ATS) framework and the prototype implementation of such variables using open source programming language Swift. Our results demonstrate the viability of the ATS framework for most real-time applications. In doing this work we have learned that the abstraction of time in software programming models still have much room for improvements from the operating system level all the way up to the programming languages and runtimes used by the developers. Sean Hamilton, Dhiman Sengupta, Rajesh K. Gupta 0001 |
ISORC | 3 |
| 2018 | CLIM: A Cross-Level Workload-Aware Timing Error Prediction Model for Functional UnitsabstractTiming errors that are caused by the timing violations of sensitized circuit paths, have emerged as an important threat to the reliability of synchronous digital circuits. To protect circuits from these timing errors, designers typically use a conservative timing margin, which leads to operational inefficiency. Existing adaptive approaches reduce such conservative margins by predicting the timing errors in advance and adjusting the timing margin adaptively. However, these error prediction approaches overlook the impact of input workload (i.e., operands) on path sensitization, thereby resulting in a loss of accuracy. The diversity of input operands leads to complex path sensitization behaviors, making them hard to represent in timing error modeling. In this paper, we propose CLIM, a cross-level workload-aware timing error prediction model for functional units (FUs). CLIM predicts whether there are timing errors in FU at two levels: bit-level and value-level. At the bit level or value level, CLIM predicts each output bit or entire output value as one of two classes: {timing correct, timing erroneous} as a function of input workload and clock period, respectively. We apply supervised learning methods to construct CLIM, by using input operands, computation history and circuit toggling as input features, as well as outputs' timing classes as labels. These training data are collected from gate-level simulations (GLS) of post place-and-route designs in TSMC 45nm process. We evaluate CLIM prediction accuracy for various FUs and compare it with baseline models. On average, CLIM exhibits 95 percent prediction accuracy at value-level, 97 percent at bit-level, and executes at a rate 173X faster than GLS. We utilize CLIM to analyze the value-level and bit-level reliability of FUs under random and real-world application workloads. At value-level, CLIM-based reliability estimation is within 2.8 percent deviation on average of detailed GLS ground truth. At bit-level, we introduce the concept of bit-level reliability specification of error-tolerant applications and compare this with the CLIM-based bit-level reliability estimation. By comparison, CLIM will classify the application quality into two classes: {acceptable, non-acceptable}. On average, 97 percent application quality classification is consistent with GLS ground truth. Xun Jiao 0002, Abbas Rahimi, Yu Jiang 0001, Jianguo Wang 0001, Hamed Fatemi, José Pineda de Gyvez, Rajesh K. Gupta 0001 |
IEEE Trans. Computers | 7 |
| 2018 | EditorialabstractAt its core, a technical publication represents a community or a community of communities that share an interest in solving commonly understood technical problems, often with an understanding of the nature of methods that must be invented. As communities go, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) represents a very diverse community spread across a large number of technical areas spanning very large scale integration (VLSI), CAD, circuits, embedded systems, formal methods, etc., and internationally spanning nearly all regions of the IEEE. Rajesh K. Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | BuildingRules: A Trigger-Action-Based System to Manage Complex Commercial BuildingsabstractModern Building Management Systems (BMSs) have been designed to automate the behavior of complex buildings, but unfortunately they do not allow occupants to customize it according to their preferences, and only the facility manager is in charge of setting the building policies. To overcome this limitation, we present BuildingRules, a trigger-action programming-based system that aims to provide occupants of commercial buildings with the possibility of specifying the characteristics of their office environment through an intuitive interface. Trigger-action programming is intuitive to use and has been shown to be effective in meeting user requirements in home environments. To extend this intuitive interface to commercial buildings, an essential step is to manage the system scalability as large number of users will express their policies. BuildingRules has been designed to scale well for large commercial buildings as it automatically detects conflicts that occur among user specified policies and it supports intelligent grouping of rules to simplify the policies across large numbers of rooms. We ensure the conflict resolution is fast for a fluid user experience by using the Z3 SMT solver. BuildingRules backend is based on RESTful web services so it can connect to various BMSs and scale well with large number of buildings. We have tested our system with 23 users across 17 days in a virtual office building, and the results we have collected prove the effectiveness and the scalability of BuildingRules. A. A. Nacci, Vincenzo Rana, Bharathan Balaji, Paola Spoletini, Rajesh K. Gupta 0001, Donatella Sciuto, Yuvraj Agarwal |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2017 | Exploiting Synchrony in Replicated State MachinesabstractWe present Timestamp Order Preserving (TOP), a replicated state machine (RSM) protocol that exploits the synchrony of networks to provide high performance. TOP uses physical timestamp of synchronized clock as a consistent total order to achieve consensus. It keeps estimating the bounds of network latency and offset of synchronized clock to deduce the commit time for each operation. It adopts speculative processing and reconciliation techniques to improve performance. To demonstrate its advantages, we implement a key-value data store that uses TOP for data replication. Through evaluations in a geo-deployed testbed, by comparing it with Primary-Copy and Quorum-Replication protocols, we demonstrate that TOP has a similar commit latency with a higher sustainable throughput. In addition, it processes operations in the order of submission timestamp, which provides a stricter form of consistency. Mulong Luo, Mani Srivastava 0001, Rajesh K. Gupta 0001 |
CLOUD | 4 |
| 2017 | QoS-Aware Scheduling of Heterogeneous Servers for Inference in Deep Neural NetworksabstractDeep neural networks (DNNs) are popular in diverse fields such as computer vision and natural language processing. DNN inference tasks are emerging as a service provided by cloud computing environments. However, cloud-hosted DNN inference faces new challenges in workload scheduling for the best Quality of Service (QoS), due to dependence on batch size, model complexity and resource allocation. This paper represents the QoS metric as a utility function of response delay and inference accuracy. We first propose a simple and effective heuristic approach that keeps low response delay and satisfies the requirement on processing throughput. Then we describe an advanced deep reinforcement learning (RL) approach that learns to schedule from experience. The RL scheduler is trained to maximize QoS, using a set of system statuses as the input to the RL policy model. Our approach performs scheduling actions only when there are free GPUs, thus reduces scheduling overhead over common RL schedulers that run at every continuous time step. We evaluate the schedulers on a simulation platform and demonstrate the advantages of RL over heuristics. Tong Yu 0001, Ole J. Mengshoel, Rajesh K. Gupta 0001 |
CIKM | 4 |
| 2017 | Mitigating multi-tenant interference on mobile offloading servers: poster abstractabstractThis work considers that multiple mobile clients offload various continuous sensing applications with end-to-end delay constraints, to a cluster of machines as the server. Contention for shared computing resources on a server can result in delay degradation and application malfunction. We present ATOMS (Accurate Timing prediction and Offloading for Mobile Systems), a framework to mitigate multi-tenant resource contention and to improve delay using a two-phase Plan-Schedule approach. The planning phase includes methods to predict future workloads from all clients, to estimate contention, and to devise offloading schedule to reduce contention. The scheduling phase dispatches arriving offloaded workload to the server machine that minimizes contention, based on the running workloads on each machine. Mulong Luo, Tong Yu 0001, Ole J. Mengshoel, Mani Srivastava 0001, Rajesh K. Gupta 0001 |
SoCC | 6 |
| 2017 | Compiler Techniques to Reduce the Synchronization Overhead of GPU Redundant MultithreadingabstractRedundant Multi-Threading (RMT) provides a potentially low cost mechanism to increase GPU reliability by replicating computation at the thread level. Prior work has shown that RMT's high performance overhead stems not only from executing redundant threads, but also from the synchronization overhead between the original and redundant threads. The overhead of inter-thread synchronization can be especially significant if the synchronization is implemented using global memory. This work presents novel compiler techniques using fingerprinting and cross-lane operations to reduce synchronization overhead for RMT on GPUs. Fingerprinting combines multiple synchronization events into one event by hashing, and cross-lane operations enable thread-level synchronization via register-level communication. This work shows that fingerprinting yields a 73.5% reduction in GPU RMT overhead while cross-lane operations reduce the overhead by 43% when compared to the state-of-the-art GPU RMT solutions on real hardware. Manish Gupta 0010, Daniel Lowell, John Kalamatianos, Steven E. Raasch, Vilas Sridharan, Dean M. Tullsen, Rajesh K. Gupta 0001 |
DAC | 7 |
| 2017 | Combining structural and timing errors in overclocked inexact speculative addersabstractWorst-case design is used in IoT devices and high performance data centers to ensure reliability, leading to a power efficiency loss. Recently, approximate computing has been proposed to trade off accuracy for efficiency. In this paper, we use Inexact Speculative Adders, which redesign the adder architecture to shorten its critical path and improve performance, but introduces controlled structural errors. On the other hand, overclocking is used to reduce conservative timing guardbands but could normally introduce catastrophic timing errors, we thus apply a supervised learning model to overclock speculative adders and predict their timing errors. We build a methodology to combine both structural and timing errors and analyze how they interplay with each other to limit the overal errors. Xun Jiao 0002, Vincent Camus, Mattia Cacciotti, Yu Jiang 0001, Christian C. Enz, Rajesh K. Gupta 0001 |
DATE | 6 |
| 2017 | SLoT: A supervised learning model to predict dynamic timing errors of functional unitsabstractDynamic timing errors (DTEs), that are caused by the timing violations of sensitized critical timing paths, have emerged as an important threat to the reliability of digital circuits. Existing approaches model the DTEs without considering the impact of input operands on dynamic path sensitization, resulting in loss of accuracy. The diversity of input operands leads to complex path sensitization behaviors, making it hard to represent in DTE modeling. In this paper, we propose SLoT, a supervised learning model to predict the output of functional units (FUs) to be one of two timing classes: {timing correct, timing erroneous} as a function of input operands and clock period. We apply random forest classification (RFC) method to construct SLoT, by using input operands, computation history and circuit toggling as input features and outputs' timing classes as labels. The outputs' timing classes are measured using gate-level simulation (GLS) of a post place-and-route design in TSMC 45nm process. For evaluation, we apply SLoT to several FUs and on average 95% predictions are consistent with GLS, which is 6.3X higher compared to the existing instruction-level model. SLoT-based reliability analysis of FUs under different datasets can achieve 0.7-4.8% average difference compared with GLS-based analysis, and execute more than 20X faster than GLS. Xun Jiao 0002, Yu Jiang 0001, Abbas Rahimi, Rajesh K. Gupta 0001 |
DATE | 4 |
| 2017 | RxRE: Throughput Optimization for High-Level Synthesis using Resource-Aware Regularity Extraction (Abstract Only)
Atieh Lotfi, Rajesh K. Gupta 0001 |
FPGA | 2 |
| 2017 | Accelerating Binarized Convolutional Neural Networks with Software-Programmable FPGAs
Ritchie Zhao, Weinan Song, Tianwei Xing, Jeng-Hau Lin, Mani Srivastava 0001, Rajesh K. Gupta 0001, Zhiru Zhang |
FPGA | 7 |
| 2017 | An assessment of vulnerability of hardware neural networks to dynamic voltage and temperature variationsabstractAs a problem solving method, neural networks have shown broad applicability from medical applications, speech recognition, and natural language processing. This success has even led to implementation of neural network algorithms into hardware. In this paper, we explore two questions: (a) to what extent microelectronic variations affects the quality of results by neural networks; and (b) if the answer to first question represents an opportunity to optimize the implementation of neural network algorithms. Regarding first question, variations are now increasingly common in aggressive process nodes and typically manifest as an increased frequency of timing errors. Combating variations - due to process and/or operating conditions - usually results in increased guardbands in circuit and architectural design, thus reducing the gains from process technology advances. Given the inherent resilience of neural networks due to adaptation of their learning parameters, one would expect the quality of results produced by neural networks to be relatively insensitive to the rising timing error rates caused by increased variations. On the contrary, using two frequently used neural networks (MLP and CNN), our results show that variations can significantly affect the inference accuracy. This paper outlines our assessment methodology and use of a cross-layer evaluation approach that extracts hardware-level errors from twenty different operating conditions and then inject such errors back to the software layer in an attempt to answer the second question posed above. Xun Jiao 0002, Mulong Luo, Jeng-Hau Lin, Rajesh K. Gupta 0001 |
ICCAD | 4 |
| 2017 | ReHLS: Resource-Aware Program Transformation Workflow for High-Level SynthesisabstractDespite considerable improvements in existing HLS tools, they still require designer interventions to provide efficient synthesis results. This manual design space exploration and code rewriting and optimization takes significant time and negates the HLS design productivity gains. To overcome this challenge, this paper uses compiler frontend as an independent preprocessing step to explore the design space and adds an automated sourceto- source transformation step before HLS. In particular, it shows how inherent regularity in applications can be used to construct a workflow that analyzes the program, explores the design space for resource optimization opportunity, and transforms the program accordingly. When the transformed program is synthesized using the HLS tool, it uses less hardware resources with similar latency comparing to the original design. The synthesis results on a modern Xilinx Virtex-7 FPGA for a diverse set of applications show that our automated transformation can reduce the design area by an average of 15.4% with less than 1% performance overhead compared to the state-of-the-art Xilinx HLS tool solutions. This automated tool reduces the design time and especially can be useful for non-expert FPGA designers. Atieh Lotfi, Rajesh K. Gupta 0001 |
ICCD | 2 |
| 2017 | Data Hub Architecture for Smart CitiesabstractToday large amount of data is generated by cities. Many of the datasets are openly available and are contributed by different sectors, government bodies and institutions. The new data can affect our understanding of the issues faced by cities and can support evidence based policies. However usage of data is limited due to difficulty in assimilating data from different sources. Open datasets often lack uniform structure which limits its analysis using traditional database systems. In this paper we present Citadel, a data hub for cities. Citadel's goal is to support end to end knowledge discovery cyber-infrastructure for effective analysis and policy support. Citadel is designed to ingest large amount of heterogeneous data and supports multiple use cases by encouraging data sharing in cities. Our poster presents the proposed features, architecture, implementation details and initial results. Jason Koh, Sandeep Singh Sandha, Bharathan Balaji, Daniel Crawl, Ilkay Altintas, Rajesh K. Gupta 0001, Mani Srivastava 0001 |
SenSys | 6 |
| 2016 | Resistive Bloom filters: From approximate membership to approximate computing with bounded errors
Vahideh Akhlaghi, Abbas Rahimi, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2016 | Grater: An approximation workflow for exploiting data-level parallelism in FPGA acceleration
Atieh Lotfi, Abbas Rahimi, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, Rajesh K. Gupta 0001 |
DATE | 5 |
| 2016 | Strategies for optimal operating point selection in timing speculative processorsabstractPerformance of timing speculative processors relies on strategies for accurate prediction of optimal operating points. In this paper, we develop an efficient process-variation-aware simulation framework and use it to evaluate a range of such timing speculation strategies. Our experiments on a timing speculative processor running applications from the MiBench benchmark suite show that, in a typical case, while a perfect timing speculation strategy can improve throughput by up to 143% over a guardbanded design, the most commonly used approach in the literature achieves only a 21.8% of the potential gains. By improving the speculation accuracy, the new strategies we propose in this paper can realize up to 35.6% of the potential gains, a throughput improvement of 50.9% over a guardbanded design. Omid Assare, Rajesh K. Gupta 0001 |
ICCD | 2 |
| 2016 | WILD: A workload-based learning model to predict dynamic delay of functional unitsabstractDynamic critical path analysis in modern processors is needed to reduce margins typically determined by the static timing analysis. Dynamic path analysis, however, is cost-prohibitive. In this paper, we propose WILD, a supervised learning model to predict dynamic delay of functional units (FUs) based on the input workload during execution. We measure the dynamic delay using switching activity generated through gate-level simulation of a post place-and-route design in TSMC 45nm process. We then look for `features' in the input data that influence dynamic path sensitization. Using these features we apply a logistic regression (LR) method to construct a predictive model trained and tested using three datasets: random, Sobel filter and Gaussian filter. We classify dynamic delay into five distinct classes. For a given test input, WILD predicts the class of output dynamic delay. On average across several FUs, 98.0% of WILD predictions are consistent with gate-level simulation. Using WILD-directed dynamic frequency scaling can improve instruction-level performance by 13%-44% compared to the state-of-the-art instruction-level timing model. Xun Jiao 0002, Yu Jiang 0001, Abbas Rahimi, Rajesh K. Gupta 0001 |
ICCD | 4 |
| 2016 | Variability Mitigation in Nanometer CMOS Integrated Systems: A Survey of Techniques From Circuits to SoftwareabstractVariation in performance and power across manufactured parts and their operating conditions is an accepted reality in modern microelectronic manufacturing processes with geometries in nanometer scales. This article surveys challenges and opportunities in identifying variations, their effects and methods to combat these variations for improved microelectronic devices. We focus on computing devices and their design at various levels to combat variability. First, we provide a review of key concepts with particular emphasis on timing errors caused by various variability sources. We consider methods to predict and prevent, detect and correct, and finally conditions under which such errors can be accepted; we also consider their implications on cost, performance and quality. We provide a comparative evaluation of methods for deployment across various layers of the system from circuits, architecture, to application software. These can be combined in various ways to achieve specific goals related to observability and controllability of the variability effects, providing means to achieve cross-layer or hybrid resilience. We then provide examples of real world resilient single-core and parallel architectures. We find that parallel architectures and parallelism in general provide the best means to combat and exploit variability to design resilient and efficient systems. Using programmable accelerator architectures such as clustered processing elements and GP-GPUs, we show how system designers can coordinate propagation of timing error information and its effects along with new techniques for memoization (i.e., spatial or temporal reuse of computation). This discussion naturally leads to use of these techniques into emerging area of “approximate computing,” and how these can be used in building resilient and efficient computing systems. We conclude with an outlook for the emerging field. Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001 |
Proc. IEEE | 3 |
| 2015 | Models, abstractions, and architectures: the missing links in cyber-physical systemsabstractBridging disparate realms of physical and cyber system components requires models and methods that enable rapid evaluation of design alternatives in cyber-physical systems (CPS). The diverse intellectual traditions of physical and mathematical sciences makes this task exceptionally hard. This paper seeks to explore potential solutions by examining specific examples of CPS applications in automobiles and smart buildings. Both smart buildings and automobiles are complex systems with embedded knowledge across several domains. We present our experiences with development of CPS applications to illustrate the challenges that arise when expertise across domains is integrated into the system, and show that creation of models, abstractions, and architectures that address these challenges are key to next generation CPS applications. Bharathan Balaji, Mohammad Abdullah Al Faruque, Nikil Dutt, Rajesh K. Gupta 0001, Yuvraj Agarwal |
DAC | 4 |
| 2015 | Task scheduling strategies to mitigate hardware variability in embedded shared memory clustersabstractManufacturing and environmental variations cause timing errors that are typically avoided by conservative design guardbands or corrected by circuit level error detection and correction. These measures incur energy and performance penalties. This paper considers methods to reduce this cost by expanding the scope of variability mitigation through the software stack. In particular, we propose workload deployment methods that reduce the likelihood of timing errors in shared memory clusters of processor cores. This and other methods are incorporated in a runtime layer in the OpenMP framework that enables parsimonious countermeasures against timing errors induced by hardware variability. The runtime system "introspectively" monitors the costs of tasks execution on various cores and transparently associates descriptive metadata with the tasks. By utilizing the characterized metadata, we propose several policies that enhance the cluster choices for scheduling tasks to cores according to measured hardware variability and system workload. We devise efficient task scheduling strategies for simultaneous management of variability and workload by exploiting centralized and distributed approaches to workload distribution. Both schedulers surpass current state-of-the-art approaches; the distributed (or the centralized) achieves on average 30% (or 17%) energy, and 17% (4%) performance improvement. Abbas Rahimi, Daniele Cesarini, Andrea Marongiu, Rajesh K. Gupta 0001, Luca Benini |
DAC | 4 |
| 2015 | Approximate associative memristive memory for energy-efficient GPUs
Abbas Rahimi, Amirali Ghofrani, Kwang-Ting Cheng, Luca Benini, Rajesh K. Gupta 0001 |
DATE | 5 |
| 2015 | Aging-Aware Compilation for GP-GPUsabstractGeneral-purpose graphic processing units (GP-GPUs) offer high computational throughput using thousands of integrated processing elements (PEs). These PEs are stressed during workload execution, and negative bias temperature instability (NBTI) adversely affects their reliability by introducing new delay-induced faults. However, the effect of these delay variations is not uniformly spread across the PEs: some are affected more—hence less reliable—than others. This variation causes significant reduction in the lifetime of GP-GPU parts. In this article, we address the problem of “wear leveling” across processing units to mitigate lifetime uncertainty in GP-GPUs. We propose innovations in the static compiled code that can improve healing in PEs and stream cores (SCs) based on their degradation status. PE healing is a fine-grained very long instruction word (VLIW) slot assignment scheme that balances the stress of instructions across the PEs within an SC. SC healing is a coarse-grained workload allocation scheme that distributes workload across SCs in GP-GPUs. Both schemes share a common property: they adaptively shift workload from less reliable units to more reliable units, either spatially or temporally. These software schemes are based on online calibration with NBTI monitoring that equalizes the expected lifetime of PEs and SCs by regenerating adaptive compiled codes to respond to the specific health state of the GP-GPUs. We evaluate the effectiveness of the proposed schemes for various OpenCL kernels from the AMD APP SDK on Evergreen and Southern Island GPU architectures. The aging-aware healthy kernels generated by the PE (or SC) healing scheme reduce NBTI-induced voltage threshold shift by 30% (77% in the case of SCs), with no (moderate) performance penalty compared to the naive kernels. Atieh Lotfi, Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2014 | Workload Shaping to Mitigate Variability in Renewable Power Use by Data CentersabstractThis paper explores the opportunity for energy saving in data centers using the flexibility from the Service Level Agreements (SLAs) and proposes a novel approach for scheduling workload that incorporates use of renewable energy sources. We investigate how much renewable power to store and how much workload to delay for increasing renewable usage while meeting latency constraints. We present an LP formulation for mitigating variability in renewable generation by dynamic deferral and give two online algorithms to determine optimal balance of workload deferral and power use. We prove the feasibility of the online algorithms and show that their worst case performances are bounded by constant factors with respect to the offline formulation. We validate our algorithms by trace-driven simulation on MapReduce workload and collected and publicly available wind and solar power generation data. Results show that the algorithms give 20-30% energy-savings compared to the naive 'follow the workload' policy. Muhammad Abdullah Adnan, Rajesh K. Gupta 0001 |
IEEE CLOUD | 2 |
| 2014 | Energy-Efficient GPGPU Architectures via Collaborative Compilation and Memristive Memory-Based ComputingabstractThousands of deep and wide pipelines working concurrently make GPGPU high power consuming parts. Energy-efficiency techniques employ voltage overscaling that increases timing sensitivity to variations and hence aggravating the energy use issues. This paper proposes a method to increase spatiotemporal reuse of computational effort by a combination of compilation and micro-architectural design. An associative memristive memory (AMM) module is integrated with the floating point units (FPUs). Together, we enable fine-grained partitioning of values and find high-frequency sets of values for the FPUs by searching the space of possible inputs, with the help of application-specific profile feedback. For every kernel execution, the compiler pre-stores these high-frequent sets of values in AMM modules -- representing partial functionality of the associated FPU-- that are concurrently evaluated over two clock cycles. Our simulation results show high hit rates with 32-entry AMM modules that enable 36% reduction in average energy use by the kernel codes. Compared to voltage overscaling, this technique enhances robustness against timing errors with 39% average energy saving. Abbas Rahimi, Amirali Ghofrani, Miguel Angel Lastras-Montaño, Kwang-Ting Cheng, Luca Benini, Rajesh K. Gupta 0001 |
DAC | 6 |
| 2014 | Temporal memoization for energy-efficient timing error recovery in GPGPUsabstractManufacturing and environmental variability lead to timing errors in computing systems that are typically corrected by error detection and correction mechanisms at the circuit level. The cost and speed of recovery can be improved by memoization-based optimization methods that exploit spatial or temporal parallelisms in suitable computing fabrics such as general-purpose graphics processing units (GPGPUs). We propose here a temporal memoization technique for use in floating-point units (FPUs) in GPGPUs that uses value locality inside data-parallel programs. The technique recalls (memorizes) the context of error-free execution of an instruction on a FPU. To enable scalable and independent recovery, a single-cycle lookup table (LUT) is tightly coupled to every FPU to maintain contexts of recent error-free executions. The LUT reuses these memorized contexts to exactly, or approximately, correct errant FP instructions based on application needs. In real-world applications, the temporal memoization technique achieves an average energy saving of 8%-28% for a wide range of timing error rates (0%-4%) and outperforms recent advances in resilient architectures. This technique also enhances robustness in the voltage overscaling regime and achieves relative average energy saving of 66 % with 11% voltage overscaling. Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2014 | Application-Adaptive Guardbanding to Mitigate Static and Dynamic VariabilityabstractTraditional application execution assumes an error-free execution hardware and environment. Such guarantees in execution are achieved by providing guardbands in the design of microelectronic processors. In reality, applications exhibit varying degrees of tolerance to error in computations. This paper proposes an adaptive guardbanding technique to combat CMOS variability for error-tolerant (probabilistic) applications as well as traditional error-intolerant applications. The proposed technique leverages a combination of accurate design time analysis and a minimally intrusive runtime technique to mitigate Process, Voltage, and Temperature (PVT) variations for a near-zero area overhead. We demonstrate our approach on a 32-bit in-order RISC processor with full post Placement and Routing (P&R) layout results in TSMC 45 nm technology. The adaptive guardbanding technique eliminates traditional guardbands on operating frequency using information from PVT variations and application-specific requirements on computational accuracy. For error-intolerant applications, we introduce the notion of Sequence-Level Vulnerability (SLV) that utilizes circuit-level vulnerability for constructing high-level software knowledge as metadata. In effect, the SLV metadata partitions sequences of integer SPARC instructions into two equivalence classes to enable the adaptive guardbanding technique to adapt the frequency simultaneously for dynamic voltage and temperature variations, as well as adapt to the different classes of the instruction sequences. The proposed technique achieves on an average$1.6 \times $speedup for error-intolerant applications compared to recent work. For probabilistic applications, the adaptive technique guarantees the error-free operation of a set of paths of the processor that always require correct timing (Vulnerable Paths) while reducing the cost of guardbanding for the rest of the paths (Invulnerable Paths). This increases the throughput of probabilistic applications upto$1.9 \times $over the traditional worst-case design. The proposed technique has 0.022% area overhead, and imposes only 0.034% and 0.031% total power overhead for intolerant and probabilistic applications respectively. Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001 |
IEEE Trans. Computers | 3 |
| 2013 | Path Consolidation for Dynamic Right-Sizing of Data Center NetworksabstractData center topologies typically consist of multirooted trees with many equal-cost paths between a given pair of hosts. Existing power optimization techniques do not utilize this property of data center networks for power proportionality. In this paper, we exploit this opportunity and show that significant energy savings can be achieved via path consolidation in the network. We present an offline formulation for the flow assignment in a data center network and develop an online algorithm by path consolidation for dynamic right-sizing of the network to save energy. To validate our algorithm, we build a flow level simulator for a data center network. Our simulation on flow traces generated from MapReduce workload shows 80% reduction in network energy consumption in data center networks and ~25% more energy savings compared to the existing techniques for saving energy in data center networks. Muhammad Abdullah Adnan, Rajesh K. Gupta 0001 |
IEEE CLOUD | 2 |
| 2013 | Aging-aware compiler-directed VLIW assignment for GPGPU architecturesabstractNegative bias temperature instability (NBTI) adversely affects the reliability of a processor by introducing new delay-induced faults. However, the effect of these delay variations is not uniformly spread across functional units and instructions: some are affected more (hence less reliable) than others. This paper proposes a NBTI-aware compiler-directed very long instruction word (VLIW) assignment scheme that uniformly distributes the stress of instructions with the aim of minimizing aging of GPGPU architecture without any performance penalty. The proposed solution is an entirely software technique based on static workload characterization and online execution with NBTI monitoring that equalizes the expected lifetime of each processing element by regenerating aging-aware healthy kernels that respond to the specific health state of GPGPU. We demonstrate our approach on AMD Evergreen architecture where iso-throughput executions of the healthy kernels reduce NBTI-induced voltage threshold shift up to 49% (11%) compared to naïve kernel executions, with (without) architectural support for power-gating. The kernel adaption flow takes average of 13 millisecond on a typical host machine thus making it suitable for practical implementation. Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001 |
DAC | 3 |
| 2013 | Utility-aware deferred load balancing in the cloud driven by dynamic pricing of electricityabstractDistributed computing resources in a cloud computing environment provides an opportunity to reduce energy and its cost by shifting loads in response to dynamically varying availability of energy. This variation in electrical power availability is represented in its dynamically changing price that can be used to drive workload deferral against performance requirements. But such deferral may cause user dissatisfaction. In this paper, we quantify the impact of deferral on user satisfaction and utilize flexibility from the service level agreements (SLAs) for deferral to adapt with dynamic price variation. We differentiate among the jobs based on their requirements for responsiveness and schedule them for energy saving while meeting deadlines and user satisfaction. Representing utility as decaying functions along with workload deferral, we make a balance between loss of user satisfaction and energy efficiency. We model delay as decaying functions and guarantee that no job violates the maximum deadline, and we minimize the overall energy cost. Our simulation on MapReduce traces show that energy consumption can be reduced by ∼15%, with such utility-aware deferred load balancing. We also found that considering utility as a decaying function gives better cost reduction than load balancing with a fixed deadline. Muhammad Abdullah Adnan, Rajesh K. Gupta 0001 |
DATE | 2 |
| 2013 | Hierarchically focused guardbanding: an adaptive approach to mitigate PVT variations and agingabstractThis paper proposes a new model of functional units for variation-induced timing errors due to PVT variations and device Aging (PVTA). The model takes into account PVTA parameter variations, clock frequency, and the physical details of Placed-and-Routed (P&R) functional units in 45nm TSMC analysis flow. Using this model and PVTA monitoring circuits, we propose Hierarchically Focused Guardbanding (HFG) as a method to adaptively mitigate PVTA variations. We demonstrate the effectiveness of HFG on GPU architecture at two granularities of observation and adaptation: (i) fine-grained instruction-level; and (ii) coarse-grained kernel-level. Using coarse-grained PVTA monitors with kernel-level adaptation, the throughput increases by 70% on average. By comparison, the instruction-by-instruction monitoring and adaptation enhances throughput by a factor of 1.8×–2.1× depending on the configuration of PVTA monitors and the type of instructions executed in the kernels. Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2013 | Variation-tolerant OpenMP tasking on tightly-coupled processor clustersabstractWe present a variation-tolerant tasking technique for tightly-coupled shared memory processor clusters that relies upon modeling advance across the hardware/software interface. This is implemented as an extension to the OpenMP 3.0 tasking programming model. Using the notion of Task-Level Vulnerability (TLV) proposed here, we capture dynamic variations caused by circuit-level variability as a high-level software knowledge. This is accomplished through a variation-aware hardware/software codesign where: (i) Hardware features variability monitors in conjunction with online per-core characterization of TLV metadata; (ii) Software supports a Task-level Errant Instruction Management (TEIM) technique to utilize TLV metadata in the runtime OpenMP task scheduler. This method greatly reduces the number of recovery cycles compared to the baseline scheduler of OpenMP [22], consequently instruction per cycle (IPC) of a 16-core processor cluster is increased up to 1.51× (1.17× on average). We evaluate the effectiveness of our approach with various number of cores (4,8,12,16), and across a wide temperature range(ΔT=90°C). Abbas Rahimi, Andrea Marongiu, Paolo Burgio, Rajesh K. Gupta 0001, Luca Benini |
DATE | 4 |
| 2013 | Minerva: Accelerating Data Analysis in Next-Generation SSDsabstractEmerging non-volatile memory (NVM) technologies have DRAM-like latency with storage-like density, offering unique capability to analyze large data sets significantly faster than flash or disk storage. However, the hybrid nature of these NVM technologies such as phase-change memory (PCM) make it difficult to use them to best advantage in the memory-storage hierarchy. These NVMs lack the fast write latency required of DRAM and are thus not suitable as DRAM equivalent on the memory bus, yet their low latency even in random access patterns is not easily exploited over an I/O bus. In this work, we describe an FPGA-based system to execute application-specific operations in the NVM controller and evaluate its performance on two microbenchmarks and a keyvalue store. Our system Minerva1extends the conventional solidstate drive (SSD) architecture to offload data or I/O intensive application code to the SSD to exploit the low latency and high internal bandwidth of NVMs. Performing computation in the FPGA-based NVM storage controller significantly reduces data traffic between the host and storage and serves as an offload engine for data analysis workloads. A runtime library enables the programmer to offload computations to the SSD without dealing with the complications of the underlying architecture and inter-controller communication management. We have implemented a prototype of Minerva on the BEE3 FPGA system. We compare the performance of Minerva to a state of the art PCIe-attached PCM-based SSD. Minerva improves performance by an order of magnitude on two microbenchmarks. Minerva based key-value store performs up to 5.2 M get operations/s and 4.0 M set operations/s which is 7.45× and 9.85× higher than the PCM-based SSD that uses the conventional I/O architecture. This huge improvement comes from the reduction of data transfer between the storage to the host and the FPGA-based data processing in the SSD. Arup De, Maya B. Gokhale, Rajesh K. Gupta 0001, Steven Swanson |
FCCM | 3 |
| 2013 | Sentinel: occupancy based HVAC actuation using existing WiFi infrastructure within commercial buildingsabstractCommercial buildings contribute to 19% of the primary energy consumption in the US, with HVAC systems accounting for 39.6% of this usage. To reduce HVAC energy use, prior studies have proposed using wireless occupancy sensors or even cameras for occupancy based actuation showing energy savings of up to 42%. However, most of these solutions require these sensors and the associated network to be designed, deployed, tested and maintained within existing buildings which is significantly costly. Bharathan Balaji, Anthony Nwokafor, Rajesh K. Gupta 0001, Yuvraj Agarwal |
SenSys | 4 |
| 2013 | From ARIES to MARS: transaction support for next-generation, solid-state drivesabstractTransaction-based systems often rely on write-ahead logging (WAL) algorithms designed to maximize performance on disk-based storage. However, emerging fast, byte-addressable, non-volatile memory (NVM) technologies (e.g., phase-change memories, spin-transfer torque MRAMs, and the memristor) present very different performance characteristics, so blithely applying existing algorithms can lead to disappointing performance. Joel Coburn, Trevor Bunker, Meir Schwarz, Rajesh K. Gupta 0001, Steven Swanson |
SOSP | 4 |
| 2013 | Underdesigned and Opportunistic Computing in Presence of Hardware VariabilityabstractMicroelectronic circuits exhibit increasing variations in performance, power consumption, and reliability parameters across the manufactured parts and across use of these parts over time in the field. These variations have led to increasing use of overdesign and guardbands in design and test to ensure yield and reliability with respect to a rigid set of datasheet specifications. This paper explores the possibility of constructing computing machines that purposely expose hardware variations to various layers of the system stack including software. This leads to the vision of underdesigned hardware that utilizes a software stack that opportunistically adapts to a sensed or modeled hardware. The envisioned underdesigned and opportunistic computing (UnO) machines face a number of challenges related to the sensing infrastructure and software interfaces that can effectively utilize the sensory data. In this paper, we outline specific sensing mechanisms that we have developed and their potential use in building UnO machines. Puneet Gupta 0001, Yuvraj Agarwal, Lara Dolecek, Nikil Dutt, Rajesh K. Gupta 0001, Rakesh Kumar 0002, Subhasish Mitra, Alexandru Nicolau, Tajana Rosing, Mani Srivastava 0001, Steven Swanson, Dennis Sylvester |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2012 | Energy Efficient Geographical Load Balancing via Dynamic Deferral of WorkloadabstractWith the increasing popularity of Cloud computing and Mobile computing, individuals, enterprises and research centers have started outsourcing their IT and computational needs to on-demand cloud services. Recently geographical load balancing techniques have been suggested for data centers hosting cloud computation in order to reduce energy cost by exploiting the electricity price differences across regions. However, these algorithms do not draw distinction among diverse requirements for responsiveness across various workloads. In this paper, we use the flexibility from the Service Level Agreements (SLAs) to differentiate among workloads under bounded latency requirements and propose a novel approach for cost savings for geographical load balancing. We investigate how much workload to be executed in each data center and how much workload to be delayed and migrated to other data centers for energy saving while meeting deadlines. We present an offline formulation for geographical load balancing problem with dynamic deferral and give online algorithms to determine the assignment of workload to the data centers and the migration of workload between data centers in order to adapt with dynamic electricity price changes. We compare our algorithms with the greedy approach and show that significant cost savings can be achieved by migration of workload and dynamic deferral with future electricity price prediction. We validate our algorithms on MapReduce traces and show that geographic load balancing with dynamic deferral can provide 20-30% cost-savings. Muhammad Abdullah Adnan, Ryo Sugihara, Rajesh K. Gupta 0001 |
IEEE CLOUD | 3 |
| 2012 | Analysis of instruction-level vulnerability to dynamic voltage and temperature variationsabstractVariation in performance and power across manufactured parts and their operating conditions is an accepted reality in aggressive CMOS processes. This paper considers challenges and opportunities in identifying this variation and methods to combat it for improved computing systems. We introduce the notion of instruction-level vulnerability (ILV) to expose variation and its effects to the software stack for use in architectural/compiler optimizations. To compute ILV, we quantify the effect of voltage and temperature variations on the performance and power of a 32-bit, RISC, in-order processor in 65 nm TSMC technology at the level of individual instructions. Results show 3.4 ns (68FO4) delay variation and 26.7x power variation among instructions, and across extreme corners. Our analysis shows that ILV is not uniform across the instruction set. In fact, ILV data partitions instructions into three equivalence classes. Based on this classification, we show how a low-overhead robustness enhancement techniques can be used to enhance performance by a factor of 1.1x-5.5x. Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2012 | Procedure hopping: a low overhead solution to mitigate variability in shared-L1 processor clustersabstractVariation in performance and power across manufactured parts and their operating conditions is a well-known issue in advanced CMOS processes. This paper proposes a resilient HW/SW architecture for shared-L1 processor clusters to combat both static and dynamic variations. We first introduce the notion of procedure-level vulnerability (PLV) to expose fast dynamic voltage variation and its effects to the software stack for use in runtime compensation. To assess PLV, we quantify the effect of full operating conditions on the dynamic voltage variation of a post-layout processor in 45nm TSMC technology. Based on our analysis, PLV shows a range of 18mV--63mV inter-corner variation among the maximum voltage droop of procedures. To exploit this variation we propose a low-cost procedure hopping technique within the processor clusters, utilizing compile time characterized metadata related to PLV. Our results show that procedure hopping avoids critical voltage droops during the execution of all procedures while incurring less than 1% latency penalty. Abbas Rahimi, Luca Benini, Rajesh K. Gupta 0001 |
ISLPED | 3 |
| 2012 | Verifying GPU kernels by test amplificationabstractWe present a novel technique for verifying properties of data parallel GPU programs via test amplification. The key insight behind our work is that we can use the technique of static information flow to amplify the result of a single test execution over the set of all inputs and interleavings that affect the property being verified. We empirically demonstrate the effectiveness of test amplification for verifying race-freedom and determinism over a large number of standard GPU kernels, by showing that the result of verifying a single dynamic execution can be amplified over the massive space of possible data inputs and thread interleavings. Alan Leung, Manish Gupta 0010, Yuvraj Agarwal, Rajesh K. Gupta 0001, Ranjit Jhala, Sorin Lerner |
PLDI | 4 |
| 2012 | Energy-efficient deadline scheduling for heterogeneous systems
Ryo Sugihara, Rajesh K. Gupta 0001 |
J. Parallel Distributed Comput. | 4 |
| 2011 | NV-Heaps: making persistent objects fast and safe with next-generation, non-volatile memoriesabstractPersistent, user-defined objects present an attractive abstraction for working with non-volatile program state. However, the slow speed of persistent storage (i.e., disk) has restricted their design and limited their performance. Fast, byte-addressable, non-volatile technologies, such as phase change memory, will remove this constraint and allow programmers to build high-performance, persistent data structures in non-volatile storage that is almost as fast as DRAM. Creating these data structures requires a system that is lightweight enough to expose the performance of the underlying memories but also ensures safety in the presence of application and system failures by avoiding familiar bugs such as dangling pointers, multiple free()s, and locking errors. In addition, the system must prevent new types of hard-to-find pointer safety bugs that only arise with persistent objects. These bugs are especially dangerous since any corruption they cause will be permanent. Joel Coburn, Adrian M. Caulfield, Ameen Akel, Laura M. Grupp, Rajesh K. Gupta 0001, Ranjit Jhala, Steven Swanson |
ASPLOS | 5 |
| 2011 | Underdesigned and Opportunistic ComputingabstractVariation in the specifications of microelectronic chips across parts and over time has been a great source of concern for the integrated circuit chip designers because of the ever-increasing guard-bands that the circuit and system designers must rely upon to ensure working parts and systems. This prompts us to look for solutions that can mitigate the effect of performance and power variability through innovations in sys-tem software. In this paper, we outline a novel, flexible hardware-software stack and interface that use adaptation in software to relax variation-induced guard-bands in hardware design. Puneet Gupta 0001, Rajesh K. Gupta 0001 |
Asian Test Symposium | 2 |
| 2011 | Understanding the role of buildings in a smart microgridabstractA `smart microgrid' refers to a distribution network for electrical energy, starting from electricity generation to its transmission and storage with the ability to respond to dynamic changes in energy supply through co-generation and demand adjustments. At the scale of a small town, a microgrid is connected to the wide-area electrical grid that may be used for `baseline' energy supply; or in the extreme case only as a storage system in a completely self-sufficient microgrid. Distributed generation, storage and intelligence are key components of a smart microgrid. In this paper, we examine the significant role that buildings play in energy use and its management in a smart microgrid. In particular, we discuss the relationship that IT equipment has on energy usage by buildings, and show that control of various building subsystems (such as IT and HVAC) can lead to significant energy savings. Using the UCSD as a prototypical smart microgrid, we discuss how buildings can be enhanced and interfaced with the smart microgrid, and demonstrate the benefits that this relationship can bring as well as the challenges in implementing this vision. Yuvraj Agarwal, Thomas Weng, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2011 | Clock Synchronization with Deterministic Accuracy Guarantee
Ryo Sugihara, Rajesh K. Gupta 0001 |
EWSN | 2 |
| 2011 | Onyx: A Prototype Phase Change Memory Storage Array
Ameen Akel, Adrian M. Caulfield, Todor I. Mollov, Rajesh K. Gupta 0001, Steven Swanson |
HotStorage | 4 |
| 2011 | Sensor localization with deterministic accuracy guaranteeabstractLocalizability of network or node is an important subproblem in sensor localization. While rigidity theory plays an important role in identifying several localizability conditions, major limitations are that the results are only applicable to generic frameworks and that the distance measurements need to be error-free. These limitations, in addition to the hardness of finding the node locations for a uniquely localizable graph, miss large portions of practical application scenarios that require sensor localization. In this paper, we describe a novel SDP-based formulation for analyzing node localizability and providing a deterministic upper bound of localization error. Unlike other optimization-based formulations for solving localization problem for the whole network, our formulation allows fine-grained evaluation on the localization accuracy per each node. Our formulation gives a sufficient condition for unique node localizability for any frameworks, i.e., not only for generic frameworks. Furthermore, we extend it for the case with measurement errors and for computing directional error bounds. We also design an iterative algorithm for large-scale networks and demonstrate the effectiveness by simulation experiments. Ryo Sugihara, Rajesh K. Gupta 0001 |
INFOCOM | 2 |
| 2011 | Duty-cycling buildings aggressively: The next frontier in HVAC control
Yuvraj Agarwal, Bharathan Balaji, Seemanta Dutta, Rajesh K. Gupta 0001, Thomas Weng |
IPSN | 4 |
| 2011 | Path Planning of Data Mules in Sensor NetworksabstractWe study the problem of planning the motion of “data mules” for collecting the data from stationary sensor nodes in wireless sensor networks. Use of data mules significantly reduces energy consumption at sensor nodes compared to commonly used multihop forwarding approaches, but has a drawback in that it increases the latency of data delivery. Optimizing the motion of data mules, including path and speed, is critical for improving the data delivery latency and making the data mule approach more useful in practice. In this article, we focus on the path selection problem: finding the optimal path of data mules so that the data delivery latency can be minimized. We formulate the path selection problem as a graph problem that is capable of expressing the benefit from larger communication range. The problem is NP-hard and we present approximation algorithms for both single-data mule case and multiple-data mules case. We further consider the case in which we have only partial knowledge of communication range, where we design semionline algorithms that improve the offline plan using online knowledge at runtime. Simulation experiments on Matlab and ns2 demonstrate that our offline and semionline algorithms produce significantly shorter path lengths and data delivery latency compared to previously proposed methods, suggesting that controlled mobility can be exploited much more effectively. Ryo Sugihara, Rajesh K. Gupta 0001 |
ACM Trans. Sens. Networks | 2 |
| 2010 | Moneta: A High-Performance Storage Array Architecture for Next-Generation, Non-volatile MemoriesabstractEmerging non-volatile memory technologies such as phase change memory (PCM) promise to increase storage system performance by a wide margin relative to both conventional disks and flash-based SSDs. Realizing this potential will require significant changes to the way systems interact with storage devices as well as a rethinking of the storage devices themselves. This paper describes the architecture of a prototype PCIe-attached storage array built from emulated PCM storage called Moneta. Moneta provides a carefully designed hardware/software interface that makes issuing and completing accesses atomic. The atomic management interface, combined with hardware scheduling optimizations, and an optimized storage stack increases performance for small, random accesses by 18x and reduces software overheads by 60%. Moneta array sustain 2.8~GB/s for sequential transfers and 541K random 4~KB~IO operations per second (8x higher than a state-of-the-art flash-based SSD). Moneta can perform a 512-byte write in 9~us (5.6x faster than the SSD). Moneta provides a harmonic mean speedup of 2.1x and a maximum speed up of 9x across a range of file system, paging, and database workloads. We also explore trade-offs in Moneta's architecture between performance, power, memory organization, and memory latency. Adrian M. Caulfield, Arup De, Joel Coburn, Todor I. Mollov, Rajesh K. Gupta 0001, Steven Swanson |
MICRO | 5 |
| 2010 | Understanding the Impact of Emerging Non-Volatile Memories on High-Performance, IO-Intensive ComputingabstractEmerging storage technologies such as flash memories, phase-change memories, and spin-transfer torque memories are poised to close the enormous performance gap between disk-based storage and main memory. We evaluate several approaches to integrating these memories into computer systems by measuring their impact on IO-intensive, database, and memory-intensive applications. We explore several options for connecting solid-state storage to the host system and find that the memories deliver large gains in sequential and random access performance, but that different system organizations lead to different performance trade-offs. The memories provide substantial application-level gains as well, but overheads in the OS, file system, and application can limit performance. As a result, fully exploiting these memories' potential will require substantial changes to application and system software. Finally, paging to fast non-volatile memories is a viable option for some applications, providing an alternative to expensive, powerhungry DRAM for supporting scientific applications with large memory footprints. Adrian M. Caulfield, Joel Coburn, Todor I. Mollov, Arup De, Ameen Akel, Jiahua He, Arun Jagatheesan, Rajesh K. Gupta 0001, Allan Snavely, Steven Swanson |
SC | 8 |
| 2010 | SleepServer: A Software-Only Approach for Reducing the Energy Consumption of PCs within Enterprise Environments
Yuvraj Agarwal, Stefan Savage, Rajesh K. Gupta 0001 |
USENIX ATC | 3 |
| 2010 | Translation Validation of High-Level SynthesisabstractThe growing complexity of systems and their implementation into silicon encourages designers to look for ways to model designs at higher levels of abstraction and then incrementally build portions of these designs - automatically or manually - from these high-level specifications. Unfortunately, this translation process itself can be buggy, which can create a mismatch between what a designer intends and what is actually implemented in the circuit. Therefore, checking if the implementation is a refinement or equivalent to its initial specification is of tremendous value. In this paper, we present an approach to automatically validate the implementation against its initial high-level specification using insights from translation validation, automated theorem proving, and relational approaches to reasoning about programs. In our experiments, we first focus on concurrent systems modeled as communicating sequential processes and show that their refinements can be validated using our approach. Next, we have applied our validation approach to a realistic scenario - a parallelizing high-level synthesis framework called Spark. We present the details of our algorithm and experimental results. Sudipta Kundu, Sorin Lerner, Rajesh K. Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | Optimal Speed Control of Mobile Node for Data Collection in Sensor NetworksabstractA data mule represents a mobile device that collects data in a sensor field by physically visiting the nodes in a sensor network. The data mule collects data when it is in the proximity of a sensor node. This can be an alternative to multihop forwarding of data when we can utilize node mobility in a sensor network. To be useful, a data mule approach needs to minimize data delivery latency. In this paper, we first formulate the problem of minimizing the latency in the data mule approach. The data mule scheduling (DMS) problem is a scheduling problem that has both location and time constraints. Then, for the 1D case of the DMS problem, we design an efficient heuristic algorithm that incorporates constraints on the data mule motion dynamics. We provide lower bounds of solutions to evaluate the quality of heuristic solutions. Through numerical experiments, we show that the heuristic algorithm runs fast and yields good solutions that are within 10 percent of the optimal solutions. Ryo Sugihara, Rajesh K. Gupta 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2010 | Speed control and scheduling of data mules in sensor networksabstractUnlike traditional multihop forwarding among stationary sensor nodes, use of mobile devices for data collection in wireless sensor networks has recently been gathering more attention. The use of mobility significantly reduces the energy consumption at sensor nodes, elongating the functional lifetime of the network. However, a drawback is an increased data delivery latency. Reducing the latency through optimizing the motion of data mules is critical for this approach to thrive. In this article, we focus on the problem of motion planning, specifically, determination of the speed of the data mule and the scheduling of the communication tasks with the sensors. We consider three models of mobility capability of the data mule to accommodate different types of vehicles. Under each mobility model, we design optimal and heuristic algorithms for different problems: single data mule case, single data mule with periodic data generation case, and multiple data mules case. We compare the performance of the heuristic algorithm with a naive algorithm and also with the multihop forwarding approach by numerical experiments. We also compare one of the optimal algorithms with a previously proposed method to see how our algorithm improves the performance and is also useful in practice. As far as we know, this study is the first of a kind that provides a systematic understanding of the motion planning problem of data mules. Ryo Sugihara, Rajesh K. Gupta 0001 |
ACM Trans. Sens. Networks | 2 |
| 2009 | LazySync: A New Synchronization Scheme for Distributed Simulation of Sensor Networks
Zhong-Yi Jin, Rajesh K. Gupta 0001 |
DCOSS | 2 |
| 2009 | Optimizing Energy-Latency Trade-Off in Sensor Networks with Controlled MobilityabstractWe consider the problem of planning path and speed of a "data mule" in a sensor network. This problem is encountered in various situations, such as modeling the motion of a data-collecting UAV (unmanned aerial vehicle) in a field of sensors for structural health monitoring. Our specific context here is use of a data mule as an alternative or supplement to multihop forwarding in a sensor network. While a data mule can reduce the energy consumption at each sensor node, it increases the latency from the time the data is generated at a node to the time the base station receives it. In this paper, we introduce the "data mule scheduling" or DMS framework that enables data mule motion planning to minimize the data delivery latency. The DMS framework is general; it can express many previously proposed problem formulations and problem settings related to data mules. We design algorithms for DMS and extend to the more general case of combined data mule and multihop forwarding to enable a flexible trade-off between energy consumption and data delivery latency. Using DMS, we can calculate the optimal way for node-to-node forwarding and data mule motion plan. Our implementation and simulation results using ns2 show nearly monotonic decrease of data delivery latency when each node can use more energy, thus vastly increasing the flexibility in the energy-latency trade-off for sensor network communications. Ryo Sugihara, Rajesh K. Gupta 0001 |
INFOCOM | 2 |
| 2009 | Improving the speed and scalability of distributed simulations of sensor networks
Zhong-Yi Jin, Rajesh K. Gupta 0001 |
IPSN | 2 |
| 2009 | Somniloquy: Augmenting Network Interfaces to Reduce PC Energy Usage
Yuvraj Agarwal, Steve Hodges 0001, Ranveer Chandra, James Scott, Paramvir Bahl, Rajesh K. Gupta 0001 |
NSDI | 6 |
| 2009 | Softspeak: Making VoIP Play Well in Existing 802.11 Deployments
Patrick Verkaik, Yuvraj Agarwal, Rajesh K. Gupta 0001, Alex C. Snoeren |
NSDI | 3 |
| 2009 | A gateway node with duty-cycled radio and processing subsystems for wireless sensor networksabstractWireless sensor nodes are increasingly being tasked with computation and communication intensive functions while still subject to constraints related to energy availability. On these embedded platforms, once all low power design techniques have been explored, duty-cycling the various subsystems remains the primary option to meet the energy and power constraints. This requires the ability to provide spurts of high MIPS and high bandwidth connections. However, due to the large overheads associated with duty-cycling the computation and communication subsystems, existing high performance sensor platforms are not efficient in supporting such an option. In this article, we present the design and optimizations taken in a wireless gateway node (WGN) that bridges data from wireless sensor networks to Wi-Fi networks in an on-demand basis. We discuss our strategies to reduce duty-cycling related costs by partitioning the system and by reducing the amount of time required to activate or deactivate the high-powered components. We compare the design choices and performance parameters with those made in the Intel Stargate platform to show the effectiveness of duty-cycling on our platform. We have built a working prototype, and the experimental results with two different power management schemes show significant reductions in latency and average power consumption compared to the Stargate . The WGN running our power-gating scheme performs about six times better in terms of average system power consumption than the Stargate running the suspend-system scheme for large working-periods where the active power dominates. For short working-periods where the transition (enable/disable) power becomes dominant, we perform up to seven times better. The comparative performance of our system is even greater when the sleep power dominates. Zhong-Yi Jin, Curt Schurgers, Rajesh K. Gupta 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2008 | Validating High-Level Synthesis
Sudipta Kundu, Sorin Lerner, Rajesh K. Gupta 0001 |
CAV | 3 |
| 2008 | Partial order reduction for scalable testing of systemC TLM designsabstractA SystemC simulation kernel consists of a deterministic implementation of the scheduler, whose specification is non-deterministic. To leverage testing of a SystemC TLM design, we focus on automatically exploring all possible behaviors of the design for a given data input. We combine static and dynamic partial order reduction techniques with SystemC semantics to intelligently explore a subset of the possible traces, while still being provably sufficient for detecting deadlocks and safety property violations. We have implemented our exploration algorithm in a framework called Satya and have applied it to a variety of examples including the TAC benchmark. Using Satya, we automatically found an assertion violation in a benchmark distributed as a part of the OSCI repository. Sudipta Kundu, Malay K. Ganai, Rajesh K. Gupta 0001 |
DAC | 3 |
| 2008 | Temperature Control of High-Performance Multi-core Platforms Using Convex OptimizationabstractWith technology advances, the number of cores integrated on a chip and their speed of operation is increasing. This, in turn is leading to a significant increase in chip temperature. Temperature gradients and hot-spots not only affect the performance of the system, but also lead to unreliable circuit operation and affect the life-time of the chip. Meeting the temperature constraints and reducing the hot-spots are critical for achieving reliable and efficient operation of complex multi-core systems. In this work, we present Pro-Temp, a convex optimization based method that pro-actively controls the temperature of the cores, while minimizing the power consumption and satisfying application performance constraints. The method guarantees that the temperature of the cores are below a user- defined threshold at all instances of operation, while also reducing the hot-spots. We perform experiments on several realistic multi-core benchmarks, which show that the proposed method guarantees that the cores never exceed the maximum temperature limit, while matching the application performance requirements. We compare this to traditional methods, where we find several temperature violations during the operation of the system. Srinivasan Murali, Almir Mutapcic, David Atienza 0001, Rajesh K. Gupta 0001, Stephen P. Boyd, Luca Benini, Giovanni De Micheli |
DATE | 4 |
| 2008 | Improved Distributed Simulation of Sensor Networks Based on Sensor Node Sleep Time
Zhong-Yi Jin, Rajesh K. Gupta 0001 |
DCOSS | 2 |
| 2008 | Improving the Data Delivery Latency in Sensor Networks with Controlled Mobility
Ryo Sugihara, Rajesh K. Gupta 0001 |
DCOSS | 2 |
| 2008 | Programming models for sensor networks: A surveyabstractSensor networks have a significant potential in diverse applications some of which are already beginning to be deployed in areas such as environmental monitoring. As the application logic becomes more complex, programming difficulties are becoming a barrier to adoption of these networks. The difficulty in programming sensor networks is not only due to their inherently distributed nature but also the need for mechanisms to address their harsh operating conditions such as unreliable communications, faulty nodes, and extremely constrained resources. Researchers have proposed different programming models to overcome these difficulties with the ultimate goal of making programming easy while making full use of available resources. In this article, we first explore the requirements for programming models for sensor networks. Then we present a taxonomy of the programming models, classified according to the level of abstractions they provide. We present an evaluation of various programming models for their responsiveness to the requirements. Our results point to promising efforts in the area and a discussion of the future directions of research in this area. Ryo Sugihara, Rajesh K. Gupta 0001 |
ACM Trans. Sens. Networks | 2 |
| 2007 | CATS: cycle accurate transaction-driven simulation with multiple processor simulatorsabstractThis paper focuses on enhancing performance of cycle accurate simulation with multiple processor simulators. Simulation performance is determined by how often simulators exchange events with one another and how accurately simulators model their behavior. Previous techniques have limited their applicability or sacrificed accuracy for performance. In this paper, we notice that inaccuracy comes from events which arrive between event exchange boundaries. To solve the problem, we propose cycle accurate transaction-driven simulation which maintains event exchange boundaries at bus transactions but compensates for accuracy. The proposed technique is implemented in a publicly available CATS framework and our experiment with 64 processors achieves 1.2M processor cycles/s (200K instructions/s) which is faster than other cycle accurate frameworks by an order of magnitude Dohyung Kim 0007, Soonhoi Ha, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2007 | Automated refinement checking of concurrent systemsabstractStepwise refinement is at the core of many approaches to synthesis and optimization of hardware and software systems. For instance, it can be used to build a synthesis approach for digital circuits from high level specifications. It can also be used for post-synthesis modification such as in Engineering Change Orders (ECOs). Therefore, checking if a system, modeled as a set of concurrent processes, is a refinement of another is of tremendous value. In this paper, we focus on concurrent systems modeled as Communicating Sequential Processes (CSP) and show their refinements can be validated using insights from translation validation, automated theorem proving and relational approaches to reasoning about programs. The novelty of our approach is that it handles infinite state spaces in a fully automated manner. We have implemented our refinement checking technique and have applied it to a variety of refinements. We present the details of our algorithm and experimental results. As an example, we were able to automatically check an infinite state space buffer refinement that cannot be checked by current state of the art tools such as FDR. We were also able to check the data part of an industrial case study on the EP2 system. Sudipta Kundu, Sorin Lerner, Rajesh K. Gupta 0001 |
ICCAD | 3 |
| 2007 | Wireless wakeups revisited: energy management for voip over wi-fi smartphonesabstractIP based telephony is rapidly gaining acceptance over traditional means of voice communication. Wireless LANs are also becoming ubiquitous due to their inherent ease of deployment and decreasing costs. In enterpriseWi-Fi environments, VoIP is a compelling application for devices such as smart phones with multiple wireless interfaces. However, the high energy consumption of Wi-Fi interfaces, especially when a device is idle,presents a significant barrier to the widespread adoption of VoIP over Wi-Fi.To address this issue, we present Cell2Notify, a practical and deployable energy management architecture that leverages the cellular radio on a smart phone to implement wakeup for the high-energy consumption Wi-Fi radio. We present detailed measurements of energy consumption on smart phone devices, and we show that Cell2Notify, can extend the battery lifetime of VoIPover Wi-Fi enabled smart phones by a factor of 1.7 to 6.4. Yuvraj Agarwal, Ranveer Chandra, Alec Wolman, Paramvir Bahl, Kevin Chin, Rajesh K. Gupta 0001 |
MobiSys | 6 |
| 2007 | Algorithms for power savingsabstractThis article examines two different mechanisms for saving power in battery-operated embedded systems. The first strategy is that the system can be placed in a sleep state if it is idle. However, a fixed amount of energy is required to bring the system back into an active state in which it can resume work. The second way in which power savings can be achieved is by varying the speed at which jobs are run. We utilize a power consumption curve P ( s ) which indicates the power consumption level given a particular speed. We assume that P ( s ) is convex, nondecreasing, and nonnegative for s ≥ 0. The problem is to schedule arriving jobs in a way that minimizes total energy use and so that each job is completed after its release time and before its deadline. We assume that all jobs can be preempted and resumed at no cost. Although each problem has been considered separately, this is the first theoretical analysis of systems that can use both mechanisms. We give an offline algorithm that is within a factor of 2 of the optimal algorithm. We also give an online algorithm with a constant competitive ratio. Sandy Irani, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
ACM Trans. Algorithms | 3 |
| 2006 | Parallel co-simulation using virtual synchronization with redundant host executionabstractIn traditional parallel co-simulation approaches, the simulation speed is heavily limited by time synchronization overhead between simulators and idle time caused by data dependency. Recent work has shown that the time synchronization overhead can be reduced significantly by predicting the next synchronization points more effectively or by separating trace-driven architecture simulation from trace generation from component simulators. The latter is known as virtual synchronization technique. In this paper, we propose redundant host execution to minimize the simulation idle time caused by data dependency in simulation models. By combining virtual synchronization and redundant host execution techniques we could make parallel execution of multiple simulators a viable solution for fast but cycle-accurate co-simulation. Experiments show about 40% performance gain over a technique which uses virtual synchronization only Dohyung Kim 0007, Soonhoi Ha, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2006 | Compositional interaction specifications for SystemCabstractSystemC is being widely used for system-level modeling of system-on-chip. When designing this class of system, one of the main challenges is to guarantee the correctness of the implementation. This can be especially difficult for designs that are composed of concurrent components with lot of interactions. Most designers use a component-based design approach, where one has an informal idea of how the design should behave, define component specifications, implement and assemble the components into a program, and then check for correctness by simulating the design with a number of testbenches. With this methodology, bugs often go undetected because when using simulation, it is very difficult to test for all possible interactions. To overcome this limitation, our goal is to establish a specification and verification methodology for SystemC. To address the scalability issue, which is a serious limiting factor in state-based verification approaches, we use the concepts of behavioral types; allowing us to effectively infer system properties from properties of its components. In this paper, we answer the following questions: (1) what is a behavioral type? (2) how are behavioral type defined? and (3) how to use the behavioral types in a compositional verification methodology Frederic Doucet, Ingolf Krüger, Rajesh K. Gupta 0001, R. K. Shyamasundar |
MEMOCODE | 3 |
| 2006 | Programming models and languages for SoC-implemented architecturesabstractSummary form only given. Microsystems implemented on chips or SoCs provide new capabilities and alternatives not easily available to architects in the past, via ISA extensions, additional data-path elements, SIMD arrays and their short vector extensions, multigrain coprocessing etc. We can, given enough design resources, build such machines. Evidence of various combinations of architectural archetypes can be seen in a flood of new machine architectures for high definition video processing, though there are significant challenges in architectural exploration. However, a central question that seems to remain unanswered, is how will the programming for such machines can keep up without waiting for the god of programming to descend. Participants will take a crack at answering the questions either from architectural side (including implementation fabrics) or from their views of the advances in programming models and methods Rajesh K. Gupta 0001 |
MEMOCODE | 1 |
| 2006 | CoolSpots: reducing the power consumption of wireless mobile devices with multiple radio interfacesabstractCoolSpots enable a wireless mobile device to automatically switch between multiple radio interfaces, such as WiFi and Bluetooth, in order to increase battery lifetime. The main contribution of this work is an exploration of the policies that enable a system to switch among these interfaces, each with diverse radio characteristics and different ranges, in order to save power - supported by detailed quantitative measurements. The system and policies do not require any changes to the mobile applications themselves, and changes required to existing infrastructure are minimal. Results are reported for a suite of commonly used applications, such as file transfer, web browsing, and streaming media, across a range of operating conditions. Experimental validation of the CoolSpot system on a mobile research platform shows substantial energy savings: more than a 50% reduction in energy consumption of the wireless subsystem is possible, with an associated increase in the effective battery lifetime. Trevor Pering, Yuvraj Agarwal, Rajesh K. Gupta 0001, Roy Want |
MobiSys | 3 |
| 2006 | Optimized Slowdown in Real-Time Task SystemsabstractSlowdown factors determine the extent of slowdown a computing system can experience based on functional and performance requirements. Dynamic voltage scaling (DVS) of a processor based on slowdown factors can lead to considerable energy savings. We address the problem of computing slowdown factors for dynamically scheduled tasks with specified deadlines. We present an algorithm to compute a near optimal constant slowdown factor based on the bisection method. As a further generalization, for the case of tasks with varying power characteristics, we present the computation of near optimal slowdown factors as a solution to convex optimization problem using the ellipsoid method. The algorithms are practically fast and have the same time complexity as the algorithms to compute the feasibility of a task set. Our simulation results show an average 20 percent energy gain over known slowdown techniques using static slowdown factors and 40 percent gain with dynamic slowdown Ravindra Jejurikar, Rajesh K. Gupta 0001 |
IEEE Trans. Computers | 2 |
| 2006 | Energy-aware task scheduling with task synchronization for embedded real-time systemsabstractSlowdown factors determine the extent of slowdown that a computing system can experience based on functional and performance requirements. Dynamic voltage scaling (DVS) of a processor based on slowdown factors can lead to considerable energy savings. This paper addresses the problem of DVS in the presence of task synchronization. Tasks synchronize to enforce mutually exclusive access to the shared resources and can be blocked by lower priority tasks. Task slowdown factors that guarantee meeting all task deadlines are computed. Both static and dynamic priority scheduling viz. rate monotonic (RM) scheduling and earliest deadline first (EDF) scheduling, respectively, are studied Ravindra Jejurikar, Rajesh K. Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Energy efficient watermarking on mobile devices using proxy-based partitioningabstractDigital watermarking embeds an imperceptible signature or watermark in a digital file containing audio, image, text, or video data. The watermark can be used to authenticate the data file and for tamper detection. It is particularly valuable in the use and exchange of digital media, such as audio and video, on emerging handheld devices. However, watermarking is computationally expensive and adds to the drain of the available energy in handheld devices. In this paper, we first analyze the energy profile of various watermarking algorithms. We also study the impact of security and image quality on energy consumption. Second, we present an approach in which we partition the watermarking embedding and extraction algorithms and migrate some tasks to a proxy server. This leads to a lower energy consumption on the handheld without compromising the security of the watermarking process. Experimental results show that executing the watermarking tasks that are partitioned between the proxy and the handheld devices, reduces the total energy consumed by 80%, and improves performance by two orders of magnitude compared to running the application on only the handheld device Arun Kejariwal, Alexandru Nicolau, Nikil Dutt, Rajesh K. Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2005 | Dynamic power management using on demand paging for networked embedded systemsabstractThe power consumption of the network interface plays a major role in determining the total operating lifetime of wireless networked embedded systems. In case of on-demand paging, a low power secondary radio is used to wake up the higher power radio, allowing the latter to sleep for longer periods of time. In this paper we present use of Bluetooth radios to serve as a paging channel for the 802.11b wireless LAN. We have implemented an on-demand paging scheme on an infrastructure based WLAN consisting of iPAQ PDAs equipped with Bluetooth radios and Cisco Aironet wireless networking cards. Our results show power saving ranging from 23% to 48% over the present 802.11b standard operating modes with negligible impact on performance. Yuvraj Agarwal, Curt Schurgers, Rajesh K. Gupta 0001 |
ASP-DAC | 3 |
| 2005 | Dynamic slack reclamation with procrastination scheduling in real-time embedded systemsabstractLeakage energy consumption is an increasing concern in current and future CMOS technology generations. Procrastination scheduling, where task execution can be delayed to maximize the duration of idle intervals, has been proposed to minimize leakage energy drain. The authors addressed dynamic slack reclamation techniques under procrastination scheduling to minimize the static and dynamic energy consumption. In addition to dynamic task slowdown, a dynamic procrastination was proposed, which seeks to extend idle intervals through slack reclamation. While using the entire slack for either slowdown or procrastination need not be the most energy efficient approach, the slack was distributed between slowdown and procrastination to exploit maximum energy savings. The simulation experiments showed that dynamic slowdown result on an average 10% energy gains over static slowdown. Dynamic procrastination extends the average sleep interval by 25%, which reduces the idle energy consumption by 15%, while meeting all timing requirements. Ravindra Jejurikar, Rajesh K. Gupta 0001 |
DAC | 2 |
| 2005 | Energy Aware Non-Preemptive Scheduling for Hard Real-Time SystemsabstractSlowdown based on dynamic voltage scaling (DVS) provides the ability to perform an energy-delay tradeoff in the system. Nonpreemptive scheduling becomes an integral part of systems where resource characteristics makes preemption undesirable or impossible. We address the problem of energy efficient scheduling of nonpreemptive tasks based on the earliest deadline first (EDF) scheduling policy. We present the stack based slowdown algorithm that builds upon the optimal feasibility test for nonpreemptive systems. We also propose a dynamic stack reclamation policy to further enhance energy savings. Simulation results show on an average 15% energy savings using static slowdown factors and 20% savings with dynamic slowdown, over known slowdown techniques. Ravindra Jejurikar, Rajesh K. Gupta 0001 |
ECRTS | 2 |
| 2005 | Improving SystemC simulation through Petri net reductionsabstractWith the growing acceptance of SystemC in co-design environments there is a need to further improve the simulation performance of complex designs. Our previous work has shown that simulation performance can be improved by carefully restructuring such designs. As a well known formal model for concurrent systems with a good balance between their expressive power and the theoretical results available for correlating structural properties with behavior, free-choice Petri nets were an ideal candidate for formalizing our restructuring technique. To do so we show how SystemC code can be mapped onto such nets followed by how such a labeled net can be reduced in a semantics preserving way. The end result is a restructured design which, as our experiments show, has improved simulation performance over the original models. Nicolae Savoiu, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
MEMOCODE | 3 |
| 2005 | Using probabilistic model checking for dynamic power managementabstractAbstract Dynamic power management (DPM) refers to the use of runtime strategies in order to achieve a tradeoff between the performance and power consumption of a system and its components. We present an approach to analysing stochastic DPM strategies using probabilistic model checking as the formal framework. This is a novel application of probabilistic model checking to the area of system design. This approach allows us to obtain performance measures of strategies by automated analytical means without expensive simulations. Moreover, one can formally establish various probabilistically quantified properties pertaining to buffer sizes, delays, energy usage etc., for each derived strategy. Gethin Norman, David Parker 0001, Marta Z. Kwiatkowska, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
Formal Aspects Comput. | 5 |
| 2005 | Line Size Adaptivity Analysis of Parameterized Loop Nests for Direct Mapped Data CacheabstractCaches are crucial components of modern processors; they allow high-performance processors to access data fast and, due to their small sizes, they enable low-power processors to save energy - by circumventing memory accesses. We examine efficient utilization of data caches in an adaptive memory hierarchy. We exploit data reuse through the static analysis of cache-line size adaptivity. We present an approach that enables the quantification of data misses with respect to cache-line size at compile-time using (parametric) equations, which model interference. Our approach aims at the analysis of perfect loop nests in scientific applications; it is applied to direct mapped cache and it is an extension and generalization of the cache miss equation (CME) proposed by Ghosh et al. (1999). Part of this analysis is implemented in a software package, STAMINA. We present analytical results in comparison with simulation-based methods and we show evidence of both the expressiveness and the practicability of the analysis. Paolo D'Alberto, Alexandru Nicolau, Alexander V. Veidenbaum, Rajesh K. Gupta 0001 |
IEEE Trans. Computers | 4 |
| 2005 | An overview of the competitive and adversarial approaches to designing dynamic power management strategiesabstractDynamic power management (DPM) refers to the problem of judicious application of various low-power techniques based on runtime conditions in an embedded system to minimize the total energy consumption. To be effective, often such decisions take into account the operating conditions and the system-level design goals. DPM has been a subject of intense research in the past decade driven by the need for low power consumption in modern embedded devices. We present a comprehensive overview of two closely related approaches to designing DPM strategies, namely, competitive analysis approach and model checking approach based on adversarial modeling. Although many other approaches exist for solving the system-level DPM problem, these two approaches are closely related and are based on a common theme. This commonality is in the fact that the underlying model is that of a competition between the system and an adversary. The environment that puts service demands on devices is viewed as an adversary, or to be in competition with the system to make it burn more energy, and the DPM strategy is employed by the system to counter that. Sandy Irani, Gaurav Singh 0006, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2005 | Energy-aware wireless systems with adaptive power-fidelity tradeoffsabstractWireless networked embedded systems, such as multimedia terminals, sensor nodes, etc., present a rich domain for making energy/performance/quality tradeoffs based on application needs, network conditions, etc. Energy awareness in these systems is the ability to perform tradeoffs between available battery energy and application quality requirements. In this paper, we show how operating system directed dynamic voltage scaling and dynamic power management can provide for such a capability. We propose a real-time scheduling algorithm that uses runtime feedback about application behavior to provide adaptive power-fidelity tradeoffs. We demonstrate our approach in the context of a static priority-based preemptive task scheduler. Simulation results show that the proposed algorithm results in significant energy savings compared to state-of-the-art dynamic voltage scaling schemes with minimal loss in system fidelity. We have implemented our scheduling algorithm into the eCos real-time operating system running on an Intel XScale-based variable voltage platform. Experimental results obtained using this platform confirm the effectiveness of our technique Vijay Raghunathan, Cristiano Pereira, Mani Srivastava 0001, Rajesh K. Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2004 | Leakage aware dynamic voltage scaling for real-time embedded systemsabstractA five-fold increase in leakage current is predicted with each technology generation. While Dynamic Voltage Scaling (DVS) is known to reduce dynamic power consumption, it also causes increased leakage energy drain by lengthening the interval over which a computation is carried out. Therefore, for minimization of the total energy, one needs to determine an operating point, called the critical speed. We compute processor slowdown factors based on the critical speed for energy minimization. Procrastination scheduling attempts to maximize the duration of idle intervals by keeping the processor in a sleep/shutdown state even if there are pending tasks, within the constraints imposed by performance requirements. Our simulation experiments show that the critical speed slowdown results in up to 5% energy gains over a leakage oblivious dynamic voltage scaling. Procrastination scheduling scheme extends the sleep intervals to up to 5 times, resulting in up to an additional 18% energy gains, while meeting all timing requirements. Ravindra Jejurikar, Cristiano Pereira, Rajesh K. Gupta 0001 |
DAC | 3 |
| 2004 | Proxy-based task partitioning of watermarking algorithms for reducing energy consumption in mobile devicesabstractDigital watermarking is a process that embeds an imperceptible signature or watermark in a digital file containing audio, image, text or video data. The watermark is later used to authenticate the data file and for tamper detection. It is particularly valuable in the use and exchange of digital media such as audio and video on emerging handheld devices. However, watermarking is computationally expensive and adds to the drain of the available energy in handheld devices. We present an approach in which we partition the watermarking embedding and extraction algorithms and migrate some tasks to a proxy server. This leads to a lower energy consumption on the handheld without compromising the security of the watermarking process. Our results show that executing watermarking partitioned between the proxy and the handheld reduces the total energy consumed by 80% over running it only on the handheld and improves performance by over two orders of magnitude. Arun Kejariwal, Alexandru Nicolau, Nikil Dutt, Rajesh K. Gupta 0001 |
DAC | 5 |
| 2004 | Energy-Aware System Design for Wireless MultimediaabstractIn this paper, we present various challenges that arise in the delivery and exchange of multimedia information to mobile devices. Specifically, we focus on techniques for maintaining QoS to end-user multimedia applications (e.g. video streaming, multimedia conferencing) while maximizing device lifetimes. In order to cope with the resource intensive nature of multimedia applications (in terms of computation, bandwidth and consequently power) and dynamic congestion levels in wireless networks, an end-to-end approach to QoS-aware power optimization is required. We discuss the trend towards such an integrated approach that couples the architectural, OS, middleware and application layers to achieve both user experience and device energy gains. We conclude with a discussion of tools for integrated system design and testing that will aid in rapid deployment of wireless multimedia. Hans Van Antwerpen, Nikil Dutt, Rajesh K. Gupta 0001, Shivajit Mohapatra, Cristiano Pereira, Nalini Venkatasubramanian, Ralph von Vignau |
DATE | 3 |
| 2004 | Network Topology Exploration of Mesh-Based Coarse-Grain Reconfigurable ArchitecturesabstractSeveral coarse-grain reconfigurable architectures proposed recently consist of a large number of processing elements (PEs) connected in a mesh-like network topology. We study the effects of three aspects of network topology exploration on the performance of applications on these architectures: (a) changing the interconnection between PEs; (b) changing the way the network topology is traversed while mapping operations to the PEs; and (c) changing the communication delays on the interconnects between PEs. We propose network topology traversal strategies that first schedule PEs that are spatially close and that have more interconnections among them. We use an interconnect aware list scheduling heuristic as a vehicle to perform the network topology exploration experiments on a set of designs derived from DSP applications. Our experimental results show that a spiral traversal strategy, coupled with a two neighbor interconnect topology leads to good performance for the DSP benchmarks considered. Our prototype framework thus provides an exploration environment for system architects to explore and tune coarse-grain reconfigurable architectures for particular application domains. Nikhil Bansal 0003, Nikil Dutt, Alexandru Nicolau, Rajesh K. Gupta 0001 |
DATE | 5 |
| 2004 | Loop Shifting and Compaction for the High-Level Synthesis of Designs with Complex Control FlowabstractEmerging embedded system applications in multimedia and image processing are characterized by complex control flow consisting of deeply nested conditionals and loops. We present a technique called loop shifting that incrementally exploits loop level parallelism across iterations by shifting and compacting operations across loop iterations. Our experimental results show that loop shifting is particularly effective for the synthesis of designs with complex control especially when resource utilization is already high and/or under tight resource constraints. In situations when further loop unrolling (or initiating another iteration of the loop body) leads to a sharp increase in the longest combinational path in the circuit and the circuit area, loop shifting is able to achieve up to 20% reduction in the input-to-output delay in the synthesized circuit. We implemented loop shifting within the SPARK parallelizing high-level synthesis framework and present results for experiments on designs derived from multimedia and image processing applications. Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau |
DATE | 3 |
| 2004 | Optimized Slowdown in Real-Time Task Systems
Ravindra Jejurikar, Rajesh K. Gupta 0001 |
ECRTS | 2 |
| 2004 | Interconnect-Aware Mapping of Applications to Coarse-Grain Reconfigurable Architectures
Nikhil Bansal 0003, Nikil Dutt, Alexandru Nicolau, Rajesh K. Gupta 0001 |
FPL | 5 |
| 2004 | Dynamic voltage scaling for systemwide energy minimization in real-time embedded systemsabstractTraditionally, dynamic voltage scaling (DVS) techniques have focused on minimizing the processor energy consumption as opposed to the entire system energy consumption. The slowdown resulting from DVS can increase the energy consumption of components like memory and network interfaces. Furthermore, the leakage power consumption is increasing with the scaling device technology and must also be taken into account. In this work, we consider energy efficient slowdown in a real-time task system. We present an algorithm to compute task slowdown factors based on the contribution of the processor leakage and standby energy consumption of the resources in the system. Our simulation experiments using randomly generated task sets show on an average 10% energy gains over traditional dynamic voltage scaling. We further combine slowdown with procrastination scheduling which increases the average energy savings to 15%. We show that our scheduling approach minimizes the total static and dynamic energy consumption of the systemwide resources. Ravindra Jejurikar, Rajesh K. Gupta 0001 |
ISLPED | 2 |
| 2004 | Procrastination scheduling in fixed priority real-time systemsabstractProcrastination scheduling has gained importance for energy efficiency due to the rapid increase in the leakage power consumption. Under procrastination scheduling, task executions are delayed to extend processor shutdown intervals, thereby reducing the idle energy consumption. We propose algorithms to compute the maximum procrastination intervals for tasks scheduled by either the fixed priority or the dual priority scheduling policy. We show that dual priority scheduling always guarantees longer shutdown intervals than fixed priority scheduling. We further combine procrastination scheduling with dynamic voltage scaling to minimize the total static and dynamic energy consumption of the system. Our simulation experiments show that the proposed algorithms can extend the sleep intervals up to 5 times while meeting the timing requirements. The results show up to 18% energy gains over dynamic voltage scaling. Ravindra Jejurikar, Rajesh K. Gupta 0001 |
LCTES | 2 |
| 2004 | Formal Refinement Checking in a System-level Design Methodology
Jean-Pierre Talpin, Paul Le Guernic, Sandeep K. Shukla, Frederic Doucet, Rajesh K. Gupta 0001 |
Fundam. Informaticae | 5 |
| 2004 | Using global code motions to improve the quality of results for high-level synthesisabstractThe quality of synthesis results for most high-level synthesis approaches is strongly affected by the choice of control flow (through conditions and loops) in the input description. This leads to a need for high-level and compiler transformations that overcome the effects of programming style on the quality of generated circuits. To address this issue, we have developed a set of speculative code-motion transformations that enable movement of operations through, beyond, and into conditionals with the objective of maximizing performance. We have implemented these code transformations, along with supporting code-motion techniques and variable renaming techniques, in a high-level synthesis research framework called Spark. Spark takes a behavioral description in ANSI-C as input and generates synthesizable register-transfer level VHDL. We present results for experiments on designs derived from three real-life multimedia and image processing applications, namely, the MPEG-1 and -2 and GNU image manipulation program applications. We find that the speculative-code motions lead to reductions between 36% and 59% in the number of states in the finite-state machine (controller complexity) and the cycles on the longest path (performance) compared with the case when only nonspeculative code motions are employed. Also, logic synthesis results show fairly constant critical path lengths (clock period) and a marginal increase in area. Nicolae Savoiu, Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2004 | Guest editorial: Special issue on networked embedded systemsabstractNo abstract available. Rajesh K. Gupta 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2004 | Coordinated parallelizing compiler optimizations and high-level synthesisabstractWe present a high-level synthesis methodology that applies a coordinated set of coarse-grain and fine-grain parallelizing transformations. The transformations are applied both during a pre-synthesis phase and during scheduling, with the objective of optimizing the results of synthesis and reducing the impact of control flow constructs on the quality of results. We first apply a set of source level presynthesis transformations that include common sub-expression elimination (CSE), copy propagation, dead code elimination and loop-invariant code motion, along with more coarse-level code restructuring transformations such as loop unrolling. We then explore scheduling techniques that use a set of aggressive speculative code motions to maximally parallelize the design by re-ordering, speculating and sometimes even duplicating operations in the design. In particular, we present a new technique called "Dynamic CSE" that dynamically coordinates CSE and code motions such as speculation and conditional speculation during scheduling. We implemented our parallelizing high-level synthesis in the SPARK framework. This framework takes a behavioral description in ANSI-C as input and generates synthesizable register-transfer level VHDL. Our results from computationally expensive portions of three moderately complex design targets, namely, MPEG-1, MPEG-2 and the GIMP image processing tool, validate the utility of our approach to the behavioral synthesis of designs with complex control flows. Rajesh K. Gupta 0001, Nikil Dutt, Alexandru Nicolau |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2003 | Typing abstractions and management in a component frameworkabstractWe consider the type inference problems in a compositional design environment where the components are automatically instantiated from pre-existing C++-based intellectual property (IP) libraries. We present a component integration language based on scripting for design specification. Our focus is architectural aspects in specification that uses aggregation- as opposed to the more commonly used inheritance- for composition of components. Our approach simplifies architectural specification by employing a type inference and type management environment. We show that the type inference problem is NP-complete. We present a heuristic based on code generation and parameterization to solve the type inference for IP selection in our C++-based composition environment. We have implemented the composition and type management in the BALBOA framework. The results show the utility of our approach. Frederic Doucet, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
ASP-DAC | 3 |
| 2003 | Formal verification - prove it or pitch itabstractDespite a number of solid advances in simulation and verification techniques over the last twenty years, semiconductor chip designs continue to see large increases in the cost of verification - both in terms of human resources and time. Most of these increases are due to the growing size and complexity of the chip designs. Many of these designs are complete systems in their own right thus enlarging the scope of the verification problem. Formal verification has held out the most promise for reducing the magnitude of the verification task. Indeed, most major microprocessor teams - at IBM, Intel and Motorola - have routinely hosted formal verification experts since the early '90s. ASIC vendors and their tool providers have been closely following these developments into a number of initiatives and new startup companies driven by that very promise of formal verification. Despite these developments, simulation continues to be the final source of signoff - if not confidence - in chip tapeouts. Why is this so? Formal verification is an important technology to be left at the margins of the validation task. Will formal verification eliminate or limit unit level verification and provide the necessary glue for a realistic validation flow? Will the testbenches be replaced by constraints and assertions? Can validation effort be reused? This panel will explore the issues related to building practical validation flows, and the technologies that the designer community can realistically look forward to materializing in their lifetimes. Rajesh K. Gupta 0001, Shishpal Rawat, Sandeep K. Shukla, Brian Bailey, Daniel K. Beece, Carl Pixley, John O'Leary, Fabio Somenzi |
DAC | 1 |
| 2003 | A survey of techniques for energy efficient on-chip communicationabstractInterconnects have been shown to be a dominant source of energy consumption in modern day System-on-Chip (SoC) designs. With a large (and growing) number of electronic systems being designed with battery considerations in mind, minimizing the energy consumed in on-chip interconnects becomes crucial. Further, the use of nanometer technologies is making it increasingly important to consider reliability issues during the design of SoC communication architectures. Continued supply voltage scaling has led to decreased noise margins, making interconnects more susceptible to noise sources such as crosstalk, power supply noise, radiation induced defects, etc. The resulting transient faults cause the interconnect to behave as an unreliable transport medium for data signals. Therefore, fault tolerant communication mechanisms, such as Automatic Repeat Request (ARQ), Forward Error Correction (FEC), etc., which have been widely used in the networking community, are likely to percolate to the SoC domain.This paper presents a survey of techniques for energy efficient on-chip communication. Techniques operating at different levels of the communication design hierarchy are described, including circuit-level techniques, such as low voltage signaling, architecture-level techniques, such as communication architecture selection and bus isolation, system-level techniques, such as communication based power management and dynamic voltage scaling for interconnects, and network-level techniques, such as error resilient encoding for packetized on-chip communication. Emerging technologies, such as Code Division Multiple Access (CDMA) based buses, and wireless interconnects are also surveyed. Vijay Raghunathan, Mani Srivastava 0001, Rajesh K. Gupta 0001 |
DAC | 3 |
| 2003 | Introspection in System-Level Language Frameworks: Meta-Level vs. Integrated
Frederic Doucet, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2003 | Dynamic Conditional Branch Balancing during the High-Level Synthesis of Control-Intensive Designs
Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau |
DATE | 3 |
| 2003 | Polychrony for Refinement-Based Design
Jean-Pierre Talpin, Paul Le Guernic, Sandeep K. Shukla, Rajesh K. Gupta 0001, Frederic Doucet |
DATE | 4 |
| 2003 | Formal Methods for Dynamic Power Management
Rajesh K. Gupta 0001, Sandy Irani, Sandeep K. Shukla |
ICCAD | 1 |
| 2003 | Interface Synthesis using Memory Mapping for an FPGA PlatformabstractSeveral system-on-chip (SoC) platforms have recently emerged that use reconfigurable logic (FPGAs) as a programmable coprocessor to reduce the computational load on the main processor core. We present an interface synthesis approach that forms part of our hardware-software codesign methodology for such an FPGA-based platform. The approach is based on a novel memory mapping algorithm that maps data used by both the hardware and the software to shared memories on the reconfigurable fabric. The memory mapping algorithm couples with a high-level synthesis tool and uses scheduling information to map variables, arrays and complex data structures to the shared memories in a way that minimizes the number of registers and multiplexers used in the hardware interface. We also present three software schemes that enable the application software to communicate with this hardware interface. We demonstrate the utility of our approach and study the trade-offs involved using a case study of the codesign of a computationally expensive portion of the MPEG-1 multimedia application on to the Altera Nios platform. Manev Luthra, Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau |
ICCD | 4 |
| 2003 | Should the space of implementation possibilities be determined by the abilities of high-level synthesis and validation?
Rajesh K. Gupta 0001, Sandeep K. Shukla |
MEMOCODE | 1 |
| 2003 | Algorithms for power savings
Sandy Irani, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
SODA | 3 |
| 2003 | BALBOA: a component-based design environment for system modelsabstractThis paper presents the BALBOA component composition framework for system-level architectural design. It has three parts: a loosely-typed component integration language (CIL); a set of C++ intellectual property (IP) component libraries; and a set of split-level interfaces (SLIs) to link the two. A CIL component interface can be mapped to many different C++ component implementations. A type-inference system maps all weakly-typed CIL interfaces to strongly typed C++ component implementations to produce an executable architectural model. Thus, this amounts to selecting IP implementations according to a set of connection constraints. The SLIs are used to select, adapt, and validate the implementation types. The advantage of using the CIL is that the design description sizes are much smaller because the runtime infrastructure automatically selects the IP and communication implementations. The type inference facilitates changes by automatically propagating them through the design structure. We show that the inference problem is NP complete and we present a heuristic solution to the problem. We bring forth a number of issues related to the automation of reusable IP composition including type- compatibility checking, split-programming, and introspective composition environment, and demonstrate their utility through design examples. Frederic Doucet, Sandeep K. Shukla, Masato Otsuka, Rajesh K. Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2003 | Online strategies for dynamic power management in systems with multiple power-saving statesabstractOnline dynamic power management (DPM) strategies refer to strategies that attempt to make power-mode-related decisions based on information available at runtime. In making such decisions, these strategies do not depend upon information of future behavior of the system, or any a priori knowledge of the input characteristics. In this paper, we present online strategies, and evaluate them based on a measure called the competitive ratio that enables a quantitative analysis of the performance of online strategies. All earlier approaches (online or predictive) have been limited to systems with two power-saving states (e.g., idle and shutdown). The only earlier approaches that handled multiple power-saving states were based on stochastic optimization. This paper provides a theoretical basis for the analysis of DPM strategies for systems with multiple power-down states, without resorting to such complex approaches. We show how a relatively simple "online learning" scheme can be used to improve the competitive ratio over deterministic strategies using the notion of "probability-based" online DPM strategies. Experimental results show that the algorithm presented here attains the best competitive ratio in comparison with other known predictive DPM algorithms. The other algorithms that come close to matching its performance in power suffer at least an additional 40% wake-up latency on average. Meanwhile, the algorithms that have comparable latency to our methods use at least 25% more power on average. Sandy Irani, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2002 | Energy aware task scheduling with task synchronization for embedded real time systemsabstractSlowdown factors determine the extent of slowdown a computing system can experience based on functional and performance requirements. Dynamic Voltage Scaling (DVS) of a processor based on slowdown factors can lead to considerable energy savings. The problem of DVS in the presence of task synchronization has not yet been addressed. We compute slowdown factors for tasks which synchronize for access to shared resources. Tasks synchronize to enforce mutually exclusive access to these resources and can be blocked by lower priority tasks. We compute static slowdown factors for the tasks which guarantee meeting all the task deadlines. Our simulation experiments show on an average 25% energy gains over the known slowdown techniques. Ravindra Jejurikar, Rajesh K. Gupta 0001 |
CASES | 2 |
| 2002 | Coordinated transformations for high-level synthesis of high performance microprocessor blocksabstractHigh performance microprocessor designs are partially characterized by functional blocks consisting of a large number of operations that are packed into very few cycles (often single-cycle) with little or no resource constraints but tight bounds on the cycle time. Extreme parallelization, conditional and speculative execution of operations is essential to meet the processor performance goals. However, this is a tedious task for which classical high-level synthesis (HLS) formulations are inadequate and thus rarely used. In this paper, we present a new methodology for application of HLS targeted to such microprocessor functional blocks that can potentially speed up the design space exploration for microprocessor designs. Our methodology consists of a coordinated set of source-level and fine-grain parallelizing compiler transformations that targets these behavioral descriptions, specifically loop constructs in them and enables efficient chaining of operations and high-level synthesis of the functional blocks. As a case study in understanding the complexity and challenges in the use of HLS, we walk the reader through the detailed design of an instruction length decoder drawn from the Pentium-family of processors. The chief contribution of this paper is formulation of a domain-specific methodology for application of high-level synthesis techniques to a domain that rarely, if ever, finds use for it. Nicolae Savoiu, Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau, Timothy Kam, Michael Kishinevsky, Shai Rotem |
DAC | 4 |
| 2002 | Profile-Based Dynamic Voltage Scheduling Using Program CheckpointsabstractDynamic voltage scaling (DVS) is a known effective mechanism for reducing CPU energy consumption without significant performance degradation. While a lot of work has been done on inter-task scheduling algorithms to implement DVS under operating system control, new research challenges exist in intra-task DVS techniques under software and compiler control. In this paper we introduce a novel intra-task DVS technique under compiler control using program checkpoints. Checkpoints are generated at compile time and indicate places in the code where the processor speed and voltage should be re-calculated. Checkpoints also carry user-defined time constraints. Our technique handles multiple intra-task performance deadlines and modulates power consumption according to a run-time power budget. We experimented with two heuristics for adjusting the clock frequency and voltage. For the particular benchmark studied, one heuristic yielded 63% more energy savings than the other. With the best of the heuristics we designed, our technique resulted in 82% energy savings over the execution of the program without employing DVS. Ilya Issenin, Radu Cornea, Rajesh K. Gupta 0001, Nikil Dutt, Alexander V. Veidenbaum, Alexandru Nicolau |
DATE | 4 |
| 2002 | An Environment for Dynamic Component Composition for Efficient Co-Design abstractThis paper describes the Balboa component integration environment that is composed of three parts: a script language interpreter, compiled C++ components, and a set of split-level interfaces to link the interpreted domain to the compiled domain. The environment applies the notion of split-level programming to relieve system engineers of software engineering concerns and to let them focus on system architecture. The script language is a Component Integration Language (CIL) because it implements a component model with introspection and loose typing capabilities. Component wrappers use split-level interfaces that implement the composition rules, dynamic type determination and type inference algorithms. Using an interface description language compiler automatically generates the split-level interfaces. The contribution of this work is two fold: an active code generation technique, and a three-layer environment that keeps the C++ components intact for reuse. We present an overview of the environment, demonstrate our approach by building three simulation models for an adaptive memory controller, and comment on code generation ratios. Frederic Doucet, Sandeep K. Shukla, Rajesh K. Gupta 0001, Masato Otsuka |
DATE | 3 |
| 2002 | Competitive Analysis of Dynamic Power Management Strategies for Systems with Multiple Power Savings StatesabstractWe present strategies for "online" dynamic power management (DPM) based on the notion of the competitive ratio that allows us to compare the effectiveness of algorithms against an optimal strategy. This paper makes two contributions: it provides a theoretical basis for the analysis of DPM strategies for systems with multiple power down states; and provides a competitive algorithm based on probabilistically generated inputs that improves the competitive ratio over deterministic strategies. Experimental results show that our probability-based DPM strategy improves the efficiency of power management over the deterministic DPM strategy by 25%, bringing the strategy to within 23% of the optimal offline DPM. Sandy Irani, Rajesh K. Gupta 0001, Sandeep K. Shukla |
DATE | 2 |
| 2002 | Automated Concurrency Re-Assignment in High Level System Models for Efficient System-Level SimulationabstractSimple and powerful modeling of concurrency and reactivity along with their efficient implementation in the simulation kernel are crucial to the overall usefulness of system level models using the C++-based modeling frameworks. However the concurrency alignment in most modeling frameworks is naturally expressed along hardware units, being supported by the various language constructs, and the system designers express concurrency in their system models by providing threads for some modules/units of the model. Our experimental analysis shows that this concurrency model leads to inefficient simulation performance, and a concurrency alignment along dataflow gives much better simulation performance, but changes the conceptual model of hardware structures. As a result, we propose an algorithmic transformation of designs written in these C++-based environments with concurrency alignment along units/modules. This transformation, provided as a compiler front-end, will re-assign the concurrency along the dataflow, as opposed to threading along concurrent hardware/software modules, keeping the functionality of the model unchanged. Such a front-end transformation strategy will relieve hardware system designers from concerns about software engineering issues such as, threading architecture, and simulation performance, while allowing them to design in the most natural manner whereas, the simulation performance can be enhanced up to almost two times as shown in our experiments. Nicolae Savoiu, Sandeep K. Shukla, Rajesh K. Gupta 0001 |
DATE | 3 |
| 2002 | Power Savings in Embedded Processors through Decode Filer CacheabstractIn embedded processors, instruction fetch and decode can consume more than 40% of processor power. An instruction filter cache can be placed between the CPU core and the instruction cache to service the instruction stream. Power savings in instruction fetch result from accesses to a small cache. In this paper, we introduce a decode filter cache to provide a decoded instruction stream. On a hit in the decode filter cache, fetching from the instruction cache and the subsequent decoding is eliminated, which results in power savings in both instruction fetch and instruction decode. We propose to classify instructions into cacheable or uncacheable depending on the decoded width. Then sectored cache design is used in the decode filter cache so that cacheable and uncacheable instructions can coexist in a decode filter cache sector. Finally, a prediction mechanism is presented to reduce the decode filter cache miss penalty. Experimental results show average 34% processor power reduction and less than 1% performance degradation. Weiyu Tang, Rajesh K. Gupta 0001, Alexandru Nicolau |
DATE | 2 |
| 2002 | Structured Component Composition Frameworks for Embedded System Design
Sandeep K. Shukla, Frederic Doucet, Rajesh K. Gupta 0001 |
HiPC | 3 |
| 2002 | An analysis of system level power management algorithms and theireffects on latencyabstractThe problem of power management for an embedded system is to reduce system level power dissipation by shutting off parts of the system when they are not being used and turning them back on when requests have to be serviced. Algorithms for this problem are online in nature; the algorithm must operate only with access to data that it has seen so far and without access to the complete data set or its characteristics. We present online algorithms to manage power for embedded systems and discuss their effects on system latency. We introduce competitive analysis as a formal framework for the evaluation of various power management algorithms. Competitive analysis does not depend on the distribution of interarrival times of requests. We present a nonadaptive online algorithm, analyze its behavior, and show that it is optimal. We also present a lower bound on the competitiveness of any adaptive algorithm. We show that no adaptive online algorithm can dissipate less than about 1.6 times the power dissipated by the optimal offline algorithm in the worst case. We also show that in order for any online algorithm to achieve this lower bound, it may have to maintain a complete history of the interarrival. times of the requests in the input sequence. Since this is not practical, we present a simple algorithm that uses only the last interarrival time to predict the arrival of the next request. Dinesh Ramanathan, Sandy Irani, Rajesh K. Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2001 | Panel: The Next HDL: If C++ is the Answer, What was the Question?abstractThe focus of this panel is on issues surrounding the use of C++ in modeling, integration of silicon IP and system-on-chip designs. In the last two years there have been several announcements promoting C++ based solutions and of multiple consortia (SystemC, Cynapps, Accellera, SpecC) that represent increasing commercial interest both from tool vendors as well as perhaps expression of genuine needs from the design houses. There are, however, serious questions about what value proposition does a C++ based design methodology bring to the IC or system designer? What has changed in the modeling technology (and/or available tools) that gives a new capability? Is synthesis the right target? or VAlidation? Tester modeling or testbench generation? This panel brings together advocates and opponents from the user community to highlight the achievements and the challenges that remain in use C++ for use in microelectronic circuits and systems. Rajesh K. Gupta 0001, Shishpal Rawat, Ingrid Verbauwhede, Gérard Berry, Ramesh Chandra, Daniel Gajski, Kris Konigsfeld, Patrick Schaumont |
DAC | 1 |
| 2001 | Speculation Techniques for High Level Synthesis of Control Intensive DesignsabstractThe quality of synthesis results for most high level synthesis approaches is strongly affected by the choice of control flow (through conditions and loops) in the input description. In this paper, we explore the effectiveness of various types of code motions, such as moving operations across conditionals, out of conditionals (speculation) and into conditionals (reverse speculation), and how they can be effectively directed by heuristics so as to lead to improved synthesis results in terms of fewer execution cycles and fewer number of states in the finite state machine controller. We also study the effects of the code motions on the area and latency of the final synthesized netlist. Based on speculative code motions, we present a novel way to perform early condition execution that leads to significant improvements in highly control-intensive designs. Overall, reductions of up to 38 \% in execution cycles are obtained with all the code motions enabled. Nicolae Savoiu, Nikil Dutt, Rajesh K. Gupta 0001, Alexandru Nicolau |
DAC | 5 |
| 2001 | Design of a Predictive Filter Cache for Energy Savings in High Performance Processor ArchitecturesabstractFilter cache has been proposed as an energy saving architectural feature. A filter cache is placed between the CPU and the instruction cache (I-cache) to provide the instruction stream. Energy savings result from accesses to a small cache. There is however loss of performance when instructions are not found in the filter cache. The majority of the energy savings from the filter cache are due to the temporal reuse of instructions in small loops. We examine subsequent fetch addresses to predict whether the next fetch address is in the filter cache dynamically. In case a miss is predicted, we reduce miss penalty by accessing the I-cache directly. Experimental results show that our next fetch prediction reduces performance penalty by more than 91% and is more energy efficient than a conventional filter cache. Average I-cache energy savings of 31 % can be achieved by our filter cache design with around 1 % performance degradation. Weiyu Tang, Rajesh K. Gupta 0001, Alexandru Nicolau |
ICCD | 2 |
| 2001 | Guest editorial reconfigurable and adaptive VLSI systems
Ranga Vemuri, Rajesh K. Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Timing driven co-design of networked embedded systemsabstractAdvances in microelectronics integration have led to emergence of tightly integrated systems with high performance network interfaces. Design of such systems especially for single chip implementation is a delicate balance of functionality and available time budget to perform the tasks. Computer-aided design tools and methodologies are needed to ensure the correctness of the design and eciency of the design process, especially for networked systems that have strict timing requirements both due to technology as well as networking needs. We present an overview of a timing-driven design methodology for networked systems, developed at the University of California, Irvine. I. Introduction Networked embedded systems (NES) are distributed embedded systems connected together using network interfaces and standardized protocols. The protocols are often implemented partly in hardware and partly in software. NES are usually designed at the highest level of abstraction as a collection of basic net... Dinesh Ramanathan, Ravindra Jejurikar, Rajesh K. Gupta 0001 |
ASP-DAC | 3 |
| 2000 | Design and implementation of a hierarchical exception handling extension to systemCabstractArticle Design and implementation of a hierarchical exception handling extension to systemC Share on Authors: Prashant Arora Department of Information and Computer, Science, University of California, Irvine, Irvine, CA Department of Information and Computer, Science, University of California, Irvine, Irvine, CAView Profile , Rajesh K. Gupta Department of Information and Computer, Science, University of California, Irvine, Irvine, CA Department of Information and Computer, Science, University of California, Irvine, Irvine, CAView Profile Authors Info & Claims CASES '00: Proceedings of the 2000 international conference on Compilers, architecture, and synthesis for embedded systemsNovember 2000 Pages 80–84https://doi.org/10.1145/354880.354892Online:01 November 2000Publication History 1citation316DownloadsMetricsTotal Citations1Total Downloads316Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Prashant Arora, Rajesh K. Gupta 0001 |
CASES | 2 |
| 2000 | Analysis of High-Level Address Code Transformations for Programmable ProcessorsabstractMemory intensive applications require considerable arithmetic for the computation and selection of the different memory access pointers. These memory address calculations often involve complex (non) linear arithmetic expressions which have to be calculated during program execution under tight timing constraints, this becoming a critical bottleneck in the overall system performance. This paper explores applicability and effectiveness of source-level optimisations (as opposed to instruction-level) for address computations in the context of multimedia. We propose and evaluate two processor-target independent source-level optimisation techniques, namely, global scope operation cost minimisation complemented with loop-invariant code hoisting, and nonlinear operator strength reduction. The transformations attempt to achieve minimal code execution within loops and reduced operator strengths. The effectiveness of the transformations is demonstrated with two real-life multimedia application kernels by comparing the improvements in the number of execution cycles, before and after applying the systematic source-level optimisations. Using state-of-the-art C compilers on several popular RISC platforms. Rajesh K. Gupta 0001, Miguel Corbalan, Francky Catthoor |
DATE | 2 |
| 2000 | System Level Online Power Management AlgorithmsabstractThe problem of power management for an embedded system is to reduce system level power dissipation by shutting off parts of the system when they are not being used and turning them back on when they are required. Algorithms for this problem are online in nature where the algorithm must operate without access to the complete data set or its characteristics. In this paper, we present online algorithms to manage power for embedded systems and provide experimental analysis to back up the theoretical results. Specifically, this paper makes four contributions. We propose an optimal online algorithm for power management. We present an analysis of algorithmic efficiency using a technique called competitive analysis which is particularly suitable for online algorithms. Using this analysis technique, we develop a lower bound for the non-adaptive version of the power management problem and show that our algorithm achieves this lower bound. Next, we explore adaptive algorithms that try to shut down the system based on historical data. We provide a lower bound for any algorithm that uses adaptive methods to manage power. We also propose an algorithm that is independent of the input data distribution, practical and usable in both hardware and software systems with guaranteed performance. Finally, we compare these algorithms with previously proposed heuristics both theoretically and experimentally. For the experiments, we model the disk drive of a laptop computer as an embedded system. The results show that the proposed algorithms perform well in practice with guaranteed bounds on their performance. Further, this paper conclusively demonstrates that to implement aggressive power management techniques for power critical subsystems, designers will have to commit greater resources such as dedicated registers and ALU units. Dinesh Ramanathan, Rajesh K. Gupta 0001 |
DATE | 2 |
| 2000 | Latency Effects of System Level Power Management AlgorithmsabstractA power management algorithm for an embedded system reduces system level power dissipation by shutting off parts of the system when they are not being used and turning them back on when they are required. Algorithms for this problem are online in nature since they must operate without knowledge of the arrival time or service requirements of future requests. In this paper, we present online algorithms to manage power for embedded systems. We perform an empirical analysis of these algorithms and give theoretical justification for the empirical results. Effective power management strategies have an adverse impact on the latency of the system for which the strategy is designed. Typically, the more aggressive the power management scheme, the greater the increase in the latency of the system. In this paper, we prove an upper bound on the additional latency of the system introduced by power management strategies. Moreover, we show that this upper bound occurs each time the system is shutdown and hence is an important system design parameter. In addition, service time and latencies have an effect on power management strategies since they alter the length and occurrences of idle periods which. We study this phenomenon experimentally, by modeling the disk drive of a laptop computer as an embedded system. The results show that if service times of arriving requests are modeled, the relative performance of algorithms can change leading to non-adaptive algorithms performing better than adaptive ones. We compare the performance of adaptive and non-adaptive power management algorithms. In particular, our experimental results show that an "immediate" shutdown strategy that shuts down the system whenever it encounters an idle period performs surprising better than sophisticated adaptive algorithms suggested in the literature. We provide an analytical explanation for the effectiveness of power management strategies. Dinesh Ramanathan, Sandy Irani, Rajesh K. Gupta 0001 |
ICCAD | 3 |
| 2000 | Interfacing Hardware and Software Using C++ Class LibrariesabstractAs chip capacity increases and system-on-a-chip becomes more than just a catch phrase, hardware and system design are being driven in new directions. Systems are designed not just as hardware, but as a tightly coupled combination of both hardware and software. C++, extended with class libraries, is emerging as the way to design such complex systems. This paper proposes methods to specify and refine designs from a purely software description at a functional level to a level where the hardware components are encapsulated as objects and their interfaces clearly defined. The entire system functionality is described in C++ using some of the commercially available class libraries like SystemC from Synopsys and Cynlib from CynApps. We propose a methodology where a designer can migrate software functionality. We also show that the software driver for the hardware device is generated as a side-effect of the interface refinement process. We demonstrate our methodology on the design of a fax machine from a purely software description of the system. We refine the design by implementing its decoder functionality in hardware and interfacing it with the software encoder. Dinesh Ramanathan, Rajesh K. Gupta 0001, Raymond Roth |
ICCD | 2 |
| 2000 | HDL presynthesis optimizations using a tabular modelabstractIn this paper, we introduce presynthesis optimizations on hardware description languages (HDLs). Presynthesis optimizations consist of two categories of tasks: 1) source-level transformations, which produce optimized behavioral HDL descriptions that lead to improved synthesis results and 2) source-level analysis, which produces information useful in the synthesis stage to improve the quality of the synthesized circuits. Presynthesis optimizations are carried out on an intermediate tabular representation called timed decision table (TDT). We have implemented the TDT-based presynthesis optimization algorithms in a software package called Pumpkin. Experiments running Pumpkin on named benchmarks show promising results. Jian Li 0061, Rajesh K. Gupta 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Efficient Algorithms for Optimum Cycle Mean and Optimum Cost to Time Ratio ProblemsabstractThe goal of this paper is to identify the most efficient algorithms for the optimum mean cycle and optimum cost to time ratio problems and compare them with the popular ones in the CAD community.These problems have numerous important applications in CAD, graph theory, discrete event system theory, and manufacturing systems.In particular, they are fundamental to the performance analysis of digital systems such as synchronous, asynchronous, dataflow, and embedded real-time systems.For instance, algorithms for these problems are used to compute the cycle period of any cyclic digital system.Without loss of generality, we discuss these algorithms in the context of the minimum mean cycle problem (MCMP).We performed a comprehensive experimental study of ten leading algorithms for MCMP.We programmed these algorithms uniformly and efficiently.We systematically compared them on a test suite composed of random graphs as well as benchmark circuits.Above all, our results provide important insight into the performance of these algorithms in practice.One of the most surprising results of this paper is that Howard's algorithm, known primarily in the stochastic control community, is by far the fastest algorithm on our test suite although the only known bound on its running time is exponential.We provide two stronger bounds on its running time. Ali Dasdan, Sandy Irani, Rajesh K. Gupta 0001 |
DAC | 3 |
| 1999 | Adapting cache line size to application behaviorabstractA cache line size has a significant effect on miss rate and memory traffic. Today's computers use a fixed line size, typically 32B, which may not be optimal for a given application. Optimal size may also change during application execution. This paper describes a cache in which the line (fetch) size is continuously adjusted by hardware based on observed application accesses to the line. The approach can improve the miss rate, even over the optimal for the fixed line size, as well as significantly reduce the memory traffic. Alexander V. Veidenbaum, Weiyu Tang, Rajesh K. Gupta 0001, Alexandru Nicolau, Xiaomei Ji |
International Conference on Supercomputing | 3 |
| 1999 | Extraction of functional regularity in datapath circuitsabstractDatapath circuits exhibit a very high degree of regularity, which is exploited by designers to generate layouts with a high density and performance as well as to reduce the overall design effort. Regularity in a datapath circuit manifests itself at functional, structural, and topological levels. Functional regularity of a circuit implies the existence of logically equivalent subcircuits-a common feature of datapath circuits. We present a new and comprehensive approach to extract functional regularity for datapath circuits from their high-level or gate-level descriptions. The key step is the generation of a large set of templates, where a template is a subcircuit with multiple instances in the circuit. Two novel template generation algorithms are presented-one for templates with a tree structure, and the other for a special class of multioutput templates, called single-principal-output-graph (SPOG) templates, where all outputs of a template are in the transitive fanin of a particular output. The set of templates generated is shown to be complete under a few simplifying, yet practical, assumptions, which is key in obtaining a desirable cover of the circuit using templates. We present a few extensions to our regularity extraction approach to demonstrate its generality; these extensions include hierarchical representation of regularity and generation of instances of user-specified templates. We show that the generation of the above two classes of templates results in good covers for datapath circuits with a regular bus structure, including several International Conference on Computer-Aided Design benchmark circuits. The regularity extracted from these circuits can be used to easily understand their structure. We have successfully used our approach to identify bit slices of very large datapath circuits from general-purpose microprocessors. Amit Chowdhary, Sudhakar Kale, Phani K. Saripella, Naresh Sehgal, Rajesh K. Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 1998 | Rate Derivation and Its Applications to Reactive, Real-Time Embedded SystemsabstractAn embedded system (the system) continuously interacts with its environment under strict timing constraints, called the external constraints, and it is important to know how these external constraints translate to time budgets, called the internal constraints, on the tasks of the system. Knowing these time budgets reduces the complexity of the system's design and validation problem and helps the designers have a simultaneous control on the system's functional as well as temporal correctness from the beginning of the design ow. The translation is carried out by rst deriving the rate of each task in the system, hence the term \\rate derivation", using the system's task structure and the rates of the input stimuli coming into the system from its environment. The derived task rates are later used to derive and validate the rest of the internal as well as external constraints. This paper proposes a general task graph model to represent the system's task structure, techniques for deriving and validating the system's timing constraints, and a hardware/software codesign methodology that puts everything together. 1 1 Ali Dasdan, Dinesh Ramanathan, Rajesh K. Gupta 0001 |
DAC | 3 |
| 1998 | An Algorithm To Determine Mutually Exclusive Operations In Behavioral DescriptionsabstractScheduling and binding are two major tasks in architectural synthesis from behavioral descriptions. The information about the mutually exclusive pairs of operations is very useful in reducing both the total delay of the schedule and the resource usage in the final circuit implementation. In this paper, we present an algorithm to identify the largest set of mutually exclusive operation pairs in behavioral descriptions. Our algorithm uses data-flow analysis on a tabular model of system functionality, and is shown to work better than the existing methods for identifying mutually exclusive operations. Jian Li 0061, Rajesh K. Gupta 0001 |
DATE | 2 |
| 1998 | A general approach for regularity extraction in datapath circuitsabstractInmajority ojhigh-performance custom ICdesigns, designers tuke advantage of the high degTee of Regularitypresent in circuits to generate eficient layouts in terms area and perfomance aswellas to Teducethe design effort.Inthis paper, we present a general and comprehensive approach to extTact functional regularity for datapath ciTcuits from their behavioralor structural HDL descn.ptions.The fundamental step is the generation ofa large set of templates, where a template as a subcircuit with multiple instances in the ciTcuit.Two novel template generation algorithms aTe presented -one fortemplates with a tree structure, andtheother foraspecial class of multi-output templates, called single-p n.ncipaloutput (single-PO) templates, wheTe all outputs of a template are in thetransitive faninof a particular output.The set of template$ generated is complete under a few simplifying, yet practical, assumptions.This is key to obtaining a desirable couerof theciTcuit using templates.Weshow that excellent cover$ are obtained for van"ous circuits, including ISCAS benchmarks.We also demonstrate that the regularitg extracted for these circuits can be used to understand their under[ging structure.We have successfully used our approach to identify bit slices of very large datapath circuits from general-purpose microprocessors. Amit Chowdhary, Sudhakar Kale, Phani K. Saripella, Naresh Sehgal, Rajesh K. Gupta 0001 |
ICCAD | 5 |
| 1998 | Faster maximum and minimum mean cycle algorithms for system-performance analysisabstractMaximum and minimum mean cycle problems are important problems with many applications in performance analysis of synchronous and asynchronous digital systems including rate analysis of embedded systems, in discrete-event systems, and in graph theory. Karp's algorithm is one of the fastest and most common algorithms for these problems. We present this paper mainly in the context of the maximum mean cycle problem. We show that Karp's algorithm processes more nodes and arcs than needed to find the maximum cycle mean of a digraph. This observation motivated us to propose a new graph-unfolding scheme that remedies this deficiency and leads to two faster algorithms with different characteristics. Theoretical analysis tells us that our algorithms always run faster than Karp's algorithm and that they are among the fastest to date. Experiments on small benchmark graphs confirm this fact for most of the graphs. These algorithms have been used in building a framework for analysis of timing constraints for embedded systems. Ali Dasdan, Rajesh K. Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | A timing-driven design and validation methodology for embedded real-time systemsabstractWe address the problem of timing constraint derivation and validation for reactive and real-time embedded systems. We assume that such a system is structured into its tasks, and the structure is modeled using a task graph. Our solution uses the timing behavior committed by the environment to the system first to derive the timing constraints on the system's internal behavior and then use them to derive and validate the timing constraints on the system's external behavior. Our solution consists of the following contributions: a generalized task graph model, a comprehensive classification of timing constraints, algorithms for derivation and validation of timing constraints of the system modeled in the generalized task graph model, a codesign methodology that combines the model and the algorithms, and the implementation of this methodology in a tool called RADHA-RATAN. The main advantages of our solution are that it simplifies the problem of ensuring timing correctness of the system by reducing the complexity of the problem from system level to task level, and that it makes the codesign methodology timing-driven in that our solution makes it possible to maintain a handle on the system's timing correctness from very early stages in the system's design flow. Ali Dasdan, Dinesh Ramanathan, Rajesh K. Gupta 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 1998 | Rate analysis for embedded systemsabstractEmbedded systems consist of interacting components that are required to deliver a specific functionality under constraints on execution rates and relative time separation of the components. In this article, we model an embedded system using concurrent processes interacting through synchronization. We assume that there are rate constraints on the execution rates of processes imposed by the designer or the environment of the system, where the execution rate of a process is the number of its executions per unit time. We address the problem of computing bounds on the execution rates of processes constituting an embedded system, and propose an interactive rate analysis framework. As part of the rate analysis framework we present an efficient algorithms for checking the consistency of the rate constraints. Bounds on the execution rate of each process are computed using an efficient algorithm based on the relationship between the execution rate of a process and the maximum mean delay cycles in the process graph. Finally, if the computed rates violate some of the rate constraints, some of the processes in the system are redesigned using information from the rate analysis step. This rate analysis framework is implemented in a tool called RATAN. We illustrate by an example how RATAN can be used in an embedded system design. Anmol Mathur, Ali Dasdan, Rajesh K. Gupta 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 1997 | A procedure for software synthesis from VHDL modelsabstractAddresses the problem of software generation from a hardware description language (HDL). In particular, we examine the issues involved in translating VHDL into C or C++ for use in system simulation and cosynthesis. Because of the concurrency supported by VHDL, and a notion of timing behavior, care must be taken to ensure behavioral correctness of the generated software. The issues involved are shown to be different in each of the application areas. The ideas set forth in this paper have been used in an efficient VHDL simulator designed to execute on multiprocessor systems. Results are presented for simulation on uniprocessor as well as multiprocessor systems. Venkatram Krishnaswamy, Rajesh K. Gupta 0001, Prithviraj Banerjee |
ASP-DAC | 2 |
| 1997 | Data-Flow Assisted Behavioral Partitioning for Embedded SystemsabstractIn this paper we present a novel compiler-directed approach tosystem-level partitioning for a given description of system functionalityin a hardware description language (HDL). The algorithmis based on a definition-use analysis of the storage in the systemmodel to ensure that the resulting portions can be implemented in aloosely-coupled multi-rate execution model with minimal synchronizationbetween the portions. The cost of the hardware-softwareinterface, in terms of amount of buffering required, is computed accuratelyas a part of the partitioning cost function using a data-flowreaching analysis. The proposed algorithm has been implementedand experimental results show up to 65% improvement in buffersizes over the min-cut partitioning algorithm. Samir Agrawal, Rajesh K. Gupta 0001 |
DAC | 2 |
| 1997 | Limited Exception Modeling and Its Use in Presynthesis Optimizationsabstract# In behavioral descriptions# statements that allow limited control jumps# suchasVerilog disable statements on named blocks# are often used for describ# ing system behavior in presence of exceptions. In this pa# per# we extend Timed Decision Tables #TDT## a tabular behavior model# to represent more general control struc# tures including exceptions. Weintroduce the notion of action sharing that allows us to reduce resource require# ments using existing high#level synthesis tools on descrip# tions with control exceptions. We present presynthesis algorithms that work on the extended TDT model and an algorithm that performs action sharing in TDT models. Our experiments on well#known HardwareC benchmarks show size reduction resulting from sharing actions in the input behavioral descriptions. 1 Introduction Behavioral descriptions in HDLs use control##ow con# structs such as conditional branches and loops. In ad# dition to the normal control #ow constructs suchas if statementorwhile loop#... Jian Li 0061, Rajesh K. Gupta 0001 |
DAC | 2 |
| 1997 | An Efficient Implementation of Reactivity for Modeling Hardware in the Scenic Design EnvironmentabstractReactivity is one of the key features of hardware description languages. We present an efficient implementation of reactivity in the Scenic framework that allows the system designer to model hardware blocks. Scenic allows the designer to use C++ to model mixed hardware--software systems with a C++ compiler and a small library and without the need of a complex event-driven run-time kernel often found embedded in hardware description languages (HDL) such as VHDL and Verilog. Moreover, Scenic hardware descriptions can be easily mapped to HDL and synthesized into hardware implementations using commercially available tools. In this paper we present Scenic's implementation of concurrency (signals and processes) and reactivity (waiting and watching). When C++ is used as an HDL, context-switching overhead can become a significant performance issue during simulation. We introduce the notion of delayed expression objects, or lambdas, to reduce context-switching. Examples and experimental results ... Stan Y. Liao, Steven W. K. Tjiang, Rajesh K. Gupta 0001 |
DAC | 3 |
| 1997 | Design technology for building wireless systems (tutorial)
Rajesh K. Gupta 0001, Mani Srivastava 0001 |
ICCAD | 1 |
| 1997 | Decomposition of timed decision tables and its use in presynthesis optimizationsabstractPresynthesis optimizations transform a behavioral HDL description into an optimized HDL description that results in improved synthesis results. We introduce the decomposition of timed decision tables (TDT), a tabular model of system behavior. The TDT decomposition is based on the kernel extraction algorithm. By experimenting using named benchmarks, we demonstrate how TDT decomposition can be used in presynthesis optimizations. Jian Li 0061, Rajesh K. Gupta 0001 |
ICCAD | 2 |
| 1997 | Architectural Adaptation for Application-Specific Locality OptimizationabstractWe propose a machine architecture that integrates programmable logic into key components of the system with the goal of customizing architectural mechanisms and policies to match an application. This approach presents an improvement over the traditional approach of exploiting programmable logic as a separate co-processor by pre-serving machine usability through software and on a traditional computer architecture by providing application-specific hardware. We present two case studies of architectural customization to enhance latency tolerance and efficiently utilize network bisection on multiprocessors for sparse matrix computations. We demonstrate that application-specific hardware and policies can provide substantial improvements in performance on a per application basis. Based on these preliminary results, we propose that an application-driven machine customization provides a promising approach to achieve high performance and combat performance fragility. Xingbin Zhang, Ali Dasdan, Martin Schulz 0001, Rajesh K. Gupta 0001, Andrew A. Chien |
ICCD | 4 |
| 1997 | Constrained software synthesis for embedded applications
Rajesh K. Gupta 0001, Giovanni De Micheli |
J. Syst. Archit. | 1 |
| 1997 | Implications of VHDL timing models on simulation and software synthesis
Venkatram Krishnaswamy, Rajesh K. Gupta 0001, Prithviraj Banerjee |
J. Syst. Archit. | 2 |
| 1997 | Hardware/software co-designabstractMost electronic systems, whether self contained or embedded, have a predominant digital component consisting of a hardware platform which executes software application programs. Hardware/software co-design means meeting system level objectives by exploiting the synergism of hardware and software through their concurrent design. Co-design problems have different flavors according to the application domain, implementation technology and design methodology. Digital hardware design has increasingly more similarities to software design. Hardware circuits are often described using modeling or programming languages, and they are validated and implemented by executing software programs, which are sometimes conceived for the specific hardware design. Current integrated circuits can incorporate one (or more) processor core(s) and memory array(s) on a single substrate. These "systems on silicon" exhibit a sizable amount of embedded software, which provides flexibility for product evolution and differentiation purposes. Thus the design of these systems requires designers to be knowledgeable in both hardware and software domains to make good design tradeoffs. The paper introduces various aspects of co-design. We highlight the commonalities and point out the differences in various co-design problems in some application areas. Co-design issues and their relationship to classical system implementation tasks are discussed to help develop a perspective on modern digital system design that relies on computer aided design (CAD) tools and methods. Giovanni De Michell, Rajesh K. Gupta 0001 |
Proc. IEEE | 2 |
| 1997 | Specification and analysis of timing constraints for embedded systemsabstractEmbedded systems consist of interacting hardware and software components that must deliver a specific functionality under constraints on relative timing of their actions. We describe operation delay and execution rate constraints, that are useful in the context of embedded systems. A delay constraint bounds the operation delay or specifies any of the thirteen possible constraints between the intervals of execution of a pair of operations. A rate constraint bounds the rate of execution of an operation and may be specified relative to the control flow in the system functionality. We present constraint propagation and analysis techniques to determine satisfaction of imposed constraints by a given system implementation. In contrast to previous purely analytical approaches on restricted models or statistical performance estimation based on runtime data, we present a static analysis in presence of conditionals and loops with the help of designer assists. The constraint analysis algorithms presented here have been implemented in a cosynthesis system, VULCAN, that allows the embedded system designer to interactively evaluate the effect of performance constraints on hardware-software implementation tradeoffs for a given functionality. We present examples to demonstrate the application and utility of the proposed techniques. Rajesh K. Gupta 0001, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1996 | Analysis of Operation Delay and Execution Rate Constraints for Embedded SystemsabstractConstraints on the delay and execution rate of operations in an embedded application are needed to ensure its timely interaction with a reactive environment. In this paper, we present a static analysis of the timing constraints satisfiability by a given system design consisting of interacting hardware and software components. We use this analysis to evaluate the effect of individual timing constraints on system design issues, such as the choice of the software runtime system, bounds on loop invocations, and the hardware-software synchronization operations. We show, by example, the use of static analysis techniques in the design of embedded systems. Rajesh K. Gupta 0001 |
DAC | 1 |
| 1996 | HDL Optimization Using Timed Decision TablesabstractSystem-level presynthesis refers to the optimization of an input HDL description that produces an optimized HDL description suitable for subsequent synthesis tasks. In this paper, we present optimization of control flow in behavioral HDL descriptions using external Don't Care conditions. The optimizations are carried out using a tabular model of system functionality, called Timed Decision Tables or TDTs. TDT based optimization presented here have been implemented in a program called PUMPKIN. Optimization results from several examples show a reduction of 3-88% in the size of synthesized hardware circuits depending upon the external Don't Care information supplied by the user. 1 Introduction Due to the maturity of optimization and synthesis tools at logic and register transfer level, system specification is increasingly being done at behavioral level using a Hardware Description Language (HDL), such as VHDL and Verilog. Though optimization can be done at all levels of system specificat... Jian Li 0061, Rajesh K. Gupta 0001 |
DAC | 2 |
| 1996 | An algorithm for synthesis of system-level interface circuitsabstractWe describe an algorithm for the synthesis and optimization of interface circuits for embedded system components such as microprocessors, memory ASIC, and network subsystems with fixed interfaces. The algorithm accepts the timing characteristics of two system components as input, and generates a combinational interface (glue logic) circuit. The algorithm consists of two parts. In the first part, we determine the direct pin-to-pin connections in the interface circuit employing a 0/1 ILP formulation to minimize wiring area and dynamic power consumption. In the second part, we determine logic subcircuits in the interface circuit, utilizing the timing diagrams of the system components. The proposed algorithm has been implemented in a software package SYNTERFACE. Experimental results are presented to demonstrate the effectiveness of the algorithm. Ki-Seok Chung, Rajesh K. Gupta 0001, C. L. Liu 0001 |
ICCAD | 2 |
| 1996 | Opportunities and pitfalls in HDL-based system designabstractThis panel discusses the complexities of system designs using textual Hardware Description Languages (HDLs) such as Verilog and VHDL. As the proliferation of circuits and systems design using HDLs continues questions arise as to whether HDL-based programming provides any real productivity gains in the design of complex integrated hardware. Are there are any alternatives, such as graphical or visual formalisms that are perhaps better suited for the task? We start the discussion by examining the relevant features of today's complex hardware systems and the requirements these impose on the modeling language and methodologies. Rajesh K. Gupta 0001, Daniel Gajski, Randy Allen, Yatin Trivedi |
ICCD | 1 |
| 1992 | Synthesis and Simulation of Digital Systems Containing Interacting Hardware and Software Components
Rajesh K. Gupta 0001, Claudionor José Nunes Coelho Jr., Giovanni De Micheli |
DAC | 1 |
| 1990 | Partitioning of Functional Models of Synchronous Digital SystemsabstractA partitioning technique is presented of functional models that are used in conjunction with high-level synthesis of digital synchronous circuits. The partitioning goal is to synthesize multi-chip systems from one behavioral description that satisfy both chip area constraints and an overall latency timing constraint. There are three major advantages to using partitioning techniques at the functional abstraction level. First, scheduling techniques can be applied concurrently to partitioning. Therefore, partitioning under timing constraints, and in particular under latency constraints, can be performed. Second, the functional model captures large hardware systems with fewer objects (than at the logic netlist abstraction level), making the partitioning algorithm more efficient. Third, hardware sharing tradeoffs can be considered. Hardware partitioning is formulated as a hypergraph partitioning problem. Algorithms for hardware partitioning are presented and experimental results are reported.> Rajesh K. Gupta 0001, Giovanni De Micheli |
ICCAD | 1 |