EDBT 2026 Demo / reviewers in the wild / expert
Tony Givargis
dblp:g/TonyGivargis · also T. D. Givargis
· DBLP profile ↗
56ranked-venue papers
11as first author
11since 2021 · last 2025
0000-0002-1608-9324ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 11 first-author · 4 since 2021Software engineering, systems software and programming languages · 9 · 1 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Always-Sparse Training by Growing Connections with Guided Stochastic ExplorationabstractThe excessive computational requirements of modern artificial neural networks (ANNs) are posing limitations on the machines that can run them. Sparsification of ANNs is often motivated by time, memory and energy savings only during model inference, yielding no benefits during training. A growing body of work is now focusing on providing the benefits of model sparsification also during training. While these methods greatly improve the training efficiency, the training algorithms yielding the most accurate models still materialize the dense weights, or compute dense gradients during training. We propose an efficient, always-sparse training algorithm with excellent scaling to larger and sparser models, supported by its linear time complexity with respect to the model width during training and inference. Moreover, our guided stochastic exploration algorithm improves over the accuracy of previous sparse training methods. We evaluate our method on the CIFAR-10/100 and ImageNet classification tasks using ResNet, VGG, and ViT models, and compare it against a range of sparsification methods3. Mike Heddes, Narayan Srinivasa, Tony Givargis, Alexandru Nicolau |
IJCNN | 3 |
| 2024 | Enhanced Detection of Transdermal Alcohol Levels Using Hyperdimensional Computing on Embedded DevicesabstractAlcohol consumption has a significant impact on individuals’ health, with even more pronounced consequences when consumption becomes excessive. One approach to promoting healthier drinking habits is implementing just-in-time interventions, where timely notifications indicating intoxication are sent during heavy drinking episodes. However, the complexity or invasiveness of an intervention mechanism may deter an individual from using it in practice. Previous research tackled this challenge using collected motion data and conventional Machine Learning (ML) algorithms to classify heavy drinking episodes, but with impractical accuracy and computational efficiency for mobile devices. Consequently, we have elected to use Hyperdimensional Computing (HDC) to design a just-in-time intervention approach that is practical for smartphones, smart wearables, and IoT deployment. HDC is a framework that has proven results in processing real-time sensor data efficiently. This approach offers several advantages, including low latency, minimal power consumption, and high parallelism. We explore various HDC encoding designs and combine them with various HDC learning models to create an optimal and feasible approach for mobile devices. Our findings indicate an accuracy rate of 89%, which represents a substantial 12% improvement over the current state-of-the-art. Manuel E. Segura, Pere Vergés, Justin Tian Jin Chen, Ramesh Arangott, Angela Kristine Garcia, Laura Garcia Reynoso, Alexandru Nicolau, Tony Givargis, Sergio Gago Masagué |
IJCNN | 8 |
| 2024 | Molecular Classification Using Hyperdimensional Graph ClassificationabstractOur work introduces an innovative approach to graph learning by leveraging Hyperdimensional Computing. Graphs serve as a widely embraced method for conveying information, and their utilization in learning has gained significant attention. This is notable in the field of chemoinformatics, where learning from graph representations plays a pivotal role. An important application within this domain involves the identification of cancerous cells across diverse molecular structures. We propose an HDC-based model that demonstrates comparable Area Under the Curve results when compared to state-of-the-art models like Graph Neural Networks (GNNs) or the Weisfieler-Lehman graph kernel (WL). Moreover, it outperforms previously proposed hyperdimensional computing graph learning methods. Furthermore, it achieves noteworthy speed enhancements, boasting a 40x acceleration in the training phase and a 15x improvement in inference time compared to GNN and WL models. This not only underscores the efficacy of the HDC-based method, but also highlights its potential for expedited and resource-efficient graph learning. Pere Vergés, Igor Nunes, Mike Heddes, Tony Givargis, Alexandru Nicolau |
IJCNN | 4 |
| 2024 | Convolution and Cross-Correlation of Count Sketches Enables Fast Cardinality Estimation of Multi-Join QueriesabstractWith the increasing rate of data generated by critical systems, estimating functions on streaming data has become essential. This demand has driven numerous advancements in algorithms designed to efficiently query and analyze one or more data streams while operating under memory constraints. The primary challenge arises from the rapid influx of new items, requiring algorithms that enable efficient incremental processing of streams in order to keep up. A prominent algorithm in this domain is the AMS sketch. Originally developed to estimate the second frequency moment of a data stream, it can also estimate the cardinality of the equi-join between two relations. Since then, two important advancements are the Count sketch, a method which significantly improves upon the sketch update time, and secondly, an extension of the AMS sketch to accommodate multi-join queries. However, combining the strengths of these methods to maintain sketches for multi-join queries while ensuring fast update times is a non-trivial task, and has remained an open problem for decades as highlighted in the existing literature. In this work, we successfully address this problem by introducing a novel sketching method which has fast updates, even for sketches capable of accurately estimating the cardinality of complex multi-join queries. We prove that our estimator is unbiased and has the same error guarantees as the AMS-based method. Our experimental results confirm the significant improvement in update time complexity, resulting in orders of magnitude faster estimates, with equal or better estimation accuracy. Mike Heddes, Igor Nunes, Tony Givargis, Alexandru Nicolau |
Proc. ACM Manag. Data | 3 |
| 2023 | An Extension to Basis-Hypervectors for Learning from Circular Data in Hyperdimensional ComputingabstractHyperdimensional Computing (HDC) is a computation framework based on random vector spaces, particularly useful for machine learning in resource-constrained environments. The encoding of information to the hyperspace is the most important stage in HDC. At its heart are basis-hypervectors, responsible for representing atomic information. We present a detailed study on basis-hypervectors, leading to broad contributions to HDC: 1) an improvement for level-hypervectors, used to encode real numbers; 2) a method to learn from circular data, an important type of information never before addressed in HDC. Results indicate that these contributions lead to considerably more accurate models for classification and regression. Igor Nunes, Mike Heddes, Tony Givargis, Alexandru Nicolau |
DAC | 3 |
| 2023 | DotHash: Estimating Set Similarity Metrics for Link Prediction and Document DeduplicationabstractMetrics for set similarity are a core aspect of several data mining tasks. To remove duplicate results in a Web search, for example, a common approach looks at the Jaccard index between all pairs of pages. In social network analysis, a much-celebrated metric is the Adamic-Adar index, widely used to compare node neighborhood sets in the important problem of predicting links. However, with the increasing amount of data to be processed, calculating the exact similarity between all pairs can be intractable. The challenge of working at this scale has motivated research into efficient estimators for set similarity metrics. The two most popular estimators, MinHash and SimHash, are indeed used in applications such as document deduplication and recommender systems where large volumes of data need to be processed. Given the importance of these tasks, the demand for advancing estimators is evident. We propose DotHash, an unbiased estimator for the intersection size of two sets. DotHash can be used to estimate the Jaccard index and, to the best of our knowledge, is the first method that can also estimate the Adamic-Adar index and a family of related metrics. We formally define this family of metrics, provide theoretical bounds on the probability of estimate errors, and analyze its empirical performance. Our experimental results indicate that DotHash is more accurate than the other estimators in link prediction and detecting duplicate documents with the same complexity and similar comparison time. Igor Nunes, Mike Heddes, Pere Vergés, Danny Abraham, Alexander V. Veidenbaum, Alexandru Nicolau, Tony Givargis |
KDD | 7 |
| 2023 | Accelerating Permute and N-Gram Operations for Hyperdimensional Learning in Embedded SystemsabstractHyperdimensional computing (HDC) is a novel computing framework that has gained significant attention for its ability to accelerate machine learning algorithms. Its fast learning and inference capabilities make it an ideal technique for various fields, including machine learning. HDC utilizes high-dimensional holographic vectors, which are vectors with independent and identically distributed dimensions, to represent information. This unique representation allows HDC to leverage highly parallelizable arithmetic operations such as bundling, binding and permute. These simple and highly optimizable operations make HDC an efficient framework for classification in embedded systems. HDC has demonstrated remarkable accuracy in learning patterns from sequenced data. In this paper, we propose a method to enhance the permute operation, which is crucial for maintaining the order of symbols or measures in real-time data. Our method enhances the efficiency of HDC's permute operations by a factor of 10×. Furthermore, by applying the same idea to n-gram encoding, we achieve a speedup of 14×, resulting in up to 26.8× speedup on a real application, compared to a state-of-the-art HDC prototyping library. To achieve this improvement, we utilized SIMD operations and shifted entire SIMD data blocks rather than individual elements. As a result, we demonstrate that real-time inference can be conducted rapidly in applications that are utilized in embedded systems with constrained computational and memory resources, such as those for recognizing emotions, gestures, and language. Pere Vergés, Igor Nunes, Mike Heddes, Tony Givargis, Alexandru Nicolau |
RTCSA | 4 |
| 2023 | Torchhd: An Open Source Python Library to Support Research on Hyperdimensional Computing and Vector Symbolic ArchitecturesabstractHyperdimensional computing (HD), also known as vector symbolic architectures (VSA), is a framework for computing with distributed representations by exploiting properties of random high-dimensional vector spaces. The commitment of the scientific community to aggregate and disseminate research in this particularly multidisciplinary area has been fundamental for its advancement. Joining these efforts, we present Torchhd, a high-performance open source Python library for HD/VSA. Torchhd seeks to make HD/VSA more accessible and serves as an efficient foundation for further research and application development. The easy-to-use library builds on top of PyTorch and features state-of-the-art HD/VSA functionality, clear documentation, and implementation examples from well-known publications. Comparing publicly available code with their corresponding Torchhd implementation shows that experiments can run up to 100x faster. Torchhd is available at: https://github.com/hyperdimensional-computing/torchhd. Mike Heddes, Igor Nunes, Pere Vergés, Denis Kleyko, Danny Abraham, Tony Givargis, Alexandru Nicolau, Alexander V. Veidenbaum |
J. Mach. Learn. Res. | 6 |
| 2022 | Hyperdimensional hashing: a robust and efficient dynamic hash tableabstractMost cloud services and distributed applications rely on hashing algorithms that allow dynamic scaling of a robust and efficient hash table. Examples include AWS, Google Cloud and BitTorrent. Consistent and rendezvous hashing are algorithms that minimize key remapping as the hash table resizes. While memory errors in large-scale cloud deployments are common, neither algorithm offers both efficiency and robustness. Hyperdimensional Computing is an emerging computational model that has inherent efficiency, robustness and is well suited for vector or hardware acceleration. We propose Hyperdimensional (HD) hashing and show that it has the efficiency to be deployed in large systems. Moreover, a realistic level of memory errors causes more than 20% mismatches for consistent hashing while HD hashing remains unaffected. Mike Heddes, Igor Nunes, Tony Givargis, Alexandru Nicolau, Alexander V. Veidenbaum |
DAC | 3 |
| 2022 | GraphHD: Efficient graph classification using hyperdimensional computingabstractHyperdimensional Computing (HDC) developed by Kanerva is a computational model for machine learning inspired by neuroscience. HDC exploits characteristics of biological neural systems such as high-dimensionality, randomness and a holographic representation of information to achieve a good balance between accuracy, efficiency and robustness. HDC models have already been proven to be useful in different learning applications, especially in resource-limited settings such as the increasingly popular Internet of Things (IoT). One class of learning tasks that is missing from the current body of work on HDC is graph classification. Graphs are among the most important forms of information representation, yet, to this day, HDC algorithms have not been applied to the graph learning problem in a general sense. Moreover, graph learning in IoT and sensor networks, with limited compute capabilities, introduce challenges to the overall design methodology. In this paper, we present GraphHD - a baseline approach for graph classification with HDC. We evaluate GraphHD on real-world graph classification problems. Our results show that when compared to the state-of-the-art Graph Neural Networks (GNNs) the proposed model achieves comparable accuracy, while training and inference times are on average$14.6\times$and$2.0 \times$faster, respectively. Igor Nunes, Mike Heddes, Tony Givargis, Alexandru Nicolau, Alexander V. Veidenbaum |
DATE | 3 |
| 2021 | Gravity: An Artificial Neural Network Compiler for Embedded ApplicationsabstractThis paper introduces the Gravity compiler. Gravity is an open source optimizing Artificial Neural Network (ANN) to ANSI C compiler with two unique design features that make it ideal for use in resource constrained embedded systems: (1) the generated ANSI C code is self-contained and void of any library or platform dependencies and (2) the generated ANSI C code is optimized for maximum performance and minimum memory usage. Moreover, Gravity is constructed as a modern compiler consisting of an intuitive input language, an expressive Intermediate Representation (IR), a mapping to a Fictitious Instruction Set Machine (FISM) and a retargetable backend, making it an ideal research tool for exploring high-performance embedded software strategies in AI and Deep-Learning applications. We validate the efficacy of Gravity by solving the MNIST handwriting digit recognition on an embedded device We measured a 300x reduction in memory, 2.5x speedup in inference and 33% speedup in training compared to TensorFlow. We also outperformed TVM, by over 2.4x in inference speed. Tony Givargis |
ASP-DAC | 1 |
| 2019 | Switching Predictive Control Using Reconfigurable State-Based ModelabstractAdvanced control methodologies have helped the development of modern vehicles that are capable of path planning and path following. For instance, Model Predictive Control (MPC) employs a predictive model to predict the behavior of the physical system for a specific time horizon in the future. An optimization problem is solved to compute optimal control actions while handling model uncertainties and nonlinearities. However, these prediction routines are computationally intensive and the computational overhead grows with the complexity of the model. Switching MPC addresses this issue by combining multiple predictive models, each with a different precision granularity. In this artcle, we proposed a novel switching predictive control method based on a model reduction scheme to achieve various model granularities for path following in autonomous vehicles. A state-based model with tunable parameters is proposed to operate as a reconfigurable predictive model of the vehicle. A runtime switching algorithm is presented that selects the best model using machine learning. We employed a metric that formulates the tradeoff between the error and computational savings due to model reduction. Our simulation results show that the use of the predictive model in the switching scheme as opposed to single granularity scheme, yields a 45% decrease in execution time in tradeoff for a small 12% loss in accuracy in prediction of future outputs and no loss of accuracy in tracking the reference trajectory. Maral Amir, Frank Vahid, Tony Givargis |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2018 | Priority Neuron: A Resource-Aware Neural Network for Cyber-Physical SystemsabstractAdvances in sensing, computation, storage and actuation technologies have entered cyber-physical systems (CPSs) into the smart era where complex control applications requiring high performance are supported. Neural networks (NNs) models are proposed as a predictive model to be used in model predictive control (MPC) applications. However, the ability to efficiently exploit resource hungry NNs in embedded resource-bound settings is a major challenge. In this paper, we propose priority neuron network (PNN), a resource-aware NNs model that can be reconfigured into smaller subnetworks at runtime. This approach enables a tradeoff between the model's computation time and accuracy based on available resources. The PNN model is memory efficient since it stores only one set of parameters to account for various subnetwork sizes. We propose a training algorithm that applies regularization techniques to constrain the activation value of neurons and assigns a priority to each one. We consider the neuron's ordinal number as our priority criteria in that the priority of the neuron is inversely proportional to its ordinal number in the layer. This imposes a relatively sorted order on the activation values. We conduct experiments to employ our PNN as the predictive model of a vehicle in MPC for path tracking. To corroborate the effectiveness of our proposed methodology, we compare it with two state-of-the-art methods for resource-aware NN design. Compared to state-of-the-art work, our approach can cut down the training time by 87% and reduce the memory storage by 75% while achieving similar accuracy. Moreover, we decrease the computation overhead for the model reduction process that searches for n neurons below a threshold, from O(n) to O(logn). Maral Amir, Tony Givargis |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Hybrid state machine model for fast model predictive control: Application to path trackingabstractCyber-Physical Systems (CPS) are composed of computing devices interacting with physical systems. Model-based design is a powerful methodology in CPS design in the implementation of control systems. For instance, Model Predictive Control (MPC) is typically implemented in CPS applications, e.g., in path tracking of autonomous vehicles. MPC deploys a model to estimate the behavior of the physical system at future time instants for a specific time horizon. Ordinary Differential Equations (ODE) are the most commonly used models to emulate the behavior of continuous-time (non-)linear dynamical systems. A complex physical model may comprise thousands of ODEs which pose scalability, performance and power consumption challenges. One approach to address these model complexity challenges are frameworks that automate the development of model-to-model transformation. In this paper, we introduce a model generation framework to transform ODE models of a physical system to Hybrid Harmonic Equivalent State (HES) Machine model equivalents. Moreover, tuning parameters are introduced to reconfigure the model and adjust its accuracy from coarse-grained time critical situations to fine-grained scenarios in which safety is paramount. Machine learning techniques are applied to adopt the model to run-time applications. We conduct experiments on a closed-loop MPC for path tracking using the vehicle dynamics model. We analyze the performance of the MPC when applying our Hybrid HES Machine model. The performance of our proposed model is compared with state-of-the-art ODE-based models, in terms of execution time and model accuracy. Our experimental results show a 32% reduction in MPC return time for 0.8% loss in model accuracy. Maral Amir, Tony Givargis |
ICCAD | 2 |
| 2016 | Towards a timing attack aware high-level synthesis of integrated circuitsabstractVariabilities in the execution time of integrated circuits are frequently exploited as a side channel attack to expose secret information of deployed systems. Standard countermeasures analyze and change the explicit timing behavior in lower level hardware description languages, but their application is time consuming and error-prone. In this paper we investigate the integration of timing attack resilience into the high-level synthesis (HLS). HLS translates programs expressed in higher level programming languages, such as C, seamlessly to synthesizable hardware. We use timing annotations of basic blocks in C to add scheduling constraints that in the synthesis process balance the execution time of security-related execution branches. We integrate our approach to the scheduling of the open source LegUp HLS tool and apply the proposed method for the asymmetric cryptography algorithms RSA and ECC. The results proof the resistance against timing attacks, with a negligible overhead in synthesis efforts, area, and run-time. Steffen Peter, Tony Givargis |
ICCD | 2 |
| 2015 | Including variability of physical models into the design automation of cyber-physical systemsabstractA good cyber-physical-systems (CPS) design methodology must conduct trade-off analysis of both the physical characteristics of the CPS as well as its cyber sub-system in a holistic manner. This paper presents a design space exploration (DSE) approach for CPSs that emphasizes the variabilities of the physical subsystem and control aspects of the system. We propose the application of parameterizable physical models and automatic recalculation of control algorithm parameters for the explored systems. The resulting parameterizable models can be applied in a systematic simulation-based DSE framework that facilitates the identification of superior system configurations. We applied the proposed design flow to a real non-linear inverted pendulum system with a range of physical and cyber settings. The results show the feasibility and effectiveness of our approach in the design of physical and control parts of CPSs. Our work supplements existing work on cyber system modeling and plays an integral part in the design automation of such systems. Hamid Mirzaei Buini, Steffen Peter, Tony Givargis |
DAC | 3 |
| 2015 | From the browser to the remote physical lab: Programming cyber-physical systemsabstractCyber Physical Systems (CPSs) integrate networked embedded computation systems with real-world physical installations. Programming of CPSs is not trivial, since CPSs combine traditional programming challenges and real-world timing, concurrency, and communication. This paper shows how a programming framework that allows students to implement and test CPS control programs in their Internet browsers, can improve both the students' learning experience and learning results. Students model and program a CPS application on a high abstraction level in a web page. This web page, provided by the instructor, invokes the student's code either together with the CPS as functional specification models in a virtual timing environment, or as component in a real-world system that interacts with a real remote physical implementation. Using the provided abstraction, students can incrementally design a CPS and experience challenges such as channel delays, model uncertainties, and real-time behavior, but without the need for complex low level programming or tools. For a CPS example system, we applied the framework in an embedded system design class. Our results show, the ability of a JavaScript-based programming and execution environment to design, program, and run CPSs on different levels of abstraction. Our results also indicate an increased approval from the students and a significantly improved understanding of modeling and programming in the class. Steffen Peter, Farshad Momtaz, Tony Givargis |
FIE | 3 |
| 2015 | Component-Based Synthesis of Embedded Systems Using Satisfiability Modulo TheoriesabstractConstraint programming solvers, such as Satisfiability Modulo Theory (SMT) solvers, are capable tools in finding preferable configurations for embedded systems from large design spaces. However, constructing SMT constraint programs is not trivial, in particular for complex systems that exhibit multiple viewpoints and models. In thisarticle we propose CoDeL: a component-based description language that allows system designers to express components as reusable building blocks of the system with their parameterizable properties, models, and interconnectivity. Systems are synthesized by allocating, connecting, and parameterizing the components to satisfy the requirements of an application. We present an algorithm that transforms component-based design spaces, expressible in CoDeL, to an SMT program, which, solved by state-of-the-art SMT solvers, determines the satisfiability of the synthesis problem, and delivers a correct-by-construction system configuration. Evaluation results for use cases in the domain of scheduling and mapping of distributed real-time processes confirm, first, the performance gain of SMT compared to traditional design space exploration approaches, second, the usability gains by expressing design problems in CoDeL, and third, the capability of the CoDeL/SMT approach to support the design of embedded systems. Steffen Peter, Tony Givargis |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2015 | Graph-Based Approaches to Placement of Processing Element Networks on FPGAs for Physical Model SimulationabstractPhysical models utilize mathematical equations to characterize physical systems like airway mechanics, neuron networks, or chemical reactions. Previous work has shown that field programmable gate arrays (FPGAs) execute physical models efficiently. To improve the implementation of physical models on FPGAs, this article leverages graph theoretic techniques to synthesize physical models onto FPGAs. The first phase maps physical model equations onto a structured virtual processing element (PE) graph using graph theoretic folding techniques. The second phase maps the structured virtual PE graph onto physical PE regions on an FPGA using graph embedding theory. A simulated annealing algorithm is introduced that can map any physical model onto an FPGA regardless of the model's underlying topology. We further extend the simulated annealing approach by leveraging existing graph drawing algorithms to generate the initial placement. Compared to previous work on physical model implementation on FPGAs, embedding increases clock frequency by 25% on average (for applicable topologies), whereas simulated annealing increases frequency by 13% on average. The embedding approach typically produces a circuit whose frequency is limited by the FPGA clock instead of routing. Additionally, complex models that could not previously be routed due to complexity were made routable when using placement constraints. Bailey Miller, Frank Vahid, Tony Givargis, Philip Brisk |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2014 | Resource Synchronization in Hierarchically Scheduled Real-Time Systems Using Preemptive Critical SectionsabstractIn this paper we outline a novel approach for accessing mutually exclusive resources in hierarchically scheduled real-time systems. Our method known as the Resource Access Control Protocol with Preemption (RACPwP) is an improved resource allocation protocol which utilizes preemptive critical sections to provide guaranteed determinism for hard real-time tasks and comparable response times for soft real-time tasks. Our experiments demonstrated that RACPwP outperforms other state-of-the-art resource access control protocols used in hierarchically scheduled systems. RACPwP was implemented as part of VxWorks and evaluated in an actual embedded application used in the aerospace industry. As a result, the response times for hard real-time tasks were improved over a traditional resource synchronization protocol. Tom Springer, Steffen Peter, Tony Givargis |
ISORC | 3 |
| 2013 | An efficient compression scheme for checkpointing of FPGA-based digital mockupsabstractThis paper outlines a transparent and nonintrusive checkpointing mechanism for use with FPGA-based digital mockups. A digital mockup is an executable model of a physical system and used for real-time test and validation of cyber-physical devices that interact with the physical system. These digital mockups are typically defined in terms of a large set of ordinary differential equations. We consider digital mockups impelemented on field-programmable gate arrays (FPGAs). A checkpoint is a snapshot of the internal state of the model at a specific point in time as captured by some controller that resides on the same FPGA. We require that the model continues uninterrupted execution during a checkpointing operation. Once a checkpoint is created, the corresponding state information is transferred from the FPGA to a host computer for visualization and other off-chip processing. We outline the architecture of a checkpointing controller that captures and transfers the state information at a desired clock cycle using an aggressive compression technique. Our compression technique achieves 90% reduction in data transferred from the FPGA to the host computer under periodic checkpointing scenarios. The checkpointing with compression yields 15-36% FPGA size overhead, versus 6-11% for checkpointing without compression. Ting-Shuo Chou, Tony Givargis, Chen Huang 0005, Bailey Miller, Frank Vahid |
ASP-DAC | 2 |
| 2013 | Exploration with upgradeable models using statistical methods for physical model emulationabstractPhysical models capture environmental phenomena such as biochemical reactions, a beating heart, or neuron synapses, using mathematical equations. Previous work has shown that physical models can execute orders of magnitude faster on FPGAs (Field-Programmable Gate Arrays) compared to desktop PCs. Different models of the same physical phenomenon may vary, with "upgraded" models being more accurate but using more FPGA area and having slower performance. We propose that design space exploration considering upgradable models can dramatically increase the useful design space. We present an analysis of the solution space for utilizing networks of processing-elements (PEs) on FPGAs to emulate physical models, implement a web-based frontend to a compiler and cycle-accurate simulator of PE networks to estimate solution metrics, and utilize design-of-experiments (DOE) statistical methods to identify Pareto points. By considering upgradeable models during the design space exploration of a human lung physical model, the solution space of possible speedup, area, and accuracy is increased by 6X, 7.3X, and 1.5X, respectively, compared to evaluating a single model. Bailey Miller, Frank Vahid, Tony Givargis |
DAC | 3 |
| 2013 | Embedding-based placement of processing element networks on FPGAs for physical model simulationabstractPhysical models utilize mathematical equations to model physical systems like airway mechanics, neuron networks, or chemical reactions. Previous work has shown that physical models can execute fast on FPGAs (field-programmable gate arrays). We introduce an approach for implementing physical models on FPGAs that applies graph theoretic techniques to make use of a physical model's natural structure--tree, ring, chain, etc.--resulting in model execution speedups. A first phase of the approach maps physical model equations to a structured virtual PE (processing element) graph using graph theoretic folding techniques. A second phase maps the structured virtual PE graph to physical PE regions on an FPGA using graph embedding theory. We also present a simulated annealing approach with custom cost and neighbor functions that can map any physical model onto an FPGA with low wire costs. Average circuit speedup improvements over previous works for various physical models are 65% using the graph embedding and 35% using the simulated annealing approach. Each approach's more efficient use of FPGA resources also enables larger models to be implemented on an FPGA device. Bailey Miller, Frank Vahid, Tony Givargis |
FPGA | 3 |
| 2013 | Automatic synthesis of physical system differential equation models to a custom network of general processing elements on FPGAsabstractFast execution of physical system models has various uses, such as simulating physical phenomena or real-time testing of medical equipment. Physical system models commonly consist of thousands of differential equations. Solving such equations using software on microprocessor devices may be slow. Several past efforts implement such models as parallel circuits on special computing devices called Field-Programmable Gate Arrays (FPGAs), demonstrating large speedups due to the excellent match between the massive fine-grained local communication parallelism common in physical models and the fine-grained parallel compute elements and local connectivity of FPGAs. However, past implementation efforts were mostly manual or ad hoc. We present the first method for automatically converting a set of ordinary differential equations into circuits on FPGAs. The method uses a general Processing Element (PE) that we developed, designed to quickly solve a set of ordinary differential equations while using few FPGA resources. The method instantiates a network of general PEs, partitions equations among the PEs to minimize communication, generates each PE's custom program, creates custom connections among PEs, and maintains synchronization of all PEs in the network. Our experiments show that the method generates a 400-PE network on a commercial FPGA that executes four different models on average 15x faster than a 3 GHz Intel processor, 30x faster than a commercial 4-core ARM, 14x faster than a commercial 6-core Texas Instruments digital signal processor, and 4.4x faster than an NVIDIA 336-core graphics processing unit. We also show that the FPGA-based approach is reasonably cost effective compared to using the other platforms. The method yields 2.1x faster circuits than a commercial high-level synthesis tool that uses the traditional method for converting behavior to circuits, while using 2x fewer lookup tables, 2x fewer hardcore multiplier (DSP) units, though 3.5x more block RAM due to being programmable. Furthermore, the method does not just generate a single fastest design, but generates a range of designs that trade off size and performance, by using different numbers of PEs. Chen Huang 0005, Frank Vahid, Tony Givargis |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2013 | Synthesis of networks of custom processing elements for real-time physical system emulationabstractEmulating a physical system in real-time or faster has numerous applications in cyber-physical system design and deployment. For example, testing of a cyber-device's software (e.g., a medical ventilator) can be done via interaction with a real-time digital emulation of the target physical system (e.g., a human's respiratory system). Physical system emulation typically involves iteratively solving thousands of ordinary differential equations (ODEs) that model the physical system. We describe an approach that creates custom processing elements (PEs) specialized to the ODEs of a particular model while maintaining some programmability, targeting implementation on field-programmable gate arrays (FPGAs). We detail the PE micro-architecture and accompanying automated compilation and synthesis techniques. Furthermore, we describe our efforts to use a high-level synthesis approach that incorporates regularity extraction techniques as an alternative FPGA-based solution, and also describe an approach using graphics processing units (GPUs). We perform experiments with five models: a Weibel lung model, a Lutchen lung model, an atrial heart model, a neuron model, and a wave model; each model consists of several thousand ODEs and targets a Xilinx Virtex 6 FPGA. Results of the experiments show that the custom PE approach achieves 4X-9X speedups (average 6.7X) versus our previous general ODE-solver PE approach, and 7X-10X speedups (average 8.7X) versus high-level synthesis, while using approximately the same or fewer FPGA resources. Furthermore, the approach achieves speedups of 18X-32X (average 26X) versus an Nvidia GTX 460 GPU, and average speedups of more than 100X compared to a six-core TI DSP processor or a four-core ARM processor, and 24X versus an Intel I7 quad core processor running at 3.06 GHz. While an FPGA implementation costs about 3X-5X more than the non-FPGA approaches, a speedup/dollar analysis shows 10X improvement versus the next best approach, with the trend of decreasing FPGA costs improving speedup/dollar in the future. Chen Huang 0005, Bailey Miller, Frank Vahid, Tony Givargis |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2012 | MEDS: Mockup Electronic Data Sheets for automated testing of cyber-physical systems using digital mockupsabstractCyber-physical systems have become more difficult to test as hardware and software complexity grows. The increased integration between computing devices and physical phenomena demands new techniques for ensuring correct operation of devices across a broad range of operating conditions. Manual test methods, which involve test personnel, require much effort and expense and lengthen a device's time to market. We describe a method for test automation of devices wherein a device is connected to a digital mockup of the physical environment, where both the device and the digital mockup are managed by PC-based software. A digital mockup consists of a behavioral model of the interacting environment, such as a medical ventilator device connected to a digital mockup of human lungs. We introduce Mockup Electronic Data Sheets (MEDS) as a method for embedding model information into the digital mockup, allowing PC software to automatically detect configurable model parameters and facilitate test automation. We summarize a case study showing the effectiveness of digital mockups and MEDS as a framework for test automation on a medical ventilator, resulting in 5× less time spent testing compared to methods requiring test personnel. Bailey Miller, Frank Vahid, Tony Givargis |
DATE | 3 |
| 2009 | FSAF: File system aware flash translation layer for NAND Flash MemoriesabstractNAND Flash Memories require Garbage Collection (GC) and Wear Leveling (WL) operations to be carried out by Flash Translation Layers (FTLs) that oversee flash management. Owing to expensive erasures and data copying, these two operations essentially determine application response times. Since file systems do not share any file deletion information with FTL, dead data is treated as valid by FTL, resulting in significant WL and GC overheads. In this work, we propose a novel method to dynamically interpret and treat dead data at the FTL level so as to reduce above overheads and improve application response times, without necessitating any changes to existing file systems. We demonstrate that our resource-efficient approach can improve application response times and memory write access times by 22% and reduce erasures by 21.6% on average. Sai Krishna Mylavarapu, Siddharth Choudhuri, Aviral Shrivastava, Jongeun Lee, Tony Givargis |
DATE | 5 |
| 2009 | Source Routing Made Practical in Embedded NetworksabstractReducing packet latency is an important requirement in embedded networks. Source routing can be used to reduce processing delay at intermediate nodes and thereby reduce the overall packet latency. However source routing is not scalable which makes it unsuitable for larger networks. The addition of the source route to every packet reduces the system good put (application level throughput). Further, source routes ignore dynamic network conditions which might lead to routing failures. In this paper, we propose strategies to counter these problems. We propose a topology encoding scheme that reduces the overhead and makes source routing scalable. We propose a lazy correction scheme that makes it take cognizance of dynamic network conditions. Through simulations on reasonably large sized network with realistic models for traffic and failure, we show that source routing is indeed usable in practical scenarios. Tony Givargis |
ICCCN | 2 |
| 2008 | Control flow optimization in loops using interval analysisabstractWe present a novel loop transformation technique, particularly well suited for optimizing embedded compilers, where an increase in compilation time is acceptable in exchange for significant performance increase. The transformation technique optimizes loops containing nested conditional blocks. Specifically, the transformation takes advantage of the fact that the Boolean value of the conditional expression, determining the true/false paths, can be statically analyzed using a novel interval analysis technique that can evaluate conditional expressions in the general polynomial form. Results from interval analysis combined with loop dependency information is used to partition the iteration space of the nested loop. In such cases, the loop nest is decomposed such as to eliminate the conditional test, thus substantially reducing the execution time. Our technique completely eliminates the conditional from the loops (unlike previous techniques) thus further facilitating the application of other optimizations and improving the overall speedup. Applying the proposed transformation technique on loop kernels taken from Mediabench, SPEC-2000, mpeg4, qsdpcm and gimp, on average we measured a 175% (1.75X) improvement of execution time when running on a SPARC processor, a 336% (4.36X) improvement of execution time when running on an Intel Core Duo processor and a 198.9% (2.98X) improvement of execution time when running on a PowerPC G5 processor. Mohammad Ali Ghodrat, Tony Givargis, Alexandru Nicolau |
CASES | 2 |
| 2007 | System Architecture for Software PeripheralsabstractSoftware peripherals (Lioupis et al., 2001) have been proposed as a design alternative to traditional peripherals. We propose a software architecture, design methodology and scheduling scheme for implementing software peripherals on general purpose processors, with fast context switch and high resolution timers. Our design flow automatically generates code for scheduling software peripherals. We demonstrate the feasibility of our proposed work by experimenting with a set of five software peripherals scheduled to execute on a MIPS processor. Our performance evaluations show that the performance impact of the software peripherals on user-level tasks is minimal (i.e., 10.11% on a 100 MHz processor) - strongly suggesting that with the right architecture, software peripherals can be efficiently accommodated in typical embedded applications. Siddharth Choudhuri, Tony Givargis |
ASP-DAC | 2 |
| 2007 | Short-Circuit Compiler Transformation: Optimizing Conditional BlocksabstractWe present the short-circuit code transformation technique, intended for embedded compilers. The transformation technique optimizes conditional blocks in high-level programs. Specifically, the transformation takes advantage of the fact that the Boolean value of the conditional expression, determining the true/false paths, can be statically analyzed to determine cases when one or the other of the true/false paths are guaranteed to execute. In such cases, code is generated to bypass the evaluation of the conditional expression. In instances when the bypass code is faster to evaluate than the conditional expression, a net performance gain is obtained. Our experiments with the Mediabench applications show that the short-circuit transformation yields a an average of 35.1% improvement in execution time for SPARC and an average of 36.3% improvement in execution time for ARM. We also measured an average of 36.4% reduction in power consumption for ARM. Mohammad Ali Ghodrat, Tony Givargis, Alexandru Nicolau |
ASP-DAC | 2 |
| 2006 | Zero cost indexing for improved processor cache performanceabstractThe increasing use of microprocessor cores in embedded systems as well as mobile and portable devices creates an opportunity for customizing the cache subsystem for improved performance. In traditional cache design, the index portion of the memory address bus consists of the K least significant bits, where K = log 2 D and D is the depth of the cache. However, in devices where the application set is known and characterized (e.g., systems that execute a fixed application set) there is an opportunity to improve cache performance by choosing a near-optimal set of bits used as index into the cache. This technique does not add any overhead in terms of area or delay. In this article, we present an efficient heuristic algorithm for selecting K index bits for improved cache performance. We show the feasibility of our algorithm by applying it to a large number of embedded system applications as well as the integer SPEC CPU 2000 benchmarks. Specifically, for data traces, we show up to 45% reduction in cache misses. Likewise, for instruction traces, we show up to 31% reduction in cache misses. When a unified data/instruction cache architecture is considered, our results show an average improvement of 14.5% for the Powerstone benchmarks and an average improvement of 15.2% for the SPEC'00 benchmarks. Tony Givargis |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2006 | Synthesis of time-constrained multitasking embedded softwareabstractIn modern embedded systems, software development plays a vital role. Many key functions are being migrated to software, aiming at a shorter time to market and easier upgrades. Multitasking is increasingly common in embedded software, and many of these tasks incorporate real-time constraints. Although multitasking simplifies coding, it demands an operating system and imposes significant overhead on the system. The use of serializing compilers, such as the Phantom compiler, allows the synthesis of a monolithic code from a multitasking C application, eliminating the need for an operating system. In this article, we introduce the synthesis of multitasking applications that execute in a timely manner. We incorporate the notion of timing constraints into the Phantom compiler, and show that our approach is effective in meeting such constraints, allowing fine-grained concurrency among the tasks. As an additional case study, we present the implementation of a software-based modem and show that real-time applications such as the modem have guaranteed performance in the serialized code generated by the Phantom compiler. André C. Nácul, Tony Givargis |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2006 | Expression equivalence checking using interval analysisabstractArithmetic expressions are the fundamental building blocks of hardware and software systems. An important problem in computational theory is to decide if two arithmetic expressions are equivalent. However, the general problem of equivalence checking, in digital computers, belongs to the NP Hard class of problems. Moreover, existing general techniques for solving this decision problem are applicable to very simple expressions and impractical when applied to more complex expressions found in programs written in high-level languages. In this paper, we propose a method for solving the arithmetic expression equivalence problem using partial evaluation. In particular, our technique is specifically designed to solve the problem of equivalence checking of arithmetic expressions obtained from high-level language descriptions of hardware/software systems. In our method, we use interval analysis to substantially prune the domain space of arithmetic expressions and limit the evaluation effort to a sufficiently limited set of subspaces. Our results show that the proposed method is fast enough to be of use in practice Mohammad Ali Ghodrat, Tony Givargis, Alexandru Nicolau |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Equivalence checking of arithmetic expressions using fast evaluationabstractArithmetic expressions are the fundamental building blocks of hardware and software systems. An important problem in computational theory is to decide if two arithmetic expressions are equivalent. However, the general problem of equivalence checking, in digital computers, belongs to the NP Hard class of problems. Moreover, existing general techniques for solving this decision problem are applicable to very simple expressions and impractical when applied to more complex expressions found in programs written in high-level languages. In this paper we propose a method for solving the arithmetic expression equivalence problem using partial evaluation. In particular, our technique is specifically designed to solve the problem of equivalence checking of arithmetic expressions obtained from high-level language descriptions of hardware/software systems, which consists of regular arithmetic operators (+, -, x) and logical operators (and, or, not). In our method, we use interval analysis to substantially prune the domain space of arithmetic expressions and limit the evaluation effort to a sufficiently limited set of subspaces. Our results show that the proposed method is fast enough to be of use in practice. Mohammad Ali Ghodrat, Tony Givargis, Alexandru Nicolau |
CASES | 2 |
| 2005 | Beep: 3D indoor positioning using audible soundabstractRapid growth in the number of wireless enabled devices has led to an increased interest in location-aware applications. The backbone of such applications is provided by a location system. In this paper we present Beep, an indoor location system that senses audible sound. The use of audible sound makes our system cheap and easily deplorable to most existing roaming devices. Unlike positioning systems using ultrasound and infrared signals, Beep does not require the user to carry any kind of specialized hardware. Our system is based on standard 3D multilateration algorithms. However, the requirement of being able to locate existing devices, whose sound cards were not designed for high-precision signaling, introduces additional challenges to the location problem. This paper describes how those problems were solved and presents experimental results. Beep works with an accuracy of about 2 feet in more than 97% cases. The paper also describes a sensor deployment strategy that requires low sensor density and consequently low installation costs. Atri Mandal, Cristina V. Lopes, Tony Givargis, Amir Haghighat, Raja Jurdak, Pierre Baldi |
CCNC | 3 |
| 2005 | LORD: A Localized, Reactive and Distributed Protocol for Node Scheduling in Wireless Sensor NetworksabstractThe lifetime of wireless sensor networks can be increased by minimizing the number of active nodes that provide complete coverage, while switching off the rest. In this paper we propose a distributed and scalable node-scheduling algorithm that conserves overall system energy by minimizing the number of active nodes, localizing the execution to the dying sensor(s), and minimizing the frequency of execution by reacting only to the occurrence of a sensing hole. This effects an increased system lifetime while maintaining coverage over an application-defined threshold value. We compare our algorithm to a network with a centralized node-scheduling algorithm. Our results show equivalent coverage degree over a wide range of sensor networks. Tony Givargis |
DATE | 2 |
| 2005 | Lightweight Multitasking Support for Embedded Systems using the Phantom Serializing CompilerabstractEmbedded software continues to play an ever increasing role in the design of complex embedded applications. In part, the elevated level of abstraction provided by a high-level programming paradigm immensely facilitates a short design cycle, fewer errors. portability, and reuse. Serializing compilers have been proposed as an alternative to traditional OS techniques, enabling a designer to develop multitasking applications without the need of OS support. In this work, we outline the inner workings of the Phantom serializing compiler and analyze the quality of the generated code with respect to memory and processing overheads. Our results show that such serializing compilers are extremely efficient, making them ideal to be used in design of highly parallel applications (e.g.. multimedia, graphics, and signal processing applications). André C. Nácul, Tony Givargis |
DATE | 2 |
| 2004 | Dynamic Voltage and Cache Reconfiguration for Low PowerabstractThis article deals about dynamic voltage and cache reconfiguration online algorithm that dynamically adapts the processor speed and the cache subsystem to the workload requirements for the purpose of saving energy. The workload is considered to be a set of tasks with real-time deadlines. Our online algorithm is invoked as part of the OS scheduler, which performs standard earliest deadline first(EDF)task scheduling first. Then, our online algorithm, determines an ideal voltage/cache configuration for the current executing task. André C. Nácul, Tony Givargis |
DATE | 2 |
| 2004 | Code partitioning for synthesis of embedded applications with phantomabstractIn a large class of embedded systems, dynamic multitasking using traditional OS techniques is infeasible because of memory and processing overheads or lack of operating systems availability for the target embedded processor. Serializing compilers have been proposed as an alternative solution, enabling a designer to develop multitasking applications without the need of OS support. A serializing compiler is a source-to-source translator that takes a POSIX compliant multitasking C program as input and generates an equivalent, embedded processor independent, single-threaded ANSI C program, to be compiled using the embedded processor-specific tool chain. Such serializing compilers work by partitioning each task into blocks of code and synthesizing a scheduler that dynamically switches among these blocks. The quality of the compiled code in terms of multitasking overhead and task latency is highly dependent on the partitioning algorithm. In this work, we give our solution to the partitioning problem in the context of serializing compilers. We show that it is possible to provide the designer with a set of Pareto-optimal solutions that trade off multitasking overhead for task latency. André C. Nácul, Tony Givargis |
ICCAD | 2 |
| 2004 | Cache optimization for embedded processor cores: An analytical approachabstractEmbedded microprocessor cores are increasingly being used in embedded and mobile devices. The software running on these embedded microprocessor cores is often a priori known; thus, there is an opportunity for customizing the cache subsystem for improved performance. In this work, we propose an efficient algorithm to directly compute cache parameters satisfying desired performance criteria. Our approach avoids simulation and exhaustive exploration, and, instead, relies on an exact algorithmic approach. We demonstrate the feasibility of our algorithm by applying it to a large number of embedded system benchmarks. Tony Givargis |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2003 | Improved indexing for cache miss reduction in embedded systemsabstractThe increasing use of microprocessor cores in embedded systems as well as mobile and portable devices creates an opportunity for customizing the cache subsystem for improved performance. In traditional cache design, the index portion of the memory address bus consists of the K least significant bits, where K=log2(D) and D is the depth of the cache. However, in devices where the application set is known and characterized (e.g., systems that execute a fixed application set) there is an opportunity to improve cache performance by choosing an optimal set of bits used as index into the cache. This technique does not add any overhead in terms of area or delay. We give an efficient heuristic algorithm for selecting K index bits for improved cache performance. We show the feasibility of our algorithm by applying it to a large number of embedded system applications as well as the integer SPEC CPU 2000 benchmarks. Tony Givargis |
DAC | 1 |
| 2003 | Analytical Design Space Exploration of Caches for Embedded Systems
Tony Givargis |
DATE | 2 |
| 2003 | Cache Optimization For Embedded Processor Cores: An Analytical Approach
Tony Givargis |
ICCAD | 2 |
| 2003 | Exploring Efficient Operating Points for Voltage Scaled Embedded Processor CoresabstractPortable and battery operated devices pose a unique design challenge in terms of performance requirements, low-power constraints, and short design cycles. Embedded soft cores, on the other hand, provide functional flexibility and guarantee rapid design and thus are gaining popularity in designing such portable and battery operated devices. To address the low power needs, dynamic voltage scaled (DVS) processors provide a new tradeoff dimension to the designer. This work proposes an application-specific design space exploration framework for selecting energy-efficient operating points in an embedded soft core. Specifically, we address the problem of selecting an appropriate number of operating voltage/frequency points and the distribution of these points along the valid voltage span of a processor, given the application that is to be executed on the processor. Furthermore, we provide a static intra-task scheduling technique that reduces energy consumption (4-20% in our experiments) even when the worst-case application execution time does not leave any slack for effective voltage scaling. We have experimentally verified our technologies on a large set of embedded benchmarks selected from MiBench, PowerStone, and MediaBench. Marcio Buss, Tony Givargis, Nikil Dutt |
RTSS | 2 |
| 2002 | Platune: a tuning framework for system-on-a-chip platformsabstractSystem-on-a-chip (SOC) platform manufacturers are increasingly adding configurable features that provide power and performance flexibility in order to increase a platform's applicability. This paper presents a framework, called Platune, for performance and power tuning of one such SOC platform. Platune is used to simulate an embedded application that is mapped onto the SOC platform and output performance and power metrics for any configuration of the SOC platform. Furthermore, Platune is used to automatically explore the large configuration space of such an SOC platform. The versatility, in terms of accuracy and speed of exploration, of Platune is demonstrated experimentally using three large benchmark examples. The power estimation techniques for processors, caches, memories, buses, and peripherals combined with the design space exploration algorithm deployed by Platune form a methodology for design-of tuning frameworks for parameterized SOC platforms in general. Tony Givargis, Frank Vahid |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2002 | System-level exploration for Pareto-optimal configurations in parameterized system-on-a-chipabstractIn this work, we provide a technique for efficiently exploring the power/performance design space of a parameterized system-on-chip (SOC) architecture to find all Pareto-optimal configurations. These Pareto-optimal configurations will represent the range of power and performance tradeoffs that are obtainable by adjusting parameter values for a fixed application that is mapped on the SOC architecture. Our approach extensively prunes the potentially large configuration space by taking advantage of parameter dependencies. We have successfully applied our technique to explore Pareto-optimal configurations of our SOC architecture for a number of applications. Tony Givargis, Frank Vahid, Jörg Henkel |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | Instruction-based system-level power evaluation of system-on-a-chip peripheral coresabstractVarious core-based power evaluation approaches for microprocessors, caches, memories and buses have been proposed in the past. We propose a new power evaluation technique that is targeted toward peripheral cores. Our approach is the first to combine for peripherals both gate-level-obtained power data with a system-level simulation model written in an object-oriented language. Our approach decomposes peripheral functionality into so-called instructions. The approach can be applied with three increasingly fast methods: system simulation, trace simulation or trace analysis. We show that our models are sufficiently accurate in order to make power-related system-level design decisions but at a computation time that is orders of magnitude faster than a gate-level simulation. Tony Givargis, Frank Vahid, Jörg Henkel |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Trace-driven system-level power evaluation of system-on-a-chip peripheral coresabstractOur earlier work for fast evaluation of power consumption of general cores in a system-on-a-chip described techniques that involved isolating high-level instructions of a core, measuring gate-level power consumption per instruction, and then annotating a system-level simulation model with the obtained data. In this work, we describe a method for speeding up the evaluation further, through the use of instruction traces and trace simulators for every core, not just microprocessor cores. Our method shows noticeable speedups at an acceptable loss of accuracy. We show that reducing trace sizes can speed up the method even further. The speedups allow for more extensive system-level power exploration and hence better optimization. Tony Givargis, Frank Vahid, Jörg Henkel |
ASP-DAC | 1 |
| 2001 | System-Level Exploration for Pareto-Optimal Configurations in Parameterized Systems-on-a-ChipabstractProvides a technique for efficiently exploring the configuration space of a parameterized system-on-a-chip (SOC) architecture to find all Pareto-optimal configurations. These configurations represent the range of meaningful power and performance tradeoffs that are obtainable by adjusting parameter values for a fixed application mapped onto the SOC architecture. The approach extensively prunes the potentially large configuration space by taking advantage of parameter dependencies. The authors have successfully incorporated the technique into the parameterized SOC tuning environment (Platune) and applied it to a number of applications. Tony Givargis, Frank Vahid, Jörg Henkel |
ICCAD | 1 |
| 2001 | Evaluating power consumption of parameterized cache and bus architectures in system-on-a-chip designsabstractArchitectures with parameterizable cache and bus can support large tradeoffs between performance and power. We provide simulation data showing the large tradeoffs by such an architecture for several applications and demonstrating that the cache and bus should be configured simultaneously to find the optimal solutions. Furthermore, we describe analytical techniques for speeding up the cache/bus power and performance evaluation by several orders of magnitude over simulation, while maintaining sufficient accuracy with respect to simulation-based approaches. Tony Givargis, Frank Vahid, Jörg Henkel |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | A hybrid approach for core-based system-level power modelingabstractReducing power consumption has become a key goal for systemon-a-chip (SOC) designs. Fast and accurate power estimation is needed early in the design process, since power reduction methods tend to have greater impact at higher abstraction levels. Unfortunately, current approaches to power estimation, which concentrate on register-transfer-level models or lower, are quite slow. Higherlevel approaches, while faster, may suffer from inaccuracy. However, the advent of cores enables a hybrid approach, described in this paper, yielding both fast and accurate estimates from high-level models. In particular, we use power estimation data obtained from the gate-level for a core’s representative input stimuli data (instructions), and we propagate this data to a higher (object-oriented) system-level model, which is parameterizable and executable. Depending on the kind of cores, various parameterizable equation or look-up table based techniques are used, resulting in self-analyzing core models. We have applied our technique to several cores of a digital camera SOC and have achieved simulation speedups of over 1000 with accuracies suitable for making reliable power-related system-level design decisions. Although we focus on power estimation, our approach can be used for estimating other metrics as well, such as performance and size. 1 Tony Givargis, Frank Vahid, Jörg Henkel |
ASP-DAC | 1 |
| 2000 | A first-step towards an architecture tuning methodology for low powerabstractWe describe an automated environment to assist a system-on-achip designer to tune a microprocessor core to a particular application program that will run on the microprocessor, and vice-versa, with the goal of reducing embedded system power consumption.We limit such tuning to modifications that do not change the microprocessor instruction set, thus avoiding the large costs that would come with such a change.Our tuning environment for the 8051 microcontroller is freely-available on the web. Greg Stitt, Frank Vahid, Tony Givargis, Roman L. Lysecky |
CASES | 3 |
| 2000 | Fast Cache and Bus Power Estimation for Parameterized System-on-a-Chip DesignabstractWe present a technique for fast estimation of the power consumed by the cache and bus sub-system of a parameterized system-on-a-chip design for a given application. The technique uses a two-step approach of first collecting intermediate data about an application using simulation, and then using equations to rapidly predict the performance and power consumption for each of thousands of possible configurations of system parameters, such as cache size and associativity and bus size and encoding. The estimations display good absolute as well as relative accuracy for various examples, and are obtained in dramatically less time than other techniques, making possible the future use of powerful search heuristics. Jörg Henkel, Tony Givargis, Frank Vahid |
DATE | 2 |
| 2000 | Techniques for Reducing Read Latency of Core Bus WrappersabstractToday's system-on-a-chip designs consist of many cores, To enable cores to be easily integrated into different systems, many propose creating cores with their internal logic separated from their wrapper. This separation may introduce extra read latency. Pre-fetching register data into register copies in the bus wrapper can reduce or eliminate this extra latency. In this paper, we introduce a technique for automatically designing a pre-fetch unit that satisfies user-imposed register-access constraints. The technique benefits from mapping the pre-fetching problem to the well-known real-time process scheduling problem. We then extend the technique to allow user-specified register interdependencies, using a Petri net model, resulting in even more efficient pre-fetch schedules. Roman L. Lysecky, Frank Vahid, Tony Givargis |
DATE | 3 |
| 1999 | Interface and cache power exploration for core-based embedded system designabstractMinimizing power consumption is of paramount importance during the design of embedded (mobile computing) systems that come as systems-on-a-chip, since interdependencies between design characteristics like power, performance and area for various system parts (cores) are becoming increasingly influential. In this scenario, interfaces play a key role, since they allow one to control/exploit these interdependencies with the aim of meeting design constraints like power. In this paper, we present a comprehensive approach to explore this impact. We consider a whole system comprising a CPU, caches, a main memory and interfaces between those cores, and we demonstrate the high impact that an adequate adaptation between core parameters and interface parameters has in terms of power consumption. We find in particular that cache parameters and the configurations of cache buses have a significant impact in this respect. In addition, we make the important observation that optimizing for performance no longer implies that power is optimized as well in deep submicron technologies. Instead, we find that, especially for newer technologies, the relative interface power contribution increases, leading to scenarios where we obtain a real power/performance tradeoff. In summary, our explorations have revealed as yet uninvestigated interdependencies that represent the first step towards future efforts to optimize/adapt interfaces and caches in core-based systems for low-power designs. Tony Givargis, Jörg Henkel, Frank Vahid |
ICCAD | 1 |