EDBT 2026 Demo / reviewers in the wild / expert
Claudio Passerone
dblp:31/6677
· DBLP profile ↗
26ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-5652-221XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LA-MTL: Latency-Aware Automated Multi-Task LearningabstractMulti-Task Learning (MTL) aims to unify a variety of tasks into a single network for improved training and inference efficiency. This is particularly attractive for real-time applications that require simultaneous execution of multiple workloads in resource-constrained embedded environments. However, most MTL approaches focus on enhancing parameters efficiency and overall tasks metrics, often lacking explicit inference latency awareness in the optimization loop. The design space exploration should not compromise on the parameters efficiency or task accuracy objectives in order to meet latency requirements. To address this, we propose LA-MTL, an automated layer-level MTL policy search that incorporates a novel analytical latency factor (ALF). By accounting for local and global latencies during the MTL policy search, we derive solutions that balance task metrics, parameters efficiency and latency constraints. LA-MTL search on ResNet34 yields solutions with up to 50% lower latency on the Jetson AGX Orin while maintaining competitive metrics in semantic segmentation and depth estimation tasks with a +/-2 p.p., on the CityScapes dataset. Additionally, we achieve a superior parameters efficiency, surpassing the state-of-theart MTL parameters reduction by over 20 p.p. Experiments on benchmark datasets (CityScapes, NYUv2) demonstrate the effectiveness of our approach across various backbones including ResNet34, MobileNetV2, and MobileOne in its expanded form. Code is available at https://github.com/shamvbs/LA-MTL.1 Shambhavi Balamuthu Sampath, Sami Sawani, Moritz Thoma, Lukas Frickenstein, Pierpaolo Morì, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Ulf Schlichtmann, Claudio Passerone, Walter Stechele |
DAC | 10 |
| 2025 | HiFi-SAGE: High Fidelity GraphSAGE-Based Latency Estimators for DNN OptimizationabstractAs deep neural networks (DNNs) are increasingly deployed on resource-constrained edge devices, optimizing and compressing them for real-time performance becomes crucial. Traditional hardware-aware DNN search methods often rely on inaccurate proxy metrics, expensive latency lookup tables, or slow hardware-in-the-Iloop (HIL) evaluations. To address this, quasi-generalized latency estimators, typically meta-learning-based, were proposed to replace HIL evaluations and accelerate the search. These come with a one-time data collection and training cost and can adapt to new hardware with few measurements. However, they still have some drawbacks: (1) They increase complexity by trying to generalize across a range of diverse hardware types; (2) They depend on handcrafted hardware descriptors, which may fail to capture hardware characteristics; (3) They often perform poorly on new, unseen hardware that significantly differs from their initial training set. To overcome these challenges, this paper turns to the more straightforward platform-specific estimators that do not require hardware descriptors and can be easily trained on any hardware. We introduce HiFi-SAGE, a high fidelity GraphSAGE-based platform-specific latency estimator. When trained from scratch on only 100 latency measurements, our novel dual-head estimator design surpasses the state-of-the-art (SoTA) on the 10% error bound metric by up to 17.4 p.p. while achieving an impressive fidelity score of 99% on the diverse LatBench dataset. We demonstrate that applying HiFi-SAGE to a genetic algorithm-based DNN compression search, achieved a Pareto front comparable to real HIL feedback with a mean absolute percentage error (MAPE) of 2.54%, 2.48%, and 4.16%, for InceptionV3, DenseNet169, and ResNet50 respectively. Compared to existing platform-specific works, the lower number of latency measurements and higher fidelity scores positions HiFi-SAGE as an attractive alternative to replace expensive HIL setups. Code is available at: https://github.com/shamvbs/HiFi-SAGE * Shambhavi Balamuthu Sampath, Leon Hecht, Moritz Thoma, Lukas Frickenstein, Pierpaolo Morì, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
DATE | 11 |
| 2025 | HotShot: A Loss-Guided Data Augmentation and Curriculum Learning Technique for the Task of Semantic SegmentationabstractSemantic segmentation is an important computer vision task that requires costly pixel-level annotations to train deep neural networks (DNNs) for. Especially for applications like autonomous driving, precise pixel-level understanding of scenes is a decisive factor between success and failure of the application. It follows that every labeled sample of an existing dataset is highly valuable and should be optimally used during training to maximize its value. This is achieved using (1) augmentation of the same labeled sample to help the model learn it in different ways, and (2) curriculum learning to introduce training samples to the model in an strategic order to ease the learning process. In this work, we present HotShot, a loss-guided cropping technique that assesses the DNN's prediction capability during the training to derive probability scores of potential cropping regions. This effectively combines augmentation and curriculum learning in one technique, where a single sample is cropped (augmentation) in regions selected based on the DNN's loss throughout the training (curriculum learning). For UperNet using a ConvNeXt-tiny backbone and DeepLabV3+ architecture using a ResNet-50 backbone, applying HotShot provides a +0.41 p.p. and +0.43 p.p. mIoU improvement over randomly cropping regions on the CityScapes and BDD100K datasets respectively. More interestingly, the analysis shows HotShot primarily boosts the classes that are most challenging for the model. For example, the rider and motorcycle classes on the BDD100K dataset improve by 163% and 129% using DeepLabV3+ with a ResNet-50 backbone. HotShot achieves improved mIoU in almost all cases and normalizes imbalances in learning challenging classes in datasets. Lukas Frickenstein, Moritz Thoma, Pierpaolo Morì, Shambhavi Balamuthu Sampath, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Christian Unger, Claudio Passerone, Walter Stechele |
IV | 9 |
| 2024 | MATAR: Multi-Quantization-Aware Training for Accurate and Fast Hardware RetargetingabstractQuantization of deep neural networks (DNNs) reduces their memory footprint and simplifies their hardware arithmetic logic, enabling efficient inference on edge devices. Different hardware targets can support different forms of quantization, e.g. full 8-bit, or 8/4/2-bit mixed-precision combinations, or fully-flexible bit-serial solutions. This makes standard quantization-aware training (QAT) of a DNN for different targets challenging, as there needs to be careful consideration of the supported quantization-levels of each target at training time. In this paper, we propose a generalized QAT solution that results in a DNN which can be retargeted to different hardware, without any retraining or prior knowledge of the hardware's supported quantization policy. First, we present the novel training scheme which makes the model aware of multiple quantization strategies. Then we demonstrate the retargeting capabilities of the resulting DNN by using a genetic algorithm to search for layer-wise, mixed-precision solutions that maximize performance and/or accuracy on the hardware target, without the need of fine-tuning. By making the DNN agnostic of the final hardware target, our method allows DNNs to be distributed to many users on different hardware platforms, without the need for sharing the training loop or dataset of the DNN developers, nor detailing the hardware capabilities ahead of time by the end-users of the efficient quantized solution. Models trained with our approach can generalize on multiple quantization policies with minimal accuracy degradation compared to target-specific quantization counterparts. Pierpaolo Morì, Moritz Thoma, Lukas Frickenstein, Shambhavi Balamuthu Sampath, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
DATE | 10 |
| 2024 | Wino Vidi Vici: Conquering Numerical Instability of 8-bit Winograd Convolution for Accurate Inference Acceleration on EdgeabstractWinograd-based convolution can reduce the total number of operations needed for convolutional neural network (CNN) inference on edge devices. Most edge hardware accelerators use low-precision, 8-bit integer arithmetic units to improve energy efficiency and latency. This makes CNN quantization a critical step before deploying the model on such an edge device. To extract the benefits of fast Winograd-based convolution and efficient integer quantization, the two approaches must be combined. Research has shown that the transform required to execute convolutions in the Winograd domain results in numerical instability and severe accuracy degradation when combined with quantization, making the two techniques incompatible on edge hardware. This paper proposes a novel training scheme to achieve efficient Winograd-accelerated, quantized CNNs. 8-bit quantization is applied to all the intermediate results of the Winograd convolution without sacrificing task-related accuracy. This is achieved by introducing clipping factors in the intermediate quantization stages as well as using the complex numerical system to improve the transform. We achieve 2.8× and 2.1× reduction in MAC operations on ResNet-20-CIFAR-10 and ResNet-18-ImageNet, respectively, with no accuracy degradation. Pierpaolo Morì, Lukas Frickenstein, Shambhavi Balamuthu Sampath, Moritz Thoma, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Christian Unger, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
WACV | 11 |
| 2023 | WinoTrain: Winograd-Aware Training for Accurate Full 8-bit Convolution AccelerationabstractEfficient inference is critical in realizing a low-power, real-time implementation of convolutional neural networks (CNNs) on compute and memory-constrained embedded platforms. Using quantization techniques and fast convolutional algorithms like Winograd, CNN inference can achieve benefits in latency and in energy consumption. Performing Winograd convolution involves (1) transforming the weights and activations to the Winograd domain, (2) performing element-wise multiplication on the transformed tensors, and (3) transforming the results back to the conventional spatial domain. Combining Winograd with quantization of all its steps results in severe accuracy degradation due to numerical instability. In this paper we propose a simple quantization-aware training technique, which quantizes all three steps of the Winograd convolution, while using a minimal number of scaling factors. Additionally, we propose an FPGA accelerator employing tiling and unrolling methods to highlight the performance benefits of using the full 8-bit quantized Winograd algorithm. We achieve 2× reduction in inference time compared to standard convolution on ResNet-18 for the ImageNet dataset, while improving the Top-1 accuracy by 55.7 p.p. compared to a standard post-training quantized Winograd variant of the network. Pierpaolo Morì, Shambhavi Balamuthu Sampath, Lukas Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Alexander Frickenstein, Walter Stechele, Claudio Passerone |
DAC | 8 |
| 2022 | Accelerating and pruning CNNs for semantic segmentation on FPGAabstractSemantic segmentation is one of the popular tasks in computer vision, providing pixel-wise annotations for scene understanding. However, segmentation-based convolutional neural networks require tremendous computational power. In this work, a fully-pipelined hardware accelerator with support for dilated convolution is introduced, which cuts down the redundant zero multiplications. Furthermore, we propose a genetic algorithm based automated channel pruning technique to jointly optimize computational complexity and model accuracy. Finally, hardware heuristics and an accurate model of the custom accelerator design enable a hardware-aware pruning framework. We achieve 2.44X lower latency with minimal degradation in semantic prediction quality (−1.98 pp lower mean intersection over union) compared to the baseline DeepLabV3+ model, evaluated on an Arria-10 FPGA. The binary files of the FPGA design, baseline and pruned models can be found in github.com/pierpaolomori/SemanticSegmentationFPGA Pierpaolo Morì, Manoj Rohit Vemparala, Nael Fasfous, Saptarshi Mitra, Sreetama Sarkar, Alexander Frickenstein, Lukas Frickenstein, Domenik Helms, Naveen Shankar Nagaraja, Walter Stechele, Claudio Passerone |
DAC | 11 |
| 2014 | 3DV - An embedded, dense stereovision-based depth mapping systemabstractThis paper describes the architecture and hardware implementation of an embedded, low-cost and low-power dense stereo reconstruction system, running at 30 fps at VGA resolution. The processing pipeline includes an initial image rectification stage, a cost generation unit based on the non-parametric census transform, a state-of-the-art Semi-Global cost optimization stage, and a final minimization and noise suppression step. The hardware implementation is based on a Xilinx ZynqTMSystem-on-Chip, which besides the FPGA provides a physical dual-core ARM CPU, which is exploited for control and to deliver output over the integrated Gigabit Ethernet connection. Gabriele Camellini, Mirko Felisa, Paolo Medici, Paolo Zani, Francesco Gregoretti, Claudio Passerone, Roberto Passerone |
Intelligent Vehicles Symposium | 6 |
| 2007 | Architecture of a Small Low-Cost SatelliteabstractThis paper presents the architecture of a small university satellite that we have developed. The main design criteria were low cost and fault tolerance, which have been achieved by using Commercial Off The Shelf components and by replicating all critical functions, while monitoring the state of the system for failures. The focus of the paper is on overall organization, design partitioning and details of the actual hardware. We show that the development of a low-cost satellite is feasible with a very limited budget. Dante Del Corso, Claudio Passerone, Leonardo Maria Reyneri, Claudio Sansoè, Marco Borri, Stefano Speretta, Maurizio Tranchero |
DSD | 2 |
| 2006 | Real time operating system modeling in a system level design environmentabstractModeling complex embedded systems requires design tools that are able to represent a variety of heterogenous components at several level of abstraction. These components often include hardware and software blocks, which play very different roles with respect to cost, performance and flexibility of the design. Early evaluation of possible design choices and trade-offs is therefore of paramount importance for an efficient implementation. One kind of component that is often neglected in the first stages of the design cycle is the operating system. However, since it sits between the software and the hardware domains, its effects on the final implementation are relevant. This paper presents a technique to model a general POSIX compliant real time operating system within the Metropolis framework. Possible models are presented and discussed with respect to their efficiency, taking into consideration the scheduling policy, the multitasking environment, interrupts and communication primitives Claudio Passerone |
ISCAS | 1 |
| 2005 | A Time Slice Based Scheduler Model for System Level DesignabstractEfficient evaluation of design choices, in terms of selection of algorithms to be implemented as hardware or software, and finding an optimal HW/SW design mix is an important requirement in the design flow of embedded systems. Time-to-market, faster upgradability and flexibility are some of the driving points to put increasing amounts of functionality as software executed on general purpose processing elements. In this scenario, dividing a monolithic task into multiple interacting tasks, and scheduling them on limited processing elements has become very important for a system designer. The paper presents an approach to model time-slice based task schedulers in the designs where the performance estimate of hardware and software models is less than time-slice accurate. The approach aims to increase the simulation efficiency of designs modeled at system level. We used Metropolis (Balarin, F. et al., IEEE Computer, vol.36, no.4, p.45-52, 2003) as our codesign environment. Luciano Lavagno, Claudio Passerone, Vishal Shah, Yosinori Watanabe |
DATE | 2 |
| 2005 | Quasi-static scheduling of independent tasks for reactive systemsabstractA reactive system must process inputs from the environment at the speed and with the delay dictated by the environment. The synthesis of reactive software from a modular concurrent specification model generates a set of concurrent tasks coordinated by an operating system. This paper presents a synthesis approach for reactive software that is aimed at minimizing the overhead introduced by the operating system and the interaction among the concurrent tasks. A formal model based on Petri nets is used to synthesize the tasks and verify the correctness of their composition. A practical application of the approach is illustrated by means of a real-life industrial example, which shows the significant impact of the approach on the performance of the system. Jordi Cortadella, Alex Kondratyev, Luciano Lavagno, Claudio Passerone, Yosinori Watanabe |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2005 | Implementation of a UMTS turbo decoder on a dynamically reconfigurable platformabstractHigh hardware design and mask production costs dictate the need to reuse an architectural platform for as many applications as possible. Embedded multimedia portable devices are required to perform in real time a huge variety of different algorithms, ranging from audio and image processing, to channel coding, to video games and java virtual machines. Dynamically reconfigurable architectures are an effective means to cope with both requirements. However, their effective and efficient use today is hindered by a lack of methodology and tools to extensively explore the hardware/software (HW/SW) design space, without requiring software developers to have a deep knowledge of the underlying architecture. This paper describes one such methodology, which extends the software programming model to the design flow for a reconfigurable processor. Its effectiveness is shown with the case study of a turbo decoder for universal mobile telecommunications systems, in which a remarkable 11X speed-up and 4X reduction of energy requirements with respect to a pure software implementation has been obtained, by mapping the more computation-intensive kernels to the reconfigurable hardware. Alberto La Rosa, Luciano Lavagno, Claudio Passerone |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2004 | Implementation of a UMTS Turbo-Decoder on a Dynamically Reconfigurable PlatformabstractModern embedded systems must execute a variety of high performance real-time tasks, such as audio and image compression, channel coding and encoding, etc. Reconfigurable platforms can effectively be used in these cases, because they allow to re-use the architecture for as many applications as possible. The paper describes the implementation of a UMTS turbo-decoder on one such platform, the XiRisc reconfigurable processor. Our goal is to test the development framework and design flow that we already developed on a real industrial example. Our results shows that, with some manual effort from the designer, very good performance improvements can be achieved, using a flow close to embedded software development. Alberto La Rosa, Claudio Passerone, Francesco Gregoretti, Luciano Lavagno |
DATE | 2 |
| 2004 | An Automated Methodology for Low Electro-Magnetic Emissions Digital Circuits DesignabstractThis paper presents an automated methodology for the design of low electromagnetic emissions digital circuits based on an optimized clock skew scheduling. The assumption that a unique clock signal must reach all memory elements of a circuit is common to all standard design flows for synchronous circuits. Unfortunately, the almost simultaneous switching of all gates in the circuit is responsible for very sharp and narrow peaks of current absorption from the power supply line. The consequent steepness of rising and falling edges of these current pulses results in significant contributions on a very wide range of frequencies and therefore can be considered the main cause for both conducted and radiated emissions in integrated circuits. A fully automated design flow has been developed integrating new tools with standard tools and allowing the designer to desynchronize a synchronous circuit. Our tools are able to read the standard delay file of a circuit, to derive a set of relative timing constraints for the clock inputs of each memory element in the circuit and then to generate a partition of the circuit into a given number of clock domains. The methodology and the tools have been tested on a 16 bit microprocessor with 5 stages of pipeline showing a dramatic reduction of the current spectrum in the range of frequencies between 100 MHz and 1 GHz. Ivan Blunno, Guy Alain Narboni, Claudio Passerone |
DSD | 3 |
| 2003 | Hardware/Software Design Space Exploration for a Reconfigurable Processor
Alberto La Rosa, Luciano Lavagno, Claudio Passerone |
DATE | 3 |
| 2002 | False Path Elimination in Quasi-Static SchedulingabstractWe have developed a technique to compute a Quasi Static Schedule of a concurrent specification for the software partition of an embedded system. Previous work did not take into account correlations among run-time values of variables, and therefore tried to find a schedule for all possible outcomes of conditional expressions. This is advantageous on one hand, because by abstracting data values one can find schedules in many cases for an originally undecidable problem. On the other hand it may lead to exploring false paths, i.e., paths that can never happen at run-time due to constraints on how the variables are updated. This affects the applicability of the approach, because it leads to an explosion in the running time and the memory requirements of the compile-time scheduler itself. Even worse, it also leads to an increase in the final code size of the generated software. In this paper we propose a semi-automatic algorithm to solve the problem of false paths: the designer identifies and tags critical expressions, and synchronization channels are automatically added to the specification to drive the search of a schedule. G. Arrigoni, L. Duchini, Claudio Passerone, Luciano Lavagno, Yosinori Watanabe |
DATE | 3 |
| 2002 | Processes, Interfaces and Platforms. Embedded Software Modeling in Metropolis
Felice Balarin, Luciano Lavagno, Claudio Passerone, Yosinori Watanabe |
EMSOFT | 3 |
| 2001 | A software development tool chain for a reconfigurable processorabstractArticle Share on A software development tool chain for a reconfigurable processor Authors: Alberto La Rosa Dipartimento di Elettronica, Politecnico di Torino, Italy Dipartimento di Elettronica, Politecnico di Torino, ItalyView Profile , Luciano Lavagno Dipartimento di Elettronica, Politecnico di Torino, Italy Dipartimento di Elettronica, Politecnico di Torino, ItalyView Profile , Claudio Passerone Dipartimento di Elettronica, Politecnico di Torino, Italy Dipartimento di Elettronica, Politecnico di Torino, ItalyView Profile Authors Info & Claims CASES '01: Proceedings of the 2001 international conference on Compilers, architecture, and synthesis for embedded systemsNovember 2001Pages 93–98https://doi.org/10.1145/502217.502232Published:16 November 2001Publication History 16citation670DownloadsMetricsTotal Citations16Total Downloads670Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Alberto La Rosa, Luciano Lavagno, Claudio Passerone |
CASES | 3 |
| 2001 | Generation of minimal size code for scheduling graphsabstractThis paper proposes a procedure for minimizing the code size of sequential programs for reactive systems. It identifies repeated code segments (a generalization of basic blocks to directed rooted trees) and finds a minimal covering of the input control flow graphs with code segments. The segments are disjunct, i.e. no two segments have the same code in common. The program is minimal in the sense that the number of code segments is minimum under the property of disjunction for the given control flow specification. The procedure makes no assumption on the target processor architecture, and is meant to be used between task synthesis algorithms from a concurrent specification and a standard compiler for the target architecture. It is aimed at optimizing the size of very large, automatically generated flat code, and extends dramatically the scope of classical common sub-expression identification techniques. The potential effectiveness of the proposed approach is demonstrated through preliminary experiments. Claudio Passerone, Yosinori Watanabe, Luciano Lavagno |
DATE | 1 |
| 2000 | Task generation and compile-time scheduling for mixed data-control embedded softwareabstractThe problem of optimal software synthesis for concurrent processes to be implemented on a single processor is addressed. The approach calls for the representation of the concurrent processes with Petri nets that give a theoretical foundation for the scheduling algorithm that sequentializes the concurrent processes and for the code generation step. The approach maximizes the amount of static scheduling to reduce the need of context switch and operating system intervention. Experimental results show the potential of our method to reduce software design time and errors. Jordi Cortadella, Alex Kondratyev, Luciano Lavagno, Marc Massot, Sandra Moral, Claudio Passerone, Yosinori Watanabe, Alberto L. Sangiovanni-Vincentelli |
DAC | 6 |
| 1999 | Computing Timed Transition Relations for Sequential Cycle-Based SimulationabstractIn this paper we address the problem of computing silent paths in an Finite State Machine (FSM). These paths are characterized by no observable activity under constant inputs, and can be used for a variety of applications, from verification, to synthesis, to simulation. First, we describe a new approach to compute the Timed Transition Relation of an FSM. Then, we concentrate on applying the methodology to simulation of reactive behaviours. In this field, we automatically extract a BDD-based behavioral model from the RT or gate level description. The behavioral model as able to "jump" in time and to avoid the simulation of internal events. Finally, we discuss a set of promising experimental results in a simulation environment under the Ptolemy simulator. Gianpiero Cabodi, Paolo Camurati, Claudio Passerone, Stefano Quer |
DATE | 3 |
| 1998 | A Case Study in Embedded System Design: An Engine Control UnitabstractA number of techniques and software tools for embedded system design have been recently proposed. However, the current practice in the designer community is heavily based on manual techniques and on past experience rather than on a rigorous approach to design. To advance the state of the art it is important to address a number of relevant design problems and solve them to demonstrate the power of the new approaches. Tullio Cuatto, Claudio Passerone, Luciano Lavagno, Attila Jurecska, Antonino Damiano, Claudio Sansoè, Alberto L. Sangiovanni-Vincentelli |
DAC | 2 |
| 1998 | Modeling reactive systems in JavaabstractWe present an application of the Java TM programming language to specify and implement reactive real-time systems. We have developed and tested a collection of classes and methods to describe concurrent modules and their asynchronous communication by means of signals. The control structures are closely patterned after those of the synchronous language Esterel , succinctly describing concurrency, sequencing and preemption. We show the user-friendliness and efficiency of the proposed technique by using an example from the automotive domain. Claudio Passerone, Claudio Sansoè, Luciano Lavagno, Patrick C. McGeer, Roberto Passerone, Alberto L. Sangiovanni-Vincentelli |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 1997 | Trade-off evaluation in embedded system design via co-simulationabstractCurrent design methodologies for embedded systems often force the designer to evaluate early in the design process architectural choices that will heavily impact the cost and performance of the final product. Examples of these choices are hardware/software partitioning, choice of the micro-controller, and choice of a run-time scheduling method. This paper describes how to help the designer in this task, by providing a flexible co-simulation environment in which these alternatives can be interactively evaluated. Claudio Passerone, Luciano Lavagno, Claudio Sansoè, Massimiliano Chiodo, Alberto L. Sangiovanni-Vincentelli |
ASP-DAC | 1 |
| 1997 | Fast Hardware/Software Co-Simulation for Virtual Prototyping and Trade-Off AnalysisabstractHardware/Software co-simulation is generally performed with separate simulation models. This makes trade-off evaluation difficult, because the models must bere-compiled whenever some architectural choice is changed. We propose a technique to simulate hardware and software that is almost cycle accurate, and uses the same model for both types of components. Only the timing information used for synchronization needs to be changed to modify the processor choice, the implementation choice, or the scheduling policy. We show how this technique can be used to decide the implementation of a real-life example, a car dashboard controller. Claudio Passerone, Luciano Lavagno, Massimiliano Chiodo, Alberto L. Sangiovanni-Vincentelli |
DAC | 1 |