EDBT 2026 Demo / reviewers in the wild / expert
Luciano Lavagno
dblp:36/1134
· DBLP profile ↗
119ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0002-9762-6522ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 96 · 11 first-author · 8 since 2021Software engineering, systems software and programming languages · 21 · 1 first-author · 2 since 2021Computer networks · 5Theory of computation · 5Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design for testability using mixed-polarity flip-flops and latchesabstractSequential circuits employing a combination of mixed-polarity flip-flops and latches allow significant improvements in clock frequency compared to useful skew and retiming. However, no work addresses the task of enabling a scan- based test on a circuit optimized with such techniques, while simultaneously minimizing the area overhead due to shadow latches used to complete the scan chain when latches are used in the design. This poses a serious limitation to the industrial application of mixed FF and latch-based techniques, since post-fabrication tests are an unavoidable step in IC production. This paper presents a macro-cell structure to enable both the exploitation of time borrowing for frequency optimization and the execution of the scan test of a design. The proposed solution requires minimal changes in the test setup and is evaluated using a recent methodology, Mix&Latch. Moreover, the work proposes modifications to Mix&Latch that allow reusing the standard cells introduced for the scan test to solve hold timing violations, avoiding additional hardware overhead. Results show that the lumped cell structure does not significantly impact frequency gains, and the ILP formulation of latch and FF type optimization can be extended to cover the DFT optimization part, ensuring only a moderate increase in area and power consumption, comparable with the DFT impact on regular FF-based designs. Lorenzo Lagostina, Jordi Cortadella, Mario R. Casu, Luciano Lavagno |
DATE | 4 |
| 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer ProgrammingabstractSkip connections have emerged as a key component of modern convolutional neural networks (CNNs) for computer vision tasks, allowing for the creation of more accurate and deeper models by addressing the vanishing gradient problem. However, the existing implementations of field-programmable gate array (FPGA)-based accelerators for ResNets and MobileNetV2 often experience decreased performance and increased computational latency due to the implementation of skip blocks. This article presents a novel framework for developing deep learning models on FPGAs that focuses on skip connections, with a unique approach to reduce buffering overhead. This results in a more efficient utilization of resources in the implementation of the skip layer. The nn2fpga compiler follows a thorough set of high-level synthesis (HLS) design principles and optimization strategies, exploiting in novel ways standard techniques to effectively map skip connection-based networks into static dataflow accelerators. To maximize throughput and efficiently use the available resources, our compiler employs a fast and effective design space exploration method based on a binary integer programming model which accurately assigns FPGA resources to the network layers, to maximize global throughput under resource constraints and then minimize resources for the achieved maximum throughput. Experimental results on the CIFAR-10 and ImageNet datasets demonstrate substantial gains in throughput ($\mathbf {3\times }$–$\mathbf {7\times }$on the past HLS-based work) for ResNet8, ResNet20, and MobileNetV2 models deployed on various Xilinx FPGA boards. Notably, MobileNetV2 deployed on the ZCU102 achieves a throughput of 2115 frame per second, representing even a 10% speedup over a state-of-the-art highly optimized manual register-transfer level implementation, showing that HLS can actually improve over manual design, thanks to the faster exploration of the design space. Roberto Bosio, Filippo Minnella, Teodoro Urso, Mario R. Casu, Luciano Lavagno, Mihai T. Lazarescu, Paolo Pasini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | SILVIA: Automated Superword-Level Parallelism Exploitation via HLS-specific LLVM Passes for Compute-Intensive FPGA AcceleratorsabstractHigh-level synthesis (HLS) aims at democratizing custom hardware acceleration with highly abstracted software-like descriptions. However, efficient accelerators still require substantial low-level hardware optimizations, defeating the HLS intent. In the context of field-programmable gate arrays, digital signal processors (DSPs) are a crucial resource that typically requires a significant optimization effort for its efficient utilization, especially when used for sub-word vectorization. This work proposes SILVIA, an open-source LLVM transformation pass that automatically identifies superword-level parallelism within an HLS design and exploits it by packing multiple operations, such as additions, multiplications, and multiply-and-adds, into a single DSP. SILVIA is integrated in the flow of the commercial AMD Vitis HLS tool and proves its effectiveness by packing multiple operations on the DSPs without any manual source-code modifications on several diverse state-of-the-art HLS designs such as convolutional neural networks and basic linear algebra subprograms accelerators, reducing the DSP utilization for additions by 70% and for multiplications and multiply-and-adds by 50% on average. Giovanni Brignone, Roberto Bosio, Fabrizio Ottati, Claudio Sansoè, Luciano Lavagno |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2024 | LESS: Low-Power Energy-Efficient Subgraph Isomorphism on FPGAabstractLow-power energy-efflcient subgraph isomorphism (LESS) is an open-source field-programmable gate array-only low-memory sub graph matching solver designed for energy efficiency. Depending on the input datagraph, the energy consumption of LESS, averaged on different diverse queries, is up to 38x and 93x lower than CPU and GPU solvers respectively. Roberto Bosio, Giovanni Brignone, Filippo Minnella, M. Usman Jamal, Luciano Lavagno |
DATE | 5 |
| 2024 | Mix & Latch: Comparison With State-of-the-Art Retiming on a RISC-V BenchmarkabstractFlip-flops (FFs) are the most commonly used sequential elements in synchronous circuits, but their timing requirements limit the operating frequency. Borrowing time with a latch-based approach can increase operating frequency, but traditional back-end optimization tools struggle to manage hold time requirements. The Mix & Latch technique achieves higher frequencies and often lower area than commercial state-of-the-art retiming by exploiting four types of synchronous sequential gates, namely, positive and negative edge-triggered flip-flops (FFs) and positive and negative transparent latches, all using a single clock tree.In this article, we first significantly accelerate the Mix & Latch flow convergence with respect to past work, by using a post-synthesis-based timing analysis that eliminates the first placement and routing needed for post-layout timing analysis. Then, by adding tolerance margins to the timing model, the pessimism is reduced to improve both convergence speed and maximum frequency. Finally, we reduce the complexity of the problem by applying the methodology only to the sequential elements belonging to critical paths. The effectiveness of Mix & Latch is then demonstrated on a RISC-V processor core from the Pulp platform using 28nm CMOS FDSOI technology. The results are compared to both the original Mix & Latch flow and a retiming performed with a state-of-the-art tool, showing a 25% frequency improvement over the original flow and 7.5% over the retiming flow. Compared to the retiming flow, we achieve comparable or lower power and area, while preserving the original registers and allowing logic equivalence checking. Lorenzo Lagostina, Filippo Minnella, Jordi Cortadella, Mario R. Casu, Mihai T. Lazarescu, Luciano Lavagno |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | A DSP shared is a DSP earned: HLS Task-Level Multi-Pumping for High-Performance Low-Resource DesignsabstractHigh-level synthesis (HLS) enhances digital hardware design productivity through a high abstraction level. Even if the HLS abstraction prevents fine-grained manual register-transfer level (RTL) optimizations, it also enables automatable optimizations that would be unfeasible or hard to automate at RTL. Specifically, we propose a task-level multi-pumping methodology to reduce resource utilization, particularly digital signal processors (DSPs), while preserving the throughput of HLS kernels modeled as dataflow graphs (DFGs) targeting field-programmable gate arrays. The methodology exploits the HLS resource sharing to automatically insert the logic for reusing the same functional unit for different operations. In addition, it relies on multi-clock DFGs to run the multi-pumped tasks at higher frequencies. The methodology scales the pipeline initiation interval (II) and the clock frequency constraints of resource-intensive tasks by a multi-pumping factor (M). The looser II allows sharing the same resource among M different operations, while the tighter clock frequency preserves the throughput. We verified that our methodology opens a new Pareto front in the throughput and resource space by applying it to open-source HLS designs using state-of-the-art commercial HLS and implementation tools by Xilinx. The multi-pumped designs require up to 40% fewer DSP resources at the same throughput as the original designs optimized for performance (i.e., running at the maximum clock frequency) and achieve up to 50% better throughput using the same DSPs as the original designs optimized for resources with a single clock. Giovanni Brignone, Mihai T. Lazarescu, Luciano Lavagno |
ICCD | 3 |
| 2022 | Fast Energy-Optimal Multikernel DNN-Like Application Allocation on Multi-FPGA PlatformsabstractPlatforms with multiple field-programmable gate arrays (FPGAs), such as Amazon Web Services (AWS) F1 instances, can efficiently accelerate multikernel pipelined applications, e.g., convolutional neural networks for machine vision tasks or transformer networks for natural language processing tasks. To reduce energy consumption when the FPGAs are underutilized, we propose a model to 1) find offline the minimum-power solution for given throughput constraints and 2) dynamically reprogram the FPGA at runtime (which is complementary to dynamic voltage and frequency scaling) to match best the workloads when they change. The offline optimization model can be solved using a mixed-integer nonlinear programming (MINLP) solver, but it can be very slow. Hence, we provide two heuristic optimization methods that improve result quality within a bounded time. We use several very large designs to demonstrate that both heuristics obtain comparable results to MINLP, when it can find the best solution, and they obtain much better results than MINLP, when it cannot find the optimum within a bounded amount of time. The heuristic methods can also be thousands of times faster than the MINLP solver. Junnan Shan, Mihai T. Lazarescu, Jordi Cortadella, Luciano Lavagno, Mario R. Casu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | CNN-on-AWS: Efficient Allocation of Multikernel Applications on Multi-FPGA PlatformsabstractMulti-FPGA platforms, like Amazon AWS F1, can run in the cloud multikernel pipelined applications, like convolutional neural networks (CNNs), with excellent performance and lower energy consumption than CPUs or GPUs. We propose a method to efficiently map these applications on multi-FPGA platforms to maximize the application throughput. Our methodology finds, for the given resources, the optimal number of parallel instances of each kernel in the pipeline and their allocation to one or more among the available FPGAs. We obtain this by formulating and solving a mixed-integer, nonlinear optimization problem, in which we model the performance of each component and the duration of the phases in which the accelerated computation can be split into, namely: 1) data transfer from a host CPU to the DDR memory of each FPGA; 2) data transfer from FPGA DDR to FPGA on-chip memory; 3) kernel computation on the FPGA; 4) data transfer from FPGA on-chip memory to FPGA DDR; and 5) data transfer from FPGA DDR to host. Finding the optimal solution using a mixed-integer nonlinear programming (MINLP) solver is often highly inefficient. Hence, we provide a fast heuristic method that according to our experiments can be much more efficient than the MINLP solver and finds comparable results. For larger problems (more CNN layers), our heuristic method can quickly find (several thousand times faster) much better solutions than the MINLP solver, even if we run the latter for a very long time. Junnan Shan, Mihai T. Lazarescu, Jordi Cortadella, Luciano Lavagno, Mario R. Casu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Exact and Heuristic Allocation of Multi-kernel Applications to Multi-FPGA PlatformsabstractFPGA-based accelerators demonstrated high energy efficiency compared to GPUs and CPUs. However, single FPGA designs may not achieve sufficient task parallelism. In this work, we optimize the mapping of high-performance multi-kernel applications, like Convolutional Neural Networks, to multi-FPGA platforms. First, we formulate the system level optimization problem, choosing within a huge design space the parallelism and number of compute units for each kernel in the pipeline. Then we solve it using a combination of Geometric Programming, producing the optimum performance solution given resource and DRAM bandwidth constraints, and a heuristic allocator of the compute units on the FPGA cluster. Junnan Shan, Mario R. Casu, Jordi Cortadella, Luciano Lavagno, Mihai T. Lazarescu |
DAC | 4 |
| 2019 | Synetgy: Algorithm-hardware Co-design for ConvNet Accelerators on Embedded FPGAsabstractUsing FPGAs to accelerate ConvNets has attracted significant attention in recent years. However, FPGA accelerator design has not leveraged the latest progress of ConvNets. As a result, the key application characteristics such as frames-per-second (FPS) are ignored in favor of simply counting GOPs, and results on accuracy, which is critical to application success, are often not even reported. In this work, we adopt an algorithm-hardware co-design approach to develop a ConvNet accelerator called Synetgy and a novel ConvNet model called DiracDeltaNet. Both the accelerator and ConvNet are tailored to FPGA requirements. DiracDeltaNet, as the name suggests, is a ConvNet with only $1\times 1$ convolutions while spatial convolutions are replaced by more efficient shift operations. DiracDeltaNet achieves competitive accuracy on ImageNet (89.0% top-5), but with 48× fewer parameters and 65× fewer OPs than VGG16. We further quantize DiracDeltaNet's weights to 1-bit and activations to 4-bits, with less than 1% accuracy loss. These quantizations exploit well the nature of FPGA hardware. In short, DiracDeltaNet's small model size, low computational OP count, ultra-low precision and simplified operators allow us to co-design a highly customized computing unit for an FPGA. We implement the computing units for DiracDeltaNet on an Ultra96 SoC system through high-level synthesis. Our accelerator's final top-5 accuracy of 88.2% on ImageNet, is higher than all the previously reported embedded FPGA accelerators. In addition, the accelerator reaches an inference speed of 96.5 FPS on the ImageNet classification task, surpassing prior works with similar accuracy by at least 16.9×. Qijing Huang 0001, Bichen Wu, Tianjun Zhang, Liang Ma 0003, Giulio Gambardella, Michaela Blott, Luciano Lavagno, Kees A. Vissers, John Wawrzynek, Kurt Keutzer |
FPGA | 8 |
| 2016 | ECOSCALE: Reconfigurable computing and runtime system for future exascale systems
Iakovos Mavroidis, Ioannis Papaefstathiou, Luciano Lavagno, Dimitrios S. Nikolopoulos, Dirk Koch, John Goodacre, Ioannis Sourdis, Vassilis Papaefstathiou, Marcello Coppola, Manuel Palomino |
DATE | 3 |
| 2016 | Designing Parameterizable Hardware IPs in a Model-Based Design Environment for High-Level SynthesisabstractModel-based hardware design allows one to map a single model to multiple hardware and/or software architectures, essentially eliminating one of the major limitations of manual coding in C or RTL. Model-based design for hardware implementation has traditionally offered a limited set of microarchitectures, which are typically suitable only for some application scenarios. In this article we illustrate how digital signal processing (DSP) algorithms can be modeled as flexible intellectual property blocks to be used within the popular Simulink model-based design environment. These blocks are written in C and are designed for both functional simulation and hardware implementation, including architectural design space exploration and hardware implementation through high-level synthesis. A key advantage of our modeling approach is that the very same bit-accurate model is used for simulation and high-level synthesis. To prove the feasibility of our proposed approach, we modeled a fast Fourier transform (FFT) algorithm and synthesized it for different DSP applications with very different performance and cost requirements. We also implemented a high-level-synthesis (HLS) intellectual property (IP) generator that can generate flexible FFT HLS-IP blocks that can be mapped to multiple micro-/macroarchitectures, to enable design space exploration as well as being used for functional simulation in the Simulink environment. Shahzad Ahmad Butt, Mehdi Roozmeh, Luciano Lavagno |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Reactive clocks with variability-tracking jitterabstractThe growing variability in nanoelectronic devices, due to uncertainties from the manufacturing process and environmental conditions (power supply, temperature, aging), requires increasing design guardbands, forcing circuits to work with conservative clock frequencies. Various schemes for clock generation based on ring oscillators and adaptive clocks have been proposed with the goal to mitigate the power and performance losses attributable to variability. However, there has been no systematic analysis to quantify the benefits of such schemes and no sign-off method has been proposed for timing correctness. This paper presents and analyzes a Reactive Clocking scheme with Variability-Tracking Jitter (RClk) that uses variability as an opportunity to reduce power by continuously adjusting the clock frequency to the varying environmental conditions, and thus, reduces guardband margins significantly. Power can be reduced between 20% and 40% at iso-performance and performance can be boosted by similar amounts at iso-power. Additionally, energy savings can be translated to substantial advantages in terms of reliability and thermal management. More importantly, the technology can be adopted with minimal modifications to conventional EDA flows. Jordi Cortadella, Luciano Lavagno, Pedro Lopez, Marc Lupon, Alberto Moreno-Conde, Antoni Roca 0001, Sachin S. Sapatnekar |
ICCD | 2 |
| 2015 | Analysis and Implementation of the Semi-Global Matching 3D Vision Algorithm Using Code Transformations and High-Level SynthesisabstractHigh-level synthesis (HLS) offers several advantages, such as faster simulation run-time and better design re-use, thanks to the higher level of abstraction. This work uses HLS to implement the Semi-Global Matching (SGM) algorithm, which is frequently used in stereo vision systems, e.g. for automotive applications. The hardware implementation is based on a Xilinx®Virtex 7 FPGA. The initial algorithmic “golden” model used very large arrays, which had to be mapped to an external DRAM and brought into the on-chip RAM of the FPGA on demand. This required both adding the memory transfer loops and inserting calls to the AXI transactors that access the DRAM through the on-chip DDR slave. Moreover, the initial single-threaded algorithm had to be parallelized, by converting the top-level sweeps of the image in eight directions into as many threads. The access to the DRAM was then managed with a centralized controller. This modified SystemC design proved to be suitable to achieve the target real-time performance. The design space was thus explored by making several fairly different micro-architectural choices. In the end, it was possible to obtain an implementation which is comparable to a very efficient (and hence very inflexible) manual RTL design that had been previously developed, including a very sophisticated fine-grained management of data and computation. Affaq Qamar, Fahad Bin Muslim, Luciano Lavagno |
VTC Spring | 3 |
| 2015 | Interactive Trace-Based Analysis Toolset for Manual Parallelization of C ProgramsabstractMassive amounts of legacy sequential code need to be parallelized to make better use of modern multiprocessor architectures. Nevertheless, writing parallel programs is still a difficult task. Automated parallelization methods can be effective both at the statement and loop levels and, recently, at the task level, but they are still restricted to specific source code constructs or application domains. We present in this article an innovative toolset that supports developers when performing manual code analysis and parallelization decisions. It automatically collects and represents the program profile and data dependencies in an interactive graphical format that facilitates the analysis and discovery of manual parallelization opportunities. The toolset can be used for arbitrary sequential C programs and parallelization patterns. Also, its program-scope data dependency tracing at runtime can complement the tools based on static code analysis and can also benefit from it at the same time. We also tested the effectiveness of the toolset in terms of time to reach parallelization decisions and of their quality. We measured a significant improvement for several real-world representative applications. Mihai T. Lazarescu, Luciano Lavagno |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Virtual Platform-Based Design Space Exploration of Power-Efficient Distributed Embedded ApplicationsabstractNetworked embedded systems are essential building blocks of a broad variety of distributed applications ranging from agriculture to industrial automation to healthcare and more. These often require specific energy optimizations to increase the battery lifetime or to operate using energy harvested from the environment. Since a dominant portion of power consumption is determined and managed by software, the software development process must have access to the sophisticated power management mechanisms provided by state-of-the-art hardware platforms to achieve the best tradeoff between system availability and reactivity. Furthermore, internode communications must be considered to properly assess the energy consumption. This article describes a design flow based on a SystemC virtual platform including both accurate power models of the hardware components and a fast abstract model of the wireless network. The platform allows both model-driven design of the application and the exploration of power and network management alternatives. These can be evaluated in different network scenarios, allowing one to exploit power optimization strategies without requiring expensive field trials. The effectiveness of the approach is demonstrated via experiments on a wireless body area network application. Parinaz Sayyah, Mihai T. Lazarescu, Sara Bocchio, Emad Samuel Malki Ebeid, Gianluca Palermo, Davide Quaglia, Alberto Rosti, Luciano Lavagno |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2014 | Energy-aware parallelization flow and toolset for C codeabstractMulticore architectures are increasingly used in embedded systems to achieve higher throughput with lower energy consumption. This trend accentuates the need to convert existing sequential code to effectively exploit the resources of these architectures. We present a parallelization flow and toolset for legacy C code that includes a performance estimation tool, a parallelization tool, and a streaming-oriented parallelization framework. These are part of the work-in-progress EU FP7 PHARAON project that aims to develop a complete set of techniques and tools to guide and assist software development for heterogeneous parallel architectures. We demonstrate the effectiveness of the use of the toolset in an experiment where we measure the parallelization quality and time for inexperienced users, and the parallelization flow and performance results for the parallelization of a practical example of a stereo vision application. Mihai T. Lazarescu, Albert Cohen 0001, Adrien Guatto, Nhat Minh Lê, Luciano Lavagno, Antoniu Pop, Andrei Sergeevich Terechko, Alexandru Sutii |
SCOPES | 5 |
| 2013 | Share with care: a quantitative evaluation of sharing approaches in high-level synthesisabstractThis paper focuses on the resource sharing problem when performing high-level synthesis. It argues that the conventionally accepted synthesis flow when resource sharing is done after scheduling is sub-optimal because it cannot account for timing penalties from resource merging. The paper describes a competitive approach when resource sharing and scheduling are performed simultaneously. It provides a quantitative evaluation of both approaches and shows that performing sharing during scheduling wins over the conventional approach in terms of quality of results. Alex Kondratyev, Luciano Lavagno, Mike Meyer, Yosinori Watanabe |
DATE | 2 |
| 2013 | EU FP7-288307 Pharaon Project: Parallel and Heterogeneous Architecture for Real-Time ApplicationsabstractIn this article, we present the work-in-progress of the EU FP7 PHARAON project, started in September 2011. The first objective of the project is the development of new techniques and tools capable to assist the designer in the development of parallel embedded systems, from executable specifications to target-specific implementation and debugging on a multicore platform. This tool chain will offer and implement several parallelization strategies, reflecting the functional and non-functional constraints of the system, and driving the designer into incremental parallelization and adaptation steps. The second objective of the project is to develop monitoring and control techniques in the middleware of the system capable to automatically adapt platform services to application requirements and therefore reduce power consumption transparently. Héctor Posadas, Eugenio Villar, Florian Broekaert, Michel Bourdellès, Albert Cohen 0001, Antoniu Pop, Nhat Minh Lê, Adrien Guatto, Mihai T. Lazarescu, Luciano Lavagno, Andrei Sergeevich Terechko, Miguel Glassée, Daniel Calvo, Eduardo de las Heras |
DSD | 10 |
| 2013 | A key management scheme for Content Centric Networking
Sarmad Ullah Khan, Thibault Cholez, Thomas Engel 0001, Luciano Lavagno |
IM | 4 |
| 2012 | Exploiting area/delay tradeoffs in high-level synthesisabstractThis paper proposes an enhanced scheduling approach for high-level synthesis, which relies on a multi-cycle behavioral timing analysis step that is performed before and during scheduling. The goal of this analysis is to accurately evaluate the criticality of operations and determine the most suitable candidate resources to implement them. The efficiency of the approach is confirmed by testing it on industrial examples, where it achieves, on average, 9% area savings after logic synthesis. Alex Kondratyev, Luciano Lavagno, Mike Meyer, Yosinori Watanabe |
DATE | 2 |
| 2012 | HEAP: A Highly Efficient Adaptive Multi-processor FrameworkabstractWriting parallel code is difficult, especially when starting from a sequential reference implementation. Our research efforts, as demonstrated in this paper, face this challenge directly by providing an innovative toolset that helps software developers profile and parallelize an existing sequential implementation, by exploiting top-level pipeline-style parallelism. The innovation of our approach is based on the facts that a) we use both automatic and profiling-driven estimates of the available parallelism, b) we refine those estimates using metric-driven verification techniques, and c) we support dynamic recovery of excessively optimistic parallelization. The proposed toolset has been utilized to find an efficient parallel code organization for a number of real-world representative applications, and a version of the toolset is provided in an open-source manner. Luciano Lavagno, Mihai T. Lazarescu, Ioannis Papaefstathiou, Andreas Brokalakis, Johan Walters, Bart Kienhuis, Florian Schäfer 0001 |
DSD | 1 |
| 2012 | FASTCUDA: Open Source FPGA Accelerator & Hardware-Software Codesign Toolset for CUDA KernelsabstractUsing FPGAs as hardware accelerators that communicate with a central CPU is becoming a common practice in the embedded design world but there is no standard methodology and toolset to facilitate this path yet. On the other hand, languages such as CUDA and OpenCL provide standard development environments for Graphical Processing Unit (GPU) programming. FASTCUDA is a platform that provides the necessary software toolset, hardware architecture, and design methodology to efficiently adapt the CUDA approach into a new FPGA design flow. With FASTCUDA, the CUDA kernels of a CUDA-based application are partitioned into two groups with minimal user intervention: those that are compiled and executed in parallel software, and those that are synthesized and implemented in hardware. A modern low power FPGA can provide the processing power (via numerous embedded micro-CPUs) and the logic capacity for both the software and hardware implementations of the CUDA kernels. This paper describes the system requirements and the architectural decisions behind the FASTCUDA approach. Iakovos Mavroidis, Ioannis Mavroidis, Ioannis Papaefstathiou, Luciano Lavagno, Mihai T. Lazarescu, Eduardo de la Torre, Florian Schäfer 0001 |
DSD | 4 |
| 2012 | Dynamic Trace-Based Data Dependency Analysis for Parallelization of C ProgramsabstractWriting parallel code is traditionally considered a difficult task, even when it is tackled from the beginning of a project. In this paper, we demonstrate an innovative toolset that faces this challenge directly. It provides the software developers with profile data and directs them to possible top-level, pipeline-style parallelization opportunities for an arbitrary sequential C program. This approach is complementary to the methods based on static code analysis and automatic code rewriting and does not impose restrictions on the structure of the sequential code or the parallelization style, even though it is mostly aimed at coarse-grained task-level parallelization. The proposed toolset has been utilized to define parallel code organizations for a number of real-world representative applications and is based on and is provided as free source. Mihai T. Lazarescu, Luciano Lavagno |
SCAM | 2 |
| 2011 | An extended framework for WSN applicationsabstractUnderstanding the real behavior of a Wireless Sensor Network application before the implementation enables engineers to reduce development time and cost because it makes possible the optimization of system resource utilization. In this context, this paper present an extension to an existing reference framework in which the designer can build network components, consider details about physical layer and communication medium, easily exchange the WSN configuration, simulate it by analyzing several parameters and finally generate code for different target hw/sw platforms. Luigi Pomante, Antonio Spinosi, Mohammad Mostafizur Rahman Mozumdar, Stefano Olivieri, Luciano Lavagno |
CCNC | 5 |
| 2011 | An Energy and Memory-Efficient Key Management Scheme for Mobile Heterogeneous Sensor NetworksabstractWireless Sensor Network (WSN) technology is being increasingly adopted in a wide variety of applications ranging from home/building and industrial automation to more safety critical applications including e-health or infrastructure monitoring. Considering mobility in the above application scenarios actually introduces additional technological challenges, especially with respect to security. The resource constrained devices should be robust to diverse security attacks and communicate securely while they are moving in the considered environment. To this aim, proper authentication and key management schemes supporting node mobility should be used. This paper presents an effective mutual authentication and key establishment scheme for heterogeneous sensor networks consisting of numerous mobile sensor nodes and only a few more powerful fixed sensor nodes. Moreover, OMNET++ simulations are used to provide a comprehensive performance evaluation of the proposed scheme. The obtained results show that the proposed solution assures better network connectivity, consumes less memory, has low communication overhead during the authentication and key establishment phase and has better network resilience against mobile nodes attacks compared with existing approaches for authentication and key establishment. Sarmad Ullah Khan, Claudio Pastrone, Luciano Lavagno, Maurizio A. Spirito |
CRiSIS | 3 |
| 2011 | Realistic performance-constrained pipelining in high-level synthesisabstractThis paper describes an approach to pipelining in high-level synthesis that modifies the control/data flow graph before and after scheduling. This enables the direct re-use of a pre-existing, timing- and area-aware non-pipelined simultaneous scheduler and binder. Such an approach ensures that the RTL output can be synthesized within the given timing and area constraints. Results from real industrial designs show the effectiveness of this approach in improving Pareto optimality with respect to area, delay and power. Alex Kondratyev, Luciano Lavagno, Mike Meyer, Yosinori Watanabe |
DATE | 2 |
| 2010 | Incremental high-level synthesisabstractThe widespread acceptance of High-level synthesis as a mainstream tool mostly depends on its tight integration with the following RTL-to-GDSII design flow. A key aspect is the handling of so-called Engineering Change Orders (ECOs), i.e. minor changes required to fix small functional bugs or meet performance requirements late in the design cycle. Traditional high-level synthesis has attempted to optimize at best the output logic. However, in the ECO scenario the goal is to implement the required change with as few modifications as possible to the RTL, logic netlist, placed netlist and layout. In this paper we show how, by judiciously changing the internal databases used by the tool to match as much as possible the original design, one can achieve minimal impact and implement ECOs in truly incremental mode, while full-blow re-synthesis would lead to massive unnecessary downstream changes. The tool essentially matches source constructs between the original and the ECO design, and copies as many synthesis decisions as possible from the original design to the ECO design. Luciano Lavagno, Alex Kondratyev, Yosinori Watanabe, Qiang Zhu 0005, Mototsugu Fujii, Mitsuru Tatesawa, Noriyasu Nakayama |
ASP-DAC | 1 |
| 2010 | Energy optimization at the MAC layer for a forest fire monitoring wireless sensor networkabstractThis paper describes several optimizations of MAC protocols that can be applied in order to satisfy the constraints that come from a real-life application. Forest fire monitoring requires very different latencies and data sizes, depending on whether it is reporting normal conditions, sending an alarm, or performing network management functions. We use MAC algorithms that extensively shut down the radio in order to save power. We show that by exploiting knowledge about the Link Quality Index and by effectively using the free time of the channel only when there is more data than usual to transmit, we manage to decrease latency and contention and increase bandwidth usage, while keeping power consumption very low. Anwar Al-Khateeb, Jun Kyoung Kim, Luciano Lavagno, Mihai T. Lazarescu |
ETFA | 3 |
| 2010 | Energy optimization framework for WSN designabstractIn this paper we introduce an energy optimization framework for wireless sensor network (WSN) design. It deals with functional design, optimization and implementation of energy aware WSNs.At the PHY layer, we provide a mechanism to optimize at design time and control at runtime the transmission power depending on the channel quality (distance, noise, fading). At the MAC layer, we use asynchronous scheduling to maximize sleep time, and we optimize the backoff time and the slot selection. At the NET layer, we optimize at design time, using closed form equations, and at runtime, using channel quality estimations, the number of nodes and hops to transmit packets across a given distance. We also discuss dynamic network routing approaches which reduce power with an acceptable error rate. Anwar Al-Khateeb, Luciano Lavagno |
IPSN | 2 |
| 2010 | MEOW: Model-based design of an energy-optimized protocol stack for wireless sensor networksabstractThe MEOW model-based design of an energy-optimized protocol stack addresses energy-aware functional design, optimization, simulation and code generation for WSNs. The goal is to satisfy functional and performance requirements, while maximizing battery lifetime. MEOW has been developed using the Matlab and Stateflow tools, which can be used in a model-based design environment to provide opportunities for co-design of the application and of the parameters of the communication stack, and then to generate code automatically for a variety of simulation and implementation platforms. MEOW allows the designer, who may not be a telecommunication protocol specialist, but is an expert of the final application domain (e.g. building automation, industrial automation, etc.), to optimize energy consumption at each network layer. It also includes cross-layer optimizations, which provide even more reduction of energy consumption and improve sharing of resources. For the Physical layer, we provide a method to optimize at design time and control at runtime the transmission power depending on the channel quality. For the MAC layer, we developed the PE-MAC algorithm, which uses asynchronous scheduling for the sleep time and optimizes the backoff delay, minimum backoff exponential and transmitted power. For the network layer, we optimize at design time, using closed form equations, and at runtime, using channel quality estimations, the number of nodes and hops to transmit packets across a given distance. We also discuss dynamic network routing approaches, which optimize battery lifetime with a limited overhead. Anwar Al-Khateeb, Luciano Lavagno |
LCN | 2 |
| 2010 | Speeding-up heuristic allocation, scheduling and binding with SAT-based abstraction/refinement techniquesabstractHardware synthesis is the process by which system-level, Register Transfer (RT)-level, or behavioral descriptions can be turned into real implementations, in terms of logic gates. Scheduling is one of the most time-consuming steps in the overall design flow, and may become much more complex when performing hardware synthesis from high-level specifications. Exploiting a single scheduling strategy on very large designs is often reductive and potentially inadequate. Furthermore, finding the “best” single candidate among all possible scheduling algorithms is practically infeasible. In this article we introduce a hybrid scheduling approach that is a preliminary step towards a comprehensive solution not yet provided by industrial or by academic solutions. Our method relies on an abstract symbolic representation of data flow nodes (operations) bound to control flow paths: it produces a more realistic lower bound during the prescheduling resource estimation step and speeds up slower but accurate heuristic scheduling techniques, thus achieving a globally improved result. Gianpiero Cabodi, Luciano Lavagno, Marco Murciano, Alex Kondratyev, Yosinori Watanabe |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2009 | Enabling adaptability through elastic clocksabstractPower and performance benefits of scaling are lost to worst case margins as uncertainty of device characteristics is increasing. Adaptive techniques can dynamically adjust the margins required to tolerate variability and recover a significant part of the benefits lost due to worst-case conditions. Additionally, the stringent timing requirements for the synthesis of low-skew clock trees involve higher power consumption, and limit the adaptability to varying operating conditions. This paper introduces an elastic clocking scheme as an adaptive technique to confront variability and provide substantial power savings by dynamically adjusting to operating conditions. The synthesis and sign-off analysis of the elastic clocks is fully automated. Changes to the design flow and sign-off analysis of elastic clocks are addressed by automation of design flow support. Emre Tuncer, Jordi Cortadella, Luciano Lavagno |
DAC | 3 |
| 2009 | A comparison of software platforms for wireless sensor networks: MANTIS, TinyOS, and ZigBeeabstractWireless sensor networks are characterized by very tight code size and power constraints and by a lack of well-established standard software development platforms such as Posix. In this article, we present a comparative study between a few fairly different such platforms, namely MANTIS, TinyOS, and ZigBee, when considering them from the application developer's perspective, that is, by focusing mostly on functional aspects, rather than on performance or code size. In other words, we compare both the tasking model used by these platforms and the API libraries they offer. Sensor network applications are basically event based, so most of the software platforms are also built on considering event handling mechanism, however some use a more traditional thread based model. In this article, we consider implementations of a simple generic application in MANTIS, TinyOS, and the Ember ZigBee development framework, with the goal of depicting major differences between these platforms, and suggesting a programming style aimed at maximizing portability between them. Mohammad Mostafizur Rahman Mozumdar, Luciano Lavagno, Laura Vanzago |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | A Symbolic Algorithm for the Synthesis of Bounded Petri Nets
Josep Carmona 0001, Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexandre Yakovlev |
Petri Nets | 5 |
| 2008 | Porting application between wireless sensor network software platforms: TinyOS, MANTIS and ZigBeeabstractApplication domains of wireless sensor networks are emerging day-by-day, so developers are creating software components at various layers by using platforms such as MANTIS, TinyOS or ZigBee. Since these platforms are quite different in terms of programming paradigm and provided APIs, porting applications between them is difficult and expensive. In this paper, we show techniques that might be used to dramatically ease the porting application code between different software platforms for WSNs. Mohammad Mostafizur Rahman Mozumdar, Francesco Gregoretti, Luciano Lavagno, Laura Vanzago |
ETFA | 3 |
| 2008 | A Framework for Modeling, Simulation and Automatic Code Generation of Sensor Network ApplicationabstractShowing functional correctness by simulation before implementation, and preserving it by automated code generation, is extremely useful to reduce the development time for an embedded application. This is even more true for wireless sensor networks, since their nodes often provide very rudimentary debugging facilities, and sufficiently large networks for realistic analysis may be expensive to deploy. While this approach, also known as model-based design, is becoming quite standard for several domains that have similar constraints as wireless sensor networks, such as automotive electronics, there is a lack of tools for this purpose in the WSN world. In order to fill this gap, in this paper we present a framework (based on Simulink, Stateflow and Embedded Coder) in which an engineer can create sensor network components (both at the application and at the protocol level) that can be used as building blocks to model, simulate and automatically generate code for different underlying platforms and operating systems. Mohammad Mostafizur Rahman Mozumdar, Francesco Gregoretti, Luciano Lavagno, Laura Vanzago, Stefano Olivieri |
SECON | 3 |
| 2007 | A Fully-Automated Desynchronization Flow for Synchronous CircuitsabstractVariability is one of the fundamental problems faced by nano-scale electronic circuits and is expected to become even worse as process technology scales. Desynchronization is a design methodology, which converts a synchronous gate-level circuit into a more robust asynchronous one. In this paper, we describe the first fully-automated desynchronization design flow, based only on contemporary synchronous EDA tools and a new point tool for performing the desynchronization transformation. The flow was used to implement, down to mask layout level, a simple pipelined processor in a 90nm industrial library. We show that the desynchronization methodology can be easily integrated into contemporary industrial EDA flows. Results, on the design implemented, indicate that desynchronized circuits exhibit increased variability tolerance and better average case performance, for a small area and power overhead. Nikolaos Andrikos, Luciano Lavagno, Davide Pandini, Christos P. Sotiriou |
DAC | 2 |
| 2007 | E2RINA: an Energy Efficient and Reliable In-Network Aggregation for Clustered Wireless Sensor NetworksabstractThe paper presents E2RINA, an aggregation algorithm for wireless sensor network applications characterized by clustered topologies, such as building automation and manufacturing plants. Thank to an efficient use of the wireless channel, E2RINA offers the robustness of the gossip-based algorithms and, at the same time, the energy performance of the faster cluster head-based algorithms. A mathematical model to predict the performance of the algorithm with respect to the free variables without the need of extensive simulations was also developed. The model was validated and the robustness of E2RINA by running a simulation model of a test case consisting of a cluster of MICA nodes. Luca Necchi, Alvise Bonivento, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli, Laura Vanzago |
WCNC | 3 |
| 2006 | Desynchronization: Synthesis of Asynchronous Circuits From Synchronous SpecificationsabstractAsynchronous implementation techniques, which measure logic delays at runtime and activate registers accordingly, are inherently more robust than their synchronous counterparts, which estimate worst case delays at design time and constrain the clock cycle accordingly. Desynchronization is a new paradigm to automate the design of asynchronous circuits from synchronous specifications, thus, permitting widespread adoption of asynchronicity without requiring special design skills or tools. In this paper, different protocols for desynchronization are first studied, and their correctness is formally proven using techniques originally developed for distributed deployment of synchronous language specifications. A taxonomy of existing protocols for asynchronous latch controllers, covering, in particular, the four-phase handshake protocols devised in the literature for micropipelines, is also provided. A new controller that exhibits provably maximal concurrency is then proposed, and the performance of desynchronized circuits is analyzed with respect to the original synchronous optimized implementation. Finally, this paper proves the feasibility and effectiveness of the proposed approach by showing its application to a set of real designs, including a complete implementation of the DLX microprocessor architecture Jordi Cortadella, Alex Kondratyev, Luciano Lavagno, Christos P. Sotiriou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | A Time Slice Based Scheduler Model for System Level DesignabstractEfficient evaluation of design choices, in terms of selection of algorithms to be implemented as hardware or software, and finding an optimal HW/SW design mix is an important requirement in the design flow of embedded systems. Time-to-market, faster upgradability and flexibility are some of the driving points to put increasing amounts of functionality as software executed on general purpose processing elements. In this scenario, dividing a monolithic task into multiple interacting tasks, and scheduling them on limited processing elements has become very important for a system designer. The paper presents an approach to model time-slice based task schedulers in the designs where the performance estimate of hardware and software models is less than time-slice accurate. The approach aims to increase the simulation efficiency of designs modeled at system level. We used Metropolis (Balarin, F. et al., IEEE Computer, vol.36, no.4, p.45-52, 2003) as our codesign environment. Luciano Lavagno, Claudio Passerone, Vishal Shah, Yosinori Watanabe |
DATE | 1 |
| 2005 | A BMC-based formulation for the scheduling problem of hardware systems
Gianpiero Cabodi, Alex Kondratyev, Luciano Lavagno, Sergio Nocco, Stefano Quer, Yosinori Watanabe |
Int. J. Softw. Tools Technol. Transf. | 3 |
| 2005 | Quasi-static scheduling of independent tasks for reactive systemsabstractA reactive system must process inputs from the environment at the speed and with the delay dictated by the environment. The synthesis of reactive software from a modular concurrent specification model generates a set of concurrent tasks coordinated by an operating system. This paper presents a synthesis approach for reactive software that is aimed at minimizing the overhead introduced by the operating system and the interaction among the concurrent tasks. A formal model based on Petri nets is used to synthesize the tasks and verify the correctness of their composition. A practical application of the approach is illustrated by means of a real-life industrial example, which shows the significant impact of the approach on the performance of the system. Jordi Cortadella, Alex Kondratyev, Luciano Lavagno, Claudio Passerone, Yosinori Watanabe |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Implementation of a UMTS turbo decoder on a dynamically reconfigurable platformabstractHigh hardware design and mask production costs dictate the need to reuse an architectural platform for as many applications as possible. Embedded multimedia portable devices are required to perform in real time a huge variety of different algorithms, ranging from audio and image processing, to channel coding, to video games and java virtual machines. Dynamically reconfigurable architectures are an effective means to cope with both requirements. However, their effective and efficient use today is hindered by a lack of methodology and tools to extensively explore the hardware/software (HW/SW) design space, without requiring software developers to have a deep knowledge of the underlying architecture. This paper describes one such methodology, which extends the software programming model to the design flow for a reconfigurable processor. Its effectiveness is shown with the case study of a turbo decoder for universal mobile telecommunications systems, in which a remarkable 11X speed-up and 4X reduction of energy requirements with respect to a pure software implementation has been obtained, by mapping the more computation-intensive kernels to the reconfigurable hardware. Alberto La Rosa, Luciano Lavagno, Claudio Passerone |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Guidelines for a graduate curriculum on embedded software and systemsabstractThe design of embedded real-time systems requires skills from multiple specific disciplines, including, but not limited to, control, computer science, and electronics. This often involves experts from differing backgrounds, who do not recognize that they address similar, if not identical, issues from complementary angles. Design methodologies are lacking in rigor and discipline so that demonstrating correctness of an embedded design, if at all possible, is a very expensive proposition that may delay significantly the introduction of a critical product. While the economic importance of embedded systems is widely acknowledged, academia has not paid enough attention to the education of a community of high-quality embedded system designers, an obvious difficulty being the need of interdisciplinarity in a period where specialization has been the target of most education systems. This paper presents the reflections that took place in the European Network of Excellence Artist leading us to propose principles and structured contents for building curricula on embedded software and systems. Paul Caspi, Alberto L. Sangiovanni-Vincentelli, Luís Almeida 0001, Albert Benveniste, Bruno Bouyssounouse, Giorgio C. Buttazzo, Ivica Crnkovic, Werner Damm, Jakob Engblom, Gerhard Fohler, Marisol García-Valls, Hermann Kopetz, Yassine Lakhnech, François Laroussinie, Luciano Lavagno, Giuseppe Lipari, Florence Maraninchi, Philipp Peti, Juan Antonio de la Puente, Norman Scaife, Joseph Sifakis, Robert de Simone, Martin Törngren, Paulo Veríssimo, Andy J. Wellings, Reinhard Wilhelm, Tim A. C. Willemse, Wang Yi 0001 |
ACM Trans. Embed. Comput. Syst. | 15 |
| 2004 | SoftContract: an Assertion-Based Software Development Process that Enables Design-by-ContractabstractThis paper discusses a model-based design flow for requirements in distributed embedded software development. Such requirements are specified using a language similar to linear temporal logic which allows one to reason about time and sequencing. They consist of assertions which must hold for a design, given some assumptions on its environment. They can be checked both during simulation and, at least for a subset, even on the target. The key contribution of the paper is the extension to the embedded software domain of assertion-based verification, and the automated generation of property-checking code in multiple target languages, from simulation, to prototyping, to final production. Jean-Yves Brunel, Marco Di Natale, Alberto Ferrari, Paolo Giusto, Luciano Lavagno |
DATE | 5 |
| 2004 | From Synchronous to Asynchronous: An Automatic ApproachabstractThis paper presents a methodology to derive asynchronous circuits from optimized synchronous circuits by replacing the clock distribution tree by a handshaking network. A case study shows the applicability of the method and the potential benefits of de-synchronizing synchronous circuits. Jordi Cortadella, Alex Kondratyev, Luciano Lavagno, Kelvin Lwin, Christos P. Sotiriou |
DATE | 3 |
| 2004 | Implementation of a UMTS Turbo-Decoder on a Dynamically Reconfigurable PlatformabstractModern embedded systems must execute a variety of high performance real-time tasks, such as audio and image compression, channel coding and encoding, etc. Reconfigurable platforms can effectively be used in these cases, because they allow to re-use the architecture for as many applications as possible. The paper describes the implementation of a UMTS turbo-decoder on one such platform, the XiRisc reconfigurable processor. Our goal is to test the development framework and design flow that we already developed on a real industrial example. Our results shows that, with some manual effort from the designer, very good performance improvements can be achieved, using a flow close to embedded software development. Alberto La Rosa, Claudio Passerone, Francesco Gregoretti, Luciano Lavagno |
DATE | 4 |
| 2004 | Coping with The Variability of Combinational Logic DelaysabstractThis paper proposes a technique for creating a combinational logic network with an output that signals when all other outputs have stabilized. The method is based on dual-rail encoding, and guarantees low timing overhead and reasonable area and power overhead. We discuss various scenarios in which completion detection can be used to measure the delay of a synchronous circuit at fabrication time or at run time, and of an asynchronous circuit at run time. We conclude by showing, on a large set of benchmarks, the effectiveness of the proposed technique. Jordi Cortadella, Alex Kondratyev, Luciano Lavagno, Christos P. Sotiriou |
ICCD | 3 |
| 2004 | Quasi-static Scheduling for Concurrent Architectures
Jordi Cortadella, Alex Kondratyev, Luciano Lavagno, Alexander Taubin, Yosinori Watanabe |
Fundam. Informaticae | 3 |
| 2004 | Designing an asynchronous microcontroller using PipefitterabstractThis paper discusses how Pipefitter, a tool chain that implements a fully automated synthesis flow for asynchronous circuits, can be used to design a simple asynchronous microcontroller. The use of register transfer level (RTL)-like Verilog hardware description languages (HDL) as the input format makes the first steps of the design flow (i.e., specification and simulation) very easy for the designer. Pipefitter directly synthesizes the control unit as a hazard-free standard cell netlist, uses a genetic algorithm to perform binding and multiplexer optimization for the datapath and allows the user to manually specify the binding. It also produces a synthesizable Verilog specification for the datapath, as well as a set of scripts driving both its synthesis and timing analysis by state-of-the-art commercial synchronous RTL and logic synthesis tools. The automated insertion of matched delays completes the logic design, and hands off the netlist to the standard cell-based layout tools. The example presented in this brief, shows how Pipefitter can be effectively used for the design of asynchronous application specific integrated circuits. Ivan Blunno, Luciano Lavagno |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Hardware/Software Design Space Exploration for a Reconfigurable Processor
Alberto La Rosa, Luciano Lavagno, Claudio Passerone |
DATE | 2 |
| 2003 | Design Space Exploration for a Wireless Protocol on a Reconfigurable Platform
Laura Vanzago, Bishnupriya Bhattacharya, Joel Cambonie, Luciano Lavagno |
DATE | 4 |
| 2003 | Guest Editorial
David T. Blaauw, Luciano Lavagno |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | False Path Elimination in Quasi-Static SchedulingabstractWe have developed a technique to compute a Quasi Static Schedule of a concurrent specification for the software partition of an embedded system. Previous work did not take into account correlations among run-time values of variables, and therefore tried to find a schedule for all possible outcomes of conditional expressions. This is advantageous on one hand, because by abstracting data values one can find schedules in many cases for an originally undecidable problem. On the other hand it may lead to exploring false paths, i.e., paths that can never happen at run-time due to constraints on how the variables are updated. This affects the applicability of the approach, because it leads to an explosion in the running time and the memory requirements of the compile-time scheduler itself. Even worse, it also leads to an increase in the final code size of the generated software. In this paper we propose a semi-automatic algorithm to solve the problem of false paths: the designer identifies and tags critical expressions, and synchronization channels are automatically added to the specification to drive the search of a schedule. G. Arrigoni, L. Duchini, Claudio Passerone, Luciano Lavagno, Yosinori Watanabe |
DATE | 4 |
| 2002 | Processes, Interfaces and Platforms. Embedded Software Modeling in Metropolis
Felice Balarin, Luciano Lavagno, Claudio Passerone, Yosinori Watanabe |
EMSOFT | 2 |
| 2002 | Designing an Asynchronous Microcontroller Using PipefitterabstractThis paper discusses how Pipefitter, a tool chain that implements a fully automated synthesis flow for asynchronous circuits, can be used to design a simple asynchronous microcontroller. The use of RTL-like Verilog HDL as the input format makes the first steps of the design flow (i.e. specification and simulation) very easy for the designer. Pipefitter directly synthesizes the control unit as a hazard-free standard cell netlist, uses a genetic algorithm to perform binding and multiplexer optimization for the data path, allows the user to manually specify the binding, and can automatically pipeline a sequential specification. It also produces a synthesizable Verilog specification for the Data Path, as well as a set of scripts driving both its synthesis and timing analysis by state-of-the-art commercial synchronous RTL and logic synthesis tools. The automated insertion of matched delays completes the logic design, and hands off the netlist to the standard cell-based layout tools. The example presented in this paper shows how Pipefitter can be effectively used for the design of asynchronous application specific integrated circuits. Ivan Blunno, Luciano Lavagno |
ICCD | 2 |
| 2002 | Models of IP's for Automotive Virtual Integration PlatformsabstractSummary form only given.The concept of virtual integration platform plays a key role in any novel methodology that is trying to address earlier validation of distributed applications in regular and faulty conditions. The methodology must rely upon libraries that model the most important features of the commonly used IP's in the automotive segment such as FlexRay, the emerging bus protocol for safety critical applications supported by BMW, Daimler-Chrysler, Philips, Bosch, and Motorola, OSEK compliant RTOSes and protocol stacks, microprocessors such as Motoro/IBM PowerPC, Infineon 167, NEC v850, Tricore, ST 10, and Janus. We believe that tools must support the easy plug and play of the IP models in a seamless way to the user. For example, it must be possible to run a fast simulation at the token level (frames) to provide insights about the best network protocol configuration within a reasonable accuracy for the estimated frame latency. Next, it must be possible to export such a configuration to (semi)-automatically configure the downstream and more refined bus protocol models for the finer grain validation step. Both steps must rely upon interchangeable IP's with clear interfaces and trade-offs between simulation speed and accuracy of the timing estimates. In this paper, we present two examples of models of IP's that can be used at two different steps in the design exploration, the token-level/cycle approximate transaction based level and the cycle accurate level. The first example is the Universal Communication Model (UCM) that captures the main common features of the most relevant bus protocols such as topology, redundancy, arbitration, etc. The model enables quick token-level simulations. The user is able to determine the communication cycle layout and bus scheduling, k-matrix, and then export it for the configuration of downstream more refined models such as the Motorola FlexRay cycle accurate transaction based model. Bus delays are as important as task execution delays and RTOS switching overheads. In the second example we introduce Janus, a multi-processor micro-controller for power train applications. The cycle approximate transaction based model of Janus can be used to assess the ECU HW/SW partitioning, in particular to quickly explore different task scheduling and allocation. Then, this model is refined and exported to configure a HW/SW co-verification tool for the cycle accurate validation of the ECU HW/SW architecture. In an example scenario, an engine control ECU is providing information about the engine (e.g. engine revolution speed) to a gear control ECU over a CAN bus (the latter typically requires precise revolution speed to operate and could also require to set the engine operation condition). In this scenario, car and subsystem makers play different roles in order to provide a virtual model of the system to validate the functionality and the performance before going to implementation. The same models can then be used to march toward implementation. Paolo Giusto, Jean-Yves Brunel, Alberto Ferrari, Eliane Fourgeau, Luciano Lavagno, Barry O'Rourke, Alberto L. Sangiovanni-Vincentelli, Emanuele Guasto |
ICCD | 5 |
| 2002 | Automotive Virtual Integration Platforms: Why's, What's, and How'sabstractIn this paper, we present the new concept of virtual integration platform for automotive electronics. The platform provides the basis for a novel methodology in which the integration of sub-systems is performed much earlier in the design cycle. As a result, cost reduction in the final implementation and in the design process can be achieved. In addition, early and repeatable fault analysis can be performed therefore easing the task of system safety proving. Paolo Giusto, Jean-Yves Brunel, Alberto Ferrari, Eliane Fourgeau, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ICCD | 5 |
| 2002 | Lazy transition systems and asynchronous circuit synthesis withrelative timing assumptionsabstractThis paper presents a design flow for timed asynchronous circuits. It introduces lazy transitions systems as a new computational model to represent the timing information required for synthesis. The notion of laziness explicitly distinguishes between the enabling and the firing of an event in a transition system. Lazy transition systems can be effectively used to model the behavior of asynchronous circuits in which relative timing assumptions can be made on the occurrence of events. These assumptions can be derived from the information known a priori about the delay of the environment and the timing characteristics of the gates that will implement the circuit. The paper presents the necessary conditions to generate circuits and a synthesis algorithm that exploits the timing assumptions for optimization. It also proposes a method for back-annotation that derives a set of sufficient timing constraints that guarantee the correctness of the circuit. Jordi Cortadella, Michael Kishinevsky, Steven M. Burns, Alex Kondratyev, Luciano Lavagno, Kenneth S. Stevens, Alexander Taubin, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2002 | Cosimulation-based power estimation for system-on-chip designabstractWe present efficient power estimation techniques for hardware-software (HW-SW) system-on-chip (SoC) designs. Our techniques are based on concurrent and synchronized execution of multiple power estimators that analyze different parts of the SoC (we refer to this as coestimation), driven by a system-level simulation master. We motivate the need for power coestimation, and demonstrate that performing independent power estimation for the various system components can lead to significant errors in the power estimates, especially for control-intensive and reactive-embedded systems. We observe that the computation time for performing power coestimation is dominated by: i) the requirement to analyze/simulate some parts of the system at lower levels of abstraction in order to obtain accurate estimates of timing and switching activity information and ii) the need to communicate between and synchronize the various simulators. Thus, a naive implementation of power coestimation may be too inefficient to be used in an iterative design exploration framework. To address this issue, we present several acceleration (speed-up) techniques for power coestimation. The acceleration techniques are energy caching, software power macro-modeling, and statistical sampling. Our speed-up techniques reduce the workload of the power estimators for the individual SoC components, as well as their communication/synchronization overhead. Experimental results indicate that the use of the proposed acceleration techniques results in significant (8/spl times/ to 87/spl times/) speed-ups in SOC power estimation time, with minimal impact on accuracy. We also show the utility of our coestimation tool to explore system-level power tradeoffs for a TCP/IP check-sum engine subsystem. Marcello Lajolo, Anand Raghunathan, Sujit Dey, Luciano Lavagno |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2001 | A software development tool chain for a reconfigurable processorabstractArticle Share on A software development tool chain for a reconfigurable processor Authors: Alberto La Rosa Dipartimento di Elettronica, Politecnico di Torino, Italy Dipartimento di Elettronica, Politecnico di Torino, ItalyView Profile , Luciano Lavagno Dipartimento di Elettronica, Politecnico di Torino, Italy Dipartimento di Elettronica, Politecnico di Torino, ItalyView Profile , Claudio Passerone Dipartimento di Elettronica, Politecnico di Torino, Italy Dipartimento di Elettronica, Politecnico di Torino, ItalyView Profile Authors Info & Claims CASES '01: Proceedings of the 2001 international conference on Compilers, architecture, and synthesis for embedded systemsNovember 2001Pages 93–98https://doi.org/10.1145/502217.502232Published:16 November 2001Publication History 16citation670DownloadsMetricsTotal Citations16Total Downloads670Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Alberto La Rosa, Luciano Lavagno, Claudio Passerone |
CASES | 2 |
| 2001 | A Hardware/Software Co-design Flow and IP Library Based of SimulinkTMabstractThis paper describes a design flow for data-dominated embedded systems. We use The Mathworks ’ Simulink environment for functional specification and algorithmic analysis. We developed a library of Simulink blocks, each parameterized by design choices such as implementation (software, analog or digital hardware, ) and numerical accuracy (resolution, S/N ratio). Each block is equipped with empirical models for cost (code size, chip area) and performance (timing, energy), based on surface fitting from actual measurements. We also developed an analysis toolbox that quickly evaluates algorithm and parameter choices performed by the designer, and presents the results for fast feedback. The chosen block netlist is then ready for implementation, by using a customization of The Mathworks ’ Real Time Workshop to generate a VHDL netlist for FPGA implementation, as well as embedded software for DSP implementation. 1. Leonardo Maria Reyneri, Franceso Cucinotta, Alessandro Serra, Luciano Lavagno |
DAC | 4 |
| 2001 | Generation of minimal size code for scheduling graphsabstractThis paper proposes a procedure for minimizing the code size of sequential programs for reactive systems. It identifies repeated code segments (a generalization of basic blocks to directed rooted trees) and finds a minimal covering of the input control flow graphs with code segments. The segments are disjunct, i.e. no two segments have the same code in common. The program is minimal in the sense that the number of code segments is minimum under the property of disjunction for the given control flow specification. The procedure makes no assumption on the target processor architecture, and is meant to be used between task synthesis algorithms from a concurrent specification and a standard compiler for the target architecture. It is aimed at optimizing the size of very large, automatically generated flat code, and extends dramatically the scope of classical common sub-expression identification techniques. The potential effectiveness of the proposed approach is demonstrated through preliminary experiments. Claudio Passerone, Yosinori Watanabe, Luciano Lavagno |
DATE | 3 |
| 2001 | System-Level Power/Performance Analysis of Portable Multimedia Systems Communicating over Wireless ChannelsabstractThis paper presents a new methodology for system-level power and performance analysis of wireless multimedia systems. More precisely, we introduce an analytical approach based on concurrent processes modeled as Stochastic Automata Networks (SANs) that can be effectively used to integrate power and performance metrics in system-level design. We show that 1) under various input traces and wireless channel conditions, the average-case behavior of a multimedia system consisting of a video encoder/decoder pair is characterized by very different probability distributions and power consumption values and 2) in order to identify the best trade-off between power and performance figures, one must take into consideration the entire environment (i.e., encoder, decoder and channel) for which the system is being designed. Compared to using simulation, our analytical technique reduces the time needed to find the steady-state behavior by orders of magnitude, with some limited loss in accuracy compared to the exact solution. We illustrate the potential of our methodology using the MPEG-2 video as the driver application. Radu Marculescu, Amit Nandi, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 3 |
| 2001 | Synchronous approach to the functional equivalence of embeddedsystem implementationsabstractDesign space exploration is the process of analyzing several functionally equivalent alternatives to determine the most suitable one. A fundamental question is whether an implementation is consistent with the high-level specification or whether two implementations are "equivalent." The synchronous assumption has made it possible to develop efficient procedures for establishing functional equivalence between different implementations in the domains of synchronous circuits and synchronous reactive systems. We extend this notion to embedded systems that do not satisfy the synchronous assumption inside their boundaries but only at the interface with the environment. Leveraging this property, we define synchronous equivalence for embedded systems that strongly resembles the concept of functional equivalence for sequential circuits. We develop efficient synchronous equivalence analysis algorithms for embedded system designs. The efficiency comes from analyzing the behavior statically on abstract representations, at a cost that some of the negative results may be false, i.e. the analysis is conservative. We develop primitives for making the representation more/less abstract, trading off complexity of the algorithms with the conservativeness of the results. We apply our analysis algorithms to an ATM switch and demonstrate that synchronous equivalence opens design exploration avenues uncharted before. Harry Hsieh, Felice Balarin, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Formal Models for Communication-Based Design
Alberto L. Sangiovanni-Vincentelli, Marco Sgroi, Luciano Lavagno |
CONCUR | 3 |
| 2000 | Task generation and compile-time scheduling for mixed data-control embedded softwareabstractThe problem of optimal software synthesis for concurrent processes to be implemented on a single processor is addressed. The approach calls for the representation of the concurrent processes with Petri nets that give a theoretical foundation for the scheduling algorithm that sequentializes the concurrent processes and for the code generation step. The approach maximizes the amount of static scheduling to reduce the need of context switch and operating system intervention. Experimental results show the potential of our method to reduce software design time and errors. Jordi Cortadella, Alex Kondratyev, Luciano Lavagno, Marc Massot, Sandra Moral, Claudio Passerone, Yosinori Watanabe, Alberto L. Sangiovanni-Vincentelli |
DAC | 3 |
| 2000 | Efficient methods for embedded system design space explorationabstractDesign space exploration is the process of analyzing several functionally equivalent alternatives to determine the most suitable one. The synchronous assumption has made it possible to develop efficient procedures for establishing functional equivalence between different implementations in the domain of synchronous circuits as well as in the domain of synchronous reactive systems. We extend this notion to embedded systems that do not satisfy the synchronous assumption inside their boundaries but only at the interface with the environment Leveraging this property, we developed efficient synchronous equivalence analysis algorithms for embedded systems with loops and architectures with multiple computational units. We demonstrate our method on an ATM switch containing many interacting components. Harry Hsieh, Felice Balarin, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
DAC | 3 |
| 2000 | Free MDD-Based Software Optimization Techniques for Embedded SystemsabstractEmbedded systems make a heavy use of software to perform real-time embedded control tasks. Embedded software is characterized by a relatively long lifetime and by tight cost, performance and safety constraints. Several super-optimization techniques for embedded softwares based on multi-valued decision diagram (MDD) representations have been described in the literature, but they all share the same basic limitation. They are based on standard ordered MDD (OMDD) packages, and hence require a used order of evaluation for the MDD variables on every execution path. Free MDDs (FMDDs) lift this limitation, and hence open up more optimization opportunities. Finding the optimal variable ordering for FMDDs is a very difficult problem. Hence in this paper we describe a heuristic procedure that performs well in practice, and is based on FMDD cost estimation applied to recursive cofactoring. Experimental results show that our new variable ordering method obtains often. Smaller embedded software than previous (sifting-based) methods. Chunghee Kim, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
DATE | 2 |
| 2000 | Efficient Power Co-Estimation Techniques for System-on-Chip DesignabstractWe present efficient power estimation techniques for HW/SW System-On-Chip (SOC) designs. Our techniques are based on concurrent and synchronized execution of multiple power estimators that analyze different parts of the SOC (we refer to this as co-estimation), driven by a system-level simulation master. We motivate the need for power co-estimation, and demonstrate that performing independent power estimation for the various system components can lead to significant errors in the power estimates, especially for control-intensive and reactive embedded systems. We observe that the computation time for performing power co-estimation is dominated by: (i) the requirement to analyze/simulate some parts of the system at lower levels of abstraction in order to obtain accurate estimates of timing and switching activity information and (ii) the need to communicate between and synchronize the various simulators. Thus, a naive implementation of power co-estimation may be too inefficient to be used in an iterative design exploration framework. To address this issue, we present several acceleration (speedup) techniques for power co-estimation. The acceleration techniques are energy caching, software power macromodeling, and statistical sampling. Our speedup techniques reduce the workload of the power estimators for the individual SOC components, as well as their communication/synchronization overhead. Experimental results indicate that the use of the proposed acceleration techniques results in significant (8/spl times/ to 87/spl times/) speedups in SOC power estimation time, with minimal impact on accuracy. We also show the utility of our co-estimation tool to explore system-level power tradeoffs for a TCP/IP network interface card sub-system and an automotive controller. Marcello Lajolo, Anand Raghunathan, Sujit Dey, Luciano Lavagno |
DATE | 4 |
| 2000 | Evaluating System Dependability in a Co-Design FrameworkabstractThe widespread adoption of embedded microprocessor-based systems for safety critical applications mandates the use of co-design tools able to evaluate system dependability at every step of the design cycle. In this paper, we describe how fault injection techniques have been integrated in an existing co-design tool and which advantages come from the availability of such an enhanced tool. The effectiveness of the proposed tool is assessed on a simple case study. Marcello Lajolo, Maurizio Rebaudengo, Matteo Sonza Reorda, Massimo Violante, Luciano Lavagno |
DATE | 5 |
| 2000 | Verification of Similar FSMs by Mixing Incremental Re-encoding, Reachability Analysis, and Combinational Checks
Stefano Quer, Gianpiero Cabodi, Paolo Camurati, Luciano Lavagno, Ellen Sentovich, Robert K. Brayton |
Formal Methods Syst. Des. | 4 |
| 2000 | Simulink-Based HW/SW Codesign of Embedded Neuro-Fuzzy SystemsabstractWe propose a semi-automatic HW/SW codesign flow for low-power and low-cost Neuro-Fuzzy embedded systems. Applications range from fast prototyping of embedded systems to high-speed simulation of Simulink models and rapid design of Neuro-Fuzzy devices. The proposed codesign flow works with different technologies and architectures (namely, software, digital and analog). We have used The Mathworks' Simulink environment for functional specification and for analysis of performance criteria such as timing (latency and throughput), power dissipation, size and cost. The proposed flow can exploit trade-offs between SW and HW as well as between digital and analog implementations, and it can generate, respectively, the C, VHDL and SKILL codes of the selected architectures. Leonardo Maria Reyneri, Marcello Chiaberge, Luciano Lavagno, Begoña del Pino |
Int. J. Neural Syst. | 3 |
| 1999 | Fast Instruction Cache Simulation Strategies in a Hardware/Software Co-Design EnvironmentabstractCache memories are one of the main factors that affect software performance, and their use is becoming increasingly common even in embedded systems. Efficient analysis of the effects of parameter variations (cache dimensions, degree of associativity, replacement policy, line size, ...) is at the same time an essential and very time-consuming aspect of embedded system design, whose complexity increases when multi-tasking and real-time aspects must be considered. We propose a new simulation-based methodology, focused on an approximate model of the cache and of the multi-tasking reactive software, that allows one to trade off smoothly between accuracy and simulation speed. In particular, we propose to accurately consider intra-task conflicts, but approximate inter-task conflicts by considering only a finite number of previous task executions. The rationale for this choice can be found in a common pattern in embedded systems, where a "normal" data flow results in a regular intra-task common flow, interrupted from time to time by some urgent event, that pessimistically can be consider as disrupting the cache behavior. The approach is conservative because re-execution of a task after a large amount of time will always be considered as not in cache, and the simulation speed-up is considerable, as shown by theoretical analysis and experimental results. Marcello Lajolo, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ASP-DAC | 2 |
| 1999 | Automatic Synthesis and Optimization of Partially Specified Asynchronous SystemsabstractA method for automating the synthesis of asynchronous control circuits from high level (CSP-like) and/or partial STG (involving only functionally critical events) specifications is presented.The method solves two key subtasks in this new, more flexible, design flow: handshake expansion, i.e. inserting reset events with maximum concurrency, and event reshuffling under interface and concurrency constraints, by means of concurrency reduction.In doing so, the algorithm optimizes the circuit both for size and performance.Experimental results show a significant increase in the solution space explored when compared to existing CSP-based or STG-based synthesis tools. Alex Kondratyev, Jordi Cortadella, Michael Kishinevsky, Luciano Lavagno, Alexandre Yakovlev |
DAC | 4 |
| 1999 | ECL: A Specification Environment for System-Level DesignabstractArticle ECL: a specification environment for system-level design Share on Authors: Luciano Lavagno Cadence Berkeley Laboratories, 2001 Addison Street, 3rd floor, Berkeley, CA Cadence Berkeley Laboratories, 2001 Addison Street, 3rd floor, Berkeley, CAView Profile , Ellen Sentovich Cadence Berkeley Laboratories, 2001 Addison Street, 3rd floor, Berkeley, CA Cadence Berkeley Laboratories, 2001 Addison Street, 3rd floor, Berkeley, CAView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 511–516https://doi.org/10.1145/309847.309989Published:01 June 1999 64citation212DownloadsMetricsTotal Citations64Total Downloads212Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Luciano Lavagno, Ellen Sentovich |
DAC | 1 |
| 1999 | Synthesis of Embedded Software Using Free-Choice Petri NetsabstractSoftware synthesis from a concurrent functional specification is a key problem in the design of embedded systems.A concurrent specification is well-suited for medium-grained partitioning.However, in order to be implemented in software, concurrent tasks need to be scheduled on a shared resource (the processor).The choice of the scheduling policy mainly depends on the specification of the system.For pure dataflow specifications, it is possible to apply a fully static scheduling technique, while for algorithms containing data-dependent control structures, like the if-then-else or while-do constructs, the dynamic behaviour of the system cannot be completely predicted at compile time and some scheduling decisions are to be made at run-time.For such applications we propose a Quasi-static scheduling (QSS) algorithm that generates a schedule in which run-time decisions are made only for data-dependent control structures.We use Free Choice Petri Nets (FCPNs), as underlying model, and define quasi-static schedulability for FCPNs.The proposed algorithm is complete, in that it can solve QSS for any FCPN that is quasi-statically schedulable.Finally, we show how to synthesize from a quasi-static schedule a C code implementation that consists of a set of concurrent tasks. Marco Sgroi, Luciano Lavagno |
DAC | 2 |
| 1999 | Fast Hardware-Software Co-simulation Using VHDL ModelsabstractWe describe a technique for hardware-software co-simulation that is almost cycle-accurate, and does nor require the use of interprocess communication for a C language interface for the software components. Software is modeled by using behavioral VHDL constructs, annotated with timing information derived from basic block-level timing estimates. Hardware is also modeled in VHDL, and can be either pre-existing intellectual property or synthesized to RTL from a functional specification. Execution of the VHDL processes modeling software tasks is coordinated by a process emulating the target RTOS behavior. The effects of changing the hardware/software partition can be quickly estimated by changing a process parameter defining its target implementation and the processor on which it is running. Bassam Tabbara, Marco Sgroi, Alberto L. Sangiovanni-Vincentelli, Enrica Filippi, Luciano Lavagno |
DATE | 5 |
| 1999 | What is the cost of delay insensitivity?abstractDeep submicron technology calls for new design techniques, in which wire and gate delays are accounted to have equal or nearly equal effect on circuit behaviour. Asynchronous speed-independent (SI) circuits, whose behaviour is only robust to gate delay variations, may be too optimistic. On the other hand, building circuits totally delay-insensitive (DI), for both gates and wires, is impractical. The paper presents an approach for automated synthesis of globally DI and locally SI circuits. It is based on order relaxation, a simple graphical transformation of a circuit's behavioural specification, for which the Signal Transition Graph, an interpreted Petri net, is used. The method is successfully tested on a set of benchmarks and a realistic design example. It proves effective showing average cost of DI interfacing at about 40% for area and 20% for speed. Hiroshi Saito, Alex Kondratyev, Jordi Cortadella, Luciano Lavagno, Alexandre Yakovlev |
ICCAD | 4 |
| 1999 | Logic decomposition of speed-independent circuitsabstractLogic decomposition is a well-known problem in logic synthesis, but it poses new challenges when targeted to speed-independent circuits. The decomposition of a gate into smaller gates must preserve not only the functional correctness of a circuit but also speed independence, i.e., hazard freedom under unbounded gate delays. This paper presents a new method for logic decomposition of speed-independent circuits that solves the problem in two major steps: (1) logic decomposition of complex gates and (2) insertion of new signals that preserve hazard freedom. The method is shown to be more general than previous approaches and its effectiveness is evaluated by experiments on a set of benchmarks. Alex Kondratyev, Jordi Cortadella, Michael Kishinevsky, Luciano Lavagno, Alexandre Yakovlev |
Proc. IEEE | 4 |
| 1999 | Synthesis of software programs for embedded control applicationsabstractSoftware components for embedded reactive real-time applications must satisfy tight code size and run-time constraints. Cooperating finite state machines provide convenient intermediate format for embedded system co-synthesis, between high-level specification languages and software or hardware implementations. We propose a software generation methodology that takes advantage of a restricted class of specifications and allows for tight control over the implementation cost. The methodology exploits several techniques from the domain of Boolean function optimization. We also describe how the simplified control/data-flow graph used as an intermediate representation can be used to accurately estimate the size and timing cost of the final executable code. Felice Balarin, Massimiliano Chiodo, Paolo Giusto, Harry Hsieh, Attila Jurecska, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli, Ellen Sentovich, Kei Suzuki |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 1999 | Decomposition and technology mapping of speed-independent circuits using Boolean relationsabstractThis paper presents a new technique for decomposition and technology mapping of speed-independent circuits. An initial circuit implementation is obtained in the form of a netlist of complex gates, which may not be available in the design library. The proposed method iteratively performs Boolean decomposition of each such gate F into a two-input combinational or sequential gate G available in the library and two gates H/sub 1/ and H/sub 2/ simpler than F, while preserving the original behavior and speed-independence of the circuit. To extract functions for H/sub 1/ and H/sub 2/ the method uses Boolean relations as opposed to the less powerful algebraic factorization approach used in previous methods. After logic decomposition, the overall library matching and optimization is carried out. Logic resynthesis, performed after speed-independent signal insertion for H/sub 1/ and H/sub 2/, allows for sharing of decomposed logic. Overall, this method is more general than the existing techniques based on restricted decomposition architectures, and thereby leads to better results in technology mapping. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Enric Pastor, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 1998 | A Case Study in Embedded System Design: An Engine Control UnitabstractA number of techniques and software tools for embedded system design have been recently proposed. However, the current practice in the designer community is heavily based on manual techniques and on past experience rather than on a rigorous approach to design. To advance the state of the art it is important to address a number of relevant design problems and solve them to demonstrate the power of the new approaches. Tullio Cuatto, Claudio Passerone, Luciano Lavagno, Attila Jurecska, Antonino Damiano, Claudio Sansoè, Alberto L. Sangiovanni-Vincentelli |
DAC | 3 |
| 1998 | Don't Care-Based BDD Minimization for Embedded SoftwareabstractThis paper explores the use of don’t cares in software synthesis for embedded systems. Embedded systems have extremely tight realtime and code/data size constraints, that make expensive optimizations desirable. We propose applying BDD minimization techniques in the presence of a don’t care set to synthesize code for extended Finite State Machines from a BDD-based representation of the FSM transition function. The don’t care set can be derived from local analysis (such as unused state codes or don’t care inputs) as well as from external information (such as impossible input patterns). We show experimental results, discuss their implications, the interaction between BDD-based minimization and dynamic variable reordering, and propose directions for future work. 1 Youpyo Hong, Peter A. Beerel, Luciano Lavagno, Ellen Sentovich |
DAC | 3 |
| 1998 | Interface synthesis: a vertical slice from digital logic to software componentsabstractInterface synthesis seeks to automate the process of interconnecting component.There are many levels of interconnection that must be considered including electrical, power, logic, register-transfer, device drivers, and higher software levels.This presentation ~cover a vertical stice of the interfacing problem from digital logic up to coordinating communications between software components.The focus Ml be within an embedded systems context where the interfacing is between procwsors and memory and peripheral blocks as is the case in system-on-a-chip design.The structure of the tutorial ~parallel the history of CAD efforts in this area.We wi~ begin with the early work in interface specMcation and logic synthesis then proceed on to the problems of interconnecting hardware to processors and their software, and ftish with purely software interfaces involving inter-process communication and protocols between multiple procwsors.At each level we ti discuss specflcation, synthesis, and verification aspecti as we~as higtight the currently avdable tools and on-going r=earch efforts. Gaetano Borriello, Luciano Lavagno, Ross B. Ortega |
ICCAD | 2 |
| 1998 | Lazy transition systems: application to timing optimization of asynchronous circuitsabstractThis paper introduces bzy Transitions Systems &zTSs).The notion of laziness exDlicitlv distinguishes behveen the enabling and the firing of an e;ent in"a transition system.LzTSS can be effectively used to model the behavior of asynchronous circuits in whicfi relative timing assumptions cm-be made on the occurrence of events.These assumptions can be derived from the information known a priori about fie de]ay of fie environment and the timing characteristics of the gates that will implement the circuit.The paper presents necessary conditions to synthesize circuits with a correct behavior under the given timing assumr3tions.Preliminary results show that significant area and performance improvements can be obtained by exploiting the extra "don't care" space implicitly provided by the Iazmess of the events. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexander Taubin, Alexandre Yakovlev |
ICCAD | 4 |
| 1998 | Deriving Petri Nets for Finite Transition SystemsabstractThis paper presents a novel method to derive a Petri net from any specification model that can be mapped into a state-based representation with arcs labeled with symbols from an alphabet of events (a Transition System, TS). The method is based on the theory of regions for Elementary Transition Systems (ETS). Previous work has shown that, for any ETS, there exists a Petri Net with minimum transition count (one transition for each label) with a reachability graph isomorphic to the original Transition System. Our method extends and implements that theory by using the following three mechanisms that provide a framework for synthesis of safe Petri nets from arbitrary TSs. First, the requirement of isomorphism is relaxed to bisimulation of TSs, thus extending the class of synthesizable TSs to a new class called Excitation-Closed Transition Systems (ECTS). Second, for the first time, we propose a method of PN synthesis for an arbitrary TS based on mapping a TS event into a set of transition labels in a PN. Third, the notion of irredundant region set is exploited, to minimize the number of places in the net without affecting its behavior. The synthesis method can derive different classes of place-irredundant Petri Nets (e.g., pure, free choice, unique choice) from the same TS, depending on the constraints imposed on the synthesis algorithm. This method has been implemented and applied in different frameworks. The results obtained from the experiments have demonstrated the wide applicability of the method. Jordi Cortadella, Michael Kishinevsky, Luciano Lavagno, Alexandre Yakovlev |
IEEE Trans. Computers | 3 |
| 1998 | Partial-scan delay fault testing of asynchronous circuitsabstractAsynchronous circuits operate correctly only under timing assumptions. Hence testing those circuits for delay faults is crucial. Previous work has shown that full-scan delay-fault testing of asynchronous circuits is feasible. In this work, we tackle the problem of partial-scan testing, which requires test-pattern generation on a sequential circuit. We show how this problem can be effectively reduced to a classical problem of stuck-at test-pattern generation for a related combinational circuit. The reduction is done in three steps. The first step reduces testing of an asynchronous sequential circuit, by using a partial-scan approach, to testing an object called an asynchronous net, in which feedback is allowed only inside asynchronous memory elements. We then decompose the problem of testing asynchronous nets into that of initializing memory elements (the second step), followed by robust path delay fault testing (the third step). We provide effective procedures to solve both the initialization and the test-pattern generation problem. The technique is complete, automated, and requires only partial scan of some memory element outputs. Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexander Saldanha, Alexander Taubin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1998 | Modeling reactive systems in JavaabstractWe present an application of the Java TM programming language to specify and implement reactive real-time systems. We have developed and tested a collection of classes and methods to describe concurrent modules and their asynchronous communication by means of signals. The control structures are closely patterned after those of the synchronous language Esterel , succinctly describing concurrency, sequencing and preemption. We show the user-friendliness and efficiency of the proposed technique by using an example from the automotive domain. Claudio Passerone, Claudio Sansoè, Luciano Lavagno, Patrick C. McGeer, Roberto Passerone, Alberto L. Sangiovanni-Vincentelli |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 1997 | Trade-off evaluation in embedded system design via co-simulationabstractCurrent design methodologies for embedded systems often force the designer to evaluate early in the design process architectural choices that will heavily impact the cost and performance of the final product. Examples of these choices are hardware/software partitioning, choice of the micro-controller, and choice of a run-time scheduling method. This paper describes how to help the designer in this task, by providing a flexible co-simulation environment in which these alternatives can be interactively evaluated. Claudio Passerone, Luciano Lavagno, Claudio Sansoè, Massimiliano Chiodo, Alberto L. Sangiovanni-Vincentelli |
ASP-DAC | 2 |
| 1997 | Disjunctive Partitioning and Partial Iterative Squaring: An Effective Approach for Symbolic Traversal of Large CircuitsabstractExtending the applicability of reachability analysis to large andreal circuits is a key issue.In fact they are still limited forthe following reasons: peak BDD size during image computation,BDD explosion for representing state sets and very highsequential depth.Following the promising trend of partitioning and problem decomposition,we present a new approach based on a disjunctivepartitioned transition relation and on an improved iterativesquaring.In this approach a Finite State Machine is decomposedand traversed one "functioning-mode" at a time bymeans of the "disjunctive" partitioned approach.The overall algorithm aims at lowering the intermediate peakBDD size pushing further reachability analysis.Experimentson a few industrial circuits containing counters and on somelarge benchmarks show the feasibility of the approach. Gianpiero Cabodi, Paolo Camurati, Luciano Lavagno, Stefano Quer |
DAC | 3 |
| 1997 | Fast Hardware/Software Co-Simulation for Virtual Prototyping and Trade-Off AnalysisabstractHardware/Software co-simulation is generally performed with separate simulation models. This makes trade-off evaluation difficult, because the models must bere-compiled whenever some architectural choice is changed. We propose a technique to simulate hardware and software that is almost cycle accurate, and uses the same model for both types of components. Only the timing information used for synchronization needs to be changed to modify the processor choice, the implementation choice, or the scheduling policy. We show how this technique can be used to decide the implementation of a real-life example, a car dashboard controller. Claudio Passerone, Luciano Lavagno, Massimiliano Chiodo, Alberto L. Sangiovanni-Vincentelli |
DAC | 2 |
| 1997 | Decomposition and technology mapping of speed-independent circuits using Boolean relationsabstractPresents a new technique for the decomposition and technology mapping of speed-independent circuits. An initial circuit implementation is obtained in the form of a netlist of complex gates, which may not be available in the design library. The proposed method iteratively performs Boolean decomposition of each such gate F into a two-input combinational or sequential gate G, which is available in the library, and two gates H/sub 1/ and H/sub 2/, which are simpler than F, while preserving the original behavior and speed-independence of the circuit. To extract functions for H/sub 1/ and H/sub 2/, the method uses Boolean relations, as opposed to the less powerful algebraic factorization approach used in previous methods. After logic decomposition, overall library matching and optimization is carried out. Logic resynthesis, performed after speed-independent signal insertion for H/sub 1/ and H/sub 2/, allows for the sharing of decomposed logic. Overall, this method is more general than existing techniques based on restricted decomposition architectures, and thereby leads to better results in technology mapping. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Enric Pastor, Alexandre Yakovlev |
ICCAD | 4 |
| 1997 | Partial scan delay fault testing of asynchronous circuitsabstractAsynchronous circuits operate correctly only under timing assumptions. Hence testing those circuits for delay faults is crucial. The paper describes a three step method to detect possible delay faults in a sequential asynchronous circuit. The delays that are to be tested must be provided by the synthesis system. By using this information a set of paths in the circuit that must be tested is identified (step 1). For these paths the circuit is made acyclic by inserting at least one scan latch in every cycle (step 2). Then test patterns are generated for these paths (step 3). These test patterns consist of setup and initialization vectors and the final test vector. We provide effective procedures to solve both the initialization and the test pattern generation problem. The latter problem is solved by reduction to a classical problem of stuck-at test pattern generation for a related combinational circuit. Finally, a heuristic is proposed to determine which state variables must become part of a scan chain, or for which input variables the positive and negative phase must be driven independently in test mode. Experimental results shows that a high level of path delay fault testability can be achieved with partial scan. Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexander Saldanha, Alexander Taubin |
ICCAD | 3 |
| 1997 | Design of embedded systems: formal models, validation, and synthesisabstractThis paper addresses the design of reactive real-time embedded systems. Such systems are often heterogeneous in implementation technologies and design styles, for example by combining hardware application-specific integrated circuits (ASICs) with embedded software. The concurrent design process for such embedded systems involves solving the specification, validation, and synthesis problems. We review the variety of approaches to these problems that have been taken. Stephen A. Edwards, Luciano Lavagno, Edward A. Lee, Alberto L. Sangiovanni-Vincentelli |
Proc. IEEE | 2 |
| 1997 | A region-based theory for state assignment in speed-independent circuitsabstractState assignment problems still need satisfactory solutions to make asynchronous circuit synthesis more practical. A well-known example of such a problem is that of complete state coding (CSC), which happens when a pair of different states in a specification has the same binary encoding. A standard way to approach state coding conflicts is to insert new state signals into the original specification in such a way that the original behavior remains intact. This paper proposes a method which improves over existing approaches by coupling generality, optimality, and efficiency. The method is based on the use of a class of "ground objects", called regions, that play the role of a bridge between state-based specifications (transition systems, TS's) and event-based specifications (signal transition graphs, STG's), We need to deal with both types of specification because designers usually prefer a timing diagram-like notation, such as STG, while optimization and cost analysis work better at the state level. A region in a transition system is a set of states that corresponds to a place in an STG (or the underlying Petri net). Regions are tightly connected with a set of properties that are to be preserved across the state encoding process, namely, 1) trace equivalence between the original and the encoded specification, and 2) implementability as a speed-independent circuit. We will build on a theoretical body of work that has shown the significance of regions for such property-preserving transformations, and describe a set of algorithms aimed at efficiently solving the encoding problem. The algorithms have been implemented in a software tool called petrify. Unlike many existing tools, petrify represents the encoded specification as an STG. This significantly improves the readability of the result (compared to a state-based description in which concurrency is represented implicitly by interleaving), and allows the designer to be more closely involved in the synthesis process. The efficiency of the method is demonstrated on a number of "difficult" examples. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexandre Yakovlev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 1996 | Formal Verification of Embedded Systems based on CFSM NetworksabstractBoth timing and functional properties are essential to characterize the correct behavior of an embedded system.Verication is in general performed either by simulation, or by bread-boarding.Given the safety requirements of such systems, a formal proof that the properties are indeed satised is highly desirable.In this paper, we present a formal veri cation methodology for embedded systems.The formal model for the behavior of the system used in POLIS is a network of Codesign Finite State Machines.This model is translated into automata, and veri ed using automatatheoretic techniques.An industrial embedded system is veri ed using the methodology.W e demonstrate that abstractions and separation of timing and functionality is crucial for the successful use of formal veri cation for this example.We also show that in POLIS abstractions and separation of timing and functionality can be done by simple syntactic modi cation of the representation of the system. Felice Balarin, Harry Hsieh, Attila Jurecska, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
DAC | 4 |
| 1996 | Methodology and Tools for State Encoding in Asynchronous Circuit SynthesisabstractThis paper proposes a state encoding method for asynchronous circuits based on the theory of regions. A region in a Transition System is a set of states that “behave uniformly” with respect to a given transition (value change of an observable signal), and is analogue to a place in a Petri net. Regions are tightly connected with a set of properties that must be preserved across the state encoding process, namely: (1) trace equivalence between the original and the encoded specification, and (2) implementability as a speed-independent circuit. We build on a theoretical body of work that has shown the significance of regions for such property-preserving transformations, and describe a set of algorithms aimed at efficiently solving the encoding problem. The algorithms have been implemented in a software tool called petrify. Unlike many existing tools, petrify represents the encoded specification as an STG, and thus allows the designer to be more closely involved in the synthesis process. The efficiency of the method is demonstrated on a number of “difficult” examples. Jordi Cortadella, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Alexandre Yakovlev |
DAC | 4 |
| 1996 | Compact and complete test set generation for multiple stuck-faultsabstractWe propose a novel procedure for testing all multiple stuck-faults in a logic circuit using two complementary algorithms. The first algorithm finds pairs of input vectors to detect the occurrence of target single stuck-faults independent of the occurrence of other faults. The second uses a sophisticated branch and bound procedure to complete the test set generation on the faults undetected by the first algorithm. The technique is complete and applies to all circuits. Experimental results presented in this paper demonstrate that compact and complete test sets can be quickly generated for standard benchmark circuits. Alok Agrawal, Alexander Saldanha, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 3 |
| 1996 | Enhancing FSM Traversal by Temporary Re-EncodingabstractSynthesis and optimization of large finite-state machines has improved dramatically over the last few years with the introduction and rapid improvement of symbolic-state manipulation techniques. The algorithms efficiently visit each reachable state in the machine while computing and storing information about these states. We propose a new technique for improving the efficacy of traversal algorithms: re-encoding the states of the machine to more efficiently represent state sets or state transitions, or to more efficiently compute the next set of states. Our technique can be embedded in existing traversal algorithms. Experiments reveal that re-encoding can indeed reduce the time and/or space required for traversal. Gianpiero Cabodi, Luciano Lavagno, Enrico Macii, Massimo Poncino, Stefano Quer, Paolo Camurati, Ellen Sentovich |
ICCD | 2 |
| 1996 | Rapid-Prototyping of Embedded Systems via Reprogrammable DevicesabstractThis paper describes a flexible board-level rapid-prototyping environment for embedded control applications. The environment is based on an APTIX board populated by Xilinx FPGA devices, a 68hcll emulator, and APTIX programmable interconnect devices. Given a design consisting of logic and of software running on a micro-controller that implement a set of tasks, the prototype is obtained by programming the FPGA devices, the micro-controller emulator and the APTIX devices. This environment being based on programmable devices offers the flexibility to perform engineering changes, the performance needed to validate complex systems and the hardware set up for field tests. The key point in our approach is the use of results of our previous research on software and hardware synthesis as well as on some commercial tools to provide the designer with fast programming data from a high level description of the algorithms to be implemented. We demonstrate the effectiveness of the approach by showing a close-to real-life example from the automotive world. Stefano Cardelli, Massimiliano Chiodo, Paolo Giusto, Attila Jurecska, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
RSP | 5 |
| 1996 | On the Models for Asynchronous Circuit Behaviour with OR Causality
Alexandre Yakovlev, Michael Kishinevsky, Alex Kondratyev, Luciano Lavagno, Marta Pietkiewicz-Koutny |
Formal Methods Syst. Des. | 4 |
| 1996 | A Unified Signal Transition Graph Model for Asynchronous Control Circuit Synthesis
Alexandre Yakovlev, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
Formal Methods Syst. Des. | 2 |
| 1995 | Synthesis of Software Programs for Embedded Control ApplicationsabstractArticle Free Access Share on Synthesis of software programs for embedded control application Authors: Massimiliano Chiodo Magneti Marelli, Italy Magneti Marelli, ItalyView Profile , Paolo Guisto Magneti Marelli, Italy Magneti Marelli, ItalyView Profile , Attila Jurecska Magneti Marelli, Italy Magneti Marelli, ItalyView Profile , Luciano Lavagno Dipartimento di Elettronica, Politecnico di Torino, Italy Dipartimento di Elettronica, Politecnico di Torino, ItalyView Profile , Ellen Sentovich Cadence Berkeley Labs, Berkeley, CA Cadence Berkeley Labs, Berkeley, CAView Profile , Harry Hsieh Department of EECS, Univ. of California, Berkeley, CA Department of EECS, Univ. of California, Berkeley, CAView Profile , Kei Suzuki Department of EECS, Univ. of California, Berkeley, CA Department of EECS, Univ. of California, Berkeley, CAView Profile , Alberto Sangiovanni-Vincentelli Department of EECS, Univ. of California, Berkeley, CA Department of EECS, Univ. of California, Berkeley, CAView Profile Authors Info & Claims DAC '95: Proceedings of the 32nd annual ACM/IEEE Design Automation ConferenceJanuary 1995 Pages 587–592https://doi.org/10.1145/217474.217594Published:01 January 1995Publication History 36citation320DownloadsMetricsTotal Citations36Total Downloads320Last 12 Months5Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Massimiliano Chiodo, Paolo Giusto, Attila Jurecska, Luciano Lavagno, Harry Hsieh, Kei Suzuki, Alberto L. Sangiovanni-Vincentelli, Ellen Sentovich |
DAC | 4 |
| 1995 | Timed Shannon Circuits: A Power-Efficient Design Style and Synthesis ToolabstractArticle Timed shared circuits: a power-efficient design style and synthesis tool Share on Authors: Luciano Lavagno Cadence Berkeley Laboratories, 1919 Addison Street, Suite 301, Berkeley-CA Cadence Berkeley Laboratories, 1919 Addison Street, Suite 301, Berkeley-CAView Profile , Patrick C. McGeer Cadence Berkeley Laboratories, 1919 Addison Street, Suite 301, Berkeley-CA Cadence Berkeley Laboratories, 1919 Addison Street, Suite 301, Berkeley-CAView Profile , Alexander Saldanha Cadence Berkeley Laboratories, 1919 Addison Street, Suite 301, Berkeley-CA Cadence Berkeley Laboratories, 1919 Addison Street, Suite 301, Berkeley-CAView Profile , Alberto L. Sangiovanni-Vincentelli Cadence Berkeley Laboratories, 1919 Addison Street, Suite 301, Berkeley-CA Cadence Berkeley Laboratories, 1919 Addison Street, Suite 301, Berkeley-CAView Profile Authors Info & Claims DAC '95: Proceedings of the 32nd annual ACM/IEEE Design Automation ConferenceJanuary 1995 Pages 254–260https://doi.org/10.1145/217474.217538Online:01 January 1995Publication History 21citation231DownloadsMetricsTotal Citations21Total Downloads231Last 12 Months1Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Luciano Lavagno, Patrick C. McGeer, Alexander Saldanha, Alberto L. Sangiovanni-Vincentelli |
DAC | 1 |
| 1995 | Synthesizing Petri nets from state-based modelsabstractThis paper presents a method to synthesize labeled Petri nets from state-based models. Although state-based models (such as finite state machines) are a powerful formalism to describe the behavior of sequential systems, they cannot explicitly express the notions of concurrency, causality and conflict Petri nets can naturally capture these notions. The proposed method in based on deriving an elementary transition system (ETS) from a specification model. Previous work has shown that for any ETS there exists a Petri net with minimum transition count (one transition for each label) with a reachability graph isomorphic to the original ETS. This paper presents the first known approach to obtain an ETS from a non-elementary TS and derive a place-irredundant Petri net. Furthermore, by imposing constraints on the synthesis method, different classes of Petri nets can be derived from the same reachability graph (pure, free choice, unique choice). This method has been implemented and efficiently applied in different frameworks: Petri net composition, synthesis of Petri nets from asynchronous circuits, and resynthesis of Petri nets. Jordi Cortadella, Michael Kishinevsky, Luciano Lavagno, Alexandre Yakovlev |
ICCAD | 3 |
| 1995 | Synthesis for testability techniques for asynchronous circuitsabstractOur goal is to synthesize hazard-free asynchronous circuits that are testable in the very stringent hazard-free robust path-delay-fault model. From a synthesis perspective producing circuits satisfying two very stringent requirements, namely, hazard-free operation and hazard-free robust path-delay-fault-testability, poses an especially exciting challenge. Here we present techniques which guarantee both hazard-free operation and hazard-free robust path-delay-fault testability, at the expense of possibly adding test inputs. We also give a set of heuristics which can improve hazard-free robust path-delay-fault testability without requiring such inputs. Finally, we present a procedure that guarantees testability in the less stringent robust gate-delay-fault model. Kurt Keutzer, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | Synthesis of hazard-free asynchronous circuits with bounded wire delaysabstractThis paper introduces a new synthesis methodology for asynchronous sequential control circuits from a high level specification, the signal transition graph (STG). The methodology is guaranteed to generate hazard-free circuits with the bounded wire-delay model, if the STG is live and has the complete state coding property. The methodology exploits knowledge of the environmental delays, speed-independence with respect to externally visible signals, and logic synthesis techniques. A proof that STG persistency is neither necessary nor sufficient for hazard-free implementation is given.> Luciano Lavagno, Kurt Keutzer, Alberto L. Sangiovanni-Vincentelli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1995 | An efficient heuristic procedure for solving the state assignment problem for event-based specificationsabstractWe propose a novel framework to solve the state assignment problem arising from the signal transition graph (STG) representation of an asynchronous circuit. We first establish a relation between STG's and finite state machines (FSM's). Then we solve the STG state assignment problem by minimizing the number of states in the corresponding FSM and by using a critical race-free state assignment technique. State signal transitions may be added to the original STG. A lower bound on the number of signals necessary to implement the STG is given. Our technique significantly increases the STG applicability as a specification for asynchronous circuits.> Luciano Lavagno, Cho W. Moon, Robert K. Brayton, Alberto L. Sangiovanni-Vincentelli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1994 | PAPRICA-3: A Real-Time Morhphological Image ProcessorabstractThe paper describes the architecture of an image coprocessor based on a linear array of elementary processors. The system has been designed for real-time vision applications and embeds a direct I/O interface to a camera device. The instruction set of the elementary processors is based on morphological match operators on a reduced 5/spl times/5 neighborhood. Two additional communication mechanisms, one fixed and one programmable, provide both global and multiple communication channels among processing elements which are not directly connected.> Alberto Broggi, Gianni Conte, G. Burzio, Luciano Lavagno, Francesco Gregoretti, Claudio Sansoè, Leonardo Maria Reyneri |
ICIP (3) | 4 |
| 1994 | A low latency asynchronous arbitration circuitabstractWe present an asynchronous circuit for an arbiter cell that can be used to construct cascaded multiway arbitration circuits. The circuit is completely speed-independent. It has a short response delay at the input request-grant handshake link due to both a) the propagation of requests in parallel with starting arbitration and b) the concurrent resetting of request-grant handshakes in different cascades of a request-grant propagation chain.> Alexandre Yakovlev, A. Petrov, Luciano Lavagno |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1992 | Solving the State Assignment Problem for Signal Transition Graphs
Luciano Lavagno, Cho W. Moon, Robert K. Brayton, Alberto L. Sangiovanni-Vincentelli |
DAC | 1 |
| 1992 | A unified signal transition graph model for asynchronous control circuit synthesisabstractBoth low-level (analysis-oriented) and high-level (specification-oriented) models for asynchronous circuits and the environment where they operate, together with strong equivalence results between the properties at the low levels, are described. One interesting side result is the precise characterization of classical static and dynamic hazards in terms of the model. Consequently the designer can check the specification and directly decide if the behavior of any implementation will depend, e.g., on the delays of the signals described by such specification.> Alexandre Yakovlev, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 2 |
| 1992 | Linear Programming for Optimum Hazard Elimination in Asynchronous CircuitsabstractIt is shown that hazards can be optimally eliminated from circuits synthesized starting with a signal transition graph (STG) specification. The proposed approach is based on a linear programming (or integer linear programming) formulation, and as such it can be solved efficiently and optimally for a variety of cost functions. Suggested cost functions optimize either the total padded delay, an estimate of the increase in area, or the maximum cycle time of the complete system. It is also shown that delay padding on all fanouts of STG signals is a necessary and sufficient condition for hazard elimination if the structure and delay of each combinational logic block cannot be changed. Experimental results indicate that the improvements obtained are well worth the added complexity of linear program solution.> Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ICCD | 1 |
| 1992 | Symbolic minimization of multilevel logic and the input encoding problemabstractTechniques for the optimization of multilevel logic with multiple-valued input variable is presented. The motivation for this is to tackle the input encoding problem in logic synthesis, where binary codes must be found for the different values that a symbolic input variable can take. It is shown how the other multilevel optimization techniques are easily extended with multiple-valued variables. These ideas have been implemented as algorithms in the program MIS-MV. The practical issues involved in the implementation of these ideas are discussed, and results of using MIS-MV for input encoding on benchmark examples presented.> Sharad Malik, Luciano Lavagno, Robert K. Brayton, Alberto L. Sangiovanni-Vincentelli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1991 | Algorithms for Synthesis of Hazard-Free Asynchronous CircuitsabstractA technique for the synthesis of asynchronous sequential circuits from a Signal Transition Graph (STG) specification is described. We give algorithms for synthesis and hazard removal, able to produce hazard-free circuits with the bounded wire-delay model, requiring the STG to be live, safe and to have the unique state coding property. A proof that, contrary to previous beliefs, STG persistency is not necessary for hazard-free implementation is given. 1 Introduction Asynchronous design is important in several applications of digital design. "Real world" interfaces and low power systems, where "lazy evaluation" style designs may extend the average life of a battery, are two examples. In addition, clock skew problems limit the performance and the flexibility of large scale synchronous systems. On the other hand asynchronous design is harder and more constrained than synchronous design, due to the hazard problem: asynchronous circuits are by definition sensitive to all signal changes, whe... Luciano Lavagno, Kurt Keutzer, Alberto L. Sangiovanni-Vincentelli |
DAC | 1 |
| 1991 | Synthesis for Testability Techniques for Asynchronous CircuitsabstractThe authors present techniques which guarantee both hazard-free operation and hazard-free robust path-delay-fault testability at the expense of possibly adding test inputs. They also give a set of heuristics which can improve hazard-free robust path-delay-fault testability without requiring such inputs. Finally, they demonstrate the effectiveness of these techniques on a set of asynchronous interface circuits gathered from industry and academia.> Kurt Keutzer, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 2 |
| 1990 | MIS-MV: Optimization of Multi-Level Logic with Multiple-Valued InputsabstractTechniques are presented for the optimization of multi-level logic with multiple-valued input variables. The motivation for this is to tackle the input encoding problem in logic synthesis, where binary codes need to be found for the different values of a symbolic input variable. Multi-level multiple-valued optimization is used to generate constraints that are used to determine the codes. The state assignment problem in sequential logic synthesis can be approximated as an input encoding problem by ignoring the next state field, which is reasonable when the primary output logic, dominates the next state logic. A novel technique is presented for extracting common factors with multiple-valued variables, and it is shown how other multi-level optimization techniques are easily extended with multiple-valued variables. These ideas have been implemented as algorithms in the MIS-MV program. Practical issues are also presented regarding implementation. Experimental results are also given.> Luciano Lavagno, Sharad Malik, Robert K. Brayton, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 1 |