VLDB 2026 Research / reviewers in the wild / expert
Mario R. Casu
dblp:43/383 · also Mario Roberto Casu
· DBLP profile ↗
40ranked-venue papers
17as first author
9since 2021 · last 2026
0000-0002-1026-0178ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 16 first-author · 8 since 2021Software engineering, systems software and programming languages · 9 · 3 first-author · 2 since 2021Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design for testability using mixed-polarity flip-flops and latchesabstractSequential circuits employing a combination of mixed-polarity flip-flops and latches allow significant improvements in clock frequency compared to useful skew and retiming. However, no work addresses the task of enabling a scan- based test on a circuit optimized with such techniques, while simultaneously minimizing the area overhead due to shadow latches used to complete the scan chain when latches are used in the design. This poses a serious limitation to the industrial application of mixed FF and latch-based techniques, since post-fabrication tests are an unavoidable step in IC production. This paper presents a macro-cell structure to enable both the exploitation of time borrowing for frequency optimization and the execution of the scan test of a design. The proposed solution requires minimal changes in the test setup and is evaluated using a recent methodology, Mix&Latch. Moreover, the work proposes modifications to Mix&Latch that allow reusing the standard cells introduced for the scan test to solve hold timing violations, avoiding additional hardware overhead. Results show that the lumped cell structure does not significantly impact frequency gains, and the ILP formulation of latch and FF type optimization can be extended to cover the DFT optimization part, ensuring only a moderate increase in area and power consumption, comparable with the DFT impact on regular FF-based designs. Lorenzo Lagostina, Jordi Cortadella, Mario R. Casu, Luciano Lavagno |
DATE | 3 |
| 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer ProgrammingabstractSkip connections have emerged as a key component of modern convolutional neural networks (CNNs) for computer vision tasks, allowing for the creation of more accurate and deeper models by addressing the vanishing gradient problem. However, the existing implementations of field-programmable gate array (FPGA)-based accelerators for ResNets and MobileNetV2 often experience decreased performance and increased computational latency due to the implementation of skip blocks. This article presents a novel framework for developing deep learning models on FPGAs that focuses on skip connections, with a unique approach to reduce buffering overhead. This results in a more efficient utilization of resources in the implementation of the skip layer. The nn2fpga compiler follows a thorough set of high-level synthesis (HLS) design principles and optimization strategies, exploiting in novel ways standard techniques to effectively map skip connection-based networks into static dataflow accelerators. To maximize throughput and efficiently use the available resources, our compiler employs a fast and effective design space exploration method based on a binary integer programming model which accurately assigns FPGA resources to the network layers, to maximize global throughput under resource constraints and then minimize resources for the achieved maximum throughput. Experimental results on the CIFAR-10 and ImageNet datasets demonstrate substantial gains in throughput ($\mathbf {3\times }$–$\mathbf {7\times }$on the past HLS-based work) for ResNet8, ResNet20, and MobileNetV2 models deployed on various Xilinx FPGA boards. Notably, MobileNetV2 deployed on the ZCU102 achieves a throughput of 2115 frame per second, representing even a 10% speedup over a state-of-the-art highly optimized manual register-transfer level implementation, showing that HLS can actually improve over manual design, thanks to the faster exploration of the design space. Roberto Bosio, Filippo Minnella, Teodoro Urso, Mario R. Casu, Luciano Lavagno, Mihai T. Lazarescu, Paolo Pasini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | STAR: Sum-Together/Apart Reconfigurable Multipliers for Precision-Scalable ML WorkloadsabstractTo achieve an optimal balance between accuracy and latency in Deep Neural Networks (DNNs), precision-scalability has become a paramount feature for hardware specialized for Machine Learning (ML) workloads. Recently, many precision-scalable (PS) multipliers and multiply-and-accumulate (MAC) units have been proposed. They are mainly divided in two categories, Sum-Apart (SA) and Sum-Together (ST), and have been always presented as alternative implementations. Instead, in this paper, we introduce for the first time a new class of PS Sum-Together/Apart Reconfigurable multipliers, which we call STAR, designed to support both SA and ST modes with a single reconfigurable architecture. STAR multipliers could be useful in MAC units of CPU or hardware accelerators, for example, enabling them to handle both 2D Convolution (in ST mode) and Depth-wise Convolution (in SA mode) with a unique PS hardware design, thus saving hardware resources. We derive four distinct STAR multiplier architectures, including two derived from the well-known Divide-and-Conquer and Sub-word Parallel SA and ST families, which support 16, 8 and 4-bit precision. We perform an extensive exploration of these architectures in terms of power, performance, and area, across a wide range of clock frequency constraints, from 0.4 to 2.0 GHz, targeting a 28-nm CMOS technology. We identify the Pareto-optimal solutions with the lowest area and power in the low-frequency, mid-frequency, and high-frequency ranges. Our findings allow designers to select the best STAR solution depending on their design target, either low-power and low-area, high performance, or balanced. Edward Manca, Luca Urbinati, Mario R. Casu |
DATE | 3 |
| 2024 | WIP: Building an Education Ecosystem for Next Generation Microelectronics Experts in Green and Circular Economy with Digitally-Supported Teaching Methods for Sustainable Chips and Applications (EU Project GreenChips-EDU)abstractThis work in progress innovative practice paper intends to report on the outline and the ongoing progress of the EU-project GreenChips-EDU, which has been started in October 2023, and intends to fundamentally redesign educational microelectronics programs especially but not limited to students and professionals. One of the major goals is the design of a new microelectronics master program to which six European universities are contributing. The contents of this program will be substantially enhanced with green electronics contents innovative teaching methods. Other work will be done in the field of a new MBA program, self-standing modules for professionals, and a new microelectronics bachelor designed by one university of applied sciences. Klaus Hofmann, Ferdinand Keil, David Riehl, Alicja Malgorzata Michalowska-Forsyth, Nikolaus Czepl, Sarah Woywod, Dominik Zupan, Mario R. Casu, Carlo Ricciardi, Massimo Violante, Mariagrazia Graziano, Yuri Ardesi, Fabrizio Mo, Dominik Berger, Sabine Sill, Volker Visotschnig, Panagiota Morfouli, Liliana Prejbeanu, Katell Morin-Allory, Cyrille Chavet, Davide Bucci, Skandar Basrour, Jean-Christophe Crebier, Nhu-Huan Nguyen, Ernesto Quisbert-Trujillo, Christian Defélix, Isabelle Corbett-Etchevers, Johannes Sturm, Jens Peter Konrath, Ulla Birnbacher, Thomas Klinger, Wolfgang Werth, Jorge Fernandes, Marcelino B. Santos, Antonio Rubio 0001, Alba Pagès-Zamora, Jordi Salazar, Beatriz Otero, J. Manuel Moreno, X. Aragones, Israel Martin, Aleix Sole, Dunja Suttnig, Julia Calabro, Floriberto Lima, Eric Jouseau, François Cerisier, Cristian Rivier, Sepp Eisenriegler, Harald Reichl, Miroslav Macan, Dubravko Kruselj, Mladen Puskaric, Mirjana Tatalovic, Vinko Zelenicic, Bernd Deutschmann |
FIE | 8 |
| 2024 | Mix & Latch: Comparison With State-of-the-Art Retiming on a RISC-V BenchmarkabstractFlip-flops (FFs) are the most commonly used sequential elements in synchronous circuits, but their timing requirements limit the operating frequency. Borrowing time with a latch-based approach can increase operating frequency, but traditional back-end optimization tools struggle to manage hold time requirements. The Mix & Latch technique achieves higher frequencies and often lower area than commercial state-of-the-art retiming by exploiting four types of synchronous sequential gates, namely, positive and negative edge-triggered flip-flops (FFs) and positive and negative transparent latches, all using a single clock tree.In this article, we first significantly accelerate the Mix & Latch flow convergence with respect to past work, by using a post-synthesis-based timing analysis that eliminates the first placement and routing needed for post-layout timing analysis. Then, by adding tolerance margins to the timing model, the pessimism is reduced to improve both convergence speed and maximum frequency. Finally, we reduce the complexity of the problem by applying the methodology only to the sequential elements belonging to critical paths. The effectiveness of Mix & Latch is then demonstrated on a RISC-V processor core from the Pulp platform using 28nm CMOS FDSOI technology. The results are compared to both the original Mix & Latch flow and a retiming performed with a state-of-the-art tool, showing a 25% frequency improvement over the original flow and 7.5% over the retiming flow. Compared to the retiming flow, we achieve comparable or lower power and area, while preserving the original registers and allowing logic equivalence checking. Lorenzo Lagostina, Filippo Minnella, Jordi Cortadella, Mario R. Casu, Mihai T. Lazarescu, Luciano Lavagno |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | HLS-based dataflow hardware architecture for Support Vector Machine in FPGAabstractImplementing fast and accurate Support Vector Machine (SVM) classifiers in embedded systems with limited compute and memory capacity and in applications with real-time constraints, such as continuous medical monitoring for anomaly detection, can be challenging and calls for low cost, low power and resource efficient hardware accelerators. In this paper, we propose a flexible FPGA-based SVM accelerator highly optimized through a dataflow architecture. Thanks to High Level Synthesis (HLS) and the dataflow method, our design is scalable and can be used for large data dimensions when there is limited on-chip memory. The hardware parallelism is adjustable and can be specified according to the available FPGA resources. The performance of different SVM kernels are evaluated in hardware. In addition, an efficient fixed-point implementation is proposed to improve the speed. We compared our design with recent SVM accelerators and achieved a minimum of $10\times$ speed-up compared to other HLS-based and $4.4\times$ compared to HDL-based designs. Mohammad Amir Mansoori, Mario R. Casu |
ISCAS | 2 |
| 2022 | A Reconfigurable Depth-Wise Convolution Module for Heterogeneously Quantized DNNsabstractIn Deep Neural Networks (DNN), the depth-wise separable convolution has often replaced the standard 2D convolution having much fewer parameters and operations. Another common technique to squeeze DNNs is heterogeneous quantization, which uses a different bitwidth for each layer. In this context we propose for the first time a novel Reconfigurable Depth-wise convolution Module (RDM), which uses multipliers that can be reconfigured to support 1, 2 or 4 operations at the same time at increasingly lower precision of the operands. We leveraged High Level Synthesis to produce five RDM variants with different channels parallelism to cover a wide range of DNNs. The comparisons with a non-configurable Standard Depth-wise convolution module (SDM) on a CMOS FDSOI 28-nm technology show a significant latency reduction for a given silicon area for the low-precision configurations. Luca Urbinati, Mario R. Casu |
ISCAS | 2 |
| 2022 | Fast Energy-Optimal Multikernel DNN-Like Application Allocation on Multi-FPGA PlatformsabstractPlatforms with multiple field-programmable gate arrays (FPGAs), such as Amazon Web Services (AWS) F1 instances, can efficiently accelerate multikernel pipelined applications, e.g., convolutional neural networks for machine vision tasks or transformer networks for natural language processing tasks. To reduce energy consumption when the FPGAs are underutilized, we propose a model to 1) find offline the minimum-power solution for given throughput constraints and 2) dynamically reprogram the FPGA at runtime (which is complementary to dynamic voltage and frequency scaling) to match best the workloads when they change. The offline optimization model can be solved using a mixed-integer nonlinear programming (MINLP) solver, but it can be very slow. Hence, we provide two heuristic optimization methods that improve result quality within a bounded time. We use several very large designs to demonstrate that both heuristics obtain comparable results to MINLP, when it can find the best solution, and they obtain much better results than MINLP, when it cannot find the optimum within a bounded amount of time. The heuristic methods can also be thousands of times faster than the MINLP solver. Junnan Shan, Mihai T. Lazarescu, Jordi Cortadella, Luciano Lavagno, Mario R. Casu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | CNN-on-AWS: Efficient Allocation of Multikernel Applications on Multi-FPGA PlatformsabstractMulti-FPGA platforms, like Amazon AWS F1, can run in the cloud multikernel pipelined applications, like convolutional neural networks (CNNs), with excellent performance and lower energy consumption than CPUs or GPUs. We propose a method to efficiently map these applications on multi-FPGA platforms to maximize the application throughput. Our methodology finds, for the given resources, the optimal number of parallel instances of each kernel in the pipeline and their allocation to one or more among the available FPGAs. We obtain this by formulating and solving a mixed-integer, nonlinear optimization problem, in which we model the performance of each component and the duration of the phases in which the accelerated computation can be split into, namely: 1) data transfer from a host CPU to the DDR memory of each FPGA; 2) data transfer from FPGA DDR to FPGA on-chip memory; 3) kernel computation on the FPGA; 4) data transfer from FPGA on-chip memory to FPGA DDR; and 5) data transfer from FPGA DDR to host. Finding the optimal solution using a mixed-integer nonlinear programming (MINLP) solver is often highly inefficient. Hence, we provide a fast heuristic method that according to our experiments can be much more efficient than the MINLP solver and finds comparable results. For larger problems (more CNN layers), our heuristic method can quickly find (several thousand times faster) much better solutions than the MINLP solver, even if we run the latter for a very long time. Junnan Shan, Mihai T. Lazarescu, Jordi Cortadella, Luciano Lavagno, Mario R. Casu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | A Machine-Learning Based Microwave Sensing Approach to Food Contaminant DetectionabstractTo detect contaminants accidentally included in packaged foods, food industries use an array of systems ranging from metal detectors to X-ray imagers. Low density plastic or glass contaminants, however, are not easily detected with standard methods. If the dielectric contrast between the packaged food and these contaminants in the microwave spectrum is sensible, Microwave Sensing (MWS) can be used as a contactless detection method, which is particularly useful when the food is already packaged. In this paper we propose using MWS combined with Machine Learning (ML). In particular, we report on experiments we did with packaged cocoa-hazelnut spread and show the accuracy of our approach. We also present an FPGA acceleration that runs the ML processing in real-time so as to keep up with the throughput of a production line. Luca Urbinati, Marco Ricci 0004, Giovanna Turvani, Jorge A. Tobon, Francesca Vipiana, Mario R. Casu |
ISCAS | 6 |
| 2019 | Exact and Heuristic Allocation of Multi-kernel Applications to Multi-FPGA PlatformsabstractFPGA-based accelerators demonstrated high energy efficiency compared to GPUs and CPUs. However, single FPGA designs may not achieve sufficient task parallelism. In this work, we optimize the mapping of high-performance multi-kernel applications, like Convolutional Neural Networks, to multi-FPGA platforms. First, we formulate the system level optimization problem, choosing within a huge design space the parallelism and number of compute units for each kernel in the pipeline. Then we solve it using a combination of Geometric Programming, producing the optimum performance solution given resource and DRAM bandwidth constraints, and a heuristic allocator of the compute units on the FPGA cluster. Junnan Shan, Mario R. Casu, Jordi Cortadella, Luciano Lavagno, Mihai T. Lazarescu |
DAC | 2 |
| 2018 | Energy-performance design exploration of a low-power microprogrammed deep-learning acceleratorabstractThis paper presents the design space exploration of a novel microprogrammable accelerator in which PEs are connected with a Network-on-Chip and benefit from low-power features enabled through a practical implementation of a Dual-Vddassignment scheme. An analytical model, fitted with postlayout data obtained with a 28nm FDSOI design kit, returns implementations with optimal energy-performance tradeoff by taking into consideration all the key design-space variables. The obtained Pareto analysis helps us infer optimization rules aimed at improving quality of design. Giulia Santoro, Mario R. Casu, Valentino Peluso, Andrea Calimera, Massimo Alioto |
DATE | 2 |
| 2018 | Design-Space Exploration of Pareto-Optimal Architectures for Deep Learning with DVFSabstractSpecialized computing engines are required to accelerate the execution of Deep Learning (DL) algorithms in an energy-efficient way. To adapt the processing throughput of these accelerators to the workload requirements while saving power, Dynamic Voltage and Frequency Scaling (DVFS) seems the natural solution. However, DL workloads need to frequently access the off-chip memory, which tends to make the performance of these accelerators memory-bound rather than computation-bound, hence reducing the effectiveness of DVFS. In this work we use a performance-power analytical model fitted on a parametrized implementation of a DL accelerator in a 28-nm FDSOI technology to explore a large design space and to obtain the Pareto points that maximize the effectiveness of DVFS in the sub-space of throughput and energy efficiency. In our model we consider the impact on performance and power of the off-chip memory using real data of a commercial low-power DRAM. Giulia Santoro, Mario R. Casu, Valentino Peluso, Andrea Calimera, Massimo Alioto |
ISCAS | 2 |
| 2017 | Power-performance assessment of different DVFS control policies in NoCs
Mario R. Casu, Paolo Giaccone |
J. Parallel Distributed Comput. | 1 |
| 2017 | Accelerators for Breast Cancer DetectionabstractAlgorithms used in microwave imaging for breast cancer detection require hardware acceleration to speed up execution time and reduce power consumption. In this article, we present the hardware implementation of two accelerators for two alternative imaging algorithms that we obtain entirely from SystemC specifications via high-level synthesis. The two algorithms present opposite characteristics that stress the design process and the capabilities of commercial HLS tools in different ways: the first is communication bound and requires overlapping and pipelining of communication and computation in order to maximize the application throughput; the second is computation bound and uses complex mathematical functions that HLS tools do not directly support. Despite these difficulties, thanks to HLS, in the span of only 4 months we were able to explore a large design space and derive about 100 implementations with different cost-performance profiles, targeting both a Field-Programmable Gate Array (FPGA) platform and a 32-nm standard-cell Application Specific Integrated Circuit (ASIC) library. In addition, we could obtain results that outperform a previous Register-Transfer Level (RTL) implementation, which confirms the remarkable progress of HLS tools. Daniele Jahier Pagliari, Mario R. Casu, Luca P. Carloni |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Rate-based vs delay-based control for DVFS in NoC
Mario R. Casu, Paolo Giaccone |
DATE | 1 |
| 2015 | Acceleration of microwave imaging algorithms for breast cancer detection via High-Level SynthesisabstractWe present the system-level design of two accelerators for two microwave imaging algorithms for breast cancer detection. The accelerators were designed in SystemC and optimized via High-Level Synthesis (HLS). The two algorithms stress the capabilities of commercial HLS tools in different ways: the first is communication-bound and requires careful pipelining of communication and computation; the second is computation-bound and requires the implementation of mathematical functions that are not properly supported by HLS tools. Still, in the span of four months we were able to design and validate about one hundred alternative implementations, targeting a Zynq SoC platform. Furthermore, we were pleased to obtain results that are superior to a previous RTL implementation, which confirms the remarkable progress of HLS tools. Daniele Jahier Pagliari, Mario R. Casu, Luca P. Carloni |
ICCD | 2 |
| 2015 | A synchronous latency-insensitive RISC for better than worst-case design
Mario R. Casu, Paolo Mantovani |
Integr. | 1 |
| 2014 | Simulation and design of an UWB imaging system for breast cancer detection
Xiaolu Guo, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni |
Integr. | 2 |
| 2014 | UWB microwave imaging for breast cancer detection: Many-core, GPU, or FPGA?abstractAn UWB microwave imaging system for breast cancer detection consists of antennas, transceivers, and a high-performance embedded system for elaborating the received signals and reconstructing breast images. In this article we focus on this embedded system. To accelerate the image reconstruction, the Beamforming phase has to be implemented in a parallel fashion. We assess its implementation in three currently available high-end platforms based on a multicore CPU, a GPU, and an FPGA, respectively. We then project the results applying technology scaling rules to future many-core CPUs, many-thread GPUs, and advanced FPGAs. We consider an optimistic case in which available resources increase according to Moore's law only, and a pessimistic case in which only a fraction of those resources are available due to a limited power budget. In both scenarios, an implementation that includes a high-end FPGA outperforms the other alternatives. Since the number of effectively usable cores in future many-cores will be power-limited, and there is a trend toward the integration of power-efficient accelerators, we conjecture that a chip consisting of a many-core section and a reconfigurable logic section will be the perfect platform for this application. Mario R. Casu, Francesco Colonna, Marco Crepaldi, Danilo Demarchi, Mariagrazia Graziano, Maurizio Zamboni |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2012 | Exploiting space diversity and Dynamic Voltage Frequency Scaling in multiplane Network-on-ChipsabstractNetwork-on-Chips (NoCs) have been proposed as a scalable solution to interconnect multiple components on a silicon chip. In this paper, we approach NoCs power optimization through Dynamic Voltage and Frequency Scaling (DVFS) under the hypothesis that two NoC planes are available, each with a different voltage supply and clock frequency. We show the high potential benefit of applying DVFS independently in each plane. We propose three strategies that allocate the traffic in the two planes to minimize power consumption. We evaluate them through a comparison with an ideal traffic allocation policy based on a linear programming technique. We show that load balancing in the two planes is not always the best policy. Indeed, in an unbalanced traffic scenario, concentrating the high-load flows in one plane and the remaining low-load flows in the other plane, is more power efficient. Andrea Bianco, Paolo Giaccone, Mario R. Casu, Nanfang Li |
GLOBECOM | 3 |
| 2011 | Coupling latency-insensitivity with variable-latency for better than worst case design: a RISC case studyabstractThe gap between worst and typical case delays is bound to increase in nanometer scale technologies due to the spread in process manufacturing parameters. To still profit from scaling, designs should tolerate worst case delays seamlessly and with a minimum performance degradation with respect to the typical case. We present a simple RISC core which tolerates worst case extra latency using the Latency-Insensitive Design approach coupled to a Variable-Latency mechanism. Stalls caused by excessive delay, by data and control hazards and by late memory access are dealt with in a uniform way. Compared to a pure worst-case approach, our design method permits to increase the core clock frequency by 23% in a 45 nm CMOS technology, without area and power penalty. Mario R. Casu, Stefano Colazzo, Paolo Mantovani |
ACM Great Lakes Symposium on VLSI | 1 |
| 2010 | A flexible UWB Transmitter for breast cancer detection imaging systemsabstractThis paper presents a flexible architecture for an integrated Ultra-Wideband (UWB) Transmitter capable of generating pulses suited for breast cancer detection imaging systems. A flexible design allows the generation of a large variety of UWB signals fully compatible with the ones used in real experiments in recent state-of-the-art. Flexibility and high degree of programmability of the mixed-signal system allow also to compensate for non-ideal effects of building blocks through a digital calibration. It is also shown how not only internal non-idealities are accounted for but also how channel and antenna responses can be compensated for through a digital pre-emphasis of UWB pulses. The circuit is designed on a 130 nm CMOS technology and simulated at transistor-level. Simulations showed 2% maximum NRMSE pulse error with respect to ideal Gaussian and Modulated and Modified Hermite Polynomial (MMHP) Matlab templates. Massimo Cutrupi, Marco Crepaldi, Mario R. Casu, Mariagrazia Graziano |
DATE | 3 |
| 2010 | MEDEA: a hybrid shared-memory/message-passing multiprocessor NoC-based architectureabstractThe shared-memory model has been adopted, both for data exchange as well as synchronization using semaphores in almost every on-chip multiprocessor implementation, ranging from general purpose chip multiprocessors (CMPs) to domain specific multi-core graphics processing units (GPUs). Low-latency synchronization is desirable but is hard to achieve in practice due to the memory hierarchy. On the contrary, an explicit exchange of synchronization tokens among the processing elements through dedicated on-chip links would be beneficial for the overall system performance. In this paper we propose the Medea NoC-based framework, a hybrid shared-memory/message-passing approach. Medea has been modeled with a fast, cycle-accurate SystemC implementation enabling a fast system exploration varying several parameters like number and types of cores, cache size and policy and NoC features. In addition, every SystemC block has its RTL counterpart for physical implementation on FPGAs and ASICs. A parallel version of the Jacobi algorithm has been used as a test application to validate the methodology. Results confirm expectations about performance and effectiveness of system exploration and design. Sergio Tota, Mario R. Casu, Massimo Ruo Roch, Luca Rostagno, Maurizio Zamboni |
DATE | 2 |
| 2009 | A mixed-signal demodulator for a low-complexity IR-UWB receiver: Methodology, simulation and design
Marco Crepaldi, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni |
Integr. | 2 |
| 2009 | A Case Study for NoC-Based Homogeneous MPSoC ArchitecturesabstractThe many-core design paradigm requires flexible and modular hardware and software components to provide the required scalability to next-generation on-chip multiprocessor architectures. A multidisciplinary approach is necessary to consider all the interactions between the different components of the design. In this paper, a complete design methodology that tackles at once the aspects of system level modeling, hardware architecture, and programming model has been successfully used for the implementation of a multiprocessor network-on-chip (NoC)-based system, the NoCRay graphic accelerator. The design, based on 16 processors, after prototyping with field-programmable gate array (FPGA), has been laid out in 90-nm technology. Post-layout results show very low power, area, as well as 500 MHz of clock frequency. Results show that an array of small and simple processors outperform a single high-end general purpose processor. Sergio Tota, Mario R. Casu, Massimo Ruo Roch, Luca Macchiarulo, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | An effective AMS top-down methodology applied to the design of a mixed-signal UWB system-on-chip
Marco Crepaldi, Mario R. Casu, Mariagrazia Graziano, Maurizio Zamboni |
DATE | 2 |
| 2006 | Implementation analysis of NoC: a MPSoC trace-driven approachabstractThis paper proposes to tackle Networks-on-chip design for MPSoC on the assumption that area and power overhead control is the primary goal. Analysis of topologies and routing strategies is performed by comparing two approaches, the wormhole and hot potato, both theoretically and using real synthesized data in 0.13μm technology. It is shown that the hot potato solution is competitive and possibly better for both occupation and dissipation, while its performance, measured by simulation on real multiprocessor traces, is not worse than the wormhole case. Sergio Tota, Mario R. Casu, Luca Macchiarulo |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | Floorplanning With Wire Pipelining in Adaptive Communication ChannelsabstractThe recent shift toward wire pipelining (WP) mandated by technological factors has attracted attention toward latency-controlled floorplanning. However, no systematic study has been published so far that takes into account block and logic-delay limitations. This paper aims at filling the gap by showing that block delay can limit and possibly prevent any real gain WP might promise. In this paper, the authors also show how a modified adaptive WP scheme, on the other hand, allows relevant gains. They built a SoC floorplanner based on the use of adaptive and nonadaptive WP, which optimizes the data rate, taking block delay into account. The results of new and old WP techniques applied on benchmarks and on an MPEG decoder are compared to the optimal results obtained when no WP is employed Mario R. Casu, Luca Macchiarulo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | A New System Design Methodology for Wire Pipelined SoCabstractWire pipelining (WP) has been proposed in order to limit the impact of increasing wire delays. In general, added pipeline elements alter the system such that architectural changes are needed to preserve functionality. We illustrate a proposal that, while allowing the use of IP blocks without modification, takes advantage of a minimal knowledge of the IP's communication profile to increase performance dramatically. We show the formal equivalence between the IP and the original system and prove the higher performance achievable through a relevant case study. Mario R. Casu, Luca Macchiarulo |
DATE | 1 |
| 2005 | Floorplan assisted data rate enhancement through wire pipelining: a real assessmentabstractThe recent shift towards wire pipelining (WP) mandated by technological factors has attracted attention towards latency-controlled floorplanning. However, no systematic study has been published so far that takes into account block and logic delay limitations. The present workaims at filling the gap by showing that blockdelay can limit and possibly prevent any real gain WP might promise. Recurring to adaptive WP schemes, on the other hand, allows relevant gains. We built floorplanner that optimizes for maximum data rate, taking into account various models of block delay, and compares them to the optimal results obtained when no wire pipelining is employed. Experiments with suitable floorplanning benchmarks and case studies are performed to substantiate theoretical intuitions. Mario R. Casu, Luca Macchiarulo |
ISPD | 1 |
| 2005 | Throughput-driven floorplanning with wire pipeliningabstractThe size of future high-performance SoC is such that the time-of-flight of wires connecting distant pins in the layout can be much higher than the clock period. In order to keep the frequency as high as possible, the wires may be pipelined. However, the insertion of flip-flops may alter the throughput of the system due to the presence of loops in the logic netlist. In this paper, we address the problem of floorplanning a large design where long interconnects are pipelined by inserting the throughput in the cost function of a tool based on simulated annealing. The results obtained on a series of benchmarks are then validated using a simple router that breaks long interconnects by suitably placing flip-flops along the wires. Mario R. Casu, Luca Macchiarulo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Implementation aspects of a transmitted-reference UWB receiverabstractAbstract In this paper, we discuss the design issues of an ultra wide band (UWB) receiver targeting a single‐chip CMOS implementation for low data‐rate applications like ad hoc wireless sensor networks. A non‐coherent transmitted‐reference (TR) receiver is chosen because of its small complexity compared to other architectures. After a brief recapitulation of the UWB fundamentals and a short discussion on the major differences between coherent and non‐coherent receivers, we discuss issues, challenges and possible design solutions. Several simulation results obtained by means of a behavioral model are presented, together with an analysis of the trade‐off between performance and complexity in an integrated circuit implementation. Copyright © 2005 John Wiley & Sons, Ltd. Mario R. Casu, Giuseppe Durisi |
Wirel. Commun. Mob. Comput. | 1 |
| 2004 | A new approach to latency insensitive designabstractLatency Insensitive Protocols have been proposed as a viable mean to speed up large Systems-on-Chip where the limit in clock frequency is given by long global wires connecting together functional blocks. In this paper we keep the philosophy of Latency Insensitive Design and show that a drastic simplification can be done that results in even no need to implement any kind of protocol. By using a scheduling algorithm for the functional blocks activation we greatly reduce the routing resources demand of the old protocol, the area occupied by the sequential elements used to pipeline long interconnects and the complexity of the gating structure used to activate the modules. Mario R. Casu, Luca Macchiarulo |
DAC | 1 |
| 2004 | Issues in Implementing Latency Insensitive ProtocolsabstractA design that works under the assumption of zero-delay connections between functional modules is modified in a latency insensitive design (LID) by encapsulating them within the wrappers ("shells") and connecting them through internally pipelined blocks ("relay stations") complying with a protocol that guarantees identity of behavior. Mario R. Casu, Luca Macchiarulo |
DATE | 1 |
| 2004 | On-Chip Transparent Wire PipeliningabstractWire pipelining has been proposed as a viable mean to break the discrepancy between decreasing gate delays and increasing wire delays in deep-submicron technologies. Far from being a straightforwardly applicable technique, this methodology requires a number of design modifications in order to insert it seamlessly in the current design flow. In this paper, we briefly survey the methods presented by other researchers in the field and then we thoroughly analyze the solutions we recently proposed, ranging from system-level wire pipelining to physical design aspects. Mario R. Casu, Luca Macchiarulo |
ICCD | 1 |
| 2004 | Floorplanning for throughputabstractLarge Systems-on-Chip (SoC) in advanced technologies run at such high frequencies that the time-of-flight of signals connecting two distant pins in the layout can be higher than the clock period. In order to avoid performance penalties wires are pipelined using latches. However the throughput of the system may be altered due to the presence of loops in the logic netlist. In this paper we address the problem of floorplanning a large design with interconnect pipelining and inserting throughput in the cost function of the floorplanning algorithm. The throughput results obtained on a series of benchmarks are then validated using a simple router that places flipflops along the nets built with an heuristical minimum rectilinear steiner tree. Mario R. Casu, Luca Macchiarulo |
ISPD | 1 |
| 2004 | An electromigration and thermal model of power wires for a priori high-level reliability predictionabstractIn this paper, a simple power-distribution electrothermal model including the interconnect self-heating is used together with a statistical model of average and rms currents of functional blocks and a high-level model of fanout distribution and interconnect wirelength. Following the 2001 SIA roadmap projections, we are able to predict a priori that the minimum width that satisfies the electromigration constraints does not scale like the minimum metal pitch in future technology nodes. As a consequence, the percentage of chip area covered by power lines is expected to increase at the expense of wiring resources unless proper countermeasures are taken. Some possible solutions are proposed in the paper. Mario R. Casu, Mariagrazia Graziano, Guido Masera, Gianluca Piccinini, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Synthesis of low-leakage PD-SOI circuits with body-biasingabstractArticle Share on Synthesis of low-leakage PD-SOI circuits with body-biasing Authors: Mario Casu Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Gianluca Piccinini Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Guido Masera Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Maurizio Zamboni Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 287–290https://doi.org/10.1145/383082.383170Online:06 August 2001Publication History 2citation166DownloadsMetricsTotal Citations2Total Downloads166Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Mario R. Casu, Gianluca Piccinini, Guido Masera, Maurizio Zamboni |
ISLPED | 1 |
| 2000 | A high accuracy-low complexity model for CMOS delaysabstractThis paper presents a new model for CMOS structures delays estimation based on a deep analysis of complex gates behavior. This approach can supply a high level of accuracy. A complex structure is reduced first to series-connected MOS, then the delay equations are applied to that reduced rate. The model is based on a time piecewise linearization so that a strongly nonlinear circuit can he solved using well known linear techniques. The delay formulas involve model parameters as MOS width functions, therefore providing routines suitable for optimization algorithms. The high level of accuracy, the low CPU time and the high degree of scaling capability are proved in the paper. These features make the model attractive for deep submicron technologies. Mario R. Casu, Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni |
ISCAS | 1 |