Matheus T. Moreira

dblp:23/3940 · also Matheus Trevisan, Matheus Trevisan Moreira · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0001-5030-9215ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 6 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 LLMs: A Driving Force in Next Generation Digital Design Automation
abstract
Innovations in generative artificial intelligence (GenAI), particularly large language models (LLMs), are poised to revolutionize silicon design automation. This paper explores the transformative potential of LLMs in automating and enhancing various tasks within the silicon design process. It reviews the current applications of LLMs and their potential to automate silicon design tasks, proposing applications, providing a qualitative analysis of the readiness of the technology to support these applications and setting directions for future research.
Matheus T. Moreira, Aram H. Markosyan, Chris Cummins, Warren Hunt, Gabriel Synnaeve, Edith Beigné
DAC1
2025 A 3D Design Methodology for Integrated Wearable SoCs: Enabling Energy Efficiency and Enhanced Performance at Iso-Area Footprint
abstract
Augmented Reality (AR) System-on-Chips (SoCs) have strict power budgets and form-factor limitations for wearable, all-day use AR glasses running high-performance applications. Limited compute and memory resources that can fit within the strict industrial design area footprint of an AR SoC, however, create performance bottlenecks for demanding workloads such as Pixel Codec Avatars (PiCA) group-calling which connects multiple users with their photorealistic representations. To alleviate this unique wearables challenge, 3D integration with hybrid-bonding technology offers energy-efficient 3D stacking of more silicon resources within the same SoC footprint. Implementing such 3D architectures, however, is another challenge as current EDA tools and flows offer limited 3D design control. In this work, we present a 3D design methodology for robust 3D clock network and datapath design using current EDA tools. To validate the proposed methodology, we implemented a 3D integrated prototype AR SoC housing a 3D-stacked Machine Learning (ML) accelerator utilizing TSMC SoIC™bonding technology. Silicon measurements demonstrate that the 3D ML accelerator enables running PiCA AR group call at 30 frames-per-second (fps) by 3D-expanding its memory resources by 4× to achieve 2× better energy-efficiency when compared to a 2D baseline accelerator at iso-footprint.
Huseyin Ekin Sumbul, Arne Symons, Lita Yang, Huichu Liu, Tony F. Wu, Matheus T. Moreira, Debabrata Mohapatra, Abhinav Agarwal, Kaushik Ravindran, Yuecheng Li, Edith Beigné
DATE6
2021 Read your Circuit: Leveraging Word Embedding to Guide Logic Optimization
abstract
To tackle the involved complexity, Electronic Design Automation (EDA) tools are broken in well-defined steps, each operating at different abstraction levels. Higher levels of abstraction shorten the flow run-time while sacrificing correlation with the physical circuit implementation. Bridging this gap between Logic Synthesis tool and Physical Design (PnR) tools is key to improve Quality of Results (QoR), while possibly shorting the time-to-market. To address this problem, in this work, we formalize logic paths as sentences, with the gates being a bag of words. Thus, we show how word embedding can be leveraged to represent generic paths and predict if a given path is likely to be critical post-PnR. We present the effectiveness of our approach, with accuracy over than 90% for our test-cases. Finally, we give a step further and introduce an intelligent and non-intrusive flow that uses this information to guide optimization. Our flow presents up to 15.53% area delay product (ADP) and 18.56% power delay product (PDP), compared to a standard flow.
Walter Lau Neto, Matheus T. Moreira, Luca G. Amarù, Cunxi Yu, Pierre-Emmanuel Gaillardon
ASP-DAC2
2021 SLAP: A Supervised Learning Approach for Priority Cuts Technology Mapping
abstract
Recently we have seen many works that leverage Machine Learning (ML) techniques in optimizing Electronic Design Automation (EDA) process. However, the uses of ML techniques are limited to learning forecasting models of existing EDA algorithms, instead of developing novel algorithms. In this work, we focus on designing an novel cut-based technology mapping algorithms assisted by ML techniques, which matches results of exhaustive cut exploration but preserving a small footprint of utilized cuts. The proposed approach has been demonstrated with a wide range of benchmarks with 24% reductions in number of cuts utilized compared to the state-of-the-art, while improving the circuit delay, and Area-Delay-Product (ADP), by average about 10%, 7%, respectively, with a 2% area penalty. Compared to the exhaustive approach, i.e., considering all the cuts, we achieve similar or better results while saving over than $2 \times $ the number of considered cuts (runtime) on average. Finally, we provide a comprehensive explanation of heuristics learned by the ML model by feature ranking.
Walter Lau Neto, Matheus T. Moreira, Luca G. Amarù, Cunxi Yu, Pierre-Emmanuel Gaillardon
DAC2
2021 Quasi Delay Insensitive FIFOs: Design Choices Exploration and Comparison
abstract
This paper explores asynchronous FIFOs design choices, more specifically FIFOs from the quasi-delay insensitive (QDI) template family. It proposes eight different asynchronous FIFO structures on a CMOS 45nm technology, using a QDI standard cell library. Structures are exercised through analog-mixed-signal simulation, ranging from nominal to subthreshold supply voltages. Follows a comparison of area, throughput and power efficiency. The experimental results allow inferring a technique for designers to select the most adequate QDI FIFO flavor for specific circuits. Insight on the experiments assesses the beneficial and/or limiting effects of using the specific cell library.
Taciano A. Rodolfo, Marcos L. L. Sartori, Matheus T. Moreira, Ney Laert Vilar Calazans
ISCAS3
2021 A High-Level Modeling Framework for Estimating Hardware Metrics of CNN Accelerators
abstract
GPUs became the reference platform for both training and inference phases of Convolutional Neural Networks (CNN) due to their tailored architecture to the CNN operators. However, GPUs are power-hungry architectures. A path to enable the deployment of CNNs in energy-constrained devices is adopting hardware accelerators for the inference phase. The design space exploration of CNNs using standard approaches, such as RTL, is limited due to their complexity. Thus, designers need frameworks enabling design space exploration that delivers accurate hardware estimation metrics to deploy CNNs. This work proposes a framework to explore CNNs design space, providing power, performance, and area (PPA) estimations. The heart of the framework is a system simulator. The system simulator front-end is TensorFlow, and the back-end is performance estimations obtained from the physical synthesis of hardware accelerators, not only from components like multipliers and adders. The first set of results evaluate the CNN accuracy using integer quantization, the accelerators PPA after physical synthesis, and the benefits of using a system simulator. These results allow a rich design space exploration, enabling selecting the best set of CNN parameters to meet the design constraints.
Leonardo Juracy, Matheus T. Moreira, Alexandre M. Amory, Alexandre F. Hampel, Fernando Gehm Moraes
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 Leveraging QDI Robustness to Simplify the Design of IoT Circuits
abstract
Internet of Things devices require innovative power efficient design techniques that ensure correct operation in harsh environments, where using synchronous design can be challenging. The timing sign-off of synchronous circuits requires analysis and optimisation under multiple corners and operating modes. Considering that energy efficient circuits demand dynamic voltage ranges and harsh environments impose significant variations, design sign-off may become prohibitively expensive. An alternative is quasi-delay-insensitive asynchronous design, which presents robustness against timing variations, simplifying timing sign-off. This paper leverages recent developments in asynchronous circuits design automation to achieve higher degrees of energy efficiency using voltage scaling, while ensuring solid robustness to variability.
Marcos L. L. Sartori, Rodrigo N. Wuerdig, Matheus T. Moreira, Sergio Bampi, Ney Laert Vilar Calazans
ISCAS3
2018 An LSSD Compliant Scan Cell for Flip-Flops
abstract
Most recent timing resilient templates are using asynchronous design techniques and integrating both flip-flops and latches in their design to enable more aggressive performance improvement and reduction in energy consumption. Despite these benefits, they impose challenges in terms of testability because both latches and flip-flops typically use different test protocols. This paper presents an optimized scan cell for flip-flops which is compatible with the protocol used by scannable latches. By using the proposed cell, it is possible to have latches and flip-flops in the same scan chain and the DfT flow fully automated by commercial EDA tools. Experimental results show that the proposed cell reduces silicon area, leakage, and dynamic power compared to the original cell.
Leonardo Juracy, Matheus T. Moreira, Felipe A. Kuentzer, Fernando Gehm Moraes, Alexandre M. Amory
ISCAS2
2018 A DfT Insertion Methodology to Scannable Q-Flop Elements
Leonardo Juracy, Matheus T. Moreira, Felipe A. Kuentzer, Alexandre M. Amory
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Post-processing of supergate networks aiming cell layout optimization
abstract
Recently, methods for switch network generation have gain relevance. The main goal of these techniques is to minimize the number of transistors in the logical arrangement. However, theses methods do not consider optimizations at layout level. In this paper, we propose a post-processing technique in a state-of-art method for network generation to improve some layout aspects such as area, delay, power and parasitic capacitance. Experiments performed over a well-known benchmark demonstrate that the proposed technique allows average gains of 7.48% and 8.48% in the cell area and wirelength, respectively. Electrical characterization results have also shown improvements for propagation delay, transition delay, leakage and switching power in 4.18%, 4.94%, 7.52% and 12.40%, in that order.
Gustavo H. Smaniotto, Regis Zanandrea, Maicon Schneider Cardoso, Renato Souza de Souza, Matheus T. Moreira, Felipe S. Marques 0001, Leomar S. da Rosa Jr.
ISCAS5
2017 Optimized Design of an LSSD Scan Cell
abstract
Type D flip-flop cell and its scannable version called Muxed-D are the most used sequential components in cell-based synchronous designs because it simplifies timing analysis and it is less susceptible to race problems. However, as technology nodes shrink, it becomes more difficult, especially for high-performance designs, to cope with a hard global timing boundary. The use of latches emerges as a possible solution to the contemporary design challenges such as clock skew/jitter, PVT variation, and low-power and high-performance designs. Moreover, latches are also gaining popularity among asynchronous and timing resilient circuits. One of the available scannable cells for latches is called level sensitivity scan-based design (LSSD). The goal of this brief is to present an open design of an optimized single-latch LSSD cell, which has better tradeoffs between propagation delay, power, energy, and silicon area than the original LSSD design, thus reducing the cost for testing latch-based designs.
Leonardo Juracy, Matheus T. Moreira, Felipe A. Kuentzer, Alexandre M. Amory
IEEE Trans. Very Large Scale Integr. Syst.2
2016 A Fine-Grain, Uniform, Energy-Efficient Delay Element for 2-Phase Bundled-Data Circuits
abstract
Contemporary digitally controlled delay elements (DEs) trade off power overheads and delay quantization error (DQE). This article proposes a new programmable DE that provides a balanced design that yields low power with moderate DQE even under process, voltage, and temperature variations. The element employs and leverages the advantages offered by a 28nm fully depleted silicon on insulator technology, using back body biasing to add an extra dimension to its programmability. To do so, a novel generic delay shift block is proposed, which enables incorporating both fine and coarse delays in a single DE that can be easily integrated into digital systems, which is an advantage over hybrid DEs that rely on analog design.
Ajay Singhvi, Matheus T. Moreira, Ramy N. Tadros, Ney Laert Vilar Calazans, Peter A. Beerel
ACM J. Emerg. Technol. Comput. Syst.2
2015 A Bundled-Data Asynchronous Circuit Synthesis Flow Using a Commercial EDA Framework
abstract
Contemporary silicon technology enables integrating billions of transistors and allows the creation of complex systems-on-chip. At the same time, strict power dissipation budgets and growing interest in high performance battery-powered devices drive the need for energy-efficient high performance circuits. Bundled-data asynchronous circuits are good candidates for high performance low power systems, as they operate with average-case delays and present reduced switching activity when compared to other asynchronous templates. The correct operation of bundled-data circuits relies on constraints that describe the timing relationships between data and control signals. However, commercial EDA frameworks do not offer an encompassing support to ensure the closure of such constraints, making implementation challenging. This paper proposes a synthesis flow to enable the description and enforcement of relative timing constraints at both logic and physical synthesis levels, using the Synopsys framework and a set of in-house scripts. Two case studies illustrate the flow: a pipelined multiplier and a network on chip input buffer FIFO, the latter comprising a non-linear pipeline and complex control circuits. Both case studies target the STMicroelectronics 28nm FDSOI technology, and validation occurs with post-layout simulations. Overall, the flow provides an automatic approach to meet relative timing constraints in a template-agnostic manner for bundled-data circuits design.
Matheus Gibiluka, Matheus T. Moreira, Ney Laert Vilar Calazans
DSD2
2014 Hardening QDI circuits against transient faults using delay-insensitive maxterm synthesis
abstract
The correct functionality of quasi-delay-insensitive asynchronous circuits can be jeopardized by the presence and propagation of transient faults. If these faults are latched, they will corrupt data validity and can make the whole circuit to stall, given the strict event ordering constraints imposed by handshaking protocols. This is particularly concerning for the delay-insensitive minterm synthesis logic style, widely adopted by asynchronous designers to implement combinatory quasi-delay-insensitive logic, because it makes extensive use of C-elements and these components are rather vulnerable to transient effects. This paper demonstrates that this logic style submits C-elements to their most vulnerable states during operation. It accordingly proposes the alternative use of the delay-insensitive maxterm synthesis for hardening QDI circuits against transient faults. The latter is a logic style based on the return-to-one 4-phase protocol. Although this style also relies on extensive usage of C-elements, the states where these components are most vulnerable are avoided. Results display improvements of over 300% in C-elements tolerance to transient faults, in the best case.
Matheus T. Moreira, Ricardo A. Guazzelli, Guilherme Heck, Ney Laert Vilar Calazans
ACM Great Lakes Symposium on VLSI1
2014 A design flow for physical synthesis of digital cells with ASTRAN
abstract
As the foundries update their advanced processes with new complex design rules and cell libraries grow in size and complexity, the cost of library development become increasingly higher. In this work we present the methodology used in ASTRAN to allow automatic layout generation of cell libraries for technologies down to 45nm from its transistor level netlist description in SPICE format. It supports non-complementary logic cells, allowing generation of any kind of transistor networks, and continuous transistor sizing. We describe our new generation flow which is currently being used to generate a library with more than 500 asynchronous cells in a 65nm process.
Adriel Ziesemer, Ricardo Augusto da Luz Reis, Matheus T. Moreira, Michel Evandro Arendt, Ney Laert Vilar Calazans
ACM Great Lakes Symposium on VLSI3
2014 Advances on the state of the art in QDI design
abstract
The ever increasing demand for more complex systems and the possibility of integrating billions of transistors in a single chip brought us to the boundaries of the synchronous paradigm capabilities. The efficient distribution of a global clock signal in a contemporary complex design can be an intricate task and even with the most modern techniques it can consume a significant portion of the total power of a contemporary chip. In fact, according to Amde et al. in [1], clock power represents in average 45% of a synchronous chip total power. Hence, as power budgets get tighter, motivated by battery-based applications demands, and performance gets over constrained by aggressive process variations, traditional design techniques prove to be unsustainable. In this scenario, asynchronous circuits emerge as a promising solution to cope with technological problems faced by synchronous designers and regain the attention of the research community.
Matheus T. Moreira, Ney Laert Vilar Calazans
VLSI-SoC1
2014 Beware the Dynamic C-Element
abstract
The C-element is a well known component of asynchronous circuits. To overcome problems of current CMOS technologies, its use has even been extended to specific domains of the synchronous paradigm, such as clock generation, clock gating, and registers. An economical implementation of this component is the dynamic C-element. Its advantages over static implementations are reduced power, transition, and propagation delays as well as lower silicon area. Yet, research evaluating its electrical behavior, functionality, and robustness is scarce. This brief presents an in-depth analysis of the dynamic C-element electrical behavior. The analysis points to a constrained nature, which can lead to undefined output logic values, as well as excessive static power consumption. The brief also proposes a technique for robust design of such components that avoids such undefined values.
Matheus T. Moreira, Fernando Gehm Moraes, Ney Laert Vilar Calazans
IEEE Trans. Very Large Scale Integr. Syst.1
2013 LiChEn: Automated Electrical Characterization of Asynchronous Standard Cell Libraries
abstract
Semi-custom design flows are a key factor for the rapid growth of integrated circuits and systems. They lower design complexity through the use of pre-designed and pre-characterized functional components called standard cells, instead of assuming that designers have to draw, place and connect each transistor. In this way, modeling of complex systems is easier. As CMOS technologies evolve into deep sub micron nodes, asynchronous techniques gain relevance in the research community, due to their ability to cope with problems that are hard to solve with the synchronous paradigm. However, several specific components required in asynchronous designs are not available in commercial standard cell libraries, which constrains asynchronous design to use approaches close to full-custom ones. This limits modularity and increases design complexity. Thus, one of the possibilities for enabling further advance of the asynchronous paradigm is the availability of asynchronous standard cell libraries. Albeit industrial tools provide reasonable support to asynchronous standard cells physical design, the characterization of these cells using standard tools is usually quite laborious. This work proposes the Library Characterization Environment (LiChEn), an open source tool applicable to automatically characterize typical asynchronous standard cells. The tool managed to successfully characterize a standard cell library with over five hundred asynchronous components.
Matheus T. Moreira, Carlos Henrique Menezes Oliveira, Ney Laert Vilar Calazans, Luciano Ost
DSD1
2013 Voltage scaling on C-elements: A speed, power and energy efficiency analysis
abstract
This work reports an evaluation of speed, energy consumption, leakage power, and silicon area tradeoffs of three different transistors topologies for C-elements, basic devices for building asynchronous circuits. The evaluation considers the devices operating under supply voltages that vary from nominal IV to 0.05V. Analog simulations provide precise measurements and the obtained results identify the lowest voltage at which each C-element operates correctly. Results suggest that operating at near-threshold voltages provides the best speed-energy and speed-leakage efficiencies. Also, they point the van Berkel topology as the most suited C-element implementation for low voltage operation, as it presents lower power and energy figures as well as higher speed, regardless of the supply voltage.
Matheus T. Moreira, Ney Laert Vilar Calazans
ICCD1
2013 BaBaNoC: An asynchronous network-on-chip described in Balsa
abstract
The downscaling of silicon technology and the possibility of building MPSoCs, make intrachip communication a mainstream research topic. NoCs are an elegant solution to provide communication scalability and modularity. NoCs are already common in MPSoC design. Moreover, new technology challenges point to a growth in the use of non-synchronous NoCs. However, the design of asynchronous infrastructures with current EDA tools is challenging. That is due to the fact that most of these tools are oriented towards synchronous design. This work proposes and evaluates a fully asynchronous NoC router based on the Balsa language and framework. The design is validates through FPGA synthesis.
Matheus T. Moreira, Felipe G. Magalhaes, Matheus Gibiluka, Fabiano Hessel, Ney Laert Vilar Calazans
RSP1