José Luís Güntzel

dblp:04/3287 · also José Luís Almada Güntzel · DBLP profile ↗
← Back
27ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-7712-869XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 A Parallel JPEG Pleno Baseline Block-Based Profile Light Field Codec Using OpenMP
abstract
Light Fields (LFs) are an image modality derived from the plenoptic function, capable of representing scenes with a richer amount of visual information, making them well-suitable for immersive applications. To efficiently compress this data-intensive modality, the Joint Photographic Experts Group (JPEG) committee created the JPEG Pleno Part 2 standard with two profiles. This work focuses on the reference encoder and decoder implementations for the Baseline Block-Based Profile (BBBP), i.e., the JPEG PLeno Model (JPLM). Our main contribution lies in the proposal and analysis of a parallel implementation of JPLM codec using OpenMP, including the coding efficiency overhead of three solutions to support parallel decoding: two based on a marker defined in Part 2, called PNT; and another non-compliant version of the entropy codec. We show that it is possible to accelerate encoding up to$12\times $when using 24threads, and decoding up to$4\times $with 10 threads. In both cases, the memory overhead remained below 20% for the tested LFs. Although the parallel encoder produces bit-exact matches when compared to the sequential version, a PNT marker is required to ensure parallel decoding, which causes an increase of about 4% in terms of BD-Rate. On the other hand, we show that by updating the arithmetic codec to avoid the PNT, this overhead reduces to at most 0.22%.
André Filipe da Silva Fernandes, Ismael Seidel, Leonardo de Sousa Marques, Arthur Scarpatto Rodrigues, José Luís Güntzel
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 Eh-DRVP: Combining placement and global routing data in a hyper-image-based DRV predictor
Sheiny Fabre Almeida, Renan Netto, Tiago Fontana, Erfan Aghaeekiasaraee, Upma Gandhi, Aysa Fakheri Tabrizi, José Luís Güntzel, Laleh Behjat, Cristina Meinhardt
Integr.7
2024 ILPGRC: ILP-Based Global Routing Optimization With Cell Movements
abstract
The placement and routing steps directly impact the circuit performance, area, power consumption, and reliability. To handle the high complexity of modern circuits, these steps are tackled separately by applying a divide-and-conquer approach. Unfortunately, due to the continuous increase of design rules complexity, the convergence of solutions can suffer from misalignment, and the effects of an unsatisfactory placement will be noticed only during routing when the placement is considered fixed. In this work, we propose the ILPGRC, an integer linear programming (ILP)-based technique that simultaneously moves cells and routes nets to optimize Global Routing. ILPGRC enables the relocation of cells that can lead to routing issues without compromising the quality concerning the number of VIAs, wirelength, and design rule violations (DRVs). We also propose a partitioning strategy named Checkered paneling, which reduces the input size of the ILP model, making this approach scalable. The Checkered paneling strategy enables the execution of multiple ILP models in parallel, providing a speedup for large circuits. Additionally, we propose a GCell cluster-based approach to legalize the solution with minimum disturbance and displacement. We evaluated our technique for the ISPD 2018 and ISPD 2019 Contests circuits within a physical synthesis flow composed of state-of-the-art place and route academic tools. The results after the detailed routing show that ILPGRC can reduce, on average, the number of VIAs by 4.69% with less than 1% impact on wirelength. Additionally, ILPGRC reduces the number of DRVs in most cases with no open nets left.
Tiago Fontana, Erfan Aghaeekiasaraee, Renan Netto, Sheiny Fabre Almeida, Upma Gandhi, Laleh Behjat, José Luís Güntzel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2023 CRP2.0: A Fast and Robust Cooperation between Routing and Placement in Advanced Technology Nodes
abstract
Traditionally, the placement and routing stages of a physical design are performed separately. Because of the additional complexities arising in advanced technology nodes, they have become more interdependent. Therefore, creating efficient cooperation between the routing and placement steps has become an important topic in Electronic Design Automation (EDA). In this article, a framework that allows cooperation between routing and placement is proposed. The main objective of the proposed framework is to improve the detailed routing solution by combining routing and placement. The core of this framework is the Cooperation between Routing and Placement (CRP2.0) 1 engine including techniques to combine routing and placement. The key contributions of CRP2.0 include an Integer Linear Programming (ILP)-based Detailed Placement (ILP-DP), net classification, and two Cost and Net Caching techniques. The efficacy of the proposed framework is evaluated on the official ACM/IEEE International Symposium on Physical Design (ISPD) 2018 and 2019 contest benchmarks. In this article, we show that by using the Cost Caching technique, the global routing runtime compared with state-of-the-art algorithms was reduced by 28.56%, on average. Moreover, numerical results show that when working with advanced technology nodes, the proposed framework can improve the detailed routing score by an average of 0.3% while only moving 0.7% of the cells, on average. The proposed engine can be employed as an add-on to the physical design flow between the global routing and detailed routing steps.
Erfan Aghaeekiasaraee, Aysa Fakheri Tabrizi, Tiago Fontana, Renan Netto, Sheiny Fabre Almeida, Upma Gandhi, José Luís Güntzel, David T. Westwick, Laleh Behjat
ACM Trans. Design Autom. Electr. Syst.7
2022 Routability-Driven Detailed Placement Using Reinforcement Learning
abstract
Technology advancements have enabled us to manufacture integrated circuits composed of a sheer number of gates onto a single chip. However, these enhancements have also introduced new challenges. In physical synthesis, the placement and routing steps have to satisfy even more complex design rules while optimizing the solution quality. However, the search for wirelength optimization may lead the placement engine to produce an infeasible routing solution, making it necessary to repeat previous steps and increase the overall project cost. Traditionally, placement algorithms estimate routability using pin density because of its low computational cost. Nonetheless, in advanced technology nodes, this has become inefficient due to more restrictive manufacturing constraints and complex standard cell layouts. Although many placement techniques propose to address routability, the problem is that these models rely on specific heuristics or designer experience. Therefore, we propose a machine learning-based framework for addressing routability during the placement step.
Sheiny Fabre Almeida, José Luís Güntzel, Laleh Behjat, Cristina Meinhardt
VLSI-SoC2
2022 Algorithm Selection Framework for Legalization Using Deep Convolutional Neural Networks and Transfer Learning
abstract
Machine learning (ML) models have been used to improve the quality of different physical design steps, such as timing analysis, clock tree synthesis, and routing. However, so far very few works have addressed the problem of algorithm selection during physical design, which can drastically reduce the computational effort of some steps. This work proposes a legalization algorithm selection framework using deep convolutional neural networks (CNNs). To extract features, we used snapshots of circuit placements and used transfer learning to train the models using pretrained weights of the Squeezenet architecture. By doing so, we can greatly reduce the training time and required data even though the pretrained weights come from a different problem. We performed extensive experimental analysis of ML models, providing details on how we chose the parameters of our model, such as CNN architecture, learning rate, and number of epochs. We evaluated the proposed framework by training a model to select between different legalization algorithms according to cell displacement and wirelength variation. The trained models achieved an average$F$-score of 0.98 when predicting cell displacement and 0.83 when predicting wirelength variation. When integrated into the physical design flow, the cell displacement model achieved the best results on 15 out of 16 designs, while the wirelength variation model achieved that for 10 out of 16 designs, being better than any individual legalization algorithm. Finally, using the proposed ML model for algorithm selection resulted in a speedup of up to$10\times $compared to running all the algorithms separately.
Renan Netto, Sheiny Fabre Almeida, Tiago Fontana, Vinicius S. Livramento, Laércio Lima Pilla, Laleh Behjat, José Luís Güntzel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 Relying on a Rate Constraint to Reduce Motion Estimation Complexity
abstract
This paper proposes a rate-based candidate elimination strategy for Motion Estimation, which is considered one of the main sources of encoder complexity. We build from findings of previous works that show that selected motion vectors are generally near the predictor to propose a solution that uses the motion vector bitrate to constrain the candidate search to a subset of the original search window, resulting in less distortion computations. The proposed method is not tied to a particular search pattern, which makes it applicable to several ME strategies. The technique was tested in the VVC reference software implementation and showed complexity reductions of over 80% at the cost of an average 0.74% increase in BD-Rate with respect to the original TZ Search algorithm in the LDP configuration.
Gabriel B. Sant'Anna, Luiz Henrique Cancellier, Ismael Seidel, Mateus Grellert, José Luís Güntzel
ICASSP5
2021 Design of Energy-Efficient Gaussian Filters by Combining Refactoring and Approximate Adders
abstract
The Gaussian image filter is a compute-intensive approach to reduce undesirable artifacts and generally serves as a pre-processing technique for emerging applications related to visual computing systems. This work evaluates alternatives for the design of power-efficient Gaussian Filters. The proposed optimization strategy combines: 1) a refactored function to minimize the arithmetic operations, and 2) a design space exploration investigating different approximation scenarios applied to the full adders. The exact version of our refactored Gaussian Filter architecture reduces the total power consumption and the circuit area by 18% and 12%, respectively compared with the baseline Gaussian Filter architecture. Moreover, the combination of different approximation levels with the refactored architecture provides design options with power reductions from 21% to 59% compared with the baseline Gaussian Filter architecture.
Marcio Monteiro, Pedro Aquino Silva, Ismael Seidel, Mateus Grellert, Leonardo Bandeira Soares, José Luís Güntzel, Cristina Meinhardt
ISCAS6
2021 Hardware-Friendly Search Patterns for the Versatile Video Coding Fractional Motion Estimation
abstract
The recently finalized Versatile Video Coding (VVC) standard brings a number of new tools to substantially improve coding efficiency, which resulted in a significant increase in complexity. Therefore, the use of such new standard in mobile devices requires not only the use of dedicated VLSI hardware architectures, but also the adoption of techniques that could reduce its complexity to meet the real-time and energy efficiency requirements. In this context, the most intensive encoding tools, such as the Fractional Motion Estimation (FME), are the first ones to be considered for optimization. Thereby, this work presents three different search patterns for the FME that are able to reduce the complexity of VVC by creating fixed search windows around the most relevant candidates. Our experiments show that these patterns result in acceptable decreases in coding efficiency, with average BD-Rates in the range of 0.56% - 0.33%, for the LD-P configuration, and 0.47% - 0.21% for the RA configuration, while allowing for significant hardware optimization. Comparing to a state-of-the-art FME hardware architecture that searches over 48 candidates the three proposed patterns can operate in higher frequencies to achieve the same throughput and require less area, leading to less power consumption and higher energy efficiency. Particularly, one of the three patterns may lead to a 60% area reduction and 53% less dynamic power consumption.
Vanio Rodrigues Filho, Marcio Monteiro, Ismael Seidel, Mateus Grellert, José Luís Güntzel
MMSP5
2021 SAD or SATD? How the Distortion Metric Impacts a Fractional Motion Estimation VLSI Architecture
abstract
Video coding systems have to deal with a number of tradeoffs. The decision of adopting a specific distortion metric in the Fractional Motion Estimation (FME) step, for instance, presents a designer with a tradeoff between energy and coding efficiency. This paper analyzes such a tradeoff considering two of the most known and used distortion metrics, the Sum of Absolute Differences (SAD) and the Sum of Absolute Transformed Differences (SATD), within a High Efficiency Video Coding (HEVC)-compatible FME hardware architecture. We show that the SATD-based FME architecture is 1.94 times larger than the SAD-based one and consumes 2.07 times more energy.
Ismael Seidel, Vanio Rodrigues Filho, Mateus Grellert, Luciano Volcan Agostini, José Luís Güntzel
MMSP5
2019 How Deep Learning Can Drive Physical Synthesis Towards More Predictable Legalization
abstract
Machine learning has been used to improve the predictability of different physical design problems, such as timing, clock tree synthesis and routing, but not for legalization. Predicting the outcome of legalization can be helpful to guide incremental placement and circuit partitioning, speeding up those algorithms. In this work we extract histograms of features and snapshots of the circuit from several regions in a way that the model can be trained independently from region size. Then, we evaluate how traditional and convolutional deep learning models use this set of features to predict the quality of a legalization algorithm without having to executing it. When evaluating the models with holdout cross validation, the best model achieves an accuracy of 80% and an F-score of at least 0.7. Finally, we used the best model to prune partitions with large displacement in a circuit partitioning strategy. Experimental results in circuits (with up to millions of cells) showed that the pruning strategy improved the maximum displacement of the legalized solution by 5% to 94%. In addition, using the machine learning model avoided from 22% to 99% of the calls to the legalization algorithm, which speeds up the pruning process by up to 3x.
Renan Netto, Sheiny Fabre Almeida, Tiago Fontana, Vinicius S. Livramento, Laércio Lima Pilla, José Luís Güntzel
ISPD6
2018 Coding- and Energy-Efficient FME Hardware Design
abstract
Hybrid video standards rely on encoding prediction residues. To improve coding efficiency of inter-frame prediction, interpolated samples may be generated in fractional positions i.e., between neighbor pixels in the reference frame. However, performing Fractional Motion Estimation (FME) increases the overall encoder complexity. Since portable mobile devices are increasingly used to capture and reproduce videos, energy-efficient FME hardware accelerators are of utmost importance. In this work, we propose and evaluate a coding- and energy-efficient hardware design strategy for FME. Such strategy addresses the main weaknesses of the architectures found in the literature. The architecture designed as case study can achieve 2160p@120fps for the HEVC 8×8 FME. We also provide an insightful area and power breakdown of the synthesized design, to drive the design of FME hardware towards further energy improvements.
Ismael Seidel, Vanio Rodrigues Filho, Luciano Volcan Agostini, José Luís Güntzel
ISCAS4
2017 How Game Engines Can Inspire EDA Tools Development: A use case for an open-source physical design library
abstract
Similarly to game engines, physical design tools must handle huge amounts of data. Although the game industry has been employing modern software development concepts such as data-oriented design, most physical design tools still relies on object-oriented design. Differently from object-oriented design, data-oriented design focuses on how data is organized in memory and can be used to solve typical object-oriented design problems. However, its adoption is not trivial because most software developers are used to think about objects' relationships rather than data organization. The entity-component design pattern can be used as an efficient alternative. It consists in decomposing a problem into a set of entities and their components (properties). This paper discusses the main data-oriented design concepts, how they improve software quality and how they can be used in the context of physical design problems. In order to evaluate this programming model, we implemented an entity-component system using the open-source library Ophidian. Experimental results for two physical design tasks show that data-oriented design is much faster than object-oriented design for problems with good data locality, while been only sightly slower for other kinds of problems.
Tiago Fontana, Renan Netto, Vinicius S. Livramento, Chrystian Guth, Sheiny Fabre Almeida, Laércio Lima Pilla, José Luís Güntzel
ISPD7
2017 Incremental Layer Assignment Driven by an External Signoff Timing Engine
abstract
Modern technologies provide wide and thick metal layers that must be wisely used to reduce the delay of critical interconnections. After global routing, incremental layer assignment can improve the circuit timing by properly selecting critical interconnect segments to be routed in the faster (but very limited) wires on upper layers. Existing techniques based on net-by-net iterative improvement may get stuck at locally-optimal solutions depending on net ordering. Recent techniques rule out such drawback through the simultaneous iterative improvement of all nets, but they unfortunately rely on objective functions that may guide the optimization off critical paths. As opposed to all reported techniques, which rely on simplified, overly pessimistic timing models, this paper proposes the decoupling of incremental layer assignment from the timing analysis and the exploitation of flow conservation conditions so as to enable the use of an external signoff timing engine. The novel technique was experimentally compared with two state-of-the art works, leading to 50% less timing violations under total negative slack metric and 35% less timing violations under worst negative slack metric with similar overhead in number of vias.
Vinicius S. Livramento, Derong Liu 0002, Salim Chowdhury, Bei Yu 0001, David Z. Pan, José Luís Güntzel, Luiz Cláudio Villar dos Santos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2016 Rate-constrained successive elimination of Hadamard-based SATDs
abstract
The efficiency improvements achieved by new video coding standards come at the cost of a huge increase in the encoder computational complexity. Paradoxically, such increasing complexity is commonly addressed by methods that have an adverse effect on coding efficiency. In this work, we propose a method to reduce the complexity of HEVC Hadamard ME, without compromising coding efficiency. Our method relies on a new metric named Absolute First Difference (AFD), which is able to eliminate impossible candidates using only a few operations. Despite its simplicity, AFD filters an average of 13.52% of SATDs in the HEVC reference model (HM), without reducing coding efficiency. Moreover, our method is able to filter up to 37.9% of SATDs for video conferencing and up to 64% for screen content.
Ismael Seidel, Luiz Henrique Cancellier, José Luís Güntzel, Luciano Volcan Agostini
ICIP3
2016 Energy-efficient SATD for beyond HEVC
abstract
State-of-the-art video coding standards adopt large block and transform sizes. Moreover, as resolutions keep growing, there is a trend in adopting even larger structures in future video encoders, resulting in higher complexity. Therefore, the design of energy-efficient architectures for variable block size distortion metrics are key to keep the energy requirements of battery devices under a reasonable budget. The Hadamard-based Sum of Absolute Transformed Differences (SATD) is used as distortion metric in several steps of encoding, increasing the overall encoding efficiency at the cost of rising complexity. In this work, we propose two main approaches for SATD calculation of N × N block sizes. One using a Transpose Buffer (TB) and another one using a Linear Buffer (LB). Furthermore, we synthesized four sizes (4 × 4 up to 32 × 32) of each main SATD architecture to evaluate their area and energy estimates. The results show a large increase in area for TB-SATD, while LB-SATD area increases in a much smaller pace. On the other hand, the TB-SATD synthesized for Low-Vdd/High-Vt show up to be the most energy-efficient architectures for all sizes.
Ismael Seidel, André Beims Bräscher, José Luís Güntzel, Luciano Volcan Agostini
ISCAS3
2016 Clock-Tree-Aware Incremental Timing-Driven Placement
abstract
The increasing impact of interconnections on overall circuit performance makes timing-driven placement (TDP) a crucial step toward timing closure. Current TDP techniques improve critical paths but overlook the impact of register placement on clock tree quality. On the other hand, register placement techniques found in the literature mainly focus on power consumption, disregarding timing and routabilty. Indeed, postponing register placement may undermine the optimization achieved by TDP, since the wiring between sequential and combinational elements would be touched. This work proposes a new approach for an effective coupling between register placement and TDP that relies on two key aspects to handle sequential and combinational elements separately: only the registers in the critical paths are touched by TDP (in practice they represent a small percentage of the total number of registers), and the shortening of clock tree wirelength can be obtained with limited variation in signal wirelength and placement density. The approach consists of two steps: (1) incremental register placement guided by a virtual clock tree to reduce clock wiring capacitance while preserving signal wirelength and density, and (2) incremental TDP to minimize the total negative slack. For the first step, we propose a novel technique that combines clock-net contraction and register clustering forces to reduce the clock wirelength. For the second step, we propose a novel Lagrangian Relaxation formulation that minimizes total negative slack for both setup and hold timing violations. To solve the formulation, we propose a TDP technique using a novel discrete search that employs a Euclidean distance to define a proper neighborhood. For the experimental evaluation of the proposed approach, we relied on the ICCAD 2014 TDP contest infrastructure and compared our results with the best results obtained from that contest in terms of timing closure, clock tree compactness, signal wirelength, and density. Assuming a long displacement constraint, our technique achieves worst and total negative slack reductions of around 24% and 26%, respectively. In addition, our approach leads to 44% shorter clock tree wirelength with negligible impact on signal wirelength and placement density. In the face of such results, the proposed coupling seems a useful approach to handle the challenges faced by contemporary physical synthesis.
Vinicius S. Livramento, Renan Netto, Chrystian Guth, José Luís Güntzel, Luiz Cláudio Villar dos Santos
ACM Trans. Design Autom. Electr. Syst.4
2015 Exploiting Non-Critical Steiner Tree Branches for Post-Placement Timing Optimization
abstract
The increasing impact of interconnections on the overall circuit performance renders physical design a crucial step to timing closure. Several techniques are used to optimize timing within the flow, such as gate sizing, buffer insertion, and timing-driven placement (TDP). Unfortunately, gate sizing and buffer insertion are not capable of modifying the length of interconnections. Although TDP is able to shorten critical interconnection by finding new legal locations for a subset of cells, it generally overlooks the impact of non-critical branches on the delay of critical cells. This work proposes a post-placement timing optimization technique to reduce the capacitive load of critical cells by shortening non-critical Steiner tree branches. To shorten such branches, our technique uses computational geometry for finding effective cell movements that consider maximum displacement constraints and macro blocks. Our experiments evaluate the capability of our technique to further reduce the timing violations from a TDP solution. We applied our technique on the solutions obtained by the top 3 teams in the ICCAD 2014 TDP Contest, where short and long displacement constraints are defined. For the short constraints, the average reductions assuming worst and total late negative slack metrics are 23% and 34%. Considering the long constraints, the average reductions are 62% and 67%. We also present extensions of our technique to tackle related physical design problems such as early violations reduction and electrical correction.
Vinicius S. Livramento, Chrystian Guth, Renan Netto, José Luís Güntzel, Luiz Cláudio Villar dos Santos
ICCAD4
2015 Timing-Driven Placement Based on Dynamic Net-Weighting for Efficient Slack Histogram Compression
abstract
Timing-driven placement (TDP) finds new legal locations for standard cells so as to minimize timing violations while preserving placement quality. Although violations may arise from unmet setup or hold constraints, most TDP approaches ignore the latter. Besides, most techniques focus on reducing the worst negative slack and let the improvements on total negative slack as a secondary goal. However, to successfully achieve timing closure, techniques must also reduce the total negative slack, which is known as slack histogram compression. This paper proposes a new Lagrangian Relaxation formulation for TDP to compress both late and early slack histograms. To solve the problem, we employ a discrete local search technique that uses the Lagrange multipliers as net-weights, which are dynamically updated using an accurate timing analyzer. To preserve placement quality, our technique uses a small fixed-size window that is anchored in the initial location of a cell. For the experimental evaluation of the proposed technique, we relied on the ICCAD 2014 TDP contest infrastructure. The results show that our technique significantly reduces the timing violations from an initial global placement. On average, late and early total negative slacks are improved by 85.03% and 42.72%, respectively, while the worst slacks are reduced by 71.55% and 34.40%. The overhead in wirelength is less than 0.1%.
Chrystian Guth, Vinicius S. Livramento, Renan Netto, Renan Fonseca, José Luís Güntzel, Luiz Cláudio Villar dos Santos
ISPD5
2014 A Hybrid Technique for Discrete Gate Sizing Based on Lagrangian Relaxation
abstract
Discrete gate sizing has attracted a lot of attention recently as the EDA industry faces the challenge of optimizing large standard cell-based circuits. The discrete nature of the problem, along with complex timing models, stringent design constraints, and ever-increasing circuit sizes, make the problem very difficult to tackle. Lagrangian Relaxation (LR) is an effective technique to handle complex constrained optimization problems and therefore has been successfully applied to solve the gate sizing problem. This article proposes an improved Lagrangian relaxation formulation for discrete gate sizing that relaxes timing, maximum gate input slew, and maximum gate output capacitance constraints. Based on such formulation, we propose a hybrid technique composed of three steps. First, a topological greedy heuristic solves the LR formulation. Such a heuristic is applied assuming a slightly increased target clock period (backoff factor) to better explore the solution space. Second, a delay recovery heuristic reestablishes the original target clock with small power overhead. Third, a power recovery heuristic explores the remaining slacks to further reduce power. Experiments on the ISPD 2012 Contest benchmarks show that our hybrid technique provides less leakage power than the state-of-the-art work for every circuit from the ISPD 2012 Contest infrastructure, achieving up to 24% less leakage. In addition, our technique achieves a much better compromise between leakage reduction and runtime, obtaining, on average, 9% less leakage power while running 8.8 times faster.
Vinicius S. Livramento, Chrystian Guth, José Luís Güntzel, Marcelo O. Johann
ACM Trans. Design Autom. Electr. Syst.3
2013 Fast and efficient lagrangian relaxation-based discrete gate sizing
abstract
Discrete gate sizing has attracted a lot of attention recently as the EDA industry faces the challenge of optimizing large standard cell-based circuits. The discreteness of the problem, along with complex timing models, stringent constraints and ever increasing circuit sizes make the problem very difficult to tackle. Lagrangian Relaxation is an effective technique to handle complex constrained optimization problems and therefore has been used for gate sizing. In this paper, we propose an improved Lagrangian Relaxation formulation for leakage power minimization that accounts for maximum gate input slew and maximum gate output capacitance in addition to the circuit timing constraints. We also present a fast topological greedy heuristic to solve the Lagrangian Relaxation Subproblem and a complementary procedure to fix the few remaining slew and capacitace violations. The experimental results, generated by using the ISPD 2012 Discrete Gate Sizing Contest infrastructure, show that our technique is able to optimize a circuit with up to 959K gates within only 51 minutes. Comparing to the ISPD Contest top three teams, our technique obtained on average 18.9%, 16.7% and 43.8% less leakage power, while being 38, 31 and 39 times faster.
Vinicius S. Livramento, Chrystian Guth, José Luís Güntzel, Marcelo O. Johann
DATE3
2013 Quality assessment of subsampling patterns for pel decimation targeting high definition video
abstract
Although pel decimation has been widely used to reduce the computational effort in video coding, there is no consensus about the optimal subsampling pattern. This paper presents an extensive analytical and statistical comparison of several different subsampling patterns using analysis of variance. The investigation includes commonly used patterns as well as seven proposed ones. The experiments were conducted on 19 video samples in a range of five resolutions (being nine videos at 1080p) and 30 bitrates by using the state-of-the-art x264 encoder running successive elimination exhaustive search. Two objective quality metrics were reported, PSNR and DSSIM, resulting in 10680 experimental points. The analysis of such amount of data allowed us to conclude that the proposed 4:3 ratio shows less than 5% in DSSIM and 1% in PSNR losses, being more than two times faster than full sampling. Compared with higher decimation, it presents a better trade-off between speedup and quality loss.
Ismael Seidel, Bruno George de Moraes, Emilio Wuerges, José Luís Güntzel
ICME4
2011 An energy-efficient 8×8 2-D DCT VLSI architecture for battery-powered portable devices
abstract
This paper presents an energy-efficient VLSI architecture for 8×8 2-D DCT, which relies on a fast and precise implementation of the LLM algorithm. The energy-efficiency is achieved by using a combinational 1-D DCT block that explores the algorithm's intrinsic parallelism and the integer constant multiplications. The target throughput of 19 Mpixels/s, which is required for VGA@30fps, is achieved by applying a 4.9 MHz clock, that corresponds only to 17.5% of the maximum clock. Synthesis results for a 350 nm technology estimate total power as 6.08 mW, and core area as 2.1 mm2. The proposed architecture shows to be at least 42% more energy efficient than the related work. To further investigate the efficiency on deep submicron technology nodes, synthesis for 90 nm and 45 nm were also performed.
Vinicius S. Livramento, Bruno George de Moraes, Brunno Abner Machado, José Luís Güntzel
ISCAS4
2007 RIC Fast Adder and its Set Tolerant Implementation in FPGAs
abstract
FPGA is currently a very important design technology to implement electronic systems due to its high logic density, its fast time-to-market and its low cost. But in order to provide high logic density FPGA devices are fabricated with nanometer CMOS technology that is becoming susceptible to radiation-induced soft errors. Among these errors, single-event transients (SETs) are those that are induced in the user's programmable logic. This paper presents a new fast adder, called RIC (Re-computing the Inverse Carry-in) and shows how this new adder architecture may be used to build SET-tolerant fast adders. Results considering FPGA-based implementation are presented.
Eduardo Mesquita, Helen Franck, Luciano Volcan Agostini, José Luís Güntzel
FPL4
2006 High throughput architecture for H.264/AVC forward transforms block
abstract
This paper presents a high throughput hardware for the complete H.264/AVC forward transforms block. There are three different transform inside this block and the presented architecture synchronizes these transforms, generating a constant processing rate in its outputs. This is an important characteristic of this architecture that was designed to be easily integrated to the other H.264/AVC blocks. The architecture does not use memory bits and the transforms in two dimensions are calculated directly, without the use of the separability property. The architecture was described in VHDL and was validated and prototyped using a Xilinx Virtex II Pro FPGA. The synthesis was directed to a VP30 FPGA and to a TSMC 0.35μm standard-cell technology. The throughputs of the T block architecture for these two different technologies reaches a processing rate higher than 120 million of samples per second, allowing its use in H.264/AVC codecs directed to HDTV.
Luciano Volcan Agostini, Roger Endrigo Carvalho Porto, Sergio Bampi, Leandro Rosa, José Luís Güntzel, Ivan Saraiva Silva
ACM Great Lakes Symposium on VLSI5
2006 High throughput multitransform and multiparallelism IP for H.264/AVC video compression standard
abstract
This paper presents the design of a high throughput multitransform and multiparallelism IP for H.264/AVC standard. This solution supports the five H.264/AVC transforms and it supports five different levels of parallelism. The proposed architecture were described in VHDL and synthesized to Altera Stratix and Xilinx Virtex-II Pro FPGAs and to TSMC 0.35/spl mu/m standard cells. The multitransform and multiparallelism architecture mapped to FPGAs could process from 124 millions to 3.2 billions of samples per second, depending on the parallelism level selected. The standard cells version could process from 218.7 millions to 3.5 billions of samples per second. These results indicate that the proposed solution presents a high flexibility and that this solution is able to be used in various H.264/AVC codecs with different performance requirements. The performance results of all experiments realized indicated that this architecture is able to be used in high definition applications, like HDTV.
Luciano Volcan Agostini, Roger Endrigo Carvalho Porto, José Luís Güntzel, Ivan Saraiva Silva, Sergio Bampi
ISCAS3
2003 A New Macro-cell Generation Strategy for three metal layer CMOS Technologies
Cristiano Lazzari, Cristiano Viana Domingues, José Luís Güntzel, Ricardo Augusto da Luz Reis
VLSI-SOC3