Luiz Cláudio Villar dos Santos

dblp:121/4418 · also Luiz C. V. dos Santos · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-9384-8347ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Multicore Environment State Representation for Agent-Directed Test Generation
abstract
A crucial step in the design of multicore systems is to validate the interaction between cores. This involves test program generation and runtime analysis. We propose a novel reinforcement learning approach to directed test generation, where an agent induces a suite of programs, which are executed in a simulation environment for a multicore. It focuses on how to recover state information from raw observations of the environment such that the agent can learn from interaction how to improve coverage for any verification task. We evaluated our state representation for different verification tasks involving 16 and 32-core ARMv8 2-level MOESI designs.
Bruno D. Miranda, Luiz M. V. Pereira, Márcio Castro 0001, Luiz Cláudio Villar dos Santos
DAC4
2025 A Canonical Test Representation for Verification of Shared-Memory Behavior in Multiprocessor Systems
abstract
The scope of this article is the design verification of a multicore chip or multichip multiprocessor by running concurrent test programs until coverage goals are reached. Interactions between multiple processors through shared memory must obey a memory consistency model, which specifies valid behaviors. We propose a canonical test-program representation that encodes primal shared-memory behaviors to be induced at runtime. It is intended as one of the main keys to the design of new test generators. We prove that our representation does not limit the search space, because it induces equivalence classes that can be completely and uniquely encoded. In particular, we show experimental evidence that our representation is also suitable to learning-based test generators, because it enables the design of effective actions. We have built a generator directed by a Reinforcement Learning agent, designed its actions based on our encoding, and compared it with three generators when targeting 32-core designs. For a given time limit, our generator reached the largest coverage and led to the fastest error diagnosis in 3/4 of the verification scenarios, despite our choice of a minimalist agent. The theoretical guarantees and the experimental evidence indicate that our representation provides proper grounds for defining effective actions, and it prevents them from either inducing redundant tests or limiting the test suite.
Bruno D. Miranda, Márcio Castro 0001, Luiz Cláudio Villar dos Santos
ACM Trans. Design Autom. Electr. Syst.3
2023 EveCheck: An Event-Driven, Scalable Algorithm for Coherent Shared Memory Verification
abstract
Cache coherence and relaxed memory consistency challenge the design verification of multicore chips. The well-known self-checking approach based on litmus tests has been successful in uncovering design errors, but it leads to limited coverage. That is why random or directed test generation approaches are used for proper coverage control, but they require a specialized checker to verify shared memory behavior. Unfortunately, the inference-based checkers that provide guarantees against false positives and false negatives are not scalable with core count (unless some memory orderings are known beforehand). On the other hand, scoreboard-based checkers are scalable, but do not offer guarantees against false diagnosis. This article proposes an event-driven algorithm that is scalable without renouncing verification quality, because it relies on direct partial order verification, on-the-fly detection of improper memory orderings, exploitation of extended observability not hampered by design optimizations, and agnostic handling of store atomicity at the implementation level. We compared the new algorithm with two scoreboard-based and one inference-based checker reported in the literature. EveCheck often reaches full error detection with half the test size required by the other checkers, has superior verification quality, and less sensitivity to test generation parameters. Its effort to detect errors was one order of magnitude inferior as compared to the inference-based checker. It did not raise false positives in correct designs, and it was able to discover all the errors studied in faulty designs.
Marleson Graf, Gabriel A. G. Andrade, Luiz Cláudio Villar dos Santos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 A Reinforcement Learning Approach to Directed Test Generation for Shared Memory Verification
abstract
Multicore chips are expected to rely on coherent shared memory. Albeit the coherence hardware can scale gracefully, the protocol state space grows exponentially with core count. That is why design verification requires directed test generation (DTG) for dynamic coverage control under the tight time constraints resulting from slow simulation and short verification budgets. Next generation EDA tools are expected to exploit Machine Learning for reaching high coverage in less time. We propose a technique that addresses DTG as a decision process and tries to find a decision-making policy for maximizing the cumulative coverage, as a result of successive actions taken by an agent. Instead of simply relying on learning, our technique builds upon the legacy from constrained random test generation (RTG). It casts DTG as coverage-driven RTG, and it explores distinct RTG engines subject to progressively tighter constraints. We compared three Reinforcement Learning generators with a state-of-the-art generator based on Genetic Programming. The experimental results show that the proper enforcement of constraints is more efficient for guiding learning towards higher coverage than simply letting the generator learn how to select the most promising memory events for increasing coverage. For a 3-level MESI 32-core design, the proposed approach led to the highest observed coverage (95.81%), and it was 2.4 times faster than the baseline generator to reach the latter's maximal coverage.
Nícolas Pfeifer, Bruno V. Zimpel, Gabriel A. G. Andrade, Luiz Cláudio Villar dos Santos
DATE4
2020 A Directed Test Generator for Shared-Memory Verification of Multicore Chip Designs
abstract
The functional verification of multicore chips requires the generation of parallel test programs able to expose design errors and ensure high coverage in less time. Albeit the coherence hardware can scale gracefully as the number of cores grows, the state space of the coherence protocol increases exponentially. That is why this article describes a directed test generation approach that exploits random test generation (RTG) for avoiding explicit enumeration of the coherence state space while memory consistency is verified. The novel approach was designed for synergy between a data-driven engine that explores neighborhoods toward higher coverage and a model-based engine that exploits constraints while driving RTG toward faster coverage evolution. As compared to a state-of-the-art data-driven generator and to a model-based generator, the proposed approach led to superior coverage evolution with time, when targeting 32-core designs relying on different protocols. For MOESI 2-level, the novel approach was from 4.8 to 18.7 faster to reach the data-driven generator's maximal coverage, and it was up to 2.7 faster to reach the model-driven generator's. For MESI 3-level, it found, in 10 to 15 min, a few errors whose detection required the data-driven generator 45 min to 7 h.
Gabriel A. G. Andrade, Marleson Graf, Nícolas Pfeifer, Luiz Cláudio Villar dos Santos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 Chaining and Biasing: Test Generation Techniques for Shared-Memory Verification
abstract
Since nondeterministic behavior is key to exposing shared-memory errors, nonsynchronized parallel programs are often used for verification and test of multicore chips. In the verification phase, however, the slow execution in a simulator requires nonconventional constraints for enabling error exposure with shorter programs. This paper proposes two novel techniques that build upon conventional random test generation for efficient shared-memory verification. The first technique exploits canonical dependence chains for constraining the random generation of instruction sequences so that the races induced at runtime are likely to raise the coverage of state transitions due to memory events conflicting at a same shared location. The second one exploits address space constraints for biasing random address assignment so that the competition of distinct shared locations for a same cache set can be controlled for raising the coverage of state transitions due to eviction events. We built generators relying on each of the proposed techniques, as well as on their combination, and we compared them to a conventional constrained random test generator for 8, 16, and 32-core architectures. Each of the four generators synthesized 1200 distinct test programs for verifying ten faulty designs derived from each of the three architectures (144 000 verification runs in total). For 32-core designs, the combination of the proposed techniques made at least 50% of the generation space capable of exposing errors, improved the median functional coverage by 44% and 83% at the two highest hierarchical levels, and reduced the average verification effort by one order of magnitude in many cases.
Gabriel A. G. Andrade, Marleson Graf, Luiz Cláudio Villar dos Santos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Spec&Check: An Approach to the Building of Shared-Memory Runtime Checkers for Multicore Chip Design Verification
abstract
Multicore architectures are likely to largely relax sequential consistency constraints on store atomicity and on the ordering between loads and stores, while preserving a coherent shared-memory abstraction. As a result, multicore chip design verification is challenged by the higher number of valid execution witnesses resulting from consistency relaxation and by the larger coherence protocol's state space induced by growing core counts. On the one hand, litmus test generation is effective in exposing consistency bugs, but their coverage of coherence events is limited. On the other hand, random test generation (RTG) leads to higher coverage, but requires specialized checkers and higher observability for detecting subtle consistency errors. Unfortunately, under RTG, no reported checker is able to handle the non-multiple-copy atomic (nMCA) stores arising from full relaxation. This paper proposes an approach that bridges that gap. It relies on an nMCA-compliant abstract specification for shared memory behavior and on an observability template for guiding the insertion of proper monitors into the design representation. We compared a checker built under the novel approach with a conventional one, when running the same test suites, each built with many programs of fixed size. For 4K-instruction programs, the conventional checker raised false positives for 1/3 of the test suites when targeting correct 32-core nMCA designs, whereas the new checker raised none. The improved verification quality resulting from the more general specification of memory behavior came at the expense of negligible overhead, and often led to effort reduction.
Marleson Graf, Olav P. Henschel, Rafael P. Alevato, Luiz Cláudio Villar dos Santos
ICCAD4
2018 Steep coverage-ascent directed test generation for shared-memory verification of multicore chips
abstract
This paper proposes a framework for functional verification of shared memory that relies on reusable coverage-driven directed test generation. It reveals a new mechanism to improve the quality of non-deterministic tests. The generator exploits general properties of coherence protocols and cache memories for better control on transition coverage, which serves as a proxy for increasing the actual coverage metric adopted in a given verification environment. Being independent of coverage metric, coherence protocol, and cache parameters, the proposed generator is reusable across quite different designs and verification environments. We report the coverage for 8, 16, and 32-core designs and the effort required for exposing nine different types of errors. The proposed technique was always able to reach similar coverage as a state-of-the-art generator, and it always did it faster above a certain threshold. For instance, when executing tests with 1K operations for verifying 32-core designs, the former reached 65% coverage around 5 times faster than the latter. Besides, we identified challenging errors that could hardly be found by the latter within one hour, but were exposed by our technique in 5 to 30 minutes.
Gabriel A. G. Andrade, Marleson Graf, Nícolas Pfeifer, Luiz Cláudio Villar dos Santos
ICCAD4
2017 Incremental Layer Assignment Driven by an External Signoff Timing Engine
abstract
Modern technologies provide wide and thick metal layers that must be wisely used to reduce the delay of critical interconnections. After global routing, incremental layer assignment can improve the circuit timing by properly selecting critical interconnect segments to be routed in the faster (but very limited) wires on upper layers. Existing techniques based on net-by-net iterative improvement may get stuck at locally-optimal solutions depending on net ordering. Recent techniques rule out such drawback through the simultaneous iterative improvement of all nets, but they unfortunately rely on objective functions that may guide the optimization off critical paths. As opposed to all reported techniques, which rely on simplified, overly pessimistic timing models, this paper proposes the decoupling of incremental layer assignment from the timing analysis and the exploitation of flow conservation conditions so as to enable the use of an external signoff timing engine. The novel technique was experimentally compared with two state-of-the art works, leading to 50% less timing violations under total negative slack metric and 35% less timing violations under worst negative slack metric with similar overhead in number of vias.
Vinicius S. Livramento, Derong Liu 0002, Salim Chowdhury, Bei Yu 0001, David Z. Pan, José Luís Güntzel, Luiz Cláudio Villar dos Santos
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2016 Chain-based pseudorandom tests for pre-silicon verification of CMP memory systems
abstract
The coherent shared-memory abstraction is expected to keep its crucial role in Chip Multiprocessors even on the scale of hundreds of cores. As a result, the growing hardware complexity to support such abstraction makes the design of the memory system proner to error. Therefore, it is crucial to check for errors in shared-memory behavior as early as possible in the design flow. Given a design representation of a memory subsystem, this paper addresses the pre-silicon verification of its expected behavior, which is captured by the axioms that specify the coherence and consistency requirements of a memory model. As opposed to typical pseudorandom generation of test programs, this paper proposes the exploitation of significant operation orderings from aggressive memory model specifications for inducing load/store sequences that are more effective in uncovering design errors. The effectiveness of the novel technique was evaluated, for 8, 16, and 32-core architectures, when synthesizing 1200 distinct test programs for verifying 8 derivative designs containing errors (9600 use cases per architecture). The synthesized tests explored 5 program sizes, 4 levels of sharing, 4 instruction mixes, and 15 random seeds. Our results show that, as compared to typical pseudorandom generation, the proposed technique is more effective in exposing design errors whose sharing-level distributions are flat. For such design errors, our technique was more effective by 30% on average.
Gabriel A. G. Andrade, Marleson Graf, Luiz Cláudio Villar dos Santos
ICCD3
2016 Clock-Tree-Aware Incremental Timing-Driven Placement
abstract
The increasing impact of interconnections on overall circuit performance makes timing-driven placement (TDP) a crucial step toward timing closure. Current TDP techniques improve critical paths but overlook the impact of register placement on clock tree quality. On the other hand, register placement techniques found in the literature mainly focus on power consumption, disregarding timing and routabilty. Indeed, postponing register placement may undermine the optimization achieved by TDP, since the wiring between sequential and combinational elements would be touched. This work proposes a new approach for an effective coupling between register placement and TDP that relies on two key aspects to handle sequential and combinational elements separately: only the registers in the critical paths are touched by TDP (in practice they represent a small percentage of the total number of registers), and the shortening of clock tree wirelength can be obtained with limited variation in signal wirelength and placement density. The approach consists of two steps: (1) incremental register placement guided by a virtual clock tree to reduce clock wiring capacitance while preserving signal wirelength and density, and (2) incremental TDP to minimize the total negative slack. For the first step, we propose a novel technique that combines clock-net contraction and register clustering forces to reduce the clock wirelength. For the second step, we propose a novel Lagrangian Relaxation formulation that minimizes total negative slack for both setup and hold timing violations. To solve the formulation, we propose a TDP technique using a novel discrete search that employs a Euclidean distance to define a proper neighborhood. For the experimental evaluation of the proposed approach, we relied on the ICCAD 2014 TDP contest infrastructure and compared our results with the best results obtained from that contest in terms of timing closure, clock tree compactness, signal wirelength, and density. Assuming a long displacement constraint, our technique achieves worst and total negative slack reductions of around 24% and 26%, respectively. In addition, our approach leads to 44% shorter clock tree wirelength with negligible impact on signal wirelength and placement density. In the face of such results, the proposed coupling seems a useful approach to handle the challenges faced by contemporary physical synthesis.
Vinicius S. Livramento, Renan Netto, Chrystian Guth, José Luís Güntzel, Luiz Cláudio Villar dos Santos
ACM Trans. Design Autom. Electr. Syst.5
2015 Exploiting Non-Critical Steiner Tree Branches for Post-Placement Timing Optimization
abstract
The increasing impact of interconnections on the overall circuit performance renders physical design a crucial step to timing closure. Several techniques are used to optimize timing within the flow, such as gate sizing, buffer insertion, and timing-driven placement (TDP). Unfortunately, gate sizing and buffer insertion are not capable of modifying the length of interconnections. Although TDP is able to shorten critical interconnection by finding new legal locations for a subset of cells, it generally overlooks the impact of non-critical branches on the delay of critical cells. This work proposes a post-placement timing optimization technique to reduce the capacitive load of critical cells by shortening non-critical Steiner tree branches. To shorten such branches, our technique uses computational geometry for finding effective cell movements that consider maximum displacement constraints and macro blocks. Our experiments evaluate the capability of our technique to further reduce the timing violations from a TDP solution. We applied our technique on the solutions obtained by the top 3 teams in the ICCAD 2014 TDP Contest, where short and long displacement constraints are defined. For the short constraints, the average reductions assuming worst and total late negative slack metrics are 23% and 34%. Considering the long constraints, the average reductions are 62% and 67%. We also present extensions of our technique to tackle related physical design problems such as early violations reduction and electrical correction.
Vinicius S. Livramento, Chrystian Guth, Renan Netto, José Luís Güntzel, Luiz Cláudio Villar dos Santos
ICCAD5
2015 Timing-Driven Placement Based on Dynamic Net-Weighting for Efficient Slack Histogram Compression
abstract
Timing-driven placement (TDP) finds new legal locations for standard cells so as to minimize timing violations while preserving placement quality. Although violations may arise from unmet setup or hold constraints, most TDP approaches ignore the latter. Besides, most techniques focus on reducing the worst negative slack and let the improvements on total negative slack as a secondary goal. However, to successfully achieve timing closure, techniques must also reduce the total negative slack, which is known as slack histogram compression. This paper proposes a new Lagrangian Relaxation formulation for TDP to compress both late and early slack histograms. To solve the problem, we employ a discrete local search technique that uses the Lagrange multipliers as net-weights, which are dynamically updated using an accurate timing analyzer. To preserve placement quality, our technique uses a small fixed-size window that is anchored in the initial location of a cell. For the experimental evaluation of the proposed technique, we relied on the ICCAD 2014 TDP contest infrastructure. The results show that our technique significantly reduces the timing violations from an initial global placement. On average, late and early total negative slacks are improved by 85.03% and 42.72%, respectively, while the worst slacks are reduced by 71.55% and 34.40%. The overhead in wirelength is less than 0.1%.
Chrystian Guth, Vinicius S. Livramento, Renan Netto, Renan Fonseca, José Luís Güntzel, Luiz Cláudio Villar dos Santos
ISPD6
2013 Reconciling real-time guarantees and energy efficiency through unlocked-cache prefetching
abstract
For real-time tasks, cache behavior must be constrained via cache locking or predicted by WCET analysis. Since the former gives up energy efficiency for predictability, this paper proposes a novel code optimization that reduces the miss rate of unlocked instruction caches and, provenly, does not increase the WCET. We optimized the 37 programs from the Mälardalen WCET benchmark for 36 cache configurations and two technologies. By exploiting software prefetching on top of on-demand fetching, we reduced the memory's contribution to the energy consumption (by 11.2%), to the average case execution time (by 10.2%), and to the WCET (by 17.4%).
Emilio Wuerges, Rômulo Silva de Oliveira, Luiz Cláudio Villar dos Santos
DAC3
2013 On-the-fly verification of memory consistency with concurrent relaxed scoreboards
abstract
Parallel programming requires the definition of shared-memory semantics by means of a consistency model, which affects how the parallel hardware is designed. Therefore, verifying the hardware compliance with a consistency model is a relevant problem, whose complexity depends on the observability of memory events. Post-silicon checkers analyze a single sequence of events per core and so do most pre-silicon checkers, although one reported method samples two sequences per core. Besides, most are post-mortem checkers requiring the whole sequence of events to be available prior to verification. On the contrary, this paper describes a novel on-the-fly technique for verifying memory consistency from an executable representation of a multicore system. To increase efficiency without hampering verification guarantees, three points are monitored per core. The sampling points are selected to be largely independent from the core's microarchitecture. The technique relies on concurrent relaxed scoreboards to check for consistency violations in each core. To check for global violations, it employs a linear order of events induced by a given test case. We prove that the technique neither indicates false negatives nor false positives when the test case exposes an error that affects the sampled sequences, making it the first on-the-fly checker with full guarantees. We compare our technique with two post-mortem checkers under 2400 scenarios for platforms with 2 to 8 cores. The results show that our technique is at least 100 times faster than a checker sampling a single sequence per processor and it needs approximately 1/4 to 3/4 of the overall verification effort required by a post-mortem checker sampling two sequences per processor.
Leandro S. Freitas, Eberle A. Rambo, Luiz Cláudio Villar dos Santos
DATE3
2012 On ESL verification of memory consistency for system-on-chip multiprocessing
abstract
Chip multiprocessing is key to Mobile and high-end Embedded Computing. It requires sophisticated multilevel hierarchies where private and shared caches coexist. It relies on hardware support to implicitly manage relaxed program order and write atomicity so as to provide well-defined shared-memory semantics (captured by the axioms of a memory consistency model) at the hardware-software interface. This paper addresses the problem of checking if an executable representation of the memory system complies with a specified consistency model. Conventional verification techniques encode the axioms as edges of a single directed graph, infer extra edges from memory traces, and indicate an error when a cycle is detected. Unlike them, we propose a novel technique that decomposes the verification problem into multiple instances of an extended bipartite graph matching problem. Since the decomposition was judiciously designed to induce independent instances, the target problem can be solved by a parallel verification algorithm. Our technique, which is proven to be complete for several memory consistency models, outperformed a conventional checker for a suite of 2400 randomly-generated use cases. On average, it found a higher percentage of faults (90%) as compared to that checker (69%) and did it, on average, 272 times faster.
Eberle A. Rambo, Olav P. Henschel, Luiz Cláudio Villar dos Santos
DATE3
2012 Efficient verification of out-of-order behaviors with relaxed scoreboards
abstract
Microarchitectures often relax order constraints to meet performance requirements. However, the design of a module handling out-of-order behaviors is error prone, since order relaxation asks for sophisticated control. Besides, its functional verification is challenging, because the module does not preserve at its output the order corresponding to its input data, violating a basic assumption of conventional scoreboards. This paper discusses the verification guarantees of three classes of dynamic checkers and experimentally compares their effectiveness and effort. Results show that a well-designed relaxed scoreboard can achieve the same effectiveness as a complete post-mortem checker with an effort similar to a conventional scoreboard's.
Leandro S. Freitas, Gabriel A. G. Andrade, Luiz Cláudio Villar dos Santos
ICCD3
2009 A novel verification technique to uncover out-of-order DUV behaviors
abstract
Post-partitioning verification has to deal with abstract data, implementation artifacts, and the order of events may not be preserved in the DUV due to the concurrency treatment in the golden model. Existing techniques are limited either by the use of greedy heuristics (jeopardizing verification guarantees) or by black-box approaches (impairing observability). This work proposes a novel white-box technique that overcomes those limitations by casting the problem as an extended bipartite graph matching. By relying on proven properties, solid verification guarantees are provided. Experimental validation was performed upon platforms built around contemporary real-life applications.
Gabriel Marcilio, Luiz Cláudio Villar dos Santos, Bruno de Carvalho Albertini, Sandro Rigo
DAC2
2009 A Multi-Model Engine for High-Level Power Estimation Accuracy Optimization
abstract
Register transfer level (RTL) power macromodeling is a mature research topic with a variety of equation and table-based approaches. Despite its maturity, macromodeling is not yet widely accepted as a de facto industrial standard for power estimation at the RT level. Each approach has many variants depending upon the parameters chosen to capture power variation. Every macromodeling technique has some intrinsic limitation affecting either its performance or its accuracy. Therefore, alternative macromodeling methods can be envisaged as part of a power modeling toolkit from which multiple models for a given component could be exploited so as to reduce the estimation errors resulting from conventional single-model approaches. This paper describes two different approaches for a new multi-model power estimation engine. The first one selects the macromodeling technique that leads to the least estimation error, for a given system component, depending on the properties of its input-vector stream. A proper selection function is built after component characterization and used during estimation. Though simple, this approach has revealed a substantial improvement in estimation accuracy. The second one builds a power estimate function that captures the correlation between individual macromodel estimates and input-stream properties. Experimental results show that our multi-model engine improves the robustness of power analysis with negligible usage overhead. Accuracy becomes seven times better on average, as compared to conventional single-model estimators, while the overall maximum estimation error is divided by 9.
Felipe Klein, Roberto Leao, Guido Araujo, Luiz Cláudio Villar dos Santos, Rodolfo Azevedo
IEEE Trans. Very Large Scale Integr. Syst.4
2008 An open-source binary utility generator
abstract
Electronic system level (ESL) modeling allows early hardware-dependent software (HDS) development. Due to broad CPU diversity and shrinking time-to-market, HDS development can neither rely on hand-retargeting binary tools, nor can it rely on pre-existent tools within standard packages. As a consequence, binary utilities which can be easily adapted to new CPU targets are of increasing interest. We present in this article a framework for automatic generation of binary utilities. It relies on two innovative ideas: platform-aware modeling and more inclusive relocation handling. Generated assemblers, linkers, disassemblers and debuggers were validated for MIPS, SPARC, PowerPC, i8051 and PIC16F84. An open-source prototype generator is available for download.
Alexandro Baldassin, Paulo Centoducatte, Sandro Rigo, Daniel C. Casarotto, Luiz Cláudio Villar dos Santos, Max R. de O. Schultz, Olinto J. V. Furtado
ACM Trans. Design Autom. Electr. Syst.5
2007 A multi-model power estimation engine for accuracy optimization
abstract
RTL power macromodeling is a mature research topic with a variety of equation and table-based approaches. Despite its maturity, macromodeling is not yet widely accepted as an industrial de facto standard for power estimation at the RT level. Each approach has many variants depending upon the parameters chosen to capture power variation. Every macromodeling technique has some intrinsic limitation affecting either its performance or its accuracy. Therefore, alternative macromodeling methods can be envisaged as part of a power modeling toolkit from which the most suitable method for a given component should be automatically selected. Thispaper describes a new multi-model power estimation engine that selects the macromodeling technique leading to the least estimation error for a given system component depending on the properties of its input-vector stream. A proper selection function is built after component characterization and used during estimation. Experimental results show that our multi-model engine improves the robustness of power analysis with negligible usage overhead. Accuracy becomes 3 times better on average, as compared to conventional single-model estimators, while the overall maximum estimation error is divided by 8.
Felipe Klein, Guido Araujo, Rodolfo Azevedo, Roberto Leao, Luiz Cláudio Villar dos Santos
ISLPED5
2000 A code-motion pruning technique for global scheduling
abstract
In the high-level synthesis of ASICs or in the code generation for ASIPs, the presence of conditionals in the behavioral description represents an obstacle to exploit parallelism. Most existing methods use greedy choices in such a way that the search space is limited by the applied heuristics. For example, they might miss opportunities to optimize across basic block boundaries when treating conditional execution. We propose a constructive method which allows generalized code motions. Scheduling and code motion are encoded in the form of a unified resource-constrained optimization problem. In our approach many alternative solutions are constructed and explored by a search algorithm, while optimal solutions are kept in the search space. Our method can cope with issues like speculative execution and code such duplication. Moreover, it can tackle constraints imposed by the advance choice of a controller, such as pipelined-control delay and limited branch capabilities. The underlying timing models support chaining and multicycling. As tasking code motion into account may lead to a larger search space, a code-motion pruning technique is presented. This pruning is proven to keep optimal solutions in the search space for cost functions in terms of schedule lengths.
Luiz Cláudio Villar dos Santos, Marc J. M. Heijligers, C. A. J. van Eijk, J. Van Eijnhoven, Jochen A. G. Jess
ACM Trans. Design Autom. Electr. Syst.1
1999 A Reordering Technique for Efficient Code Motion
abstract
Article Free Access Share on A reordering technique for efficient code motion Authors: Luiz C. V. dos Santos Design Automation Section, Eindhoven University of Technology, P.O. Box 513, 5600 MB Eindhoven, The Netherlands Design Automation Section, Eindhoven University of Technology, P.O. Box 513, 5600 MB Eindhoven, The NetherlandsView Profile , Jochen A. G. Jess Design Automation Section, Eindhoven University of Technology, P.O. Box 513, 5600 MB Eindhoven, The Netherlands Design Automation Section, Eindhoven University of Technology, P.O. Box 513, 5600 MB Eindhoven, The NetherlandsView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 296–299https://doi.org/10.1145/309847.309935Published:01 June 1999Publication History 17citation240DownloadsMetricsTotal Citations17Total Downloads240Last 12 Months15Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Luiz Cláudio Villar dos Santos, Jochen A. G. Jess
DAC1
1999 Exploiting State Equivalence on the Fly while Applying Code Motion and Speculation
abstract
Emerging design problems are prompting the use of code motion and speculation in high-level synthesis to shorten schedules and meet tight time-constraints. Unfortunately, they may increase the number of states to an extent not always affordable for embedded systems. We propose a new technique that not only leads to less states, but also speeds up scheduling. Equivalent states are predicted and merged while building the finite state machine. Experiments indicate that flexible code motions can be used, since our technique restrains state expansion.
Luiz Cláudio Villar dos Santos, Jochen A. G. Jess
DATE1