EDBT 2026 Demo / reviewers in the wild / expert
Rafael Ruiz-Sautua
dblp:52/1714
· DBLP profile ↗
11ranked-venue papers
5as first author
0since 2021 · last 2009
0009-0003-7022-259XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Electronic design automation · 97% Cloud and datacenter computing · 3% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.2 | 3 | 2009 | Frequent-Pattern-Guided Multilevel Decomposition of Behavioral Specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009 Exploiting Bit-Level Delay Calculations to Soften Read-After-Write Dependences in Behavioral Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Bitwise scheduling to balance the computational cost of behavioral specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 2 | 2007 | Exploiting Bit-Level Delay Calculations to Soften Read-After-Write Dependences in Behavioral Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Bitwise scheduling to balance the computational cost of behavioral specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Electronic design automation › design optimization
area optimization |
0.1 | 1 | 2009 | Frequent-Pattern-Guided Multilevel Decomposition of Behavioral Specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009 |
Electronic design automation
logic synthesis |
0.1 | 1 | 2007 | Exploiting Bit-Level Delay Calculations to Soften Read-After-Write Dependences in Behavioral Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Electronic design automation › logic synthesis
datapath optimization |
0.0 | 1 | 2009 | Frequent-Pattern-Guided Multilevel Decomposition of Behavioral Specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009 |
Cloud and datacenter computing
resource allocation |
0.0 | 1 | 2006 | Bitwise scheduling to balance the computational cost of behavioral specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Methods — techniques the papers use, named apart from their topics
frequent-pattern-guided decomposition · 0.1data flow graph analysis · 0.1bit-level delay calculation · 0.1operation fragmentation · 0.1bit-level algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Performance-driven scheduling of behavioural specifications
María C. Molina, Rafael Ruiz-Sautua, Pedro Garcia-Repetto, Jose Manuel Mendias |
Integr. | 2 |
| 2009 | Frequent-Pattern-Guided Multilevel Decomposition of Behavioral SpecificationsabstractConventional high-level synthesis algorithms treat specification operations as atomic elements that are executed in one or several consecutive cycles and over one functional unit. However, in most specifications, there exist different types, representations, and widths of operations that, handled at different decomposition levels, may produce better designs. In this way, most arithmetic operations can be decomposed into smaller operations, applying several arithmetical properties. Different decompositions can be performed in order to improve the performance or reduce the area or power consumption. In this paper, we propose a pattern-based design methodology that is able to treat every operation at its most appropriate decomposition level. It produces reduced datapaths while meeting specified time constraints. In comparison to conventional algorithms, the amount of area saved averages 40%. María C. Molina, Rafael Ruiz-Sautua, Pedro Garcia-Repetto, Román Hermida |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Exploiting Internal Operation Patterns during the High-Level Synthesis of Time-Constrained CircuitsabstractConventional high-level synthesis algorithms treat specification operations as atomic elements that are executed in one or several consecutive cycles and over one functional unit. However, in most specifications there exist different operations, in function of their type, representation, and width that handled at different decomposition levels may produce better designs. In this way, most arithmetic operations can be decomposed into smaller operations applying several arithmetical properties. Different decompositions can be performed, in function of the pursued objective: performance improvement, area reduction, or power consumption reduction. In this paper we propose a pattern-based design methodology able to treat every operation at its most appropriate decomposition level. It produces reduced datapaths while meeting the specified time constraints. In comparison to conventional algorithms the amount of area saved averages 40%. Pedro Garcia-Repetto, María C. Molina, Rafael Ruiz-Sautua, Guillermo Botella Juan |
DSD | 3 |
| 2007 | Area optimization of multi-cycle operators in high-level synthesisabstractConventional high-level synthesis algorithms usually employ multi-cycle operators to reduce the cycle length in order to improve the circuit performance. These operators need several cycles to execute one operation, but the entire functional unit is not used in any cycle. Additionally, the execution of operations over wider multi-cycle operators is unfeasible if their results must be available in a smaller number of cycles than the functional unit delay. This obliges to add new functional resources to the datapath even if multi-cycle operators are idle when the execution of the operation begins. In this paper a new design technique to overcome the restricted reusability of multi-cycle operators is presented. It reduces the area of these functional units allowing their internal reuse when executing one operation. It also expands the possibilities of common hardware sharing as it allows the partial use of multicycle operators to calculate narrower operations faster than the functional unit delay. This technique is applied as an optimization phase at the end of the high-level synthesis process, and can optimize the circuits synthesized by any high-level synthesis tool María C. Molina, Rafael Ruiz-Sautua, Jose Manuel Mendias, Román Hermida |
DATE | 2 |
| 2007 | Exploiting Bit-Level Delay Calculations to Soften Read-After-Write Dependences in Behavioral SynthesisabstractConventional high-level synthesis (HLS) algorithms are very conservative when dealing with read-after-write (RAW) dependences, the execution of one operation is allowed once all its predecessors have been calculated. However, in the execution of arithmetic operations, some bits are required later than others, and some bits are produced earlier than others. This paper proposes a presynthesis optimization algorithm that relaxes RAW dependences, taking advantage of this feature for a more efficient HLS of data flow graphs formed by additions, multiplications, and logic operations. The presented preprocessor analyzes the critical path at bit granularity and splits the arithmetic operations into subword fragments. These fragments become the input to any regular HLS tool to speed up circuit execution times through scheduling in different cycles of the fragments obtained from the same original operation. This way, the execution of one operation may begin before the calculus of its predecessors has been completed. This becomes feasible when the execution of the predecessor has begun in the selected cycle or in a previous one, and even if it will finish in a posterior cycle. The experimental results that were carried out show that implementations obtained from the optimized specification are, on the average, 70% faster, with only slight variations in the data path area. Rafael Ruiz-Sautua, María C. Molina, Jose Manuel Mendias |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Pre-synthesis optimization of multiplications to improve circuit performanceabstractConventional high-level synthesis uses the worst case delay to relate all inputs to all outputs of an operation. This is a very conservative approximation of reality, especially in arithmetic operations (where some bits are required later than others and some bits are produced earlier than others). This paper proposes a pre-synthesis optimization algorithm that takes advantage of this feature for more efficient high-level synthesis of data-flow graphs formed by additions and multiplications. The presented pre-processor analyzes the critical path at bit-granularity and splits the arithmetic operations into sub-words fragments. In particular, some of the specification multiplications are broken up into several smaller multiplications, additions, and other operations of three new types specially defined to reduce the clock cycle duration. These fragments become the input to any regular high-level synthesis tool to speed up circuit execution times. The experimental results carried out show that implementations obtained from the optimized specification are on average 70% faster and in most cases substantial area reductions are also achieved Rafael Ruiz-Sautua, María C. Molina, Jose Manuel Mendias, Román Hermida |
DATE | 1 |
| 2006 | Bitwise scheduling to balance the computational cost of behavioral specificationsabstractConventional scheduling algorithms try to balance the number of operations of every different type executed per cycle. However, in most cases, a uniform distribution is not reachable, and thus, some hardware (HW) waste appears. This situation becomes worse when heterogeneous specifications (those formed by operations with different data formats and widths) are synthesized. Our proposal is an innovative bit-level algorithm able to minimize this HW waste. In order to obtain uniform distributions of the computational cost of operations among cycles, it successively transforms specification operations into sets of smaller ones, which are then scheduled independently. As a consequence, some specification operations may be executed during a set of nonconsecutive cycles, and over several functional units. In combination with allocation algorithms able to guarantee the bit-level reuse of HW resources, our approach produces circuits with substantially smaller area than conventional implementations. Due to the fragmentation of operations, in the proposed implementations, the type, number, and width of HW resources are, in general, independent of the type, number, and width of both specification operations and variables. Additionally, the clock-cycle length is also reduced in most circuits. María C. Molina, Rafael Ruiz-Sautua, Jose Manuel Mendias, Román Hermida |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Arrival time aware scheduling to minimize clock cycle lengthabstractConventional scheduling algorithms usually adjust the clock cycle duration to the execution time of the longest operations. This results in large slack times wasted in those cycles with faster operations. To reduce the wasted times multi-cycle and chaining techniques have been employed. The scheduling algorithm presented in this paper goes one step further. It breaks up some of the specification operations and schedule several data-dependent operation fragments in the same cycle. In consequence, some of the specification operations are executed during several cycles (non-necessarily consecutive ones), and in every execution cycle some result bits are calculated. Thus the execution of one operation may start even if its predecessors have not finished yet. In the experimental results carried out, the proposed algorithm improves circuit performance above 70% on average, with slight increments in the datapath area. Rafael Ruiz-Sautua, María C. Molina, Jose Manuel Mendias, Román Hermida |
ASP-DAC | 1 |
| 2005 | Behavioural Transformation to Improve Circuit Performance in High-Level SynthesisabstractEarly scheduling algorithms usually adjusted the clock cycle duration to the execution time of the slowest operation. This resulted in large slack times wasted in those cycles executing faster operations. To reduce the wasted times multi-cycle and chaining techniques have been employed. While these techniques have produced successful designs, their effectiveness are often limited due to the area increment that may derive from chaining, and the extra latencies that may derive from multicycling. In this paper we present an optimization method that solves the time-constrained scheduling problem by transforming behavioural specifications into new ones whose subsequent synthesis substantially improves circuit performance. Our proposal breaks up some of the specification operations, allowing their execution during several possibly unconsecutive cycles, and also the calculation of several data-dependent operation fragments in the same cycle. To do so, it takes into account the circuit latency and the execution time of every specification operation. The experimental results carried out show that circuits obtained from the optimized specification are on average 60% faster than those synthesized from the original specification, with only slight increments in the circuit area. Rafael Ruiz-Sautua, María C. Molina, Jose Manuel Mendias, Román Hermida |
DATE | 1 |
| 2005 | Performance-driven read-after-write dependencies softening in high-level synthesisabstractConventional scheduling algorithms usually produce schedules whose cycle lengths are influenced by the operations latency. Operations are assigned to one or several consecutive cycles (multicycle operators) and one operation cannot begin until all its predecessors have finished, except for operations chaining. This technique allows the execution in the same cycle of several data-dependent operations, being some bits of the chained operations calculated in parallel. Chaining one operation to its predecessor requires the completion of both in the selected cycle. This paper presents a less restrictive technique based on the softening of the read-after-write dependencies among operations, which allows beginning the execution of one operation before the calculus of its predecessors has been completed. This becomes feasible when the execution of the predecessor has begun in the selected cycle or in a previous one, and even if it finishes in a posterior cycle. This design technique is applied before synthesis to transform behavioural specifications into new ones whose subsequent synthesis substantially improves circuit performance. Rafael Ruiz-Sautua, María C. Molina, Jose Manuel Mendias, Román Hermida |
ICCAD | 1 |
| 2004 | Behavioural Bitwise Scheduling Based on Computational Effort BalancingabstractConventional synthesis algorithms schedule multiple precision specifications by balancing the number of operations of every different type and width executed per cycle. However, totally balanced schedules are not always possible and therefore some hardware waste appears. In this paper a heuristic scheduling algorithm to minimize this hardware waste is presented. It successively transforms specification operations into sets of smaller ones until the most uniform distribution of the computational effort of operations among cycles is reached. In the schedules proposed some operations are executed during a set of non-consecutive cycles. María C. Molina, Rafael Ruiz-Sautua, Jose Manuel Mendias, Román Hermida |
DATE | 2 |