Margarida F. Jacome

dblp:75/1137 · DBLP profile ↗
← Back
34ranked-venue papers
9as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 9 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
15 papers
Processor architecture and microarchitecture · 32% Embedded and real-time systems · 16% Electronic design automation · 12%
Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 89% Debugging and program repair · 11%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
instruction scheduling
0.132005
Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
High-Quality Operation Binding for Clustered VLIW Datapaths · DAC 2001
Clustered VLIW Architectures with Predicated Switching · DAC 2001
Electronic design automation
high-level synthesis
0.112007
Defect-Aware High-Level Synthesis Targeted at Reconfigurable Nanofabrics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Hardware reliability and fault tolerance
defect tolerance
0.122007
Defect tolerant probabilistic design paradigm for nanotechnologies · DAC 2004
Defect-Aware High-Level Synthesis Targeted at Reconfigurable Nanofabrics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Processor architecture and microarchitecture
instruction-level parallelism
0.122005
Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Clustered VLIW Architectures with Predicated Switching · DAC 2001
Processor architecture and microarchitecture
speculation
0.122005
Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Clustered VLIW Architectures with Predicated Switching · DAC 2001
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
clustered VLIW
0.122001
High-Quality Operation Binding for Clustered VLIW Datapaths · DAC 2001
Clustered VLIW Architectures with Predicated Switching · DAC 2001
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.122001
High-Quality Operation Binding for Clustered VLIW Datapaths · DAC 2001
Clustered VLIW Architectures with Predicated Switching · DAC 2001
Compilers and program optimization
predicated execution
0.112005
Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Processor architecture and microarchitecture › instruction-level parallelism
compiler-controlled speculative execution
0.112005
Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Memory systems › on-chip memory › embedded memory
embedded memory architecture
0.112005
Xtream-fit: an energy-delay efficient data memory subsystem for embedded media processing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Hardware reliability and fault tolerance
transient fault tolerance
0.112005
High performance computing on fault-prone nanotechnologies: novel microarchitecture techniques exploiting reliability-delay trade-offs · DAC 2005
Emerging computing paradigms
nanotechnology
0.012004
Defect tolerant probabilistic design paradigm for nanotechnologies · DAC 2004
Embedded and real-time systems › embedded software
embedded software performance analysis
0.012003
Embedded Architect: A Tool for Early Performance Evaluation of Embedded Software · ICSE 2003
Memory systems › memory architecture
embedded system memory
0.012003
Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing · DAC 2003
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff
0.012003
Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing · DAC 2003
Memory systems › on-chip memory
scratchpad memory
0.012003
Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing · DAC 2003
Compilers and program optimization
code size reduction
0.012002
RS-FDRA: A register-sensitive software pipelining algorithm for embedded VLIW processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002
Compilers and program optimization › instruction scheduling
software pipelining
0.012002
RS-FDRA: A register-sensitive software pipelining algorithm for embedded VLIW processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002
Embedded and real-time systems › embedded processor
embedded processor design
0.012002
Application-specific clustered VLIW datapaths: early exploration on a parameterized design space · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
VLIW processor
0.012002
Application-specific clustered VLIW datapaths: early exploration on a parameterized design space · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002
Debugging and program repair › fault localization
predicate switching
0.012001
Clustered VLIW Architectures with Predicated Switching · DAC 2001
Electronic design automation
yield analysis
0.012007
Defect-Aware High-Level Synthesis Targeted at Reconfigurable Nanofabrics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Performance modeling and evaluation
probabilistic performance analysis
0.011998
Hierarchical Algorithms for Assessing Probabilistic Constraints on System Performance · DAC 1998
Processor architecture and microarchitecture › instruction set architecture
EPIC architecture
0.012005
Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Emerging computing paradigms › nanotechnology
nanoscale computing
0.012005
High performance computing on fault-prone nanotechnologies: novel microarchitecture techniques exploiting reliability-delay trade-offs · DAC 2005
Hardware reliability and fault tolerance › defect tolerance
defect-tolerant design
0.012004
Defect tolerant probabilistic design paradigm for nanotechnologies · DAC 2004
Electronic design automation
design methodology
0.012004
Defect tolerant probabilistic design paradigm for nanotechnologies · DAC 2004
Embedded and real-time systems › embedded system design
component-based embedded systems
0.012003
Architecture-level performance evaluation of component-based embedded systems · DAC 2003
Electronic design automation
design space exploration
0.012003
Embedded Architect: A Tool for Early Performance Evaluation of Embedded Software · ICSE 2003
Embedded and real-time systems
multimedia processing
0.012003
Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing · DAC 2003

Methods — techniques the papers use, named apart from their topics

predicated switching · 0.2static single assignment · 0.1task-based execution model · 0.1prefetching · 0.1dynamic energy conservation · 0.1design space exploration · 0.1static performance evaluation · 0.1probabilistic design space exploration · 0.1dynamic programming · 0.1static speculation algorithm · 0.1speculative execution · 0.1pareto optimization · 0.0force-directed retiming · 0.0compiler transformation · 0.0
YearPublicationVenuePosition
2009 Compiler Controlled Speculation for Power Aware ILP Extraction in Dataflow Architectures
Muhammad Umar Farooq 0003, Lizy Kurian John, Margarida F. Jacome
HiPEAC3
2007 Global Optimization of Compositional Systems
abstract
Embedded systems typically consist of a composition of a set of hardware and software IP modules. Each module is heavily optimized by itself. However, when these modules are composed together, significant additional opportunities for optimizations are introduced because only a subset of the entire functionality is actually used. We propose COSE-a technique to jointly optimize such designs. We use symbolic execution to compute invariants in each component of the design. We propagate these invariants as constraints to other modules using global flow analysis of the composition of the design. This captures optimizations that go beyond, and are qualitatively different than, those achievable by compiler optimization techniques such as common subexpression elimination, which are localized. We again employ static analysis techniques to perform optimizations subject to these constraints. We implemented COSE in the Metropolis platform and achieved significant optimizations using reasonable computational resources.
Fadi A. Zaraket, John Pape, Adnan Aziz, Margarida F. Jacome, Sarfraz Khurshid
FMCAD4
2007 An RFID-Based Platform Supporting Context-Aware Computing in Complex Spaces
abstract
Ubiquitous computing promises to both assist us in everyday tasks and enhance our capabilities. Key elements towards fulfilling this goal are exploiting the physical and logical context in which computation occurs, in order to scope the interaction between users and applications. In this paper we describe an RFID-based platform allowing mobile entities to transparently associate with ubiquitous applications running within complex physical spaces. Entity- application associations occur only for applications within a physical space whose services are within predefined sets specified by an entity for that space type. Mobile entities in our platform are uniquely identified by a temporary ID, and further characterized by a set of attributes describing the above-mentioned set of services. Computation is mediated through the exchange of protocol messages guarded by such attributes. Furthermore, relevant application state is distributed on each mobile entity through a set of messaging boards, enabling a targeted form of communication and cooperation among ubiquitous applications. In this paper we report on our experience experimenting with this platform. Our initial results indicate that this platform is suitable for current RFID technology and exhibits low-cost, scalability and privacy.
Ayis Ziotopoulos, Margarida F. Jacome, Gustavo de Veciana
MDM2
2007 Self-Imposed Temporal Redundancy: An Efficient Technique to Enhance the Reliability of Pipelined Functional Units
abstract
Temporal redundancy (TR) improves the reliability of computational functional units (FUs). However, it can guarantee detection of transient errors only, and may have a substantial power and area overhead. In this paper we present self-imposed temporal redundancy (SITR), a form of TR that can be applied to pipelined FUs and does not suffer from the aforementioned problems. A SITR-enhanced FU forces redundant computations to fire in consecutive cycles and requires a single additional cycle for the second computation and the comparison of the two results. We evaluate the power and area overhead of SITR and conclude that is always smaller than that of standard TR and that it does not depend on the FU complexity. We also use SITR to improve the reliability of the execution datapath of a simple out-of-order engine, typical of that used in high reliability embedded systems and future many-core architectures. Our simulations show that SITR outperforms TR, especially in FP applications. When the number of integer ALUs is larger than the machine width, the performance penalty of SITR is consistently less than 10%.
Elias Mizan, Tileli Amimeur, Margarida F. Jacome
SBAC-PAD3
2007 Defect-Aware High-Level Synthesis Targeted at Reconfigurable Nanofabrics
abstract
Entering the nanometer era, a major challenge to current design methodologies and tools is how to effectively address the high defect densities projected for nanoelectronic technologies. To this end, a reconfiguration-based defect-avoidance methodology for defect-prone nanofabrics was proposed. It judiciously architects the nanofabric, using probabilistic considerations, such that a very large number of alternative implementations can be mapped into it, enabling defects to be circumvented at configuration time, in a scalable way. Building on this foundation, in this paper, a synthesis framework aimed at implementing this new design paradigm is proposed. A key novelty of the approach with respect to traditional high-level synthesis (HLS) is that, rather than carefully optimizing a single (“deterministic”) solution, the goal is to simultaneously synthesize a large family of alternative solutions, so as to meet the required probability of successful configuration, or yield, while maximizing the average performance of the family of synthesized solutions. Experimental results generated for a set of representative benchmark kernels, assuming different defect regimes and target yields, empirically show that the proposed algorithms can effectively explore the complex probabilistic design space associated with this new class of HLS problems.
Margarida F. Jacome
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 RAS-NANO: a reliability-aware synthesis framework for reconfigurable nanofabrics
abstract
Entering the nanometer era, a major challenge to current design methodologies and tools is to effectively address the high defect densities projected for nanotechnologies. To this end, we proposed a reconfiguration-based defect-avoidance methodology for defect-prone nanofabrics. It judiciously architects the nanofabric, using probabilistic considerations, such that a very large number of alternative implementations can be mapped into it, enabling defects to be circumvented at configuration time in a scalable way. Building on this foundation, in this paper we propose a synthesis framework aimed at implementing this new design paradigm. A key novelty of our approach with respect to traditional high level synthesis is that, rather than carefully optimizing a single (`deterministic') solution, our goal is to simultaneously synthesize a large family of alternative solutions, so as to meet the required probability of successful configuration, or yield, while maximizing the family's average performance. Experimental results generated for a set of representative benchmark kernels, assuming different defect regimes and target yields, empirically show that our proposed algorithms can effectively explore the complex probabilistic design space associated with this new class of high level synthesis problems
Margarida F. Jacome
DATE2
2005 High performance computing on fault-prone nanotechnologies: novel microarchitecture techniques exploiting reliability-delay trade-offs
abstract
Device and interconnect fabrics at the nanoscale will have a density of defects and susceptibility to transient faults far exceeding those of current silicon technologies. In this paper we introduce a new performance optimization dimension at the microarchitecture level which can mitigate overheads introduced by fault tolerance. This is achieved by directly exposing reliability versus delay design trade-offs while incorporating novel forms of speculation which use faster but less reliable versions of a microarchitecture's performance critical components. Based on a parameterized microarchitecture, we exhibit the benefits of optimizing these tradeoffs.
Andrey V. Zykov, Elias Mizan, Margarida F. Jacome, Gustavo de Veciana, Ajay Subramanian
DAC3
2005 Predicated switching - optimizing speculation on EPIC machines
abstract
Explicitly parallel instruction computing (EPIC) processors are a very attractive platform for many of today's multimedia and communications applications. In particular, clustered EPIC machines can take aggressive advantage of the available instruction-level parallelism, while maintaining high energy-delay efficiency. However, multicluster machines are more challenging to compile to than centralized machines. In this paper, we propose a novel compiler-directed speculation technique called predicated switching (PS) that can be applied to both centralized and multicluster EPIC machines. The two novel contributions in PS are: 1) a compiler transformation, denoted static single assignment-predicated switching, that leverages required data transfers between clusters for performance gains and 2) a static speculation algorithm to decide which specific kernel operations should actually be speculated in the final code, so as to maximize execution performance on the target processor. Experimental results performed on a representative set of time critical kernels compiled for a number of target machines show that, when compared to "resource-unaware" speculation techniques, PS improves performance with respect to at least one of the baselines in 80% of the cases by up to 38%. Moreover, we show that code size and register pressure are not adversely affected by our technique.
Satish Pillai, Margarida F. Jacome
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 Xtream-fit: an energy-delay efficient data memory subsystem for embedded media processing
abstract
Due to the critical role played by data memory subsystems in the performance and energy efficiency of embedded systems, the design of energy-efficient data memory architectures has received considerable attention in recent years. In this paper, we propose a novel special-purpose data memory subsystem called Xtream-Fit which is aimed at achieving high energy-delay efficiency for streaming media applications. A key novelty of Xtream-Fit is that it exposes a single customization parameter, thus enabling a very simple and yet effective design space exploration methodology. A second key contribution of this paper is the ability to achieve very high energy-delay efficiency through a synergistic combination of: 1) special purpose memory subsystem components, namely, a streaming memory and a scratch-pad memory and (2) a novel task-based execution model that exposes/enhances opportunities for efficient prefetching, and aggressive dynamic energy conservation techniques targeting on-chip and off-chip memory components. Extensive experimental results show that Xtream-Fit reduces the energy-delay product by 22% to 61%, as compared to general-purpose memory subsystems enhanced with state of the art cache decay and SDRAM power-mode control policies.
Margarida F. Jacome
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2004 Defect tolerant probabilistic design paradigm for nanotechnologies
abstract
Recent successes in the development and self-assembly of nanoelectronic devices suggest that the ability to manufacture dense nanofabrics is on the near horizon. However, the tremendous increase in device density of nanoelectronics will be accompanied by a substantial increase in hard and soft faults, posing a major challenge to current design methodologies and tools. In this paper we propose a novel probabilistic design paradigm for defective but reconfigurable nanofabrics. The new design goal is to devise an appropriate structural/behavioral decomposition which improves scalability by constraining the reconfiguration process, while meeting a desired probability of successful instantiation, i.e, yield. Our approach not only addresses the scalability problem in configuring dense nanofabrics subject to defects, but gives a rich framework in which critical trade-offs among performance, yield, and per chip cost can be explored. We present a concrete instance of the approach and show extensive experimental results supporting these claims.
Margarida F. Jacome, Gustavo de Veciana, Stephen Bijansky
DAC1
2003 Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing
abstract
In this paper we propose a novel special-purpose data memory subsystem, called Xtream-Fit, aimed at achieving high energy-delay efficiency for streaming media applications. A key novelty of Xtream-Fit is that it exposes a single customization parameter, thus enabling a very simple and yet effective design space exploration methodology. A second key contribution of this work is the ability to achieve very high energy-delay efficiency through a synergistic combination of: (1) special purpose memory subsystem components, namely, a Streaming Memory and Scratch-Pad Memory; and (2) a novel task-based execution model that exposes/enhances opportunities for efficient prefetching, and aggressive dynamic energy conservation techniques targeting on-chip and off-chip memory components. Extensive experimental results show that Xtream-Fit reduces energy-delay product by 46% to 83%, as compared to general-purpose memory subsystems enhanced with state of the art Cache Decay and SDRAM power mode control policies.
Margarida F. Jacome
DAC2
2003 Architecture-level performance evaluation of component-based embedded systems
abstract
A static performance evaluation technique is proposed to support early, architecture-level design space exploration for component-based embedded systems. The novel contribution is the use of a designer-specified evaluation scenario to identify a characteristic subset of system functionality that serves as a context for a rapid performance evaluation between candidate architectures. Fidelity is demonstrated with a case study that compares performance estimates of several candidate architectures to measurements from respective implementations.
Jeffry T. Russell, Margarida F. Jacome
DAC2
2003 Compiler-Directed ILP Extraction for Clustered VLIW/EPIC Machines: Predication, Speculation and Modulo Scheduling
Satish Pillai, Margarida F. Jacome
DATE2
2003 Embedded Architect: A Tool for Early Performance Evaluation of Embedded Software
abstract
Embedded Architect is a design automation tool that embodies a static performance evaluation technique to support early, architecture-level design space exploration for component-based embedded systems. A static control flow characterization, called an evaluation scenario, is specified based on an incremental refinement of software source code, from which a pseudo-trace of operations is generated in combination with architecture mapping and several component parameters, a software performance metric is estimated The novel contribution is the implementation of a tool that automates specification of an evaluation scenario, which sets the context for a rapid performance evaluation of distinct candidate architectures.
Jeffry T. Russell, Margarida F. Jacome
ICSE2
2003 Special issue on power-aware embedded computing
abstract
article Share on Special issue on power-aware embedded computing Editors: Margarida Jacome View Profile , Francky Catthoor View Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 2Issue 3August 2003 pp 251–254https://doi.org/10.1145/860176.860177Published:01 August 2003Publication History 1citation1,203DownloadsMetricsTotal Citations1Total Downloads1,203Last 12 Months4Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Margarida F. Jacome, Francky Catthoor
ACM Trans. Embed. Comput. Syst.1
2002 Scenario-based software characterization as a contingency to traditional program profiling
abstract
Program profiling is common way to characterize program behavior based on representative input. Some software, especially in embedded systems, cannot be profiled do to lack of tools or problems introduced by instrumentation of the code. As an alternative to traditionally profiling, a static analysis technique is proposed that allows a designer to characterize the flow of control of software.Operating on a flow graph representation of software, the proposed technique assists an expert designer in the specification of one or more representative scenarios. A scenario defines a specific flow of control that corresponds to a typical behavior of the system. The efficacy of the technique is demonstrated with two experiments: a comparison to traditional profiling and application to real embedded operating system software for which traditional profiling is not possible.
Jeffry T. Russell, Margarida F. Jacome
CASES2
2002 RS-FDRA: A register-sensitive software pipelining algorithm for embedded VLIW processors
abstract
The paper proposes a novel software-pipelining algorithm, Register-Sensitive Force-Directed Retiming Algorithm (RS-FDRA), suitable for optimizing compilers targeting embedded very large instruction word processors. The key difference between RS-FDRA and previous approaches is that this algorithm can handle code-size constraints along with latency and resource constraints. This capability enables the exploration of Pareto "optimal" points with respect to code size and performance. RS-FDRA can also minimize the increase in register pressure typically incurred by software pipelining. This ability is critical since the need to insert spill code may result in significant performance degradation. Extensive experimental results are presented demonstrating that the extended set of optimization goals and constraints supported by RS-FDRA enables a thorough compiler-assisted exploration of tradeoffs among performance, code size, and register requirements for time-critical segments of embedded software components.
Cagdas Akturan, Margarida F. Jacome
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 Application-specific clustered VLIW datapaths: early exploration on a parameterized design space
abstract
Specialized clustered very large instruction word (VLIW) processors combined with effective compilation techniques enable aggressive exploitation of the high instruction-level parallelism inherent in many embedded media applications, while unlocking a variety of possible performance/cost tradeoffs. In this work, the authors propose a methodology to support early design space exploration of clustered VLIW datapaths, in the context of a specific target application. They argue that, due to the large size and complexity of the design space, the early design space exploration phase should consider only design space parameters that have a first-order impact on two key physical figures of merit: clock rate and power dissipation. These parameters were found to be: maximum cluster capacity, number of clusters, and bus (interconnect) capacity. Experimental validation of their design space exploration algorithm shows that a thorough exploration of the complex design space can be performed very efficiently in this abstract parameterized design space.
Viktor S. Lapinskii, Margarida F. Jacome, Gustavo de Veciana
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 Cluster assignment for high-performance embedded VLIW processors
abstract
Clustering is an effective method to increase the available parallelism in VLIW datapaths without incurring severe penalties associated with a large number of register file ports. Efficient utilization of a clustered datapath requires careful binding/assignment of operations to clusters. The article proposes a binding algorithm that effectively explores trade-offs between in-cluster operation serialization and delays associated with data transfers between clusters. Extensive experimental evidence is provided showing that the algorithm generates high quality solutions for representative kernels, with up to 33% improvement over a state-of-the-art binding algorithm.
Viktor S. Lapinskii, Margarida F. Jacome, Gustavo de Veciana
ACM Trans. Design Autom. Electr. Syst.2
2001 Clustered VLIW Architectures with Predicated Switching
abstract
In order to meet the high throughput requirements of applications exhibiting high ILP, VLIW ASIPs may increasingly include large numbers of functional units(FUs). Unfortunately, ”switching“ data through register files shared by large numbers of FUs quickly becomes a dominant cost/ performance factor suggesting that clustering smaller number of FUs around local register files may be beneficial even if data transfers are required among clusters. With such machines in mind, we propose a compiler transformation, predicated switching, which enables aggressive speculation while leveraging the penalties associated with inter-cluster communication to achieve gains in performance. Based on representative benchmarks, we demonstrate that this novel technique is particularly suitable for application specific clustered machines aimed at supporting high ILP as compared to state-of-the-art approaches.
Margarida F. Jacome, Gustavo de Veciana, Satish Pillai
DAC1
2001 High-Quality Operation Binding for Clustered VLIW Datapaths
abstract
Clustering is an effective method to increase the available parallelism in VLIW datapaths without incurring severe penalties associated with large number of register file ports. Efficient utilization of a clustered datapath requires careful binding of operations to clusters. The paper proposes a binding algorithm that effectively explores tradeoffs between in-cluster operation serialization and delays associated with data transfers between clusters. Extensive experimental evidence is provided showing that the algorithm generates high quality solutions for basic blocks, with up to 29% improvement over a state-of-the-art advanced binding algorithm.
Viktor S. Lapinskii, Margarida F. Jacome, Gustavo de Veciana
DAC2
2001 CALiBeR: A Software Pipelining Algorithm for Clustered Embedded VLIW Processors
abstract
In this paper, we describe a software pipelining framework, CALiBeR (cluster aware load balancing retiming algorithm), suitable for compilers targeting clustered embedded VLIW processors. CALiBeR can be effectively used by embedded system designers to explore different code optimization alternatives, i.e. it can assist the generation of high-quality customized retiming solutions for desired program memory size and throughput requirements, while minimizing register pressure. An extensive set of experimental results is presented, considering several representative benchmark loop kernels and a wide variety of clustered datapath configurations, demonstrating that our algorithm compares favorably with one of the best state-of-the-art algorithms, achieving up to 50% improvement in performance and up to 47% improvement in register requirements.
Cagdas Akturan, Margarida F. Jacome
ICCAD2
2000 A new technique for estimating lower bounds on latency for high level synthesis
abstract
In this paper we present a novel and fast estimation technique that produces tight latency lower bounds for Data Flow Graphs representing time critical segments of the application of interest. Our proposed technique can be used to compute a tighter earliest scheduling step for nodes (operations) in the Data Flow Graph and thus be used to improve the result quality of any technique requiring the computation of such ASAP values.
Helvio P. Peixoto, Margarida F. Jacome
ACM Great Lakes Symposium on VLSI2
2000 Exploring Performance Tradeoffs for Clustered VLIW ASIPs
abstract
VLIW ASIPs provide an attractive solution for increasingly pervasive real-time multimedia and signal processing embedded applications. In this paper we propose an algorithm to support trade-off exploration during the early phases of the design/specialization of VLIW ASIPs with clustered datapaths. For purposes of an early exploration step, we define a parameterized family of clustered datapaths D(m,n), where m and n denote interconnect capacity and cluster capacity constraints on the family. Given a kernel, the proposed algorithm explores the space of feasible clustered datapaths and returns: a datapath configuration; a binding and scheduling for the operations; and a corresponding estimate for the best achievable latency over the specified family. Moreover, we show how the parameters m and n, as well as a target latency optionally specified by the designer, can be used to effectively explore trade-offs among delay, power/energy, and latency. Extensive empirical evidence is provided showing that the proposed approach is strikingly effective at attacking this complex optimization problem.
Margarida F. Jacome, Gustavo de Veciana, Viktor S. Lapinskii
ICCAD1
2000 Symbolic Binding for Clustered VLIW ASIPs
abstract
The paper proposes a symbolic framework to address the binding problem for embedded VLIW ASIPs. Alternative objective functions as well as trade-offs relevant to the binding phase of code generation for embedded processors are presented and discussed. Experimental results obtained for a number of benchmarks extracted from the literature empirically demonstrate the promise of our approach.
Satish Pillai, Margarida F. Jacome
ICCD2
1999 The Design Space Layer: Supporting Early Design Space Exploration for Core-Based Designs
abstract
A novel library layer, called the "design space layer," is proposed, aimed at supporting both IP-based and traditional "in-house" design methodologies, during early design space exploration. Strategies for effectively pruning the large design spaces characteristic of system-on-a-chip designs, and for transparently retrieving information on cores adequate for implementing the system components, are supported by the proposed layer. The layer is self-documented and highly compartmentalized into hierarchies of classes of design objects, and is thus easily scalable. A design space layer developed for encryption applications is presented and discussed in some detail.
Margarida F. Jacome, Helvio P. Peixoto, Ander Royo, Juan Carlos López 0001
DATE1
1999 Lower bound on latency for VLIW ASIP datapaths
abstract
Traditional lower bound estimates on latency for dataflow graphs assume no data transfer delays. While such approaches can generate tight lower bounds for datapaths with a centralized register file, the results may be uninformative for datapaths with distributed register file structures that are characteristic of VLIW ASIPs (very large instruction word application-specific instruction set processors). In this paper, we propose a latency bound that accounts for such data transfer delays. The novelty of our approach lies in constructing the "window dependency graph" and bounds associated with the problem which capture delay penalties due to operation serialization and/or data moves among distributed register files. Through a set of benchmark examples, we show that the bound is competitive with state-of-the-art approaches. Moreover, our experiments show that the approach can aid an iterative improvement algorithm in determining good functional unit assignments-a key step in code generation for VLIW ASIPs.
Margarida F. Jacome, Gustavo de Veciana
ICCAD1
1998 Hierarchical Algorithms for Assessing Probabilistic Constraints on System Performance
abstract
We propose an algorithm for assessing probabilistic performance constraints for systems including components with uncertain delays. We make a case for designing systems based on a probabilistic relaxation of performance constraints, as this has the potential for resulting in lower silicon area and/or power consumption. We consider a concrete example, an MPEG decoder, for which we discuss modeling and assessment of probabilistic throughput constraints.
Gustavo de Veciana, Margarida F. Jacome, Jian-Huei Guo
DAC2
1998 A Methodology for Task Based Partitioning and Scheduling of Dynamically Reconfigurable Systems
abstract
Taking maximum advantage of dynamic reconfiguration in the implementation of digital systems poses a number of challenging research problems. Specifically, techniques are needed to partition the system behavioral description into segments of computation (or "scheduling units"), and to define a reconfiguration schedule with respect to those units, so as to maximize the performance of the dynamically reconfigurable system, subject to the area constraints of the FPGA. We propose a methodology to: (1) perform a coarse-grained partitioning of the system behavioral description into a set of tasks, (2) determine which sub-set of tasks is to remain resident in the FPGA, and which sub-set is to be non resident, (3) generate a reconfiguration schedule for the non-resident tasks by specifying when such tasks should be loaded on to and erased from the FPGA.
Pedro Merino 0001, Margarida F. Jacome, Juan Carlos López 0001
FCCM2
1998 Software power estimation and optimization for high performance, 32-bit embedded processors
abstract
A software energy estimation model is presented for a family of high performance, integrated, 32-bit embedded RISC processors. This model is significantly less complex than previous models, and yet is demonstrated to accurately predict energy consumption to within 8% with 99% confidence based on physical measurements. Factors such as operating frequency, source/destination registers, and operand values are explored. In view of this model, previously proposed optimizations are evaluated for potential energy savings. We conclude that, for the class of processors under discussion, a good optimizing compiler that minimizes execution time will simultaneously minimize energy consumption.
Jeffry T. Russell, Margarida F. Jacome
ICCD2
1997 Algorithm and architecture-level design space exploration using hierarchical data flows
abstract
Incorporating algorithm and architecture level design space exploration in the early phases of the design process can have a dramatic impact on the area, speed, and power consumption of the resulting systems. This paper proposes a framework for supporting system-level design space exploration and discusses the three fundamental issues involved in effectively supporting such an early design space exploration: definition of an adequate level of abstraction; definition of good fidelity system-level metrics; and definition of mechanisms for automating the exploration process. The first issue, the definition of an adequate level of abstraction is then addressed in detail. Specifically, an algorithm-level model, an architecture-level model, and a set of operations on these models, are proposed, aiming at efficiently supporting an early, aggressive system-level design space exploration. A discussion on work in progress in the other two topics, metrics and automation, concludes the paper.
Helvio P. Peixoto, Margarida F. Jacome
ASAP2
1996 A formal basis for design process planning and management
abstract
In this paper we present a formalism that allows for a complete and general characterization of design disciplines and for a unified representation of arbitrarily complex design processes taking place in the context of these disciplines. This formalism has been used as the basis for the development of several prototype CAD meta-tools that offer effective design process planning and management services.
Margarida F. Jacome, Stephen W. Director
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1994 A formal basis for design process planning and management
Margarida F. Jacome, Stephen W. Director
ICCAD1
1992 Design Process Management for CAD Frameworks
Margarida F. Jacome, Stephen W. Director
DAC1