VLDB 2026 Research / reviewers in the wild / expert
Margarida F. Jacome
dblp:75/1137
· DBLP profile ↗
34ranked-venue papers
9as first author
0since 2021 · last 2009
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 9 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
15 papers |
Processor architecture and microarchitecture · 32% Embedded and real-time systems · 16% Electronic design automation · 12% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 89% Debugging and program repair · 11% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
instruction scheduling |
0.1 | 3 | 2005 | Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005 High-Quality Operation Binding for Clustered VLIW Datapaths · DAC 2001 Clustered VLIW Architectures with Predicated Switching · DAC 2001 |
Electronic design automation
high-level synthesis |
0.1 | 1 | 2007 | Defect-Aware High-Level Synthesis Targeted at Reconfigurable Nanofabrics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Hardware reliability and fault tolerance
defect tolerance |
0.1 | 2 | 2007 | Defect tolerant probabilistic design paradigm for nanotechnologies · DAC 2004 Defect-Aware High-Level Synthesis Targeted at Reconfigurable Nanofabrics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 2 | 2005 | Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005 Clustered VLIW Architectures with Predicated Switching · DAC 2001 |
Processor architecture and microarchitecture
speculation |
0.1 | 2 | 2005 | Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005 Clustered VLIW Architectures with Predicated Switching · DAC 2001 |
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
clustered VLIW |
0.1 | 2 | 2001 | High-Quality Operation Binding for Clustered VLIW Datapaths · DAC 2001 Clustered VLIW Architectures with Predicated Switching · DAC 2001 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.1 | 2 | 2001 | High-Quality Operation Binding for Clustered VLIW Datapaths · DAC 2001 Clustered VLIW Architectures with Predicated Switching · DAC 2001 |
Compilers and program optimization
predicated execution |
0.1 | 1 | 2005 | Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005 |
Processor architecture and microarchitecture › instruction-level parallelism
compiler-controlled speculative execution |
0.1 | 1 | 2005 | Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005 |
Memory systems › on-chip memory › embedded memory
embedded memory architecture |
0.1 | 1 | 2005 | Xtream-fit: an energy-delay efficient data memory subsystem for embedded media processing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005 |
Hardware reliability and fault tolerance
transient fault tolerance |
0.1 | 1 | 2005 | High performance computing on fault-prone nanotechnologies: novel microarchitecture techniques exploiting reliability-delay trade-offs · DAC 2005 |
Emerging computing paradigms
nanotechnology |
0.0 | 1 | 2004 | Defect tolerant probabilistic design paradigm for nanotechnologies · DAC 2004 |
Embedded and real-time systems › embedded software
embedded software performance analysis |
0.0 | 1 | 2003 | Embedded Architect: A Tool for Early Performance Evaluation of Embedded Software · ICSE 2003 |
Memory systems › memory architecture
embedded system memory |
0.0 | 1 | 2003 | Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing · DAC 2003 |
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff |
0.0 | 1 | 2003 | Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing · DAC 2003 |
Memory systems › on-chip memory
scratchpad memory |
0.0 | 1 | 2003 | Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing · DAC 2003 |
Compilers and program optimization
code size reduction |
0.0 | 1 | 2002 | RS-FDRA: A register-sensitive software pipelining algorithm for embedded VLIW processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002 |
Compilers and program optimization › instruction scheduling
software pipelining |
0.0 | 1 | 2002 | RS-FDRA: A register-sensitive software pipelining algorithm for embedded VLIW processors · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002 |
Embedded and real-time systems › embedded processor
embedded processor design |
0.0 | 1 | 2002 | Application-specific clustered VLIW datapaths: early exploration on a parameterized design space · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002 |
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
VLIW processor |
0.0 | 1 | 2002 | Application-specific clustered VLIW datapaths: early exploration on a parameterized design space · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2002 |
Debugging and program repair › fault localization
predicate switching |
0.0 | 1 | 2001 | Clustered VLIW Architectures with Predicated Switching · DAC 2001 |
Electronic design automation
yield analysis |
0.0 | 1 | 2007 | Defect-Aware High-Level Synthesis Targeted at Reconfigurable Nanofabrics · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Performance modeling and evaluation
probabilistic performance analysis |
0.0 | 1 | 1998 | Hierarchical Algorithms for Assessing Probabilistic Constraints on System Performance · DAC 1998 |
Processor architecture and microarchitecture › instruction set architecture
EPIC architecture |
0.0 | 1 | 2005 | Predicated switching - optimizing speculation on EPIC machines · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005 |
Emerging computing paradigms › nanotechnology
nanoscale computing |
0.0 | 1 | 2005 | High performance computing on fault-prone nanotechnologies: novel microarchitecture techniques exploiting reliability-delay trade-offs · DAC 2005 |
Hardware reliability and fault tolerance › defect tolerance
defect-tolerant design |
0.0 | 1 | 2004 | Defect tolerant probabilistic design paradigm for nanotechnologies · DAC 2004 |
Electronic design automation
design methodology |
0.0 | 1 | 2004 | Defect tolerant probabilistic design paradigm for nanotechnologies · DAC 2004 |
Embedded and real-time systems › embedded system design
component-based embedded systems |
0.0 | 1 | 2003 | Architecture-level performance evaluation of component-based embedded systems · DAC 2003 |
Electronic design automation
design space exploration |
0.0 | 1 | 2003 | Embedded Architect: A Tool for Early Performance Evaluation of Embedded Software · ICSE 2003 |
Embedded and real-time systems
multimedia processing |
0.0 | 1 | 2003 | Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processing · DAC 2003 |
Methods — techniques the papers use, named apart from their topics
predicated switching · 0.2static single assignment · 0.1task-based execution model · 0.1prefetching · 0.1dynamic energy conservation · 0.1design space exploration · 0.1static performance evaluation · 0.1probabilistic design space exploration · 0.1dynamic programming · 0.1static speculation algorithm · 0.1speculative execution · 0.1pareto optimization · 0.0force-directed retiming · 0.0compiler transformation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Compiler Controlled Speculation for Power Aware ILP Extraction in Dataflow Architectures
Muhammad Umar Farooq 0003, Lizy Kurian John, Margarida F. Jacome |
HiPEAC | 3 |
| 2007 | Global Optimization of Compositional SystemsabstractEmbedded systems typically consist of a composition of a set of hardware and software IP modules. Each module is heavily optimized by itself. However, when these modules are composed together, significant additional opportunities for optimizations are introduced because only a subset of the entire functionality is actually used. We propose COSE-a technique to jointly optimize such designs. We use symbolic execution to compute invariants in each component of the design. We propagate these invariants as constraints to other modules using global flow analysis of the composition of the design. This captures optimizations that go beyond, and are qualitatively different than, those achievable by compiler optimization techniques such as common subexpression elimination, which are localized. We again employ static analysis techniques to perform optimizations subject to these constraints. We implemented COSE in the Metropolis platform and achieved significant optimizations using reasonable computational resources. Fadi A. Zaraket, John Pape, Adnan Aziz, Margarida F. Jacome, Sarfraz Khurshid |
FMCAD | 4 |
| 2007 | An RFID-Based Platform Supporting Context-Aware Computing in Complex SpacesabstractUbiquitous computing promises to both assist us in everyday tasks and enhance our capabilities. Key elements towards fulfilling this goal are exploiting the physical and logical context in which computation occurs, in order to scope the interaction between users and applications. In this paper we describe an RFID-based platform allowing mobile entities to transparently associate with ubiquitous applications running within complex physical spaces. Entity- application associations occur only for applications within a physical space whose services are within predefined sets specified by an entity for that space type. Mobile entities in our platform are uniquely identified by a temporary ID, and further characterized by a set of attributes describing the above-mentioned set of services. Computation is mediated through the exchange of protocol messages guarded by such attributes. Furthermore, relevant application state is distributed on each mobile entity through a set of messaging boards, enabling a targeted form of communication and cooperation among ubiquitous applications. In this paper we report on our experience experimenting with this platform. Our initial results indicate that this platform is suitable for current RFID technology and exhibits low-cost, scalability and privacy. Ayis Ziotopoulos, Margarida F. Jacome, Gustavo de Veciana |
MDM | 2 |
| 2007 | Self-Imposed Temporal Redundancy: An Efficient Technique to Enhance the Reliability of Pipelined Functional UnitsabstractTemporal redundancy (TR) improves the reliability of computational functional units (FUs). However, it can guarantee detection of transient errors only, and may have a substantial power and area overhead. In this paper we present self-imposed temporal redundancy (SITR), a form of TR that can be applied to pipelined FUs and does not suffer from the aforementioned problems. A SITR-enhanced FU forces redundant computations to fire in consecutive cycles and requires a single additional cycle for the second computation and the comparison of the two results. We evaluate the power and area overhead of SITR and conclude that is always smaller than that of standard TR and that it does not depend on the FU complexity. We also use SITR to improve the reliability of the execution datapath of a simple out-of-order engine, typical of that used in high reliability embedded systems and future many-core architectures. Our simulations show that SITR outperforms TR, especially in FP applications. When the number of integer ALUs is larger than the machine width, the performance penalty of SITR is consistently less than 10%. Elias Mizan, Tileli Amimeur, Margarida F. Jacome |
SBAC-PAD | 3 |
| 2007 | Defect-Aware High-Level Synthesis Targeted at Reconfigurable NanofabricsabstractEntering the nanometer era, a major challenge to current design methodologies and tools is how to effectively address the high defect densities projected for nanoelectronic technologies. To this end, a reconfiguration-based defect-avoidance methodology for defect-prone nanofabrics was proposed. It judiciously architects the nanofabric, using probabilistic considerations, such that a very large number of alternative implementations can be mapped into it, enabling defects to be circumvented at configuration time, in a scalable way. Building on this foundation, in this paper, a synthesis framework aimed at implementing this new design paradigm is proposed. A key novelty of the approach with respect to traditional high-level synthesis (HLS) is that, rather than carefully optimizing a single (“deterministic”) solution, the goal is to simultaneously synthesize a large family of alternative solutions, so as to meet the required probability of successful configuration, or yield, while maximizing the average performance of the family of synthesized solutions. Experimental results generated for a set of representative benchmark kernels, assuming different defect regimes and target yields, empirically show that the proposed algorithms can effectively explore the complex probabilistic design space associated with this new class of HLS problems. Margarida F. Jacome |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | RAS-NANO: a reliability-aware synthesis framework for reconfigurable nanofabricsabstractEntering the nanometer era, a major challenge to current design methodologies and tools is to effectively address the high defect densities projected for nanotechnologies. To this end, we proposed a reconfiguration-based defect-avoidance methodology for defect-prone nanofabrics. It judiciously architects the nanofabric, using probabilistic considerations, such that a very large number of alternative implementations can be mapped into it, enabling defects to be circumvented at configuration time in a scalable way. Building on this foundation, in this paper we propose a synthesis framework aimed at implementing this new design paradigm. A key novelty of our approach with respect to traditional high level synthesis is that, rather than carefully optimizing a single (`deterministic') solution, our goal is to simultaneously synthesize a large family of alternative solutions, so as to meet the required probability of successful configuration, or yield, while maximizing the family's average performance. Experimental results generated for a set of representative benchmark kernels, assuming different defect regimes and target yields, empirically show that our proposed algorithms can effectively explore the complex probabilistic design space associated with this new class of high level synthesis problems Margarida F. Jacome |
DATE | 2 |
| 2005 | High performance computing on fault-prone nanotechnologies: novel microarchitecture techniques exploiting reliability-delay trade-offsabstractDevice and interconnect fabrics at the nanoscale will have a density of defects and susceptibility to transient faults far exceeding those of current silicon technologies. In this paper we introduce a new performance optimization dimension at the microarchitecture level which can mitigate overheads introduced by fault tolerance. This is achieved by directly exposing reliability versus delay design trade-offs while incorporating novel forms of speculation which use faster but less reliable versions of a microarchitecture's performance critical components. Based on a parameterized microarchitecture, we exhibit the benefits of optimizing these tradeoffs. Andrey V. Zykov, Elias Mizan, Margarida F. Jacome, Gustavo de Veciana, Ajay Subramanian |
DAC | 3 |
| 2005 | Predicated switching - optimizing speculation on EPIC machinesabstractExplicitly parallel instruction computing (EPIC) processors are a very attractive platform for many of today's multimedia and communications applications. In particular, clustered EPIC machines can take aggressive advantage of the available instruction-level parallelism, while maintaining high energy-delay efficiency. However, multicluster machines are more challenging to compile to than centralized machines. In this paper, we propose a novel compiler-directed speculation technique called predicated switching (PS) that can be applied to both centralized and multicluster EPIC machines. The two novel contributions in PS are: 1) a compiler transformation, denoted static single assignment-predicated switching, that leverages required data transfers between clusters for performance gains and 2) a static speculation algorithm to decide which specific kernel operations should actually be speculated in the final code, so as to maximize execution performance on the target processor. Experimental results performed on a representative set of time critical kernels compiled for a number of target machines show that, when compared to "resource-unaware" speculation techniques, PS improves performance with respect to at least one of the baselines in 80% of the cases by up to 38%. Moreover, we show that code size and register pressure are not adversely affected by our technique. Satish Pillai, Margarida F. Jacome |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Xtream-fit: an energy-delay efficient data memory subsystem for embedded media processingabstractDue to the critical role played by data memory subsystems in the performance and energy efficiency of embedded systems, the design of energy-efficient data memory architectures has received considerable attention in recent years. In this paper, we propose a novel special-purpose data memory subsystem called Xtream-Fit which is aimed at achieving high energy-delay efficiency for streaming media applications. A key novelty of Xtream-Fit is that it exposes a single customization parameter, thus enabling a very simple and yet effective design space exploration methodology. A second key contribution of this paper is the ability to achieve very high energy-delay efficiency through a synergistic combination of: 1) special purpose memory subsystem components, namely, a streaming memory and a scratch-pad memory and (2) a novel task-based execution model that exposes/enhances opportunities for efficient prefetching, and aggressive dynamic energy conservation techniques targeting on-chip and off-chip memory components. Extensive experimental results show that Xtream-Fit reduces the energy-delay product by 22% to 61%, as compared to general-purpose memory subsystems enhanced with state of the art cache decay and SDRAM power-mode control policies. Margarida F. Jacome |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Defect tolerant probabilistic design paradigm for nanotechnologiesabstractRecent successes in the development and self-assembly of nanoelectronic devices suggest that the ability to manufacture dense nanofabrics is on the near horizon. However, the tremendous increase in device density of nanoelectronics will be accompanied by a substantial increase in hard and soft faults, posing a major challenge to current design methodologies and tools. In this paper we propose a novel probabilistic design paradigm for defective but reconfigurable nanofabrics. The new design goal is to devise an appropriate structural/behavioral decomposition which improves scalability by constraining the reconfiguration process, while meeting a desired probability of successful instantiation, i.e, yield. Our approach not only addresses the scalability problem in configuring dense nanofabrics subject to defects, but gives a rich framework in which critical trade-offs among performance, yield, and per chip cost can be explored. We present a concrete instance of the approach and show extensive experimental results supporting these claims. Margarida F. Jacome, Gustavo de Veciana, Stephen Bijansky |
DAC | 1 |
| 2003 | Xtream-Fit: an energy-delay efficient data memory subsystem for embedded media processingabstractIn this paper we propose a novel special-purpose data memory subsystem, called Xtream-Fit, aimed at achieving high energy-delay efficiency for streaming media applications. A key novelty of Xtream-Fit is that it exposes a single customization parameter, thus enabling a very simple and yet effective design space exploration methodology. A second key contribution of this work is the ability to achieve very high energy-delay efficiency through a synergistic combination of: (1) special purpose memory subsystem components, namely, a Streaming Memory and Scratch-Pad Memory; and (2) a novel task-based execution model that exposes/enhances opportunities for efficient prefetching, and aggressive dynamic energy conservation techniques targeting on-chip and off-chip memory components. Extensive experimental results show that Xtream-Fit reduces energy-delay product by 46% to 83%, as compared to general-purpose memory subsystems enhanced with state of the art Cache Decay and SDRAM power mode control policies. Margarida F. Jacome |
DAC | 2 |
| 2003 | Architecture-level performance evaluation of component-based embedded systemsabstractA static performance evaluation technique is proposed to support early, architecture-level design space exploration for component-based embedded systems. The novel contribution is the use of a designer-specified evaluation scenario to identify a characteristic subset of system functionality that serves as a context for a rapid performance evaluation between candidate architectures. Fidelity is demonstrated with a case study that compares performance estimates of several candidate architectures to measurements from respective implementations. Jeffry T. Russell, Margarida F. Jacome |
DAC | 2 |
| 2003 | Compiler-Directed ILP Extraction for Clustered VLIW/EPIC Machines: Predication, Speculation and Modulo Scheduling
Satish Pillai, Margarida F. Jacome |
DATE | 2 |
| 2003 | Embedded Architect: A Tool for Early Performance Evaluation of Embedded SoftwareabstractEmbedded Architect is a design automation tool that embodies a static performance evaluation technique to support early, architecture-level design space exploration for component-based embedded systems. A static control flow characterization, called an evaluation scenario, is specified based on an incremental refinement of software source code, from which a pseudo-trace of operations is generated in combination with architecture mapping and several component parameters, a software performance metric is estimated The novel contribution is the implementation of a tool that automates specification of an evaluation scenario, which sets the context for a rapid performance evaluation of distinct candidate architectures. Jeffry T. Russell, Margarida F. Jacome |
ICSE | 2 |
| 2003 | Special issue on power-aware embedded computingabstractarticle Share on Special issue on power-aware embedded computing Editors: Margarida Jacome View Profile , Francky Catthoor View Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 2Issue 3August 2003 pp 251–254https://doi.org/10.1145/860176.860177Published:01 August 2003Publication History 1citation1,203DownloadsMetricsTotal Citations1Total Downloads1,203Last 12 Months4Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Margarida F. Jacome, Francky Catthoor |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2002 | Scenario-based software characterization as a contingency to traditional program profilingabstractProgram profiling is common way to characterize program behavior based on representative input. Some software, especially in embedded systems, cannot be profiled do to lack of tools or problems introduced by instrumentation of the code. As an alternative to traditionally profiling, a static analysis technique is proposed that allows a designer to characterize the flow of control of software.Operating on a flow graph representation of software, the proposed technique assists an expert designer in the specification of one or more representative scenarios. A scenario defines a specific flow of control that corresponds to a typical behavior of the system. The efficacy of the technique is demonstrated with two experiments: a comparison to traditional profiling and application to real embedded operating system software for which traditional profiling is not possible. Jeffry T. Russell, Margarida F. Jacome |
CASES | 2 |
| 2002 | RS-FDRA: A register-sensitive software pipelining algorithm for embedded VLIW processorsabstractThe paper proposes a novel software-pipelining algorithm, Register-Sensitive Force-Directed Retiming Algorithm (RS-FDRA), suitable for optimizing compilers targeting embedded very large instruction word processors. The key difference between RS-FDRA and previous approaches is that this algorithm can handle code-size constraints along with latency and resource constraints. This capability enables the exploration of Pareto "optimal" points with respect to code size and performance. RS-FDRA can also minimize the increase in register pressure typically incurred by software pipelining. This ability is critical since the need to insert spill code may result in significant performance degradation. Extensive experimental results are presented demonstrating that the extended set of optimization goals and constraints supported by RS-FDRA enables a thorough compiler-assisted exploration of tradeoffs among performance, code size, and register requirements for time-critical segments of embedded software components. Cagdas Akturan, Margarida F. Jacome |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | Application-specific clustered VLIW datapaths: early exploration on a parameterized design spaceabstractSpecialized clustered very large instruction word (VLIW) processors combined with effective compilation techniques enable aggressive exploitation of the high instruction-level parallelism inherent in many embedded media applications, while unlocking a variety of possible performance/cost tradeoffs. In this work, the authors propose a methodology to support early design space exploration of clustered VLIW datapaths, in the context of a specific target application. They argue that, due to the large size and complexity of the design space, the early design space exploration phase should consider only design space parameters that have a first-order impact on two key physical figures of merit: clock rate and power dissipation. These parameters were found to be: maximum cluster capacity, number of clusters, and bus (interconnect) capacity. Experimental validation of their design space exploration algorithm shows that a thorough exploration of the complex design space can be performed very efficiently in this abstract parameterized design space. Viktor S. Lapinskii, Margarida F. Jacome, Gustavo de Veciana |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | Cluster assignment for high-performance embedded VLIW processorsabstractClustering is an effective method to increase the available parallelism in VLIW datapaths without incurring severe penalties associated with a large number of register file ports. Efficient utilization of a clustered datapath requires careful binding/assignment of operations to clusters. The article proposes a binding algorithm that effectively explores trade-offs between in-cluster operation serialization and delays associated with data transfers between clusters. Extensive experimental evidence is provided showing that the algorithm generates high quality solutions for representative kernels, with up to 33% improvement over a state-of-the-art binding algorithm. Viktor S. Lapinskii, Margarida F. Jacome, Gustavo de Veciana |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2001 | Clustered VLIW Architectures with Predicated SwitchingabstractIn order to meet the high throughput requirements of applications exhibiting high ILP, VLIW ASIPs may increasingly include large numbers of functional units(FUs). Unfortunately, ”switching“ data through register files shared by large numbers of FUs quickly becomes a dominant cost/ performance factor suggesting that clustering smaller number of FUs around local register files may be beneficial even if data transfers are required among clusters. With such machines in mind, we propose a compiler transformation, predicated switching, which enables aggressive speculation while leveraging the penalties associated with inter-cluster communication to achieve gains in performance. Based on representative benchmarks, we demonstrate that this novel technique is particularly suitable for application specific clustered machines aimed at supporting high ILP as compared to state-of-the-art approaches. Margarida F. Jacome, Gustavo de Veciana, Satish Pillai |
DAC | 1 |
| 2001 | High-Quality Operation Binding for Clustered VLIW DatapathsabstractClustering is an effective method to increase the available parallelism in VLIW datapaths without incurring severe penalties associated with large number of register file ports. Efficient utilization of a clustered datapath requires careful binding of operations to clusters. The paper proposes a binding algorithm that effectively explores tradeoffs between in-cluster operation serialization and delays associated with data transfers between clusters. Extensive experimental evidence is provided showing that the algorithm generates high quality solutions for basic blocks, with up to 29% improvement over a state-of-the-art advanced binding algorithm. Viktor S. Lapinskii, Margarida F. Jacome, Gustavo de Veciana |
DAC | 2 |
| 2001 | CALiBeR: A Software Pipelining Algorithm for Clustered Embedded VLIW ProcessorsabstractIn this paper, we describe a software pipelining framework, CALiBeR (cluster aware load balancing retiming algorithm), suitable for compilers targeting clustered embedded VLIW processors. CALiBeR can be effectively used by embedded system designers to explore different code optimization alternatives, i.e. it can assist the generation of high-quality customized retiming solutions for desired program memory size and throughput requirements, while minimizing register pressure. An extensive set of experimental results is presented, considering several representative benchmark loop kernels and a wide variety of clustered datapath configurations, demonstrating that our algorithm compares favorably with one of the best state-of-the-art algorithms, achieving up to 50% improvement in performance and up to 47% improvement in register requirements. Cagdas Akturan, Margarida F. Jacome |
ICCAD | 2 |
| 2000 | A new technique for estimating lower bounds on latency for high level synthesisabstractIn this paper we present a novel and fast estimation technique that produces tight latency lower bounds for Data Flow Graphs representing time critical segments of the application of interest. Our proposed technique can be used to compute a tighter earliest scheduling step for nodes (operations) in the Data Flow Graph and thus be used to improve the result quality of any technique requiring the computation of such ASAP values. Helvio P. Peixoto, Margarida F. Jacome |
ACM Great Lakes Symposium on VLSI | 2 |
| 2000 | Exploring Performance Tradeoffs for Clustered VLIW ASIPsabstractVLIW ASIPs provide an attractive solution for increasingly pervasive real-time multimedia and signal processing embedded applications. In this paper we propose an algorithm to support trade-off exploration during the early phases of the design/specialization of VLIW ASIPs with clustered datapaths. For purposes of an early exploration step, we define a parameterized family of clustered datapaths D(m,n), where m and n denote interconnect capacity and cluster capacity constraints on the family. Given a kernel, the proposed algorithm explores the space of feasible clustered datapaths and returns: a datapath configuration; a binding and scheduling for the operations; and a corresponding estimate for the best achievable latency over the specified family. Moreover, we show how the parameters m and n, as well as a target latency optionally specified by the designer, can be used to effectively explore trade-offs among delay, power/energy, and latency. Extensive empirical evidence is provided showing that the proposed approach is strikingly effective at attacking this complex optimization problem. Margarida F. Jacome, Gustavo de Veciana, Viktor S. Lapinskii |
ICCAD | 1 |
| 2000 | Symbolic Binding for Clustered VLIW ASIPsabstractThe paper proposes a symbolic framework to address the binding problem for embedded VLIW ASIPs. Alternative objective functions as well as trade-offs relevant to the binding phase of code generation for embedded processors are presented and discussed. Experimental results obtained for a number of benchmarks extracted from the literature empirically demonstrate the promise of our approach. Satish Pillai, Margarida F. Jacome |
ICCD | 2 |
| 1999 | The Design Space Layer: Supporting Early Design Space Exploration for Core-Based DesignsabstractA novel library layer, called the "design space layer," is proposed, aimed at supporting both IP-based and traditional "in-house" design methodologies, during early design space exploration. Strategies for effectively pruning the large design spaces characteristic of system-on-a-chip designs, and for transparently retrieving information on cores adequate for implementing the system components, are supported by the proposed layer. The layer is self-documented and highly compartmentalized into hierarchies of classes of design objects, and is thus easily scalable. A design space layer developed for encryption applications is presented and discussed in some detail. Margarida F. Jacome, Helvio P. Peixoto, Ander Royo, Juan Carlos López 0001 |
DATE | 1 |
| 1999 | Lower bound on latency for VLIW ASIP datapathsabstractTraditional lower bound estimates on latency for dataflow graphs assume no data transfer delays. While such approaches can generate tight lower bounds for datapaths with a centralized register file, the results may be uninformative for datapaths with distributed register file structures that are characteristic of VLIW ASIPs (very large instruction word application-specific instruction set processors). In this paper, we propose a latency bound that accounts for such data transfer delays. The novelty of our approach lies in constructing the "window dependency graph" and bounds associated with the problem which capture delay penalties due to operation serialization and/or data moves among distributed register files. Through a set of benchmark examples, we show that the bound is competitive with state-of-the-art approaches. Moreover, our experiments show that the approach can aid an iterative improvement algorithm in determining good functional unit assignments-a key step in code generation for VLIW ASIPs. Margarida F. Jacome, Gustavo de Veciana |
ICCAD | 1 |
| 1998 | Hierarchical Algorithms for Assessing Probabilistic Constraints on System PerformanceabstractWe propose an algorithm for assessing probabilistic performance constraints for systems including components with uncertain delays. We make a case for designing systems based on a probabilistic relaxation of performance constraints, as this has the potential for resulting in lower silicon area and/or power consumption. We consider a concrete example, an MPEG decoder, for which we discuss modeling and assessment of probabilistic throughput constraints. Gustavo de Veciana, Margarida F. Jacome, Jian-Huei Guo |
DAC | 2 |
| 1998 | A Methodology for Task Based Partitioning and Scheduling of Dynamically Reconfigurable SystemsabstractTaking maximum advantage of dynamic reconfiguration in the implementation of digital systems poses a number of challenging research problems. Specifically, techniques are needed to partition the system behavioral description into segments of computation (or "scheduling units"), and to define a reconfiguration schedule with respect to those units, so as to maximize the performance of the dynamically reconfigurable system, subject to the area constraints of the FPGA. We propose a methodology to: (1) perform a coarse-grained partitioning of the system behavioral description into a set of tasks, (2) determine which sub-set of tasks is to remain resident in the FPGA, and which sub-set is to be non resident, (3) generate a reconfiguration schedule for the non-resident tasks by specifying when such tasks should be loaded on to and erased from the FPGA. Pedro Merino 0001, Margarida F. Jacome, Juan Carlos López 0001 |
FCCM | 2 |
| 1998 | Software power estimation and optimization for high performance, 32-bit embedded processorsabstractA software energy estimation model is presented for a family of high performance, integrated, 32-bit embedded RISC processors. This model is significantly less complex than previous models, and yet is demonstrated to accurately predict energy consumption to within 8% with 99% confidence based on physical measurements. Factors such as operating frequency, source/destination registers, and operand values are explored. In view of this model, previously proposed optimizations are evaluated for potential energy savings. We conclude that, for the class of processors under discussion, a good optimizing compiler that minimizes execution time will simultaneously minimize energy consumption. Jeffry T. Russell, Margarida F. Jacome |
ICCD | 2 |
| 1997 | Algorithm and architecture-level design space exploration using hierarchical data flowsabstractIncorporating algorithm and architecture level design space exploration in the early phases of the design process can have a dramatic impact on the area, speed, and power consumption of the resulting systems. This paper proposes a framework for supporting system-level design space exploration and discusses the three fundamental issues involved in effectively supporting such an early design space exploration: definition of an adequate level of abstraction; definition of good fidelity system-level metrics; and definition of mechanisms for automating the exploration process. The first issue, the definition of an adequate level of abstraction is then addressed in detail. Specifically, an algorithm-level model, an architecture-level model, and a set of operations on these models, are proposed, aiming at efficiently supporting an early, aggressive system-level design space exploration. A discussion on work in progress in the other two topics, metrics and automation, concludes the paper. Helvio P. Peixoto, Margarida F. Jacome |
ASAP | 2 |
| 1996 | A formal basis for design process planning and managementabstractIn this paper we present a formalism that allows for a complete and general characterization of design disciplines and for a unified representation of arbitrarily complex design processes taking place in the context of these disciplines. This formalism has been used as the basis for the development of several prototype CAD meta-tools that offer effective design process planning and management services. Margarida F. Jacome, Stephen W. Director |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1994 | A formal basis for design process planning and management
Margarida F. Jacome, Stephen W. Director |
ICCAD | 1 |
| 1992 | Design Process Management for CAD Frameworks
Margarida F. Jacome, Stephen W. Director |
DAC | 1 |