Erika Gunadi

dblp:93/4603 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 53% Hardware reliability and fault tolerance · 27% Energy-efficient computing · 20%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
out-of-order execution
0.222011
CRIB: consolidated rename, issue, and bypass · ISCA 2011
Physical Register Inlining · ISCA 2004
Processor architecture and microarchitecture › out-of-order execution
register renaming
0.222011
CRIB: consolidated rename, issue, and bypass · ISCA 2011
Physical Register Inlining · ISCA 2004
Energy-efficient computing
dynamic power reduction
0.112011
CRIB: consolidated rename, issue, and bypass · ISCA 2011
Processor architecture and microarchitecture
instruction issue logic
0.112011
CRIB: consolidated rename, issue, and bypass · ISCA 2011
Energy-efficient computing › low-power design
low-power processor design
0.112011
CRIB: consolidated rename, issue, and bypass · ISCA 2011
Hardware reliability and fault tolerance
aging
0.112010
Combating Aging with the Colt Duty Cycle Equalizer · MICRO 2010
Hardware reliability and fault tolerance › aging
aging mitigation
0.112010
Combating Aging with the Colt Duty Cycle Equalizer · MICRO 2010
Hardware reliability and fault tolerance › aging
bias temperature instability
0.112010
Combating Aging with the Colt Duty Cycle Equalizer · MICRO 2010
Processor architecture and microarchitecture › register file
physical register management
0.012004
Physical Register Inlining · ISCA 2004
Processor architecture and microarchitecture
register file
0.012004
Physical Register Inlining · ISCA 2004
Processor architecture and microarchitecture
instruction-level parallelism
0.012011
CRIB: consolidated rename, issue, and bypass · ISCA 2011
Processor architecture and microarchitecture
instruction scheduling
0.012004
Physical Register Inlining · ISCA 2004

Methods — techniques the papers use, named apart from their topics

pseudorandom indexing · 0.1operand swapping · 0.1duty cycle equalization · 0.1microarchitectural simulation · 0.0
YearPublicationVenuePosition
2013 REEL: Reducing effective execution latency of floating point operations
abstract
The height of the dynamic dependence graph of a program, as executed by a processor, determines the minimum bound on the execution time. This height can be decreased by reducing the effective execution latency of operations that form dependence chains in the graph. In this paper, we propose a technique called REEL to reduce overall latency of chains of dependent floating point (FP) operations by increasing the throughput of computation. REEL comprises of a high-throughput floating point unit (HFP) that allows early issue of an FP Add that is dependent on another FP Add or FP Multiply. This is complemented by instruction scheduler modifications that allow early issue of dependent FP Adds, and a novel checker logic that corrects any precision errors. Unlike conventional static operation fusion, like fused Multiply-Add (FMA), there are no changes to the instruction set to enable utilization of the new hardware, and no recompilation is necessary. Furthermore, unlike ISA-level FMA, our technique produces results that are bit compatible while boosting performance of Add-Add dependence pairs in addition to Multiply-Add pairs. Our evaluation of REEL using CFP2006 benchmarks shows an average performance gain of 7.6% and maximum performance gain of 17% while consuming 1.2% lower energy.
Vignyan Reddy Kothinti Naresh, Syed Zohaib Gilani, Erika Gunadi, Nam Sung Kim, Michael J. Schulte, Mikko H. Lipasti
ISLPED3
2011 CRIB: consolidated rename, issue, and bypass
abstract
Conventional high-performance processors utilize register renaming and complex broadcast-based scheduling logic to steer instructions into a small number of heavily-pipelined execution lanes. This requires multiple complex structures and repeated dependency resolution, imposing a significant dynamic power overhead. This paper advocates in-place execution of instructions, a power-saving, pipeline-free approach that consolidates rename, issue, and bypass logic into one structure---the CRIB---while simultaneously eliminating the need for a multiported register file, instead storing architected state in a simple rank of latches. CRIB achieves the high IPC of an out-of-order machine while keeping the execution core clean, simple, and low power. The datapath within a CRIB structure is purely combinational, eliminating most of the clocked elements in the core while keeping a fully synchronous yet high-frequency design. Experimental results match the IPC and cycle time of a baseline out-of-order design while reducing dynamic energy consumption by more than 60% in affected structures.
Erika Gunadi, Mikko H. Lipasti
ISCA1
2010 Combating Aging with the Colt Duty Cycle Equalizer
abstract
Bias temperature instability, hot-carrier injection, and gate-oxide wear out will cause severe lifetime degradation in the performance and the reliability of future CMOS devices. The design guard band to counter these negative effects will be too expensive, largely due to the worst-case behavior induced by the uneven utilization of devices on the chip. To mitigate these effects over a chip's lifetime, this paper proposes Colt, a simple yet holistic scheme to balance the utilization of devices in a processor by equalizing the duty cycle ratio of circuits'internal nodes and the usage frequency of devices. Colt relies on alternating true-and complement-mode operations to equalize the duty cycle ratio of signals (thus the utilization of devices) in most data path and storage devices. Colt also employs a pseudorandom indexing scheme to balance the usage of entries in storage structures that often exhibit highly uneven utilization of entries. Finally, an operand-swapping scheme equalizes utilization of the left and right operand data paths. The proposed mechanisms impose trivial overhead in area, complexity, power, and performance, while recapturing 27% of aging-induced performance degradation and improving meantime to failure by an estimated 40%.
Erika Gunadi, Abhishek A. Sinkar, Nam Sung Kim, Mikko H. Lipasti
MICRO1
2007 A position-insensitive finished store buffer
abstract
This paper presents the finished store buffer (or FSB), an alternative and position-insensitive approach for building a scalable store buffer for an out-of-order processor. Exploiting the fact that only a small portion of in-flight stores are done executing (i.e. finished) and waiting for retirement, we are able to build a much smaller and more scalable store buffer. Our study shows that we only need at most half of the number of entries in a conventional store queue if we buffer only the stores that have finished execution. Entries in the store buffer are allocated at issue and disallocated on retirement. A clever encoder circuit is used to provide positional searches without an explicitly positional queue structure. While reducing the access latency and power consumption significantly, our technique has virtually no detrimental effect on per-cycle performance (IPC).
Erika Gunadi, Mikko H. Lipasti
ICCD1
2007 Power-aware operand delivery
abstract
Based on operand delivery, existing microprocessors can be categorized into architected register file (ARF) or physical register file (PRF) machines, both with or without payload RAM (PL). Though many previous generation microprocessors use a PRF without PL, the trend of newer microprocessors targeting lower power environments seem to be moving towards ARF with PL. We quantitatively analyze power consumption of different machine styles: ARF with PL, ARF without PL, PRF with PL, and PRF only machine. Our result shows that PRF without PL consumes the least amount of power and is fundamentally the best approach for building power-aware out-of-order microprocessors.
Erika Gunadi, Mikko H. Lipasti
ISLPED1
2004 Physical Register Inlining
abstract
Physical register access time increases the delay between scheduling and execution in modern out-of-order processors. As the number of physical registers increases, this delay grows, forcing designers to employ register files with multicycle access. This paper advocates more efficient utilization of a fewer number of physical registers in order to reduce the access time of the physical register file. Register values with few significant bits are stored in the rename map using physical register inlining, a scheme analogous to inlining of operand fields in data structures. Specifically, whenever a register value can be expressed with fewer bits than the register map would need to specify a physical register number, the value is stored directly in the map, avoiding the indirection, and saving space in the physical register file. Not surprisingly, we find that a significant portion of all register operands can be stored in the map in this fashion, and describe straightforward microarchitectural extensions that correctly implement physical register inlining. We find that physical register inlining performs well, particularly in processors that are register-constrained.
Mikko H. Lipasti, Brian R. Mestan, Erika Gunadi
ISCA3