Xavier Vera

dblp:34/3559 · DBLP profile ↗
← Back
42ranked-venue papers
12as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 9 first-authorSoftware engineering, systems software and programming languages · 19 · 6 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
21 papers
Hardware reliability and fault tolerance · 27% Electronic design automation · 23% Processor architecture and microarchitecture · 17%
Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 100%

Topics — the 30 heaviest of 53, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware reliability and fault tolerance
soft errors
0.652016
A Case for Acoustic Wave Detectors for Soft-Errors · IEEE Trans. Computers 2016
Selective replication: A lightweight technique for soft errors · ACM Trans. Comput. Syst. 2009
Reducing Soft Errors through Operand Width Aware Policies · IEEE Trans. Dependable Secur. Comput. 2009
Electronic design automation › hardware verification and test › design validation
post-silicon validation
0.322013
Deconfigurable microprocessor architectures for silicon debug acceleration · ISCA 2013
Hardware/software-based diagnosis of load-store queues using expandable activity logs · HPCA 2011
Memory systems
cache
0.262010
IATAC: a smart predictor to turn-off L2 cache lines · ACM Trans. Archit. Code Optim. 2005
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
Efficient and Accurate Analytical Modeling of Whole-Program Data Cache Behavior · IEEE Trans. Computers 2004
Electronic design automation › hardware verification and test
online testing
0.222010
Microarchitectural Online Testing for Failure Detection in Memory Order Buffers · IEEE Trans. Computers 2010
End-to-end register data-flow continuous self-test · ISCA 2009
Electronic design automation › hardware verification and test
post-silicon debug
0.212013
Deconfigurable microprocessor architectures for silicon debug acceleration · ISCA 2013
Hardware reliability and fault tolerance › soft errors
soft error detection
0.112012
Setting an error detection infrastructure with low cost acoustic wave detectors · ISCA 2012
Electronic design automation › hardware verification and test › design for testability
built-in self-test
0.112011
Implementing End-to-End Register Data-Flow Continuous Self-Test · IEEE Trans. Computers 2011
Electronic design automation › hardware verification and test
hardware verification
0.112011
Accelerating microprocessor silicon validation by exposing ISA diversity · MICRO 2011
Processor architecture and microarchitecture
load/store queue
0.112011
Hardware/software-based diagnosis of load-store queues using expandable activity logs · HPCA 2011
Processor architecture and microarchitecture › memory system microarchitecture
memory ordering
0.112011
Hardware/software-based diagnosis of load-store queues using expandable activity logs · HPCA 2011
Electronic design automation › hardware verification and test › test generation
random test generation
0.112011
Accelerating microprocessor silicon validation by exposing ISA diversity · MICRO 2011
Electronic design automation › hardware verification and test
silicon validation
0.112011
Accelerating microprocessor silicon validation by exposing ISA diversity · MICRO 2011
Energy-efficient computing › power-performance tradeoff
energy-delay product optimization
0.112010
High-Performance low-vcc in-order core · HPCA 2010
Processor architecture and microarchitecture › microprocessor design › processor core design
in-order core
0.112010
High-Performance low-vcc in-order core · HPCA 2010
Energy-efficient computing
low-voltage operation
0.112010
High-Performance low-vcc in-order core · HPCA 2010
Energy-efficient computing
voltage scaling
0.112010
High-Performance low-vcc in-order core · HPCA 2010
Hardware reliability and fault tolerance
error detection and correction
0.112009
Reducing Soft Errors through Operand Width Aware Policies · IEEE Trans. Dependable Secur. Comput. 2009
Processor architecture and microarchitecture › out-of-order execution
out-of-order processor
0.112009
Selective replication: A lightweight technique for soft errors · ACM Trans. Comput. Syst. 2009
Distributed systems › replication › partial replication
selective replication
0.112009
Selective replication: A lightweight technique for soft errors · ACM Trans. Comput. Syst. 2009
Performance modeling and evaluation › cache performance modeling
data cache analysis
0.122004
Efficient and Accurate Analytical Modeling of Whole-Program Data Cache Behavior · IEEE Trans. Computers 2004
Data Caches in Multitasking Hard Real-Time Systems · RTSS 2003
Memory systems › cache
cache behavior
0.122004
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
Let's Study Whole-Program Cache Behaviour Analytically · HPCA 2002
Memory systems › cache management
cache locking
0.122003
Data cache locking for higher program predictability · SIGMETRICS 2003
Data Caches in Multitasking Hard Real-Time Systems · RTSS 2003
Memory systems
cache management
0.122003
Data cache locking for higher program predictability · SIGMETRICS 2003
Data Caches in Multitasking Hard Real-Time Systems · RTSS 2003
Compilers and program optimization › memory optimization
data locality optimization
0.132005
An accurate cost model for guiding data locality transformations · ACM Trans. Program. Lang. Syst. 2005
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
Let's Study Whole-Program Cache Behaviour Analytically · HPCA 2002
Hardware reliability and fault tolerance › aging
transistor aging
0.112007
Penelope: The NBTI-Aware Processor · MICRO 2007
Compilers and program optimization › memory optimization
data layout transformation
0.112005
An accurate cost model for guiding data locality transformations · ACM Trans. Program. Lang. Syst. 2005
Compilers and program optimization › loop optimization
loop tiling
0.112005
An accurate cost model for guiding data locality transformations · ACM Trans. Program. Lang. Syst. 2005
Energy-efficient computing › leakage power reduction
cache leakage reduction
0.112005
IATAC: a smart predictor to turn-off L2 cache lines · ACM Trans. Archit. Code Optim. 2005
Performance modeling and evaluation › cache performance modeling
cache miss equation
0.012004
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
Performance modeling and evaluation
cache performance modeling
0.012004
Efficient and Accurate Analytical Modeling of Whole-Program Data Cache Behavior · IEEE Trans. Computers 2004

Methods — techniques the papers use, named apart from their topics

simulation · 0.2acoustic wave detection · 0.2arithmetic codes · 0.2parity · 0.2ECC · 0.2random test programs · 0.2failure triage · 0.2expandable logging · 0.1error detection · 0.1diagnosis algorithm · 0.1genetic algorithm · 0.1cost model · 0.1sampling · 0.0polyhedral analysis · 0.0tiling · 0.0static cache analysis · 0.0padding · 0.0cache partitioning · 0.0
YearPublicationVenuePosition
2020 Inside Tiger Lake: Intel's Next Generation Mobile Client CPU
abstract
This article consists only of a collection of slides from the author's conference presentation.
Xavier Vera
Hot Chips Symposium1
2016 A Case for Acoustic Wave Detectors for Soft-Errors
abstract
The continuing decrease in dimensions and operating voltage of transistors has increased their sensitivity against radiation phenomena, making soft errors an important challenge in future microprocessors. New techniques for detecting errors in the logic and memories that allow meeting the desired failure rate are key to keep harnessing the benefits of Moore's law. This paper proposes a low-cost dynamic particle strike detection mechanism based on acoustic wave detectors. Our results show that the proposed mechanism can protect the whole chip, including both the logic and the memory arrays, and detect all the soft errors caused by particle strikes with minimal hardware overhead and performance cost.
Gaurang Upasani, Xavier Vera, Antonio González 0001
IEEE Trans. Computers2
2014 Framework for economical error recovery in embedded cores
abstract
The vulnerability of the current and future processors towards transient errors caused by particle strikes is expected to increase rapidly because of exponential growth rate of on-chip transistors, the lower voltages and the shrinking feature size. This encourages innovation in the direction of finding new techniques for providing robustness in logic and memories that allow meeting the desired failures in-time (FIT) budget in future chip multiprocessors (CMPs) present in embedded systems. In embedded systems two aspects of robustness, error detection and containment, are of paramount importance. This paper proposes a light-weight and scalable architecture that uses acoustic wave detectors for error detection and contains errors at the core level. We show how selectively applying error containment can reduce the number of detectors required for error containment. We observe that by using 17 detectors we can achieve error containment coverage of 97.8%.
Gaurang Upasani, Xavier Vera, Antonio González 0001
IOLTS2
2014 Avoiding core's DUE & SDC via acoustic wave detectors and tailored error containment and recovery
abstract
The trend of downsizing transistors and operating voltage scaling has made the processor chip more sensitive against radiation phenomena making soft errors an important challenge. New reliability techniques for handling soft errors in the logic and memories that allow meeting the desired failures-in-time (FIT) target are key to keep harnessing the benefits of Moore's law. The failure to scale the soft error rate caused by particle strikes, may soon limit the total number of cores that one may have running at the same time. This paper proposes a light-weight and scalable architecture to eliminate silent data corruption errors (SDC) and detected unrecoverable errors (DUE) of a core. The architecture uses acoustic wave detectors for error detection. We propose to recover by confining the errors in the cache hierarchy, allowing us to deal with the relatively long detection latencies. Our results show that the proposed mechanism protects the whole core (logic, latches and memory arrays) incurring performance overhead as low as 0.60%.
Gaurang Upasani, Xavier Vera, Antonio González 0001
ISCA2
2013 Capturing vulnerability variations for register files
abstract
Soft error rates are estimated based on worst-case architectural vulnerability factor (AVF). Therefore, it makes tracking real-time accurate AVF very attractive to computer designers: more accurate AVF numbers will allow turning on more features at runtime while keeping the promised SDC and DUE rates. This paper presents a hardware mechanism based on linear regressions to estimate the AVF (SDC and DUE) of the register file for out-of-order cores. Our results show that we are able to have a high correlation factor at low cost.
Javier Carretero, Enric Herrero, Matteo Monchiero, Tanausú Ramírez, Xavier Vera
DATE5
2013 Reducing DUE-FIT of caches by exploiting acoustic wave detectors for error recovery
abstract
Cosmic radiation induced soft errors have emerged as a key challenge in computer system design. The exponential increase in the transistor count will drive the per chip fault rate sky high. New techniques for detecting errors in the logic and memories that allow meeting the desired failures in-time (FIT) budget in future chip multiprocessors (CMPs) are essential. Among the two major contributors towards soft error rate, silent data corruption (SDC) and detected unrecoverable error (DUE), DUE is the largest. Moreover, processors can experience a super-linear increase in DUE when the size of the write-back cache is doubled. This paper targets the DUE problem in write-back data caches. We analyze the cost of protection against single bit and multi-bit upsets into caches. Our results show that the proposed mechanism can reduce the DUE to “0” with minimum area, power and performance overheads.
Gaurang Upasani, Xavier Vera, Antonio González 0001
IOLTS2
2013 Deconfigurable microprocessor architectures for silicon debug acceleration
abstract
The share of silicon debug in the overall microprocessor chips development cycle is rapidly expanding due to the ever growing design complexity and the limited efficiency of pre-silicon validation methods. Massive application of short random test programs on the prototype microprocessor chips is one of the most effective parts of silicon debug. However, a major bottleneck and source of "noise" in this phase is that large numbers of random test programs fail due to the same or similar design bugs. This redundant behavior adds long delays in the debug flow since each failing random program must be separately examined, although it does not usually bring new debug information. The development of effective techniques that detect dominant modes of failure among random programs and triage them into common categories eliminate redundant debug sessions and significantly boost silicon debug.
Nikos Foutris, Dimitris Gizopoulos, Xavier Vera, Antonio González 0001
ISCA3
2012 Setting an error detection infrastructure with low cost acoustic wave detectors
abstract
The continuing decrease in dimensions and operating voltage of transistors has increased their sensitivity against radiation phenomena making soft errors an important challenge in future chip multiprocessors (CMPs). Hence, new techniques for detecting errors in the logic and memories that allow meeting the desired failures-in-time (FIT) budget in CMPs are required. This paper proposes a low-cost dynamic particle strike detection mechanism through acoustic wave detectors. Our results show that our mechanism can protect both the logic and the memory arrays. As a case study, we also show how this technique can be combined with error codes to protect the last-level cache at low cost.
Gaurang Upasani, Xavier Vera, Antonio González 0001
ISCA2
2011 Architectures for online error detection and recovery in multicore processors
abstract
The huge investment in the design and production of multicore processors may be put at risk because the emerging highly miniaturized but unreliable fabrication technologies will impose significant barriers to the life-long reliable operation of future chips. Extremely complex, massively parallel, multi-core processor chips fabricated in these technologies will become more vulnerable to: (a) environmental disturbances that produce transient (or soft) errors, (b) latent manufacturing defects as well as aging/wearout phenomena that produce permanent (or hard) errors, and (c) verification inefficiencies that allow important design bugs to escape in the system. In an effort to cope with these reliability threats, several research teams have recently proposed multicore processor architectures that provide low-cost dependability guarantees against hardware errors and design bugs. This paper focuses on dependable multicore processor architectures that integrate solutions for online error detection, diagnosis, recovery, and repair during field operation. It discusses taxonomy of representative approaches and presents a qualitative comparison based on: hardware cost, performance overhead, types of faults detected, and detection latency. It also describes in more detail three recently proposed effective architectural approaches: a software-anomaly detection technique (SWAT), a dynamic verification technique (Argus), and a core salvaging methodology.
Dimitris Gizopoulos, Mihalis Psarakis, Sarita V. Adve, Pradeep Ramachandran, Siva Kumar Sastry Hari, Daniel J. Sorin, Albert Meixner, Arijit Biswas, Xavier Vera
DATE9
2011 Hardware/software-based diagnosis of load-store queues using expandable activity logs
abstract
The increasing device count and design complexity are posing significant challenges to post-silicon validation. Bug diagnosis is the most difficult step during post-silicon validation. Limited reproducibility and low testing speeds are common limitations in current testing techniques. Moreover, low observability defies full-speed testing approaches. Modern solutions like on-chip trace buffers alleviate these issues, but are unable to store long activity traces. As a consequence, the cost of post-Si validation now represents a large fraction of the total design cost. This work describes a hybrid post-Si approach to validate a modern load-store queue. We use an effective error detection mechanism and an expandable logging mechanism to observe the microarchitectural activity for long periods of time, at processor full-speed. Validation is performed by analyzing the log activity by means of a diagnosis algorithm. Correct memory ordering is checked to root the cause of errors.
Javier Carretero, Xavier Vera, Jaume Abella 0001, Tanausú Ramírez, Matteo Monchiero, Antonio González 0001
HPCA2
2011 New reliability mechanisms in memory design for sub-22nm technologies
abstract
The TRAMS (Terascale Reliable Adaptive MEMORY Systems) project addresses in an evolutionary way the ultimate CMOS scaling technologies and paves the way for revolutionary, most promising beyond-CMOS technologies. In this abstract we show the significant variability levels of future 18 and 13nm device bulk-CMOS technologies as well as its dramatic effect on the yield of memory cells, and what kind of circuit solution would be required to maintain the current yield level. Later, we discuss the impact of errors at the system level, and different approaches at system level to adapt the heterogeneous systems to user's requirements.
Nivard Aymerich, A. Asenov, Andrew R. Brown, Ramon Canal, Binjie Cheng, Joan Figueras, Antonio González 0001, Enric Herrero, S. Markov, Miguel Corbalan, Peyman Pouyan, Tanausú Ramírez, Antonio Rubio 0001, Elena I. Vatajelu, Xavier Vera, Xingsheng Wang, Paul Zuber
IOLTS15
2011 Accelerating microprocessor silicon validation by exposing ISA diversity
abstract
Microprocessor design validation is a time consuming and costly task that tends to be a bottleneck in the release of new architectures. The validation step that detects the vast majority of design bugs is the one that stresses the silicon prototypes by applying huge numbers of random tests. Despite its bug detection capability, this step is constrained by extreme computing needs for random tests simulation to extract the bug-free memory image for comparison with the actual silicon image.
Nikos Foutris, Dimitris Gizopoulos, Mihalis Psarakis, Xavier Vera, Antonio González 0001
MICRO4
2011 Implementing End-to-End Register Data-Flow Continuous Self-Test
abstract
While Moore's Law predicts the ability of semiconductor industry to engineer smaller and more efficient transistors and circuits, there are serious issues not contemplated in that law. One concern is the verification effort of modern computing systems, which has grown to dominate the cost of system design. On the other hand, technology scaling leads to burn-in phase out. As a result, in-the-field error rate may increase due to both actual errors and latent defects. Whereas data can be protected with arithmetic codes, there is a lack of cost-effective mechanisms for control logic. This paper presents a light-weight microarchitectural mechanism that ensures that data consumed through registers are correct. The structures protected include the issue queue logic and the data associated (i.e., tags and control signals), input multiplexors, rename data, replay logic, register free-list and release logic, and register file logic. Our results show a coverage around 90 percent for the targeted structures with a cost in power and area of about four percent, and without impact in performance.
Javier Carretero, Pedro Chaparro, Xavier Vera, Jaume Abella 0001, Antonio González 0001
IEEE Trans. Computers3
2010 The split register file
abstract
Technology scaling requires lowering Vcc due to power constraints. Unfortunately, permanent faulty bit rates grow due to the higher impact of process variations at low Vcc, especially in the register file whose critical timing limits circuit optimizations. This paper proposes a novel register file design based on splitting registers and discarding faulty blocks to increase the number of registers available. By increasing the number of registers available higher performance can be obtained and yield increases because a larger number of processors reaches the minimum number of registers required to operate.
Jaume Abella 0001, Javier Carretero, Pedro Chaparro, Xavier Vera
DATE4
2010 High-Performance low-vcc in-order core
abstract
Power density grows in new technology nodes, thus requiring Vcc to scale especially in mobile platforms where energy is critical. This paper presents a novel approach to decrease Vcc while keeping operating frequency high. Our mechanism is referred to as immediate read after write (IRAW) avoidance. We propose an implementation of the mechanism for an Intel®SilverthorneTMin-order core. Furthermore, we show that our mechanism can be adapted dynamically to provide the highest performance and lowest energy-delay product (EDP) at each Vcc level. Results show that IRAW avoidance increases operating frequency by 57% at 500mV and 99% at 400mV with negligible area and power overhead (below 1%), which translates into large speedups (48% at 500mV and 90% at 400mV) and EDP reductions (0.61 EDP at 500mV and 0.33 at 400mV).
Jaume Abella 0001, Pedro Chaparro, Xavier Vera, Javier Carretero, Antonio González 0001
HPCA3
2010 MT-SBST: Self-test optimization in multithreaded multicore architectures
abstract
Instruction-based or software-based self-testing (SBST) is a scalable functional testing paradigm that has gained increasing acceptance in testing of single-threaded uniprocessors. Recent computer architecture trends towards chip multiprocessing and multithreading have raised new challenges in the test process. In this paper, we present a novel self-test optimization strategy for multithreaded, multicore microprocessor architectures and apply it to both manufacturing testing (execution from on-chip cache memory) and post-silicon validation (execution from main memory) setups. The proposed self-test program execution optimization aims to: (a) take maximum advantage of the available execution parallelism provided by multiple threads and multiple cores, (b) preserve the high fault coverage that single-thread execution provides for the processor components, and (c) enhance the fault coverage of the thread-specific control logic of the multithreaded multiprocessor. The proposed multithreaded (MT) SBST methodology generates an efficient multithreaded version of the test program and schedules the resulting test threads into the hardware threads of the processor to reduce the overall test execution time and on the same time to increase the overall fault coverage. We demonstrate our methodology in the OpenSPARC T1 processor model which integrates eight CPU cores, each one supporting four hardware threads. MT-SBST methodology and scheduling algorithm significantly speeds up self-test time at both the core level (3.6 times) and the processor level (6.0 times) against single-threaded execution, while at the same time it improves the overall fault coverage. Compared with straightforward multithreaded execution, it reduces the self-test time at both the core level and the processor level by 33% and 20%, respectively. Overall, MT-SBST reaches more than 91% stuck-at fault coverage for the functional units and 88% for the entire chip multiprocessor, a total of more than 1.5M logic gates.
Nikos Foutris, Mihalis Psarakis, Dimitris Gizopoulos, Andreas Apostolakis, Xavier Vera, Antonio González 0001
ITC5
2010 VCTA: A Via-Configurable Transistor Array regular fabric
abstract
Layout regularity is introduced progressively by integrated circuit manufacturers to reduce the increasing systematic process variations in the deep sub-micron era. In this paper we focus on a scenario where layout regularity must be pushed to the limit to deal with severe systematic process variations in future technology nodes. With this objective, we propose and evaluate a new regular layout style called Via-Configurable Transistor Array (VCTA) that maximizes regularity at device and interconnect levels. In order to assess VCTA maximum layout regularity tradeoffs, we implement 32-bit adders in the 90 nm technology node for VCTA and compare them with implementations that make use of standard cells. For this purpose we study the impact of photolithography proximity and coma effects on channel length variations, and the impact of shallow trench isolation mechanical stress on threshold voltage variations. We demonstrate that both variations, that are important sources of energy and delay circuit variability, are minimized through VCTA regularity.
Marc Pons 0001, Francesc Moll, Antonio Rubio 0001, Jaume Abella 0001, Xavier Vera, Antonio González 0001
VLSI-SoC5
2010 Microarchitectural Online Testing for Failure Detection in Memory Order Buffers
abstract
Technology scaling leads to burn-in phase out and higher postsilicon test complexity, which increases in-the-field failure rate due to both latent defects and actual errors, respectively. As a consequence, current reliability qualification methods will likely be infeasible. Microarchitecture knowledge of application runtime behavior offers a possibility to have low-cost continuous online testing techniques detect hard errors in the field. Whereas data can be protected with redundancy (like parity or ECC), there is a lack of mechanism for control logic. This paper proposes a microarchitectural approach for validating that the memory order buffer logic works correctly. Our design relies on a small cache-like structure that keeps track of the last store to each cached address. Each load is checked to have obtained the data from the youngest older producing store. We present three different implementations of this idea, offering different trade-offs for error coverage, performance overhead, and design complexity.
Javier Carretero, Xavier Vera, Pedro Chaparro, Jaume Abella 0001
IEEE Trans. Computers2
2009 DFx for massively multiprocessors
Xavier Vera
IOLTS1
2009 Online error detection and correction of erratic bits in register files
abstract
Aggressive voltage scaling needed for low power in each new process generation causes large deviations in the threshold voltage of minimally sized devices of the 6T SRAM cell. Gate oxide scaling can cause large transient gate leakage (a trap in the gate oxide), which is known as the erratic bits phenomena. Register file protection is necessary to prevent errors from quickly spreading to different parts of the system, which may cause applications to crash or silent data corruption. This paper proposes a simple and cost-effective mechanism that increases the resiliency of the register files to erratic bits. Our mechanism detects those registers that have erratic bits, recovers from the error and quarantines the faulty register. After the quarantine period, it is able to detect whether they are fully operational with low overhead.
Xavier Vera, Jaume Abella 0001, Javier Carretero, Pedro Chaparro, Antonio González 0001
IOLTS1
2009 End-to-end register data-flow continuous self-test
abstract
While Moore's Law predicts the ability of semi-conductor industry to engineer smaller and more efficient transistors and circuits, there are serious issues not contemplated in that law. One concern is the verification effort of modern computing systems, which has grown to dominate the cost of system design. On the other hand, technology scaling leads to burn-in phase out. As a result, in-the-field error rate may increase due to both actual errors and latent defects. Whereas data can be protected with arithmetic codes (like parity or ECC), there is a lack of cost-effective mechanisms for control logic.
Javier Carretero, Pedro Chaparro, Xavier Vera, Jaume Abella 0001, Antonio González 0001
ISCA3
2009 Low Vccmin fault-tolerant cache with highly predictable performance
abstract
Transistors per area unit double in every new technology node. However, the electric field density and power demand grow if Vcc is not scaled. Therefore, Vcc must be scaled in pace with new technology nodes to prevent excessive degradation and keep power demand within reasonable limits. Unfortunately, low Vcc operation exacerbates the effect of variations and decreases noise and stability margins, increasing the likelihood of errors in SRAM memories such as caches. Those errors translate into performance loss and performance variation across different cores, which is especially undesirable in a multi-core processor.
Jaume Abella 0001, Javier Carretero, Pedro Chaparro, Xavier Vera, Antonio González 0001
MICRO4
2009 Reducing Soft Errors through Operand Width Aware Policies
abstract
Soft errors are an important challenge in contemporary microprocessors. Particle hits on the components of a processor are expected to create an increasing number of transient errors with each new microprocessor generation. In this paper, we propose simple mechanisms that effectively reduce the vulnerability to soft errors in a processor. Our designs are generally motivated by the fact that many of the produced and consumed values in the processors are narrow and their upper order bits are meaningless. Soft errors caused by any particle strike to these higher order bits can be avoided by simply identifying these narrow values. Alternatively, soft errors can be detected or corrected on the narrow values by replicating the vulnerable portion of the value inside the storage space provided for the upper order bits of these operands. As a faster but less fault tolerant alternative to ECC and parity, we offer a variety of schemes that make use of narrow values and analyze their efficiency in reducing soft error vulnerability of different data-holding components of a processor. On average, techniques that make use of the narrowness of the values can provide 49 percent error detection, 45 percent error correction, or 27 percent error avoidance coverage for single bit upsets in the first level data cache across all Spec2K. In other structures such as the immediate field of the issue queue, an average error detection rate of 64 percent is achieved.
Oguz Ergin, Osman S. Unsal, Xavier Vera, Antonio González 0001
IEEE Trans. Dependable Secur. Comput.3
2009 Selective replication: A lightweight technique for soft errors
abstract
Soft errors are an important challenge in contemporary microprocessors. Modern processors have caches and large memory arrays protected by parity or error detection and correction codes. However, today's failure rate is dominated by flip flops, latches, and the increasing sensitivity of combinational logic to particle strikes. Moreover, as Chip Multi-Processors (CMPs) become ubiquitous, meeting the FIT budget for new designs is becoming a major challenge. Solutions based on replicating threads have been explored deeply; however, their high cost in performance and energy make them unsuitable for current designs. Moreover, our studies based on a typical configuration for a modern processor show that focusing on the top 5 most vulnerable structures can provide up to 70% reduction in FIT rate. Therefore, full replication may overprotect the chip by reducing the FIT much below budget. We propose Selective Replication , a lightweight-reconfigurable mechanism that achieves a high FIT reduction by protecting the most vulnerable instructions with minimal performance and energy impact. Low performance degradation is achieved by not requiring additional issue slots and reissuing instructions only during the time window between when they are retirable and they actually retire. Coverage can be reconfigured online by replicating only a subset of the instructions (the most vulnerable ones). Instructions' vulnerability is estimated based on the area they occupy and the time they spend in the issue queue. By changing the vulnerability threshold, we can adjust the trade-off between coverage and performance loss. Results for an out-of-order processor configured similarly to Intel® Core™ Micro-Architecture show that our scheme can achieve over 65% FIT reduction with less than 4% performance degradation with small area and complexity overhead.
Xavier Vera, Jaume Abella 0001, Javier Carretero, Antonio González 0001
ACM Trans. Comput. Syst.1
2008 Issue system protection mechanisms
abstract
Multi-core microprocessors require reducing the FIT (failures-in-time) rate per core drastically to enable a larger number of cores within a FIT budget. Since large arrays like caches and register flies are typically protected with either ECC or parity, the issue system becomes as one of the largest contributors to the core's FIT rate. Soft-errors are an important concern in contemporary microprocessors. Particle hits on the components of a processor are expected to create an increasing number of transient errors in each new microprocessor generation. In addition, the number of hard-errors in the field is expected to grow as burn-in becomes less effective. Moreover, the continuous device shrinking increases the likelihood of in-the-field failures due to rather small defects exacerbated by degradation. This paper proposes on-line mechanisms to detect and recover to a consistent state, classify and confine in-the-field errors in the issue system of both in-order and out-of-order cores. Such mechanisms provide high coverage at a small cost.
Pedro Chaparro, Jaume Abella 0001, Javier Carretero, Xavier Vera
ICCD4
2008 On-Line Failure Detection and Confinement in Caches
abstract
Technology scaling leads to burn-in phase out and increasing post-silicon test complexity, which increases in-the-field error rate due to both latent defects and actual errors. As a consequence, there is an increasing need for continuous on-line testing techniques to cope with hard errors in the field. Similarly, those techniques are needed for detecting soft errors in logic, whose error rate is expected to raise in future technologies. Cache memories, which occupy most of the area of the chip, are typically protected with parity or ECC, but most of the wires as well as some combinational blocks remain unprotected against both soft and hard errors. This paper presents a set of techniques to detect and confine hard and soft errors in cache memories in combination with parity/ECC at very low cost. By means of hard signatures in data rows and error tracking, faults can be detected, classified properly and confined for hardware reconfiguration.
Jaume Abella 0001, Pedro Chaparro, Xavier Vera, Javier Carretero, Antonio González 0001
IOLTS3
2008 On-line Failure Detection in Memory Order Buffers
abstract
Technology scaling leads to burn-in phase out and higher post-silicon test complexity, which increases in-the-field error rate due to both latent defects and actual errors respectively. As a consequence, current reliability qualification methods will likely be infeasible. Microarchitecture knowledge of application runtime behavior offers a possibility to have low-cost continuous online testing techniques to cope with hard errors in the field. Whereas data can be protected with redundancy (like parity or ECC), there is a lack of mechanisms for control logic. This paper proposes a microarchitectural approach for validating that the memory order buffer logic works correctly.
Javier Carretero, Xavier Vera, Pedro Chaparro, Jaume Abella 0001
ITC2
2007 Fuse: A Technique to Anticipate Failures due to Degradation in ALUs
abstract
This paper proposes the fuse, a technique to anticipate failures due to degradation in any ALU (arithmetic logic unit), and particularly in an adder. The fuse consists of a replica of the weakest transistor in the adder and the circuitry required to measure its degradation. By mimicking the behavior of the replicated transistor the fuse anticipates the failure short before the first failure in the adder appears, and hence, data corruption and program crashes can be avoided. Our results show that the fuse anticipates the failure in more than 99.9% of the cases after 96.6% of the lifetime, even for pessimistic random within-die variations.
Jaume Abella 0001, Xavier Vera, Osman S. Unsal, Oguz Ergin, Antonio González 0001
IOLTS2
2007 Surviving to Errors in Multi-Core Environments
abstract
In this paper, the authors present a global view of the issues outlined above as well as some directions to address them. First, the most important sources of failure (SOF) are presented as well as their impact on CMOS technology. Then, techniques and key parameters to measure degradation due to different SOF are introduced and microarchitectural approaches to mitigate degradation are outlined. The problem of error detection and anticipation is illustrated as well as pros and cons of different types of mechanisms to perform such detection and anticipation. Finally, we illustrate the whole picture where performance and reliability must be traded carefully. We point out some directions to use the information about the detected errors and the amount of degradation of each component to configure the multi-core in such a way that performance is maximized without compromising reliability.
Xavier Vera, Jaume Abella 0001
IOLTS1
2007 Penelope: The NBTI-Aware Processor
abstract
Transistors consist of lower number of atoms with every technology generation. Such atoms may be displaced due to the stress caused by high temperature, frequency and current, leading to failures. NBTI (negative bias temperature instability) is one of the most important sources of failure affecting transistors. NBTI degrades PMOS transistors whenever the voltage at the gate is negative (logic input "0"). The main consequence is a reduction in the maximum operating frequency and an increase in the minimum supply voltage of storage structures to cope for the degradation. Many PMOS transistors affected by NBTI can be found in both combinational and storage blocks since they observe a "0 " at their gates most of the time. This paper proposes and evaluates the design of Penelope, an NBTI-aware processor. We propose (i) generic strategies to mitigate degradation in both combinational and storage blocks, (ii) specific techniques to protect individual blocks by applying the global strategies, and (Hi) a metric to assess the benefits of reduced degradation and the overheads in performance and power.
Jaume Abella 0001, Xavier Vera, Antonio González 0001
MICRO2
2007 Data cache locking for tight timing calculations
abstract
Caches have become increasingly important with the widening gap between main memory and processor speeds. Small and fast cache memories are designed to bridge this discrepancy. However, they are only effective when programs exhibit sufficient data locality. In addition, caches are a source of unpredictability, resulting in programs sometimes behaving in a different way than expected. Detailed information about the number of cache misses and their causes allows us to predict cache behavior and to detect bottlenecks. Small modifications in the source code may change memory patterns, thereby altering the cache behavior. Code transformations, which take the cache behavior into account, might result in a high cache performance improvement. However, cache memory behavior is very hard to predict, thus making the task of optimizing and timing cache behavior very difficult. This article proposes and evaluates a new compiler framework that times cache behavior for multitasking systems. Our method explores the use of cache partitioning and dynamic cache locking to provide worst-case performance estimates in a safe and tight way for multitasking systems. We use cache partitioning, which divides the cache among tasks to eliminate intertask cache interferences. We combine static cache analysis and cache-locking mechanisms to ensure that all intratask conflicts, and consequently, memory access times, are exactly predictable. The results of our experiments demonstrate the capability of our framework to describe cache behavior at compile time. We compare our timing approach with a system equipped with a nonpartitioned, but statically, locked data cache. Our method outperforms static cache locking for all analyzed task sets under various cache architectures, demonstrating that our fully predictable scheme does not compromise the performance of the transformed programs.
Xavier Vera, Björn Lisper, Jingling Xue
ACM Trans. Embed. Comput. Syst.1
2006 Empowering a helper cluster through data-width aware instruction selection policies
abstract
Narrow values that can be represented by less number of bits than the full machine width occur very frequently in programs. On the other hand, clustering mechanisms enable cost- and performance-effective scaling of processor back-end features. Those attributes can be combined synergistically to design special clusters operating on narrow values (a.k.a. helper cluster), potentially providing performance benefits. We complement a 32-bit monolithic processor with a low-complexity 8-bit helper cluster. Then, in our main focus, we propose various ideas to select suitable instructions to execute in the data-width based clusters. We add data-width information as another instruction steering decision metric and introduce new data-width based selection algorithms which also consider dependency, inter-cluster communication and load imbalance. Utilizing those techniques, the performance of a wide range of workloads are substantially increased; helper cluster achieves an average speedup of 11% for a wide range of 412 apps. When focusing on integer applications, the speedup can be as high as 22% on average
Osman S. Unsal, Oguz Ergin, Xavier Vera, Antonio González 0001
IPDPS3
2005 IATAC: a smart predictor to turn-off L2 cache lines
abstract
As technology evolves, power dissipation increases and cooling systems become more complex and expensive. There are two main sources of power dissipation in a processor: dynamic power and leakage. Dynamic power has been the most significant factor, but leakage will become increasingly significant in future. It is predicted that leakage will shortly be the most significant cost as it grows at about a 5× rate per generation. Thus, reducing leakage is essential for future processor design. Since large caches occupy most of the area, they are one of the leakiest structures in the chip and hence, a main source of energy consumption for future processors.This paper introduces IATAC (inter-access time per access count), a new hardware technique to reduce cache leakage for L2 caches. IATAC dynamically adapts the cache size to the program requirements turning off cache lines whose content is not likely to be reused. Our evaluation shows that this approach outperforms all previous state-of-the-art techniques. IATAC turns off 65% of the cache lines across different L2 cache configurations with a very small performance degradation of around 2%.
Jaume Abella 0001, Antonio González 0001, Xavier Vera, Michael F. P. O'Boyle
ACM Trans. Archit. Code Optim.3
2005 An accurate cost model for guiding data locality transformations
abstract
Caches have become increasingly important with the widening gap between main memory and processor speeds. Small and fast cache memories are designed to bridge this discrepancy. However, they are only effective when programs exhibit sufficient data locality.The performance of the memory hierarchy can be improved by means of data and loop transformations. Tiling is a loop transformation that aims at reducing capacity misses by shortening the reuse distance. Padding is a data layout transformation targeted to reduce conflict misses.This article presents an accurate cost model that describes misses across different hierarchy levels and considers the effects of other hardware components such as branch predictors. The cost model drives the application of tiling and padding transformations. We combine the cost model with a genetic algorithm to compute the tile and pad factors that enhance the program performance.To validate our strategy, we ran experiments for a set of benchmarks on a large set of modern architectures. Our results show that this scheme is useful to optimize programs' performance. When compared to previous approaches, we observe that with a reasonable compile-time overhead, our approach gives significant performance improvements for all studied kernels on all architectures.
Xavier Vera, Jaume Abella 0001, Josep Llosa, Antonio González 0001
ACM Trans. Program. Lang. Syst.1
2004 Efficient and Accurate Analytical Modeling of Whole-Program Data Cache Behavior
abstract
Data caches are a key hardware means to bridge the gap between processor and memory speeds, but only for programs that exhibit sufficient data locality in their memory accesses. Thus, a method for evaluating cache performance is required to both determine quantitatively cache misses and to guide data cache optimizations. Existing analytical models for data cache optimizations target mainly isolated perfect loop nests. We present an analytical model that is capable of statically analyzing not only loop nest fragments, but also complete numerical programs with regular and compile-time predictable memory accesses. Central to the whole-program approach are abstract call inlining, memory access vectors, and parametric reuse analysis, which allow the reuse and interference both within and across loop nests to be quantified precisely in a unified framework. Based on the framework, the cache misses of a program are specified using mathematical formulas and the miss ratio is predicted from these formulas based on statistical sampling techniques. Our experimental results using kernels and whole programs indicate accurate cache miss estimates in a substantially shorter amount of time (typically, several orders of magnitude faster) than simulation.
Jingling Xue, Xavier Vera
IEEE Trans. Computers2
2004 A fast and accurate framework to analyze and optimize cache memory behavior
abstract
The gap between processor and main memory performance increases every year. In order to overcome this problem, cache memories are widely used. However, they are only effective when programs exhibit sufficient data locality. Compile-time program transformations can significantly improve the performance of the cache. To apply most of these transformations, the compiler requires a precise knowledge of the locality of the different sections of the code, both before and after being transformed.Cache miss equations (CMEs) allow us to obtain an analytical and precise description of the cache memory behavior for loop-oriented codes. Unfortunately, a direct solution of the CMEs is computationally intractable due to its NP-complete nature.This article proposes a fast and accurate approach to estimate the solution of the CMEs. We use sampling techniques to approximate the absolute miss ratio of each reference by analyzing a small subset of the iteration space. The size of the subset, and therefore the analysis time, is determined by the accuracy selected by the user. In order to reduce the complexity of the algorithm to solve CMEs, effective mathematical techniques have been developed to analyze the subset of the iteration space that is being considered. These techniques exploit some properties of the particular polyhedra represented by CMEs.
Xavier Vera, Nerina Bermudo, Josep Llosa, Antonio González 0001
ACM Trans. Program. Lang. Syst.1
2003 Code Tiling for Improving the Cache Performance of PDE Solvers
abstract
For SOR-like PDE solvers, loop tiling either helps little in improving data locality or hurts their performance. We present a novel compiler technique called code tiling for generating fast tiled codes for these solvers on uniprocessors with a memory hierarchy. Code tiling combines loop tiling with a new array layout transformation called data tiling in such a way that a significant amount of cache misses that would otherwise be present in tiled codes are eliminated. Compared to nine existing loop tiling algorithms, our technique delivers impressive performance speedups (faster by factors of 1.55-2.62) and smooth performance curves across a range of problem sizes on representative machine architectures. The synergy of loop tiling and data tiling allows us to find a problem-size-independent tile size that minimises a cache miss objective function independently of the problem size parameters. This "one-size-fits-all" scheme makes our approach attractive for designing fast SOR solvers without having to generate a multitude of versions specialised for different problem sizes.
Qingguang Huang, Jingling Xue, Xavier Vera
ICPP3
2003 Data Caches in Multitasking Hard Real-Time Systems
abstract
Data caches are essential in modern processors, bridging the widening gap between main memory and processor speeds. However, they yield very complex performance models, which make it hard to bound execution times tightly. This paper contributes a new technique to obtain predictability in preemptive multitasking systems in the presence of data caches. We explore the use of cache partitioning, dynamic cache locking, and static cache analysis to provide worst-case performance estimates in a safe and tight way. Cache partitioning divides the cache among tasks to eliminate inter-task cache interferences. We combine static cache analysis and cache locking mechanisms to ensure that all intra-task conflicts, and consequently, memory access times, are exactly predictable. To minimize the performance degradation due to cache partitioning and locking, two strategies are employed. First, the cache is loaded with data likely to be accessed so that their cache utilization is maximized. Second, compiler optimizations such as tiling and padding are applied in order to reduce cache replacement misses. Experimental results show that this scheme is fully predictable, without compromising the performance of the transformed programs. Our method outperforms static cache locking for all analyzed task sets under various cache architectures, with a CPU utilization reduction ranging between 3.8 and 20.0 times for a high performance system.
Xavier Vera, Björn Lisper, Jingling Xue
RTSS1
2003 Data cache locking for higher program predictability
abstract
Caches have become increasingly important with the widening gap between main memory and processor speeds. However, they are a source of unpredictability due to their characteristics, resulting in programs behaving in a different way than expected.Cache locking mechanisms adapt caches to the needs of real-time systems. Locking the cache is a solution that trades performance for predictability: at a cost of generally lower performance, the time of accessing the memory becomes predictable.This paper combines compile-time cache analysis with data cache locking to estimate the worst-case memory performance (WCMP) in a safe, tight and fast way. In order to get predictable cache behavior, we first lock the cache for those parts of the code where the static analysis fails. To minimize the performance degradation, our method loads the cache, if necessary, with data likely to be accessed.Experimental results show that this scheme is fully predictable, without compromising the performance of the transformed program. When compared to an algorithm that assumes compulsory misses when the state of the cache is unknown, our approach eliminates all overestimation for the set of benchmarks, giving an exact WCMP of the transformed program without any significant decrease in performance.
Xavier Vera, Björn Lisper, Jingling Xue
SIGMETRICS1
2002 Let's Study Whole-Program Cache Behaviour Analytically
abstract
Based on a new characterisation of data reuse across multiple loop nests, we preset a method, a prototyping implementation and some experimental results for analysing the cache behaviour of whole programs with regular computations. Validation against cache simulation using real codes shows the efficiency and accuracy of our method. The largest program, we have analysed, Applu from SPECfP95, has 3868 lines, 16 subroutines and 2565 references. In the case of a 32KB cache with a 32B line size, our method obtains the miss ratio with an absolute error of about 0.80% in about 128 seconds while the simulator used runs for nearly 5 hours on a 933MHz Pentium. III PC. Our method can be used to guide compiler locality optimisations and improve cache simulation performance.
Xavier Vera, Jingling Xue
HPCA1
2000 A Fast and Accurate Approach to Analyze Cache Memory Behavior (Research Note)
Xavier Vera, Josep Llosa, Antonio González 0001, Nerina Bermudo
Euro-Par1
2000 An efficient solver for Cache Miss Equations
abstract
Cache Miss Equations (CME) (S. Ghosh et al., 1997) is a method that accurately describes the cache behavior by means of polyhedra. Even though the computation cost of generating CME is a linear function of the number of references, solving them is a very time consuming task and thus trying to study a whole program may be infeasible. The paper presents effective techniques that exploit some properties of the particular polyhedra generated by CME. Such techniques reduce the complexity of the algorithm to solve CME, which results in a significant speedup when compared with traditional methods. In particular, the proposed approach does not require the computation of the vertices of each polyhedron, which has an exponential complexity.
Nerina Bermudo, Xavier Vera, Antonio González 0001, Josep Llosa
ISPASS2