Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Enric Herrero

dblp:29/3485 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
0since 2021 · last 2013
0000-0001-7837-3593ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 88% Processor architecture and microarchitecture · 12%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache management › storage caching
cooperative caching
0.322012
Distributed Cooperative Caching: An Energy Efficient Memory Scheme for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Elastic cooperative caching: an autonomous dynamically adaptive memory hierarchy for chip multiprocessors · ISCA 2010
Memory systems
DRAM
0.212013
Thread Row Buffers: Improving Memory Performance Isolation and Throughput in Multiprogrammed Environments · IEEE Trans. Computers 2013
Memory systems › DRAM › DRAM microarchitecture
row buffer management
0.212013
Thread Row Buffers: Improving Memory Performance Isolation and Throughput in Multiprogrammed Environments · IEEE Trans. Computers 2013
Memory systems › memory hierarchy
cache hierarchy
0.112012
Distributed Cooperative Caching: An Energy Efficient Memory Scheme for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Memory systems
cache
0.112010
Elastic cooperative caching: an autonomous dynamically adaptive memory hierarchy for chip multiprocessors · ISCA 2010
Memory systems
memory hierarchy
0.112010
Elastic cooperative caching: an autonomous dynamically adaptive memory hierarchy for chip multiprocessors · ISCA 2010
Processor architecture and microarchitecture
chip multiprocessor
0.122012
Distributed Cooperative Caching: An Energy Efficient Memory Scheme for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Elastic cooperative caching: an autonomous dynamically adaptive memory hierarchy for chip multiprocessors · ISCA 2010
Processor architecture and microarchitecture › many-core architecture
tiled microarchitecture
0.122012
Distributed Cooperative Caching: An Energy Efficient Memory Scheme for Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2012
Elastic cooperative caching: an autonomous dynamically adaptive memory hierarchy for chip multiprocessors · ISCA 2010
Memory systems › memory controller
memory scheduling
0.012013
Thread Row Buffers: Improving Memory Performance Isolation and Throughput in Multiprogrammed Environments · IEEE Trans. Computers 2013

Methods — techniques the papers use, named apart from their topics

service partitioning · 0.2simulation · 0.1dynamic adaptation · 0.1
YearPublicationVenuePosition
2013 Capturing vulnerability variations for register files
abstract
Soft error rates are estimated based on worst-case architectural vulnerability factor (AVF). Therefore, it makes tracking real-time accurate AVF very attractive to computer designers: more accurate AVF numbers will allow turning on more features at runtime while keeping the promised SDC and DUE rates. This paper presents a hardware mechanism based on linear regressions to estimate the AVF (SDC and DUE) of the register file for out-of-order cores. Our results show that we are able to have a high correlation factor at low cost.
Javier Carretero, Enric Herrero, Matteo Monchiero, Tanausú Ramírez, Xavier Vera
DATE2
2013 Thread Row Buffers: Improving Memory Performance Isolation and Throughput in Multiprogrammed Environments
abstract
The widespread adoption of chip multiprocessors in recent years has increased the number of applications simultaneously accessing DRAM memories. Therefore, memory access patterns have also changed and this has reduced row buffer locality significantly, degrading performance and energy efficiency. Furthermore, concurrent execution of applications also has shown the need of performance isolation among threads in the memory controller to enforce a quality of service in virtualized environments. Existing DRAM memories, however, enforce a tradeoff between throughput and isolation. To solve these problems, this paper proposes the addition of Thread Row Buffers (TRBs) to DRAM memories. TRBs keep an active row per thread, thereby increasing DRAM efficiency by avoiding alternate accesses to a limited number of rows and allowing the implementation of a memory scheduler not bound to the throughput-isolation tradeoff. Thread Row Buffers with Service Partitioning (TRB-SP) increase the row hit-rate by 38 percent with respect to FR-FCFS and by 11 percent with respect to Cache DRAM. This, in turn, increases overall performance by 17 and 7 percent, respectively. TRB-SP is also able to reduce the standard deviation of the memory access time of an application by 40 percent over FR-FCFS, 31 percent over PAR-BS, and 42 percent over Cache DRAM.
Enric Herrero, José González 0002, Ramon Canal, Dean M. Tullsen
IEEE Trans. Computers1
2012 Distributed Cooperative Caching: An Energy Efficient Memory Scheme for Chip Multiprocessors
abstract
Current trends in CMPs indicate that the core count will increase in the near future. One of the main performance limiters of these forthcoming microarchitectures is the latency and high demand of the on-chip network and the off-chip memory communication. One of the main trade-offs when searching an optimal cache hierarchy is the sharing degree of cache space and its on-die distribution. Several techniques have appeared recently that optimize these parameters to get a better performance. This work provides some insight in the most promising configurations for tiled microarchitectures and shows the advantages and limitations of each of them in terms of performance and energy efficiency. This paper extends previous works by providing a complete study that evaluates different network topologies, single and multithreaded benchmarks, and single and multiprogrammed execution. In all these studies, the Distributed Cooperative Caching shows to be a promising alternative to traditional configurations for chip multiprocessors, providing a scalable and energy efficient solution.
Enric Herrero, José González 0002, Ramon Canal
IEEE Trans. Parallel Distributed Syst.1
2011 New reliability mechanisms in memory design for sub-22nm technologies
abstract
The TRAMS (Terascale Reliable Adaptive MEMORY Systems) project addresses in an evolutionary way the ultimate CMOS scaling technologies and paves the way for revolutionary, most promising beyond-CMOS technologies. In this abstract we show the significant variability levels of future 18 and 13nm device bulk-CMOS technologies as well as its dramatic effect on the yield of memory cells, and what kind of circuit solution would be required to maintain the current yield level. Later, we discuss the impact of errors at the system level, and different approaches at system level to adapt the heterogeneous systems to user's requirements.
Nivard Aymerich, A. Asenov, Andrew R. Brown, Ramon Canal, Binjie Cheng, Joan Figueras, Antonio González 0001, Enric Herrero, S. Markov, Miguel Corbalan, Peyman Pouyan, Tanausú Ramírez, Antonio Rubio 0001, Elena I. Vatajelu, Xavier Vera, Xingsheng Wang, Paul Zuber
IOLTS8
2010 Power-Efficient Spilling Techniques for Chip Multiprocessors
Enric Herrero, José González 0002, Ramon Canal
Euro-Par (1)1
2010 Elastic cooperative caching: an autonomous dynamically adaptive memory hierarchy for chip multiprocessors
abstract
Next generation tiled microarchitectures are going to be limited by off-chip misses and by on-chip network usage. Furthermore, these platforms will run an heterogeneous mix of applications with very different memory needs, leading to significant optimization opportunities. Existing adaptive memory hierarchies use either centralized structures that limit the scalability or software based resource allocation that increases programming complexity.
Enric Herrero, José González 0002, Ramon Canal
ISCA1
2008 Distributed cooperative caching
abstract
This paper presents the Distributed Cooperative Caching, a scalable and energy-efficient scheme to manage chip multiprocessor (CMP) cache resources. The proposed configuration is based in the Cooperative Caching framework [3] but it is intended for large scale CMPs. Both centralized and distributed configurations have the advantage of combining the benefits of private and shared caches. In our proposal, the Coherence Engine has been redesigned to allow its partitioning and thus, eliminate the size constraints imposed by the duplication of all tags. At the same time, a global replacement mechanism has been added to improve the usage of cache space. Our framework uses several Distributed Coherence Engines spread across all the nodes to improve scalability. The distribution permits a better balance of the network traffic over the entire chip avoiding bottlenecks and increasing performance for a 32-core CMP by 21% over a traditional shared memory configuration and by 57% over the Cooperative Caching scheme.
Enric Herrero, José González 0002, Ramon Canal
PACT1