Pablo Montesinos

dblp:89/6993 · also Pablo Montesinos-Ortego · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorSecurity and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Parallel and multicore computing · 48% Processor architecture and microarchitecture · 16% Memory systems · 14%
Software engineering, system software, and programming languages
4 papers
Operating systems · 56% Concurrent programming · 44%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Operating systems
mobile systems
0.212013
ZOOMM: a parallel web browser engine for multicore mobile devices · PPoPP 2013
Processor architecture and microarchitecture › debugging support
hardware-assisted deterministic replay
0.112009
Capo: a software-hardware interface for practical deterministic multiprocessor replay · ASPLOS 2009
Parallel and multicore computing › concurrent programming
concurrency bugs
0.112008
DeLorean: Recording and Deterministically Replaying Shared-Memory Multiprocessor Execution Effciently · ISCA 2008
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay
0.112008
DeLorean: Recording and Deterministically Replaying Shared-Memory Multiprocessor Execution Effciently · ISCA 2008
Concurrent programming › synchronization
data-centric synchronization
0.112007
Colorama: Architectural Support for Data-Centric Synchronization · HPCA 2007
Concurrent programming
synchronization
0.112007
Colorama: Architectural Support for Data-Centric Synchronization · HPCA 2007
Memory systems › memory consistency
memory consistency model
0.112007
BulkSC: bulk enforcement of sequential consistency · ISCA 2007
Memory systems › memory consistency › memory consistency model
sequential consistency
0.112007
BulkSC: bulk enforcement of sequential consistency · ISCA 2007
Parallel and multicore computing
transactional memory
0.112007
Colorama: Architectural Support for Data-Centric Synchronization · HPCA 2007
Distributed systems
fault tolerance
0.112006
ReViveI/O: efficient handling of I/O in highly-available rollback-recovery servers · HPCA 2006
Distributed systems › fault tolerance
rollback recovery
0.112006
ReViveI/O: efficient handling of I/O in highly-available rollback-recovery servers · HPCA 2006
Operating systems › resource management › process management
multiprogramming
0.012009
Capo: a software-hardware interface for practical deterministic multiprocessor replay · ASPLOS 2009
Processor architecture and microarchitecture
atomic block execution
0.012008
DeLorean: Recording and Deterministically Replaying Shared-Memory Multiprocessor Execution Effciently · ISCA 2008
Processor architecture and microarchitecture › multithreading
speculative multithreading
0.012008
DeLorean: Recording and Deterministically Replaying Shared-Memory Multiprocessor Execution Effciently · ISCA 2008
Concurrent programming
parallel programming models
0.012007
Colorama: Architectural Support for Data-Centric Synchronization · HPCA 2007
Processor architecture and microarchitecture
chip multiprocessor
0.012007
BulkSC: bulk enforcement of sequential consistency · ISCA 2007
Operating systems › fault tolerance
checkpoint and rollback
0.012006
ReViveI/O: efficient handling of I/O in highly-available rollback-recovery servers · HPCA 2006

Methods — techniques the papers use, named apart from their topics

resource preloading · 0.3concurrency management · 0.3software-hardware interface · 0.2hardware extension · 0.1critical section inference · 0.1simulation · 0.1logging · 0.1
YearPublicationVenuePosition
2013 ZOOMM: a parallel web browser engine for multicore mobile devices
abstract
We explore the challenges in expressing and managing concurrency in browsers on mobile devices. Browsers are complex applications that implement multiple standards, need to support legacy behavior, and are highly dynamic and interactive. We present ZOOMM, a highly concurrent web browser engine prototype and show how concurrency is effectively exploited at different levels: speed up computation performance, preload network resources, and preprocess resources outside the critical path of page loading. On a dual-core Android mobile device we demonstrate that ZOOMM is two times faster than the native WebKit based browser when loading the set of pages defined in the Vellamo benchmark.
Calin Cascaval, Seth Fowler, Pablo Montesinos, Wayne Piekarski, Mehrdad Reshadi, Behnam Robatmili, Michael Weber 0002, Vrajesh Bhavsar
PPoPP3
2009 Capo: a software-hardware interface for practical deterministic multiprocessor replay
abstract
While deterministic replay of parallel programs is a powerful technique, current proposals have shortcomings. Specifically, software-based replay systems have high overheads on multiprocessors, while hardware-based proposals focus only on basic hardware-level mechanisms, ignoring the overall replay system. To be practical, hardware-based replay systems need to support an environment with multiple parallel jobs running concurrently -- some being recorded, others being replayed and even others running without recording or replay. Moreover, they need to manage limited-size log buffers.
Pablo Montesinos, Matthew Hicks, Samuel T. King, Josep Torrellas
ASPLOS1
2008 DeLorean: Recording and Deterministically Replaying Shared-Memory Multiprocessor Execution Effciently
abstract
Support for deterministic replay of multithreaded execution can greatly help in finding concurrency bugs. For highest effectiveness, replay schemes should (i) record at production-run speed, (ii) keep their logging requirements minute, and (iii) replay at a speed similar to that of the initial execution. In this paper, we propose a new substrate for deterministic replay that provides substantial advances along these axes. In our proposal, processors execute blocks of instructions atomically, as in transactional memory or speculative multithreading, and the system only needs to record the commit order of these blocks. We call our scheme DeLorean. Our results show that DeLorean records execution at a speed similar to that of release consistency (RC) execution and replays at about 82% of its speed. In contrast, most current schemes only record at the speed of Sequential Consistency (SC) execution. Moreover, DeLorean only needs 7.5% of the log size needed by a state-of-the-art scheme. Finally, DeLorean can be configured to need only 0.6% of the log size of the state-of-the-art scheme at the cost of recording at 86% of RCpsilas execution speed - still faster than SC. In this configuration, the log of an 8-processor 5-GHz machine is estimated to be only about 20GB per day.
Pablo Montesinos, Luis Ceze, Josep Torrellas
ISCA1
2007 Using Register Lifetime Predictions to Protect Register Files against Soft Errors
abstract
To increase the resistance of register files to soft errors, this paper presents the ParShield architecture. ParShield is based on two observations: (i) the data in a register is only useful for a small fraction of the register's lifetime, and (ii) not all registers are equally vulnerable. ParShield selectively protects registers by generating, storing, and checking the ECCs of only the most vulnerable registers while they contain useful data. In addition, it stores a parity bit for all the registers, re-using the ECC circuitry for parity generation and checking. ParShield has no SDC AVF and a small average DUE AVF of 0.040 and 0.010 for the integer and floating-point register files, respectively. ParShield consumes on average only 81% and 78% of the power of a design with full ECC for the SPECint and SPECfp applications, respectively. Finally, ParShield has no performance impact and little area requirements.
Pablo Montesinos, Wei Liu 0014, Josep Torrellas
DSN1
2007 Colorama: Architectural Support for Data-Centric Synchronization
abstract
With the advent of ubiquitous multi-core architectures, a major challenge is to simplify parallel programming. One way to tame one of the main sources of programming complexity, namely synchronization, is transactional memory (TM). However, we argue that TM does not go far enough, since the programmer still needs nonlocal reasoning to decide where to place transactions in the code. A significant improvement to the art is data-centric synchronization (DCS), where the programmer uses local reasoning to assign synchronization constraints to data. Based on these, the system automatically infers critical sections and inserts synchronization operations. This paper proposes novel architectural support to make DCS feasible, and describes its programming model and interface. The proposal, called Colorama, needs only modest hardware extensions, supports general-purpose, pointer-based languages such as C/C++ and, in our opinion, can substantially simplify the task of writing new parallel programs
Luis Ceze, Pablo Montesinos, Christoph von Praun, Josep Torrellas
HPCA2
2007 BulkSC: bulk enforcement of sequential consistency
abstract
While Sequential Consistency (SC) is the most intuitive memory consistency model and the one most programmers likely assume, current multiprocessors do not support it. Instead, they support more relaxed models that deliver high performance. SC implementations are considered either too slow or -- when they can match the performance of relaxed models -- too difficult to implement.
Luis Ceze, James Tuck 0001, Pablo Montesinos, Josep Torrellas
ISCA3
2006 ReViveI/O: efficient handling of I/O in highly-available rollback-recovery servers
abstract
The increasing demand for reliable computers has led to proposals for hardware-assisted rollback of memory state. Such approach promises major reductions in mean time to repair (MTTR). The benefits are especially compelling for database servers, where existing recovery software typically leads to downtimes of tens of minutes. Unfortunately, adoption of such proposals is hindered by the lack of efficient mechanisms for I/O recovery. This paper presents and evaluates ReViveI/O, a scheme for I/O undo and redo that is compatible with mechanisms for hardware-assisted rollback of memory state. We have fully implemented a Linux-based prototype that shows that low-overhead, low-MTTR recovery of I/O is feasible. For 20-120 ms between checkpoints, a throughput-oriented workload such as TPC-C has negligible overhead. Moreover, for 50 ms or less between checkpoints, the response time of a latency-bound workload such as WebStone remains tolerable. In all cases, the recovery time of ReViveI/O is practically negligible. The result is a cost-effective highly-available server.
Jun Nakano, Pablo Montesinos, Kourosh Gharachorloo, Josep Torrellas
HPCA2