Cristiano Pereira

dblp:36/6883 · also Cristiano L. Pereira · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14Software engineering, systems software and programming languages · 14

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
8 papers
Concurrent programming · 33% Debugging and program repair · 30% Software testing · 23%
Computer architecture, parallel and distributed computing, and storage systems
8 papers
Processor architecture and microarchitecture · 44% Memory systems · 15% Parallel and multicore computing · 15%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Debugging and program repair
fault localization
0.322017
AsyncClock: Scalable Inference of Asynchronous Event Causality · ASPLOS 2017
Recording shared memory dependencies using strata · ASPLOS 2006
Software testing
concurrency testing
0.322013
Selective mutation testing for concurrent code · ISSTA 2013
Maple: a coverage-driven testing tool for multithreaded programs · OOPSLA 2012
Concurrent programming
concurrency bugs
0.312017
AsyncClock: Scalable Inference of Asynchronous Event Causality · ASPLOS 2017
Concurrent programming
concurrency bug detection
0.222014
Race detection for event-driven mobile applications · PLDI 2014
Maple: a coverage-driven testing tool for multithreaded programs · OOPSLA 2012
Parallel and multicore computing › parallel computing › parallel program debugging
deterministic replay
0.232013
QuickRec: prototyping an intel architecture extension for record and replay of multithreaded programs · ISCA 2013
CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011
Architecting a chunk-based memory race recorder in modern CMPs · MICRO 2009
Debugging and program repair › record and replay
deterministic replay
0.222013
Cyrus: unintrusive application-level record-replay for replay parallelism · ASPLOS 2013
Recording shared memory dependencies using strata · ASPLOS 2006
Processor architecture and microarchitecture
chip multiprocessor
0.222011
CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011
Architecting a chunk-based memory race recorder in modern CMPs · MICRO 2009
Processor architecture and microarchitecture › multicore design
memory race recording
0.222011
CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011
Architecting a chunk-based memory race recorder in modern CMPs · MICRO 2009
Processor architecture and microarchitecture
multicore design
0.222011
CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011
Architecting a chunk-based memory race recorder in modern CMPs · MICRO 2009
Debugging and program repair
root cause analysis
0.212015
Failure sketching: a technique for automated root cause diagnosis of in-production failures · SOSP 2015
Concurrent programming › concurrency bug detection
data race detection
0.212014
Race detection for event-driven mobile applications · PLDI 2014
Program analysis › concurrent program analysis
event-race detection
0.212014
Race detection for event-driven mobile applications · PLDI 2014
Concurrent programming
multithreaded debugging
0.212013
Cyrus: unintrusive application-level record-replay for replay parallelism · ASPLOS 2013
Software testing
mutation testing
0.212013
Selective mutation testing for concurrent code · ISSTA 2013
Processor architecture and microarchitecture › debugging support
hardware-assisted deterministic replay
0.212013
QuickRec: prototyping an intel architecture extension for record and replay of multithreaded programs · ISCA 2013
Software testing › concurrency testing
thread interleaving coverage
0.112012
Maple: a coverage-driven testing tool for multithreaded programs · OOPSLA 2012
Memory systems
cache coherence
0.122009
Architecting a chunk-based memory race recorder in modern CMPs · MICRO 2009
Offline symbolic analysis for multi-processor execution replay · MICRO 2009
Hardware reliability and fault tolerance
execution replay
0.112009
Offline symbolic analysis for multi-processor execution replay · MICRO 2009
Memory systems › cache coherence › cache coherence protocol
snoopy coherence
0.112009
Architecting a chunk-based memory race recorder in modern CMPs · MICRO 2009
Distributed systems › fault tolerance
failure diagnosis
0.112015
Failure sketching: a technique for automated root cause diagnosis of in-production failures · SOSP 2015
Distributed systems
fault tolerance
0.112015
Failure sketching: a technique for automated root cause diagnosis of in-production failures · SOSP 2015
Network security › intrusion detection and prevention › intrusion detection
intrusion analysis
0.012013
Cyrus: unintrusive application-level record-replay for replay parallelism · ASPLOS 2013
Parallel and multicore computing › thread-level parallelism
multithreaded applications
0.012013
QuickRec: prototyping an intel architecture extension for record and replay of multithreaded programs · ISCA 2013
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.012004
Leakage aware dynamic voltage scaling for real-time embedded systems · DAC 2004
Energy-efficient computing
power management
0.012004
Leakage aware dynamic voltage scaling for real-time embedded systems · DAC 2004
Embedded and real-time systems
real-time scheduling
0.012004
Leakage aware dynamic voltage scaling for real-time embedded systems · DAC 2004
Memory systems › memory consistency
memory consistency model
0.012011
CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011
Memory systems › memory consistency › memory consistency model
total store order
0.012011
CoreRacer: a practical memory race recorder for multicore x86 TSO processors · MICRO 2011
Processor architecture and microarchitecture
debugging support
0.012006
Recording shared memory dependencies using strata · ASPLOS 2006
Embedded and real-time systems
real-time embedded systems
0.012004
Leakage aware dynamic voltage scaling for real-time embedded systems · DAC 2004

Methods — techniques the papers use, named apart from their topics

happens-before analysis · 0.6failure sketching · 0.4automated diagnosis · 0.4hardware prototyping · 0.3dynamic race detection · 0.2selective mutation · 0.2mutation testing · 0.2schedule control · 0.1coverage-driven testing · 0.1simulation · 0.1hardware race recording · 0.1strata logging · 0.1symbolic analysis · 0.1
YearPublicationVenuePosition
2017 AsyncClock: Scalable Inference of Asynchronous Event Causality
abstract
Asynchronous programming model is commonly used in mobile systems and Web 2.0 environments. Asynchronous race detectors use algorithms that are an order of magnitude performance and space inefficient compared to conventional data race detectors. We solve this problem by identifying and addressing two important problems in reasoning about causality between asynchronous events.
Chun-Hung Hsiao, Satish Narayanasamy, Essam Muhammad Idris Khan, Cristiano Pereira, Gilles Pokam
ASPLOS4
2015 Failure Sketches: A Better Way to Debug
Baris Kasikci, Cristiano Pereira, Gilles Pokam, Benjamin Schubert, Madan Musuvathi, George Candea
HotOS2
2015 Failure sketching: a technique for automated root cause diagnosis of in-production failures
abstract
Developers spend a lot of time searching for the root causes of software failures. For this, they traditionally try to reproduce those failures, but unfortunately many failures are so hard to reproduce in a test environment that developers spend days or weeks as ad-hoc detectives. The shortcomings of many solutions proposed for this problem prevent their use in practice.
Baris Kasikci, Benjamin Schubert, Cristiano Pereira, Gilles Pokam, George Candea
SOSP3
2014 DrDebug: Deterministic Replay based Cyclic Debugging with Dynamic Slicing
Harish Patil, Cristiano Pereira, Gregory Lueck, Rajiv Gupta 0001, Iulian Neamtiu
CGO3
2014 Race detection for event-driven mobile applications
abstract
Mobile systems commonly support an event-based model of concurrent programming. This model, used in popular platforms such as Android, naturally supports mobile devices that have a rich array of sensors and user input modalities. Unfortunately, most existing tools for detecting concurrency errors of parallel programs focus on a thread-based model of concurrency. If one applies such tools directly to an event-based program, they work poorly because they infer false dependencies between unrelated events handled sequentially by the same thread.
Chun-Hung Hsiao, Cristiano Pereira, Jie Yu 0016, Gilles Pokam, Satish Narayanasamy, Peter M. Chen, Ziyun Kong, Jason Flinn
PLDI2
2013 Concurrent predicates: A debugging technique for every parallel programmer
abstract
To reduce the complexity of debugging multithreaded programs, researchers have developed many techniques that automatically detect bugs that arise from shared memory errors. These techniques can identify a wide range of bugs, but it can be challenging for a programmer to reproduce a specific bug that he or she is interested in using such techniques. This is because these techniques were not intended for individual bug reproduction but rather an exploratory search for possible bugs. To address this concern we present concurrent predicates (CPs) and concurrent predicate expressions (CPEs), which allow programmers to single out a specific bug by specifying the schedule and program state that must be satisfied for the bug to be reproduced. We present the recipes, that is, the mechanical processes, we use to reproduce data races, atomicity violations, and deadlocks with CP and CPE. We then show how these recipes apply to the diagnosis and reproduction of bugs from 13 handcrafted bugs, five real-world application bugs from RADBench, and three previously unresolved bugs from TBoost.STM, which now includes the fixes we generated using CP and CPE.
Justin Emile Gottschlich, Gilles Pokam, Cristiano Pereira, Youfeng Wu
PACT3
2013 Cyrus: unintrusive application-level record-replay for replay parallelism
abstract
Architectures for deterministic record-replay (R&R) of multithreaded code are attractive for program debugging, intrusion analysis, and fault-tolerance uses. However, very few of the proposed designs have focused on maximizing replay speed -- a key enabling property of these systems. The few efforts that focus on replay speed require intrusive hardware or software modifications, or target whole-system R&R rather than the more useful application-level R&R.
Nima Honarmand, Nathan Dautenhahn, Josep Torrellas, Samuel T. King, Gilles Pokam, Cristiano Pereira
ASPLOS6
2013 QuickRec: prototyping an intel architecture extension for record and replay of multithreaded programs
abstract
There has been significant interest in hardware-assisted deterministic Record and Replay (RnR) systems for multithreaded programs on multiprocessors. However, no proposal has implemented this technique in a hardware prototype with full operating system support. Such an implementation is needed to assess RnR practicality.
Gilles Pokam, Klaus Danne, Cristiano Pereira, Rolf Kassa, Tim Kranich, Shiliang Hu, Justin Emile Gottschlich, Nima Honarmand, Nathan Dautenhahn, Samuel T. King, Josep Torrellas
ISCA3
2013 Selective mutation testing for concurrent code
abstract
Concurrent code is becoming increasingly important with the advent of multi-cores, but testing concurrent code is challenging. Researchers are developing new testing techniques and test suites for concurrent code, but evaluating these techniques and test suites often uses a small number of real or manually seeded bugs.
Milos Gligoric 0001, Lingming Zhang 0001, Cristiano Pereira, Gilles Pokam
ISSTA3
2012 PinADX: an interface for customizable debugging with dynamic instrumentation
abstract
Dynamic binary instrumentation systems have become popular frameworks for building custom program analysis tools. For example, Pin [8], Valgrind [9], and DynamoRIO [5] have been used to build a variety of memory checking, thread checking, cache simulation, and profiling tools. However, there has been little emphasis on building custom debugging tools with these frameworks. We introduce PinADX, a set of debugging extensions to the Pin instrumentation framework that allow easy development of custom debugger features on both Linux and Windows. This paper presents the PinADX API, demonstrates its use in a simple debugger extension, discusses the implementation details in the Pin instrumentation system, and shows that the performance is better than a native debugger for some typical debugging tasks. We believe that dynamic instrumentation technology is a powerful tool for building advanced debugging features and that an easy to use API will greatly facilitate the development of such features.
Gregory Lueck, Harish Patil, Cristiano Pereira
CGO3
2012 Maple: a coverage-driven testing tool for multithreaded programs
abstract
Testing multithreaded programs is a hard problem, because it is challenging to expose those rare interleavings that can trigger a concurrency bug. We propose a new thread interleaving coverage-driven testing tool called Maple that seeks to expose untested thread interleavings as much as possible. It memoizes tested interleavings and actively seeks to expose untested interleavings for a given test input to increase interleaving coverage. We discuss several solutions to realize the above goal. First, we discuss a coverage metric based on a set of interleaving idioms. Second, we discuss an online technique to predict untested interleavings that can potentially be exposed for a given test input. Finally, the predicted untested interleavings are exposed by actively controlling the thread schedule while executing for the test input. We discuss our experiences in using the tool to expose several known and unknown bugs in real-world applications such as Apache and MySQL.
Jie Yu 0016, Satish Narayanasamy, Cristiano Pereira, Gilles Pokam
OOPSLA3
2011 CoreRacer: a practical memory race recorder for multicore x86 TSO processors
abstract
Shared memory multiprocessors are difficult to program because of the non-deterministic ways in which the memory operations from different threads interleave. To address this issue, many hardware-based memory race recorders have been proposed that efficiently log an ordering of the shared memory interleavings between threads for deterministic replay. These approaches are challenging to integrate into current processors because they change the cache subsystem or the coherence protocol, and they mostly support a sequentially consistent memory model.
Gilles Pokam, Cristiano Pereira, Shiliang Hu, Ali-Reza Adl-Tabatabai, Justin Emile Gottschlich, Youfeng Wu
MICRO2
2010 PinPlay: a framework for deterministic replay and reproducible analysis of parallel programs
abstract
Analysis of parallel programs is hard mainly because their behavior changes from run to run. We present an execution capture and deterministic replay system that enables repeatable analysis of parallel programs. Our goal is to provide an easy-to-use framework for capturing, deterministically replaying, and analyzing execution of large programs with reasonable runtime and disk usage. Our system, called PinPlay, is based on the popular Pin dynamic instrumentation system hence is very easy to use. PinPlay extends the capability of Pin-based analysis by providing a tool for capturing one execution instance of a program (as log files called pinballs) and by allowing Pin-based tools to run off the captured execution. Most Pintools can be trivially modified to work off pinballs thus doing their usual analysis but with a guaranteed repeatability. Furthermore, the capture/replay works across operating systems (Windows to Linux) as the pinball format is independent of the operating system. We have used PinPlay to analyze and deterministically debug large parallel programs running trillions of instructions. This paper describes the design of PinPlay and its applications for analyses such as simulation point selection, tracing, and debugging.
Harish Patil, Cristiano Pereira, Mack Stallcup, Gregory Lueck, James Cownie
CGO2
2009 Offline symbolic analysis for multi-processor execution replay
abstract
Ability to replay a program's execution on a multi-processor system can significantly help parallel programming. To replay a shared-memory multi-threaded program, existing solutions record its program input (I/O, DMA, etc.) and the shared-memory dependencies between threads. Prior processor based record-and-replay solutions are efficient, but they require non-trivial modifications to the coherency protocol and the memory sub-system for recording the shared-memory dependencies.
Mahmoud Said, Satish Narayanasamy, Zijiang Yang 0006, Cristiano Pereira
MICRO5
2009 Architecting a chunk-based memory race recorder in modern CMPs
abstract
Prior work on HW support for memory race recording piggybacks time stamps on coherence messages and logs the outcome of memory races using point-to-point or chunk-based approaches. These memory race recorder (MRR) techniques are effective, but they require modifications to the cache coherence protocol that can hurt performance. In addition, prior work has mostly focused on directory coherence and considered only CMP systems with single-level cache hierarchies. Most modern CMP systems shipped today, however, implement snoop coherence and feature multilevel cache hierarchies. To be practical, a MRR must target CMPs with multilevel caches, mitigate the coherence overhead due to piggybacking, and emphasize on replay speed to broaden applicability of deterministic replay. This paper contributes three new solutions for making chunk-based MRR practical for modern CMPs. We show that MRR interactions with a cache hierarchy can degrade performance and present a novel mechanism that mitigates this degradation. We propose new mechanisms for snoop-based caches that eliminate coherence traffic overhead due to piggybacking. We finally propose new techniques for improving replay speed and introduce a novel framework for evaluating the replay speed potential of MRR designs.
Gilles Pokam, Cristiano Pereira, Klaus Danne, Rolf Kassa, Ali-Reza Adl-Tabatabai
MICRO2
2006 Recording shared memory dependencies using strata
abstract
Significant time is spent by companies trying to reproduce and fix bugs. BugNet and FDR are recent architecture proposals that provide architecture support for deterministic replay debugging. They focus on continuously recording information about the program's execution, which can be communicated back to the developer. Using that information, the developer can deterministically replay the program's execution to reproduce and fix the bugs.In this paper, we propose using Strata to efficiently capture the shared memory dependencies. A stratum creates a time layer across all the logs for the running threads, which separates all the memory operations executed before and after the stratum. A strata log allows us to determine all the shared memory dependencies during replay and thereby supports deterministic replay debugging for multi-threaded programs.
Satish Narayanasamy, Cristiano Pereira, Brad Calder
ASPLOS2
2006 Software Profiling for Deterministic Replay Debugging of User Code
Satish Narayanasamy, Cristiano Pereira, Brad Calder
SoMeT2
2005 Energy-aware wireless systems with adaptive power-fidelity tradeoffs
abstract
Wireless networked embedded systems, such as multimedia terminals, sensor nodes, etc., present a rich domain for making energy/performance/quality tradeoffs based on application needs, network conditions, etc. Energy awareness in these systems is the ability to perform tradeoffs between available battery energy and application quality requirements. In this paper, we show how operating system directed dynamic voltage scaling and dynamic power management can provide for such a capability. We propose a real-time scheduling algorithm that uses runtime feedback about application behavior to provide adaptive power-fidelity tradeoffs. We demonstrate our approach in the context of a static priority-based preemptive task scheduler. Simulation results show that the proposed algorithm results in significant energy savings compared to state-of-the-art dynamic voltage scaling schemes with minimal loss in system fidelity. We have implemented our scheduling algorithm into the eCos real-time operating system running on an Intel XScale-based variable voltage platform. Experimental results obtained using this platform confirm the effectiveness of our technique
Vijay Raghunathan, Cristiano Pereira, Mani Srivastava 0001, Rajesh K. Gupta 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2004 Leakage aware dynamic voltage scaling for real-time embedded systems
abstract
A five-fold increase in leakage current is predicted with each technology generation. While Dynamic Voltage Scaling (DVS) is known to reduce dynamic power consumption, it also causes increased leakage energy drain by lengthening the interval over which a computation is carried out. Therefore, for minimization of the total energy, one needs to determine an operating point, called the critical speed. We compute processor slowdown factors based on the critical speed for energy minimization. Procrastination scheduling attempts to maximize the duration of idle intervals by keeping the processor in a sleep/shutdown state even if there are pending tasks, within the constraints imposed by performance requirements. Our simulation experiments show that the critical speed slowdown results in up to 5% energy gains over a leakage oblivious dynamic voltage scaling. Procrastination scheduling scheme extends the sleep intervals to up to 5 times, resulting in up to an additional 18% energy gains, while meeting all timing requirements.
Ravindra Jejurikar, Cristiano Pereira, Rajesh K. Gupta 0001
DAC2
2004 Energy-Aware System Design for Wireless Multimedia
abstract
In this paper, we present various challenges that arise in the delivery and exchange of multimedia information to mobile devices. Specifically, we focus on techniques for maintaining QoS to end-user multimedia applications (e.g. video streaming, multimedia conferencing) while maximizing device lifetimes. In order to cope with the resource intensive nature of multimedia applications (in terms of computation, bandwidth and consequently power) and dynamic congestion levels in wireless networks, an end-to-end approach to QoS-aware power optimization is required. We discuss the trend towards such an integrated approach that couples the architectural, OS, middleware and application layers to achieve both user experience and device energy gains. We conclude with a discussion of tools for integrated system design and testing that will aid in rapid deployment of wireless multimedia.
Hans Van Antwerpen, Nikil Dutt, Rajesh K. Gupta 0001, Shivajit Mohapatra, Cristiano Pereira, Nalini Venkatasubramanian, Ralph von Vignau
DATE5