Gilberto Contreras

dblp:57/6257 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 46% Energy-efficient computing · 16% Performance modeling and evaluation · 16%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
load balancing
0.112011
Parallelization libraries: Characterizing and reducing overheads · ACM Trans. Archit. Code Optim. 2011
Parallel and multicore computing
parallel libraries
0.112011
Parallelization libraries: Characterizing and reducing overheads · ACM Trans. Archit. Code Optim. 2011
Parallel and multicore computing
parallel programming runtimes
0.112011
Parallelization libraries: Characterizing and reducing overheads · ACM Trans. Archit. Code Optim. 2011
Electronic design automation › high-level synthesis
scheduling
0.112011
Parallelization libraries: Characterizing and reducing overheads · ACM Trans. Archit. Code Optim. 2011
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.112006
Live, Runtime Phase Monitoring and Prediction on Real Systems with Application to Dynamic Power Management · MICRO 2006
Performance modeling and evaluation › workload characterization › program behavior analysis
phase analysis
0.112006
Live, Runtime Phase Monitoring and Prediction on Real Systems with Application to Dynamic Power Management · MICRO 2006
Energy-efficient computing
power management
0.112006
Live, Runtime Phase Monitoring and Prediction on Real Systems with Application to Dynamic Power Management · MICRO 2006
Performance modeling and evaluation
workload characterization
0.112006
Live, Runtime Phase Monitoring and Prediction on Real Systems with Application to Dynamic Power Management · MICRO 2006
Processor architecture and microarchitecture
chip multiprocessor
0.012011
Parallelization libraries: Characterizing and reducing overheads · ACM Trans. Archit. Code Optim. 2011
Processor architecture and microarchitecture
branch prediction
0.012006
Live, Runtime Phase Monitoring and Prediction on Real Systems with Application to Dynamic Power Management · MICRO 2006

Methods — techniques the papers use, named apart from their topics

occupancy-based task stealing · 0.1criticality-guided task stealing · 0.1performance counters · 0.1branch predictor design · 0.1
YearPublicationVenuePosition
2011 Parallelization libraries: Characterizing and reducing overheads
abstract
Creating efficient, scalable dynamic parallel runtime systems for chip multiprocessors (CMPs) requires understanding the overheads that manifest at high core counts and small task sizes. In this article, we assess these overheads on Intel's Threading Building Blocks (TBB) and OpenMP. First, we use real hardware and simulations to detail various scheduler and synchronization overheads. We find that these can amount to 47% of TBB benchmark runtime and 80% of OpenMP benchmark runtime. Second, we propose load balancing techniques such as occupancy-based and criticality-guided task stealing, to boost performance. Overall, our study provides valuable insights for creating robust, scalable runtime libraries.
Abhishek Bhattacharjee, Gilberto Contreras, Margaret Martonosi
ACM Trans. Archit. Code Optim.2
2008 Full-system chip multiprocessor power evaluations using FPGA-based emulation
abstract
The design process for chip multiprocessors (CMPs) requires extremely long simulation times to explore performance, power, and thermal issues, particularly when operating system (OS) effects are included. In response, our novel FPGA-based emulation methodology models a full CMP design including applications and an OS. Activity counters programmed into the cores feed per-component microarchitectural power models. These models achieve under 10 % error compared to detailed gate-level simulations. Our method retains software flexibility, but offers up to 35 × speedup compared to full-system software simulations. We present our approach by emulating a 2-core Leon3 cache-coherent multiprocessor running Linux and parallel benchmarks. In an example case study, our emulated system uses activity counts (a proxy for temperature) to guide process migration between the CMP cores. Overall, this paper’s methodology makes possible detailed power and thermal studies of CMPs and their operating systems.
Abhishek Bhattacharjee, Gilberto Contreras, Margaret Martonosi
ISLPED2
2007 The XTREM power and performance simulator for the Intel XScale core: Design and experiences
abstract
Managing power concerns in microprocessors has become a pressing research problem across the domains of computer architecture, CAD, and compilers. As a result, several parameterized cycle-level power simulators have been introduced. While these simulators can be quite useful for microarchitectural studies, their generality limits how accurate they can be for any one chip family. Furthermore, their hardware focus means that they do not explicitly enable studying the interaction of different software layers, such as Java applications and their underlying runtime system software. This paper describes and evaluates XTREM, a power-simulation tool tailored for the Intel XScale microarchitecture. In building XTREM, our goals were to develop a microarchitecture simulator that, while still offering size parameterizations for cache and other structures, more accurately reflected a realistic processor pipeline. We present a detailed set of validations based on multimeter power measurements and hardware performance counter sampling. XTREM exhibits an average performance error of only 6.5% and an even smaller average power error: 4%. The paper goes on to present an application study enabled by the simulator. Namely, we use XTREM to produce an energy consumption breakdown for Java CDC and CLDC applications. Our simulator measurements indicate that a large percentage of the total energy consumption (up to 35%) is devoted to the virtual machine's support functions.
Gilberto Contreras, Margaret Martonosi, Jinzhan Peng, Guei-Yuan Lueh, Roy Dz-Ching Ju
ACM Trans. Embed. Comput. Syst.1
2006 Live, Runtime Phase Monitoring and Prediction on Real Systems with Application to Dynamic Power Management
abstract
Computer architecture has experienced a major paradigm shift from focusing only on raw performance to considering power-performance efficiency as the defining factor of the emerging systems. Along with this shift has come increased interest in workload characterization. This interest fuels two closely related areas of research. First, various studies explore the properties of workload variations and develop methods to identify and track different execution behavior, commonly referred to as "phase analysis". Second, a large complementary set of research studies dynamic, on-the-fly system management techniques that can adaptively respond to these differences in application behavior. Both of these lines of work have produced very interesting and widely useful results. Thus far, however, there exists only a weak link between these conceptually related areas, especially for real-system studies. Our work aims to strengthen this link by demonstrating a real-system implementation of a runtime phase predictor that works cooperatively with on-the-fly dynamic management. We describe a fully-functional deployed system that performs accurate phase predictions on running applications. The key insight of our approach is to draw from prior branch predictor designs to create a phase history table that guides predictions. To demonstrate the value of our approach, we implement a prototype system that uses it to guide dynamic voltage and frequency scaling. Our runtime phase prediction methodology achieves above 90% prediction accuracies for many of the experimented benchmarks. For highly variable applications, our approach can reduce mispredictions by more than 6X over commonly-used statistical approaches. Dynamic frequency and voltage scaling, when guided by our runtime phase predictor, achieves energy-delay product improvements as high as 34% for benchmarks with non-negligible variability, on average 7% better than previous methods and 18% better than a baseline unmanaged system
Canturk Isci, Gilberto Contreras, Margaret Martonosi
MICRO2
2005 Power prediction for intel XScale processors using performance monitoring unit events
abstract
This paper demonstrates a first-order, linear power estimation model ha uses performance counters to estimate run-time CPU and memory power consumption of the Intel PXA255 processor. Our model uses a set of power weights that map hardware performance counter values to processor and memory power consumption. Power weights are derived offline once per processor voltage and frequency configuration using parameter estimation echniques. They can be applied in a dynamic voltage/frequency scaling environment by setting six descriptive parameters. We have tested our model using a wide selection of benchmarks including SPEC2000, Java CDC and Java CLDC programming environments. The accuracy is quite good; average estimated power consumption is within 4% of he measured average CPU power consumption. We believe such power estimation schemes can serve as a foundation for intelligent, power-aware embedded systems tha dynamically adapt to the device's power consumption
Gilberto Contreras, Margaret Martonosi
ISLPED1
2004 XTREM: a power simulator for the Intel XScale® core
abstract
Managing power concerns in icroprocessors has become a pressing research problem across the domains of computer architecture, CAD, and compilers. As a result, several parameterized cycle-level power simulators have been introduced. While these simulators can be quite useful for microarchitectural studies, their generality limits how accurate they can be for any one chip family. Furthermore, their hardware focus means that they do not explicitly enable studying the interaction of different software layers, such as Java applications and their underlying Runtime system software.This paper describes and evaluates XTREM, a power simulation tool tailored for the Intel XScale icroarchitecture. In building XTREM, our goals were to develop a icroarchitecture simulator that, while still offering size parameterizations for cache, TLB, etc., more accurately reflected a realistic processor pipeline. We present a detailed set of validations based on ultimeter power measurements and hardware performance counter sampling. Based on these validations across a wide range of stressmarks, Java benchmarks, and non-Java benchmarks, XTREM has an average performance error of only 6.5% and an even smaller average power error: 4%. The paper goes on to present a selection of application studies enabled by the simulator. For example, presenting power behavior vs. time for selected embedded C and Java CLDC benchmarks, we can make power distinctions between the two programming domains as well as distinguishing Java application (JITted code) power from Java Runtime system power. We also study how the Intel XScale core 's power consumption varies for different data activity factors, creating power swings as large as 50mW for a 200Mhz core. We are planning to release XTREM for wider use, and feel that it offers a useful step forward for compiler and embedded software designers.
Gilberto Contreras, Margaret Martonosi, Jinzhan Peng, Roy Dz-Ching Ju, Guei-Yuan Lueh
LCTES1