Christopher I. Rodrigues

dblp:61/3460 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
GPUs and heterogeneous computing · 33% Parallel and multicore computing · 26% High-performance computing · 19%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
code generation
0.212016
Efficient kernel synthesis for performance portable programming · MICRO 2016
Hardware accelerators and domain-specific architectures
kernel generation
0.212016
Efficient kernel synthesis for performance portable programming · MICRO 2016
High-performance computing › performance engineering
performance portability
0.212016
Efficient kernel synthesis for performance portable programming · MICRO 2016
Parallel and multicore computing › parallel programming models › structured parallelism
algorithmic skeletons
0.212014
Triolet: a programming system that unifies algorithmic skeleton interfaces for high-performance cluster computing · PPoPP 2014
Memory systems › cache management › cache insertion policy
cache bypassing
0.212014
Adaptive Cache Management for Energy-Efficient GPU Computing · MICRO 2014
High-performance computing
cluster computing
0.212014
Triolet: a programming system that unifies algorithmic skeleton interfaces for high-performance cluster computing · PPoPP 2014
Parallel and multicore computing › parallel programming models
distributed memory programming
0.212014
Triolet: a programming system that unifies algorithmic skeleton interfaces for high-performance cluster computing · PPoPP 2014
GPUs and heterogeneous computing › GPU memory management
GPU cache management
0.212014
Adaptive Cache Management for Energy-Efficient GPU Computing · MICRO 2014
Parallel and multicore computing
parallel programming models
0.212014
Triolet: a programming system that unifies algorithmic skeleton interfaces for high-performance cluster computing · PPoPP 2014
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.212014
Adaptive Cache Management for Energy-Efficient GPU Computing · MICRO 2014
GPUs and heterogeneous computing › GPU kernel optimization
CUDA optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
GPUs and heterogeneous computing
GPU performance optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
GPUs and heterogeneous computing
GPU programming
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
Memory systems
memory access optimization
0.112008
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA · PPoPP 2008
GPUs and heterogeneous computing
heterogeneous architecture
0.112016
Efficient kernel synthesis for performance portable programming · MICRO 2016
Compilers and program optimization › loop transformation
loop fusion
0.112014
Triolet: a programming system that unifies algorithmic skeleton interfaces for high-performance cluster computing · PPoPP 2014
GPUs and heterogeneous computing › GPU architecture
energy-efficient GPU design
0.112014
Adaptive Cache Management for Energy-Efficient GPU Computing · MICRO 2014

Methods — techniques the papers use, named apart from their topics

design space exploration · 0.5composition algorithm · 0.5task decomposition · 0.4scheduling · 0.4communication optimization · 0.4runtime contention detection · 0.2reuse distance analysis · 0.2predictor for active warp count · 0.2memory coalescing · 0.1latency hiding · 0.1
YearPublicationVenuePosition
2016 Efficient kernel synthesis for performance portable programming
abstract
The diversity of microarchitecture designs in heterogeneous computing systems allows programs to achieve high performance and energy efficiency, but results in substantial software re-development cost for each type or generation of hardware. To mitigate this cost, a performance portable programming system is required. One fundamental difference between architectures that makes performance portability challenging is the hierarchical organization of their computing elements. To address this challenge, we introduce TANGRAM, a kernel synthesis framework that composes architecture-neutral computations and composition rules into high-performance kernels customized for different architectural hierarchies. TANGRAM is based on an extensible architectural model that can be used to specify a variety of architectures. This model is coupled with a generic design space exploration and composition algorithm that can generate multiple composition plans for any specified architecture. A custom code generator then compiles these plans for the target architecture while performing various optimizations such as data placement and tuning. We show that code synthesized by TANGRAM for different types and generations of devices achieves no less than 70% of the performance of highly optimized vendor libraries such as Intel MKL and NVIDIA CUBLAS/CUSPARSE.
Li-Wen Chang, Izzat El Hajj, Christopher I. Rodrigues, Juan Gómez-Luna, Wen-Mei W. Hwu
MICRO3
2014 Adaptive Cache Management for Energy-Efficient GPU Computing
abstract
With the SIMT execution model, GPUs can hide memory latency through massive multithreading for many applications that have regular memory access patterns. To support applications with irregular memory access patterns, cache hierarchies have been introduced to GPU architectures to capture temporal and spatial locality and mitigate the effect of irregular accesses. However, GPU caches exhibit poor efficiency due to the mismatch of the throughput-oriented execution model and its cache hierarchy design, which limits system performance and energy-efficiency. The massive amount of memory requests generated by GPU scause cache contention and resource congestion. Existing CPUcache management policies that are designed for multicoresystems, can be suboptimal when directly applied to GPUcaches. We propose a specialized cache management policy for GPGPUs. The cache hierarchy is protected from contention by the bypass policy based on reuse distance. Contention and resource congestion are detected at runtime. To avoid oversaturatingon-chip resources, the bypass policy is coordinated with warp throttling to dynamically control the active number of warps. We also propose a simple predictor to dynamically estimate the optimal number of active warps that can take full advantage of the cache space and on-chip resources. Experimental results show that cache efficiency is significantly improved and on-chip resources are better utilized for cache sensitive benchmarks. This results in a harmonic mean IPC improvement of 74% and 17% (maximum 661% and 44% IPCimprovement), compared to the baseline GPU architecture and optimal static warp throttling, respectively.
Xuhao Chen 0001, Li-Wen Chang, Christopher I. Rodrigues, Zhiying Wang 0003, Wen-Mei W. Hwu
MICRO3
2014 Triolet: a programming system that unifies algorithmic skeleton interfaces for high-performance cluster computing
abstract
Functional algorithmic skeletons promise a high-level programming interface for distributed-memory clusters that free developers from concerns of task decomposition, scheduling, and communication. Unfortunately, prior distributed functional skeleton frameworks do not deliver performance comparable to that achievable in a low-level distributed programming model such as C with MPI and OpenMP, even when used in concert with high-performance array libraries. There are several causes: they do not take advantage of shared memory on each cluster node; they impose a fixed partitioning strategy on input data; and they have limited ability to fuse loops involving skeletons that produce a variable number of outputs per input.
Christopher I. Rodrigues, Thomas B. Jablin, Abdul Dakkak, Wen-Mei W. Hwu
PPoPP1
2013 Scalable SIMD-parallel memory allocation for many-core machines
Xiaohuang Huang, Christopher I. Rodrigues, Ian Buck, Wen-Mei W. Hwu
J. Supercomput.2
2008 Program optimization space pruning for a multithreaded gpu
abstract
Program optimization for highly-parallel systems has historically been considered an art, with experts doing much of the performance tuning by hand. With the introduction of inexpensive, single-chip, massively parallel platforms, more developers will be creating highly-parallel applications for these platforms, who lack the substantial experience and knowledge needed to maximize their performance. This creates a need for more structured optimization methods with means to estimate their performance effects. Furthermore these methods need to be understandable by most programmers. This paper shows the complexity involved in optimizing applications for one such system and one relatively simple methodology for reducing the workload involved in the optimization process.
Shane Ryoo, Christopher I. Rodrigues, Sam S. Stone, Sara S. Baghsorkhi, Sain-Zee Ueng, John A. Stratton, Wen-Mei W. Hwu
CGO2
2008 Optimization principles and application performance evaluation of a multithreaded GPU using CUDA
abstract
GPUs have recently attracted the attention of many application developers as commodity data-parallel coprocessors. The newest generations of GPU architecture provide easier programmability and increased generality while maintaining the tremendous memory bandwidth and computational power of traditional GPUs. This opportunity should redirect efforts in GPGPU research from ad hoc porting of applications to establishing principles and strategies that allow efficient mapping of computation to graphics hardware. In this work we discuss the GeForce 8800 GTX processor's organization, features, and generalized optimization strategies. Key to performance on this platform is using massive multithreading to utilize the large number of cores and hide global memory latency. To achieve this, developers face the challenge of striking the right balance between each thread's resource usage and the number of simultaneously active threads. The resources to manage include the number of registers and the amount of on-chip memory used per thread, number of threads per multiprocessor, and global memory bandwidth. We also obtain increased performance by reordering accesses to off-chip memory to combine requests to the same or contiguous memory locations and apply classical optimizations to reduce the number of executed operations. We apply these strategies across a variety of applications and domains and achieve between a 10.5X to 457X speedup in kernel codes and between 1.16X to 431X total application speedup.
Shane Ryoo, Christopher I. Rodrigues, Sara S. Baghsorkhi, Sam S. Stone, David Blair Kirk, Wen-Mei W. Hwu
PPoPP2
2008 Program optimization carving for GPU computing
Shane Ryoo, Christopher I. Rodrigues, Sam S. Stone, John A. Stratton, Sain-Zee Ueng, Sara S. Baghsorkhi, Wen-Mei W. Hwu
J. Parallel Distributed Comput.2