Timothy Creech

dblp:138/4157 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 90% Embedded and real-time systems · 10%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
dependence analysis
0.212015
Affine Parallelization Using Dependence and Cache Analysis in a Binary Rewriter · IEEE Trans. Parallel Distributed Syst. 2015
Parallel and multicore computing › parallel programming models
automatic parallelization
0.212015
Affine Parallelization Using Dependence and Cache Analysis in a Binary Rewriter · IEEE Trans. Parallel Distributed Syst. 2015
Parallel and multicore computing › parallel scheduling
malleable task scheduling
0.212013
Efficient multiprogramming for multicores with SCAF · MICRO 2013
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.212013
Efficient multiprogramming for multicores with SCAF · MICRO 2013
Embedded and real-time systems › worst-case execution time analysis
cache analysis
0.112015
Affine Parallelization Using Dependence and Cache Analysis in a Binary Rewriter · IEEE Trans. Parallel Distributed Syst. 2015
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
0.012013
Efficient multiprogramming for multicores with SCAF · MICRO 2013

Methods — techniques the papers use, named apart from their topics

dependence analysis · 0.4cache analysis · 0.4feedback-based allocation · 0.2
YearPublicationVenuePosition
2015 Affine Parallelization Using Dependence and Cache Analysis in a Binary Rewriter
abstract
Today, nearly all general-purpose computers are parallel, but nearly all software running on them is serial. Bridging this disconnect by manually rewriting source code in parallel is prohibitively expensive. Hence, automatic parallelization is an attractive alternative. We present a method to perform automatic parallelization in a binary rewriter. The input to the binary rewriter is the serial binary executable program and the output is a parallel binary executable. The advantages of parallelization in a binary rewriter versus a compiler include: (i) applicability to legacy binaries whose source is not available; (ii) applicability to library code that is often distributed only as binary; (iii) usability by end-user of the software; and (iv) applicability to assembly-language programs. Adapting existing parallelizing compiler methods that work on source code to binaries is a significant challenge. This is primarily because symbolic and array index information used by existing parallelizing compilers to take parallelizing decisions is not available from a binary. We show how to adapt existing parallelization methods to binaries without using symbolic and array index information to achieve equivalent source parallelization from binaries. Results using our x86 binary rewriter called SecondWrite are presented in detail for the dense-matrix and regular Polybench benchmark suite. For these, the average speedup from our method for binaries is 27.95X versus 28.22X from source on 24 threads, compared to the input serial binary. Super-linear speedups are possible due to cache optimizations. In addition our results show that our method is robust and has correctly parallelized much larger binaries from the SPEC 2006 and OMP 2001 benchmark suites, totaling over 2 million source lines of code (SLOC). Good speedups result on the subset of those benchmarks that have affine parallelism in them; this subset exceeds 100,000 SLOC. This robustness is unusual even for the latest leading source parallelizers, many of which are famously fragile.
Aparna Kotha, Kapil Anand, Timothy Creech, Khaled Elwazeer, Matthew Smithson, Greeshma Yellareddy, Rajeev Barua
IEEE Trans. Parallel Distributed Syst.3
2014 Affine Parallelization of Loops with Run-Time Dependent Bounds from Binaries
Aparna Kotha, Kapil Anand, Timothy Creech, Khaled Elwazeer, Matthew Smithson, Rajeev Barua
ESOP3
2013 Efficient multiprogramming for multicores with SCAF
abstract
As hardware becomes increasingly parallel and the availability of scalable parallel software improves, the problem of managing multiple multithreaded applications (processes) becomes important. Malleable processes, which can vary the number of threads used as they run, enable sophisticated and flexible resource management. Although many existing applications parallelized for SMPs with parallel runtimes are in fact already malleable, deployed run-time environments provide no interface nor any strategy for intelligently allocating hardware threads or even preventing oversubscription. Work up until SCAF either depends upon profiling applications ahead of time in order to make good decisions about allocations, or does not account for process efficiency at all. This paper presents the Scheduling and Allocation with Feedback (SCAF) system, a drop-in runtime solution which supports existing malleable applications in making intelligent allocation decisions based on observed efficiency without any paradigm change, changes to semantics, program modification, offline profiling, or even recompilation. Our existing implementation can control most unmodified OpenMP applications. Other malleable threading libraries can also easily be supported with small modifications, without requiring application modification.
Timothy Creech, Aparna Kotha, Rajeev Barua
MICRO1