Timothy G. Armstrong

dblp:76/7185 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 4 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 54% Performance modeling and evaluation · 17% GPUs and heterogeneous computing · 13%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 59% Graph data management · 31% Web and social media mining · 9%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 79% Operating systems · 21%

Topics — the 16 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
0.222009
EvaluatIR: an online tool for evaluating and comparing IR systems · SIGIR 2009
Has adhoc retrieval improved since 1994? · SIGIR 2009
Compilers and program optimization
parallelizing compiler
0.212014
Compiler Techniques for Massively Scalable Implicit Task Parallelism · SC 2014
GPUs and heterogeneous computing
GPU computing
0.212014
Design and evaluation of the gemtc framework for GPU-enabled many-task computing · HPDC 2014
Parallel and multicore computing
parallel programming runtimes
0.212014
Design and evaluation of the gemtc framework for GPU-enabled many-task computing · HPDC 2014
Parallel and multicore computing › parallel programming models
task parallelism
0.212014
Compiler Techniques for Massively Scalable Implicit Task Parallelism · SC 2014
Parallel and multicore computing
task scheduling
0.212014
Design and evaluation of the gemtc framework for GPU-enabled many-task computing · HPDC 2014
Performance modeling and evaluation
benchmarking
0.212013
LinkBench: a database benchmark based on the Facebook social graph · SIGMOD Conference 2013
Performance modeling and evaluation › benchmarking
database system benchmarking
0.212013
LinkBench: a database benchmark based on the Facebook social graph · SIGMOD Conference 2013
Parallel and multicore computing › parallel programming models
dataflow programming
0.212013
Swift/T: scalable data flow programming for many-task applications · PPoPP 2013
High-performance computing › high-throughput computing
many-task computing
0.212013
Swift/T: scalable data flow programming for many-task applications · PPoPP 2013
Parallel and multicore computing
parallel programming models
0.212013
Swift/T: scalable data flow programming for many-task applications · PPoPP 2013
Information retrieval › retrieval models
ad-hoc retrieval
0.112009
Has adhoc retrieval improved since 1994? · SIGIR 2009
GPUs and heterogeneous computing › heterogeneous supercomputing
accelerator-based supercomputing
0.112014
Design and evaluation of the gemtc framework for GPU-enabled many-task computing · HPDC 2014
Web and social media mining
social network analysis
0.012013
LinkBench: a database benchmark based on the Facebook social graph · SIGMOD Conference 2013
Memory systems
processing-in-memory
0.012013
Parallelizing the execution of sequential scripts · SC 2013
Information retrieval › evaluation
test collection
0.012009
Has adhoc retrieval improved since 1994? · SIGIR 2009

Methods — techniques the papers use, named apart from their topics

intermediate representation · 0.4compiler transformation · 0.4functional file management · 0.3collective file movement · 0.3swift parallel dataflow language · 0.2data flow language implementation · 0.2online tool · 0.1benchmark comparison · 0.1
YearPublicationVenuePosition
2015 Toward Interlanguage Parallel Scripting for Distributed-Memory Scientific Computing
abstract
Scripting languages such as Python and R have been widely adopted as tools for the productive development of scientific software because of the power and expressiveness of the languages and available libraries. However, deploying scripted applications on large-scale parallel computer systems such as the IBM Blue Gene/Q or Cray XE6 is a challenge because of issues including operating system limitations, interoperability challenges, parallel filesystem overheads due to the small file system accesses common in scripted approaches, and other issues. We present here a new approach to these problems in which the Swift scripting system is used to integrate high-level scripts written in Python, R, and Tcl, with native code developed in C, C++, and Fortran, by linking Swift to the library interfaces to the script interpreters. In this approach, Swift handles data management, movement, and marshaling among distributed-memory processes without direct user manipulation of low-level communication libraries such as MPI. We present a technique to efficiently launch scripted applications on large-scale supercomputers using a hierarchical programming model.
Justin M. Wozniak, Timothy G. Armstrong, Ketan Maheshwari, Daniel S. Katz, Michael Wilde, Ian T. Foster
CLUSTER2
2015 Porting Ordinary Applications to Blue Gene/Q Supercomputers
abstract
Efficiently porting ordinary applications to Blue Gene/Q supercomputers is a significant challenge. Codes are often originally developed without considering advanced architectures and related tool chains. Science needs frequently lead users to want to run large numbers of relatively small jobs (often called many-task computing, an ensemble, or a workflow), which can conflict with supercomputer configurations. In this paper, we discuss techniques developed to execute ordinary applications over leadership class supercomputers. We use the high-performance Swift parallel scripting framework and build two workflow execution techniques -- sub-jobs and main-wrap. The sub-jobs technique, built on top of the IBM Blue Gene/Q resource manager Cobalt's sub-block jobs, lets users submit multiple, independent, repeated smaller jobs within a single larger resource block. The main-wrap technique is a scheme that enables C/C++ programs to be defined as functions that are wrapped by a high-performance Swift wrapper and that are invoked as a Swift script. We discuss the needs, benefits, technicalities, and current limitations of these techniques. We further discuss the real-world science enabled by these techniques and the results obtained.
Ketan Maheshwari, Justin M. Wozniak, Timothy G. Armstrong, Daniel S. Katz, T. Andrew Binkowski, Xiaoliang Zhong, Olle Heinonen, Dmitry Karpeyev, Michael Wilde
e-Science3
2014 Compiler Optimization for Extreme-Scale Scripting
abstract
The data-driven task parallelism execution model can support parallel programming models that are well suited for large-scale distributed-memory parallel computing, for example, simulations and analysis pipelines running on clusters and clouds. We describe a novel compiler intermediate representation and optimizations for this execution model, including adaptions of standard techniques alongside novel techniques. These techniques are applied to Swift/T, a high-level scripting language for flexible data flow composition of functions, which may be serial or use lower-level parallel programming models such as MPI and OpenMP. This paper presents preliminary results, indicating that our compiler optimizations reduce communication overhead by 70% to 93% on distributed-memory systems.
Timothy G. Armstrong, Justin M. Wozniak, Michael Wilde, Ian T. Foster
CCGRID1
2014 Design and evaluation of the gemtc framework for GPU-enabled many-task computing
abstract
We present the design and first performance and usability evaluation of GeMTC, a novel execution model and runtime system that enables accelerators to be programmed with many concurrent and independent tasks of potentially short or variable duration. With GeMTC, a broad class of such "many-task" applications can leverage the increasing number of accelerated and hybrid high-end computing systems. GeMTC overcomes the obstacles to using GPUs in a many-task manner by scheduling and launching independent tasks on hardware designed for SIMD-style vector processing. We demonstrate the use of a high-level MTC programming model (the Swift parallel dataflow language) to run tasks on many accelerators and thus provide a high-productivity programming model for the growing number of supercomputers that are accelerator-enabled. While still in an experimental stage, GeMTC can already support tasks of fine (subsecond) granularity and execute concurrent heterogeneous tasks on 86,000 independent GPU warps spanning 2.7M GPU threads on the Blue Waters supercomputer.
Scott J. Krieder, Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Benjamin Grimmer, Ian T. Foster, Ioan Raicu
HPDC3
2014 Compiler Techniques for Massively Scalable Implicit Task Parallelism
abstract
Swift/T is a high-level language for writing concise, deterministic scripts that compose serial or parallel codes implemented in lower-level programming models into large-scale parallel applications. It executes using a data-driven task parallel execution model that is capable of orchestrating millions of concurrently executing asynchronous tasks on homogeneous or heterogeneous resources. Producing code that executes efficiently at this scale requires sophisticated compiler transformations: poorly optimized code inhibits scaling with excessive synchronization and communication. We present a comprehensive set of compiler techniques for data-driven task parallelism, including novel compiler optimizations and intermediate representations. We report application benchmark studies, including unbalanced tree search and simulated annealing, and demonstrate that our techniques greatly reduce communication overhead and enable extreme scalability, distributing up to 612 million dynamically load balanced tasks per second at scales of up to 262,144 cores without explicit parallelism, synchronization, or load balancing in application code.
Timothy G. Armstrong, Justin M. Wozniak, Michael Wilde, Ian T. Foster
SC1
2013 Swift/T: Large-Scale Application Composition via Distributed-Memory Dataflow Processing
abstract
Many scientific applications are conceptually built up from independent component tasks as a parameter study, optimization, or other search. Large batches of these tasks may be executed on high-end computing systems, however, the coordination of the independent processes, their data, and their data dependencies is a significant scalability challenge. Many problems must be addressed, including load balancing, data distribution, notifications, concurrent programming, and linking to existing codes. In this work, we present Swift/T, a programming language and runtime that enables the rapid development of highly concurrent, task-parallel applications. Swift/Tis composed of several enabling technologies to address scalability challenges, offers a high-level optimizing compiler for user programming and debugging, and provides tools for binding user code in C/C++/Fortran into a logical script. In this work, we describe the Swift/T solution and present scaling results from the IBM Blue Gene/Pand Blue Gene/Q.
Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Ewing L. Lusk, Ian T. Foster
CCGRID2
2013 Swift/T: scalable data flow programming for many-task applications
abstract
Swift/T, a novel programming language implementation for highly scalable data flow programs, is presented.
Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Ewing L. Lusk, Ian T. Foster
PPoPP2
2013 Dataflow coordination of data-parallel tasks via MPI 3.0
abstract
Scientific applications are often complex collections of many large-scale tasks. Mature tools exist for describing task-parallel workflows consisting of serial tasks, and a variety of tools exist for programming a single data-parallel operation. However, few tools cover the intersection of these two models. In this work, we extend the load balancing library ADLB to support parallel tasks. We demonstrate how applications can easily be composed of parallel tasks using Swift dataflow scripts, which are compiled to ADLB programs with performance comparable to hand-coded equivalents. By combining this framework with data-parallel analysis libraries, we are able to dynamically execute many instances of a parallel data analysis application in support of a parameter exploration workload.
Justin M. Wozniak, Tom Peterka, Timothy G. Armstrong, James Dinan, Ewing L. Lusk, Michael Wilde, Ian T. Foster
EuroMPI3
2013 Parallelizing the execution of sequential scripts
abstract
Scripting is often used in science to create applications via the composition of existing programs. Parallel scripting systems allow the creation of such applications, but each system introduces the need to adopt a somewhat specialized programming model. We present an alternative scripting approach, AMFS Shell, that lets programmers express parallel scripting applications via minor extensions to existing sequential scripting languages, such as Bash, and then execute them in-memory on large-scale computers. We define a small set of commands between the scripts and a parallel scripting runtime system, so that programmers can compose their scripts in a familiar scripting language. The underlying AMFS implements both collective (fast file movement) and functional (transformation based on content) file management. Tasks are handled by AMFS's built-in execution engine. AMFS Shell is expressive enough for a wide range of applications, and the framework can run such applications efficiently on large-scale computers.
Zhao Zhang 0007, Daniel S. Katz, Timothy G. Armstrong, Justin M. Wozniak, Ian T. Foster
SC3
2013 LinkBench: a database benchmark based on the Facebook social graph
abstract
Database benchmarks are an important tool for database researchers and practitioners that ease the process of making informed comparisons between different database hardware, software and configurations. Large scale web services such as social networks are a major and growing database application area, but currently there are few benchmarks that accurately model web service workloads.
Timothy G. Armstrong, Vamsi Ponnekanti, Dhruba Borthakur, Mark Callaghan
SIGMOD Conference1
2013 Turbine: A Distributed-memory Dataflow Engine for High Performance Many-task Applications
abstract
Efficiently utilizing the rapidly increasing concurrency of multi-petaflop computing systems is a significant programming challenge. One approach is to structure applications with an upper layer of many loosely coupled coarse-grained tasks, each comp
Justin M. Wozniak, Timothy G. Armstrong, Ketan Maheshwari, Ewing L. Lusk, Daniel S. Katz, Michael Wilde, Ian T. Foster
Fundam. Informaticae2
2009 Improvements that don't add up: ad-hoc retrieval results since 1998
abstract
The existence and use of standard test collections in information retrieval experimentation allows results to be compared between research groups and over time. Such comparisons, however, are rarely made. Most researchers only report results from their own experiments, a practice that allows lack of overall improvement to go unnoticed. In this paper, we analyze results achieved on the TREC Ad-Hoc, Web, Terabyte, and Robust collections as reported in SIGIR (1998--2008) and CIKM (2004--2008). Dozens of individual published experiments report effectiveness improvements, and often claim statistical significance. However, there is little evidence of improvement in ad-hoc retrieval technology over the past decade. Baselines are generally weak, often being below the median original TREC system. And in only a handful of experiments is the score of the best TREC automatic run exceeded. Given this finding, we question the value of achieving even a statistically significant result over a weak baseline. We propose that the community adopt a practice of regular longitudinal comparison to ensure measurable progress, or at least prevent the lack of it from going unnoticed. We describe an online database of retrieval runs that facilitates such a practice.
Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel
CIKM1
2009 Has adhoc retrieval improved since 1994?
abstract
Evaluation forums such as TREC allow systematic measurement and comparison of information retrieval techniques. The goal is consistent improvement, based on reliable comparison of the effectiveness of different approaches and systems. In this paper we report experiments to determine whether this goal has been achieved. We ran five publicly available search systems, in a total of seventeen different configurations, against nine TREC adhoc-style collections, spanning 1994 to 2005. These runsets were then used as a benchmark for reassessing the relative effectiveness of the original TREC runs for those collections. Surprisingly, there appears to have been no overall improvement in effectiveness for either median or top-end TREC submissions, even after allowing for several possible confounds. We therefore question whether the effectiveness of adhoc information retrieval has improved over the past decade and a half.
Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel
SIGIR1
2009 EvaluatIR: an online tool for evaluating and comparing IR systems
abstract
No abstract available.
Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel
SIGIR1