EDBT 2026 Demo / reviewers in the wild / expert
Timothy G. Armstrong
dblp:76/7185
· DBLP profile ↗
14ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 4 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Parallel and multicore computing · 54% Performance modeling and evaluation · 17% GPUs and heterogeneous computing · 13% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 59% Graph data management · 31% Web and social media mining · 9% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 79% Operating systems · 21% |
Topics — the 16 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
evaluation |
0.2 | 2 | 2009 | EvaluatIR: an online tool for evaluating and comparing IR systems · SIGIR 2009 Has adhoc retrieval improved since 1994? · SIGIR 2009 |
Compilers and program optimization
parallelizing compiler |
0.2 | 1 | 2014 | Compiler Techniques for Massively Scalable Implicit Task Parallelism · SC 2014 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2014 | Design and evaluation of the gemtc framework for GPU-enabled many-task computing · HPDC 2014 |
Parallel and multicore computing
parallel programming runtimes |
0.2 | 1 | 2014 | Design and evaluation of the gemtc framework for GPU-enabled many-task computing · HPDC 2014 |
Parallel and multicore computing › parallel programming models
task parallelism |
0.2 | 1 | 2014 | Compiler Techniques for Massively Scalable Implicit Task Parallelism · SC 2014 |
Parallel and multicore computing
task scheduling |
0.2 | 1 | 2014 | Design and evaluation of the gemtc framework for GPU-enabled many-task computing · HPDC 2014 |
Performance modeling and evaluation
benchmarking |
0.2 | 1 | 2013 | LinkBench: a database benchmark based on the Facebook social graph · SIGMOD Conference 2013 |
Performance modeling and evaluation › benchmarking
database system benchmarking |
0.2 | 1 | 2013 | LinkBench: a database benchmark based on the Facebook social graph · SIGMOD Conference 2013 |
Parallel and multicore computing › parallel programming models
dataflow programming |
0.2 | 1 | 2013 | Swift/T: scalable data flow programming for many-task applications · PPoPP 2013 |
High-performance computing › high-throughput computing
many-task computing |
0.2 | 1 | 2013 | Swift/T: scalable data flow programming for many-task applications · PPoPP 2013 |
Parallel and multicore computing
parallel programming models |
0.2 | 1 | 2013 | Swift/T: scalable data flow programming for many-task applications · PPoPP 2013 |
Information retrieval › retrieval models
ad-hoc retrieval |
0.1 | 1 | 2009 | Has adhoc retrieval improved since 1994? · SIGIR 2009 |
GPUs and heterogeneous computing › heterogeneous supercomputing
accelerator-based supercomputing |
0.1 | 1 | 2014 | Design and evaluation of the gemtc framework for GPU-enabled many-task computing · HPDC 2014 |
Web and social media mining
social network analysis |
0.0 | 1 | 2013 | LinkBench: a database benchmark based on the Facebook social graph · SIGMOD Conference 2013 |
Memory systems
processing-in-memory |
0.0 | 1 | 2013 | Parallelizing the execution of sequential scripts · SC 2013 |
Information retrieval › evaluation
test collection |
0.0 | 1 | 2009 | Has adhoc retrieval improved since 1994? · SIGIR 2009 |
Methods — techniques the papers use, named apart from their topics
intermediate representation · 0.4compiler transformation · 0.4functional file management · 0.3collective file movement · 0.3swift parallel dataflow language · 0.2data flow language implementation · 0.2online tool · 0.1benchmark comparison · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Toward Interlanguage Parallel Scripting for Distributed-Memory Scientific ComputingabstractScripting languages such as Python and R have been widely adopted as tools for the productive development of scientific software because of the power and expressiveness of the languages and available libraries. However, deploying scripted applications on large-scale parallel computer systems such as the IBM Blue Gene/Q or Cray XE6 is a challenge because of issues including operating system limitations, interoperability challenges, parallel filesystem overheads due to the small file system accesses common in scripted approaches, and other issues. We present here a new approach to these problems in which the Swift scripting system is used to integrate high-level scripts written in Python, R, and Tcl, with native code developed in C, C++, and Fortran, by linking Swift to the library interfaces to the script interpreters. In this approach, Swift handles data management, movement, and marshaling among distributed-memory processes without direct user manipulation of low-level communication libraries such as MPI. We present a technique to efficiently launch scripted applications on large-scale supercomputers using a hierarchical programming model. Justin M. Wozniak, Timothy G. Armstrong, Ketan Maheshwari, Daniel S. Katz, Michael Wilde, Ian T. Foster |
CLUSTER | 2 |
| 2015 | Porting Ordinary Applications to Blue Gene/Q SupercomputersabstractEfficiently porting ordinary applications to Blue Gene/Q supercomputers is a significant challenge. Codes are often originally developed without considering advanced architectures and related tool chains. Science needs frequently lead users to want to run large numbers of relatively small jobs (often called many-task computing, an ensemble, or a workflow), which can conflict with supercomputer configurations. In this paper, we discuss techniques developed to execute ordinary applications over leadership class supercomputers. We use the high-performance Swift parallel scripting framework and build two workflow execution techniques -- sub-jobs and main-wrap. The sub-jobs technique, built on top of the IBM Blue Gene/Q resource manager Cobalt's sub-block jobs, lets users submit multiple, independent, repeated smaller jobs within a single larger resource block. The main-wrap technique is a scheme that enables C/C++ programs to be defined as functions that are wrapped by a high-performance Swift wrapper and that are invoked as a Swift script. We discuss the needs, benefits, technicalities, and current limitations of these techniques. We further discuss the real-world science enabled by these techniques and the results obtained. Ketan Maheshwari, Justin M. Wozniak, Timothy G. Armstrong, Daniel S. Katz, T. Andrew Binkowski, Xiaoliang Zhong, Olle Heinonen, Dmitry Karpeyev, Michael Wilde |
e-Science | 3 |
| 2014 | Compiler Optimization for Extreme-Scale ScriptingabstractThe data-driven task parallelism execution model can support parallel programming models that are well suited for large-scale distributed-memory parallel computing, for example, simulations and analysis pipelines running on clusters and clouds. We describe a novel compiler intermediate representation and optimizations for this execution model, including adaptions of standard techniques alongside novel techniques. These techniques are applied to Swift/T, a high-level scripting language for flexible data flow composition of functions, which may be serial or use lower-level parallel programming models such as MPI and OpenMP. This paper presents preliminary results, indicating that our compiler optimizations reduce communication overhead by 70% to 93% on distributed-memory systems. Timothy G. Armstrong, Justin M. Wozniak, Michael Wilde, Ian T. Foster |
CCGRID | 1 |
| 2014 | Design and evaluation of the gemtc framework for GPU-enabled many-task computingabstractWe present the design and first performance and usability evaluation of GeMTC, a novel execution model and runtime system that enables accelerators to be programmed with many concurrent and independent tasks of potentially short or variable duration. With GeMTC, a broad class of such "many-task" applications can leverage the increasing number of accelerated and hybrid high-end computing systems. GeMTC overcomes the obstacles to using GPUs in a many-task manner by scheduling and launching independent tasks on hardware designed for SIMD-style vector processing. We demonstrate the use of a high-level MTC programming model (the Swift parallel dataflow language) to run tasks on many accelerators and thus provide a high-productivity programming model for the growing number of supercomputers that are accelerator-enabled. While still in an experimental stage, GeMTC can already support tasks of fine (subsecond) granularity and execute concurrent heterogeneous tasks on 86,000 independent GPU warps spanning 2.7M GPU threads on the Blue Waters supercomputer. Scott J. Krieder, Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Benjamin Grimmer, Ian T. Foster, Ioan Raicu |
HPDC | 3 |
| 2014 | Compiler Techniques for Massively Scalable Implicit Task ParallelismabstractSwift/T is a high-level language for writing concise, deterministic scripts that compose serial or parallel codes implemented in lower-level programming models into large-scale parallel applications. It executes using a data-driven task parallel execution model that is capable of orchestrating millions of concurrently executing asynchronous tasks on homogeneous or heterogeneous resources. Producing code that executes efficiently at this scale requires sophisticated compiler transformations: poorly optimized code inhibits scaling with excessive synchronization and communication. We present a comprehensive set of compiler techniques for data-driven task parallelism, including novel compiler optimizations and intermediate representations. We report application benchmark studies, including unbalanced tree search and simulated annealing, and demonstrate that our techniques greatly reduce communication overhead and enable extreme scalability, distributing up to 612 million dynamically load balanced tasks per second at scales of up to 262,144 cores without explicit parallelism, synchronization, or load balancing in application code. Timothy G. Armstrong, Justin M. Wozniak, Michael Wilde, Ian T. Foster |
SC | 1 |
| 2013 | Swift/T: Large-Scale Application Composition via Distributed-Memory Dataflow ProcessingabstractMany scientific applications are conceptually built up from independent component tasks as a parameter study, optimization, or other search. Large batches of these tasks may be executed on high-end computing systems, however, the coordination of the independent processes, their data, and their data dependencies is a significant scalability challenge. Many problems must be addressed, including load balancing, data distribution, notifications, concurrent programming, and linking to existing codes. In this work, we present Swift/T, a programming language and runtime that enables the rapid development of highly concurrent, task-parallel applications. Swift/Tis composed of several enabling technologies to address scalability challenges, offers a high-level optimizing compiler for user programming and debugging, and provides tools for binding user code in C/C++/Fortran into a logical script. In this work, we describe the Swift/T solution and present scaling results from the IBM Blue Gene/Pand Blue Gene/Q. Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Ewing L. Lusk, Ian T. Foster |
CCGRID | 2 |
| 2013 | Swift/T: scalable data flow programming for many-task applicationsabstractSwift/T, a novel programming language implementation for highly scalable data flow programs, is presented. Justin M. Wozniak, Timothy G. Armstrong, Michael Wilde, Daniel S. Katz, Ewing L. Lusk, Ian T. Foster |
PPoPP | 2 |
| 2013 | Dataflow coordination of data-parallel tasks via MPI 3.0abstractScientific applications are often complex collections of many large-scale tasks. Mature tools exist for describing task-parallel workflows consisting of serial tasks, and a variety of tools exist for programming a single data-parallel operation. However, few tools cover the intersection of these two models. In this work, we extend the load balancing library ADLB to support parallel tasks. We demonstrate how applications can easily be composed of parallel tasks using Swift dataflow scripts, which are compiled to ADLB programs with performance comparable to hand-coded equivalents. By combining this framework with data-parallel analysis libraries, we are able to dynamically execute many instances of a parallel data analysis application in support of a parameter exploration workload. Justin M. Wozniak, Tom Peterka, Timothy G. Armstrong, James Dinan, Ewing L. Lusk, Michael Wilde, Ian T. Foster |
EuroMPI | 3 |
| 2013 | Parallelizing the execution of sequential scriptsabstractScripting is often used in science to create applications via the composition of existing programs. Parallel scripting systems allow the creation of such applications, but each system introduces the need to adopt a somewhat specialized programming model. We present an alternative scripting approach, AMFS Shell, that lets programmers express parallel scripting applications via minor extensions to existing sequential scripting languages, such as Bash, and then execute them in-memory on large-scale computers. We define a small set of commands between the scripts and a parallel scripting runtime system, so that programmers can compose their scripts in a familiar scripting language. The underlying AMFS implements both collective (fast file movement) and functional (transformation based on content) file management. Tasks are handled by AMFS's built-in execution engine. AMFS Shell is expressive enough for a wide range of applications, and the framework can run such applications efficiently on large-scale computers. Zhao Zhang 0007, Daniel S. Katz, Timothy G. Armstrong, Justin M. Wozniak, Ian T. Foster |
SC | 3 |
| 2013 | LinkBench: a database benchmark based on the Facebook social graphabstractDatabase benchmarks are an important tool for database researchers and practitioners that ease the process of making informed comparisons between different database hardware, software and configurations. Large scale web services such as social networks are a major and growing database application area, but currently there are few benchmarks that accurately model web service workloads. Timothy G. Armstrong, Vamsi Ponnekanti, Dhruba Borthakur, Mark Callaghan |
SIGMOD Conference | 1 |
| 2013 | Turbine: A Distributed-memory Dataflow Engine for High Performance Many-task ApplicationsabstractEfficiently utilizing the rapidly increasing concurrency of multi-petaflop computing systems is a significant programming challenge. One approach is to structure applications with an upper layer of many loosely coupled coarse-grained tasks, each comp Justin M. Wozniak, Timothy G. Armstrong, Ketan Maheshwari, Ewing L. Lusk, Daniel S. Katz, Michael Wilde, Ian T. Foster |
Fundam. Informaticae | 2 |
| 2009 | Improvements that don't add up: ad-hoc retrieval results since 1998abstractThe existence and use of standard test collections in information retrieval experimentation allows results to be compared between research groups and over time. Such comparisons, however, are rarely made. Most researchers only report results from their own experiments, a practice that allows lack of overall improvement to go unnoticed. In this paper, we analyze results achieved on the TREC Ad-Hoc, Web, Terabyte, and Robust collections as reported in SIGIR (1998--2008) and CIKM (2004--2008). Dozens of individual published experiments report effectiveness improvements, and often claim statistical significance. However, there is little evidence of improvement in ad-hoc retrieval technology over the past decade. Baselines are generally weak, often being below the median original TREC system. And in only a handful of experiments is the score of the best TREC automatic run exceeded. Given this finding, we question the value of achieving even a statistically significant result over a weak baseline. We propose that the community adopt a practice of regular longitudinal comparison to ensure measurable progress, or at least prevent the lack of it from going unnoticed. We describe an online database of retrieval runs that facilitates such a practice. Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel |
CIKM | 1 |
| 2009 | Has adhoc retrieval improved since 1994?abstractEvaluation forums such as TREC allow systematic measurement and comparison of information retrieval techniques. The goal is consistent improvement, based on reliable comparison of the effectiveness of different approaches and systems. In this paper we report experiments to determine whether this goal has been achieved. We ran five publicly available search systems, in a total of seventeen different configurations, against nine TREC adhoc-style collections, spanning 1994 to 2005. These runsets were then used as a benchmark for reassessing the relative effectiveness of the original TREC runs for those collections. Surprisingly, there appears to have been no overall improvement in effectiveness for either median or top-end TREC submissions, even after allowing for several possible confounds. We therefore question whether the effectiveness of adhoc information retrieval has improved over the past decade and a half. Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel |
SIGIR | 1 |
| 2009 | EvaluatIR: an online tool for evaluating and comparing IR systemsabstractNo abstract available. Timothy G. Armstrong, Alistair Moffat, William Webber, Justin Zobel |
SIGIR | 1 |