EDBT 2026 Demo / reviewers in the wild / expert
Timothy Shaffer
dblp:145/4361 · also Tim Shaffer, Timothy R. Shaffer
· DBLP profile ↗
13ranked-venue papers
6as first author
4since 2021 · last 2023
0000-0002-7028-5246ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 3 first-author · 1 since 2021Systems, architecture and hardware · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Landlord: Coordinating Dynamic Software Environments to Reduce Container SprawlabstractContainers provide customizable software environments that are independent from the system on which they are deployed. Online services for task execution must often generate containers on the fly to meet user-generated requests. However, as the number of users grows and container environments are changed and updated over time, there is an explosion in the number of containers that must be managed, despite the fact that there is significant overlap among many of the containers in use. We analyze a trace of container launches on the public Binder service and demonstrate the performance and resource usage issues associated with container sprawl. We presentLandlord, an algorithm that coalesces related container environments, and show that it can improve container reuse and reduce the number of container builds required in the Binder trace by 40%. We perform a sensitivity analysis ofLandlordusing randomized synthetic workloads on a high-energy physics (HEP) software repository and demonstrate thatLandlordshows benefits for container management across a wide range of usage patterns. Finally, we compareLandlordto offline clustering, and observe that the continuous churn in software necessitates an online approach. Timothy Shaffer, Thanh Son Phung, Kyle Chard, Douglas Thain |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Infusing pub-sub storage with transactionsabstractThe need to support new features in existing storage systems is an ongoing concern for storage developers. So is the desire to develop next generation storage systems that can adopt newly developed feature improvements with relative ease. Extending storage systems is challenging because of the inherent complexity of their codebases and the need to ensure that the storage state does not become corrupt or inconsistent when enabling new features. In this work, we examine a new storage architecture, FDMI, that uses the well-established publish-subscribe model for extending the feature set of a host storage system using plugins. A central mechanism in FDMI is transactional coupling. With transactional coupling, the subscribed plugin can either create new transactions that execute asynchronously following the successful completion of the precipitating event or can participate in the pending transaction and control whether the precipitating event itself will or will not be committed. We further create a classification of transactional mechanisms as well as possible desired plugin functionality and explore the matrix of these two classifications to create a new model for faster, safer distributed storage development. Liana V. Rodriguez, John Bent, Timothy Shaffer, Raju Rangaswami |
HotStorage | 3 |
| 2021 | An Empirical Study of Package Dependencies and Lifetimes in Binder Python ContainersabstractContainers are widely used in scientific applications as they provide greater precision and flexibility in controlling nearly every aspect of the software environment. They can also be easily shared, enabling researchers to run with the same environment on different host systems—an important requirement for scientific reproducibility. In this work, we studied logs of container launches from Binder, a publicly accessible online service for executing Git repositories. Binder dynamically builds and deploys containers following a recipe stored in the repository. These logs capture usage over several years, and include nearly 14 million container launches of around 70,000 unique repositories. To gain more insight about the types of containers and software environments in use for container-based scientific computing services, we downloaded the user-provided recipe repositories referenced in the logs and captured the software specifications and repository metadata. We discovered a number of interesting trends that may be of interest to site administrators provisioning container-based infrastructure for scientific computing in Python. Based on this analysis, we found that automatically generated containers present unique management challenges that are not well handled by straightforward caching. We used historical metadata on package releases from Pip and Conda to quantify difficulties in keeping previously built containers up to date with changing external dependencies. Finally, we proposed several management strategies for reducing infrastructure costs and improving user experience when managing a large container-based service, and back-tested these strategies against Binder launch activity and historical package metadata to demonstrate the value of dependency-oriented container management. Timothy Shaffer, Kyle Chard, Douglas Thain |
e-Science | 1 |
| 2021 | Lightweight Function Monitors for Fine-Grained Management in Large Scale Python ApplicationsabstractPython has become a widely used programming language for research, not only for small one-off analyses, but also for complex application pipelines running at supercomputer-scale. Modern parallel programming frameworks for Python present users with a more granular unit of management than traditional Unix processes and batch submissions: the Python function. We review the challenges involved in running native Python functions at scale, and present techniques for dynamically determining a minimal set of dependencies and for assembling a lightweight function monitor (LFM) that captures the software environment and manages resources at the granularity of single functions. We evaluate these techniques in a range of environments, from campus cluster to supercomputer, and show that our advanced dependency management planning and dynamic resource management methods provide superior performance and utilization relative to coarser-grained management approaches, achieving several-fold decrease in execution time for several large Python applications. Timothy Shaffer, Zhuozhao Li, Benjamín Tovar, Yadu N. Babuji, T. J. Dasso, Zoe Surma, Kyle Chard, Ian T. Foster, Douglas Thain |
IPDPS | 1 |
| 2020 | Autoscaling High-Throughput Workloads on Container OrchestratorsabstractHigh-throughput computing (HTC) workloads seek to complete as many jobs as possible over a long period of time. Such workloads require efficient execution of many parallel jobs and can occupy a large number of resources for a long time. As a result, full utilization is the normal state of an HTC facility. The widespread use of container orchestrators eases the deployment of HTC frameworks across different platforms, which also provides an opportunity to scale up HTC workloads with almost infinite resources on the public cloud. However, the autoscaling mechanisms of container orchestrators are primarily designed to support latency-sensitive microservices, and result in unexpected behavior when presented with HTC workloads. In this paper, we design a feedback autoscaler, High Throughput Autoscaler (HTA), that leverages the unique characteristics of the HTC workload to autoscales the resource pools used by HTC workloads on container orchestrators. HTA takes into account a reference input, the real-time status of the jobs' queue, as well as two feedback inputs, resource consumption of jobs, and the resource initialization time of the container orchestrator. We implement HTA using the Makeflow workload manager, Work Queue job scheduler, and the Kubernetes cluster manager. We evaluate its performance on both CPU-bound and IO-bound workloads. The evaluation results show that, by using HTA, we improve resource utilization by 5.6× with a slight increase in execution time (about 15%) for a CPU-bound workload, and shorten the workload execution time by up to 3.65× for an IO-bound workload. Chao Zheng 0002, Nathaniel Kremer-Herman, Timothy Shaffer, Douglas Thain |
CLUSTER | 3 |
| 2020 | Solving the Container Explosion Problem for Distributed High Throughput ComputingabstractContainer technologies are seeing wider use at advanced computing facilities for managing highly complex applications that must execute at multiple sites. However, in a distributed high throughput computing setting, the unrestricted use of containers can result in the container explosion problem. If a new container image is generated for each variation of a job dispatched to a site, shared storage is soon exceeded. On the other hand, if a single large container image is used to meet multiple needs, the size of that container may become a problem for storage and transport. To address this problem, we observe that many containers have an internal structure generated by a structured package manager, and this information could be used to strategically combine and share container images. We develop Landlord to exploit this property and evaluate its performance through a combination of simulation studies and empirical measurement of high energy physics applications. Timothy Shaffer, Nicholas L. Hazekamp, Jakob Blomer, Douglas Thain |
IPDPS | 1 |
| 2018 | A First Look at the JX Workflow LanguageabstractScientific workflows are typically expressed as a graph of logical tasks, each one representing a single program along with its input and output files. This poster introduces JX (JSON eXtended), a declarative language that can express complex workloads as an assembly of sub-graphs that can be partitioned in flexible ways. We present a case study of using JX to represent complex workflows for the Lifemapper biodiversity project. We evaluate partitioning approaches across several computing environments, including ND-Condor, IU-Jetstream, and SDSC-Comet, and show that a coarse partitioning results in faster turnaround times, reduced data transfer, and lower master utilization across all three systems. Timothy Shaffer, Kyle M. D. Sweeney, Nathaniel Kremer-Herman, Douglas Thain |
eScience | 1 |
| 2017 | Eye movements in software traceability link recovery
Bonita Sharif, John Meinken, Timothy Shaffer, Huzefa H. Kagdi |
Empir. Softw. Eng. | 3 |
| 2017 | Eye gaze and interaction contexts for change tasks - Observations and potential
Katja Kevic, Braden Walters, Timothy Shaffer, Bonita Sharif, David C. Shepherd, Thomas Fritz 0001 |
J. Syst. Softw. | 3 |
| 2015 | Eye-Tracking Metrics in Software EngineeringabstractEye-tracking studies are getting more prevalent in software engineering. Researchers often use different metrics when publishing their results in eye-tracking studies. Even when the same metrics are used, they are given different names, causing difficulties in comparing studies. To encourage replications and facilitate advancing the state of the art, it is important that the metrics used by researchers be clearly and consistently defined in the literature. There is therefore a need for a survey of eye-tracking metrics to support the (future) goal of standardizing eye-tracking metrics. This paper seeks to bring awareness to the use of different metrics along with practical suggestions on using them. It compares and contrasts various eye-tracking metrics used in software engineering. It also provides definitions for common metrics and discusses some metrics that the software engineering community might borrow from other fields. Zohreh Sharafi, Timothy Shaffer, Bonita Sharif, Yann-Gaël Guéhéneuc |
APSEC | 2 |
| 2015 | Tracing software developers' eyes and interactions for change tasksabstractWhat are software developers doing during a change task? While an answer to this question opens countless opportunities to support developers in their work, only little is known about developers' detailed navigation behavior for realistic change tasks. Most empirical studies on developers performing change tasks are limited to very small code snippets or are limited by the granularity or the detail of the data collected for the study. In our research, we try to overcome these limitations by combining user interaction monitoring with very fine granular eye-tracking data that is automatically linked to the underlying source code entities in the IDE. In a study with 12 professional and 10 student developers working on three change tasks from an open source system, we used our approach to investigate the detailed navigation of developers for realistic change tasks. The results of our study show, amongst others, that the eye tracking data does indeed capture different aspects than user interaction data and that developers focus on only small parts of methods that are often related by data flow. We discuss our findings and their implications for better developer tool support. Katja Kevic, Braden Walters, Timothy Shaffer, Bonita Sharif, David C. Shepherd, Thomas Fritz 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2015 | iTrace: enabling eye tracking on software artifacts within the IDE to support software engineering tasksabstractThe paper presents iTrace, an Eclipse plugin that implicitly records developers' eye movements while they work on change tasks. iTrace is the first eye tracking environment that makes it possible for researchers to conduct eye tracking studies on large software systems. An overview of the design and architecture is presented along with features and usage scenarios. iTrace is designed to support a variety of eye trackers. The design is flexible enough to record eye movements on various types of software artifacts (Java code, text/html/xml documents, diagrams), as well as IDE user interface elements. The plugin has been successfully used for software traceability tasks and program comprehension tasks. iTrace is also applicable to other tasks such as code summarization and code recommendations based on developer eye movements. A short video demonstration is available at https://youtu.be/3OUnLCX4dXo. Timothy Shaffer, Jenna DiVincenzo, Braden Walters, Sebastian C. Müller, Michael Falcone, Bonita Sharif |
ESEC/SIGSOFT FSE | 1 |
| 2014 | Capturing software traceability links from developers' eye gazesabstractThe paper presents a novel approach for recovering software traceability links from developers' eye gazes. An eye tracker is used to capture eye gazes while developers perform software maintenance tasks within the Eclipse IDE. An algorithm is presented that establishes a set of traceability links from the eye-gaze data of several developer sessions. A preliminary study assesses the feasibility and validity of the approach. The links generated by the approach were validated by another set of developers. Results indicate that our algorithm achieves strong recall when developers accurately perform bug-localization tasks. Braden Walters, Timothy Shaffer, Bonita Sharif, Huzefa H. Kagdi |
ICPC | 2 |