VLDB 2026 Research / reviewers in the wild / expert
Connor Scully-Allison
dblp:231/5693 · also Connor Francis Scully-Allison
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-5990-6179ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSystems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue VisualizationabstractDomain-specific visualizations sometimes focus on narrow, albeit important, tasks for one group of users. This focus limits the utility of a visualization to other groups working with the same data. While tasks elicited from other groups can present a design pitfall if not disambiguated, they also present a design opportunity-namely, the development of visualizations that support multiple groups. This development choice presents a trade-off of broadening the scope but limiting support for the more narrow tasks of any one group, which in some cases can enhance the overall utility of the visualization. We investigate this scenario through a design study where we develop Guidepost, a notebook-embedded visualization of data that helps scientists assess compute wait times, machine learning researchers understand prediction accuracy, and system maintainers analyze usage trends. We adapt the use of personas for visualization design from existing literature in the HCI and design domains, applying them to categorize tasks based on their uniqueness across stakeholder personas. Under this model, tasks shared between all groups should be supported by interactive visualizations and tasks unique to each group can be deferred to scripting with notebook-embedded visualization design. We evaluate our visualization through real-world case studies and a task-focused evaluation with nine participants. We observe that together, Guidepost's visual encodings, interactions, and export capabilities support the tasks of our differing personas. Connor Scully-Allison, Kevin Menear, Kristi Potter, Andrew M. McNutt, Katherine E. Isaacs, Dmitry Duplyakin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Evaluating Communication Pattern Representations in Execution Trace Gantt ChartsabstractGantt charts are frequently used to explore execution traces of large-scale parallel programs. In these visualizations, each parallel processor is assigned a row showing the computation state of a processor at a particular time. Lines are drawn between rows to show communication between these processors. When drawn to align equivalent calls across rows, visual patterns can emerge reflecting communication behavior of the executing code. However, though these patterns have the same definition at any scale, they can be obscured by the density of rendered lines when displaying more than a few hundred processors. We seek to understand the effectiveness of various strategies for recognizing these patterns in Gantt charts. Specifically, we conduct a pre-registered user study comparing recognition of patterns when viewing all processors, a subset of processors, or a set of abstracted glyphs overlaid on the chart. We find that all strategies have limitations when scaling, motivating further designs. Our results further indicate that for simple patterns, the glyphs are more effective in general pattern recognition while the zoomed subsets provide nuance to specific characteristics, such as offsets, in patterns. These results suggest the development of a combined approach may be appropriate to enable pattern comprehension in large-scale Gantt charts. Connor Scully-Allison, Katherine E. Isaacs |
VISSOFT | 1 |
| 2024 | Design Concerns for Integrated Scripting and Interactive Visualization in Notebook EnvironmentsabstractInteractive visualization can support fluid exploration but is often limited to predetermined tasks. Scripting can support a vast range of queries but may be more cumbersome for free-form exploration. Embedding interactive visualization in scripting environments, such as computational notebooks, provides an opportunity to leverage the strengths of both direct manipulation and scripting. We investigate interactive visualization design methodology, choices, and strategies under this paradigm through a design study of calling context trees used in performance analysis, a field which exemplifies typical exploratory data analysis workflows with Big Data and hard to define problems. We first produce a formal task analysis assigning tasks to graphical or scripting contexts based on their specificity, frequency, and suitability. We then design a notebook-embedded interactive visualization and validate it with intended users. In a follow-up study, we present participants with multiple graphical and scripting interaction modes to elicit feedback about notebook-embedded visualization design, finding consensus in support of the interaction model. We report and reflect on observations regarding the process and design implications for combining visualization and scripting in notebooks. Connor Scully-Allison, Ian Lumsden, Katy Williams, Jesse Bartels, Michela Taufer, Stephanie Brink, Abhinav Bhatele, Olga Pearce, Katherine E. Isaacs |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Thicket: Seeing the Performance Experiment Forest for the Individual Run TreesabstractThicket is an open-source Python toolkit for Exploratory Data Analysis (EDA) of multi-run performance experiments. It enables an understanding of optimal performance configuration for large-scale application codes. Most performance tools focus on a single execution (e.g., single platform, single measurement tool, single scale). Thicket bridges the gap to convenient analysis in multi-dimensional, multi-scale, multi-architecture, and multi-tool performance datasets by providing an interface for interacting with the performance data. Thicket has a modular structure composed of three components. The first component is a data structure for multi-dimensional performance data, which is composed automatically on the portable basis of call trees, and accommodates any subset of dimensions present in the dataset. The second is the metadata, enabling distinction and sub-selection of dimensions in performance data. The third is a dimensionality reduction mechanism, enabling analysis such as computing aggregated statistics on a given data dimension. Extensible mechanisms are available for applying analyses (e.g., top-down on Intel CPUs), data science techniques (e.g., K-means clustering from scikit-learn), modeling performance (e.g., Extra-P), and interactive visualization. We demonstrate the power and flexibility of Thicket through two case studies, first with the open-source RAJA Performance Suite on CPU and GPU clusters and another with a large physics simulation run on both a traditional HPC cluster and an AWS Parallel Cluster instance. Stephanie Brink, Michael McKinsey, David Böhme, Connor Scully-Allison, Ian Lumsden, W. Daryl Hawkins, Treece Burgess, Vanessa Lama, Jakob Lüttgau, Katherine E. Isaacs, Michela Taufer, Olga Pearce |
HPDC | 4 |
| 2023 | Traveler: Navigating Task Parallel Traces for Performance AnalysisabstractUnderstanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace-a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. To address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler, an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library. Sayef Azad Sakin, Alex Bigelow, R. Tohid, Connor Scully-Allison, Carlos Scheidegger, Steven R. Brandt, Kevin A. Huck, Hartmut Kaiser, Katherine E. Isaacs |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Enabling Call Path Querying in Hatchet to Identify Performance Bottlenecks in Scientific ApplicationsabstractAs computational science applications benefit from larger-scale, more heterogeneous high performance computing (HPC) systems, the process of studying their performance becomes increasingly complex. The performance data analysis library Hatchet provides some insights into this complexity, but is currently limited in its analysis capabilities. Missing capabilities include the handling of relational caller-callee data captured by HPC profilers. To address this shortcoming, we augment Hatchet with a Call Path Query Language that leverages relational data in the performance analysis of scientific applications. Specifically, our Query Language enables data reduction using call path pattern matching. We demonstrate the effectiveness of our Query Language in identifying performance bottlenecks and enhancing Hatchet's analysis capabilities through three case studies. In the first case study, we compare the performance of sequential and multi-threaded versions of the graph alignment application Fido. In doing so, we identify the existence of large memory inefficiencies in both versions. In the second case study, we examine the performance of MPI calls in the linear algebra mini-application AMG2013 when using MVAPICH and Spectrum-MPI. In doing so, we identify hidden performance losses in specific MPI functions. In the third case study, we illustrate the use of our Query Language in Hatchet's interactive visualization. In doing so, we show that our Query Language enables a simple and intuitive way to massively reduce profiling data. Ian Lumsden, Jakob Lüttgau, Vanessa Lama, Connor Scully-Allison, Stephanie Brink, Katherine E. Isaacs, Olga Pearce, Michela Taufer |
e-Science | 4 |
| 2022 | PERMAVOST '22: Workshop on Performance EngineeRing, Modelling, Analysis, and VisualizatiOn StrategyabstractModern software engineering is getting increasingly complicated. Especially in the HPC field, we are dealing with cutting edge infrastructure and a novel problem with unprecedented scale. The ability to monitor and analyze the performance of such applications and infrastructure is imperative for the future of improvement, design, and maintenance. In the current era, the writing and maintenance of these applications have ceased to be the job solely of computer scientists and has grown to encompass a wide variety of experts in mathematics, science, and other engineering disciplines. The fact that many developers from these disciplines have not received a formal education in computer science and rely increasingly on the tools created by computer scientists to analyze and optimize their code shows that there's a need for a forum to work together. Radita Liem, Ana Veroneze Solorzano, Connor Scully-Allison |
HPDC | 3 |
| 2019 | Factors Influencing The Human Preferred Interaction DistanceabstractNonverbal interactions are a key component of human communication. Since robots have become significant by trying to get close to human beings, it is important that they follow social rules governing the use of space. Prior research has conceptualized personal space as physical zones which are based on static distances. This work examined how preferred interaction distance can change given different interaction scenarios. We conducted a user study using three different robot heights. We also examined the difference in preferred interaction distance when a robot approaches a human and, conversely, when a human approaches a robot. Factors included in quantitative analysis are the participants' gender, robot's height, and method of approach. Subjective measures included human comfort and perceived safety. The results obtained through this study shows that robot height, participant gender and method of approach were significant factors influencing measured proxemic zones and accordingly participant comfort. Subjective data showed that experiment respondents regarded robots in a more favorable light following their participation in this study. Furthermore, the NAO was perceived most positively by respondents according to various metrics and the PR2 Tall, most negatively. Vineeth Rajamohan, Connor Scully-Allison, Sergiu M. Dascalu, David Feil-Seifer |
RO-MAN | 2 |
| 2018 | Near Real-time Autonomous Quality Control for Streaming Environmental Sensor DataabstractIn this paper, we present a novel and accessible approach to time-series data validation: the Near-Real Time Autonomous Quality Control (NRAQC) system. The design, implementation, and impacts of this software are explored in detail within this paper. This software system, created in close conference with environmental scientists, leverages microservice design patterns employed for high volume web applications to develop a contemporary solution to the problem of data quality control with streaming sensor data. Through a comparative analysis between NRAQC and the GCE Toolbox, we argue that the web based deployment of QC software enhances accessibility to crucial tools required to make a robust and useful data product from raw measurements. Additionally, a key innovation of the NRAQC platform is its positive impact on modern data management practices and quality data dissemination. Connor Scully-Allison, Vinh D. Le, Eric Fritzinger, Scotty Strachan, Frederick C. Harris Jr., Sergiu M. Dascalu |
KES | 1 |