David Koop

dblp:54/1720 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0002-4422-6162ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
5 papers
Visualization and visual analytics · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 50% Data models and query languages · 50%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics › visualization theory
visualization pipeline
1.122026
Analyzing Notebook Histories to Understand Data Visualization Workflows · IEEE Trans. Vis. Comput. Graph. 2026
VisComplete: Automating Suggestions for Visualization Pipelines · IEEE Trans. Vis. Comput. Graph. 2008
Empirical software engineering
mining software repositories
1.012026
Analyzing Notebook Histories to Understand Data Visualization Workflows · IEEE Trans. Vis. Comput. Graph. 2026
Visualization and visual analytics
multi-view visualization
0.622022
Towards Systematic Design Considerations for Visualizing Cross-View Data Relationships · IEEE Trans. Vis. Comput. Graph. 2022
Querying and Creating Visualizations by Analogy · IEEE Trans. Vis. Comput. Graph. 2007
Visualization and visual analytics
visualization design
0.612022
Towards Systematic Design Considerations for Visualizing Cross-View Data Relationships · IEEE Trans. Vis. Comput. Graph. 2022
Visualization and visual analytics › visual analytics
provenance
0.522019
Enhancing Web-based Analytics Applications through Provenance · IEEE Trans. Vis. Comput. Graph. 2019
Querying and Creating Visualizations by Analogy · IEEE Trans. Vis. Comput. Graph. 2007
Visualization and visual analytics › visual analytics › interactive visual analysis
collaborative visual analytics
0.112019
Enhancing Web-based Analytics Applications through Provenance · IEEE Trans. Vis. Comput. Graph. 2019
Data models and query languages › query interface
query by example
0.112008
Querying and re-using workflows with VsTrails · SIGMOD Conference 2008
Data integration and cleaning › data provenance
workflow provenance
0.112008
Querying and re-using workflows with VsTrails · SIGMOD Conference 2008
Visualization and visual analytics
visualization generation
0.112007
Querying and Creating Visualizations by Analogy · IEEE Trans. Vis. Comput. Graph. 2007
Visualization and visual analytics › visual analytics
visual analytics system
0.012007
Querying and Creating Visualizations by Analogy · IEEE Trans. Vis. Comput. Graph. 2007

Methods — techniques the papers use, named apart from their topics

version history analysis · 2.0framework comparison · 2.0brushing and linking · 0.6streaming data · 0.4client-side javascript · 0.4subgraph matching · 0.1provenance-enabled workflow refinement · 0.1consensus-based prediction · 0.1analogy-based workflow refinement · 0.1provenance metadata · 0.1analogy-based recommendation · 0.1
YearPublicationVenuePosition
2026 Distribution-Informed Overviews for Quantitative Point Values
David Koop
PacificVis1
2026 Analyzing Notebook Histories to Understand Data Visualization Workflows
abstract
Visualization design is often a demanding process that involves trying different encodings, exploring different data transformations, and refining details. While there have been important studies of these workflows, the lower-level, code-intensive practices remain underexplored. Exploratory notebook tools have allowed designers to rapidly iterate on visualizations. We use publicly-available version histories of notebooks to study how users work in these environments, observing both how they build new visualizations from existing templates or previous work and how they refine visualizations over time. We examine the interplay between data manipulation and visualization, and classify the types of changes made when updating visual encodings. We also analyze the impact of different frameworks by comparing two code-oriented libraries and a chart wizard. Finally, we examine how interactions with notebooks have changed over the years, including after the widespread availability of AI. These analyses help us understand how users iterate to produce visualizations over time using different frameworks.
David Koop, Colin Brown, Hamed Alhoori, Maoyuan Sun
IEEE Trans. Vis. Comput. Graph.1
2025 iTrace: Interactive tracing of Cross-View Data Relationships
abstract
Exploring data relations across multiple views has been a common task in many domains such as bioinformatics, cybersecurity, and healthcare. To support this, various techniques (e.g., visual links and brushing & linking) are used to show related visual elements across views via lines and highlights. However, understanding the relations using these techniques, when many related elements are scattered, can be difficult due to spatial distance and complexity. To address this, we present iTrace, an interactive visualization technique to effectively trace cross-view data relationships. iTrace leverages the concept of interactive focus transitions, which allows users to see and directly manipulate their focus as they navigate between views. By directing the user’s attention through smooth transitions between related elements, iTrace makes it easier to follow data relationships. We demonstrate the effectiveness of iTrace with a user study, and we conclude with a discussion of how iTrace can be broadly used to enhance data exploration in various types of visualizations.
Abdul Rahman Shaikh, Maoyuan Sun, Hamed Alhoori, Jian Zhao 0010, David Koop
Graphics Interface6
2024 Navigating the Landscape of Reproducible Research: A Predictive Modeling Approach
abstract
The reproducibility of scientific articles is central to the advancement of science. Despite this importance, evaluating reproducibility remains challenging due to the scarcity of ground truth data. Predictive models can address this limitation by streamlining the tedious evaluation process. Typically, a paper's reproducibility is inferred based on the availability of artifacts such as code, data, or supplemental information, often without extensive empirical investigation. To address these issues, we utilized artifacts of papers as fundamental units to develop a novel, dual-spectrum framework that focuses on author-centric and external-agent perspectives. We used the author-centric spectrum, followed by the external-agent spectrum, to guide a structured, model-based approach to quantify and assess reproducibility. We explored the interdependencies between different factors influencing reproducibility and found that linguistic features such as readability and lexical diversity are strongly correlated with papers achieving the highest statuses on both spectrums. Our work provides a model-driven pathway for evaluating the reproducibility of scientific research.
Akhil Pandey Akella, Sagnik Ray Choudhury, David Koop, Hamed Alhoori
CIKM3
2022 Towards Systematic Design Considerations for Visualizing Cross-View Data Relationships
abstract
Due to the scale of data and the complexity of analysis tasks, insight discovery often requires coordinating multiple visualizations (views), with each view displaying different parts of data or the same data from different perspectives. For example, to analyze car sales records, a marketing analyst uses a line chart to visualize the trend of car sales, a scatterplot to inspect the price and horsepower of different cars, and a matrix to compare the transaction amounts in types of deals. To explore related information across multiple views, current visual analysis tools heavily rely on brushing and linking techniques, which may require a significant amount of user effort (e.g., many trial-and-error attempts). There may be other efficient and effective ways of displaying cross-view data relationships to support data analysis with multiple views, but currently there are no guidelines to address this design challenge. In this article, we present systematic design considerations for visualizing cross-view data relationships, which leverages descriptive aspects of relationships and usable visual context of multi-view visualizations. We discuss pros and cons of different designs for showing cross-view data relationships, and provide a set of recommendations for helping practitioners make design decisions.
Maoyuan Sun, Akhil Namburi, David Koop, Jian Zhao 0010, Tianyi Li 0008, Haeyong Chung
IEEE Trans. Vis. Comput. Graph.3
2019 Enhancing Web-based Analytics Applications through Provenance
abstract
Visual analytics systems continue to integrate new technologies and leverage modern environments for exploration and collaboration, making tools and techniques available to a wide audience through web browsers. Many of these systems have been developed with rich interactions, offering users the opportunity to examine details and explore hypotheses that have not been directly encoded by a designer. Understanding is enhanced when users can replay and revisit the steps in the sensemaking process, and in collaborative settings, it is especially important to be able to review not only the current state but also what decisions were made along the way. Unfortunately, many web-based systems lack the ability to capture such reasoning, and the path to a result is transient, forgotten when a user moves to a new view. This paper explores the requirements to augment existing client-side web applications with support for capturing, reviewing, sharing, and reusing steps in the reasoning process. Furthermore, it considers situations where decisions are made with streaming data, and the insights gained from revisiting those choices when more data is available. It presents a proof of concept, the Shareable Interactive Manipulation Provenance framework (SIMProv.js), that addresses these requirements in a modern, client-side JavaScript library, and describes how it can be integrated with existing frameworks.
Akhilesh Camisetty, Chaitanya Chandurkar, Maoyuan Sun, David Koop
IEEE Trans. Vis. Comput. Graph.4
2013 Visual summaries for graph collections
abstract
Graphs can be used to represent a variety of information, from molecular structures to biological pathways to computational workflows. With a growing volume of data represented as graphs, the problem of understanding and analyzing the variations in a collection of graphs is of increasing importance. We present an algorithm to compute a single summary graph that efficiently encodes an entire collection of graphs by finding and merging similar nodes and edges. Instead of only merging nodes and edges that are exactly the same, we use domain-specific comparison functions to collapse similar nodes and edges which allows us to generate more compact representations of the collection. In addition, we have developed methods that allow users to interactively control the display of these summary graphs. These interactions include the ability to highlight individual graphs in the summary, control the succinctness of the summary, and explicitly define when specific nodes should or should not be merged. We show that our approach to generating and interacting with graph summaries leads to a better understanding of a graph collection by allowing users to more easily identify common substructures and key differences between graphs.
David Koop, Juliana Freire, Cláudio T. Silva
PacificVis1
2010 Bridging Workflow and Data Provenance Using Strong Links
David Koop, Emanuele Santos, Bela Bauer, Matthias Troyer, Juliana Freire, Cláudio T. Silva
SSDBM1
2009 Using Workflow Medleys to Streamline Exploratory Tasks
Emanuele Santos, David Koop, Huy T. Vo, Erik W. Anderson, Juliana Freire, Cláudio T. Silva
SSDBM2
2008 Using Mediation to Achieve Provenance Interoperability (Extended Abstract)
abstract
Provenance is essential in scientific experiments. It contains information that is key to preserving the data and to determine it's quality and authorship. In complex experiments and analyses, where multiple tools are used to derive data products, provenance captured by these tools must be combined in order to determine the complete lineage of the derived products. We propose a mediator-based architecture to integrate provenance information from multiple sources, which contains two key components: a global mediated schema that is general and capable of representing provenance information represented in different models; and an expressive query interface that supports complex queries over provenance information spread over multiple different sources. We also present a case study where we show how this model was applied to integrate provenance from three provenance-enabled systems.
Tommy Ellkvist, David Koop, Juliana Freire, Cláudio T. Silva, Lena Strömbäck
eScience2
2008 Querying and re-using workflows with VsTrails
abstract
We show how work flow systems can be augmented to leverage provenance information to enhance usability. In particular, we will demonstrate new mechanisms and intuitive user interfaces designed to allow users to query work flows by example and to refine work flows by analogies. These techniques are implemented in VisTrails, an open-source provenance-enabled scientific work flow system that can be combined with a wide range of tools, libraries, and visualization systems. We will show di erent scenarios where these techniques can be used to simplify the notoriously hard tasks of creating and refining work flows.
Carlos Scheidegger, Huy T. Vo, David Koop, Juliana Freire, Cláudio T. Silva
SIGMOD Conference3
2008 Examining Statistics of Workflow Evolution Provenance: A First Study
Lauro Didier Lins, David Koop, Erik W. Anderson, Steven P. Callahan, Emanuele Santos, Carlos Scheidegger, Juliana Freire, Cláudio T. Silva
SSDBM2
2008 Special Issue: The First Provenance Challenge
abstract
Abstract The first Provenance Challenge was set up in order to provide a forum for the community to understand the capabilities of different provenance systems and the expressiveness of their provenance representations. To this end, a functional magnetic resonance imaging workflow was defined, which participants had to either simulate or run in order to produce some provenance representation, from which a set of identified queries had to be implemented and executed. Sixteen teams responded to the challenge, and submitted their inputs. In this paper, we present the challenge workflow and queries, and summarize the participants' contributions. Copyright © 2007 John Wiley & Sons, Ltd.
Luc Moreau 0001, Bertram Ludäscher, Ilkay Altintas, Roger S. Barga, Shawn Bowers, Steven P. Callahan, George Chin, Ben Clifford, Shirley Cohen, Sarah Cohen Boulakia, Susan B. Davidson, Ewa Deelman, Luciano A. Digiampietri, Ian T. Foster, Juliana Freire, James Frew, Joe Futrelle, Tara Gibson, Yolanda Gil, Carole A. Goble, Jennifer Golbeck, Paul Groth, David A. Holland, Jihie Kim, David Koop, Ales Krenek, Timothy M. McPhillips, Gaurang Mehta, Simon Miles, Dominic Metzger, Steve Munroe, James D. Myers, Beth Plale, Norbert Podhorszki, Varun Ratnakar, Emanuele Santos, Carlos Scheidegger, Karen Schuchardt, Margo I. Seltzer, Yogesh L. Simmhan, Cláudio T. Silva, Peter Slaughter, Eric G. Stephan, Robert Stevens 0001, Daniele Turi, Huy T. Vo, Michael Wilde, Jun Zhao 0003, Yong Zhao 0009
Concurr. Comput. Pract. Exp.26
2008 Tackling the Provenance Challenge one layer at a time
abstract
Abstract VisTrails is a new workflow and provenance management system that provides support for scientific data exploration and visualization. Whereas workflows have been traditionally used to automate repetitive tasks, for applications that are exploratory in nature, change is the norm. VisTrails uses a new change‐based provenance mechanism, which was designed to handle rapidly evolving workflows. It uniformly and automatically captures provenance information for data products and for the evolution of the workflows used to generate these products. In this paper, we describe how the VisTrails provenance data are organized in layers and present a first approach for querying this data that we developed to tackle the Provenance Challenge queries. Copyright © 2007 John Wiley & Sons, Ltd.
Carlos Scheidegger, David Koop, Emanuele Santos, Huy T. Vo, Steven P. Callahan, Juliana Freire, Cláudio T. Silva
Concurr. Comput. Pract. Exp.2
2008 VisComplete: Automating Suggestions for Visualization Pipelines
abstract
Building visualization and analysis pipelines is a large hurdle in the adoption of visualization and workflow systems by domain scientists. In this paper, we propose techniques to help users construct pipelines by consensus--automatically suggesting completions based on a database of previously created pipelines. In particular, we compute correspondences between existing pipeline subgraphs from the database, and use these to predict sets of likely pipeline additions to a given partial pipeline. By presenting these predictions in a carefully designed interface, users can create visualizations and other data products more efficiently because they can augment their normal work patterns with the suggested completions. We present an implementation of our technique in a publicly-available, open-source scientific workflow system and demonstrate efficiency gains in real-world situations.
David Koop, Carlos Scheidegger, Steven P. Callahan, Juliana Freire, Cláudio T. Silva
IEEE Trans. Vis. Comput. Graph.1
2007 Querying and Creating Visualizations by Analogy
abstract
While there have been advances in visualization systems, particularly in multi-view visualizations and visual exploration, the process of building visualizations remains a major bottleneck in data exploration. We show that provenance metadata collected during the creation of pipelines can be reused to suggest similar content in related visualizations and guide semi-automated changes. We introduce the idea of query-by-example in the context of an ensemble of visualizations, and the use of analogies as first-class operations in a system to guide scalable interactions. We describe an implementation of these techniques in VisTrails, a publicly-available, open-source system.
Carlos Scheidegger, Huy T. Vo, David Koop, Juliana Freire, Cláudio T. Silva
IEEE Trans. Vis. Comput. Graph.3