Justin Talbot

dblp:71/2250 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
10 papers
Visualization and visual analytics · 93% Rendering · 7%
Databases, data mining, and information retrieval
5 papers
Database system architecture and tuning · 82% Information retrieval · 14% Knowledge graphs · 2%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 56% Parallel and multicore computing · 44%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 77% User interface design and tools · 23%

Topics — the 22 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning
database design
0.612022
Statistical Schema Learning with Occam's Razor · SIGMOD Conference 2022
Database system architecture and tuning › database design
schema normalization
0.612022
Statistical Schema Learning with Occam's Razor · SIGMOD Conference 2022
Visualization and visual analytics
graphical perception
0.532014
Four Experiments on the Perception of Bar Charts · IEEE Trans. Vis. Comput. Graph. 2014
An Empirical Model of Slope Ratio Comparisons · IEEE Trans. Vis. Comput. Graph. 2012
Arc Length-Based Aspect Ratio Selection · IEEE Trans. Vis. Comput. Graph. 2011
Visualization and visual analytics › visualization design
aspect ratio selection
0.322012
An Empirical Model of Slope Ratio Comparisons · IEEE Trans. Vis. Comput. Graph. 2012
Arc Length-Based Aspect Ratio Selection · IEEE Trans. Vis. Comput. Graph. 2011
Visualization and visual analytics › multi-view visualization
small multiples
0.212016
Automatic Selection of Partitioning Variables for Small Multiple Displays · IEEE Trans. Vis. Comput. Graph. 2016
Visualization and visual analytics
interactive data exploration
0.222009
Vispedia: on-demand data integration for interactive visualization and exploration · SIGMOD Conference 2009
Vispedia: Interactive Visual Exploration of Wikipedia Data via Search-Based Integration · IEEE Trans. Vis. Comput. Graph. 2008
Visualization and visual analytics
automatic graphic design
0.112010
An Extension of Wilkinson's Algorithm for Positioning Tick Labels on Axes · IEEE Trans. Vis. Comput. Graph. 2010
Distributed systems › stream processing
continuous query
0.112010
Online aggregation and continuous query support in MapReduce · SIGMOD Conference 2010
Parallel and multicore computing › data-parallel programming
mapreduce
0.112010
Online aggregation and continuous query support in MapReduce · SIGMOD Conference 2010
Visualization and visual analytics
visual analytics
0.112009
Vispedia: on-demand data integration for interactive visualization and exploration · SIGMOD Conference 2009
Human-AI interaction
interactive machine learning
0.112009
EnsembleMatrix: interactive visualization to support machine learning with multiple classifiers · CHI 2009
Information retrieval › interactive information retrieval
exploratory search
0.112008
Structuring collections with Scatter/Gather extensions · SIGIR 2008
Information retrieval › document retrieval › interactive document retrieval
scatter/gather
0.112008
Structuring collections with Scatter/Gather extensions · SIGIR 2008
Visualization and visual analytics
data visualization
0.112008
Vispedia: Interactive Visual Exploration of Wikipedia Data via Search-Based Integration · IEEE Trans. Vis. Comput. Graph. 2008
Visualization and visual analytics › information visualization
heterogeneous data visualization
0.112007
Visualization of Heterogeneous Data · IEEE Trans. Vis. Comput. Graph. 2007
Rendering
global illumination
0.112005
Energy redistribution path tracing · ACM Trans. Graph. 2005
Rendering
monte carlo integration
0.112005
Energy redistribution path tracing · ACM Trans. Graph. 2005
Rendering › ray tracing
path tracing
0.112005
Energy redistribution path tracing · ACM Trans. Graph. 2005
Distributed systems
fault tolerance
0.012010
Online aggregation and continuous query support in MapReduce · SIGMOD Conference 2010
Data mining › clustering
document clustering
0.012008
Structuring collections with Scatter/Gather extensions · SIGIR 2008
Information retrieval › interactive information retrieval
sensemaking
0.012008
Structuring collections with Scatter/Gather extensions · SIGIR 2008
Data models and query languages › semistructured data
RDF data
0.012007
Visualization of Heterogeneous Data · IEEE Trans. Vis. Comput. Graph. 2007

Methods — techniques the papers use, named apart from their topics

unsupervised learning · 0.6occam's razor · 0.6permutation test · 0.2non-parametric testing · 0.2keyword query · 0.2graph search · 0.2perceptual experiment · 0.2user experiment · 0.1empirical modeling · 0.1arc length minimization · 0.1search algorithm · 0.1pipelining · 0.1optimization · 0.1user study · 0.1ensemble learning · 0.1search-based integration · 0.1path search · 0.1schema matching · 0.1
YearPublicationVenuePosition
2022 Statistical Schema Learning with Occam's Razor
abstract
A judiciously normalized database schema can increase data interpretability, reduce data size, and improve data integrity. However, real world data sets are often stored or shared in a denormalized state. We examine the problem of automatically creating a good schema for a denormalized table, approaching it as an unsupervised machine learning problem which must learn an optimal schema from the data. This differs from past rule-based approaches that focus on normalization into a canonical form. We define a principled schema optimization criterion, based on Occam's razor, that is robust to noise and extensible---allowing users to easily specify desirable properties of the resulting schema. We develop an efficient learning algorithm for this criterion and empirically demonstrate that it is 3 to 100 times faster than previous work and produces higher quality schemas with 1/5th the errors.
Justin Talbot, Daniel Ting
SIGMOD Conference1
2016 Automatic Selection of Partitioning Variables for Small Multiple Displays
abstract
Effective small multiple displays are created by partitioning a visualization on variables that reveal interesting conditional structure in the data. We propose a method that automatically ranks partitioning variables, allowing analysts to focus on the most promising small multiple displays. Our approach is based on a randomized, non-parametric permutation test, which allows us to handle a wide range of quality measures for visual patterns defined on many different visualization types, while discounting spurious patterns. We demonstrate the effectiveness of our approach on scatterplots of real-world, multidimensional datasets.
Anushka Anand, Justin Talbot
IEEE Trans. Vis. Comput. Graph.2
2014 Four Experiments on the Perception of Bar Charts
abstract
Bar charts are one of the most common visualization types. In a classic graphical perception paper, Cleveland & McGill studied how different bar chart designs impact the accuracy with which viewers can complete simple perceptual tasks. They found that people perform substantially worse on stacked bar charts than on aligned bar charts, and that comparisons between adjacent bars are more accurate than between widely separated bars. However, the study did not explore why these differences occur. In this paper, we describe a series of follow-up experiments to further explore and explain their results. While our results generally confirm Cleveland & McGill's ranking of various bar chart configurations, we provide additional insight into the bar chart reading task and the sources of participants' errors. We use our results to propose new hypotheses on the perception of bar charts.
Justin Talbot, Vidya Setlur, Anushka Anand
IEEE Trans. Vis. Comput. Graph.1
2012 Riposte: a trace-driven compiler and parallel VM for vector code in R
abstract
There is a growing utilization gap between modern hardware and modern programming languages for data analysis.Due to power and other constraints, recent processor design has sought improved performance through increased SIMD and multi-core parallelism. At the same time, high-level, dynamically-typed languages for data analysis have become popular. These languages emphasize ease of use and high productivity, but have, in general, low performance and limited support for exploiting hardware parallelism.
Justin Talbot, Zach DeVito, Pat Hanrahan
PACT1
2012 An Empirical Model of Slope Ratio Comparisons
abstract
Comparing slopes is a fundamental graph reading task and the aspect ratio chosen for a plot influences how easy these comparisons are to make. According to Banking to 45°, a classic design guideline first proposed and studied by Cleveland et al., aspect ratios that center slopes around 45° minimize errors in visual judgments of slope ratios. This paper revisits this earlier work. Through exploratory pilot studies that expand Cleveland et al.'s experimental design, we develop an empirical model of slope ratio estimation that fits more extreme slope ratio judgments and two common slope ratio estimation strategies. We then run two experiments to validate our model. In the first, we show that our model fits more generally than the one proposed by Cleveland et al. and we find that, in general, slope ratio errors are not minimized around 45°. In the second experiment, we explore a novel hypothesis raised by our model: that visible baselines can substantially mitigate errors made in slope judgments. We conclude with an application of our model to aspect ratio selection.
Justin Talbot, John Gerth, Pat Hanrahan
IEEE Trans. Vis. Comput. Graph.1
2011 Arc Length-Based Aspect Ratio Selection
abstract
The aspect ratio of a plot has a dramatic impact on our ability to perceive trends and patterns in the data. Previous approaches for automatically selecting the aspect ratio have been based on adjusting the orientations or angles of the line segments in the plot. In contrast, we recommend a simple, effective method for selecting the aspect ratio: minimize the arc length of the data curve while keeping the area of the plot constant. The approach is parameterization invariant, robust to a wide range of inputs, preserves visual symmetries in the data, and is a compromise between previously proposed techniques. Further, we demonstrate that it can be effectively used to select the aspect ratio of contour plots. We believe arc length should become the default aspect ratio selection method.
Justin Talbot, John Gerth, Pat Hanrahan
IEEE Trans. Vis. Comput. Graph.1
2010 Online aggregation and continuous query support in MapReduce
abstract
MapReduce is a popular framework for data-intensive distributed computing of batch jobs. To simplify fault tolerance, the output of each MapReduce task and job is materialized to disk before it is consumed. In this demonstration, we describe a modified MapReduce architecture that allows data to be pipelined between operators. This extends the MapReduce programming model beyond batch processing, and can reduce completion times and improve system utilization for batch jobs as well. We demonstrate a modified version of the Hadoop MapReduce framework that supports online aggregation, which allows users to see "early returns" from a job as it is being computed. Our Hadoop Online Prototype (HOP) also supports continuous queries, which enable MapReduce programs to be written for applications such as event monitoring and stream processing. HOP retains the fault tolerance properties of Hadoop, and can run unmodified user-defined MapReduce programs.
Tyson Condie, Neil Conway, Peter Alvaro, Joseph M. Hellerstein, John Gerth, Justin Talbot, Khaled Elmeleegy, Russell Sears
SIGMOD Conference6
2010 An Extension of Wilkinson's Algorithm for Positioning Tick Labels on Axes
abstract
The non-data components of a visualization, such as axes and legends, can often be just as important as the data itself. They provide contextual information essential to interpreting the data. In this paper, we describe an automated system for choosing positions and labels for axis tick marks. Our system extends Wilkinson’s optimization-based labeling approach to create a more robust, full-featured axis labeler. We define an expanded space of axis labelings by automatically generating additional nice numbers as needed and by permitting the extreme labels to occur inside the data range. These changes provide flexibility in problematic cases, without degrading quality elsewhere. We also propose an additional optimization criterion, legibility, which allows us to simultaneously optimize over label formatting, font size, and orientation. To solve this revised optimization problem, we describe the optimization function and an efficient search algorithm. Finally, we compare our method to previous work using both quantitative and qualitative metrics. This paper is a good example of how ideas from automated graphic design can be applied to information visualization.
Justin Talbot, Sharon Lin, Pat Hanrahan
IEEE Trans. Vis. Comput. Graph.1
2009 EnsembleMatrix: interactive visualization to support machine learning with multiple classifiers
abstract
Machine learning is an increasingly used computational tool within human-computer interaction research. While most researchers currently utilize an iterative approach to refining classifier models and performance, we propose that ensemble classification techniques may be a viable and even preferable alternative. In ensemble learning, algorithms combine multiple classifiers to build one that is superior to its components. In this paper, we present EnsembleMatrix, an interactive visualization system that presents a graphical view of confusion matrices to help users understand relative merits of various classifiers. EnsembleMatrix allows users to directly interact with the visualizations in order to explore and build combination models. We evaluate the efficacy of the system and the approach in a user study. Results show that users are able to quickly combine multiple classifiers operating on multiple feature sets to produce an ensemble classifier with accuracy that approaches best-reported performance classifying images in the CalTech-101 dataset.
Justin Talbot, Bongshin Lee, Ashish Kapoor, Desney S. Tan
CHI1
2009 Vispedia: on-demand data integration for interactive visualization and exploration
abstract
Wikipedia is an example of the large, collaborative, semi-structured data sets emerging on the Web. Typically, before these data sets can be used, they must transformed into structured tables via data integration. We present Vispedia, a Web-based visualization system which incorporates data integration into an iterative, interactive data exploration and analysis process. This reduces the upfront cost of using heterogeneous data sets like Wikipedia. Vispedia is driven by a keyword-query-based integration interface implemented using a fast graph search. The search occurs interactively over DBpedia's semantic graph of Wikipedia, without depending on the existence of a structured ontology. This combination of data integration and visualization enables a broad class of non-expert users to more effectively use the semi-structured data available on the Web.
Bryan Chan 0001, Justin Talbot, Leslie Wu, Nathan Sakunkoo, Mike Cammarano, Pat Hanrahan
SIGMOD Conference2
2008 Structuring collections with Scatter/Gather extensions
abstract
A major component of sense-making is organizing--grouping, labeling, and summarizing--the data at hand in order to form a useful mental model, a necessary precursor to identifying missing information and to reasoning about the data. Previous work has shown the Scatter/Gather model to be useful in exploratory activities that occur when users encounter unknown document collections. However, the topic structure communicated by Scatter/Gather is closely tied to the behavior of the underlying clustering algorithm; this structure may not reflect the mental model most applicable to the information need. In this paper we describe the initial design of a mixed-initiative information structuring tool that leverages aspects of the well-studied Scatter/Gather model but permits the user to impose their own desired structure when necessary.
Omar Alonso, Justin Talbot
SIGIR2
2008 Vispedia: Interactive Visual Exploration of Wikipedia Data via Search-Based Integration
abstract
Wikipedia is an example of the collaborative, semi-structured data sets emerging on the Web. These data sets have large, non-uniform schema that require costly data integration into structured tables before visualization can begin. We present Vispedia, a Web-based visualization system that reduces the cost of this data integration. Users can browse Wikipedia, select an interesting data table, then use a search interface to discover, integrate, and visualize additional columns of data drawn from multiple Wikipedia articles. This interaction is supported by a fast path search algorithm over DBpedia, a semantic graph extracted from Wikipedia's hyperlink structure. Vispedia can also export the augmented data tables produced for use in traditional visualization systems. We believe that these techniques begin to address the "long tail" of visualization by allowing a wider audience to visualize a broader class of data. We evaluated this system in a first-use formative lab study. Study participants were able to quickly create effective visualizations for a diverse set of domains, performing data integration as needed.
Bryan Chan 0001, Leslie Wu, Justin Talbot, Mike Cammarano, Pat Hanrahan
IEEE Trans. Vis. Comput. Graph.3
2007 Visualization of Heterogeneous Data
abstract
Both the Resource Description Framework (RDF), used in the semantic web, and Maya Viz u-forms represent data as a graph of objects connected by labeled edges. Existing systems for flexible visualization of this kind of data require manual specification of the possible visualization roles for each data attribute. When the schema is large and unfamiliar, this requirement inhibits exploratory visualization by requiring a costly up-front data integration step. To eliminate this step, we propose an automatic technique for mapping data attributes to visualization attributes. We formulate this as a schema matching problem, finding appropriate paths in the data model for each required visualization attribute in a visualization template.
Mike Cammarano, Xin Dong 0001, Bryan Chan 0001, Jeff Klingner, Justin Talbot, Alon Y. Halevy, Pat Hanrahan
IEEE Trans. Vis. Comput. Graph.5
2006 Two Stage Importance Sampling for Direct Lighting
David Cline, Parris K. Egbert, Justin Talbot, David L. Cardon
Rendering Techniques3
2005 Importance Resampling for Global Illumination
Justin Talbot, David Cline, Parris K. Egbert
Rendering Techniques1
2005 Energy redistribution path tracing
abstract
We present Energy Redistribution (ER) sampling as an unbiased method to solve correlated integral problems. ER sampling is a hybrid algorithm that uses Metropolis sampling-like mutation strategies in a standard Monte Carlo integration setting, rather than resorting to an intermediate probability distribution step. In the context of global illumination, we present Energy Redistribution Path Tracing (ERPT). Beginning with an inital set of light samples taken from a path tracer, ERPT uses path mutations to redistribute the energy of the samples over the image plane to reduce variance. The result is a global illumination algorithm that is conceptually simpler than Metropolis Light Transport (MLT) while retaining its most powerful feature, path mutation. We compare images generated with the new technique to standard path tracing and MLT.
David Cline, Justin Talbot, Parris K. Egbert
ACM Trans. Graph.2