VLDB 2026 Research / reviewers in the wild / expert
Zhicheng Liu 0001
dblp:53/1020-1
· DBLP profile ↗
45ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-1015-2759ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 20 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VisAnatomy: An SVG Chart Corpus with Fine-Grained Semantic LabelsabstractChart corpora, which comprise data visualizations and their semantic labels, are crucial for advancing visualization research. However, the labels in most existing corpora are high-level (e.g., chart types), hindering their utility for broader applications in the era of AI. In this paper, we contribute VisAnatomy, a corpus containing 942 real-world SVG charts produced by over 50 tools, encompassing 40 chart types and featuring structural and stylistic design variations. Each chart is augmented with multi-level fine-grained labels on its semantic components, including each graphical element's type, role, and position, hierarchical groupings of elements, group layouts, and visual encodings. In total, VisAnatomy provides labels for more than 383k graphical elements. We demonstrate the richness of the semantic labels by comparing VisAnatomy with existing corpora. We illustrate its usefulness through four applications: semantic role inference for SVG elements, chart semantic decomposition, chart type classification, and content navigation for accessibility. Finally, we discuss research opportunities to further improve VisAnatomy. Chen Chen 0080, Hannah K. Bako, Peihong Yu, John Hooker, Jeffrey Joyal, Simon C. Wang, Samuel Kim, Jessica Wu, Aoxue Ding, Lara Sandeep, Alex Chen, Chayanika Sinha, Zhicheng Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 13 |
| 2025 | Comparing Native and Non-native English Speakers' Behaviors in Collaborative Writing through Visual AnalyticsabstractUnderstanding collaborative writing dynamics between native speakers (NS) and non-native speakers (NNS) is critical for enhancing collaboration quality and team inclusivity. In this paper, we partnered with communication researchers to develop visual analytics solutions for comparing NS and NNS behaviors in 162 writing sessions across 27 teams. The primary challenges in analyzing writing behaviors are data complexity and the uncertainties introduced by automated methods. In response, we present \textsc{COALA}, a novel visual analytics tool that improves model interpretability by displaying uncertainties in author clusters, generating behavior summaries using large language models, and visualizing writing-related actions at multiple granularities. We validated the effectiveness of \textsc{COALA} through user studies with domain experts (N=2+2) and researchers with relevant experience (N=8). We present the insights discovered by participants using \textsc{COALA}, suggest features for future AI-assisted collaborative writing tools, and discuss the broader implications for analyzing collaborative processes beyond writing. Yuexi Chen, Yimin Xiao, Kazi Tasnim Zinat, Naomi Yamashita, Ge Gao 0001, Zhicheng Liu 0001 |
CHI | 6 |
| 2025 | Uncovering Causal Relation Shifts in Event Sequences Under Out-of-Domain Interventions
Kazi Tasnim Zinat, Zhicheng Liu 0001 |
ICANN (3) | 5 |
| 2025 | Unveiling How Examples Shape Visualization Design OutcomesabstractVisualization designers (e.g., journalists or data analysts) often rely on examples to explore the space of possible designs, yet we have little insight into how examples shape data visualization design outcomes. While the effects of examples have been studied in other disciplines, such as web design or engineering, the results are not readily applicable to visualization due to inconsistencies in findings and challenges unique to visualization design. Towards bridging this gap, we conduct an exploratory experiment involving 32 data visualization designers focusing on the influence of five factors (timing, quantity, diversity, data topic similarity, and data schema similarity) on objectively measurable design outcomes (e.g., numbers of designs and idea transfers). Our quantitative analysis shows that when examples are introduced after initial brainstorming, designers curate examples with topics less similar to the dataset they are working on and produce more designs with a high variation in visualization components. Also, designers copy more ideas from examples with higher data schema similarities. Our qualitative analysis of participants' thought processes provides insights into why designers incorporate examples into their designs, revealing potential factors that have not been previously investigated. Finally, we discuss how our results inform how designers may use examples during design ideation as well as future research on quantifying designs and supporting example-based visualization design. All supplemental materials are available in our OSF repo. Hannah K. Bako, Xinyi Liu 0001, Grace Ko, Hyemi Song, Leilani Battle, Zhicheng Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Manipulable Semantic Components: A Computational Representation of Data Visualization ScenesabstractVarious data visualization applications such as reverse engineering and interactive authoring require a vocabulary that describes the structure of visualization scenes and the procedure to manipulate them. A few scene abstractions have been proposed, but they are restricted to specific applications for a limited set of visualization types. A unified and expressive model of data visualization scenes for different applications has been missing. To fill this gap, we present Manipulable Semantic Components (MSC), a computational representation of data visualization scenes, to support applications in scene understanding and augmentation. MSC consists of two parts: a unified object model describing the structure of a visualization scene in terms of semantic components, and a set of operations to generate and modify the scene components. We demonstrate the benefits of MSC in three applications: visualization authoring, visualization deconstruction and reuse, and animation specification. Zhicheng Liu 0001, Chen Chen 0080, John Hooker |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | A Multi-Level Task Framework for Event Sequence AnalysisabstractDespite the development of numerous visual analytics tools for event sequence data across various domains, including but not limited to healthcare, digital marketing, and user behavior analysis, comparing these domain-specific investigations and transferring the results to new datasets and problem areas remain challenging. Task abstractions can help us go beyond domain-specific details, but existing visualization task abstractions are insufficient for event sequence visual analytics because they primarily focus on multivariate datasets and often overlook automated analytical techniques. To address this gap, we propose a domain-agnostic multi-level task framework for event sequence analytics, derived from an analysis of 58 papers that present event sequence visualization systems. Our framework consists of four levels: objective, intent, strategy, and technique. Overall objectives identify the main goals of analysis. Intents comprises five high-level approaches adopted at each analysis step: augment data, simplify data, configure data, configure visualization, and manage provenance. Each intent is accomplished through a number of strategies, for instance, data simplification can be achieved through aggregation, summarization, or segmentation. Finally, each strategy can be implemented by a set of techniques depending on the input and output components. We further show that each technique can be expressed through a quartet of action-input-output-criteria. We demonstrate the framework's descriptive power through case studies and discuss its similarities and differences with previous event sequence task taxonomies. Kazi Tasnim Zinat, Saimadhav Naga Sakhamuri, Aaron Sun Chen, Zhicheng Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksabstractMixed-media tutorials, which integrate videos, images, text, and diagrams to teach procedural skills, offer more browsable alternatives than timeline-based videos. However, manually creating such tutorials is tedious, and existing automated solutions are often restricted to a particular domain. While AI models hold promise, it is unclear how to effectively harness their powers, given the multi-modal data involved and the vast landscape of models. We present TutoAI, a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasks. First, we distill common tutorial components by surveying existing work; then, we present an approach to identify, assemble, and evaluate AI models for component extraction; finally, we propose guidelines for designing user interfaces (UI) that support tutorial creation based on AI-generated components. We show that TutoAI has achieved higher or similar quality compared to a baseline model in preliminary user studies. Yuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng Liu 0001 |
CHI | 4 |
| 2024 | Evaluating the Semantic Profiling Abilities of LLMs for Natural Language Utterances in Data VisualizationabstractAutomatically generating data visualizations in response to human utterances on datasets necessitates a deep semantic understanding of the utterance, including implicit and explicit references to data attributes, visualization tasks, and necessary data preparation steps. Natural Language Interfaces (NLIs) for data visualization have explored ways to infer such information, yet challenges persist due to inherent uncertainty in human speech. Recent advances in Large Language Models (LLMs) provide an avenue to address these challenges, but their ability to extract the relevant semantic information remains unexplored. In this study, we evaluate four publicly available LLMs (GPT-4, Gemini-Pro, Llama3, and Mixtral), investigating their ability to comprehend utterances even in the presence of uncertainty and identify the relevant data context and visual tasks. Our findings reveal that LLMs are sensitive to uncertainties in utterances. Despite this sensitivity, they are able to extract the relevant data context. However, LLMs struggle with inferring visualization tasks. Based on these results, we highlight future research directions on using LLMs for visualization generation. Our supplementary materials have been shared on GitHub: https://github.com/hdi-umd/Semantic_Profiling_LLM_Evaluation. Hannah K. Bako, Arshnoor Bhutani, Xinyi Liu 0001, Kwesi A. Cobbina, Zhicheng Liu 0001 |
IEEE VIS | 5 |
| 2024 | (Dis)placed Contributions: Uncovering Hidden Hurdles to Collaborative Writing Involving Non-Native Speakers, Native Speakers, and AI-Powered Editing ToolsabstractContent creation today often takes place via collaborative writing. A longstanding interest of CSCW research lies in understanding and promoting the coordination between co-writers. However, little attention has been paid to individuals who write in their non-native language and to co-writer groups involving them. We present a mixed-method study that fills the above gap. Our participants included 32 co-writer groups, each consisting of one native speaker (NS) of English and one non-native speaker (NNS) with limited proficiency. They performed collaborative writing adopting two different workflows: half of the groups began with NNSs taking the first editing turn and half had NNSs act after NSs. Our data revealed a 'late-mover disadvantage' exclusively experienced by NNSs: an NNS's ideational contributions to the joint document were suppressed when their editing turn was placed after an NS's turn, as opposed to ahead of it. Surprisingly, editing help provided by AI-powered tools did not exempt NNSs from being disadvantaged. Instead, it triggered NSs' overestimation of NNSs' English proficiency and agency displayed in the writing, introducing unintended tensions into the collaboration. These findings shed light on the fair assessment and effective promotion of a co-writer's contributions in language diverse settings. In particular, they underscore the necessity of disentangling contributions made to the ideational, expressional, and lexical aspects of the joint writing. Yimin Xiao, Yuewen Chen, Naomi Yamashita, Yuexi Chen, Zhicheng Liu 0001, Ge Gao 0001 |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | Mystique: Deconstructing SVG Charts for Layout ReuseabstractTo facilitate the reuse of existing charts, previous research has examined how to obtain a semantic understanding of a chart by deconstructing its visual representation into reusable components, such as encodings. However, existing deconstruction approaches primarily focus on chart styles, handling only basic layouts. In this paper, we investigate how to deconstruct chart layouts, focusing on rectangle-based ones, as they cover not only 17 chart types but also advanced layouts (e.g., small multiples, nested layouts). We develop an interactive tool, called Mystique, adopting a mixed-initiative approach to extract the axes and legend, and deconstruct a chart's layout into four semantic components: mark groups, spatial relationships, data encodings, and graphical constraints. Mystique employs a wizard interface that guides chart authors through a series of steps to specify how the deconstructed components map to their own data. On 150 rectangle-based SVG charts, Mystique achieves above 85% accuracy for axis and legend extraction and 96% accuracy for layout deconstruction. In a chart reproduction study, participants could easily reuse existing charts on new datasets. We discuss the current limitations of Mystique and future research directions. Chen Chen 0080, Bongshin Lee, Yunhai Wang, Yunjeong Chang, Zhicheng Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | WhatsNext: Guidance-enriched Exploratory Data Analysis with Interactive, Low-Code NotebooksabstractComputational notebooks such as Jupyter are popular for exploratory data analysis and insight finding. Despite the module-based structure, notebooks visually appear as a single thread of interleaved cells containing text, code, visualizations, and tables, which can be unorganized and obscure users' data analysis workflow. Furthermore, users with limited coding expertise may struggle to quickly engage in the analysis process. In this work, we design and implement an interactive notebook framework, WhatsNext, with the goal of supporting low-code visual data exploration with insight-based user guidance. In particular, we (1) re-design a standard notebook cell to include a recommendation panel that suggests possible next-step exploration questions or analysis actions to take, and (2) create an interactive, dynamic tree visualization that reflects the analytic dependencies between notebook cells to make it easy for users to see the structure of the data exploration threads and trace back to previous steps. Chen Chen 0080, Jane Hoffswell, Shunan Guo, Ryan Rossi, Gromit Yeuk-Yin Chan, Fan Du, Eunyee Koh, Zhicheng Liu 0001 |
VL/HCC | 8 |
| 2023 | DocDancer: Authoring Ultra-Responsive Documents with Layout GenerationabstractResponsive design enhances user experience by adapting layout and content to different display factors. However, existing tools for authoring responsive design primarily adapt to screen width only, and they either rely on predefined templates or require significant manual effort. To expedite responsive design creation, we introduce an authoring tool called DocDancer. DocDancer supports creating ultra-responsive documents where both layout and content adapt to multiple factors (screen width, font properties, customer segments, etc), and provides layout suggestions based on user-provided content and popular responsive patterns. A comparative user study with 16 participants shows that authoring responsive documents in DocDancer takes significantly less time and effort than a commercial tool. Yuexi Chen, Zhicheng Liu 0001, Chris Tensmeyer, Niklas Elmqvist, Vlad I. Morariu |
VL/HCC | 2 |
| 2023 | The State of the Art in Creating Visualization Corpora for Automated Chart AnalysisabstractAbstract We present a state‐of‐the‐art report on visualization corpora in automated chart analysis research. We survey 56 papers that created or used a visualization corpus as the input of their research techniques or systems. Based on a multi‐level task taxonomy that identifies the goal, method, and outputs of automated chart analysis, we examine the property space of existing chart corpora along five dimensions: format, scope, collection method, annotations, and diversity. Through the survey, we summarize common patterns and practices of creating chart corpora, identify research gaps and opportunities, and discuss the desired properties of future benchmark corpora and the required tools to create them. Chen Chen 0080, Zhicheng Liu 0001 |
Comput. Graph. Forum | 2 |
| 2023 | A Comparative Evaluation of Visual Summarization Techniques for Event SequencesabstractAbstract Real‐world event sequences are often complex and heterogeneous, making it difficult to create meaningful visualizations using simple data aggregation and visual encoding techniques. Consequently, visualization researchers have developed numerous visual summarization techniques to generate concise overviews of sequential data. These techniques vary widely in terms of summary structures and contents, and currently there is a knowledge gap in understanding the effectiveness of these techniques. In this work, we present the design and results of an insight‐based crowdsourcing experiment evaluating three existing visual summarization techniques: CoreFlow, SentenTree, and Sequence Synopsis. We compare the visual summaries generated by these techniques across three tasks, on six datasets, at six levels of granularity. We analyze the effects of these variables on summary quality as rated by participants and completion time of the experiment tasks. Our analysis shows that Sequence Synopsis produces the highest‐quality visual summaries for all three tasks, but understanding Sequence Synopsis results also takes the longest time. We also find that the participants evaluate visual summary quality based on two aspects: content and interpretability. We discuss the implications of our findings on developing and evaluating new visual summarization techniques. Kazi Tasnim Zinat, Jinhua Yang, Arjun Gandhi, Nistha Mitra, Zhicheng Liu 0001 |
Comput. Graph. Forum | 5 |
| 2023 | Understanding how Designers Find and Use Data Visualization ExamplesabstractExamples are useful for inspiring ideas and facilitating implementation in visualization design. However, there is little understanding of how visualization designers use examples, and how computational tools may support such activities. In this paper, we contribute an exploratory study of current practices in incorporating visualization examples. We conducted semi-structured interviews with 15 university students and 15 professional designers. Our analysis focus on two core design activities: searching for examples and utilizing examples. We characterize observed strategies and tools for performing these activities, as well as major challenges that hinder designers' current workflows. In addition, we identify themes that cut across these two activities: criteria for determining example usefulness, curation practices, and design fixation. Given our findings, we discuss the implications for visualization design and authoring tools and highlight critical areas for future research. Hannah K. Bako, Xinyi Liu 0001, Leilani Battle, Zhicheng Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Leveraging Text-Chart Links to Support Authoring of Data-Driven Articles with VizFlowabstractData-driven articles—i.e., articles featuring text and supporting charts—play a key role in communicating information to the public. New storytelling formats like scrollytelling apply compelling dynamics to these articles to help walk readers through complex insights, but are challenging to craft. In this work, we investigate ways to support authors of data-driven articles using such storytelling forms via a text-chart linking strategy. From formative interviews with 6 authors and an assessment of 43 scrollytelling stories, we built VizFlow, a prototype system that uses text-chart links to support a range of dynamic layouts. We validate our text-chart linking approach via an authoring study with 12 participants using VizFlow, and a reading study with 24 participants comparing versions of the same article with different VizFlow intervention levels. Assessments showed our approach enabled a rapid and expressive authoring experience, and informed key design recommendations for future efforts in the space. Nicole Sultanum, Fanny Chevalier, Zoya Bylinskii, Zhicheng Liu 0001 |
CHI | 4 |
| 2021 | Data Animator: Authoring Expressive Animated Data GraphicsabstractAnimation helps viewers follow transitions in data graphics. When authoring animations that incorporate data, designers must carefully coordinate the behaviors of visual objects such as entering, exiting, merging and splitting, and specify the temporal rhythms of transition through staging and staggering. We present Data Animator, a system for authoring animated data graphics without programming. Data Animator leverages the Data Illustrator framework to analyze and match objects between two static visualizations, and generates automated transitions by default. Designers have the flexibility to interpret and adjust the matching results through a visual interface. Data Animator also supports the division of a complex animation into stages through hierarchical keyframes, and uses data attributes to stagger the start time and vary the speed of animating objects through a novel timeline interface. We validate Data Animator’s expressiveness via a gallery of examples, and evaluate its usability in a re-creation study with designers. John Thompson 0002, Zhicheng Liu 0001, John T. Stasko |
CHI | 2 |
| 2020 | Techniques for Flexible Responsive Visualization DesignabstractResponsive visualizations adapt to effectively present information based on the device context. Such adaptations are essential for news content that is increasingly consumed on mobile devices. However, existing tools provide little support for responsive visualization design. We analyze a corpus of 231 responsive news visualizations and discuss formative interviews with five journalists about responsive visualization design. These interviews motivate four central design guidelines: enable simultaneous cross-device edits, facilitate device-specific customization, show cross-device previews, and support propagation of edits. Based on these guidelines, we present a prototype system that allows users to preview and edit multiple visualization versions simultaneously. We demonstrate the utility of the system features by recreating four real-world responsive visualizations from our corpus. Jane Hoffswell, Wilmot Li, Zhicheng Liu 0001 |
CHI | 3 |
| 2020 | Data-driven Multi-level Segmentation of Image Editing LogsabstractAutomatic segmentation of logs for creativity tools such as image editing systems could improve their usability and learnability by supporting such interaction use cases as smart history navigation or recommending alternative design choices. We propose a multi-level segmentation model that works for many image editing tasks including poster creation, portrait retouching, and special effect creation. The lowest-level chunks of logged events are computed using a support vector machine model and higher-level chunks are built on top of these, at a level of granularity that can be customized for specific use cases. Our model takes into account features derived from four event attributes collected in realistically complex Photoshop sessions with expert users: command, timestamp, image content, and artwork layer. We present a detailed analysis of the relevance of each feature and evaluate the model using both quantitative performance metrics and qualitative analysis of sample sessions. Zipeng Liu, Zhicheng Liu 0001, Tamara Munzner |
CHI | 2 |
| 2020 | Understanding the Design Space and Authoring Paradigms for Animated Data GraphicsabstractAbstract Creating expressive animated data graphics often requires designers to possess highly specialized programming skills. Alternatively, the use of direct manipulation tools is popular among animation designers, but these tools have limited support for generating graphics driven by data. Our goal is to inform the design of next‐generation animated data graphic authoring tools. To understand the composition of animated data graphics, we survey real‐world examples and contribute a description of the design space. We characterize animated transitions based on object, graphic, data, and timing dimensions. We synthesize the primitives from the object, graphic, and data dimensions as a set of 10 transition types, and describe how timing primitives compose broader pacing techniques. We then conduct an ideation study that uncovers how people approach animation creation with three authoring paradigms: keyframe animation, procedural animation, and presets & templates. Our analysis shows that designers have an overall preference for keyframe animation. However, we find evidence that an authoring tool should combine these three paradigms as designers’ preferences depend on the characteristics of the animated transition design and the authoring task. Based on these findings, we contribute guidelines and design considerations for developing future animated data graphic authoring tools. John Thompson 0002, Zhicheng Liu 0001, Wilmot Li, John T. Stasko |
Comput. Graph. Forum | 2 |
| 2020 | Critical Reflections on Visualization Authoring SystemsabstractAn emerging generation of visualization authoring systems support expressive information visualization without textual programming. As they vary in their visualization models, system architectures, and user interfaces, it is challenging to directly compare these systems using traditional evaluative methods. Recognizing the value of contextualizing our decisions in the broader design space, we present critical reflections on three systems we developed -Lyra, Data Illustrator, and Charticulator. This paper surfaces knowledge that would have been daunting within the constituent papers of these three systems. We compare and contrast their (previously unmentioned) limitations and trade-offs between expressivity and learnability. We also reflect on common assumptions that we made during the development of our systems, thereby informing future research directions in visualization authoring systems. Arvind Satyanarayan, Bongshin Lee, Donghao Ren, Jeffrey Heer, John T. Stasko, John Thompson 0002, Matthew Brehmer, Zhicheng Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2019 | Visualizing Uncertainty and Alternatives in Event Sequence PredictionsabstractData analysts apply machine learning and statistical methods to timestamped event sequences to tackle various problems but face unique challenges when interpreting the results. Especially in event sequence prediction, it is difficult to convey uncertainty and possible alternative paths or outcomes. In this work, informed by interviews with five machine learning practitioners, we iteratively designed a novel visualization for exploring event sequence predictions of multiple records where users are able to review the most probable predictions and possible alternatives alongside uncertainty information. Through a controlled study with 18 participants, we found that users are more confident in making decisions when alternative predictions are displayed and they consider the alternatives more when deciding between two options with similar top predictions. Shunan Guo, Fan Du, Sana Malik, Eunyee Koh, Sungchul Kim, Zhicheng Liu 0001, Donghyun Kim 0007, Hongyuan Zha, Nan Cao 0001 |
CHI | 6 |
| 2019 | Interactive Repair of Tables Extracted from PDF Documents on Mobile DevicesabstractPDF documents often contain rich data tables that offer opportunities for dynamic reuse in new interactive applications. We describe a pipeline for extracting, analyzing, and parsing PDF tables based on existing machine learning and rule-based techniques. Implementing and deploying this pipeline on a corpus of 447 documents with 1,171 tables results in only 11 tables that are correctly extracted and parsed. To improve the results of automatic table analysis, we first present a taxonomy of errors that arise in the analysis pipeline and discuss the implications of cascading errors on the user experience. We then contribute a system with two sets of lightweight interaction techniques (gesture and toolbar), for viewing and repairing extraction errors in PDF tables on mobile devices. In an evaluation with 17 users involving both a phone and a tablet, participants effectively repaired common errors in 10 tables, with an average time of about 2 minutes per table. Jane Hoffswell, Zhicheng Liu 0001 |
CHI | 2 |
| 2019 | Trust and Recall of Information across Varying Degrees of Title-Visualization MisalignmentabstractVisualizations are emerging as a means of spreading digital misinformation. Prior work has shown that visualization interpretation can be manipulated through slanted titles that favor only one side of the visual story, yet people still think the visualization is impartial. In this work, we study whether such effects continue to exist when titles and visualizations exhibit greater degrees of misalignment: titles whose message differs from the visually cued message in the visualization, and titles whose message contradicts the visualization. We found that although titles with a contradictory slant triggered more people to identify bias compared to titles with a miscued slant, visualizations were persistently perceived as impartial by the majority. Further, people's recall of the visualization's message more frequently aligned with the titles than the visualization. Based on these results, we discuss the potential of leveraging textual components to detect and combat visual-based misinformation with text-based slants. Ha Kyung Kong, Zhicheng Liu 0001, Karrie Karahalios |
CHI | 2 |
| 2019 | Understanding Visual Cues in Visualizations Accompanied by Audio NarrationsabstractIt is often assumed that visual cues, which highlight specific parts of a visualization to guide the audience's attention, facilitate visualization storytelling and presentation. This assumption has not been systematically studied. We present an in-lab experiment and a Mechanical Turk study to examine the effects of integral and separable visual cues on the recall and comprehension of visualizations that are accompanied by audio narration. Eye-tracking data in the in-lab experiment confirm that cues helped the viewers focus on relevant parts of the visualization faster. We found that in general, visual cues did not have a significant effect on learning outcomes, but for specific cue techniques (e.g. glow) or specific chart types (e.g heatmap), cues significantly improved comprehension. Based on these results, we discuss how presenters might select visual cues depending on the role of the cues and the visualization type. Ha Kyung Kong, Zhicheng Liu 0001, Karrie Karahalios |
CHI | 3 |
| 2019 | Elastic Documents: Coupling Text and Tables through Contextual Visualizations for Enhanced Document ReadingabstractToday's data-rich documents are often complex datasets in themselves, consisting of information in different formats such as text, figures, and data tables. These additional media augment the textual narrative in the document. However, the static layout of a traditional for-print document often impedes deep understanding of its content because of the need to navigate to access content scattered throughout the text. In this paper, we seek to facilitate enhanced comprehension of such documents through a contextual visualization technique that couples text content with data tables contained in the document. We parse the text content and data tables, cross-link the components using a keyword-based matching algorithm, and generate on-demand visualizations based on the reader's current focus within a document. We evaluate this technique in a user study comparing our approach to a traditional reading experience. Results from our study show that (1) participants comprehend the content better with tighter coupling of text and data, (2) the contextual visualizations enable participants to develop better summaries that capture the main data-rich insights within the document, and (3) overall, our method enables participants to develop a more detailed understanding of the document content. Sriram Karthik Badam, Zhicheng Liu 0001, Niklas Elmqvist |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | MAQUI: Interweaving Queries and Pattern Mining for Recursive Event Sequence ExplorationabstractExploring event sequences by defining queries alone or by using mining algorithms alone is often not sufficient to support analysis. Analysts often interweave querying and mining in a recursive manner during event sequence analysis: sequences extracted as query results are used for mining patterns, patterns generated are incorporated into a new query for segmenting the sequences, and the resulting segments are mined or queried again. To support flexible analysis, we propose a framework that describes the process of interwoven querying and mining. Based on this framework, we developed MAQUI, a Mining And Querying User Interface that enables recursive event sequence exploration. To understand the efficacy of MAQUI, we conducted two case studies with domain experts. The findings suggest that the capability of interweaving querying and mining helps the participants articulate their questions and gain novel insights from their data. Po-Ming Law, Zhicheng Liu 0001, Sana Malik, Rahul C. Basole |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | Frames and Slants in Titles of Visualizations on Controversial TopicsabstractSlanted framing in news article titles induce bias and influence recall. While recent studies found that viewers focus extensively on titles when reading visualizations, the impact of titles in visualization remains underexplored. We study frames in visualization titles, and how the slanted framing of titles and the viewer's pre-existing attitude impact recall, perception of bias, and change of attitude. When asked to compose visualization titles, people used five existing news frames, an open-ended frame, and a statistics frame. We found that the slant of the title influenced the perceived main message of a visualization, with viewers deriving opposing messages from the same visualization. The results did not show any significant effect on attitude change. We highlight the danger of subtle statistics frames and viewers' unwarranted conviction of the neutrality of visualizations. Finally, we present a design implication for the generation of visualization titles and one for the viewing of titles. Ha Kyung Kong, Zhicheng Liu 0001, Karrie Karahalios |
CHI | 2 |
| 2018 | Data Illustrator: Augmenting Vector Design Tools with Lazy Data Binding for Expressive Visualization AuthoringabstractBuilding graphical user interfaces for visualization authoring is challenging as one must reconcile the tension between flexible graphics manipulation and procedural visualization generation based on a graphical grammar or declarative languages. To better support designers' workflows and practices, we propose Data Illustrator, a novel visualization framework. In our approach, all visualizations are initially vector graphics; data binding is applied when necessary and only constrains interactive manipulation to that data bound property. The framework augments graphic design tools with new concepts and operators, and describes the structure and generation of a variety of visualizations. Based on the framework, we design and implement a visualization authoring system. The system extends interaction techniques in modern vector design tools for direct manipulation of visualization configurations and parameters. We demonstrate the expressive power of our approach through a variety of examples. A qualitative study shows that designers can use our framework to compose visualizations. Zhicheng Liu 0001, John Thompson 0002, Alan Wilson 0004, Mira Dontcheva, James Delorey, Sam Grigg, Bernard Kerr, John T. Stasko |
CHI | 1 |
| 2018 | VisIRR: A Visual Analytics System for Information Retrieval and Recommendation for Large-Scale Document DataabstractIn this article, we present an interactive visual information retrieval and recommendation system, called VisIRR, for large-scale document discovery. VisIRR effectively combines the paradigms of (1) a passive pull through query processes for retrieval and (2) an active push that recommends items of potential interest to users based on their preferences. Equipped with an efficient dynamic query interface against a large-scale corpus, VisIRR organizes the retrieved documents into high-level topics and visualizes them in a 2D space, representing the relationships among the topics along with their keyword summary. In addition, based on interactive personalized preference feedback with regard to documents, VisIRR provides document recommendations from the entire corpus, which are beyond the retrieved sets. Such recommended documents are visualized in the same space as the retrieved documents, so that users can seamlessly analyze both existing and newly recommended ones. This article presents novel computational methods, which make these integrated representations and fast interactions possible for a large-scale document corpus. We illustrate how the system works by providing detailed usage scenarios. Additionally, we present preliminary user study results for evaluating the effectiveness of the system. Jaegul Choo, Hannah Kim 0001, Edward Clarkson, Zhicheng Liu 0001, Fuxin Li, Hanseung Lee, Ramakrishnan Kannan, Charles D. Stolper, John T. Stasko, Haesun Park |
ACM Trans. Knowl. Discov. Data | 4 |
| 2017 | Identifying Frequent User Tasks from Application LogsabstractIn the light of continuous growth in log analytics, application logs remain a valuable source to understand and analyze patterns in user behavior. Today, almost every major software company employs analysts to reveal user insights from log data. To understand the tasks and challenges of the analysts, we conducted a background study with a group of analysts from a major software company. A fundamental analytics objective that we recognized through this study involves identifying frequent user tasks from application logs. More specifically, analysts are interested in identifying operation groups that represent meaningful tasks performed by many users inside applications. This is challenging, primarily because of the nature of modern application logs, which are long, noisy and consist of events from high-cardinality set. In this paper, we address these challenges to design a novel frequent pattern ranking technique that extracts frequent user tasks from application logs. Our experimental study shows that our proposed technique significantly outperforms state of the art for real-world data. Himel Dev, Zhicheng Liu 0001 |
IUI | 2 |
| 2017 | Internal and External Visual Cue Preferences for Visualizations in PresentationsabstractAbstract Presenters, such as analysts briefing to an executive committee, often use visualizations to convey information. In these cases, providing clear visual guidance is important to communicate key concepts without confusion. This paper explores visual cues that guide attention to a particular area of a visualization. We developed a visual cue taxonomy distinguishing internal from external cues, designed a web tool based on the taxonomy, and conducted a user study with 24 participants to understand user preferences in choosing visual cues. Participants perceived internal cues (e.g., transparency, brightness, and magnification) as the most useful visual cues and often combined them with other internal or external cues to emphasize areas of focus for their audience. Interviews also revealed that the choice of visual cues depends on not only the chart type, but also the presentation setting, the audience, and the function cues are serving. Considering the complexity of choosing visual cues, we provide design implications for improving the organization, consistency, and integration of visual cues within existing workflows. Ha Kyung Kong, Zhicheng Liu 0001, Karrie Karahalios |
Comput. Graph. Forum | 2 |
| 2017 | CoreFlow: Extracting and Visualizing Branching Patterns from Event SequencesabstractAbstract Event sequence datasets with high event cardinality and long sequences are difficult to visualize and analyze. In particular, it is hard to generate a high level visual summary of paths and volume of flow. Existing approaches of mining and visualizing frequent sequential patterns look promising, but have limitations in terms of scalability, interpretability and utility. We propose CoreFlow, a technique that automatically extracts and visualizes branching patterns in event sequences. CoreFlow constructs a tree by recursively applying a three‐step procedure: rank events, divide sequences into groups, and trim sequences by the chosen event. The resulting tree contains key events as nodes, and links represent aggregated flows between key events. Based on CoreFlow, we have developed an interactive system for event sequence analysis. Our approach can compute branching patterns for millions of events in a few seconds, with improved interpretability of extracted patterns compared to previous work. We also present case studies of using the system in three different domains and discuss success and failure cases of applying CoreFlow to real‐world analytic problems. These case studies call forth future research on metrics and models to evaluate the quality of visual summaries of event sequences. Zhicheng Liu 0001, Bernard Kerr, Mira Dontcheva, Justin Grover, Matthew Hoffman 0001, Alan Wilson 0004 |
Comput. Graph. Forum | 1 |
| 2017 | Data-Driven Guides: Supporting Expressive Design for Information GraphicsabstractIn recent years, there is a growing need for communicating complex data in an accessible graphical form. Existing visualization creation tools support automatic visual encoding, but lack flexibility for creating custom design; on the other hand, freeform illustration tools require manual visual encoding, making the design process time-consuming and error-prone. In this paper, we present Data-Driven Guides (DDG), a technique for designing expressive information graphics in a graphic design environment. Instead of being confined by predefined templates or marks, designers can generate guides from data and use the guides to draw, place and measure custom shapes. We provide guides to encode data using three fundamental visual encoding channels: length, area, and position. Users can combine more than one guide to construct complex visual structures and map these structures to data. When underlying data is changed, we use a deformation technique to transform custom shapes using the guides as the backbone of the shapes. Our evaluation shows that data-driven guides allow users to create expressive and more accurate custom data-driven graphics. Eston Schweickart, Zhicheng Liu 0001, Mira Dontcheva, Wilmot Li, Jovan Popovic, Hanspeter Pfister |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | Patterns and Sequences: Interactive Exploration of Clickstreams to Understand Common Visitor PathsabstractModern web clickstream data consists of long, high-dimensional sequences of multivariate events, making it difficult to analyze. Following the overarching principle that the visual interface should provide information about the dataset at multiple levels of granularity and allow users to easily navigate across these levels, we identify four levels of granularity in clickstream analysis: patterns, segments, sequences and events. We present an analytic pipeline consisting of three stages: pattern mining, pattern pruning and coordinated exploration between patterns and sequences. Based on this approach, we discuss properties of maximal sequential patterns, propose methods to reduce the number of patterns and describe design considerations for visualizing the extracted sequential patterns and the corresponding raw sequences. We demonstrate the viability of our approach through an analysis scenario and discuss the strengths and limitations of the methods based on user feedback. Zhicheng Liu 0001, Mira Dontcheva, Matthew Hoffman 0001, Seth Walker, Alan Wilson 0004 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | MatrixWave: Visual Comparison of Event Sequence DataabstractEvent sequence data analysis is common in many domains, including web and software development, transportation, and medical care. Few have investigated visualization techniques for comparative analysis of multiple event sequence datasets. Grounded in the real-world characteristics of web clickstream data, we explore visualization techniques for comparison of two clickstream datasets collected on different days or from users with different demographics. Through iterative design with web analysts, we designed MatrixWave, a matrix-based representation that allows analysts to get an overview of differences in traffic patterns and interactively explore paths through the website. We use color to encode differences and size to offer context over traffic volume. User feedback on MatrixWave is positive. Our study participants made fewer errors with MatrixWave and preferred it over the more familiar Sankey diagram. Jian Zhao 0010, Zhicheng Liu 0001, Mira Dontcheva, Aaron Hertzmann, Alan Wilson 0004 |
CHI | 2 |
| 2015 | Learning style similarity for searching infographics
Babak Saleh, Mira Dontcheva, Aaron Hertzmann, Zhicheng Liu 0001 |
Graphics Interface | 4 |
| 2015 | DataTone: Managing Ambiguity in Natural Language Interfaces for Data VisualizationabstractAnswering questions with data is a difficult and time-consuming process. Visual dashboards and templates make it easy to get started, but asking more sophisticated questions often requires learning a tool designed for expert analysts. Natural language interaction allows users to ask questions directly in complex programs without having to learn how to use an interface. However, natural language is often ambiguous. In this work we propose a mixed-initiative approach to managing ambiguity in natural language interfaces for data visualization. We model ambiguity throughout the process of turning a natural language query into a visualization and use algorithmic disambiguation coupled with interactive ambiguity widgets. These widgets allow the user to resolve ambiguities by surfacing system decisions at the point where the ambiguity matters. Corrections are stored as constraints and influence subsequent queries. We have implemented these ideas in a system, DataTone. In a comparative study, we find that DataTone is easy to learn and lets users ask questions without worrying about syntax and proper question form. Mira Dontcheva, Eytan Adar, Zhicheng Liu 0001, Karrie Karahalios |
UIST | 4 |
| 2014 | sTrack: Secure Tracking in Community SurveillanceabstractWe present sTrack, a system that can track objects across multiple cameras without sharing any visual information between two cameras except whether an object was seen by both. To achieve this challenging privacy goal, we leverage recent advances in secure two-party computation and multi-camera tracking. We derive a new distance metric learning technique that is more suited for secure computation. Compared to the existing methods, our technique has lower complexity in secure computation without sacrificing the tracking accuracy. We implement it using a new Boolean circuit for secure tracking. Experiments using real datasets show that the performance overhead of secure tracking is low, adding only a few seconds over non-private tracking. Chun-Te Chu, Jaeyeon Jung, Zhicheng Liu 0001, Ratul Mahajan |
ACM Multimedia | 3 |
| 2014 | The Effects of Interactive Latency on Exploratory Visual AnalysisabstractTo support effective exploration, it is often stated that interactive visualizations should provide rapid response times. However, the effects of interactive latency on the process and outcomes of exploratory visual analysis have not been systematically studied. We present an experiment measuring user behavior and knowledge discovery with interactive visualizations under varying latency conditions. We observe that an additional delay of 500 ms incurs significant costs, decreasing user activity and data set coverage. Analyzing verbal data from think-aloud protocols, we find that increased latency reduces the rate at which users make observations, draw generalizations and generate hypotheses. Moreover, we note interaction effects in which initial exposure to higher latencies leads to subsequently reduced performance in a low-latency setting. Overall, increased latency causes users to shift exploration strategy, in turn affecting performance. We discuss how these results can inform the design of interactive analysis tools. Zhicheng Liu 0001, Jeffrey Heer |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2013 | imMens: Real-time Visual Querying of Big DataabstractAbstract Data analysts must make sense of increasingly large data sets, sometimes with billions or more records. We present methods for interactive visualization of big data, following the principle that perceptual and interactive scalability should be limited by the chosen resolution of the visualized data, not the number of records. We first describe a design space of scalable visual summaries that use data reduction methods (such as binned aggregation or sampling) to visualize a variety of data types. We then contribute methods for interactive querying (e.g., brushing & linking) among binned plots through a combination of multivariate data tiles and parallel query processing. We implement our techniques in imMens, a browser‐based visual analysis system that uses WebGL for data processing and rendering on the GPU. In benchmarks imMens sustains 50 frames‐per‐second brushing & linking among dozens of visualizations, with invariant performance on data sizes ranging from thousands to billions of records. Zhicheng Liu 0001, Biye Jiang, Jeffrey Heer |
Comput. Graph. Forum | 1 |
| 2013 | Combining Computational Analyses and Interactive Visualization for Document Exploration and Sensemaking in JigsawabstractInvestigators across many disciplines and organizations must sift through large collections of text documents to understand and piece together information. Whether they are fighting crime, curing diseases, deciding what car to buy, or researching a new field, inevitably investigators will encounter text documents. Taking a visual analytics approach, we integrate multiple text analysis algorithms with a suite of interactive visualizations to provide a flexible and powerful environment that allows analysts to explore collections of documents while sensemaking. Our particular focus is on the process of integrating automated analyses with interactive visualizations in a smooth and fluid manner. We illustrate this integration through two example scenarios: an academic researcher examining InfoVis and VAST conference papers and a consumer exploring car reviews while pondering a purchase decision. Finally, we provide lessons learned toward the design and implementation of visual analytics systems for document exploration and understanding. Carsten Görg, Zhicheng Liu 0001, Jaeyeon Kihm, Jaegul Choo, Haesun Park, John T. Stasko |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | Mental Models, Visual Reasoning and Interaction in Information Visualization: A Top-down PerspectiveabstractAlthough previous research has suggested that examining the interplay between internal and external representations can benefit our understanding of the role of information visualization (InfoVis) in human cognitive activities, there has been little work detailing the nature of internal representations, the relationship between internal and external representations and how interaction is related to these representations. In this paper, we identify and illustrate a specific kind of internal representation, mental models, and outline the high-level relationships between mental models and external visualizations. We present a top-down perspective of reasoning as model construction and simulation, and discuss the role of visualization in model based reasoning. From this perspective, interaction can be understood as active modeling for three primary purposes: external anchoring, information foraging, and cognitive offloading. Finally we discuss the implications of our approach for design, evaluation and theory development. Zhicheng Liu 0001, John T. Stasko |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2009 | SellTrend: Inter-Attribute Visual Analysis of Temporal Transaction DataabstractWe present a case study of our experience designing SellTrend, a visualization system for analyzing airline travel purchase requests. The relevant transaction data can be characterized as multi-variate temporal and categorical event sequences, and the chief problem addressed is how to help company analysts identify complex combinations of transaction attributes that contribute to failed purchase requests. SellTrend combines a diverse set of techniques ranging from time series visualization to faceted browsing and historical trend analysis in order to help analysts make sense of the data. We believe that the combination of views and interaction capabilities in SellTrend provides an innovative approach to this problem and to other similar types of multivariate, temporally driven transaction data analysis. Initial feedback from company analysts confirms the utility and benefits of the system. Zhicheng Liu 0001, John T. Stasko, Timothy Sullivan |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2008 | Distributed Cognition as a Theoretical Framework for Information VisualizationabstractEven though information visualization (InfoVis) research has matured in recent years, it is generally acknowledged that the field still lacks supporting, encompassing theories. In this paper, we argue that the distributed cognition framework can be used to substantiate the theoretical foundation of InfoVis. We highlight fundamental assumptions and theoretical constructs of the distributed cognition approach, based on the cognitive science literature and a real life scenario. We then discuss how the distributed cognition framework can have an impact on the research directions and methodologies we take as InfoVis researchers. Our contributions are as follows. First, we highlight the view that cognition is more an emergent property of interaction than a property of the human mind. Second, we argue that a reductionist approach to study the abstract properties of isolated human minds may not be useful in informing InfoVis design. Finally we propose to make cognition an explicit research agenda, and discuss the implications on how we perform evaluation and theory building. Zhicheng Liu 0001, Nancy J. Nersessian, John T. Stasko |
IEEE Trans. Vis. Comput. Graph. | 1 |