VLDB 2026 Research / reviewers in the wild / expert
Shunan Guo
dblp:210/5384
· DBLP profile ↗
31ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0001-5355-8399ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Narrix: Remixing Narrative Strategies from Examples for Story Writing
Chao Zhang 0082, Shunan Guo, Abe Davis, Eunyee Koh |
CHI | 2 |
| 2026 | SlideSAVR: Enabling Live Analysis during Data Presentations via Multimodal Sketching and Voice InputabstractAbstract Interpersonal communication in data science can yield sought‐after insights, but presentation environments are often not conducive for live analysis, forcing the process to move offline. Through a formative survey with 16 participants, we identified both technical (e.g., complexity of tools) and psychological (e.g., pressure of programming during presentation) factors constraining live data analysis. To enable live analysis, we present SlideSAVR, a data‐driven presentation assistant that leverages sketching and voice inputs in live discussion to support collaborative data analysis during presentations. Powered by an agentic framework that flexibly defines augmentation rules, updates slide content dynamically to match the live context, and automates backend computations, SlideSAVR enables fluid audience‐presenter interaction and reduces the need for offline reanalysis and follow‐up communication. We demonstrate SlideSAVR's ability to support a range of tasks through nine representative use cases. We further evaluate the system's accuracy and computation time across different settings, showing that SlideSAVR can reliably perform diverse tasks when provided with both sketch and voice inputs. Chang Han, Md Mehrab Tanjim, Shunan Guo, Christine Dierk, Katherine E. Isaacs, Jane Hoffswell |
Comput. Graph. Forum | 3 |
| 2025 | OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language ModelsabstractAs multi-turn dialogues with large language models (LLMs) grow longer and more complex, how can users better evaluate and review progress on their conversational goals?We present OnGoal, an LLM chat interface that helps users better manage goal progress.OnGoal provides real-time feedback on goal alignment through LLM-assisted evaluation, explanations for evaluation results with examples, and overviews of goal progression over time, enabling users to navigate complex dialogues more effectively.Through a Adam Coscia, Shunan Guo, Eunyee Koh, Alex Endert |
UIST | 2 |
| 2025 | ConvoMap: Interactive Visualizations for Exploring Complex Conversations in Multi-Agent Systemsabstract—Following the rapid emergence of large language models, Multi-Agent Systems (MASs) became a promising approach for accomplishing complex tasks. In MASs, multiple autonomous agents with predetermined roles collaborate by dividing responsibilities. However, MAS developers often struggle to understand and diagnose agents’ behavior from thousands of inter-agent messages across multiple complex conversations. To identify key requirements and challenges related to evaluating, debugging, and managing MASs, we conducted a formative study with six MAS developers. We then introduce ConvoMap, a prototype that addresses a key challenge of MAS development-understanding agents’ behaviors across multiple conversations. ConvoMap integrates automated qualitative coding to enable multi-level inspection of agents’ behavior. ConvoMap can then visualize hundreds of MAS conversations by representing messages as points on a 2D map that encode their semantic meanings and interactions between agents. To better support navigation and deeper analysis, ConvoMap provides topic overviews and highlights relevant text segments. A comparison study showed that ConvoMap helped to understand agents’ behavior more accurately than the baseline. Ashley Ge Zhang, Victor S. Bursztyn, Gromit Yeuk-Yin Chan, Shunan Guo, Eunyee Koh, Steve Oney, Jane Hoffswell |
VL/HCC | 4 |
| 2025 | Utilizing Provenance as an Attribute for Visual Data Analysis: A Design Probe With ProvenanceLensabstractAnalytic provenance can be visually encoded to help users track their ongoing analysis trajectories, recall past interactions, and inform new analytic directions. Despite its significance, provenance is often hardwired into analytics systems, affording limited user control and opportunities for self-reflection. We thus propose modeling provenance as an attribute that is available to users during analysis. We demonstrate this concept by modeling two provenance attributes that track the recency and frequency of user interactions with data. We integrate these attributes into a visual data analysis system prototype, ProvenanceLens, wherein users can visualize their interaction recency and frequency by mapping them to encoding channels (e.g., color, size) or applying data transformations (e.g., filter, sort). Using ProvenanceLens as a design probe, we conduct an exploratory study with sixteen users to investigate how these provenance-tracking affordances are utilized for both decision-making and self-reflection. We find that users can accurately and confidently answer questions about their analysis, and we show that mismatches between the user's mental model and the provenance encodings can be surprising, thereby prompting useful self-reflection. We also report on the user strategies surrounding these affordances, and reflect on their intuitiveness and effectiveness in representing provenance. Arpit Narechania, Shunan Guo, Eunyee Koh, Alex Endert, Jane Hoffswell |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Litforager: Exploring Multimodal Literature Foraging Strategies in Immersive SensemakingabstractExploring and comprehending relevant academic literature is a vital yet challenging task for researchers, especially given the rapid expansion in research publications. This task fundamentally involves sensemaking-interpreting complex, scattered information sources to build understanding. While emerging immersive analytics tools have shown cognitive benefits like enhanced spatial memory and reduced mental load, they predominantly focus on information synthesis (e.g., organizing known documents). In contrast, the equally important information foraging phase-discovering and gathering relevant literature-remains underexplored within immersive environments, hindering a complete sensemaking workflow. To bridge this gap, we introduce LitForager, an interactive literature exploration tool designed to facilitate information foraging of research literature within an immersive sensemaking workflow using network-based visualizations and multimodal interactions. Developed with WebXR and informed by a formative study with researchers, LitForager supports exploration guidance, spatial organization, and seamless transition through a 3D literature network. An observational user study with 15 researchers demonstrated LitForager's effectiveness in supporting fluid foraging strategies and spatial sensemaking through its multimodal interface. Haoyang Yang, Elliott H. Faa, Shunan Guo, Polo Chau, Yalong Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | CompositingVis: Exploring Interactions for Creating Composite Visualizations in Immersive EnvironmentsabstractComposite visualization represents a widely embraced design that combines multiple visual representations to create an integrated view. However, the traditional approach of creating composite visualizations in immersive environments typically occurs asynchronously outside of the immersive space and is carried out by experienced experts. In this work, we aim to empower users to participate in the creation of composite visualization within immersive environments through embodied interactions. This could provide a flexible and fluid experience with immersive visualization and has the potential to facilitate understanding of the relationship between visualization views. We begin with developing a design space of embodied interactions to create various types of composite visualizations with the consideration of data relationships. Drawing inspiration from people's natural experience of manipulating physical objects, we design interactions based on the combination of 3D manipulations in immersive environments. Building upon the design space, we present a series of case studies showcasing the interaction to create different kinds of composite visualizations in virtual reality. Subsequently, we conduct a user study to evaluate the usability of the derived interaction techniques and user experience of creating composite visualizations through embodied interactions. We find that empowering users to participate in composite visualizations through embodied interactions enables them to flexibly leverage different visualization views for understanding and communicating the relationships between different views, which underscores the potential of several future application scenarios. Qian Zhu 0010, Tao Lu 0013, Shunan Guo, Xiaojuan Ma, Yalong Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | "The Data Says Otherwise" - Towards Automated Fact-checking and Communication of Data ClaimsabstractFact-checking data claims requires data evidence retrieval and analysis, which can become tedious and intractable when done manually. This work presents Aletheia, an automated fact-checking prototype designed to facilitate data claims verification and enhance data evidence communication. For verification, we utilize a pre-trained LLM to parse the semantics for evidence retrieval. To effectively communicate the data evidence, we design representations in two forms: data tables and visualizations, tailored to various data fact types. Additionally, we design interactions that showcase a real-world application of these techniques. We evaluate the performance of two core NLP tasks with a curated dataset comprising 400 data claims and compare the two representation forms regarding viewers’ assessment time, confidence, and preference via a user study with 20 participants. The evaluation offers insights into the feasibility and bottlenecks of using LLMs for data fact-checking tasks, potential advantages and disadvantages of using visualizations over data tables, and design recommendations for presenting data evidence. Yu Fu 0010, Shunan Guo, Jane Hoffswell, Victor S. Bursztyn, Ryan Rossi, John T. Stasko |
UIST | 2 |
| 2024 | Representing Charts as Text for Language Models: An In-Depth Study of Question Answering for Bar ChartsabstractMachine Learning models for chart-grounded Q&A (CQA) often treat charts as images, but performing CQA on pixel values has proven challenging. We thus investigate a resource overlooked by current ML-based approaches: the declarative documents describing how charts should visually encode data (i.e., chart specifications). In this work, we use chart specifications to enhance language models (LMs) for chart-reading tasks, such that the resulting system can robustly understand language for CQA. Through a case study with 359 bar charts, we test novel fine tuning schemes on both GPT-3 and T5 using a new dataset curated for two CQA tasks: question-answering and visual explanation generation. Our text-only approaches strongly outperform vision-based GPT-4 on explanation generation (99% vs. 63% accuracy), and show promising results for question-answering (57–67% accuracy). Through in-depth experiments, we also show that our text-only approaches are mostly robust to natural language variation. Victor S. Bursztyn, Jane Hoffswell, Eunyee Koh, Shunan Guo |
IEEE VIS | 4 |
| 2024 | Socrates: Data Story Generation via Adaptive Machine-Guided Elicitation of User FeedbackabstractVisual data stories can effectively convey insights from data, yet their creation often necessitates intricate data exploration, insight discovery, narrative organization, and customization to meet the communication objectives of the storyteller. Existing automated data storytelling techniques, however, tend to overlook the importance of user customization during the data story authoring process, limiting the system's ability to create tailored narratives that reflect the user's intentions. We present a novel data story generation workflow that leverages adaptive machine-guided elicitation of user feedback to customize the story. Our approach employs an adaptive plug-in module for existing story generation systems, which incorporates user feedback through interactive questioning based on the conversation history and dataset. This adaptability refines the system's understanding of the user's intentions, ensuring the final narrative aligns with their goals. We demonstrate the feasibility of our approach through the implementation of an interactive prototype: Socrates. Through a quantitative user study with 18 participants that compares our method to a state-of-the-art data story generation algorithm, we show that Socrates produces more relevant stories with a larger overlap of insights compared to human-generated stories. We also demonstrate the usability of Socrates via interviews with three data analysts and highlight areas of future work. Guande Wu, Shunan Guo, Jane Hoffswell, Gromit Yeuk-Yin Chan, Ryan Rossi, Eunyee Koh |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | DataCockpit: A Toolkit for Data Lake Navigation and Monitoring Utilizing Quality and Usage InformationabstractModern organizations amass their datasets into centralized repositories called data lakes, affording analytics as needed. The resultant scale and complexity of these data lakes, however, can make data navigation and monitoring challenging for users. We present DataCockpit, a Python toolkit that leverages datasets, usage logs, and associated meta-data to provision data usage and quality characteristics. DataCockpit computes these characteristics for each attribute (e.g., number of times it was queried for subsequent use in downstream applications) and record (e.g., number of non-missing, valid values) and aggregates them at the level of datasets. We develop a visual monitoring tool, powered by DataCockpit, and demonstrate how it can assist data / system administrators as well as end-users to effectively navigate and monitor a data lake. DataCockpit and the monitoring tool are available as open source software for developers to build custom monitoring applications on top of data lakes. Arpit Narechania, Surya Chakraborty, Shivam Agarwal, Atanu R. Sinha, Ryan Rossi, Fan Du, Jane Hoffswell, Shunan Guo, Eunyee Koh, Alex Endert, Shamkant B. Navathe |
IEEE Big Data | 8 |
| 2023 | On Chatbots for Visual Exploratory Data AnalysisabstractAnalyzing data and creating effective visualizations often requires extensive domain expertise. For users with less experience, it can be difficult to know how to get started with exploratory data analysis (EDA) and how to approach the code. Chatbots can reduce the gap between analysis outcomes and user expectations by leveraging multi-turn conversations to provide a more natural interface between the user and computer-agent. To inform the design of future visual EDA chatbots, we conduct a survey and interview study with ten potential users. Our results suggest that users want a visual EDA chatbot that can make exploratory data analysis easier, while also augmenting their knowledge of visualization and analysis techniques. Between the initial survey and post-interview questionnaire, we saw increased optimism overall for the usefulness and anticipated analytic ease of visual EDA chatbots. Based on these results, we identify four key design guidelines: future visual EDA chatbots should (1) understand the user’s data and intent, (2) respond with useful visualizations, (3) leverage the history of the visualizations and data, and (4) produce verifiable and shareable analysis processes. Brodrick Stigall, Ryan Rossi, Jane Hoffswell, Xiang Chen 0010, Shunan Guo, Fan Du, Eunyee Koh, Kelly Caine |
IEEE Big Data | 5 |
| 2023 | DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data PreparationabstractSelecting relevant data subsets from large, unfamiliar datasets can be difficult. We address this challenge by modeling and visualizing two kinds of auxiliary information: (1) quality – the validity and appropriateness of data required to perform certain analytical tasks; and (2) usage – the historical utilization characteristics of data across multiple users. Through a design study with 14 data workers, we integrate this information into a visual data preparation and analysis tool, DataPilot. DataPilot presents visual cues about “the good, the bad, and the ugly” aspects of data and provides graphical user interface controls as interaction affordances, guiding users to perform subset selection. Through a study with 36 participants, we investigate how DataPilot helps users navigate a large, unfamiliar tabular dataset, prepare a relevant subset, and build a visualization dashboard. We find that users selected smaller, effective subsets with higher quality and usage, and with greater success and confidence. Arpit Narechania, Fan Du, Atanu R. Sinha, Ryan Rossi, Jane Hoffswell, Shunan Guo, Eunyee Koh, Shamkant B. Navathe, Alex Endert |
CHI | 6 |
| 2023 | De-Stijl: Facilitating Graphics Design with Interactive 2D Color Palette RecommendationabstractSelecting a proper color palette is critical in crafting a high-quality graphic design to gain visibility and communicate ideas effectively. To facilitate this process, we propose De-Stijl, an intelligent and interactive color authoring tool to assist novice designers in crafting harmonic color palettes, achieving quick design iterations, and fulfilling design constraints. Through De-Stijl, we contribute a novel 2D color palette concept that allows users to intuitively perceive color designs in context with their proportions and proximities. Further, De-Stijl implements a holistic color authoring system that supports 2D palette extraction, theme-aware and spatial-sensitive color recommendation, and automatic graphical elements (re)colorization. We evaluated De-Stijl through an in-lab user study by comparing the system with existing industry standard tools, followed by in-depth user interviews. Quantitative and qualitative results demonstrate that De-Stijl is effective in assisting novice design practitioners to quickly colorize graphic designs and easily deliver several alternatives. Xinyu Shi 0002, Ziqi Zhou 0003, Jing Wen Zhang, Ali Neshati, Anjul Kumar Tyagi, Ryan Rossi, Shunan Guo, Fan Du, Jian Zhao 0010 |
CHI | 7 |
| 2023 | Direct Embedding of Temporal Network Edges via Time-Decayed Line Graphs
Sudhanshu Chanpuriya, Ryan Rossi, Sungchul Kim, Tong Yu 0001, Jane Hoffswell, Nedim Lipka, Shunan Guo, Cameron Musco |
ICLR | 7 |
| 2023 | WhatsNext: Guidance-enriched Exploratory Data Analysis with Interactive, Low-Code NotebooksabstractComputational notebooks such as Jupyter are popular for exploratory data analysis and insight finding. Despite the module-based structure, notebooks visually appear as a single thread of interleaved cells containing text, code, visualizations, and tables, which can be unorganized and obscure users' data analysis workflow. Furthermore, users with limited coding expertise may struggle to quickly engage in the analysis process. In this work, we design and implement an interactive notebook framework, WhatsNext, with the goal of supporting low-code visual data exploration with insight-based user guidance. In particular, we (1) re-design a standard notebook cell to include a recommendation panel that suggests possible next-step exploration questions or analysis actions to take, and (2) create an interactive, dynamic tree visualization that reflects the analytic dependencies between notebook cells to make it easy for users to see the structure of the data exploration threads and trace back to previous steps. Chen Chen 0080, Jane Hoffswell, Shunan Guo, Ryan Rossi, Gromit Yeuk-Yin Chan, Fan Du, Eunyee Koh, Zhicheng Liu 0001 |
VL/HCC | 3 |
| 2023 | Evaluating the Use of Uncertainty Visualisations for Imputations of Data Missing At Random in ScatterplotsabstractMost real-world datasets contain missing values yet most exploratory data analysis (EDA) systems only support visualising data points with complete cases. This omission may potentially lead the user to biased analyses and insights. Imputation techniques can help estimate the value of a missing data point, but introduces additional uncertainty. In this work, we investigate the effects of visualising imputed values in charts using different ways of representing data imputations and imputation uncertainty-no imputation, mean, 95% confidence intervals, probability density plots, gradient intervals, and hypothetical outcome plots. We focus on scatterplots, which is a commonly used chart type, and conduct a crowdsourced study with 202 participants. We measure users' bias and precision in performing two tasks-estimating average and detecting trend-and their self-reported confidence in performing these tasks. Our results suggest that, when estimating averages, uncertainty representations may reduce bias but at the cost of decreasing precision. When estimating trend, only hypothetical outcome plots may lead to a small probability of reducing bias while increasing precision. Participants in every uncertainty representation were less certain about their response when compared to the baseline. The findings point towards potential trade-offs in using uncertainty encodings for datasets with a large number of missing values. This paper and the associated analysis materials are available at: https://osf.io/q4y5r/. Abhraneel Sarma, Shunan Guo, Jane Hoffswell, Ryan Rossi, Fan Du, Eunyee Koh, Matthew Kay 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Cicero: A Declarative Grammar for Responsive VisualizationabstractDesigning responsive visualizations can be cast as applying transformations to a source view to render it suitable for a different screen size. However, designing responsive visualizations is often tedious as authors must manually apply and reason about candidate transformations. We present Cicero, a declarative grammar for concisely specifying responsive visualization transformations which paves the way for more intelligent responsive visualization authoring tools. Cicero’s flexible specifier syntax allows authors to select visualization elements to transform, independent of the source view’s structure. Cicero encodes a concise set of actions to encode a diverse set of transformations in both desktop-first and mobile-first design processes. Authors can ultimately reuse design-agnostic transformations across different visualizations. To demonstrate the utility of Cicero, we develop a compiler to an extended version of Vega-Lite, and provide principles for our compiler. We further discuss the incorporation of Cicero into responsive visualization authoring tools, such as a design recommender. Hyeok Kim, Ryan Rossi, Fan Du, Eunyee Koh, Shunan Guo, Jessica Hullman, Jane Hoffswell |
CHI | 5 |
| 2022 | VisGNN: Personalized Visualization Recommendationvia Graph Neural NetworksabstractIn this work, we develop a Graph Neural Network (GNN) framework for the problem of personalized visualization recommendation. The GNN-based framework first represents the large corpus of datasets and visualizations from users as a large heterogeneous graph. Then, it decomposes a visualization into its data and visual components, and then jointly models each of them as a large graph to obtain embeddings of the users, attributes (across all datasets in the corpus), and visual-configurations. From these user-specific embeddings of the attributes and visual-configurations, we can predict the probability of any visualization arising from a specific user. Finally, the experiments demonstrated the effectiveness of using graph neural networks for automatic and personalized recommendation of visualizations to specific users based on their data and visual (design choice) preferences. To the best of our knowledge, this is the first such work to develop and leverage GNNs for this problem. Fayokemi Ojo, Ryan Rossi, Jane Hoffswell, Shunan Guo, Fan Du, Sungchul Kim, Chang Xiao 0001, Eunyee Koh |
WWW | 4 |
| 2022 | Survey on Visual Analysis of Event Sequence DataabstractEvent sequence data record series of discrete events in the time order of occurrence. They are commonly observed in a variety of applications ranging from electronic health records to network logs, with the characteristics of large-scale, high-dimensional and heterogeneous. This high complexity of event sequence data makes it difficult for analysts to manually explore and find patterns, resulting in ever-increasing needs for computational and perceptual aids from visual analytics techniques to extract and communicate insights from event sequence datasets. In this paper, we review the state-of-the-art visual analytics approaches, characterize them with our proposed design space, and categorize them based on analytical tasks and applications. From our review of relevant literature, we have also identified several remaining research challenges and future research opportunities. Shunan Guo, Zhuochen Jin, Smiti Kaul, David Gotz, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Interpretable Anomaly Detection in Event Sequences via Sequence Matching and Visual ComparisonabstractAnomaly detection is a common analytical task that aims to identify rare cases that differ from the typical cases that make up the majority of a dataset. When analyzing event sequence data, the task of anomaly detection can be complex because the sequential and temporal nature of such data results in diverse definitions and flexible forms of anomalies. This, in turn, increases the difficulty in interpreting detected anomalies. In this article, we propose a visual analytic approach for detecting anomalous sequences in an event sequence dataset via an unsupervised anomaly detection algorithm based on Variational AutoEncoders. We further compare the anomalous sequences with their reconstructions and with the normal sequences through a sequence matching algorithm to identify event anomalies. A visual analytics system is developed to support interactive exploration and interpretations of anomalies through novel visualization designs that facilitate the comparison between anomalous sequences and normal sequences. Finally, we quantitatively evaluate the performance of our anomaly detection algorithm, demonstrate the effectiveness of our system through case studies, and report feedback collected from study participants. Shunan Guo, Zhuochen Jin, Qing Chen 0001, David Gotz, Hongyuan Zha, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | A Design Space for Applying the Freytag's Pyramid Structure to Data StoriesabstractData stories integrate compelling visual content to communicate data insights in the form of narratives. The narrative structure of a data story serves as the backbone that determines its expressiveness, and it can largely influence how audiences perceive the insights. Freytag's Pyramid is a classic narrative structure that has been widely used in film and literature. While there are continuous recommendations and discussions about applying Freytag's Pyramid to data stories, little systematic and practical guidance is available on how to use Freytag's Pyramid for creating structured data stories. To bridge this gap, we examined how existing practices apply Freytag's Pyramid by analyzing stories extracted from 103 data videos. Based on our findings, we proposed a design space of narrative patterns, data flows, and visual communications to provide practical guidance on achieving narrative intents, organizing data facts, and selecting visual design techniques through story creation. We evaluated the proposed design space through a workshop with 25 participants. Results show that our design space provides a clear framework for rapid storyboarding of data stories with Freytag's Pyramid. Leni Yang, Xingyu Lan, Shunan Guo, Yang Shi 0007, Huamin Qu, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Vinci: An Intelligent Graphic Design System for Generating Advertising PostersabstractAdvertising posters are a commonly used form of information presentation to promote a product. Producing advertising posters often takes much time and effort of designers when confronted with abundant choices of design elements and layouts. This paper presents Vinci, an intelligent system that supports the automatic generation of advertising posters. Given the user-specified product image and taglines, Vinci uses a deep generative model to match the product image with a set of design elements and layouts for generating an aesthetic poster. The system also integrates online editing-feedback that supports users in editing the posters and updating the generated results with their design preference. Through a series of user studies and a Turing test, we found that Vinci can generate posters as good as human designers and that the online editing-feedback improves the efficiency in poster modification. Shunan Guo, Zhuochen Jin, Fuling Sun, Zhaorui Li, Yang Shi 0007, Nan Cao 0001 |
CHI | 1 |
| 2021 | Visual Causality Analysis of Event Sequence DataabstractCausality is crucial to understanding the mechanisms behind complex systems and making decisions that lead to intended outcomes. Event sequence data is widely collected from many real-world processes, such as electronic health records, web clickstreams, and financial transactions, which transmit a great deal of information reflecting the causal relations among event types. Unfortunately, recovering causalities from observational event sequences is challenging, as the heterogeneous and high-dimensional event variables are often connected to rather complex underlying event excitation mechanisms that are hard to infer from limited observations. Many existing automated causal analysis techniques suffer from poor explainability and fail to include an adequate amount of human knowledge. In this paper, we introduce a visual analytics method for recovering causalities in event sequence data. We extend the Granger causality analysis algorithm on Hawkes processes to incorporate user feedback into causal model refinement. The visualization system includes an interactive causal analysis framework that supports bottom-up causal exploration, iterative causal verification and refinement, and causal comparison through a set of novel visualizations and interactions. We report two forms of evaluation: a quantitative evaluation of the model improvements resulting from the user-feedback mechanism, and a qualitative evaluation through case studies in different application domains to demonstrate the usefulness of the system. Zhuochen Jin, Shunan Guo, Daniel Weiskopf, David Gotz, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | CarePre: An Intelligent Clinical Decision Assistance SystemabstractClinical decision support systems are widely used to assist with medical decision making. However, clinical decision support systems typically require manually curated rules and other data that are difficult to maintain and keep up to date. Recent systems leverage advanced deep learning techniques and electronic health records to provide a more timely and precise result. Many of these techniques have been developed with a common focus on predicting upcoming medical events. However, although the prediction results from these approaches are promising, their value is limited by their lack of interpretability. To address this challenge, we introduce CarePre, an intelligent clinical decision assistance system. The system extends a state-of-the-art deep learning model to predict upcoming diagnosis events for a focal patient based on his or her historical medical records. The system includes an interactive framework together with intuitive visualizations designed to support diagnosis, treatment outcome analysis, and the interpretation of the analysis results. We demonstrate the effectiveness and usefulness of the CarePre system by reporting results from a quantities evaluation of the prediction algorithm, two case studies, and interviews with senior physicians and pulmonologists. Zhuochen Jin, Shuyuan Cui, Shunan Guo, David Gotz, Jimeng Sun 0001, Nan Cao 0001 |
ACM Trans. Comput. Heal. | 3 |
| 2019 | Visual Anomaly Detection in Event Sequence DataabstractAnomaly detection is a common analytical task that aims to identify rare cases that differ from the typical cases that make up the majority of a dataset. When applied to the analysis of event sequence data, the task of anomaly detection can be complex because the sequential and temporal nature of such data results in diverse definitions and flexible forms of anomalies. This, in turn, increases the difficulty in interpreting detected anomalies. In this paper, we propose an unsupervised anomaly detection algorithm based on Variational AutoEncoders (VAE) to estimate underlying normal progressions for each given sequence represented as occurrence probabilities of events along the sequence progression. Events in violation of their occurrence probability are identified as abnormal. We also introduce a visualization system, EventThread3 (ET3, to support interactive exploration and interpretations of anomalies within the context of normal sequence progressions in the dataset through comprehensive one-to-many sequence comparison. Finally, we quantitatively evaluate the performance of our anomaly detection algorithm and demonstrate the effectiveness of our system through a case study. Shunan Guo, Zhuochen Jin, Qing Chen 0001, David Gotz, Hongyuan Zha, Nan Cao 0001 |
IEEE BigData | 1 |
| 2019 | Visualizing Uncertainty and Alternatives in Event Sequence PredictionsabstractData analysts apply machine learning and statistical methods to timestamped event sequences to tackle various problems but face unique challenges when interpreting the results. Especially in event sequence prediction, it is difficult to convey uncertainty and possible alternative paths or outcomes. In this work, informed by interviews with five machine learning practitioners, we iteratively designed a novel visualization for exploring event sequence predictions of multiple records where users are able to review the most probable predictions and possible alternatives alongside uncertainty information. Through a controlled study with 18 participants, we found that users are more confident in making decisions when alternative predictions are displayed and they consider the alternatives more when deciding between two options with similar top predictions. Shunan Guo, Fan Du, Sana Malik, Eunyee Koh, Sungchul Kim, Zhicheng Liu 0001, Donghyun Kim 0007, Hongyuan Zha, Nan Cao 0001 |
CHI | 1 |
| 2019 | Visual Progression Analysis of Event Sequence DataabstractEvent sequence data is common to a broad range of application domains, from security to health care to scholarly communication. This form of data captures information about the progression of events for an individual entity (e.g., a computer network device; a patient; an author) in the form of a series of time-stamped observations. Moreover, each event is associated with an event type (e.g., a computer login attempt, or a hospital discharge). Analyses of event sequence data have been shown to help reveal important temporal patterns, such as clinical paths resulting in improved outcomes, or an understanding of common career trajectories for scholars. Moreover, recent research has demonstrated a variety of techniques designed to overcome methodological challenges such as large volumes of data and high dimensionality. However, the effective identification and analysis of latent stages of progression, which can allow for variation within different but similarly evolving event sequences, remain a significant challenge with important real-world motivations. In this paper, we propose an unsupervised stage analysis algorithm to identify semantically meaningful progression stages as well as the critical events which help define those stages. The algorithm follows three key steps: (1) event representation estimation, (2) event sequence warping and alignment, and (3) sequence segmentation. We also present a novel visualization system, ET2, which interactively illustrates the results of the stage analysis algorithm to help reveal evolution patterns across stages. Finally, we report three forms of evaluation for ET2: (1) case studies with two real-world datasets, (2) interviews with domain expert users, and (3) a performance evaluation on the progression analysis algorithm and the visualization design. Shunan Guo, Zhuochen Jin, David Gotz, Fan Du, Hongyuan Zha, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | ECGLens: Interactive Visual Exploration of Large Scale ECG Data for Arrhythmia DetectionabstractThe Electrocardiogram (ECG) is commonly used to detect arrhythmias. Traditionally, a single ECG observation is used for diagnosis, making it difficult to detect irregular arrhythmias. Recent technology developments, however, have made it cost-effective to collect large amounts of raw ECG data over time. This promises to improve diagnosis accuracy, but the large data volume presents new challenges for cardiologists. This paper introduces ECGLens, an interactive system for arrhythmia detection and analysis using large-scale ECG data. Our system integrates an automatic heartbeat classification algorithm based on convolutional neural network, an outlier detection algorithm, and a set of rich interaction techniques. We also introduce A-glyph, a novel glyph designed to improve the readability and comparison of ECG signals. We report results from a comprehensive user study showing that A-glyph improves the efficiency in arrhythmia detection, and demonstrate the effectiveness of ECGLens in arrhythmia detection through two expert interviews. Shunan Guo, Nan Cao 0001, David Gotz, Aiwen Xu, Huamin Qu, Zhenjie Yao 0001, Yixin Chen 0001 |
CHI | 2 |
| 2018 | Anomaly detection in spatiotemporal data via regularized non-negative tensor analysis
Chaoguang Lin, Qiuhan Zhu, Shunan Guo, Zhuochen Jin, Yu-Ru Lin, Nan Cao 0001 |
Data Min. Knowl. Discov. | 3 |
| 2018 | EventThread: Visual Summarization and Stage Analysis of Event Sequence DataabstractEvent sequence data such as electronic health records, a person's academic records, or car service records, are ordered series of events which have occurred over a period of time. Analyzing collections of event sequences can reveal common or semantically important sequential patterns. For example, event sequence analysis might reveal frequently used care plans for treating a disease, typical publishing patterns of professors, and the patterns of service that result in a well-maintained car. It is challenging, however, to visually explore large numbers of event sequences, or sequences with large numbers of event types. Existing methods focus on extracting explicitly matching patterns of events using statistical analysis to create stages of event progression over time. However, these methods fail to capture latent clusters of similar but not identical evolutions of event sequences. In this paper, we introduce a novel visualization system named EventThread which clusters event sequences into threads based on tensor analysis and visualizes the latent stage categories and evolution patterns by interactively grouping the threads by similarity into time-specific clusters. We demonstrate the effectiveness of EventThread through usage scenarios in three different application domains and via interviews with an expert user. Shunan Guo, Rongwen Zhao, David Gotz, Hongyuan Zha, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |