VLDB 2026 Research / reviewers in the wild / expert
Nan Cao 0001
dblp:66/5146-1
· DBLP profile ↗
127ranked-venue papers
13as first author
60since 2021 · last 2026
0000-0003-1316-7515ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 10 first-author · 34 since 2021Databases, data management, data science and information retrieval · 32 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 27 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 25 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DensiCrafter: Physically-Constrained Generation and Fabrication of Self-Supporting Hollow StructuresabstractThe rise of 3D generative models has enabled automatic 3D geometry and texture synthesis from multimodal inputs (e.g., text or images). However, these methods often ignore physical constraints and manufacturability considerations. In this work, we address the challenge of producing 3D designs that are both lightweight and self-supporting. We present DensiCrafter, a framework for generating lightweight, self-supporting 3D hollow structures by optimizing the density field. Starting from coarse voxel grids produced by Trellis, we interpret these as continuous density fields to optimize and introduce three differentiable, physically constrained, and simulation-free loss terms. Additionally, a mass regularization penalizes unnecessary material, while a restricted optimization domain preserves the outer surface. Our method seamlessly integrates with pretrained Trellis-based models (e.g., Trellis, DSO) without any architectural changes. In extensive evaluations, we achieve up to 43% reduction in material mass on the text-to-3D task. Compared to state-of-the-art baselines, our method could improve the stability and maintain high geometric fidelity. Real-world 3D-printing experiments confirm that our hollow designs can be reliably fabricated and could be self-supporting. Shengqi Dang, Fu Chai, Nan Cao 0001 |
AAAI | 6 |
| 2026 | Designed to Spread: A Generative Approach to Enhance Information DiffusionabstractSocial media has fundamentally transformed how people access information and form social connections, with content expression playing a critical role in driving information diffusion. While prior research has focused largely on network structures and tipping point identification, it provides limited tools for automatically generating content tailored for virality within a specific audience. To fill this gap, we propose the novel task of Diffusion-Oriented Content Generation (DOCG) and introduce an information enhancement algorithm for generating content optimized for diffusion. Our method includes an influence indicator that enables content-level diffusion assessment without requiring access to network topology, and an information editor that employs reinforcement learning to explore interpretable editing strategies. The editor leverages generative models to produce semantically faithful, audience-aware textual or visual content. Experiments on real-world social media datasets and user study demonstrate that our approach significantly improves diffusion effectiveness while preserving the core semantics of the original content. Ziqing Qian, Jiaying Lei, Shengqi Dang, Nan Cao 0001 |
AAAI | 4 |
| 2026 | Beyond Input-Output: Rethinking Creativity through Design-by-Analogy in Human-AI CollaborationabstractWhile the proliferation of foundation models has significantly boosted individual productivity, it also introduces a potential challenge: the homogenization of creative content [39]. In response, we revisit Design-by-Analogy (DbA), a cognitively grounded approach that fosters novel solutions by mapping inspiration across domains. However, prevailing perspectives often restrict DbA to early ideation or specific data modalities, while reducing AI-driven design to simplified input–output pipelines. Such conceptual limitations inadvertently foster widespread design fixation. To address this, we expand the understanding of DbA by embedding it into the entire creative process, thereby demonstrating its capacity to mitigate such fixation. Through a systematic review of 85 studies, we identify six forms of representation and classify techniques across seven stages of the creative process. We further discuss three major application domains: creative industries, intelligent manufacturing, and education and services, demonstrating DbA’s practical relevance. Building on this synthesis, we frame DbA as a mediating technology for human-AI collaboration and outline the potential opportunities and inherent risks for advancing creativity support in HCI and design research. Xuechen Li 0003, Nan Cao 0001, Qing Chen 0001 |
CHI | 3 |
| 2026 | From Touch to Change: Understanding Public Engagement in Data Physicalization for Social GoodabstractData physicalization, which encodes data in physical form, has been increasingly used to engage the public with issues of social good. While public engagement is often invoked as a motivation or expected outcome, it has not been systematically examined as a design objective. This gap raises two key challenges: what characterizes engagement in data physicalization for social good (Phys4Good), and how it can be effectively designed. In this work, we address these challenges by first curating a corpus of 45 Phys4Good projects and deriving a design space structured around a modified three-act framework comprising Stage, Encounter, and Impact. We then conducted semi-structured interviews with designers of eight projects to identify recurring challenges and strategies for fostering engagement. Finally, we demonstrated the effectiveness of our design space and strategies through a case study, which showed that they can guide designers in structuring engagement, anticipating barriers, and creating more impactful Phys4Good experiences. Yechun Peng, Runxi Wu, Nan Cao 0001, Yang Shi 0007 |
CHI | 3 |
| 2026 | Vistoryteller: Designing Data Stories with LLM Agent-Based Generation and Interactive User ControlabstractData stories that combine data, visualizations, and prose are widely used for communication, decision making, and persuasion, but producing them typically requires coordinated effort across specialized roles such as analysts, scripters, and designers, which is time consuming and difficult to manage. Existing AI-assisted methods generally treat storytelling as a single-agent task and offer only coarse, global controls, limiting an author’s ability to preserve and shape their communication intention over the course of a narrative. In this work, we present Vistoryteller, a multi-agent authoring system that models the division of labor found in human teams by assigning specialized large language model agents to complementary roles and orchestrating their interactions to generate cohesive, intention-aligned data stories. Vistoryteller supports fine-grained authorial control through two complementary mechanisms: a sketch-based tension-flow control for specifying how thematic emphasis and narrative tension should evolve, and a conversational interface for issuing localized directives to individual agents or to the team. We evaluate Vistoryteller with two controlled experiments and a qualitative user study. Results show that Vistoryteller generates narratives that align more closely with user intentions, preserve coherence across agent contributions, and surface diverse and expressive insights. Yang Shi 0007, Chuyi Zheng, Nan Cao 0001 |
IUI | 5 |
| 2026 | InScribe: Intelligent Augmentation of Natural-Language Statements with Data Facts
Chuer Chen, Danqing Shi, Shixiong Cao, Nan Cao 0001 |
PacificVis | 5 |
| 2026 | Urania: Visualizing Data Analysis Pipelines for Natural Language-Based Data ExplorationabstractExploratory Data Analysis (EDA) is an essential yet tedious process for examining a new dataset. To facilitate it, Natural Language Interfaces (NLIs) can help people intuitively explore the dataset via data-oriented questions. However, existing NLIs primarily focus on providing accurate answers to questions, with few offering explanations or presentations of the data analysis pipeline used to uncover the answer. Such presentations are crucial for EDA as they enhance the interpretability and reliability of the answer, while also helping users understand the analysis process and derive insights. To fill this gap, we introduce Urania, a natural language interactive system that can visualize the data analysis pipelines used to resolve input questions. It integrates a NLI that allows users to explore data via questions, and a novel data-aware question decomposition algorithm that resolves each input question into a data analysis pipeline. This pipeline is visualized in the form of a datamation, with animated presentations of analysis operations and their corresponding data changes. Through two quantitative experiments and expert interviews, we demonstrated that our data-aware question decomposition algorithm shows competitive performance compared to existing techniques in terms of execution accuracy, and that Urania can help people explore datasets better. In the end, we discuss the observations from the studies and the potential future works. Xiaoyu Qi, Haoyang Li 0015, Jing Zhang 0001, Danqing Shi, Qing Chen 0001, Daniel Weiskopf, Nan Cao 0001 |
ACM Trans. Interact. Intell. Syst. | 8 |
| 2026 | FreeShell: A Context-Free 4D Printing Technique for Fabricating Complex 3D Triangle Mesh ShellsabstractFreeform thin-shell surfaces are critical in various fields, but their fabrication is complex and costly. Traditional methods are wasteful and require custom molds, while 3D printing needs extensive support structures and post-processing. Thermal shrinkage actuated 4D printing is an effective method for fabricating 3D shell. However, existing research faces issues related to precise deformation and limited robustness. Addressing these issues is challenging due to three key factors: (1) Difficulty in finding a universal method to control deformation across different materials; (2) Variability in deformation influenced by factors such as printing speed, layer thickness, and heating temperature; (3) Environmental factors affecting the deformation process. To overcome these challenges, we introduce FreeShell, a robust 4D printing technique that uses thermal shrinkage to create precise 3D shells. This method prints triangular tiles connected by shrinkable connectors using a single material. Upon heating, the connectors shrink, moving the tiles to form the desired 3D shape, simplifying fabrication and reducing material and environment dependency. An optimized mesh layout algorithm computes suitable printing structures that satisfy the defined structural objectives. FreeShell demonstrates its effectiveness through various examples and experiments, showcasing precision, robustness, and strength, representing advancement in fabricating complex freeform surfaces. Shengqi Dang, Xuejiao Ma, Nan Cao 0001 |
ACM Trans. Graph. | 4 |
| 2026 | IDEA: Automated Design Space Exploration for Visualization DesignabstractDesign spaces serve as a conceptual framework that enables designers to explore feasible solutions, yet their lack of computational formalization limits integration with AI-driven design automation. To address this, we introduce a structured design space model that formalizes design spaces with orthogonal dimensions and discrete elements, making them machine-interpretable and executable. Building on this model, we present IDEA, a fully automated design exploration framework to generate effective outcomes based on user requirements and design spaces. Specifically, IDEA leverages large language models (LLMs) for constraint generation, incorporates a constraint-guided Monte Carlo Tree Search (MCTS) algorithm to explore the space, and instantiates abstract decisions into domain-specific implementations. We evaluate IDEA in two design scenarios: data-driven article and standard visualization, supported by comparativeratings, expert interviews, and quantitative experiments. Results demonstrate the IDEA's adaptability across domains and its capability to produce high-quality designs. Chuer Chen, Xiaoke Yan, Xiaoyu Qi, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | How We Map Possibilities: Understanding Design Spaces for VisualizationabstractDesign spaces serve as conceptual frameworks that enable systematic exploration of possibilities and constraints for particular design problems. Despite growing recognition of their importance in visualization research, the community faces two main challenges: characterizing what constitute a design space, given the lack of consensus on its definition, and determining how to construct these spaces in the absence of established methodologies. To address the challenges, we first conducted a literature review of visualization design space research, identifying three distinct research threads. Focusing on the thread that views design spaces as multi-dimensional frameworks, we refined our corpus to 49 papers and developed a unified conceptualization of design spaces. Building on this foundation, we proposed a systematic approach to design space construction, synthesized from an analysis of practices spanning five phases: exploration, data collection, creation, evaluation, and communication. Zichun Dai, Yechun Peng, Nan Cao 0001, Yang Shi 0007 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | ChartBlender: An Interactive System for Authoring and Synchronizing Visualization Charts in VideoabstractEmbedded data visualizations have emerged as a powerful narrative medium for conveying complex information within video footage. However, creating such content remains labor-intensive, as existing workflows rely on manual frame-by-frame adjustments to ensure spatial and temporal consistency. To address these challenges, we present ChartBlender, an interactive authoring system designed to streamline the creation, embedding, and automatic synchronization of data visualizations within video scenes. We develop a tracking pipeline that supports both object and camera tracking, ensuring robust alignment of visualizations with dynamic video content. To maintain visual clarity and aesthetic coherence, we also explore the design space of video-suited visualizations and develop a library of customizable templates optimized for video embedding. We evaluated ChartBlender through two controlled experiments and expert interviews with five domain experts. Results show that our system enables accurate synchronization and accelerates the production of data-driven videos. Chenpu Li, Ruoyan Chen, Chuer Chen, Shengqi Dang, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | EmotiCrafter: Text-to-Emotional-Image Generation Based on Valence-Arousal Model
Shengqi Dang, Long Ling, Ziqing Qian, Nanxuan Zhao, Nan Cao 0001 |
ICCV | 6 |
| 2025 | Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge GraphabstractLarge language models (LLMs) have demonstrated exceptional performance across a wide variety of domains. Nonetheless, generalist LLMs continue to fall short in reasoning tasks necessitating specialized knowledge, e.g., emotional sociology and medicine. Prior investigations into specialized LLMs focused on domain-specific training, which entails substantial efforts in domain data acquisition and model parameter fine-tuning. To address these challenges, this paper proposes the Way-to-Specialist (WTS) framework, which synergizes retrieval-augmented generation with knowledge graphs (KGs) to enhance the specialized capability of LLMs in the absence of specialized training. In distinction to existing paradigms that merely utilize external knowledge from general KGs or static domain KGs to prompt LLM for enhanced domain-specific reasoning, WTS proposes an innovative ''LLM↻KG'' paradigm, which achieves bidirectional enhancement between specialized LLM and domain knowledge graph (DKG). The proposed paradigm encompasses two closely coupled components: the DKG-Augmented LLM and the LLM-Assisted DKG Evolution. The former retrieves question-relevant domain knowledge from DKG and uses it to prompt LLM to enhance the reasoning capability for domain-specific tasks; the latter leverages LLM to generate new domain knowledge from processed tasks and use it to evolve DKG. WTS closes the loop between DKG-Augmented LLM and LLM-Assisted DKG Evolution, enabling continuous improvement in the domain specialization as it progressively answers and learns from domain-specific questions. We validate the performance of WTS on 7 datasets (e.g., TweetQA, ChatDoctor5k) spanning 6 domains, e.g., emotional sociology, medical, ect. The experimental results show that WTS surpasses the previous SOTA in 5 specialized domains, and achieves a maximum performance improvement of 11.3%. Yutong Zhang 0003, Lixing Chen, Shenghong Li 0001, Nan Cao 0001, Yang Shi 0007, Jiaxin Ding 0001, Pan Zhou 0001, Yang Bai 0010 |
KDD (1) | 4 |
| 2025 | ViviClay: Designing and Fabricating Ceramics with Animation Effects on Physical Surfaces
Guanhong Liu, Jingxin Ye, Qiaoqiao Jin, Xuechen Li 0003, Yang Shi 0007, Qing Chen 0001, Nan Cao 0001 |
UIST | 8 |
| 2025 | Double Tap for This Post: Understanding the Communication of Data Visualization on Social MediaabstractData visualizations are increasingly used by news outlets on social media to communicate insights to a broad audience. However, little is known about how readers interact with and respond to data visualizations in these quick-consumption environments. In this work, we introduce a conceptual model that categorizes visualization reading that leads to the communication effect of likes on Instagram. The model was developed through a grounded theory analysis of the statements explaining the reasoning behind the likes of visualization, which were recorded from a preliminary study. Informed by coding the statements from two dimensions including scopes and design patterns concerning visualization, our model consists of three levels: depicting the "look" of a visualization (e.g., artistic style and color scheme); interpreting the "flesh and bones" of a visualization (e.g., visualization and narrative); and elucidating the "heart and soul" of a visualization (e.g., insights and conclusion). We also conducted an online crowdsourcing user study with 200 participants to demonstrate how our model can be applied to improve the communication of visualization by comparing the three levels. Yang Shi 0007, Yechun Peng, Jieying Ding, Xingyu Lan, Nan Cao 0001 |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2025 | MV-Crafter: An Intelligent System for Music-Guided Video GenerationabstractMusic videos, as a prevalent form of multimedia entertainment, deliver engaging audio-visual experiences to audiences and have gained immense popularity among singers and fans. Creators can express their interpretations of music naturally through visual elements. However, the creation process of music video demands proficiency in script design, video shooting, and music-video synchronization, posing significant challenges for non-professionals. Previous work has designed automated music video generation frameworks. However, they suffer from complexity in input and poor output quality. In response, we present MV-Crafter, a system capable of producing high-quality music videos with synchronized music-video rhythm and style. Our approach involves three technical modules that simulate the human creation process: the script generation module, video generation module, and music-video synchronization module. MV-Crafter leverages a large language model to generate scripts considering the musical semantics. To address the challenge of synchronizing short video clips with music of varying lengths, we propose a dynamic beat-matching algorithm and visual envelope-induced warping method to ensure precise, monotonic music-video synchronization. Besides, we design a user-friendly interface to simplify the creation process with intuitive editing features. Extensive experiments have demonstrated that MV-Crafter provides an effective solution for improving the quality of generated music videos. Chuer Chen, Shengqi Dang, Nanxuan Zhao, Yang Shi 0007, Nan Cao 0001 |
ACM Trans. Interact. Intell. Syst. | 6 |
| 2025 | Chart2Vec: A Universal Embedding of Context-Aware VisualizationsabstractThe advances in AI-enabled techniques have accelerated the creation and automation of visualizations in the past decade. However, presenting visualizations in a descriptive and generative format remains a challenge. Moreover, current visualization embedding methods focus on standalone visualizations, neglecting the importance of contextual information for multi-view visualizations. To address this issue, we propose a new representation model, Chart2Vec, to learn a universal embedding of visualizations with context-aware information. Chart2Vec aims to support a wide range of downstream visualization tasks such as recommendation and storytelling. Our model considers both structural and semantic information of visualizations in declarative specifications. To enhance the context-aware capability, Chart2Vec employs multi-task learning on both supervised and unsupervised tasks concerning the cooccurrence of visualizations. We evaluate our method through an ablation study, a user study, and a quantitative comparison. The results verified the consistency of our embedding method with human cognition and showed its advantages over existing methods. Qing Chen 0001, Ruishi Zou, Wei Shuai, Jiazhe Wang, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Visual Analysis of Multi-Outcome Causal GraphsabstractWe introduce a visual analysis method for multiple causal graphs with different outcome variables, namely, multi-outcome causal graphs. Multi-outcome causal graphs are important in healthcare for understanding multimorbidity and comorbidity. To support the visual analysis, we collaborated with medical experts to devise two comparative visualization techniques at different stages of the analysis process. First, a progressive visualization method is proposed for comparing multiple state-of-the-art causal discovery algorithms. The method can handle mixed-type datasets comprising both continuous and categorical variables and assist in the creation of a fine-tuned causal graph of a single o utcome. Second, a comparative graph layout technique and specialized visual encodings are devised for the quick comparison of multiple causal graphs. In our visual analysis approach, analysts start by building individual causal graphs for each outcome variable, and then, multi-outcome causal graphs are generated and visualized with our comparative technique for analyzing differences and commonalities of these causal graphs. Evaluation includes quantitative measurements on benchmark datasets, a case study with a medical expert, and expert user studies with real-world health research data. Mengjie Fan, Jinlu Yu, Daniel Weiskopf, Nan Cao 0001, Huai-Yu Wang, Liang Zhou 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Leveraging Foundation Models for Crafting Narrative Visualization: A SurveyabstractNarrative visualization transforms data into engaging stories, making complex information accessible to a broad audience. Foundation models, with their advanced capabilities such as natural language processing, content generation, and multimodal integration, hold substantial potential for enriching narrative visualization. Recently, a collection of techniques have been introduced for crafting narrative visualizations based on foundation models from different aspects. We build our survey upon 66 articles to study how foundation models can progressively engage in this process and then propose a reference model categorizing the reviewed literature into four essential phases: Analysis, Narration, Visualization, and Interaction. Furthermore, we identify eight specific tasks (e.g., Insight Extraction and Authoring) where foundation models are applied across these stages to facilitate the creation of visual narratives. Detailed descriptions, related literature, and reflections are presented for each task. To make it a more impactful and informative experience for diverse readers, we discuss key research problems and provide the strengths and weaknesses in each task to guide people in identifying and seizing opportunities while navigating challenges in this field. Shixiong Cao, Yang Shi 0007, Qing Chen 0001, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Active Gaze Labeling: Visualization for Trust BuildingabstractAreas of interest (AOIs) are well-established means of providing semantic information for visualizing, analyzing, and classifying gaze data. However, the usual manual annotation of AOIs is time-consuming and further impaired by ambiguities in label assignments. To address these issues, we present an interactive labeling approach that combines visualization, machine learning, and user-centered explainable annotation. Our system provides uncertainty-aware visualization to build trust in classification with an increasing number of annotated examples. It combines specifically designed EyeFlower glyphs, dimensionality reduction, and selection and exploration techniques in an integrated workflow. The approach is versatile and hardware-agnostic, supporting video stimuli from stationary and unconstrained mobile eye tracking alike. We conducted an expert review to assess labeling strategies and trust building. Maurice Koch, Nan Cao 0001, Daniel Weiskopf, Kuno Kurzhals |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Complex Surface Fabrication Via Developable Surface Approximation: A SurveyabstractComplex surfaces are commonly observed in various applications and have significant value in enhancing comfort, aesthetics, and functionality. However, their fabrication often involves complex and costly processes. To simplify the fabrication difficulty, significant research has focused on using 3D developable surfaces to approximate target 3D surfaces. This process involves converting target 3D surfaces into developable surfaces and then flattening them into 2D patterns. Since the geometric and topological diversity of target surfaces, this task is both comprehensive and intricate, encompassing multiple aspects from design to fabrication. In this paper, we review relevant technologies and methods in fabrication processes, classify them, and summarize a pipeline from design to fabrication. This provides a comprehensive introduction to the field for researchers and practitioners. Through the analysis of relevant literature, we also discuss some of the research challenges and future research opportunities. Nan Cao 0001, Yang Shi 0007 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Beyond Numbers: Creating Analogies to Enhance Data Comprehension and Communication with Generative AIabstractUnfamiliar measurements usually hinder readers from grasping the scale of the numerical data, understanding the content, and feeling engaged with the context. To enhance data comprehension and communication, we leverage analogies to bridge the gap between abstract data and familiar measurements. In this work, we first conduct semi-structured interviews with design experts to identify design problems and summarize design considerations. Then, we collect an analogy dataset of 138 cases from various online sources. Based on the collected dataset, we characterize a design space for creating data analogies. Next, we build a prototype system, AnalogyMate, that automatically suggests data analogies, their corresponding design solutions, and generated visual representations powered by generative AI. The study results show the usefulness of AnalogyMate in aiding the creation process of data analogies and the effectiveness of data analogy in enhancing data comprehension and communication. Qing Chen 0001, Wei Shuai, Jiyao Zhang, Zhida Sun, Nan Cao 0001 |
CHI | 5 |
| 2024 | Personalizing Products with Stylized Head Portraits for Self-ExpressionabstractPersonalizing products aesthetically or functionally can help users increase personal relevance and support self-expression. However, using non-abstract personal data such as head portraits for product personalization has been understudied. While recent advances in Artificial Intelligence have enabled generating stylized head portraits, these images also raise concerns about lack of control, artificiality, and ethics, which potentially limit their broader use. In this work, we present PicMe, a design support tool that converts user face photos into stylized head portraits as vector graphics that can be used to personalize products. To enable style transfer, PicMe leverages a deep-learning-based algorithm trained on an extended open-source illustration dataset of characters in a cartoonish and minimalistic style. We evaluated PicMe through two experiments and a user study. The results of our evaluation showed that PicMe can help create personalized head portraits that support self-expression. Yang Shi 0007, Yechun Peng, Shengqi Dang, Nanxuan Zhao, Nan Cao 0001 |
CHI | 5 |
| 2024 | Understanding and Automating Graphical Annotations on Animated ScatterplotsabstractScatterplots are commonly used in various contexts, from scientific publications to infographics for the general public. However, not everyone is able to read them, and even experts may struggle to notice some important information such as overlapping clusters or temporal changes. To address these issues, a computational approach for annotating scatterplots has been developed. This approach involves various forms of annotation, including drawing lines to show correlations, circling areas to show clusters, and indicating movement with arrows. The approach is based on a study that identified common annotation strategies used by people to annotate scatterplots. These strategies are distilled into an automated method for generating graphical annotations on scatterplots. The method involves a problem formulation using a Markov Decision Process and a model for making annotation decisions. The model generates step-by-step graphical annotations by analyzing data insights and observing the chart. The final result conveys a narrative that is easy to understand and allows for the conveyance of temporal changes in the data. The study results suggest that the method can generate understandable and functional annotations that are comparable to those created by human experts. This approach can potentially reduce the time and effort required to read scatterplots, making it a useful tool for data visualization novices. Danqing Shi, Antti Oulasvirta, Tino Weinkauf, Nan Cao 0001 |
PacificVis | 4 |
| 2024 | Talk2Data: A Natural Language Interface for Exploratory Visual Analysis via Question DecompositionabstractThrough a natural language interface (NLI) for exploratory visual analysis, users can directly “ask” analytical questions about the given tabular data. This process greatly improves user experience and lowers the technical barriers of data analysis. Existing techniques focus on generating a visualization from a concrete question. However, complex questions, requiring multiple data queries and visualizations to answer, are frequently asked in data exploration and analysis, which cannot be easily solved with the existing techniques. To address this issue, in this article, we introduce Talk2Data, a natural language interface for exploratory visual analysis that supports answering complex questions. It leverages an advanced deep-learning model to resolve complex questions into a series of simple questions that could gradually elaborate on the users’ requirements. To present answers, we design a set of annotated and captioned visualizations to represent the answers in a form that supports interpretation and narration. We conducted an ablation study and a controlled user study to evaluate the Talk2Data’s effectiveness and usefulness. Danqing Shi, Mingjuan Guo, Yanqiu Wu 0001, Nan Cao 0001, Qing Chen 0001 |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2024 | Calliope-Net: Automatic Generation of Graph Data Facts via Annotated Node-Link DiagramsabstractGraph or network data are widely studied in both data mining and visualization communities to review the relationship among different entities and groups. The data facts derived from graph visual analysis are important to help understand the social structures of complex data, especially for data journalism. However, it is challenging for data journalists to discover graph data facts and manually organize correlated facts around a meaningful topic due to the complexity of graph data and the difficulty to interpret graph narratives. Therefore, we present an automatic graph facts generation system, Calliope-Net, which consists of a fact discovery module, a fact organization module, and a visualization module. It creates annotated node-link diagrams with facts automatically discovered and organized from network data. A novel layout algorithm is designed to present meaningful and visually appealing annotated graphs. We evaluate the proposed system with two case studies and an in-lab user study. The results show that Calliope-Net can benefit users in discovering and understanding graph data facts with visually pleasing annotated visualizations. Qing Chen 0001, Wei Shuai, Guande Wu, Zhe Xu 0007, Hanghang Tong, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | How Does Automation Shape the Process of Narrative Visualization: A Survey of ToolsabstractIn recent years, narrative visualization has gained much attention. Researchers have proposed different design spaces for various narrative visualization genres and scenarios to facilitate the creation process. As users' needs grow and automation technologies advance, increasingly more tools have been designed and developed. In this study, we summarized six genres of narrative visualization (annotated charts, infographics, timelines & storylines, data comics, scrollytelling & slideshow, and data videos) based on previous research and four types of tools (design spaces, authoring tools, ML/AI-supported tools and ML/AI-generator tools) based on the intelligence and automation level of the tools. We surveyed 105 papers and tools to study how automation can progressively engage in visualization design and narrative processes to help users easily create narrative visualizations. This research aims to provide an overview of current research and development in the automation involvement of narrative visualization tools. We discuss key research problems in each category and suggest new opportunities to encourage further research in the related domain. Qing Chen 0001, Shixiong Cao, Jiazhe Wang, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Affective Visualization Design: Leveraging the Emotional Impact of DataabstractIn recent years, more and more researchers have reflected on the undervaluation of emotion in data visualization and highlighted the importance of considering human emotion in visualization design. Meanwhile, an increasing number of studies have been conducted to explore emotion-related factors. However, so far, this research area is still in its early stages and faces a set of challenges, such as the unclear definition of key concepts, the insufficient justification of why emotion is important in visualization design, and the lack of characterization of the design space of affective visualization design. To address these challenges, first, we conducted a literature review and identified three research lines that examined both emotion and data visualization. We clarified the differences between these research lines and kept 109 papers that studied or discussed how data visualization communicates and influences emotion. Then, we coded the 109 papers in terms of how they justified the legitimacy of considering emotion in visualization design (i.e., why emotion is important) and identified five argumentative perspectives. Based on these papers, we also identified 61 projects that practiced affective visualization design. We coded these design projects in three dimensions, including design fields (where), design tasks (what), and design methods (how), to explore the design space of affective visualization design. Xingyu Lan, Yanqiu Wu 0001, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Supporting Guided Exploratory Visual Analysis on Time Series Data with Reinforcement LearningabstractThe exploratory visual analysis (EVA) of time series data uses visualization as the main output medium and input interface for exploring new data. However, for users who lack visual analysis expertise, interpreting and manipulating EVA can be challenging. Thus, providing guidance on EVA is necessary and two relevant questions need to be answered. First, how to recommend interesting insights to provide a first glance at data and help develop an exploration goal. Second, how to provide step-by-step EVA suggestions to help identify which parts of the data to explore. In this work, we present a reinforcement learning (RL)-based system, Visail, which generates EVA sequences to guide the exploration of time series data. As a user uploads a time series dataset, Visail can generate step-by-step EVA suggestions, while each step is visualized as an annotated chart combined with textual descriptions. The RL-based algorithm uses exploratory data analysis knowledge to construct the state and action spaces for the agent to imitate human analysis behaviors in data exploration tasks. In this way, the agent learns the strategy of generating coherent EVA sequences through a well-designed network. To evaluate the effectiveness of our system, we conducted an ablation study, a user study, and two case studies. The results of our evaluation suggested that Visail can provide effective guidance on supporting EVA on time series data. Yang Shi 0007, Bingchang Chen, Zhuochen Jin, Xiaohan Jiao, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2024 | : A Visual Analytics Approach for Understanding the Dual Frontiers of Science and TechnologyabstractScience has long been viewed as a key driver of economic growth and rising standards of living. Knowledge about how scientific advances support marketplace inventions is therefore essential for understanding the role of science in propelling real-world applications and technological progress. The increasing availability of large-scale datasets tracing scientific publications and patented inventions and the complex interactions among them offers us new opportunities to explore the evolving dual frontiers of science and technology at an unprecedented level of scale and detail. However, we lack suitable visual analytics approaches to analyze such complex interactions effectively. Here we introduce InnovationInsights, an interactive visual analysis system for researchers, research institutions, and policymakers to explore the complex linkages between science and technology, and to identify critical innovations, inventors, and potential partners. The system first identifies important associations between scientific papers and patented inventions through a set of statistical measures introduced by our experts from the field of the Science of Science. A series of visualization views are then used to present these associations in the data context. In particular, we introduce the Interplay Graph to visualize patterns and insights derived from the data, helping users effectively navigate citation relationships between papers and patents. This visualization thereby helps them identify the origins of technical inventions and the impact of scientific research. We evaluate the system through two case studies with experts followed by expert interviews. We further engage a premier research institution to test-run the system, helping its institution leaders to extract new insights for innovation. Through both the case studies and the engagement project, we find that our system not only meets our original goals of design, allowing users to better identify the sources of technical inventions and to understand the broad impact of scientific research; it also goes beyond these purposes to enable an array of new applications for researchers and research institutions, ranging from identifying untapped innovation potential within an institution to forging new collaboration opportunities between science and industry. Yifang Wang 0001, Yifan Qian, Xiaoyu Qi, Nan Cao 0001, Dashun Wang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Bring Clipart to LifeabstractThe development of face editing has been boosted since the birth of StyleGAN. While previous works have explored different interactive methods, such as sketching and exemplar photos, they have been limited in terms of expressiveness and generality. In this paper, we propose a new interaction method by guiding the editing with abstract clipart, composed of a set of simple semantic parts, allowing users to control across face photos with simple clicks. However, this is a challenging task given the large domain gap between colorful face photos and abstract clipart with limited data. To solve this problem, we introduce a frame-work called ClipFaceShop1built on top of StyleGAN. The key idea is to take advantage of $\mathcal{W} +$ latent code encoded rich and disentangled visual features, and create a new lightweight selective feature adaptor to predict a modifiable path toward the target output photo. Since no pairwise labeled data exists for training, we design a set of losses to provide supervision signals for learning the modifiable path. Experimental results show that ClipFaceShop generates realistic and faithful face photos, sharing the same facial attributes as the reference clipart. We demonstrate that ClipFaceShop supports clipart in diverse styles, even in form of a free-hand sketch. Nanxuan Zhao, Shengqi Dang, Hexun Lin, Yang Shi 0007, Nan Cao 0001 |
ICCV | 5 |
| 2023 | Understanding Design Collaboration Between Designers and Artificial Intelligence: A Systematic Literature ReviewabstractRecent interest in design through the artificial intelligence (AI) lens is rapidly increasing. Designers, as a special user group interacting with AI, have received more attention in the Human-Computer Interaction community. Prior work has discussed emerging challenges that persist in designing for AI. However, few systematic reviews focus on AI for design to understand how designers and AI can augment each other's complementary strengths in design collaboration. In this work, we conducted a landscape analysis of AI for design, via a systematic literature review of 93 papers. The analysis first provides a bird's eye view of overall patterns in this area. The analysis also reveals three themes interpreted from the paper corpus associated with AI for design, including AI assisting designers, designers assisting AI, and characterizing designer-AI collaboration. We discuss the implications of our findings and suggested methodological proposals to guide HCI toward research and practices that center on collaborative creativity. Yang Shi 0007, Xiaohan Jiao, Nan Cao 0001 |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2023 | Adversarial Attacks on Multi-Network Mining: Problem Definition and Fast SolutionsabstractMulti-sourced networks naturally appear in many application domains, ranging from bioinformatics, social networks, neuroscience to management. Although state-of-the-art offers rich models and algorithms to find various patterns when input networks are given, it has largely remained nascent on how vulnerable the mining results are due to the adversarial attacks. In this paper, we address the problem of attacking multi-network mining through the way of deliberately perturbing the networks to alter the mining results. The key idea of the proposed method admiring is effective and efficient influence functions on the Sylvester equation defined over the input networks, which plays a central and unifying role in various multi-network mining tasks. The proposed algorithms bear three main advantages, including (1) effectiveness, being able to accurately quantify the rate of change of the mining results in response to attacks; (2) efficiency, scaling linearly with more than 100 times speed-up over the straight-forward implementation without any quality loss; and (3) generality, being applicable to a variety of multi-network mining tasks (e.g., graph kernel, network alignment, cross-network node similarity) with different attacking strategies (e.g., edge/node removal, attribute alteration). Qinghai Zhou, Liangyue Li, Nan Cao 0001, Lei Ying 0001, Hanghang Tong |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Diverse Interaction Recommendation for Public Users Exploring Multi-view Visualization using Deep LearningabstractInteraction is an important channel to offer users insights in interactive visualization systems. However, which interaction to operate and which part of data to explore are hard questions for public users facing a multi-view visualization for the first time. Making these decisions largely relies on professional experience and analytic abilities, which is a huge challenge for non-professionals. To solve the problem, we propose a method aiming to provide diverse, insightful, and real-time interaction recommendations for novice users. Building on the Long-Short Term Memory Model (LSTM) structure, our model captures users' interactions and visual states and encodes them in numerical vectors to make further recommendations. Through an illustrative example of a visualization system about Chinese poets in the museum scenario, the model is proven to be workable in systems with multi-views and multiple interaction types. A further user study demonstrates the method's capability to help public users conduct more insightful and diverse interactive explorations and gain more accurate data insights. Yusheng Qi, Yang Shi 0007, Qing Chen 0001, Nan Cao 0001, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Breaking the Fourth Wall of Data Stories through InteractionabstractInteraction is increasingly integrating into data stories to support data exploration and explanation. Interaction can also be combined with the narrative device, breaking the fourth wall (BTFW), to build a deeper connection between readers and data stories. BTFW interaction directly addresses readers by requiring their input. Such user input is then integrated into the narrative or visuals of data stories to encourage readers to inspect the stories more closely. In this work, we explore the design patterns of BTFW interaction commonly used in data stories. Six design patterns were identified through the analysis of 58 high-quality data stories collected from a range of online sources. Specifically, the data stories were categorized using a coding framework, including the input of BTFW interaction provided by readers and the output of BTFW interaction generated by data stories to respond to the input. To explore the benefits as well as concerns of using BTFW interaction, we conducted a three-session user study including the reading, interview, and recall sessions. The results of our user study suggested that BTFW interaction has a positive impact on self-story connection, user engagement, and information recall. We also discussed design implications to address the possible negative effects on the interactivity-comprehensibility balance, information privacy, and the learning curve of interaction brought by BTFW interaction. Yang Shi 0007, Xiaohan Jiao, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Supporting Expressive and Faithful Pictorial Visualization Design with Visual Style TransferabstractPictorial visualizations portray data with figurative messages and approximate the audience to the visualization. Previous research on pictorial visualizations has developed authoring tools or generation systems, but their methods are restricted to specific visualization types and templates. Instead, we propose to augment pictorial visualization authoring with visual style transfer, enabling a more extensible approach to visualization design. To explore this, our work presents Vistylist, a design support tool that disentangles the visual style of a source pictorial visualization from its content and transfers the visual style to one or more intended pictorial visualizations. We evaluated Vistylist through a survey of example pictorial visualizations, a controlled user study, and a series of expert interviews. The results of our evaluation indicated that Vistylist is useful for creating expressive and faithful pictorial visualizations. Yang Shi 0007, Siji Chen, Mengdi Sun, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Erato: Cooperative Data Story Editing via Fact InterpolationabstractAs an effective form of narrative visualization, visual data stories are widely used in data-driven storytelling to communicate complex insights and support data understanding. Although important, they are difficult to create, as a variety of interdisciplinary skills, such as data analysis and design, are required. In this work, we introduce Erato, a human-machine cooperative data story editing system, which allows users to generate insightful and fluent data stories together with the computer. Specifically, Erato only requires a number of keyframes provided by the user to briefly describe the topic and structure of a data story. Meanwhile, our system leverages a novel interpolation algorithm to help users insert intermediate frames between the keyframes to smooth the transition. We evaluated the effectiveness and usefulness of the Erato system via a series of evaluations including a Turing test, a controlled user study, a performance validation, and interviews with three expert users. The evaluation results showed that the proposed interpolation technique was able to generate coherent story content and help users create data stories more efficiently. Mengdi Sun, Ligan Cai, Weiwei Cui 0001, Yanqiu Wu 0001, Yang Shi 0007, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Negative Emotions, Positive Outcomes? Exploring the Communication of Negativity in Serious Data StoriesabstractRecent work has highlighted that emotion is key to the user experience with data stories. However, limited attention has been paid to negative emotions specifically. This work investigates the outcomes of negative emotions in the context of serious data stories and examines how they can be augmented by design methods from the perspectives of both storytellers and viewers. First, we conducted a workshop with 9 data story experts to understand the possible benefits of eliciting negative emotions in serious data stories and 19 potential design methods that contribute to negative emotions. Based on the findings from the workshop, we then conducted a lab study with 35 participants to explore the outcomes of eliciting negative emotions as well as the effectiveness of the design methods. The results indicated that negative emotions mainly facilitated contemplative experiences and long-term memory. Besides, the design methods showed varied effectiveness in augmenting negative emotions and being recalled. Xingyu Lan, Yanqiu Wu 0001, Yang Shi 0007, Qing Chen 0001, Nan Cao 0001 |
CHI | 5 |
| 2022 | iNet: visual analysis of irregular transition in multivariate dynamic networks
Dongming Han, Jiacheng Pan, Rusheng Pan, Dawei Zhou 0003, Nan Cao 0001, Jingrui He, Mingliang Xu 0001, Wei Chen 0001 |
Frontiers Comput. Sci. | 5 |
| 2022 | Development and validation of a deep learning model to predict the survival of patients in ICUabstractBACKGROUND: Patients in the intensive care unit (ICU) are often in critical condition and have a high mortality rate. Accurately predicting the survival probability of ICU patients is beneficial to timely care and prioritizing medical resources to improve the overall patient population survival. Models developed by deep learning (DL) algorithms show good performance on many models. However, few DL algorithms have been validated in the dimension of survival time or compared with traditional algorithms. METHODS: Variables from the Early Warning Score, Sequential Organ Failure Assessment Score, Simplified Acute Physiology Score II, Acute Physiology and Chronic Health Evaluation (APACHE) II, and APACHE IV models were selected for model development. The Cox regression, random survival forest (RSF), and DL methods were used to develop prediction models for the survival probability of ICU patients. The prediction performance was independently evaluated in the MIMIC-III Clinical Database (MIMIC-III), the eICU Collaborative Research Database (eICU), and Shanghai Pulmonary Hospital Database (SPH). RESULTS: Forty variables were collected in total for model development. 83 943 participants from 3 databases were included in the study. The New-DL model accurately stratified patients into different survival probability groups with a C-index of >0.7 in the MIMIC-III, eICU, and SPH, performing better than the other models. The calibration curves of the models at 3 and 10 days indicated that the prediction performance was good. A user-friendly interface was developed to enable the model's convenience. CONCLUSIONS: Compared with traditional algorithms, DL algorithms are more accurate in predicting the survival probability during ICU hospitalization. This novel model can provide reliable, individualized survival probability prediction. Hai Tang, Zhuochen Jin, Jiajun Deng, Yunlang She, Yifan Zhong, Weiyan Sun, Yijiu Ren, Nan Cao 0001, Chang Chen 0006 |
J. Am. Medical Informatics Assoc. | 8 |
| 2022 | ColorCook: Augmenting Color Design for Dashboarding with Domain-Associated PalettesabstractVisualization dashboards serve as an information presentation that uses a tiled layout of key metrics visualized in charts for collaborative decision-making. Existing work has developed tools and techniques for computational color design. Much of these efforts have focused on selecting effective color palettes for independent charts while few attempts have been made to support the expressive color design of multiple coordinated charts in dashboards. In this work, we describe ColorCook, an interactive system that helps design expressive and effective dashboard colorings using domain-associated palettes. ColorCook employs an integrated color workflow for dashboarding, consisting of color selection, assignment, and adjustment. We evaluated ColorCook through a crowdsourcing experiment and a user study. The results of our evaluation indicated that ColorCook is useful for effective and expressive color design. Yang Shi 0007, Siji Chen, Nan Cao 0001 |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2022 | Visual Analytics of Anomalous User Behaviors: A SurveyabstractWith the pervasive use of information technologies, the increasing availability of data provides new opportunities for understanding user behaviors. Unearthing anomalies in user behavior is of particular importance as it helps signal harmful incidents such as network intrusions, terrorist activities, and financial frauds. In this article, we survey state-of-the-art research work in visual analytics of anomalous user behaviors and classify them into four application domains, which are social interaction, travel, network communication, and financial transaction. We further examine the research work in each category in terms of data types, visualization techniques, and interactive analysis methods. We hope that our survey can provide systematic guidelines for researchers and practitioners to find effective solutions to their research problems in specific application domains. Finally, we discuss trends of academic interest over the past decades and suggest potential directions across visual analytics of these user behaviors for future research. Yang Shi 0007, Yuyin Liu, Hanghang Tong, Jingrui He, Nan Cao 0001 |
IEEE Trans. Big Data | 6 |
| 2022 | Guest Editors' Introduction: Special Section on IEEE PacificVis 2022abstractThis special section of the IEEE Transactions on Visualization and Computer Graphics (IEEE TVCG) presents the five most highly rated papers from the 2022 IEEE Pacific Visualization Symposium (IEEE PacificVis). This year, IEEE PacificVis was scheduled to be hosted by the University of Tsukuba and held in Tsukuba, Japan, from April 11 to 14, 2022. IEEE PacificVis, sponsored by the IEEE Visualization and Graphics Technical Committee (VGTC), aims to foster greater exchange between visualization researchers and practitioners, especially in the Asia-Pacific region. This forum has grown to be a truly international event, attracting submissions and attendees from many countries in the Asia-Pacific, Europe, America, and beyond. Thus, IEEE PacificVis is serving the additional purposes of sharing the latest advances in visualization with researchers and practitioners in the region and introducing research developments in the region to the broader international visualization research community. Nan Cao 0001, Timo Ropinski, Jian Zhao 0010 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | VizLinter: A Linter and Fixer Framework for Data VisualizationabstractDespite the rising popularity of automated visualization tools, existing systems tend to provide direct results which do not always fit the input data or meet visualization requirements. Therefore, additional specification adjustments are still required in real-world use cases. However, manual adjustments are difficult since most users do not necessarily possess adequate skills or visualization knowledge. Even experienced users might create imperfect visualizations that involve chart construction errors. We present a framework, VizLinter, to help users detect flaws and rectify already-built but defective visualizations. The framework consists of two components, (1) a visualization linter, which applies well-recognized principles to inspect the legitimacy of rendered visualizations, and (2) a visualization fixer, which automatically corrects the detected violations according to the linter. We implement the framework into an online editor prototype based on Vega-Lite specifications. To further evaluate the system, we conduct an in-lab user study. The results prove its effectiveness and efficiency in identifying and fixing errors for data visualizations. Qing Chen 0001, Fuling Sun, Zui Chen, Jiazhe Wang, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Survey on Visual Analysis of Event Sequence DataabstractEvent sequence data record series of discrete events in the time order of occurrence. They are commonly observed in a variety of applications ranging from electronic health records to network logs, with the characteristics of large-scale, high-dimensional and heterogeneous. This high complexity of event sequence data makes it difficult for analysts to manually explore and find patterns, resulting in ever-increasing needs for computational and perceptual aids from visual analytics techniques to extract and communicate insights from event sequence datasets. In this paper, we review the state-of-the-art visual analytics approaches, characterize them with our proposed design space, and categorize them based on analytical tasks and applications. From our review of relevant literature, we have also identified several remaining research challenges and future research opportunities. Shunan Guo, Zhuochen Jin, Smiti Kaul, David Gotz, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Interpretable Anomaly Detection in Event Sequences via Sequence Matching and Visual ComparisonabstractAnomaly detection is a common analytical task that aims to identify rare cases that differ from the typical cases that make up the majority of a dataset. When analyzing event sequence data, the task of anomaly detection can be complex because the sequential and temporal nature of such data results in diverse definitions and flexible forms of anomalies. This, in turn, increases the difficulty in interpreting detected anomalies. In this article, we propose a visual analytic approach for detecting anomalous sequences in an event sequence dataset via an unsupervised anomaly detection algorithm based on Variational AutoEncoders. We further compare the anomalous sequences with their reconstructions and with the normal sequences through a sequence matching algorithm to identify event anomalies. A visual analytics system is developed to support interactive exploration and interpretations of anomalies through novel visualization designs that facilitate the comparison between anomalous sequences and normal sequences. Finally, we quantitatively evaluate the performance of our anomaly detection algorithm, demonstrate the effectiveness of our system through case studies, and report feedback collected from study participants. Shunan Guo, Zhuochen Jin, Qing Chen 0001, David Gotz, Hongyuan Zha, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Improving Visualization Interpretation Using CounterfactualsabstractComplex, high-dimensional data is used in a wide range of domains to explore problems and make decisions. Analysis of high-dimensional data, however, is vulnerable to the hidden influence of confounding variables, especially as users apply ad hoc filtering operations to visualize only specific subsets of an entire dataset. Thus, visual data-driven analysis can mislead users and encourage mistaken assumptions about causality or the strength of relationships between features. This work introduces a novel visual approach designed to reveal the presence of confounding variables via counterfactual possibilities during visual data analysis. It is implemented in CoFact, an interactive visualization prototype that determines and visualizes counterfactual subsets to better support user exploration of feature relationships. Using publicly available datasets, we conducted a controlled user study to demonstrate the effectiveness of our approach; the results indicate that users exposed to counterfactual visualizations formed more careful judgments about feature-to-outcome relationships. Smiti Kaul, David Borland, Nan Cao 0001, David Gotz |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Kineticharts: Augmenting Affective Expressiveness of Charts in Data Stories with Animation DesignabstractData stories often seek to elicit affective feelings from viewers. However, how to design affective data stories remains under-explored. In this work, we investigate one specific design factor, animation, and present Kineticharts, an animation design scheme for creating charts that express five positive affects: joy, amusement, surprise, tenderness, and excitement. These five affects were found to be frequently communicated through animation in data stories. Regarding each affect, we designed varied kinetic motions represented by bar charts, line charts, and pie charts, resulting in 60 animated charts for the five affects. We designed Kineticharts by first conducting a need-finding study with professional practitioners from data journalism and then analyzing a corpus of affective motion graphics to identify salient kinetic patterns. We evaluated Kineticharts through two user studies. The results suggest that Kineticharts can accurately convey affects, and improve the expressiveness of data stories, as well as enhance user engagement without hindering data comprehension compared to the animation design from DataClips, an authoring tool for data videos. Xingyu Lan, Yang Shi 0007, Yanqiu Wu 0001, Xiaohan Jiao, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | A Design Space for Applying the Freytag's Pyramid Structure to Data StoriesabstractData stories integrate compelling visual content to communicate data insights in the form of narratives. The narrative structure of a data story serves as the backbone that determines its expressiveness, and it can largely influence how audiences perceive the insights. Freytag's Pyramid is a classic narrative structure that has been widely used in film and literature. While there are continuous recommendations and discussions about applying Freytag's Pyramid to data stories, little systematic and practical guidance is available on how to use Freytag's Pyramid for creating structured data stories. To bridge this gap, we examined how existing practices apply Freytag's Pyramid by analyzing stories extracted from 103 data videos. Based on our findings, we proposed a design space of narrative patterns, data flows, and visual communications to provide practical guidance on achieving narrative intents, organizing data facts, and selecting visual design techniques through story creation. We evaluated the proposed design space through a workshop with 25 participants. Results show that our design space provides a clear framework for rapid storyboarding of data stories with Freytag's Pyramid. Leni Yang, Xingyu Lan, Shunan Guo, Yang Shi 0007, Huamin Qu, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2021 | Communicating with Motion: A Design Space for Animated Visual Narratives in Data VideosabstractData videos are a genre of narrative visualization that communicates stories by combining data visualization and motion graphics. While data videos are increasingly gaining popularity, few systematic reviews or structured analyses exist for their design. In this work, we introduce a design space for animated visual narratives in data videos. The design space combines a dimension for animation techniques that are frequently used to facilitate data communication with one for visual narrative strategies served by such animation techniques to support story presentation. We derived our design space from the analysis of 82 high-quality data videos collected from online sources. We conducted a workshop with 20 participants to evaluate the effectiveness of our design space. Qualitative and quantitative feedback suggested that our design space is inspirational and useful for designing and creating data videos. Yang Shi 0007, Xingyu Lan, Zhaorui Li, Nan Cao 0001 |
CHI | 5 |
| 2021 | Vinci: An Intelligent Graphic Design System for Generating Advertising PostersabstractAdvertising posters are a commonly used form of information presentation to promote a product. Producing advertising posters often takes much time and effort of designers when confronted with abundant choices of design elements and layouts. This paper presents Vinci, an intelligent system that supports the automatic generation of advertising posters. Given the user-specified product image and taglines, Vinci uses a deep generative model to match the product image with a set of design elements and layouts for generating an aesthetic poster. The system also integrates online editing-feedback that supports users in editing the posters and updating the generated results with their design preference. Through a series of user studies and a Turing test, we found that Vinci can generate posters as good as human designers and that the online editing-feedback improves the efficiency in poster modification. Shunan Guo, Zhuochen Jin, Fuling Sun, Zhaorui Li, Yang Shi 0007, Nan Cao 0001 |
CHI | 7 |
| 2021 | Understanding Narrative Linearity for Telling Expressive Time-Oriented StoriesabstractCreating expressive narrative visualization often requires choosing a well-planned narrative order that invites the audience in. The narrative can either follow the linear order of story events (chronology), or deviate from linearity (anachronies). While evidence exists that anachronies in novels and films can enhance story expressiveness, little is known about how they can be incorporated into narrative visualization. To bridge this gap, this work introduces the idea of narrative linearity to visualization and investigates how different narrative orders affect the expressiveness of time-oriented stories. First, we conducted preliminary interviews with seven experts to clarify the motivations and challenges of manipulating narrative linearity in time-oriented stories. Then, we analyzed a corpus of 80 time-oriented stories and identified six most salient patterns of narrative orders. Next, we conducted a crowdsourcing study with 221 participants. Results indicated that anachronies have the potential to make time-oriented stories more expressive without hindering comprehensibility. Xingyu Lan, Nan Cao 0001 |
CHI | 3 |
| 2021 | Deep Co-Attention Network for Multi-View Subspace LearningabstractMany real-world applications involve data from multiple modalities and thus exhibit the view heterogeneity. For example, user modeling on social media might leverage both the topology of the underlying social network and the content of the users’ posts; in the medical domain, multiple views could be X-ray images taken at different poses. To date, various techniques have been proposed to achieve promising results, such as canonical correlation analysis based methods, etc. In the meanwhile, it is critical for decision-makers to be able to understand the prediction results from these methods. For example, given the diagnostic result that a model provided based on the X-ray images of a patient at different poses, the doctor needs to know why the model made such a prediction. However, state-of-the-art techniques usually suffer from the inability to utilize the complementary information of each view and to explain the predictions in an interpretable manner. Lecheng Zheng, Hongxia Yang, Nan Cao 0001, Jingrui He |
WWW | 4 |
| 2021 | Attent: Active Attributed Network AlignmentabstractNetwork alignment finds node correspondences across multiple networks, where the alignment accuracy is of crucial importance because of its profound impact on downstream applications. The vast majority of existing works focus on how to best utilize the topology and attribute information of the input networks as well as the anchor links when available. Nonetheless, it has not been well studied on how to boost the alignment performance through actively obtaining high-quality and informative anchor links, with a few exceptions. The sparse literature on active network alignment introduces the human in the loop to label some seed node correspondence (i.e., anchor links), which are informative from the perspective of querying the most uncertain node given few potential matchings. However, the direct influence of the intrinsic network attribute information on the alignment results has largely remained unknown. In this paper, we tackle this challenge and propose an active network alignment method (Attent) to identify the best nodes to query. The key idea of the proposed method is to leverage effective and efficient influence functions defined over the alignment solution to evaluate the goodness of the candidate nodes for query. Our proposed query strategy bears three distinct advantages, including (1) effectiveness, being able to accurately quantify the influence of the candidate nodes on the alignment results; (2) efficiency, scaling linearly with 15 − 17 × speed-up over the straight-forward implementation without any quality loss; (3) generality, consistently improving alignment performance of a variety of network alignment algorithms. Qinghai Zhou, Liangyue Li, Xintao Wu, Nan Cao 0001, Lei Ying 0001, Hanghang Tong |
WWW | 4 |
| 2021 | AutoClips: An Automatic Approach to Video Generation from Data FactsabstractAbstract Data videos, a storytelling genre that visualizes data facts with motion graphics, are gaining increasing popularity among data journalists, non‐profits, and marketers to communicate data to broad audiences. However, crafting a data video is often time‐consuming and asks for various domain knowledge such as data visualization, animation design, and screenwriting. Existing authoring tools usually enable users to edit and compose a set of templates manually, which still cost a lot of human effort. To further lower the barrier of creating data videos, this work introduces a new approach, AutoClips, which can automatically generate data videos given the input of a sequence of data facts. We built AutoClips through two stages. First, we constructed a fact‐driven clip library where we mapped ten data facts to potential animated visualizations respectively by analyzing 230 online data videos and conducting interviews. Next, we constructed an algorithm that generates data videos from data facts through three steps: selecting and identifying the optimal clip for each of the data facts, arranging the clips into a coherent video, and optimizing the duration of the video. The results from two user studies indicated that the data videos generated by AutoClips are comprehensible, engaging, and have comparable quality with human‐made videos. Danqing Shi, F. Sun, Xingyu Lan, David Gotz, Nan Cao 0001 |
Comput. Graph. Forum | 6 |
| 2021 | Graph Ranking Auditing: Problem Definition and Fast SolutionsabstractRanking on graphs is a centerpiece in many high-impact application domains, such as information retrieval, recommender systems, team management, neuroscience and many more. PageRank, along with many of its variants, is widely used across these application domains thanks to its mathematical elegance and the superior performance. Although PageRank and its variants are effective in ranking nodes on graphs, they often lack an efficient and effective way to audit the ranking results in terms of the input graph structure, e.g., which node or edge in the graph contributes most to the top-1 ranked node; which subgraph plays a crucial role in generating the overall ranking result? In this paper, we propose to audit graph ranking by finding the influential graph elements (e.g., edges, nodes, attributes, and subgraphs) regarding their impact on the ranking results. First, we formulate graph ranking auditing problem as quantifying the influence of graph elements on the ranking results. Second, we show that our formulation can be applied to a variety of graph structures. Third, we propose effective and efficient algorithms to find the top-k influential edges/nodes/subgraph. Finally, we perform extensive empirical evaluations on real-world datasets to demonstrate that the proposed methods (Aurora) provide intuitive auditing results with linear scalability. Jian Kang 0008, Nan Cao 0001, Yinglong Xia, Wei Fan 0001, Hanghang Tong |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Guest Editors' Introduction: Special Section on IEEE PacificVis 2021abstractThis special section of the IEEE Transactions on Visualization and Computer Graphics (IEEE TVCG) presents the five most highly rated papers from the 2021 IEEE Pacific Visualization Symposium (IEEE PacificVis). This year, IEEE PacificVis was scheduled to be hosted by Tianjin University and held in Tianjin, China, from April 19 to 22, 2021. IEEE PacificVis, sponsored by the IEEE Visualization and Graphics Technical Committee (VGTC), aims to foster greater exchange between visualization researchers and practitioners, especially in the Asia-Pacific region. This forum has grown to be a truly international event, attracting submissions and attendees from many countries in the Asia-Pacific and Europe, America, and beyond. Thus, IEEE PacificVis is serving the additional purposes of sharing the latest advances in visualization with researchers and practitioners in the region and introducing research developments in the region to the broader international visualization research community. Nan Cao 0001, Holger Theisel, Chaoli Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | Visual Causality Analysis of Event Sequence DataabstractCausality is crucial to understanding the mechanisms behind complex systems and making decisions that lead to intended outcomes. Event sequence data is widely collected from many real-world processes, such as electronic health records, web clickstreams, and financial transactions, which transmit a great deal of information reflecting the causal relations among event types. Unfortunately, recovering causalities from observational event sequences is challenging, as the heterogeneous and high-dimensional event variables are often connected to rather complex underlying event excitation mechanisms that are hard to infer from limited observations. Many existing automated causal analysis techniques suffer from poor explainability and fail to include an adequate amount of human knowledge. In this paper, we introduce a visual analytics method for recovering causalities in event sequence data. We extend the Granger causality analysis algorithm on Hawkes processes to incorporate user feedback into causal model refinement. The visualization system includes an interactive causal analysis framework that supports bottom-up causal exploration, iterative causal verification and refinement, and causal comparison through a set of novel visualizations and interactions. We report two forms of evaluation: a quantitative evaluation of the model improvements resulting from the user-feedback mechanism, and a qualitative evaluation through case studies in different application domains to demonstrate the usefulness of the system. Zhuochen Jin, Shunan Guo, Daniel Weiskopf, David Gotz, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Smile or Scowl? Looking at Infographic Design Through the Affective LensabstractInfographics are frequently promoted for their ability to communicate data to audiences affectively. To facilitate the creation of affect-stirring infographics, it is important to characterize and understand people's affective responses to infographics and derive practical design guidelines for designers. To address these research questions, we first conducted two crowdsourcing studies to identify 12 infographic-associated affective responses and collect user feedback explaining what triggered affective responses in infographics. Then, by coding the user feedback, we present a taxonomy of design heuristics that exemplifies the affect-related design factors in infographics. We evaluated the design heuristics with 15 designers. The results showed that our work supports assessing the affective design in infographics and facilitates the ideation and creation of affective infographics. Xingyu Lan, Yang Shi 0007, Yueyao Zhang, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | Calliope: Automatic Visual Data Story Generation from a SpreadsheetabstractVisual data stories shown in the form of narrative visualizations such as a poster or a data video, are frequently used in data-oriented storytelling to facilitate the understanding and memorization of the story content. Although useful, technique barriers, such as data analysis, visualization, and scripting, make the generation of a visual data story difficult. Existing authoring tools rely on users' skills and experiences, which are usually inefficient and still difficult. In this paper, we introduce a novel visual data story generating system, Calliope, which creates visual data stories from an input spreadsheet through an automatic process and facilities the easy revision of the generated story based on an online story editor. Particularly, Calliope incorporates a new logic-oriented Monte Carlo tree search algorithm that explores the data space given by the input spreadsheet to progressively generate story pieces (i.e., data facts) and organize them in a logical order. The importance of data facts is measured based on information theory, and each data fact is visualized in a chart and captioned by an automatically generated description. We evaluate the proposed technique through three example stories, two controlled experiments, and a series of interviews with 10 domain experts. Our evaluation shows that Calliope is beneficial to efficient visual data story generation. Danqing Shi, Fuling Sun, Yang Shi 0007, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | EmoG: Supporting the Sketching of Emotional Expressions for StoryboardingabstractStoryboarding is an important ideation technique that uses sequential art to depict important scenarios of user experience. Existing data-driven support for storyboarding focuses on constructing user stories, but fail to address its benefit as a graphic narrative device. Instead, we propose to develop a data-driven design support tool that increases the expressiveness of user stories by facilitating sketching storyboards. To explore this, we focus on supporting the sketching of emotional expressions of characters in storyboards. In this paper, we present EmoG, an interactive system that generates sketches of characters with emotional expressions based on input strokes from the user. We evaluated EmoG with 21 participants in a controlled user study. The results showed that our tool has significantly better performance in usefulness, ease of use, and quality of results than the baseline system. Yang Shi 0007, Nan Cao 0001, Xiaojuan Ma, Siji Chen |
CHI | 2 |
| 2020 | CarePre: An Intelligent Clinical Decision Assistance SystemabstractClinical decision support systems are widely used to assist with medical decision making. However, clinical decision support systems typically require manually curated rules and other data that are difficult to maintain and keep up to date. Recent systems leverage advanced deep learning techniques and electronic health records to provide a more timely and precise result. Many of these techniques have been developed with a common focus on predicting upcoming medical events. However, although the prediction results from these approaches are promising, their value is limited by their lack of interpretability. To address this challenge, we introduce CarePre, an intelligent clinical decision assistance system. The system extends a state-of-the-art deep learning model to predict upcoming diagnosis events for a focal patient based on his or her historical medical records. The system includes an interactive framework together with intuitive visualizations designed to support diagnosis, treatment outcome analysis, and the interpretation of the analysis results. We demonstrate the effectiveness and usefulness of the CarePre system by reporting results from a quantities evaluation of the prediction algorithm, two case studies, and interviews with senior physicians and pulmonologists. Zhuochen Jin, Shuyuan Cui, Shunan Guo, David Gotz, Jimeng Sun 0001, Nan Cao 0001 |
ACM Trans. Comput. Heal. | 6 |
| 2020 | RCAnalyzer: visual analytics of rare categories in dynamic networksabstractA dynamic network refers to a graph structure whose nodes and/or links dynamically change over time. Existing visualization and analysis techniques focus mainly on summarizing and revealing the primary evolution patterns of the network structure. Little work focuses on detecting anomalous changing patterns in the dynamic network, the rare occurrence of which could damage the development of the entire structure. In this study, we introduce the first visual analysis system RCAnalyzer designed for detecting rare changes of sub-structures in a dynamic network. The proposed system employs a rare category detection algorithm to identify anomalous changing structures and visualize them in the context to help oracles examine the analysis results and label the data. In particular, a novel visualization is introduced, which represents the snapshots of a dynamic network in a series of connected triangular matrices. Hierarchical clustering and optimal tree cut are performed on each matrix to illustrate the detected rare change of nodes and links in the context of their surrounding structures. We evaluate our technique via a case study and a user study. The evaluation results verify the effectiveness of our system. Jiacheng Pan, Dongming Han, Fangzhou Guo, Dawei Zhou 0003, Nan Cao 0001, Jingrui He, Mingliang Xu 0001, Wei Chen 0001 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2019 | AI-Sketcher : A Deep Generative Model for Producing High-Quality SketchesabstractSketch drawings play an important role in assisting humans in communication and creative design since ancient period. This situation has motivated the development of artificial intelligence (AI) techniques for automatically generating sketches based on user input. Sketch-RNN, a sequence-to-sequence variational autoencoder (VAE) model, was developed for this purpose and known as a state-of-the-art technique. However, it suffers from limitations, including the generation of lowquality results and its incapability to support multi-class generations. To address these issues, we introduced AI-Sketcher, a deep generative model for generating high-quality multiclass sketches. Our model improves drawing quality by employing a CNN-based autoencoder to capture the positional information of each stroke at the pixel level. It also introduces an influence layer to more precisely guide the generation of each stroke by directly referring to the training data. To support multi-class sketch generation, we provided a conditional vector that can help differentiate sketches under various classes. The proposed technique was evaluated based on two large-scale sketch datasets, and results demonstrated its power in generating high-quality sketches. Nan Cao 0001, Yang Shi 0007 |
AAAI | 1 |
| 2019 | Visual Anomaly Detection in Event Sequence DataabstractAnomaly detection is a common analytical task that aims to identify rare cases that differ from the typical cases that make up the majority of a dataset. When applied to the analysis of event sequence data, the task of anomaly detection can be complex because the sequential and temporal nature of such data results in diverse definitions and flexible forms of anomalies. This, in turn, increases the difficulty in interpreting detected anomalies. In this paper, we propose an unsupervised anomaly detection algorithm based on Variational AutoEncoders (VAE) to estimate underlying normal progressions for each given sequence represented as occurrence probabilities of events along the sequence progression. Events in violation of their occurrence probability are identified as abnormal. We also introduce a visualization system, EventThread3 (ET3, to support interactive exploration and interpretations of anomalies within the context of normal sequence progressions in the dataset through comprehensive one-to-many sequence comparison. Finally, we quantitatively evaluate the performance of our anomaly detection algorithm and demonstrate the effectiveness of our system through a case study. Shunan Guo, Zhuochen Jin, Qing Chen 0001, David Gotz, Hongyuan Zha, Nan Cao 0001 |
IEEE BigData | 6 |
| 2019 | Visualizing Uncertainty and Alternatives in Event Sequence PredictionsabstractData analysts apply machine learning and statistical methods to timestamped event sequences to tackle various problems but face unique challenges when interpreting the results. Especially in event sequence prediction, it is difficult to convey uncertainty and possible alternative paths or outcomes. In this work, informed by interviews with five machine learning practitioners, we iteratively designed a novel visualization for exploring event sequence predictions of multiple records where users are able to review the most probable predictions and possible alternatives alongside uncertainty information. Through a controlled study with 18 participants, we found that users are more confident in making decisions when alternative predictions are displayed and they consider the alternatives more when deciding between two options with similar top predictions. Shunan Guo, Fan Du, Sana Malik, Eunyee Koh, Sungchul Kim, Zhicheng Liu 0001, Donghyun Kim 0007, Hongyuan Zha, Nan Cao 0001 |
CHI | 9 |
| 2019 | ADMIRING: Adversarial Multi-network MiningabstractMulti-sourced networks naturally appear in many application domains, ranging from bioinformatics, social networks, neuroscience to management. Although state-of-the-art offers rich models and algorithms to find various patterns when input networks are given, it has largely remained nascent on how vulnerable the mining results are due to the adversarial attacks. In this paper, we address the problem of attacking multi-network mining through the way of deliberately perturbing the networks to alter the mining results. The key idea of the proposed method (Admiring) is effective influence functions on the Sylvester equation defined over the input networks, which plays a central and unifying role in various multi-network mining tasks. The proposed algorithms bear two main advantages, including (1) effectiveness, being able to accurately quantify the rate of change of the mining results in response to attacks; and (2) generality, being applicable to a variety of multi-network mining tasks ( e.g., graph kernel, network alignment, cross-network node similarity) with different attacking strategies (e.g., edge/node removal, attribute alteration). Qinghai Zhou, Liangyue Li, Nan Cao 0001, Lei Ying 0001, Hanghang Tong |
ICDM | 3 |
| 2019 | Interactive Context-Aware Anomaly Detection Guided by User FeedbackabstractAutomatic anomaly detection techniques have been extensively used to support decision making in abnormal situations. However, existing approaches are limited in their capacity of effectively identifying anomalies due to the complexity of the real-world environment, the uncertainty of the data input, and the unavailability of ground truth. In this paper, we propose an interactive context-aware anomaly detection algorithm framework that incorporates human judgment in searching for anomalous regions within a large geographic environment. In specific, our framework, 1) estimates a focal region and detect anomalous situations in real time, through which the user can observe and analyze suspicious entities, 2) leverages user feedback to refine results and guide further analysis, and 3) tolerates potential fault feedback provided by the users and resignal dubious anomalous points. Based on the framework, we propose two algorithm implementations, respectively, employ Bayes’ theorem and metric learning. We demonstrate the effectiveness of the proposed framework and corresponding implementations through two controlled user studies and a case study with a domain expert. Yang Shi 0007, Maoran Xu, Rongwen Zhao, Sherry Tongshuang Wu, Nan Cao 0001 |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2019 | ACM TIST Special Issue on Visual Analyticsabstracteditorial Free Access Share on ACM TIST Special Issue on Visual Analytics Authors: Nan Cao Tongji University Tongji UniversityView Profile , Steffen Koch University of Stuttgart University of StuttgartView Profile , David Gotz University of North Carolina at Chapel Hill University of North Carolina at Chapel HillView Profile , Editor: Yingcai Wu State Key Lab of CAD8CG Zhejiang University State Key Lab of CAD8CG Zhejiang UniversityView Profile Authors Info & Claims ACM Transactions on Intelligent Systems and TechnologyVolume 10Issue 1January 2019 Article No.: 1pp 1–4https://doi.org/10.1145/3277019Published:13 December 2018Publication History 0citation404DownloadsMetricsTotal Citations0Total Downloads404Last 12 Months43Last 6 weeks10 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteView all FormatsPDF Nan Cao 0001, Steffen Koch 0001, David Gotz, Yingcai Wu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2019 | Visual Progression Analysis of Event Sequence DataabstractEvent sequence data is common to a broad range of application domains, from security to health care to scholarly communication. This form of data captures information about the progression of events for an individual entity (e.g., a computer network device; a patient; an author) in the form of a series of time-stamped observations. Moreover, each event is associated with an event type (e.g., a computer login attempt, or a hospital discharge). Analyses of event sequence data have been shown to help reveal important temporal patterns, such as clinical paths resulting in improved outcomes, or an understanding of common career trajectories for scholars. Moreover, recent research has demonstrated a variety of techniques designed to overcome methodological challenges such as large volumes of data and high dimensionality. However, the effective identification and analysis of latent stages of progression, which can allow for variation within different but similarly evolving event sequences, remain a significant challenge with important real-world motivations. In this paper, we propose an unsupervised stage analysis algorithm to identify semantically meaningful progression stages as well as the critical events which help define those stages. The algorithm follows three key steps: (1) event representation estimation, (2) event sequence warping and alignment, and (3) sequence segmentation. We also present a novel visualization system, ET2, which interactively illustrates the results of the stage analysis algorithm to help reveal evolution patterns across stages. Finally, we report three forms of evaluation for ET2: (1) case studies with two real-world datasets, (2) interviews with domain expert users, and (3) a performance evaluation on the progression analysis algorithm and the visualization design. Shunan Guo, Zhuochen Jin, David Gotz, Fan Du, Hongyuan Zha, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2019 | A Semantic-Based Method for Visualizing Large Image CollectionsabstractInteractive visualization of large image collections is important and useful in many applications, such as personal album management and user profiling on images. However, most prior studies focus on using low-level visual features of images, such as texture and color histogram, to create visualizations without considering the more important semantic information embedded in images. This paper proposes a novel visual analytic system to analyze images in a semantic-aware manner. The system mainly comprises two components: a semantic information extractor and a visual layout generator. The semantic information extractor employs an image captioning technique based on convolutional neural network (CNN) to produce descriptive captions for images, which can be transformed into semantic keywords. The layout generator employs a novel co-embedding model to project images and the associated semantic keywords to the same 2D space. Inspired by the galaxy metaphor, we further turn the projected 2D space to a galaxy visualization of images, in which semantic keywords and images are visually encoded as stars and planets. Our system naturally supports multi-scale visualization and navigation, in which users can immediately see a semantic overview of an image collection and drill down for detailed inspection of a certain group of images. Users can iteratively refine the visual layout by integrating their domain knowledge into the co-embedding process. Two task-based evaluations are conducted to demonstrate the effectiveness of our system. Xiao Xie, Xiwen Cai, Junpei Zhou, Nan Cao 0001, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | EnsembleLens: Ensemble-based Visual Exploration of Anomaly Detection Algorithms with Multidimensional DataabstractThe results of anomaly detection are sensitive to the choice of detection algorithms as they are specialized for different properties of data, especially for multidimensional data. Thus, it is vital to select the algorithm appropriately. To systematically select the algorithms, ensemble analysis techniques have been developed to support the assembly and comparison of heterogeneous algorithms. However, challenges remain due to the absence of the ground truth, interpretation, or evaluation of these anomaly detectors. In this paper, we present a visual analytics system named EnsembleLens that evaluates anomaly detection algorithms based on the ensemble analysis process. The system visualizes the ensemble processes and results by a set of novel visual designs and multiple coordinated contextual views to meet the requirements of correlation analysis, assessment and reasoning of anomaly detection algorithms. We also introduce an interactive analysis workflow that dynamically produces contextualized and interpretable data summaries that allow further refinements of exploration results based on user feedback. We demonstrate the effectiveness of EnsembleLens through a quantitative evaluation, three case studies with real-world data and interviews with two domain experts. Meng Xia 0002, Xing Mu, Yun Wang 0012, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | A Visual Analytics Approach for Equipment Condition Monitoring in Smart Factories of Process IndustryabstractMonitoring equipment conditions is of great value in manufacturing, which can not only reduce unplanned downtime by early detecting anomalies of equipment but also avoid unnecessary routine maintenance. With the coming era of Industry 4.0 (or industrial internet), more and more assets and machines in plants are equipped with various sensors and information systems, which brings an unprecedented opportunity to capture large-scale and fine-grained data for effective on-line equipment condition monitoring. However, due to the lack of systematic methods, analysts still find it challenging to carry out efficient analyses and extract valuable information from the mass volume of data collected, especially for process industry (e.g., a petrochemical plant) with complex manufacturing procedures. In this paper, we report the design and implementation of an interactive visual analytics system, which helps managers and operators at manufacturing sites leverage their domain knowledge and apply substantial human judgements to guide the automated analytical approaches, thus generating understandable and trustable results for real-world applications. Our system integrates advanced analytical algorithms (e.g., Gaussian mixture model with a Bayesian framework) and intuitive visualization designs to provide a comprehensive and adaptive semi-supervised solution to equipment condition monitoring. The example use cases based on a real-world manufacturing dataset and interviews with domain experts demonstrate the effectiveness of our system. Yixian Zheng, Nan Cao 0001 |
PacificVis | 5 |
| 2018 | Local Partition in Rich GraphsabstractLocal graph partitioning is a key graph mining tool that allows researchers to identify small groups of interrelated nodes (e.g., people) and their connective edges (e.g., interactions). As local graph partitioning focuses primarily on the graph structure (vertices and edges), it often fails to consider the additional information contained in the attributes. We propose a scalable algorithm to improve local graph partitioning by taking into account both the graph structure and attributes. Experimental results show that our proposed AttriPart algorithm finds up to 1.6× denser local partitions, while running approximately 43× faster than traditional local partitioning techniques (PageRank-Nibble). Scott Freitas, Nan Cao 0001, Yinglong Xia, Polo Chau, Hanghang Tong |
IEEE BigData | 2 |
| 2018 | AURORA: Auditing PageRank on Large GraphsabstractRanking on large-scale graphs plays a fundamental role in many high-impact application domains, ranging from information retrieval, recommender systems, sports team management, biology to neuroscience and many more. PageRank, together with many of its random walk based variants, has become one of the most well-known and widely used algorithms, due to its mathematical elegance and the superior performance across a variety of application domains. Important as it might be, state-of-the-art lacks an intuitive way to explain the ranking results by PageRank (or its variants), e.g., why it thinks the returned top-k webpages are the most important ones in the entire graph; why it gives a higher rank to actor John than actor Smith in terms of their relevance w.r.t. a particular movie? In order to answer these questions, this paper proposes a paradigm shift for PageRank, from identifying which nodes are most important to understanding why the ranking algorithm gives a particular ranking result. We formally define the PageRank auditing problem, whose central idea is to identify a set of key graph elements (e.g., edges, nodes, subgraphs) with the highest influence on the ranking results. We formulate it as an opti-mization problem and propose a family of effective and scalable algorithms (Aurora) to solve it. Our algorithms measure the influence of graph elements and incrementally select influential elements w.r.t. their gradients over the ranking results. We perform extensive empirical evaluations on real-world datasets, which demonstrate that the proposed methods (Aurora) provide intuitive explanations with a linear scalability. Jian Kang 0008, Nan Cao 0001, Yinglong Xia, Wei Fan 0001, Hanghang Tong |
IEEE BigData | 3 |
| 2018 | ECGLens: Interactive Visual Exploration of Large Scale ECG Data for Arrhythmia DetectionabstractThe Electrocardiogram (ECG) is commonly used to detect arrhythmias. Traditionally, a single ECG observation is used for diagnosis, making it difficult to detect irregular arrhythmias. Recent technology developments, however, have made it cost-effective to collect large amounts of raw ECG data over time. This promises to improve diagnosis accuracy, but the large data volume presents new challenges for cardiologists. This paper introduces ECGLens, an interactive system for arrhythmia detection and analysis using large-scale ECG data. Our system integrates an automatic heartbeat classification algorithm based on convolutional neural network, an outlier detection algorithm, and a set of rich interaction techniques. We also introduce A-glyph, a novel glyph designed to improve the readability and comparison of ECG signals. We report results from a comprehensive user study showing that A-glyph improves the efficiency in arrhythmia detection, and demonstrate the effectiveness of ECGLens in arrhythmia detection through two expert interviews. Shunan Guo, Nan Cao 0001, David Gotz, Aiwen Xu, Huamin Qu, Zhenjie Yao 0001, Yixin Chen 0001 |
CHI | 3 |
| 2018 | X-Rank: Explainable Ranking in Complex Multi-Layered NetworksabstractIn this paper we present a web-based prototype for an explainable ranking algorithm in multi-layered networks, incorporating both network topology and knowledge information. While traditional ranking algorithms such as PageRank and HITS are important tools for exploring the underlying structure of networks, they have two fundamental limitations in their efforts to generate high accuracy rankings. First, they are primarily focused on network topology, leaving out additional sources of information (e.g. attributes, knowledge). Secondly, most algorithms do not provide explanations to the end-users on why the algorithm gives the specific ranking results, hindering the usability of the ranking information. We developed Xrank, an explainable ranking tool, to address these drawbacks. Empirical results indicate that our explainable ranking method not only improves ranking accuracy, but facilitates user understanding of the ranking by exploring the top influential elements in multi-layered networks. The web-based prototype (Xrank: http://www.x-rank.net) is currently online - we believe it will assist both researchers and practitioners looking to explore and exploit multi-layered network data. Jian Kang 0008, Scott Freitas, Haichao Yu, Yinglong Xia, Nan Cao 0001, Hanghang Tong |
CIKM | 5 |
| 2018 | Extra: explaining team recommendation in networksabstractState-of-the-art in network science of teams offers effective recommendation methods to answer questions like who is the best replacement, what is the best team expansion strategy, but lacks intuitive ways to explain why the optimization algorithm gives the specific recommendation for a given team optimization scenario. To tackle this problem, we develop an interactive prototype system, Extra, as the first step towards addressing such a sense-making challenge, through the lens of the underlying network where teams embed, to explain the team recommendation results. The main advantages are (1) Algorithm efficacy: we propose an effective and fast algorithm to explain random walk graph kernel, the central technique for networked team recommendation; (2) Intuitive visual explanation: we present intuitive visual analysis of the recommendation results, which can help users better understand the rationality of the underlying team recommendation algorithm. Qinghai Zhou, Liangyue Li, Nan Cao 0001, Norbou Buchler, Hanghang Tong |
RecSys | 3 |
| 2018 | Anomaly detection in spatiotemporal data via regularized non-negative tensor analysis
Chaoguang Lin, Qiuhan Zhu, Shunan Guo, Zhuochen Jin, Yu-Ru Lin, Nan Cao 0001 |
Data Min. Knowl. Discov. | 6 |
| 2018 | Guest Editorial: Special Issue on Interactive Visual Analysis of Human and Crowd BehaviorsabstractThe analysis of human behaviors has impacted many social and commercial domains. How could interactive visual analytic systems be used to further provide behavioral insights? This editorial introduction features emerging research trend related to this question. The four articles accepted for this special issue represent recent progress: they identify research challenges arising from analysis of human and crowd behaviors, and present novel methods in visual analysis to address those challenges and help make behavioral data more useful. Yu-Ru Lin, Nan Cao 0001 |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2018 | Voila: Visual Anomaly Detection and Monitoring with Streaming Spatiotemporal DataabstractThe increasing availability of spatiotemporal data continuously collected from various sources provides new opportunities for a timely understanding of the data in their spatial and temporal context. Finding abnormal patterns in such data poses significant challenges. Given that there is often no clear boundary between normal and abnormal patterns, existing solutions are limited in their capacity of identifying anomalies in large, dynamic and heterogeneous data, interpreting anomalies in their multifaceted, spatiotemporal context, and allowing users to provide feedback in the analysis loop. In this work, we introduce a unified visual interactive system and framework, Voila, for interactively detecting anomalies in spatiotemporal data collected from a streaming data source. The system is designed to meet two requirements in real-world applications, i.e., online monitoring and interactivity. We propose a novel tensor-based anomaly analysis algorithm with visualization and interaction design that dynamically produces contextualized, interpretable data summaries and allows for interactively ranking anomalous patterns based on user input. Using the "smart city" as an example scenario, we demonstrate the effectiveness of the proposed framework through quantitative evaluation and qualitative case studies. Nan Cao 0001, Chaoguang Lin, Qiuhan Zhu, Yu-Ru Lin, Xian Teng, Xidao Wen |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | EventThread: Visual Summarization and Stage Analysis of Event Sequence DataabstractEvent sequence data such as electronic health records, a person's academic records, or car service records, are ordered series of events which have occurred over a period of time. Analyzing collections of event sequences can reveal common or semantically important sequential patterns. For example, event sequence analysis might reveal frequently used care plans for treating a disease, typical publishing patterns of professors, and the patterns of service that result in a well-maintained car. It is challenging, however, to visually explore large numbers of event sequences, or sequences with large numbers of event types. Existing methods focus on extracting explicitly matching patterns of events using statistical analysis to create stages of event progression over time. However, these methods fail to capture latent clusters of similar but not identical evolutions of event sequences. In this paper, we introduce a novel visualization system named EventThread which clusters event sequences into threads based on tensor analysis and visualizes the latent stage categories and evolution patterns by interactively grouping the threads by similarity into time-specific clusters. We demonstrate the effectiveness of EventThread through usage scenarios in three different application domains and via interviews with an expert user. Shunan Guo, Rongwen Zhao, David Gotz, Hongyuan Zha, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | RCLens: Interactive Rare Category Exploration and IdentificationabstractRare category identification is an important task in many application domains, ranging from network security, to financial fraud detection, to personalized medicine. These are all applications which require the discovery and characterization of sets of rare but structurally-similar data entities which are obscured within a larger but structurally different dataset. This paper introduces RCLens, a visual analytics system designed to support user-guided rare category exploration and identification. RCLens adopts a novel active learning-based algorithm to iteratively identify more accurate rare categories in response to user-provided feedback. The algorithm is tightly integrated with an interactive visualization-based interface which supports a novel and effective workflow for rare category identification. This paper (1) defines RCLens' underlying active-learning algorithm; (2) describes the visualization and interaction designs, including a discussion of how the designs support user-guided rare category identification; and (3) presents results from an evaluation demonstrating RCLens' ability to support the rare category identification process. Hanfei Lin, David Gotz, Fan Du, Jingrui He, Nan Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | StreamExplorer: A Multi-Stage System for Visually Exploring Events in Social StreamsabstractAnalyzing social streams is important for many applications, such as crisis management. However, the considerable diversity, increasing volume, and high dynamics of social streams of large events continue to be significant challenges that must be overcome to ensure effective exploration. We propose a novel framework by which to handle complex social streams on a budget PC. This framework features two components: 1) an online method to detect important time periods (i.e., subevents), and 2) a tailored GPU-assisted Self-Organizing Map (SOM) method, which clusters the tweets of subevents stably and efficiently. Based on the framework, we present StreamExplorer to facilitate the visual analysis, tracking, and comparison of a social stream at three levels. At a macroscopic level, StreamExplorer uses a new glyph-based timeline visualization, which presents a quick multi-faceted overview of the ebb and flow of a social stream. At a mesoscopic level, a map visualization is employed to visually summarize the social stream from either a topical or geographical aspect. At a microscopic level, users can employ interactive lenses to visually examine and explore the social stream from different perspectives. Two case studies and a task-based evaluation are used to demonstrate the effectiveness and usefulness of StreamExplorer.Analyzing social streams is important for many applications, such as crisis management. However, the considerable diversity, increasing volume, and high dynamics of social streams of large events continue to be significant challenges that must be overcome to ensure effective exploration. We propose a novel framework by which to handle complex social streams on a budget PC. This framework features two components: 1) an online method to detect important time periods (i.e., subevents), and 2) a tailored GPU-assisted Self-Organizing Map (SOM) method, which clusters the tweets of subevents stably and efficiently. Based on the framework, we present StreamExplorer to facilitate the visual analysis, tracking, and comparison of a social stream at three levels. At a macroscopic level, StreamExplorer uses a new glyph-based timeline visualization, which presents a quick multi-faceted overview of the ebb and flow of a social stream. At a mesoscopic level, a map visualization is employed to visually summarize the social stream from either a topical or geographical aspect. At a microscopic level, users can employ interactive lenses to visually examine and explore the social stream from different perspectives. Two case studies and a task-based evaluation are used to demonstrate the effectiveness and usefulness of StreamExplorer. Yingcai Wu, Chen Zhu-Tian, Guodao Sun, Xiao Xie, Nan Cao 0001, Shixia Liu, Weiwei Cui 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Steering data quality with visual analytics: The complexity challengeabstractData quality management, especially data cleansing, has been extensively studied for many years in the areas of data management and visual analytics. In the paper, we first review and explore the relevant work from the research areas of data management, visual analytics and human-computer interaction. Then for different types of data such as multimedia data, textual data, trajectory data, and graph data, we summarize the common methods for improving data quality by leveraging data cleansing techniques at different analysis stages. Based on a thorough analysis, we propose a general visual analytics framework for interactively cleansing data. Finally, the challenges and opportunities are analyzed and discussed in the context of data and humans. Shixia Liu, Gennady L. Andrienko, Yingcai Wu, Nan Cao 0001, Liu Jiang, Conglei Shi, Yu-Shuen Wang, Seok-Hee Hong 0001 |
Vis. Informatics | 4 |
| 2017 | MobiSeg: Interactive region segmentation using heterogeneous mobility dataabstractWith the acceleration of urbanization and modern civilization, more and more complex regions are formed in urban area. Although understanding these regions could provide huge insights to facilitate valuable applications for urban planning and business intelligence, few methods have been developed to effectively capture the rapid transformation of urban regions. In recent years, the widely applied location-acquisition technologies offer a more effective way to capture the dynamics of a city through analyzing people's movement activities based on mobility data. However, several challenges exist, including data sparsity and difficulties in result understanding and validation. To tackle these challenges, in this paper, we propose MobiSeg, an interactive visual analytics system, which supports the exploration of people's movement activities to segment the urban area into regions sharing similar activity patterns. A joint analysis is conducted on three types of heterogeneous mobility data (i.e., taxi trajectories, metro passenger RFID card data, and telco data), which can complement each other and provide a full picture of people's activities in a region. In addition, advanced analytical algorithms (e.g., non-negative matrix factorization (NMF) based method to capture latent activity patterns, as well as metric learning to calibrate and supervise the underlying analysis) and novel visualization designs are integrated into our system to provide a comprehensive solution to region segmentation in urban areas. We demonstrate the effectiveness of our system via case studies with real-world datasets and qualitative interviews with domain experts. Yixian Zheng, Nan Cao 0001, Haipeng Zeng, Bing Ni, Huamin Qu, Lionel M. Ni |
PacificVis | 3 |
| 2017 | iSphere: Focus+Context Sphere Visualization for Interactive Large Graph ExplorationabstractInteractive exploration plays a critical role in large graph visualization. Existing techniques, such as zoom-and-pan on a 2D plane and hyperbolic browser facilitate large graph exploration by showing both the details of a focal area and its surrounding context that guides the exploration process. However, existing techniques for large graph exploration are limited in either providing too little context or presenting graphs with too much distortion. In this paper, we propose a novel focus+context technique, iSphere, to address the limitation. iSphere maps a large graph onto a Riemann Sphere that better preserves graph structures and shows greater context information. We conduct extensive experiment studies on different graph exploration tasks under various conditions. The results show that iSphere performs the best in task completion time compared to the baseline techniques in link and path exploration tasks. This research also contributes to understanding large graph exploration on small screens. Fan Du, Nan Cao 0001, Yu-Ru Lin, Hanghang Tong |
CHI | 2 |
| 2017 | Rapid Analysis of Network ConnectivityabstractThis research focuses on accelerating the computational time of two base network algorithms (k-simple shortest paths and minimum spanning tree for a subset of nodes)---cornerstones behind a variety of network connectivity mining tasks---with the goal of rapidly finding networkpathways andtrees using a set of user-specific query nodes. To facilitate this process we utilize: (1) multi-threaded algorithm variations, (2) network re-use for subsequent queries and (3) a novel algorithm, Key Neighboring Vertices (KNV), to reduce the network search space. The proposed KNV algorithm serves a dual purpose: (a) to reduce the computation time for algorithmic analysis and (b) to identify key vertices in the network (\textit ). Empirical results indicate this combination of techniques significantly improves the baseline performance of both algorithms. We have also developed a web platform utilizing the proposed network algorithms to enable researchers and practitioners to both visualize and interact with their datasets (PathFinder: http://www.path-finder.io. Scott Freitas, Hanghang Tong, Nan Cao 0001, Yinglong Xia |
CIKM | 3 |
| 2017 | Video-based Evanescent, Anonymous, Asynchronous Social Interaction: Motivation and Adaption to MediumabstractDanmaku is an emerging socio-digital media paradigm that puts anonymous, asynchronous user-generated scrolling comments on videos. (How) can danmaku afford the illusion and realization of social interactions, if at all possible given its interactional incoherence nature? To answer this question, we collect Chinese danmaku users' reflection on their motivations to use this social service and explore the actual practices that meet the needs. According to a preliminary danmaku usage survey, users consider it as an information seeking and emotion venting channel. Through archival analysis of real-world data, we find that danmaku commentaries are relatively short, video-centric, saturated with emotions, and similar in syntactic and semantic features. Users have developed a set of mechanisms adapted to the medium, to leverage such text-based messages to foster interpersonal and hyperpersonal communication for sharing of facts, thoughts, and feelings. Xiaojuan Ma, Nan Cao 0001 |
CSCW | 2 |
| 2017 | FIRST: Fast Interactive Attributed Subgraph MatchingabstractAttributed subgraph matching is a powerful tool for explorative mining of large attributed networks. In many applications (e.g., network science of teams, intelligence analysis, finance informatics), the user might not know what exactly s/he is looking for, and thus require the user to constantly revise the initial query graph based on what s/he finds from the current matching results. A major bottleneck in such an interactive matching scenario is the efficiency, as simply rerunning the matching algorithm on the revised query graph is computationally prohibitive. In this paper, we propose a family of effective and efficient algorithms (FIRST) to support interactive attributed subgraph matching. There are two key ideas behind the proposed methods. The first is to recast the attributed subgraph matching problem as a cross-network node similarity problem, whose major computation lies in solving a Sylvester equation for the query graph and the underlying data graph. The second key idea is to explore the smoothness between the initial and revised queries, which allows us to solve the new/updated Sylvester equation incrementally, without re-solving it from scratch. Experimental results show that our method can achieve (1) up to 16x speed-up when applying on networks with 6M$+$ nodes; (2) preserving more than 90% accuracy compared with existing methods; and (3) scales linearly with respect to the size of the data graph. Boxin Du, Nan Cao 0001, Hanghang Tong |
KDD | 3 |
| 2017 | Is the Whole Greater Than the Sum of Its Parts?abstractThe PART-WHOLE relationship routinely finds itself in many disciplines, ranging from collaborative teams, crowdsourcing, autonomous systems to networked systems. From the algorithmic perspective, the existing work has primarily focused on predicting the outcomes of the whole and parts, by either separate models or linear joint models, which assume the outcome of the parts has a linear and independent effect on the outcome of the whole. In this paper, we propose a joint predictive method named PAROLE to simultaneously and mutually predict the part and whole outcomes. The proposed method offers two distinct advantages over the existing work. First (Model Generality), we formulate joint PART-WHOLE outcome prediction as a generic optimization problem, which is able to encode a variety of complex relationships between the outcome of the whole and parts, beyond the linear independence assumption. Second (Algorithm Efficacy), we propose an effective and efficient block coordinate descent algorithm, which is able to find the coordinate-wise optimum with a linear complexity in both time and space. Extensive empirical evaluations on real-world datasets demonstrate that the proposed PAROLE (1) leads to consistent prediction performance improvement by modeling the non-linear part-whole relationship as well as part-part interdependency, and (2) scales linearly in terms of the size of the training dataset. Liangyue Li, Hanghang Tong, Yong Wang 0021, Conglei Shi, Nan Cao 0001, Norbou Buchler |
KDD | 5 |
| 2017 | Discovering rare categories from graph streams
Dawei Zhou 0003, Arun Karthikeyan, Kangyang Wang, Nan Cao 0001, Jingrui He |
Data Min. Knowl. Discov. | 4 |
| 2017 | Adaptive Contextualization Methods for Combating Selection Bias during High-Dimensional VisualizationabstractLarge and high-dimensional real-world datasets are being gathered across a wide range of application disciplines to enable data-driven decision making. Interactive data visualization can play a critical role in allowing domain experts to select and analyze data from these large collections. However, there is a critical mismatch between the very large number of dimensions in complex real-world datasets and the much smaller number of dimensions that can be concurrently visualized using modern techniques. This gap in dimensionality can result in high levels of selection bias that go unnoticed by users. The bias can in turn threaten the very validity of any subsequent insights. This article describes Adaptive Contextualization (AC), a novel approach to interactive visual data selection that is specifically designed to combat the invisible introduction of selection bias. The AC approach (1) monitors and models a user’s visual data selection activity, (2) computes metrics over that model to quantify the amount of selection bias after each step, (3) visualizes the metric results, and (4) provides interactive tools that help users assess and avoid bias-related problems. This article expands on an earlier article presented at ACM IUI 2016 [16] by providing a more detailed review of the AC methodology and additional evaluation results. David Gotz, Shun Sun, Nan Cao 0001, Rita Kundu, Anne-Marie Meyer |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2017 | Enhancing Team Composition in Professional Networks: Problem Definitions and Fast SolutionsabstractIn this paper, we study ways to enhance the composition of teams based on new requirements in a collaborative environment. We focus on recommending team members who can maintain the team's performance by minimizing changes to the team's skills and social structure. Our recommendations are based on computing team-level similarity, which includes skill similarity, structural similarity as well as the synergy between the two. Current heuristic approaches are one-dimensional and not comprehensive, as they consider the two aspects independently. To formalize team-level similarity, we adopt the notion of graph kernel of attributed graphs to encompass the two aspects and their interaction. To tackle the computational challenges, we propose a family of fast algorithms by (a) designing effective pruning strategies, and (b) exploring the smoothness between the existing and the new team structures. Extensive empirical evaluations on real world datasets validate the effectiveness and efficiency of our algorithms. Liangyue Li, Hanghang Tong, Nan Cao 0001, Kate Ehrlich, Yu-Ru Lin, Norbou Buchler |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2017 | Evaluation of Graph Sampling: A Visualization PerspectiveabstractGraph sampling is frequently used to address scalability issues when analyzing large graphs. Many algorithms have been proposed to sample graphs, and the performance of these algorithms has been quantified through metrics based on graph structural properties preserved by the sampling: degree distribution, clustering coefficient, and others. However, a perspective that is missing is the impact of these sampling strategies on the resultant visualizations. In this paper, we present the results of three user studies that investigate how sampling strategies influence node-link visualizations of graphs. In particular, five sampling strategies widely used in the graph mining literature are tested to determine how well they preserve visual features in node-link diagrams. Our results show that depending on the sampling strategy used different visual features are preserved. These results provide a complimentary view to metric evaluations conducted in the graph mining literature and provide an impetus to conduct future visualization studies. Nan Cao 0001, Daniel Archambault, Qiaomu Shen, Huamin Qu, Weiwei Cui 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | Interactive visual co-cluster analysis of bipartite graphsabstractA bipartite graph models the relation between two different types of entities. It is applicable, for example, to describe persons' affiliations to different social groups or their association with subjects such as topics of interest. In these applications, it is important to understand the connectivity patterns among the entities in the bipartite graph. For the example of a bipartite relation between persons and their topics of interest, people may form groups based on their common interests, and the topics also can be grouped or categorized based on the interested audiences. Co-clustering methods can identify such connectivity patterns and find clusters within the two types of entities simultaneously. In this paper, we propose an interactive visualization design that incorporates co-clustering methods to facilitate the identification of node clusters formed by their common connections in a bipartite graph. Besides highlighting the automatically detected node clusters and the connections among them, the visual interface also provides visual cues for evaluating the homogeneity of the bipartite connections in a cluster, identifying potential outliers, and analyzing the correlation of node attributes with the cluster structure. The interactive visual interface allows users to flexibly adjust the node grouping to incorporate their prior knowledge of the domain, either by direct manipulation (i.e., splitting and merging the clusters), or by providing explicit feedback on the cluster quality, based on which the system will learn a parametrization of the co-clustering algorithm to better align with the users' notion of node similarity. To demonstrate the utility of the system, we present two example usage scenarios on real world datasets. Nan Cao 0001, Huamin Qu, John T. Stasko |
PacificVis | 2 |
| 2016 | TelcoFlow: Visual exploration of collective behaviors based on telco dataabstractCollective behavior is an important concept defined to capture behavioral patterns emerged among the crowd spontaneously. In social science, people's behaviors can be regarded as temporal transitions between a set of typical states (e.g., home and work) which are always associated with certain locations. This fact leads to an interesting research topic in developing ways to explore people's collective behavior patterns through movement analysis, which is our focus in this paper. In recent years, massive volumes of spatiotemporal data generated by mobile phones, called telco data, bring an unprecedented opportunity to study collective behaviors in terms of large coverage and fine-grained resolution. However, distilling valuable collective behavior patterns from the large scale of telco data is not an easy task. The challenge is rooted in two aspects, including the data uncertainty as well as the lack of methods to characterize, compare and understand dynamic crowd behaviors, which triggers the use of visual analytics to take full advantage of machines' computational power as well as human's domain knowledge and cognitive abilities. In this paper, we propose TelcoFlow, a comprehensive visual analytics system which incorporates advanced quantitative analyses (e.g., statebased behavior model) and intuitive visualizations (e.g., an extended flow view embedded with state glyphs) to support an efficient and in-depth analysis of collective behaviors based on telco data. Case studies with a real-world dataset and expert interviews are carried out to demonstrate the effectiveness of our system for analysts to gain insights into collective behaviors and facilitate various analytical tasks. Yixian Zheng, Haipeng Zeng, Nan Cao 0001, Huamin Qu, Mingxuan Yuan, Lionel M. Ni |
IEEE BigData | 4 |
| 2016 | TEAMOPT: Interactive Team Optimization in Big NetworksabstractThe science of team science is a rapidly emerging research field that studies strategies to understand and enhance the process and outcomes of collaborative, team-based research. An interesting research question we address in this work is how to maintain and optimize the team performance should certain changes happen to the team. In particular, we take the network approach to understanding the teams and consider optimizing the teams with several operations (e.g., replacement, expansion, shrinkage). We develop TEAMOPT, a system to assist users in optimizing the team performance interactively to support the changes to a team. TEAMOPT takes as input a large network of individuals (e.g., co-author network of researchers) and is able to assist users in assembling a team with specific requirements and optimizing the team in response to the changes made to the team. It is effective in finding the best candidates, and interactive with users' feedback in the loop. The system is developed using HTML5, JavaScript, D3.js (front-end) and Python CGI (back-end). A prototype system is already deployed. We will invite the audience to experiment with our TEAMOPT in terms of its effectiveness, efficiency and applicability to various scenarios. Liangyue Li, Hanghang Tong, Nan Cao 0001, Kate Ehrlich, Yu-Ru Lin, Norbou Buchler |
CIKM | 3 |
| 2016 | Adaptive Contextualization: Combating Bias During High-Dimensional Visualization and Data SelectionabstractLarge and high-dimensional real-world datasets are being gathered across a wide range of application disciplines to enable data-driven decision making. Interactive data visualization can play a critical role in allowing domain experts to select and analyze data from these large collections. However, there is a critical mismatch between the very large number of dimensions in complex real-world datasets and the much smaller number of dimensions that can be concurrently visualized using modern techniques. This gap in dimensionality can result in high levels of selection bias that go unnoticed by users. The bias can in turn threaten the very validity of any subsequent insights. In this paper, we present Adaptive Contextualization (AC), a novel approach to interactive visual data selection that is specifically designed to combat the invisible introduction of selection bias. Our approach (1) monitors and models a user's visual data selection activity, (2) computes metrics over that model to quantify the amount of selection bias after each step, (3) visualizes the metric results, and (4) provides interactive tools that help users assess and avoid bias-related problems. We also share results from a user study which demonstrate the effectiveness of our technique. David Gotz, Shun Sun, Nan Cao 0001 |
IUI | 3 |
| 2016 | Guest Editorial: Visual Analytics in Multimedia - Opportunities and Research ChallengesabstractThe ten papers in this special section are devoted to the topic of visual analytics, an emerging research direction that focuses on data exploration and analysis with a seamless integration of interaction, visualization, and analysis. Nan Cao 0001, Yingcai Wu, David Gotz, D. Kiem, Y.-P. Tan |
IEEE Trans. Multim. | 1 |
| 2016 | A Survey on Visual Analytics of Social Media DataabstractThe unprecedented availability of social media data offers substantial opportunities for data owners, system operators, solution providers, and end users to explore and understand social dynamics. However, the exponential growth in the volume, velocity, and variability of social media data prevents people from fully utilizing such data. Visual analytics, which is an emerging research direction, has received considerable attention in recent years. Many visual analytics methods have been proposed across disciplines to understand large-scale structured and unstructured social media data. This objective, however, also poses significant challenges for researchers to obtain a comprehensive picture of the area, understand research challenges, and develop new techniques. In this paper, we present a comprehensive survey to characterize this fast-growing area and summarize the state-of-the-art techniques for analyzing social media data. In particular, we classify existing techniques into two categories: gathering information and understanding user behaviors. We aim to provide a clear overview of the research area through the established taxonomy. We then explore the design space and identify the research trends. Finally, we discuss challenges and open questions for future studies. Yingcai Wu, Nan Cao 0001, David Gotz, Yap-Peng Tan, Daniel A. Keim |
IEEE Trans. Multim. | 2 |
| 2016 | UnTangle Map: Visual Analysis of Probabilistic Multi-Label DataabstractData with multiple probabilistic labels are common in many situations. For example, a movie may be associated with multiple genres with different levels of confidence. Despite their ubiquity, the problem of visualizing probabilistic labels has not been adequately addressed. Existing approaches often either discard the probabilistic information, or map the data to a low-dimensional subspace where their associations with original labels are obscured. In this paper, we propose a novel visual technique, UnTangle Map, for visualizing probabilistic multi-labels. In our proposed visualization, data items are placed inside a web of connected triangles, with labels assigned to the triangle vertices such that nearby labels are more relevant to each other. The positions of the data items are determined based on the probabilistic associations between items and labels. UnTangle Map provides both (a) an automatic label placement algorithm, and (b) adaptive interactions that allow users to control the label positioning for different information needs. Our work makes a unique contribution by providing an effective way to investigate the relationship between data items and their probabilistic labels, as well as the relationships among labels. Our user study suggests that the visualization effectively helps users discover emergent patterns and compare the nuances of probabilistic information in the data labels. Nan Cao 0001, Yu-Ru Lin, David Gotz |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | TargetVue: Visual Analysis of Anomalous User Behaviors in Online Communication SystemsabstractUsers with anomalous behaviors in online communication systems (e.g. email and social medial platforms) are potential threats to society. Automated anomaly detection based on advanced machine learning techniques has been developed to combat this issue; challenges remain, though, due to the difficulty of obtaining proper ground truth for model training and evaluation. Therefore, substantial human judgment on the automated analysis results is often required to better adjust the performance of anomaly detection. Unfortunately, techniques that allow users to understand the analysis results more efficiently, to make a confident judgment about anomalies, and to explore data in their context, are still lacking. In this paper, we propose a novel visual analysis system, TargetVue, which detects anomalous users via an unsupervised learning model and visualizes the behaviors of suspicious users in behavior-rich context through novel visualization designs and multiple coordinated contextual views. Particularly, TargetVue incorporates three new ego-centric glyphs to visually summarize a user's behaviors which effectively present the user's communication activities, features, and social interactions. An efficient layout method is proposed to place these glyphs on a triangle grid, which captures similarities among users and facilitates comparisons of behaviors of different users. We demonstrate the power of TargetVue through its application in a social bot detection challenge using Twitter data, a case study based on email records, and an interview with expert users. Our evaluation shows that TargetVue is beneficial to the detection of users with anomalous communication behaviors. Nan Cao 0001, Conglei Shi, Wan-Yi Sabrina Lin, Jie Lu 0002, Yu-Ru Lin, Ching-Yung Lin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | g-Miner: Interactive Visual Group Mining on Multivariate GraphsabstractWith the rapid growth of rich network data available through various sources such as social media and digital archives,there is a growing interest in more powerful network visual analysis tools and methods. The rich information about the network nodes and links can be represented as multivariate graphs, in which the nodes are accompanied with attributes to represent the properties of individual nodes. An important task often encountered in multivariate network analysis is to uncover link structure with groups, e.g., to understand why a person fits a specific job or certain role in a social group well.The task usually involves complex considerations including specific requirement of node attributes and link structure, and hence a fully automatic solution is typically not satisfactory.In this work, we identify the design challenges for min-ing groups with complex criteria and present an interactive system, "g-Miner," that enables visual mining of groups on multivariate graph data. We demonstrate the effectiveness of our system through case study and in-depth expert inter-views. This work contributes to understanding the design of systems for leveraging users' knowledge progressively with algorithmic capacity for tackling massive heterogeneous information. Nan Cao 0001, Yu-Ru Lin, Liangyue Li, Hanghang Tong |
CHI | 1 |
| 2015 | Trajectory Bundling for Animated TransitionsabstractAnimated transition has been a popular design choice for smoothly switching between different visualization views or layouts, in which movement trajectories are created as cues for tracking objects during location shifting. Tracking moving objects, however, becomes difficult when their movement paths overlap or the number of tracking targets increases. We propose a novel design to facilitate tracking moving objects in animated transitions. Instead of simply animating an object along a straight line, we create "bundled" movement trajectories for a group of objects that have spatial proximity and share similar moving directions. To study the effect of bundled trajectories, we untangle variations due to different aspects of tracking complexity in a comprehensive controlled user study. The results indicate that using bundled trajectories is particularly effective when tracking more targets (six vs. three targets) or when the object movement involves a high degree of occlusion or deformation. Based on the study, we discuss the advantages and limitations of the new technique, as well as provide design implications. Fan Du, Nan Cao 0001, Jian Zhao 0010, Yu-Ru Lin |
CHI | 2 |
| 2015 | Rare Category Detection on Time-Evolving GraphsabstractRare category detection(RCD) is an important topicin data mining, focusing on identifying the initial examples fromrare classes in imbalanced data sets. This problem becomes more challenging when the data is presented as time-evolving graphs, as used in synthetic ID detection and insider threat detection. Most existing techniques for RCD are designed for static data sets, thus not suitable for time-evolving RCD applications. To address this challenge, in this paper, we first proposetwo incremental RCD algorithms, SIRD and BIRD. They arebuilt upon existing density-based techniques for RCD, andincrementally update the detection models, which provide 'timeflexible' RCD. Furthermore, based on BIRD, we propose amodified version named BIRD-LI to deal with the cases wherethe exact priors of the minority classes are not available. Wealso identify a critical task in RCD named query distribution. Itaims to allocate the limited budget into multiple time steps, suchthat the initial examples from the rare classes are detected asearly as possible with the minimum labeling cost. The proposedincremental RCD algorithms and various query distributionstrategies are evaluated empirically on both synthetic and real data. Dawei Zhou 0003, Kangyang Wang, Nan Cao 0001, Jingrui He |
ICDM | 3 |
| 2015 | MindMiner: A Mixed-Initiative Interface for Interactive Distance Metric Learning
Xiangmin Fan, Youming Liu, Nan Cao 0001, Jason I. Hong |
INTERACT (2) | 3 |
| 2015 | Replacing the Irreplaceable: Fast Algorithms for Team Member RecommendationabstractIn this paper, we study the problem of TEAM MEMBER REPLACEMENT -- given a team of people embedded in a social network working on the same task, find a good candidate to best replace a team member who becomes unavailable to perform the task for certain reason (e.g., conflicts of interests or resource capacity). Prior studies in teamwork have suggested that a good team member replacement should bring synergy to the team in terms of having both skill matching and structure matching. However, existing techniques either do not cover both aspects or consider the two aspects independently. In this work, we propose a novel problem formulation using the concept of graph kernels that takes into account the interaction of both skill and structure matching requirements. To tackle the computational challenges, we propose a family of fast algorithms by (a) designing effective pruning strategies, and (b) exploring the smoothness between the existing and the new team structures. We conduct extensive experimental evaluations and user studies on real world datasets to demonstrate the effectiveness and efficiency. Our algorithms (a) perform significantly better than the alternative choices in terms of both precision and recall and (b) scale sub-linearly. Liangyue Li, Hanghang Tong, Nan Cao 0001, Kate Ehrlich, Yu-Ru Lin, Norbou Buchler |
WWW | 3 |
| 2015 | A relative similarity based method for interactive patient risk prediction
Buyue Qian, Xiang Wang 0001, Nan Cao 0001, Yu-Gang Jiang 0001 |
Data Min. Knowl. Discov. | 3 |
| 2014 | UnTangle: Visual Mining for Data with Uncertain Multi-labels via Triangle MapabstractData with multiple uncertain labels are common in many situations. For examples, a movie may be associated with multiple genres with different levels of confidence, and a protein sequence may be probabilistically assigned to several structural subcategories. Despite their ubiquity, the problem of visualizing uncertain labels has not been adequately addressed. Existing approaches often either discard the uncertainty information, or map the data to a low-dimensional subspace where their associations with multiple labels are obscured. In this paper, we propose a novel visual mining technique, UnTangle, for visualizing uncertain multi-labels. In our proposed visualization, data items are placed inside a web of connected triangles, with labels assigned to the triangle vertices such that nearby labels are more relevant to each other. The positions of the data items are determined based on the probabilistic associations between items and labels. UnTangle provides both (a) an automatic label placement algorithm, and (b) adaptive interaction mechanisms that allow users to control the label positioning for different visual queries. Our work makes a unique contribution by providing an effective way to investigate the relationship between data items and their uncertain labels, as well as the relationships among labels. Our user study suggests that the visualization effectively helps users discover emergent patterns and compare the nuances of uncertainty information in the data labels. Yu-Ru Lin, Nan Cao 0001, David Gotz |
ICDM | 2 |
| 2014 | Visual Analysis of Uncertainty in Trajectories
Nan Cao 0001, Siyuan Liu 0001, Lionel M. Ni, Xiaoru Yuan, Huamin Qu |
PAKDD (1) | 2 |
| 2014 | Exploring the associations between drug side-effects and therapeutic indications
Fei Wang 0001, Ping Zhang 0016, Nan Cao 0001, Jianying Hu, Robert Sorrentino |
J. Biomed. Informatics | 3 |
| 2014 | Learning Multiple Relative Attributes With Humans in the LoopabstractSemantic attributes have been recognized as a more spontaneous manner to describe and annotate image content. It is widely accepted that image annotation using semantic attributes is a significant improvement to the traditional binary or multiclass annotation due to its naturally continuous and relative properties. Though useful, existing approaches rely on an abundant supervision and high-quality training data, which limit their applicability. Two standard methods to overcome small amounts of guidance and low-quality training data are transfer and active learning. In the context of relative attributes, this would entail learning multiple relative attributes simultaneously and actively querying a human for additional information. This paper addresses the two main limitations in existing work: 1) it actively adds humans to the learning loop so that minimal additional guidance can be given and 2) it learns multiple relative attributes simultaneously and thereby leverages dependence amongst them. In this paper, we formulate a joint active learning to rank framework with pairwise supervision to achieve these two aims, which also has other benefits such as the ability to be kernelized. The proposed framework optimizes over a set of ranking functions (measuring the strength of the presence of attributes) simultaneously and dependently on each other. The proposed pairwise queries take the form of which one of these two pictures is more natural? These queries can be easily answered by humans. Extensive empirical study on real image data sets shows that our proposed method, compared with several state-of-the-art methods, achieves superior retrieval performance while requires significantly less human inputs. Buyue Qian, Xiang Wang 0001, Nan Cao 0001, Yu-Gang Jiang 0001, Ian Davidson |
IEEE Trans. Image Process. | 3 |
| 2014 | #FluxFlow: Visual Analysis of Anomalous Information Spreading on Social MediaabstractWe present FluxFlow, an interactive visual analysis system for revealing and analyzing anomalous information spreading in social media. Everyday, millions of messages are created, commented, and shared by people on social media websites, such as Twitter and Facebook. This provides valuable data for researchers and practitioners in many application domains, such as marketing, to inform decision-making. Distilling valuable social signals from the huge crowd's messages, however, is challenging, due to the heterogeneous and dynamic crowd behaviors. The challenge is rooted in data analysts' capability of discerning the anomalous information behaviors, such as the spreading of rumors or misinformation, from the rest that are more conventional patterns, such as popular topics and newsworthy events, in a timely fashion. FluxFlow incorporates advanced machine learning algorithms to detect anomalies, and offers a set of novel visualization designs for presenting the detected threads for deeper analysis. We evaluated FluxFlow with real datasets containing the Twitter feeds captured during significant events such as Hurricane Sandy. Through quantitative measurements of the algorithmic performance and qualitative interviews with domain experts, the results show that the back-end anomaly detection model is effective in identifying anomalous retweeting threads, and its front-end interactive visualizations are intuitive and useful for analysts to discover insights in data and comprehend the underlying analytical model. Jian Zhao 0010, Nan Cao 0001, Yale Song, Yu-Ru Lin, Christopher Collins 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | Fast Pairwise Query Selection for Large-Scale Active Learning to RankabstractPair wise learning to rank algorithms (such as Rank SVM) teach a machine how to rank objects given a collection of ordered object pairs. However, their accuracy is highly dependent on the abundance of training data. To address this limitation and reduce annotation efforts, the framework of active pair wise learning to rank was introduced recently. However, in such a framework the number of possible query pairs increases quadratic ally with the number of instances. In this work, we present the first scalable pair wise query selection method using a layered (two-step) hashing framework. The first step relevance hashing aims to retrieve the strongly relevant or highly ranked points, and the second step uncertainty hashing is used to nominate pairs whose ranking is uncertain. The proposed framework aims to efficiently reduce the search space of pair wise queries and can be used with any pair wise learning to rank algorithm with a linear ranking function. We evaluate our approach on large-scale real problems and show it has comparable performance to exhaustive search. The experimental results demonstrate the effectiveness of our approach, and validate the efficiency of hashing in accelerating the search of massive pair wise queries. Buyue Qian, Xiang Wang 0001, Jun Wang 0006, Nan Cao 0001, Weifeng Zhi, Ian Davidson |
ICDM | 5 |
| 2013 | Visual Analysis of Set Relations in a GraphabstractAbstract Many applications can be modeled as a graph with additional attributes attached to the nodes. For example, a graph can be used to model the relationship of people in a social media website or a bibliographical dataset. Meanwhile, additional information is often available, such as the topics people are interested in and the music they listen to. Based on this additional information, different set relationships may exist among people. Revealing the set relationships in a network can help people gain social insight and better understand their roles within a community. In this paper, we present a visualization system for exploring set relations in a graph. Our system is designed to reveal three different relationships simultaneously: the social relationship of people, the set relationship among people's items of interest, and the similarity relationship of the items. We propose two novel visualization designs: a) a glyph‐based visualization to reveal people's set relationships in the context of their social networks; b) an integration of visual links and a contour map to show people and their items of interest which are clustered into different groups. The effectiveness of the designs has been demonstrated by the case studies on two representative datasets including one from a social music service website and another from an academic collaboration network. Fan Du, Nan Cao 0001, Conglei Shi, Hong Zhou 0004, Huamin Qu |
Comput. Graph. Forum | 3 |
| 2012 | Whisper: Tracing the Spatiotemporal Process of Information Diffusion in Real TimeabstractWhen and where is an idea dispersed? Social media, like Twitter, has been increasingly used for exchanging information, opinions and emotions about events that are happening across the world. Here we propose a novel visualization design, "Whisper", for tracing the process of information diffusion in social media in real time. Our design highlights three major characteristics of diffusion processes in social media: the temporal trend, social-spatial extent, and community response of a topic of interest. Such social, spatiotemporal processes are conveyed based on a sunflower metaphor whose seeds are often dispersed far away. In Whisper, we summarize the collective responses of communities on a given topic based on how tweets were retweeted by groups of users, through representing the sentiments extracted from the tweets, and tracing the pathways of retweets on a spatial hierarchical layout. We use an efficient flux line-drawing algorithm to trace multiple pathways so the temporal and spatial patterns can be identified even for a bursty event. A focused diffusion series highlights key roles such as opinion leaders in the diffusion process. We demonstrate how our design facilitates the understanding of when and where a piece of information is dispersed and what are the social responses of the crowd, for large-scale events including political campaigns and natural disasters. Initial feedback from domain experts suggests promising use for today's information consumption and dispersion in the wild. Nan Cao 0001, Yu-Ru Lin, David Lazer, Shixia Liu, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | SolarMap: Multifaceted Visual Analytics for Topic ExplorationabstractDocuments in rich text corpora often contain multiple facets of information. For example, an article from a medical document collection might consist of multifaceted information about symptoms, treatments, causes, diagnoses, prognoses, and preventions. Thus, documents in the collection may have different relations across each of these various facets. Topic analysis and exploration for such multi-relational corpora is a challenging visual analytic task. This paper presents Solar Map, a multifaceted visual analytic technique for visually exploring topics in multi-relational data. Solar Map simultaneously visualizes the topic distribution of the underlying entities from one facet together with keyword distributions that convey the semantic definition of each cluster along a secondary facet. Solar Map combines several visual techniques including 1) topic contour clusters and interactive multifaceted keyword topic rings, 2) a global layout optimization algorithm that aligns each topic cluster with its corresponding keywords, and 3) 2) an optimal temporal network segmentation and layout method that renders temporal evolution of clusters. Finally, the paper concludes with two case studies and quantitative user evaluation which show the power of the Solar Map technique. Nan Cao 0001, David Gotz, Jimeng Sun 0001, Yu-Ru Lin, Huamin Qu |
ICDM | 1 |
| 2011 | ImpactWheel: Visual Analysis of the Impact of Online NewsabstractOnline news usually describes various events over multiple topics. Some of them may generate great impact and affection on other events, organizations or people. For example, a bankruptcy news about a big company may generate a great impact on other companies. Detecting this kind of impact helps users better to understand the affection of a specified event and its epidemic. Powerful text mining techniques have been developed to help users to detect topic trends of news articles. However, there is a lack of effective analysis tools that analyze and reveal the news impact in an intuitive approach. In this paper, we introduce Impact Wheel, an explorative visual analysis system for topic driven news impact detection. We describe two unique aspects of Impact Wheel, including 1) topic driven impact analysis and 2) interactive rich context visualization. Experiments on performance evaluation show that our proposed approach outperforms the two baseline methods on topic driven impact analysis. In addition, we demonstrate the power of the Impact Wheel system through a case study, which shows the benefits of this work, especially in support of rich topic data analysis. Wei Wei 0013, Nan Cao 0001, Jon Atle Gulla, Huamin Qu |
Web Intelligence | 2 |
| 2011 | DICON: Interactive Visual Analysis of Multidimensional ClustersabstractClustering as a fundamental data analysis technique has been widely used in many analytic applications. However, it is often difficult for users to understand and evaluate multidimensional clustering results, especially the quality of clusters and their semantics. For large and complex data, high-level statistical information about the clusters is often needed for users to evaluate cluster quality while a detailed display of multidimensional attributes of the data is necessary to understand the meaning of clusters. In this paper, we introduce DICON, an icon-based cluster visualization that embeds statistical information into a multi-attribute display to facilitate cluster interpretation, evaluation, and comparison. We design a treemap-like icon to represent a multidimensional cluster, and the quality of the cluster can be conveniently evaluated with the embedded statistical information. We further develop a novel layout algorithm which can generate similar icons for similar clusters, making comparisons of clusters easier. User interaction and clutter reduction are integrated into the system to help users more effectively analyze and refine clustering results for large datasets. We demonstrate the power of DICON through a user study and a case study in the healthcare domain. Our evaluation shows the benefits of the technique, especially in support of complex multidimensional cluster analysis. Nan Cao 0001, David Gotz, Jimeng Sun 0001, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2010 | ContexTour: Contextual Contour Analysis on Dynamic Multi-relational ClusteringabstractHuge amounts of rich context social network data are generated everyday from various applications such as FaceBook and Twitter. These data involve multiple social relations which are community-driven and dynamic in nature. The complex interplay of these characteristics poses tremendous challenges on the users who try to understand the underlying patterns in the social media. We introduce an exploratory analytical framework, ContexTour, which generates visual representations for exploring multiple dimensions of community activities, including relevant topics, representative users and the community-generated content, as well as their evolutions. ContexTour consists of two novel and complementary components: (1) Dynamic Relational Clustering (DRC) that efficiently tracks the community evolution and smoothly adapts to the community changes, and (2) Dynamic Network Contour-map (DNC) that visualizes the community activities and evolutions in various dimensions. In our experiments, we demonstrate ContexTour through case studies on the DBLP dataset. The visual results capture interesting and reasonable evolution in Computer Science research communities. Quantitatively, we show 85-165X performance gain of our DRC algorithm over the baseline method. Yu-Ru Lin, Jimeng Sun 0001, Nan Cao 0001, Shixia Liu |
SDM | 3 |
| 2010 | FacetAtlas: Multifaceted Visualization for Rich Text CorporaabstractDocuments in rich text corpora usually contain multiple facets of information. For example, an article about a specific disease often consists of different facets such as symptom, treatment, cause, diagnosis, prognosis, and prevention. Thus, documents may have different relations based on different facets. Powerful search tools have been developed to help users locate lists of individual documents that are most related to specific keywords. However, there is a lack of effective analysis tools that reveal the multifaceted relations of documents within or cross the document clusters. In this paper, we present FacetAtlas, a multifaceted visualization technique for visually analyzing rich text corpora. FacetAtlas combines search technology with advanced visual analytical tools to convey both global and local patterns simultaneously. We describe several unique aspects of FacetAtlas, including (1) node cliques and multifaceted edges, (2) an optimized density map, and (3) automated opacity pattern enhancement for highlighting visual patterns, (4) interactive context switch between facets. In addition, we demonstrate the power of FacetAtlas through a case study that targets patient education in the health care domain. Our evaluation shows the benefits of this work, especially in support of complex multifaceted data analysis. Nan Cao 0001, Jimeng Sun 0001, Yu-Ru Lin, David Gotz, Shixia Liu, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2009 | HiMap: Adaptive visualization of large-scale online social networksabstractVisualizing large-scale online social network is a challenging yet essential task. This paper presents HiMap, a system that visualizes it by clustered graph via hierarchical grouping and summarization. HiMap employs a novel adaptive data loading technique to accurately control the visual density of each graph view, and along with the optimized layout algorithm and the two kinds of edge bundling methods, to effectively avoid the visual clutter commonly found in previous social network visualization tools. HiMap also provides an integrated suite of interactions to allow the users to easily navigate the social map with smooth and coherent view transitions to keep their momentum. Finally, we confirm the effectiveness of HiMap algorithms through graph-travesal based evaluations. Lei Shi 0002, Nan Cao 0001, Shixia Liu, Weihong Qian, Jimeng Sun 0001, Ching-Yung Lin |
PacificVis | 2 |
| 2009 | SmallBlue: Social Network Analysis for Expertise Search and Collective IntelligenceabstractSmallBlue is a social networking application that unlocks the valuable business intelligence of 'who knows what?', 'who knows whom?' and 'who knows what about whom' within an organization, without requiring explicit involvement of individuals. The aim of SmallBlue is to locate knowledgeable colleagues, communities, and knowledge networks in companies. The suite also helps users manage their personal networks, and reach out to their extended network (the friends of their friends) to find and access expertise and information. Ching-Yung Lin, Nan Cao 0001, Shixia Liu, Spiros Papadimitriou, Jimeng Sun 0001, Xifeng Yan |
ICDE | 2 |
| 2009 | MultiVis: Content-Based Social Network Exploration through Multi-way Visual AnalysisabstractWith the explosion of social media, scalability becomes a key challenge. There are two main aspects of the problems that arise: 1) data volume: how to manage and analyze huge datasets to efficiently extract patterns, 2) data understanding: how to facilitate understanding of the patterns by users? To address both aspects of the scalability challenge, we present a hybrid approach that leverages two complementary disciplines, data mining and information visualization. In particular, we propose 1) an analytic data model for content-based networks using tensors; 2) an efficient high-order clustering framework for analyzing the data; 3) a scalable context-sensitive graph visualization to present the clusters. We evaluate the proposed methods using both synthetic and real datasets. In terms of computational efficiency, the proposed methods are an order of magnitude faster compared to the baseline. In terms of effectiveness, we present several case studies of real corporate social networks. Jimeng Sun 0001, Spiros Papadimitriou, Ching-Yung Lin, Nan Cao 0001, Shixia Liu, Weihong Qian |
SDM | 4 |
| 2008 | Interactive Visual Analysis of the NSF Funding InformationabstractThis paper presents an interactive visualization toolkit for navigating and analyzing the National Science Foundation (NSF) funding information. Our design builds upon the treemap layout and the stacked graph to contribute customized techniques for visually navigating and interacting with the hierarchical data of NSF programs and proposals, supporting visual search and analysis, and allowing the user to make informed decision. In this visualization toolkit, we propose two visualization techniques to simplify the navigation of the hierarchical data: 2.5 Dimensional treemaps to make the hierarchical structure more easily to be recognized, and labeled treemap to help the user to get a clear overview of the content of the structure and make the internal area of rectangles correspond to the weights of the data set. Furthermore, an incremental layout method is adopted to handle information on a large scale. The improved treemap visualization will help to visually analyze the static funding data and the stacked graph is utilized to analyze the time-series data. Through these visual analysis techniques, research trends of NSF, popular NSF programs are quickly identified. The primary contribution is a demonstration of novel ways to effectively present and analyze NSF funding data. Shixia Liu, Nan Cao 0001 |
PacificVis | 2 |
| 2006 | The visual funding navigator: analysis of the NSF funding informationabstractThis paper presents an interactive visualization toolkit for navigating and analyzing the National Science Foundation (NSF) funding information. Our design builds upon an improved 2.5D treemap layout and the stacked graph to contribute customized techniques for visually navigating and interacting with the hierarchical data of NSF programs and proposals. Furthermore, an incremental layout method is adopted to handle information on a large scale. The improved treemap visualization will help to visually analyze the static funding related data and the stacked graph is utilized to analyze the time-series data. Through these visual analysis techniques, research trends of NSF, popular NSF programs are quickly identified. Shixia Liu, Nan Cao 0001, Hui Su |
CIKM | 2 |