Weiwei Cui 0001

dblp:34/2915-1 · DBLP profile ↗
← Back
54ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0003-0870-7628ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 7 first-author · 15 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
YearPublicationVenuePosition
2026 RCInvestigator: Towards Better Investigation of Anomaly Root Causes in Cloud Computing Systems
abstract
Root cause analysis (RCA) is critical for maintaining the availability and efficiency of cloud computing systems. However, identifying root causes from the large-scale, high-dimensional monitoring data generated by these complex environments is a significant challenge. Current approaches often rely on time-consuming manual analysis to ensure flexibility and reliability, while recent automated methods lack the crucial insights provided by domain experts. To bridge this gap, we propose RCInvestigator, a visual analytics system that facilitates interactive root cause investigation by establishing a tight collaboration between human experts and machine analysis. Our approach addresses three key challenges: a) modeling databases for the root cause investigation, b) inferring root causes from large-scale time series, and c) building comprehensible investigation results. We demonstrate the effectiveness and utility of RCInvestigator through two real-world case studies, which received positive feedback from domain experts.
Yunfan Zhou, Shandan Zhou, Weiwei Cui 0001, Qingwei Lin, Thomas Moscibroda, Di Weng, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.5
2025 ReSpark: Leveraging Previous Data Reports as References to Generate New Reports with LLMs
abstract
Figure 1: ReSpark Overview.ReSpark helps users generate new data reports by reusing existing ones.It begins by ranking reports based on their alignment with the target dataset to assist in selecting a reference report (a).The selected report is then segmented into a sequence of analysis segments (b).For each segment, ReSpark extracts the analysis objective (c) and adapts it to the target dataset (d1).It then generates the corresponding code and charts (d2), followed by textual insights (d3).These elements are finally composed into a new report.
Sitong Pan, Weiwei Cui 0001, Dazhen Deng, Yingcai Wu
UIST5
2025 Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables
abstract
Data cleaning is a long-standing challenge in data management. While powerful logic and statistical algorithms have been developed to detect and repair data errors in tables, existing algorithms predominantly rely on domain-experts to first manually specify data-quality constraints specific to a given table, before data cleaning algorithms can be applied. In this work, we observe that there is an important class of data-quality constraints that we call Semantic-Domain Constraints, which can be reliably inferred and automatically applied to any tables, without requiring domain-experts to manually specify on a per-table basis. We develop a principled framework to systematically learn such constraints from table corpora using large-scale statistical tests, which can further be distilled into a core set of constraints using our optimization framework, with provable quality guarantees. Extensive evaluations show that this new class of constraints can be used to both (1) directly detect errors on real tables in the wild, and (2) augment existing expert-driven data-cleaning techniques as a new class of complementary constraints. Our code and data are available at https://github.com/qixuchen/AutoTest for future research.
Qixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui 0001, Dongmei Zhang 0001, Surajit Chaudhuri
Proc. ACM Manag. Data4
2025 Relation-Driven Query of Multiple Time Series
abstract
Querying time series based on their relations is a crucial part of multiple time series analysis. By retrieving and understanding time series relations, analysts can easily detect anomalies and validate hypotheses in complex time series datasets. However, current relation extraction approaches, including knowledge- and data-driven ones, tend to be laborious and do not support heterogeneous relations. By conducting a formative study with 11 experts, we concluded six time series relations, including correlation, causality, similarity, lag, arithmetic, and meta, and summarized three pain points in querying time series involving these relations. We proposed RelaQ, an interactive system that supports the time series query via relation specifications. RelaQ allows users to intuitively specify heterogeneous relations when querying multiple time series, understand the query results based on a scalable, multi-level visualization, and explore possible relations beyond the existing queries. RelaQ is evaluated with two cases and a user study with 12 participants, showing promising effectiveness and usability.
Zikun Deng, Weiwei Cui 0001, Di Weng, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.4
2025 ChartGPT: Leveraging LLMs to Generate Charts From Abstract Natural Language
abstract
The use of natural language interfaces (NLIs) to create charts is becoming increasingly popular due to the intuitiveness of natural language interactions. One key challenge in this approach is to accurately capture user intents and transform them to proper chart specifications. This obstructs the wide use of NLI in chart generation, as users' natural language inputs are generally abstract (i.e., ambiguous or under-specified), without a clear specification of visual encodings. Recently, pre-trained large language models (LLMs) have exhibited superior performance in understanding and generating natural language, demonstrating great potential for downstream tasks. Inspired by this major trend, we propose ChartGPT, generating charts from abstract natural language inputs. However, LLMs are struggling to address complex logic problems. To enable the model to accurately specify the complex parameters and perform operations in chart generation, we decompose the generation process into a step-by-step reasoning pipeline, so that the model only needs to reason a single and specific sub-task during each run. Moreover, LLMs are pre-trained on general datasets, which might be biased for the task of chart generation. To provide adequate visualization knowledge, we create a dataset consisting of abstract utterances and charts and improve model performance through fine-tuning. We further design an interactive interface for ChartGPT that allows users to check and modify the intermediate outputs of each step. The effectiveness of the proposed system is evaluated through quantitative evaluations and a user study.
Weiwei Cui 0001, Dazhen Deng, Xinjing Yi, Yurun Yang, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.2
2024 Understanding Nonlinear Collaboration between Human and AI Agents: A Co-design Framework for Creative Design
abstract
Creative design is a nonlinear process where designers generate diverse ideas in the pursuit of an open-ended goal and converge towards consensus through iterative remixing. In contrast, AI-powered design tools often employ a linear sequence of incremental and precise instructions to approximate design objectives. Such operations violate customary creative design practices and thus hinder AI agents’ ability to complete creative design tasks. To explore better human-AI co-design tools, we first summarize human designers’ practices through a formative study with 12 design experts. Taking graphic design as a representative scenario, we formulate a nonlinear human-AI co-design framework and develop a proof-of-concept prototype, OptiMuse. We evaluate OptiMuse and validate the nonlinear framework through a comparative study. We notice a subconscious change in people’s attitudes towards AI agents, shifting from perceiving them as mere executors to regarding them as opinionated colleagues. This shift effectively fostered the exploration and reflection processes of individual designers.
Renzhong Li, Junxiu Tang, Tan Tang, Haotian Li 0001, Weiwei Cui 0001, Yingcai Wu
CHI6
2024 Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table Representations
abstract
Spreadsheets are widely recognized as the most popular end-user programming tools, which blend the power of formula-based computation, with an intuitive table-based interface. Today, spreadsheets are used by billions of users to manipulate tables, most of whom are neither database experts nor professional programmers. Despite the success of spreadsheets, authoring complex formulas remains challenging, as non-technical users need to look up and understand non-trivial formula syntax. To address this pain point, we leverage the observation that there is often an abundance of similar-looking spreadsheets in the same organization, which not only have similar data, but also share similar computation logic encoded as formulas. We develop an Auto-Formula system that can accurately predict formulas that users want to author in a target spreadsheet cell, by learning and adapting formulas that already exist in similar spreadsheets, using contrastive-learning techniques inspired by "similar-face recognition" from compute vision. Extensive evaluations on over 2K test formulas extracted from real enterprise spreadsheets show the effectiveness of Auto-Formula over alternatives. Our benchmark data is available at https://github.com/microsoft/Auto-Formula to facilitate future research.
Sibei Chen, Yeye He, Weiwei Cui 0001, Ju Fan, Dongmei Zhang 0001, Surajit Chaudhuri
Proc. ACM Manag. Data3
2024 Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks
abstract
Language models, such as GPT-3 and ChatGPT, demonstrate remarkable abilities to follow diverse human instructions and perform a wide range of tasks, using instruction fine-tuning. However, when we test language models with a range of basic table-understanding tasks, we observe that today's language models are still sub-optimal in many table-related tasks, likely because they are pre-trained predominantly on one-dimensional natural-language texts, whereas relational tables are two-dimensional objects. In this work, we propose a new "\emphtable fine-tuning '' paradigm, where we continue to train/fine-tune language models like GPT-3.5 and ChatGPT, using diverse table-tasks synthesized from real tables as training data, which is analogous to "instruction fine-tuning'', but with the goal of enhancing language models' ability to understand tables and perform table tasks. We show that our resulting \sys models demonstrate: (1) better table-understanding capabilities, by consistently outperforming the vanilla GPT-3.5 and ChatGPT, on a wide range of table tasks (data transformation, data cleaning, data profiling, data imputation, table-QA, etc.), including tasks that are completely holdout and unseen during training, and (2) strong generalizability, in its ability to respond to diverse human instructions to perform new and unseen table-tasks, in a manner similar to GPT-3.5 and ChatGPT. Our code and data have been released at https://github.com/microsoft/Table-GPT for future research.
Peng Li 0062, Yeye He, Dror Yashar, Weiwei Cui 0001, Danielle Rifinski Fainman, Dongmei Zhang 0001, Surajit Chaudhuri
Proc. ACM Manag. Data4
2024 Causality-Based Visual Analysis of Questionnaire Responses
abstract
As the final stage of questionnaire analysis, causal reasoning is the key to turning responses into valuable insights and actionable items for decision-makers. During the questionnaire analysis, classical statistical methods (e.g., Differences-in-Differences) have been widely exploited to evaluate causality between questions. However, due to the huge search space and complex causal structure in data, causal reasoning is still extremely challenging and time-consuming, and often conducted in a trial-and-error manner. On the other hand, existing visual methods of causal reasoning face the challenge of bringing scalability and expert knowledge together and can hardly be used in the questionnaire scenario. In this work, we present a systematic solution to help analysts effectively and efficiently explore questionnaire data and derive causality. Based on the association mining algorithm, we dig question combinations with potential inner causality and help analysts interactively explore the causal sub-graph of each question combination. Furthermore, leveraging the requirements collected from the experts, we built a visualization tool and conducted a comparative study with the state-of-the-art system to show the usability and efficiency of our system.
Renzhong Li, Weiwei Cui 0001, Xiao Xie, Rui Ding 0001, Yun Wang 0012, Hong Zhou 0004, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.2
2024 NL2Color: Refining Color Palettes for Charts with Natural Language
abstract
Choice of color is critical to creating effective charts with an engaging, enjoyable, and informative reading experience. However, designing a good color palette for a chart is a challenging task for novice users who lack related design expertise. For example, they often find it difficult to articulate their abstract intentions and translate these intentions into effective editing actions to achieve a desired outcome. In this work, we present NL2Color, a tool that allows novice users to refine chart color palettes using natural language expressions of their desired outcomes. We first collected and categorized a dataset of 131 triplets, each consisting of an original color palette of a chart, an editing intent, and a new color palette designed by human experts according to the intent. Our tool employs a large language model (LLM) to substitute the colors in original palettes and produce new color palettes by selecting some of the triplets as few-shot prompts. To evaluate our tool, we conducted a comprehensive two-stage evaluation, including a crowd-sourcing study ( N=71) and a within-subjects user study ( N=12). The results indicate that the quality of the color palettes revised by NL2Color has no significantly large difference from those designed by human experts. The participants who used NL2Color obtained revised color palettes to their satisfaction in a shorter period and with less effort.
Chuhan Shi, Weiwei Cui 0001, Chengzhong Liu, Chengbo Zheng, Qiong Luo 0001, Xiaojuan Ma
IEEE Trans. Vis. Comput. Graph.2
2023 Visually Analysing the Fairness of Clustered Federated Learning with Non-IID Data
abstract
As a popular distributed privacy-preserving machine learning approach, federated learning (FL) trains a global model across numerous clients without directly sharing their private training data. A fundamental challenge of FL is non-independently and identically distributed (non-IID) data, which can be addressed by clustered FL (CFL) that explores inherent partitions among clients to capture heterogeneous local data distributions. Existing CFL methods, however, tend to optimise the mean accuracy across all clients, overlooking the important issue of fairness, that is, uniformity of clients' performance. Besides, prior FL- related visual analytics (VA) frameworks are developed based on the standard FL scenario of single global model and are hence inapplicable to the CFL case of multiple global (cluster) models. To fill the research gap, this study proposes a novel VA framework named cCFLvis to illustrate and analyse variation of fairness during the CFL learning process. A new metric, change of absolute deviation from the mean (ΔAD), is introduced to quantify such variation. Client embeddings are generated by contrastive learning-based dimensionality reduction techniques to facilitate 2-d visualisation of client disparity. Visual cues can be drawn to help understand fairness variation and to offer clues for improving learning outcomes. cCFL vis is a general framework applicable to various CFL methods. Its effectiveness is showcased in a simple demonstration and two case studies, each conducted with a distinct representative CFL method and a different federated non-IID dataset.
Weiwei Cui 0001, Bin B. Zhu
IJCNN2
2023 Auto-Validate by-History: Auto-Program Data Quality Constraints to Validate Recurring Data Pipelines
abstract
Data pipelines are widely employed in modern enterprises to power a variety of Machine-Learning (ML) and Business-Intelligence (BI) applications. Crucially, these pipelines are recurring (e.g., daily or hourly) in production settings to keep data updated so that ML models can be re-trained regularly, and BI dashboards refreshed frequently. However, data quality (DQ) issues can often creep into recurring pipelines because of upstream schema and data drift over time. As modern enterprises operate thousands of recurring pipelines, today data engineers have to spend substantial efforts to manually monitor and resolve DQ issues, as part of their DataOps and MLOps practices.
Dezhan Tu, Yeye He, Weiwei Cui 0001, Shi Han, Dongmei Zhang 0001, Surajit Chaudhuri
KDD3
2023 Revisiting the Design Patterns of Composite Visualizations
abstract
Composite visualization is a popular design strategy that represents complex datasets by integrating multiple visualizations in a meaningful and aesthetic layout, such as juxtaposition, overlay, and nesting. With this strategy, numerous novel designs have been proposed in visualization publications to accomplish various visual analytic tasks. However, there is a lack of understanding of design patterns of composite visualization, thus failing to provide holistic design space and concrete examples for practical use. In this article, we opted to revisit the composite visualizations in IEEE VIS publications and answered what and how visualizations of different types are composed together. To achieve this, we first constructed a corpus of composite visualizations from the publications and analyzed common practices, such as the pattern distributions and co-occurrence of visualization types. From the analysis, we obtained insights into different design patterns on the utilities and their potential pros and cons. Furthermore, we discussed usage scenarios of our taxonomy and corpus and how future research on visualization composition can be conducted on the basis of this study.
Dazhen Deng, Weiwei Cui 0001, Xiyu Meng, Mengye Xu, Yu Liao, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.2
2023 VisImages: A Fine-Grained Expert-Annotated Visualization Dataset
abstract
Images in visualization publications contain rich information, e.g., novel visualization designs and implicit design patterns of visualizations. A systematic collection of these images can contribute to the community in many aspects, such as literature analysis and automated tasks for visualization. In this paper, we build and make public a dataset, VisImages, which collects 12,267 images with captions from 1,397 papers in IEEE InfoVis and VAST. Built upon a comprehensive visualization taxonomy, the dataset includes 35,096 visualizations and their bounding boxes in the images. We demonstrate the usefulness of VisImages through three use cases: 1) investigating the use of visualizations in the publications with VisImages Explorer, 2) training and benchmarking models for visualization classification, and 3) localizing visualizations in the visual analytics systems automatically.
Dazhen Deng, Yihong Wu 0003, Xinhuan Shu, Jiang Wu 0012, Siwei Fu, Weiwei Cui 0001, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.6
2023 Aesthetics++: Refining Graphic Designs by Exploring Design Principles and Human Preference
abstract
During the creation of graphic designs, individuals inevitably spend a lot of time and effort on adjusting visual attributes (e.g., positions, colors, and fonts) of elements to make them more aesthetically pleasing. It is a trial-and-error process, requires repetitive edits, and relies on good design knowledge. In this work, we seek to alleviate such difficulty by automatically suggesting aesthetic improvements, i.e., taking an existing design as the input and generating a refined version with improved aesthetic quality as the output. This goal presents two challenges: proposing a refined design based on the user-given one, and assessing whether the new design is better aesthetically. To cope with these challenges, we propose a design principle-guided candidate generation stage and a data-driven candidate evaluation stage. In the candidate generation stage, we generate candidate designs by leveraging design principles as the guidance to make changes around the existing design. In the candidate evaluation stage, we learn a ranking model upon a dataset that can reflect humans' aesthetic preference, and use it to choose the most aesthetically pleasing one from the generated candidates. We implement a prototype system on presentation slides and demonstrate the effectiveness of our approach through quantitative analysis, sample results, and user studies.
Wenyuan Kong, Zhaoyun Jiang, Shizhao Sun, Zhuoning Guo, Weiwei Cui 0001, Ting Liu 0002, Jianguang Lou, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.5
2023 Erato: Cooperative Data Story Editing via Fact Interpolation
abstract
As an effective form of narrative visualization, visual data stories are widely used in data-driven storytelling to communicate complex insights and support data understanding. Although important, they are difficult to create, as a variety of interdisciplinary skills, such as data analysis and design, are required. In this work, we introduce Erato, a human-machine cooperative data story editing system, which allows users to generate insightful and fluent data stories together with the computer. Specifically, Erato only requires a number of keyframes provided by the user to briefly describe the topic and structure of a data story. Meanwhile, our system leverages a novel interpolation algorithm to help users insert intermediate frames between the keyframes to smooth the transition. We evaluated the effectiveness and usefulness of the Erato system via a series of evaluations including a Turing test, a controlled user study, a performance validation, and interviews with three expert users. The evaluation results showed that the proposed interpolation technique was able to generate coherent story content and help users create data stories more efficiently.
Mengdi Sun, Ligan Cai, Weiwei Cui 0001, Yanqiu Wu 0001, Yang Shi 0007, Nan Cao 0001
IEEE Trans. Vis. Comput. Graph.3
2022 ChartStamp: Robust Chart Embedding for Real-World Applications
abstract
Deep learning-based image embedding methods are typically designed for natural images and may not work for chart images due to their homogeneous regions, which lack variations to hide data both robustly and imperceptibly. In this paper, we propose ChartStamp, the first chart embedding method that is robust to real-world printing and displaying (printed on paper and displayed on screen, respectively, and then captured with a camera) while maintaining a good perceptual quality. ChartStamp hides 100, 1,000, or 10,000 raw bits into a chart image, depending on the designated robustness to printing, displaying, or JPEG. To ensure perceptual quality, it introduces a new perceptual model to guide embedding to insensitive regions of a chart image and a smoothness loss to ensure smoothness of the embedding residual in homogeneous regions. ChartStamp applies a distortion layer approximating designated real-world manipulations to train a model robust to these manipulations. Our experimental evaluation indicates that ChartStamp achieves the robustness and embedding capacity on chart images similar to their state-of-the-art counterparts on natural images. Our user studies indicate that ChartStamp achieves better perceptual quality than existing robust chart embedding methods and that our perceptual model outperforms the existing perceptual model.
Jiayun Fu, Bin B. Zhu, Yayi Zou, Weiwei Cui 0001, Yun Wang 0012, Dongmei Zhang 0001, Xiaojing Ma 0002, Hai Jin 0001
ACM Multimedia6
2022 A Mixed-Initiative Approach to Reusing Infographic Charts
abstract
Infographic bar charts have been widely adopted for communicating numerical information because of their attractiveness and memorability. However, these infographics are often created manually with general tools, such as PowerPoint and Adobe Illustrator, and merely composed of primitive visual elements, such as text blocks and shapes. With the absence of chart models, updating or reusing these infographics requires tedious and error-prone manual edits. In this paper, we propose a mixed-initiative approach to mitigate this pain point. On one hand, machines are adopted to perform precise and trivial operations, such as mapping numerical values to shape attributes and aligning shapes. On the other hand, we rely on humans to perform subjective and creative tasks, such as changing embellishments or approving the edits made by machines. We encapsulate our technique in a PowerPoint add-in prototype and demonstrate the effectiveness by applying our technique on a diverse set of infographic bar chart examples.
Weiwei Cui 0001, Jinpeng Wang 0001, Yun Wang 0012, Chin-Yew Lin, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.1
2022 AI4VIS: Survey on Artificial Intelligence Approaches for Data Visualization
abstract
Visualizations themselves have become a data format. Akin to other data formats such as text and images, visualizations are increasingly created, stored, shared, and (re-)used with artificial intelligence (AI) techniques. In this survey, we probe the underlying vision of formalizing visualizations as an emerging data format and review the recent advance in applying AI techniques to visualization data (AI4VIS). We define visualization data as the digital representations of visualizations in computers and focus on data visualization (e.g., charts and infographics). We build our survey upon a corpus spanning ten different fields in computer science with an eye toward identifying important common interests. Our resulting taxonomy is organized around WHAT is visualization data and its representation, WHY and HOW to apply AI to visualization data. We highlight a set of common tasks that researchers apply to the visualization data and present a detailed discussion of AI approaches developed to accomplish those tasks. Drawing upon our literature review, we discuss several important research questions surrounding the management and exploitation of visualization data, as well as the role of AI in support of those processes. We make the list of surveyed papers and related material available online at.
Aoyu Wu, Yun Wang 0012, Xinhuan Shu, Dominik Moritz, Weiwei Cui 0001, Dongmei Zhang 0001, Huamin Qu
IEEE Trans. Vis. Comput. Graph.5
2021 Learning to Automate Chart Layout Configurations Using Crowdsourced Paired Comparison
abstract
We contribute a method to automate parameter configurations for chart layouts by learning from human preferences. Existing charting tools usually determine the layout parameters using predefined heuristics, producing sub-optimal layouts. People can repeatedly adjust multiple parameters (e.g., chart size, gap) to achieve visually appealing layouts. However, this trial-and-error process is unsystematic and time-consuming, without a guarantee of improvement. To address this issue, we develop Layout Quality Quantifier (LQ2), a machine learning model that learns to score chart layouts from paired crowdsourcing data. Combined with optimization techniques, LQ2 recommends layout parameters that improve the charts’ layout quality. We apply LQ2 on bar charts and conduct user studies to evaluate its effectiveness by examining the quality of layouts it produces. Results show that LQ2 can generate more visually appealing layouts than both laypeople and baselines. This work demonstrates the feasibility and usages of quantifying human preferences and aesthetics for chart layouts.
Aoyu Wu, Liwenhan Xie, Bongshin Lee, Yun Wang 0012, Weiwei Cui 0001, Huamin Qu
CHI5
2021 Animated Presentation of Static Infographics with InfoMotion
abstract
Abstract By displaying visual elements logically in temporal order, animated infographics can help readers better understand layers of information expressed in an infographic. While many techniques and tools target the quick generation of static infographics, few support animation designs. We propose InfoMotion that automatically generates animated presentations of static infographics. We first conduct a survey to explore the design space of animated infographics. Based on this survey, InfoMotion extracts graphical properties of an infographic to analyze the underlying information structures; then, animation effects are applied to the visual elements in the infographic in temporal order to present the infographic. The generated animations can be used in data videos or presentations. We demonstrate the utility of InfoMotion with two example applications, including mixed‐initiative animation authoring and animation recommendation. To further understand the quality of the generated animations, we conduct a user study to gather subjective feedback on the animations generated by InfoMotion.
Yun Wang 0012, Ray Huang, Weiwei Cui 0001, Dongmei Zhang 0001
Comput. Graph. Forum4
2021 Chartem: Reviving Chart Images with Data Embedding
abstract
In practice, charts are widely stored as bitmap images. Although easily consumed by humans, they are not convenient for other uses. For example, changing the chart style or type or a data value in a chart image practically requires creating a completely new chart, which is often a time-consuming and error-prone process. To assist these tasks, many approaches have been proposed to automatically extract information from chart images with computer vision and machine learning techniques. Although they have achieved promising preliminary results, there are still a lot of challenges to overcome in terms of robustness and accuracy. In this paper, we propose a novel alternative approach called Chartem to address this issue directly from the root. Specifically, we design a data-embedding schema to encode a significant amount of information into the background of a chart image without interfering human perception of the chart. The embedded information, when extracted from the image, can enable a variety of visualization applications to reuse or repurpose chart images. To evaluate the effectiveness of Chartem, we conduct a user study and performance experiments on Chartem embedding and extraction algorithms. We further present several prototype applications to demonstrate the utility of Chartem.
Jiayun Fu, Bin B. Zhu, Weiwei Cui 0001, Yun Wang 0012, Dongmei Zhang 0001, Xiaojing Ma 0002
IEEE Trans. Vis. Comput. Graph.3
2021 Retrieve-Then-Adapt: Example-based Automatic Generation for Proportion-related Infographics
abstract
Infographic is a data visualization technique which combines graphic and textual descriptions in an aesthetic and effective manner. Creating infographics is a difficult and time-consuming process which often requires significant attempts and adjustments even for experienced designers, not to mention novice users with limited design expertise. Recently, a few approaches have been proposed to automate the creation process by applying predefined blueprints to user information. However, predefined blueprints are often hard to create, hence limited in volume and diversity. In contrast, good infogrpahics have been created by professionals and accumulated on the Internet rapidly. These online examples often represent a wide variety of design styles, and serve as exemplars or inspiration to people who like to create their own infographics. Based on these observations, we propose to generate infographics by automatically imitating examples. We present a two-stage approach, namely retrieve-then-adapt. In the retrieval stage, we index online examples by their visual elements. For a given user information, we transform it to a concrete query by sampling from a learned distribution about visual elements, and then find appropriate examples in our example library based on the similarity between example indexes and the query. For a retrieved example, we generate an initial drafts by replacing its content with user information. However, in many cases, user information cannot be perfectly fitted to retrieved examples. Therefore, we further introduce an adaption stage. Specifically, we propose a MCMC-like approach and leverage recursive neural networks to help adjust the initial draft and improve its visual appearance iteratively, until a satisfactory result is obtained. We implement our approach on widely-used proportion-related infographics, and demonstrate its effectiveness by sample results and expert reviews.
Chunyao Qian, Shizhao Sun, Weiwei Cui 0001, Jian-Guang Lou, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.3
2020 Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements
abstract
Combining data content with visual embellishments, infographics can effectively deliver messages in an engaging and memorable manner. Various authoring tools have been proposed to facilitate the creation of infographics. However, creating a professional infographic with these authoring tools is still not an easy task, requiring much time and design expertise. Therefore, these tools are generally not attractive to casual users, who are either unwilling to take time to learn the tools or lacking in proper design expertise to create a professional infographic. In this paper, we explore an alternative approach: to automatically generate infographics from natural language statements. We first conducted a preliminary study to explore the design space of infographics. Based on the preliminary study, we built a proof-of-concept system that automatically converts statements about simple proportion-related statistics to a set of infographics with pre-designed styles. Finally, we demonstrated the usability and usefulness of the system through sample results, exhibits, and expert reviews.
Weiwei Cui 0001, Xiaoyu Zhang 0014, Yun Wang 0012, Bei Chen 0008, Lei Fang 0004, Jian-Guang Lou, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.1
2020 DeepDrawing: A Deep Learning Approach to Graph Drawing
abstract
Node-link diagrams are widely used to facilitate network explorations. However, when using a graph drawing technique to visualize networks, users often need to tune different algorithm-specific parameters iteratively by comparing the corresponding drawing results in order to achieve a desired visual effect. This trial and error process is often tedious and time-consuming, especially for non-expert users. Inspired by the powerful data modelling and prediction capabilities of deep learning techniques, we explore the possibility of applying deep learning techniques to graph drawing. Specifically, we propose using a graph-LSTM-based approach to directly map network structures to graph drawings. Given a set of layout examples as the training dataset, we train the proposed graph-LSTM-based model to capture their layout characteristics. Then, the trained model is used to generate graph drawings in a similar style for new networks. We evaluated the proposed approach on two special types of layouts (i.e., grid layouts and star layouts) and two general types of layouts (i.e., ForceAtlas2 and PivotMDS) in both qualitative and quantitative ways. The results provide support for the effectiveness of our approach. We also conducted a time cost assessment on the drawings of small graphs with 20 to 50 nodes. We further report the lessons we learned and discuss the limitations and future work.
Yong Wang 0021, Zhihua Jin, Qianwen Wang 0001, Weiwei Cui 0001, Tengfei Ma 0001, Huamin Qu
IEEE Trans. Vis. Comput. Graph.4
2020 DataShot: Automatic Generation of Fact Sheets from Tabular Data
abstract
Fact sheets with vivid graphical design and intriguing statistical insights are prevalent for presenting raw data. They help audiences understand data-related facts effectively and make a deep impression. However, designing a fact sheet requires both data and design expertise and is a laborious and time-consuming process. One needs to not only understand the data in depth but also produce intricate graphical representations. To assist in the design process, we present DataShot which, to the best of our knowledge, is the first automated system that creates fact sheets automatically from tabular data. First, we conduct a qualitative analysis of 245 infographic examples to explore general infographic design space at both the sheet and element levels. We identify common infographic structures, sheet layouts, fact types, and visualization styles during the study. Based on these findings, we propose a fact sheet generation pipeline, consisting of fact extraction, fact composition, and presentation synthesis, for the auto-generation workflow. To validate our system, we present use cases with three real-world datasets. We conduct an in-lab user study to understand the usage of our system. Our evaluation results show that DataShot can efficiently generate satisfactory fact sheets to support further customization and data presentation.
Yun Wang 0012, Zhida Sun, Weiwei Cui 0001, Xiaojuan Ma, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.4
2019 Oui! Outlier Interpretation on Multi-dimensional Data via Visual Analytics
abstract
Abstract Outliers, the data instances that do not conform with normal patterns in a dataset, are widely studied in various domains, such as cybersecurity, social analysis, and public health. By detecting and analyzing outliers, users can either gain insights into abnormal patterns or purge the data of errors. However, different domains usually have different considerations with respect to outliers. Understanding the defining characteristics of outliers is essential for users to select and filter appropriate outliers based on their domain requirements. Unfortunately, most existing work focuses on the efficiency and accuracy of outlier detection, neglecting the importance of outlier interpretation. To address these issues, we propose Oui, a visual analytic system that helps users understand, interpret, and select the outliers detected by various algorithms. We also present a usage scenario on a real dataset and a qualitative user study to demonstrate the effectiveness and usefulness of our system.
Weiwei Cui 0001, Huamin Qu, Dongmei Zhang 0001
Comput. Graph. Forum2
2019 DeepTracker: Visualizing the Training Process of Convolutional Neural Networks
abstract
Deep Convolutional Neural Networks (CNNs) have achieved remarkable success in various fields. However, training an excellent CNN is practically a trial-and-error process that consumes a tremendous amount of time and computer resources. To accelerate the training process and reduce the number of trials, experts need to understand what has occurred in the training process and why the resulting CNN behaves as it does. However, current popular training platforms, such as TensorFlow, only provide very little and general information, such as training/validation errors, which is far from enough to serve this purpose. To bridge this gap and help domain experts with their training tasks in a practical environment, we propose a visual analytics system, DeepTracker, to facilitate the exploration of the rich dynamics of CNN training processes and to identify the unusual patterns that are hidden behind the huge amount of information in training log. Specifically, we combine a hierarchical index mechanism and a set of hierarchical small multiples to help experts explore the entire training log from different levels of detail. We also introduce a novel cube-style visualization to reveal the complex correlations among multiple types of heterogeneous training data, including neuron weights, validation images, and training iterations. Three case studies are conducted to demonstrate how DeepTracker provides its users with valuable knowledge in an industry-level CNN training process; namely, in our case, training ResNet-50 on the ImageNet dataset. We show that our method can be easily applied to other state-of-the-art “very deep” CNN models.
Dongyu Liu, Weiwei Cui 0001, Yuxiao Guo 0001, Huamin Qu
ACM Trans. Intell. Syst. Technol.2
2019 iStoryline: Effective Convergence to Hand-drawn Storylines
abstract
Storyline visualization techniques have progressed significantly to generate illustrations of complex stories automatically. However, the visual layouts of storylines are not enhanced accordingly despite the improvement in the performance and extension of its application area. Existing methods attempt to achieve several shared optimization goals, such as reducing empty space and minimizing line crossings and wiggles. However, these goals do not always produce optimal results when compared to hand-drawn storylines. We conducted a preliminary study to learn how users translate a narrative into a hand-drawn storyline and check whether the visual elements in hand-drawn illustrations can be mapped back to appropriate narrative contexts. We also compared the hand-drawn storylines with storylines generated by the state-of-the-art methods and found they have significant differences. Our findings led to a design space that summarizes 1) how artists utilize narrative elements and 2) the sequence of actions artists follow to portray expressive and attractive storylines. We developed iStoryline, an authoring tool for integrating high-level user interactions into optimization algorithms and achieving a balance between hand-drawn storylines and automatic layouts. iStoryline allows users to create novel storyline visualizations easily according to their preferences by modifying the automatically generated layouts. The effectiveness and usability of iStoryline are studied with qualitative evaluations.
Tan Tang, Sadia Rubab, Jiewen Lai, Weiwei Cui 0001, Lingyun Yu 0001, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.4
2019 Narvis: Authoring Narrative Slideshows for Introducing Data Visualization Designs
abstract
Visual designs can be complex in modern data visualization systems, which poses special challenges for explaining them to the non-experts. However, few if any presentation tools are tailored for this purpose. In this study, we present Narvis, a slideshow authoring tool designed for introducing data visualizations to non-experts. Narvis targets two types of end users: teachers, experts in data visualization who produce tutorials for explaining a data visualization, and students, non-experts who try to understand visualization designs through tutorials. We present an analysis of requirements through close discussions with the two types of end users. The resulting considerations guide the design and implementation of Narvis. Additionally, to help teachers better organize their introduction slideshows, we specify a data visualization as a hierarchical combination of components, which are automatically detected and extracted by Narvis. The teachers craft an introduction slideshow through first organizing these components, and then explaining them sequentially. A series of templates are provided for adding annotations and animations to improve efficiency during the authoring process. We evaluate Narvis through a qualitative analysis of the authoring experience, and a preliminary evaluation of the generated slideshows.
Qianwen Wang 0001, Zhen Li 0044, Siwei Fu, Weiwei Cui 0001, Huamin Qu
IEEE Trans. Vis. Comput. Graph.4
2019 iForest: Interpreting Random Forests via Visual Analytics
abstract
As an ensemble model that consists of many independent decision trees, random forests generate predictions by feeding the input to internal trees and summarizing their outputs. The ensemble nature of the model helps random forests outperform any individual decision tree. However, it also leads to a poor model interpretability, which significantly hinders the model from being used in fields that require transparent and explainable predictions, such as medical diagnosis and financial fraud detection. The interpretation challenges stem from the variety and complexity of the contained decision trees. Each decision tree has its unique structure and properties, such as the features used in the tree and the feature threshold in each tree node. Thus, a data input may lead to a variety of decision paths. To understand how a final prediction is achieved, it is desired to understand and compare all decision paths in the context of all tree structures, which is a huge challenge for any users. In this paper, we propose a visual analytic system aiming at interpreting random forest models and predictions. In addition to providing users with all the tree information, we summarize the decision paths in random forests, which eventually reflects the working mechanism of the model and reduces users' mental burden of interpretation. To demonstrate the effectiveness of our system, two usage scenarios and a qualitative user study are conducted.
Dik Lun Lee, Weiwei Cui 0001
IEEE Trans. Vis. Comput. Graph.4
2018 How Do Ancestral Traits Shape Family Trees Over Generations?
abstract
Whether and how does the structure of family trees differ by ancestral traits over generations? This is a fundamental question regarding the structural heterogeneity of family trees for the multi-generational transmission research. However, previous work mostly focuses on parent-child scenarios due to the lack of proper tools to handle the complexity of extending the research to multi-generational processes. Through an iterative design study with social scientists and historians, we develop TreeEvo that assists users to generate and test empirical hypotheses for multi-generational research. TreeEvo summarizes and organizes family trees by structural features in a dynamic manner based on a traditional Sankey diagram. A pixel-based technique is further proposed to compactly encode trees with complex structures in each Sankey Node. Detailed information of trees is accessible through a space-efficient visualization with semantic zooming. Moreover, TreeEvo embeds Multinomial Logit Model (MLM) to examine statistical associations between tree structure and ancestral traits. We demonstrate the effectiveness and usefulness of TreeEvo through an in-depth case-study with domain experts using a real-world dataset (containing 54,128 family trees of 126,196 individuals).
Siwei Fu, Hao Dong 0008, Weiwei Cui 0001, Jian Zhao 0010, Huamin Qu
IEEE Trans. Vis. Comput. Graph.3
2018 StreamExplorer: A Multi-Stage System for Visually Exploring Events in Social Streams
abstract
Analyzing social streams is important for many applications, such as crisis management. However, the considerable diversity, increasing volume, and high dynamics of social streams of large events continue to be significant challenges that must be overcome to ensure effective exploration. We propose a novel framework by which to handle complex social streams on a budget PC. This framework features two components: 1) an online method to detect important time periods (i.e., subevents), and 2) a tailored GPU-assisted Self-Organizing Map (SOM) method, which clusters the tweets of subevents stably and efficiently. Based on the framework, we present StreamExplorer to facilitate the visual analysis, tracking, and comparison of a social stream at three levels. At a macroscopic level, StreamExplorer uses a new glyph-based timeline visualization, which presents a quick multi-faceted overview of the ebb and flow of a social stream. At a mesoscopic level, a map visualization is employed to visually summarize the social stream from either a topical or geographical aspect. At a microscopic level, users can employ interactive lenses to visually examine and explore the social stream from different perspectives. Two case studies and a task-based evaluation are used to demonstrate the effectiveness and usefulness of StreamExplorer.Analyzing social streams is important for many applications, such as crisis management. However, the considerable diversity, increasing volume, and high dynamics of social streams of large events continue to be significant challenges that must be overcome to ensure effective exploration. We propose a novel framework by which to handle complex social streams on a budget PC. This framework features two components: 1) an online method to detect important time periods (i.e., subevents), and 2) a tailored GPU-assisted Self-Organizing Map (SOM) method, which clusters the tweets of subevents stably and efficiently. Based on the framework, we present StreamExplorer to facilitate the visual analysis, tracking, and comparison of a social stream at three levels. At a macroscopic level, StreamExplorer uses a new glyph-based timeline visualization, which presents a quick multi-faceted overview of the ebb and flow of a social stream. At a mesoscopic level, a map visualization is employed to visually summarize the social stream from either a topical or geographical aspect. At a microscopic level, users can employ interactive lenses to visually examine and explore the social stream from different perspectives. Two case studies and a task-based evaluation are used to demonstrate the effectiveness and usefulness of StreamExplorer.
Yingcai Wu, Chen Zhu-Tian, Guodao Sun, Xiao Xie, Nan Cao 0001, Shixia Liu, Weiwei Cui 0001
IEEE Trans. Vis. Comput. Graph.7
2018 SkyLens: Visual Analysis of Skyline on Multi-Dimensional Data
abstract
Skyline queries have wide-ranging applications in fields that involve multi-criteria decision making, including tourism, retail industry, and human resources. By automatically removing incompetent candidates, skyline queries allow users to focus on a subset of superior data items (i.e., the skyline), thus reducing the decision-making overhead. However, users are still required to interpret and compare these superior items manually before making a successful choice. This task is challenging because of two issues. First, people usually have fuzzy, unstable, and inconsistent preferences when presented with multiple candidates. Second, skyline queries do not reveal the reasons for the superiority of certain skyline points in a multi-dimensional space. To address these issues, we propose SkyLens, a visual analytic system aiming at revealing the superiority of skyline points from different perspectives and at different scales to aid users in their decision making. Two scenarios demonstrate the usefulness of SkyLens on two datasets with a dozen of attributes. A qualitative study is also conducted to show that users can efficiently accomplish skyline understanding and comparison tasks with SkyLens.
Weiwei Cui 0001, Xinnan Du, Yong Wang 0021, Dik Lun Lee, Huamin Qu
IEEE Trans. Vis. Comput. Graph.3
2017 Visual Analysis of MOOC Forums with iForum
abstract
Discussion forums of Massive Open Online Courses (MOOC) provide great opportunities for students to interact with instructional staff as well as other students. Exploration of MOOC forum data can offer valuable insights for these staff to enhance the course and prepare the next release. However, it is challenging due to the large, complicated, and heterogeneous nature of relevant datasets, which contain multiple dynamically interacting objects such as users, posts, and threads, each one including multiple attributes. In this paper, we present a design study for developing an interactive visual analytics system, called iForum, that allows for effectively discovering and understanding temporal patterns in MOOC forums. The design study was conducted with three domain experts in an iterative manner over one year, including a MOOC instructor and two official teaching assistants. iForum offers a set of novel visualization designs for presenting the three interleaving aspects of MOOC forums (i.e., posts, users, and threads) at three different scales. To demonstrate the effectiveness and usefulness of iForum, we describe a case study involving field experts, in which they use iForum to investigate real MOOC forum data for a course on JAVA programming.
Siwei Fu, Jian Zhao 0010, Weiwei Cui 0001, Huamin Qu
IEEE Trans. Vis. Comput. Graph.3
2017 NameClarifier: A Visual Analytics System for Author Name Disambiguation
abstract
In this paper, we present a novel visual analytics system called NameClarifier to interactively disambiguate author names in publications by keeping humans in the loop. Specifically, NameClarifier quantifies and visualizes the similarities between ambiguous names and those that have been confirmed in digital libraries. The similarities are calculated using three key factors, namely, co-authorships, publication venues, and temporal information. Our system estimates all possible allocations, and then provides visual cues to users to help them validate every ambiguous case. By looping users in the disambiguation process, our system can achieve more reliable results than general data mining models for highly ambiguous cases. In addition, once an ambiguous case is resolved, the result is instantly added back to our system and serves as additional cues for all the remaining unidentified names. In this way, we open up the black box in traditional disambiguation processes, and help intuitively and comprehensively explain why the corresponding classifications should hold. We conducted two use cases and an expert review to demonstrate the effectiveness of NameClarifier.
Qiaomu Shen, Sherry Tongshuang Wu, Huamin Qu, Weiwei Cui 0001
IEEE Trans. Vis. Comput. Graph.6
2017 Evaluation of Graph Sampling: A Visualization Perspective
abstract
Graph sampling is frequently used to address scalability issues when analyzing large graphs. Many algorithms have been proposed to sample graphs, and the performance of these algorithms has been quantified through metrics based on graph structural properties preserved by the sampling: degree distribution, clustering coefficient, and others. However, a perspective that is missing is the impact of these sampling strategies on the resultant visualizations. In this paper, we present the results of three user studies that investigate how sampling strategies influence node-link visualizations of graphs. In particular, five sampling strategies widely used in the graph mining literature are tested to determine how well they preserve visual features in node-link diagrams. Our results show that depending on the sampling strategy used different visual features are preserved. These results provide a complimentary view to metric evaluations conducted in the graph mining literature and provide an impetus to conduct future visualization studies.
Nan Cao 0001, Daniel Archambault, Qiaomu Shen, Huamin Qu, Weiwei Cui 0001
IEEE Trans. Vis. Comput. Graph.6
2016 Online Visual Analytics of Text Streams
abstract
We present an online visual analytics approach to helping users explore and understand hierarchical topic evolution in high-volume text streams. The key idea behind this approach is to identify representative topics in incoming documents and align them with the existing representative topics that they immediately follow (in time). To this end, we learn a set of streaming tree cuts from topic trees based on user-selected focus nodes. A dynamic Bayesian network model has been developed to derive the tree cuts in the incoming topic trees to balance the fitness of each tree cut and the smoothness between adjacent tree cuts. By connecting the corresponding topics at different times, we are able to provide an overview of the evolving hierarchical topics. A sedimentation-based visualization has been designed to enable the interactive analysis of streaming text data from global patterns to local details. We evaluated our method on real-world datasets and the results are generally favorable.
Shixia Liu, Jialun Yin, Xiting Wang, Weiwei Cui 0001, Kelei Cao, Jian Pei 0001
IEEE Trans. Vis. Comput. Graph.4
2016 PieceStack: Toward Better Understanding of Stacked Graphs
abstract
Stacked graphs have been widely adopted in various fields, because they are capable of hierarchically visualizing a set of temporal sequences as well as their aggregation. However, because of visual illusion issues, connections between overly-detailed individual layers and overly-generalized aggregation are intercepted. Consequently, information in this area has yet to be fully excavated. Thus, we present PieceStack in this paper, to reveal the relevance of stacked graphs in understanding intrinsic details of their displayed shapes. This new visual analytic design interprets the ways through which aggregations are generated with individual layers by interactively splitting and re-constructing the stacked graphs. A clustering algorithm is designed to partition stacked graphs into sub-aggregated pieces based on trend similarities of layers. We then visualize the pieces with augmented encoding to help analysts decompose and explore the graphs with respect to their interests. Case studies and a user study are conducted to demonstrate the usefulness of our technique in understanding the formation of stacked graphs.
Sherry Tongshuang Wu, Yingcai Wu, Conglei Shi, Huamin Qu, Weiwei Cui 0001
IEEE Trans. Vis. Comput. Graph.5
2014 Let It Flow: A Static Method for Exploring Dynamic Graphs
abstract
Research into social network analysis has shown that graph metrics, such as degree and closeness, are often used to summarize structural changes in a dynamic graph. However there have been few visual analytics approaches that have been proposed to help analysts study graph evolutions in the context of graph metrics. In this paper, we present a novel approach, called GraphFlow, to visualize dynamic graphs. In contrast to previous approaches that provide users with an animated visualization, GraphFlow offers a static flow visualization that summarizes the graph metrics of the entire graph and its evolution over time. Our solution supports the discovery of high-level patterns that are difficult to identify in an animation or in individual static representations. In addition, GraphFlow provides users with a set of interactions to create filtered views. These views allow users to investigate why a particular pattern has occurred. We showcase the versatility of GraphFlow using two different datasets and describe how it can help users gain insights into complex dynamic graphs.
Weiwei Cui 0001, Xiting Wang, Shixia Liu, Nathalie Henry Riche, Tara M. Madhyastha, Kwan-Liu Ma, Baining Guo
PacificVis1
2014 How Hierarchical Topics Evolve in Large Text Corpora
abstract
Using a sequence of topic trees to organize documents is a popular way to represent hierarchical and evolving topics in text corpora. However, following evolving topics in the context of topic trees remains difficult for users. To address this issue, we present an interactive visual text analysis approach to allow users to progressively explore and analyze the complex evolutionary patterns of hierarchical topics. The key idea behind our approach is to exploit a tree cut to approximate each tree and allow users to interactively modify the tree cuts based on their interests. In particular, we propose an incremental evolutionary tree cut algorithm with the goal of balancing 1) the fitness of each tree cut and the smoothness between adjacent tree cuts; 2) the historical and new information related to user interests. A time-based visualization is designed to illustrate the evolving topics over time. To preserve the mental map, we develop a stable layout algorithm. As a result, our approach can quickly guide users to progressively gain profound insights into evolving hierarchical topics. We evaluate the effectiveness of the proposed method on Amazon's Mechanical Turk and real-world news data. The results show that users are able to successfully analyze evolving topics in text data.
Weiwei Cui 0001, Shixia Liu, Zhuofeng Wu 0002
IEEE Trans. Vis. Comput. Graph.1
2014 A survey on information visualization: recent advances and challenges
Shixia Liu, Weiwei Cui 0001, Yingcai Wu, Mengchen Liu
Vis. Comput.2
2012 Watch the Story Unfold with TextWheel: Visualization of Large-Scale News Streams
abstract
Keyword-based searching and clustering of news articles have been widely used for news analysis. However, news articles usually have other attributes such as source, author, date and time, length, and sentiment which should be taken into account. In addition, news articles and keywords have complicated macro/micro relations, which include relations between news articles (i.e., macro relation), relations between keywords (i.e., micro relation), and relations between news articles and keywords (i.e., macro-micro relation). These macro/micro relations are time varying and pose special challenges for news analysis. In this article we present a visual analytics system for news streams which can bring multiple attributes of the news articles and the macro/micro relations between news streams and keywords into one coherent analytical context, all the while conveying the dynamic natures of news streams. We introduce a new visualization primitive called TextWheel which consists of one or multiple keyword wheels, a document transportation belt, and a dynamic system which connects the wheels and belt. By observing the TextWheel and its content changes, some interesting patterns can be detected. We use our system to analyze several news corpora related to some major companies and the results demonstrate the high potential of our method.
Weiwei Cui 0001, Huamin Qu, Hong Zhou 0004, Wenbin Zhang 0007, Steven Skiena
ACM Trans. Intell. Syst. Technol.1
2012 RankExplorer: Visualization of Ranking Changes in Large Time Series Data
abstract
For many applications involving time series data, people are often interested in the changes of item values over time as well as their ranking changes. For example, people search many words via search engines like Google and Bing every day. Analysts are interested in both the absolute searching number for each word as well as their relative rankings. Both sets of statistics may change over time. For very large time series data with thousands of items, how to visually present ranking changes is an interesting challenge. In this paper, we propose RankExplorer, a novel visualization method based on ThemeRiver to reveal the ranking changes. Our method consists of four major components: 1) a segmentation method which partitions a large set of time series curves into a manageable number of ranking categories; 2) an extended ThemeRiver view with embedded color bars and changing glyphs to show the evolution of aggregation values related to each ranking category over time as well as the content changes in each ranking category; 3) a trend curve to show the degree of ranking changes over time; 4) rich user interactions to support interactive exploration of ranking changes. We have applied our method to some real time series data and the case studies demonstrate that our method can reveal the underlying patterns related to ranking changes which might otherwise be obscured in traditional visualizations.
Conglei Shi, Weiwei Cui 0001, Shixia Liu, Wei Chen 0001, Huamin Qu
IEEE Trans. Vis. Comput. Graph.2
2011 Tracking and Connecting Topics via Incremental Hierarchical Dirichlet Processes
abstract
Much research has been devoted to topic detection from text, but one major challenge has not been addressed: revealing the rich relationships that exist among the detected topics. Finding such relationships is important since many applications are interested in how topics come into being, how they develop, grow, disintegrate, and finally disappear. In this paper, we present a novel method that reveals the connections between topics discovered from the text data. Specifically, our method focuses on how one topic splits into multiple topics, and how multiple topics merge into one topic. We adopt the hierarchical Dirichlet process (HDP) model, and propose an incremental Gibbs sampling algorithm to incrementally derive and refine the labels of clusters. We then characterize the splitting and merging patterns among clusters based on how labels change. We propose a global analysis process that focuses on cluster splitting and merging, and a finer granularity analysis process that helps users to better understand the content of the clusters and the evolution patterns. We also develop a visualization process to present the results.
Zekai Gao, Yangqiu Song, Shixia Liu, Haixun Wang, Yang Chen 0048, Weiwei Cui 0001
ICDM7
2011 Visual analysis of people's mobility pattern from mobile phone data
abstract
The large amount of phone call records from mobile operators in a city can inform us how many people are present in any given area and how many are entering or leaving. Each phone call record usually contains the caller and callee IDs, date and time, and the base station where the phone calls are made. As mobile phones are widely used in our daily life, many human behaviors can be revealed by analyzing mobile phone data. In this paper, we propose a comprehensive visual analysis system which can be used to analyze the population's mobility patterns from millions of phone call records. Our system consists of three major components: 1) visual analysis of user groups in a base station; 2) visual analysis of the mobility patterns on different user groups making phone calls in certain base stations; 3) visual analysis of handoff phone call records. Some well-established visualization techniques such as parallel coordinates and pixel-based representations have been integrated into our system. We also develop a novel visualization schemes, Voronoi-diagram-based visual encoding to reveal the unique features of mobile phone data. We have applied our system to real mobile phone data collected in a large city and obtained some interesting findings regarding people's mobility pattern.
Jiansu Pu, Huamin Qu, Weiwei Cui 0001, Siyuan Liu 0001, Lionel M. Ni
VINCI4
2011 TextFlow: Towards Better Understanding of Evolving Topics in Text
abstract
Understanding how topics evolve in text data is an important and challenging task. Although much work has been devoted to topic analysis, the study of topic evolution has largely been limited to individual topics. In this paper, we introduce TextFlow, a seamless integration of visualization and topic mining techniques, for analyzing various evolution patterns that emerge from multiple topics. We first extend an existing analysis technique to extract three-level features: the topic evolution trend, the critical event, and the keyword correlation. Then a coherent visualization that consists of three new visual components is designed to convey complex relationships between them. Through interaction, the topic mining model and visualization can communicate with each other to help users refine the analysis result and gain insights into the data progressively. Finally, two case studies are conducted to demonstrate the effectiveness and usefulness of TextFlow in helping users understand the major topic evolution patterns in time-varying text data.
Weiwei Cui 0001, Shixia Liu, Conglei Shi, Yangqiu Song, Zekai Gao, Huamin Qu, Xin Tong 0001
IEEE Trans. Vis. Comput. Graph.1
2010 Context preserving dynamic word cloud visualization
abstract
In this paper, we introduce a visualization method that couples a trend chart with word clouds to illustrate temporal content evolutions in a set of documents. Specifically, we use a trend chart to encode the overall semantic evolution of document content over time. In our work, semantic evolution of a document collection is modeled by varied significance of document content, represented by a set of representative keywords, at different time points. At each time point, we also use a word cloud to depict the representative keywords. Since the words in a word cloud may vary one from another over time (e.g., words with increased importance), we use geometry meshes and an adaptive force-directed model to lay out word clouds to highlight the word differences between any two subsequent word clouds. Our method also ensures semantic coherence and spatial stability of word clouds over time. Our work is embodied in an interactive visual analysis system that helps users to perform text analysis and derive insights from a large collection of documents. Our preliminary evaluation demonstrates the usefulness and usability of our work.
Weiwei Cui 0001, Yingcai Wu, Shixia Liu, Furu Wei, Michelle X. Zhou, Huamin Qu
PacificVis1
2010 OpinionSeer: Interactive Visualization of Hotel Customer Feedback
abstract
The rapid development of Web technology has resulted in an increasing number of hotel customers sharing their opinions on the hotel services. Effective visual analysis of online customer opinions is needed, as it has a significant impact on building a successful business. In this paper, we present OpinionSeer, an interactive visualization system that could visually analyze a large collection of online hotel customer reviews. The system is built on a new visualization-centric opinion mining technique that considers uncertainty for faithfully modeling and analyzing customer opinions. A new visual representation is developed to convey customer opinions by augmenting well-established scatterplots and radial visualization. To provide multiple-level exploration, we introduce subjective logic to handle and organize subjective opinions with degrees of uncertainty. Several case studies illustrate the effectiveness and usefulness of OpinionSeer on analyzing relationships among multiple data dimensions and comparing opinions of different groups. Aside from data on hotel customer feedback, OpinionSeer could also be applied to visually analyze customer opinions on other products or services.
Yingcai Wu, Furu Wei, Shixia Liu, Norman Au, Weiwei Cui 0001, Hong Zhou 0004, Huamin Qu
IEEE Trans. Vis. Comput. Graph.5
2009 Splatting the Lines in Parallel Coordinates
abstract
Abstract In this paper, we propose a novel splatting framework for clutter reduction and pattern revealing in parallel coordinates. Our framework consists of two major components: a polyline splatter for cluster detection and a segment splatter for clutter reduction. The cluster detection is performed by splatting the lines one by one into the parallel coordinates plots, and for each splatted line we enhance its neighboring lines and suppress irrelevant ones. To reduce visual clutter caused by line crossings and overlappings in the clustered results, we provide a segment splatter which represents each polyline by one segment and splats these segments with different speeds, colors, and lengths from the leftmost axis to the rightmost axis. Users can interactively control both the polyline splatting and the segment splatting processes to emphasize the features they are interested in. The experimental results demonstrate that our framework can effectively reveal some hidden patterns in parallel coordinates.
Hong Zhou 0004, Weiwei Cui 0001, Huamin Qu, Yingcai Wu, Xiaoru Yuan, Wei Zhuo 0001
Comput. Graph. Forum2
2009 Focus+Context Route Zooming and Information Overlay in 3D Urban Environments
abstract
In this paper we present a novel focus+context zooming technique, which allows users to zoom into a route and its associated landmarks in a 3D urban environment from a 45-degree bird's-eye view. Through the creative utilization of the empty space in an urban environment, our technique can informatively reveal the focus region and minimize distortions to the context buildings. We first create more empty space in the 2D map by broadening the road with an adapted seam carving algorithm. A grid-based zooming technique is then used to enlarge the landmarks to reclaim the created empty space and thus reduce distortions to the other parts. Finally, an occlusion-free route visualization scheme adaptively scales the buildings occluding the route to make the route always visible to users. Our method can be conveniently integrated into Google Earth and Virtual Earth to provide seamless route zooming and help users better explore a city and plan their tours. It can also be used in other applications such as information overlay to a virtual city.
Huamin Qu, Haomian Wang, Weiwei Cui 0001, Yingcai Wu, Ming-Yuen Chan
IEEE Trans. Vis. Comput. Graph.3
2008 Energy-Based Hierarchical Edge Clustering of Graphs
abstract
Effectively visualizing complex node-link graphs which depict relationships among data nodes is a challenging task due to the clutter and occlusion resulting from an excessive amount of edges. In this paper, we propose a novel energy-based hierarchical edge clustering method for node-link graphs. Taking into the consideration of the graph topology, our method first samples graph edges into segments using Delaunay triangulation to generate the control points, which are then hierarchically clustered by energy-based optimization. The edges are grouped according to their positions and directions to improve comprehensibility through abstraction and thus reduce visual clutter. The experimental results demonstrate the effectiveness of our proposed method in clustering edges and providing good high level abstractions of complex graphs.
Hong Zhou 0004, Xiaoru Yuan, Weiwei Cui 0001, Huamin Qu, Baoquan Chen
PacificVis3
2008 Visual Clustering in Parallel Coordinates
abstract
Abstract Parallel coordinates have been widely applied to visualize high‐dimensional and multivariate data, discerning patterns within the data through visual clustering. However, the effectiveness of this technique on large data is reduced by edge clutter. In this paper, we present a novel framework to reduce edge clutter, consequently improving the effectiveness of visual clustering. We exploit curved edges and optimize the arrangement of these curved edges by minimizing their curvature and maximizing the parallelism of adjacent edges. The overall visual clustering is improved by adjusting the shape of the edges while keeping their relative order. The experiments on several representative datasets demonstrate the effectiveness of our approach.
Hong Zhou 0004, Xiaoru Yuan, Huamin Qu, Weiwei Cui 0001, Baoquan Chen
Comput. Graph. Forum4
2008 Geometry-Based Edge Clustering for Graph Visualization
abstract
Graphs have been widely used to model relationships among data. For large graphs, excessive edge crossings make the display visually cluttered and thus difficult to explore. In this paper, we propose a novel geometry-based edge-clustering framework that can group edges into bundles to reduce the overall edge crossings. Our method uses a control mesh to guide the edge-clustering process; edge bundles can be formed by forcing all edges to pass through some control points on the mesh. The control mesh can be generated at different levels of detail either manually or automatically based on underlying graph patterns. Users can further interact with the edge-clustering results through several advanced visualization techniques such as color and opacity enhancement. Compared with other edge-clustering methods, our approach is intuitive, flexible, and efficient. The experiments on some large graphs demonstrate the effectiveness of our method.
Weiwei Cui 0001, Hong Zhou 0004, Huamin Qu, Pak Chung Wong, Xiaoming Li 0001
IEEE Trans. Vis. Comput. Graph.1