VLDB 2026 Research / reviewers in the wild / expert
Shixia Liu
dblp:22/904
· DBLP profile ↗
108ranked-venue papers
18as first author
35since 2021 · last 2026
0000-0003-4499-6420ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 71 · 10 first-author · 29 since 2021Databases, data management, data science and information retrieval · 25 · 7 first-authorArtificial intelligence and machine learning · 19 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unpacking Visual Metaphors in Infographics: A Design Space
Yukai Guo, Lanxi Xiao, Xinhuan Shu, Bongshin Lee, Shixia Liu |
CHI | 6 |
| 2026 | Revealing the Gap: Visual Comparison of Large-Scale Datasets via Multi-Scale Density Difference MapabstractVisual comparison of high-dimensional machine learning datasets helps practitioners identify gaps in data coverage, diagnose distribution shifts, and understand their potential influence on downstream tasks such as classification and object detection. However, the commonly used density map often blurs details and is computationally expensive. We present DiffGrid, a grid-based tool for comparing differences in large datasets. A regularized, grid-based density difference visualization method is developed to enable multi-level analysis of the differences. Interactive zooming and image labels are provided for efficiently exploring differences from overview to detail. We demonstrate the practical value of DiffGrid with two case studies, comparing coresets with full datasets and comparing synthetic infographics with real ones, and validate its effectiveness and usefulness with a quantitative experiment and a user study. Xinyuan Guo, Shixia Liu |
CHI | 4 |
| 2026 | OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI CollaborationabstractOpen-world object detection (OWOD) extends traditional object detection to identifying both known and unknown object, necessitating continuous model adaptation as new annotations emerge. Current approaches face significant limitations: 1) data-hungry training due to reliance on a large number of crowdsourced annotations, 2) susceptibility to "partial feature overfitting," and 3) limited flexibility due to required model architecture modifications. To tackle these issues, we present OW-CLIP, a visual analytics system that provides curated data and enables data-efficient OWOD model incremental training. OW-CLIP implements plug-and-play multimodal prompt tuning tailored for OWOD settings and introduces a novel "Crop-Smoothing" technique to mitigate partial feature overfitting. To meet the data requirements for the training methodology, we propose dual-modal data refinement methods that leverage large language models and cross-modal similarity for data generation and filtering. Simultaneously, we develope a visualization interface that enables users to explore and deliver high-quality annotations-including class-specific visual feature phrases and fine-grained differentiated images. Quantitative evaluation demonstrates that OW-CLIP achieves competitive performance at 89% of state-of-the-art performance while requiring only 3.8% self-generated data, while outperforming SOTA approach when trained with equivalent data volumes. A case study shows the effectiveness of the developed method and the improved annotation quality of our visualization system. Junwen Duan, Ziyao Kang, Shixia Liu, Jiazhi Xia |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | BiasField: Interactive Bias Probing of Machine Learning DatasetsabstractBias in machine learning datasets occurs when certain attributes are unfairly associated, e.g., serious males being mostly linked with law enforcement officers in job-related image datasets. Training models on biased datasets will degrade model performance and lead to fairness issues, particularly for underrepresented groups. Existing bias detection methods mainly focus on explicit biases associated with predefined attributes (e.g., gender and ethnicity) while overlooking implicit biases associated with subtler, dataset-specific attributes (e.g., facial expressions and attire). To address this gap, we present BiasField, an interactive tool that offers a closed-loop workflow for detecting, analyzing, and mitigating bias. Central to BiasField is the adaptive detection of both explicit and implicit biases, a process facilitated by the automatic extraction of the dataset-specific attributes. It then employs a plant-growth metaphor to visualize these biases, enabling structured analysis to identify similar biases and track how they strengthen with additional attributes. Finally, confirmed biases are mitigated through targeted generative data augmentation. A user study, two case studies, and an expert study are conducted to demonstrate its capability to detect, analyze, and mitigate complex biases. Zhen Li 0044, Weikai Yang, Xinhuan Shu, Jiangning Zhu, Hui Zhang 0013, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | RouteFlow: Trajectory-Aware Animated Transitions
Xinyuan Guo, Xinhuan Shu, Lanxi Xiao, Lingyun Yu 0001, Shixia Liu |
CHI | 6 |
| 2025 | Structural-Entropy-Based Sample Selection for Efficient and Effective LearningabstractSample selection improves the efficiency and effectiveness of machine learning models by providing informative and representative samples. Typically, samples can be modeled as a sample graph, where nodes are samples and edges represent their similarities. Most existing methods are based on local information, such as the training difficulty of samples, thereby overlooking global information, such as connectivity patterns. This oversight can result in suboptimal selection because global information is crucial for ensuring that the selected samples well represent the structural properties of the graph. To address this issue, we employ structural entropy to quantify global information and losslessly decompose it from the whole graph to individual nodes using the Shapley value. Based on the decomposition, we present $\textbf{S}$tructural-$\textbf{E}$ntropy-based sample $\textbf{S}$election ($\textbf{SES}$), a method that integrates both global and local information to select informative and representative samples. SES begins by constructing a $k$NN-graph among samples based on their similarities. It then measures sample importance by combining structural entropy (global metric) with training difficulty (local metric). Finally, SES applies importance-biased blue noise sampling to select a set of diverse and representative samples. Comprehensive experiments on three learning scenarios --- supervised learning, active learning, and continual learning --- clearly demonstrate the effectiveness of our method. Tianchi Xie, Jiangning Zhu, Guozu Ma, Minzhi Lin, Wei Chen 0001, Weikai Yang, Shixia Liu |
ICLR | 7 |
| 2025 | InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic ChartsabstractUnderstanding infographic charts with design-driven visual elements (e.g., pictograms, icons) requires both visual recognition and reasoning, posing challenges for multimodal large language models (MLLMs). However, existing visual question answering benchmarks fall short in evaluating these capabilities of MLLMs due to the lack of paired plain charts and visual-element-based questions. To bridge this gap, we introduce InfoChartQA, a benchmark for evaluating MLLMs on infographic chart understanding. It includes 5,642 pairs of infographic and plain charts, each sharing the same underlying data but differing in visual presentations. We further design visual-element-based questions to capture their unique visual designs and communicative intent. Evaluation of 20 MLLMs reveals a substantial performance decline on infographic charts, particularly for visual-element-based questions related to metaphors. The paired infographic and plain charts enable fine-grained error analysis and ablation studies, which highlight new opportunities for advancing MLLMs in infographic chart understanding. We release InfoChartQA at https://github.com/CoolDawnAnt/InfoChartQA. Tianchi Xie, Minzhi Lin, Mengchen Liu, Changjian Chen, Shixia Liu |
NeurIPS | 6 |
| 2025 | Dynamic Color Assignment for Hierarchical DataabstractAssigning discriminable and harmonic colors to samples according to their class labels and spatial distribution can generate attractive visualizations and facilitate data exploration. However, as the number of classes increases, it is challenging to generate a high-quality color assignment result that accommodates all classes simultaneously. A practical solution is to organize classes into a hierarchy and then dynamically assign colors during exploration. However, existing color assignment methods fall short in generating high-quality color assignment results and dynamically aligning them with hierarchical structures. To address this issue, we develop a dynamic color assignment method for hierarchical data, which is formulated as a multi-objective optimization problem. This method simultaneously considers color discriminability, color harmony, and spatial distribution at each hierarchical level. By using the colors of parent classes to guide the color assignment of their child classes, our method further promotes both consistency and clarity across hierarchical levels. We demonstrate the effectiveness of our method in generating dynamic color assignment results with quantitative experiments and a user study. Jiashu Chen, Weikai Yang, Zelin Jia, Lanxi Xiao, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | RuleExplorer: A Scalable Matrix Visualization for Understanding Tree Ensemble ClassifiersabstractThe high performance of tree ensemble classifiers benefits from a large set of rules, which, in turn, makes the models hard to understand. To improve interpretability, existing methods extract a subset of rules for approximation using model reduction techniques. However, by focusing on the reduced rule set, these methods often lose fidelity and ignore anomalous rules that, despite their infrequency, play crucial roles in real-world applications. This paper introduces a scalable visual analysis method to explain tree ensemble classifiers that contain tens of thousands of rules. The key idea is to address the issue of losing fidelity by adaptively organizing the rules as a hierarchy rather than reducing them. To ensure the inclusion of anomalous rules, we develop an anomaly-biased model reduction method to prioritize these rules at each hierarchical level. Synergized with this hierarchical organization of rules, we develop a matrix-based hierarchical visualization to support exploration at different levels of detail. Our quantitative experiments and case studies demonstrate how our method fosters a deeper understanding of both common and anomalous rules, thereby enhancing interpretability without sacrificing comprehensiveness. Zhen Li 0044, Weikai Yang, Jun Yuan 0003, Jing Wu 0004, Changjian Chen, Yao Ming, Fan Yang 0094, Hui Zhang 0013, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2025 | Hierarchical Fuzzy-Cluster-Aware Grid Layout for Large-Scale DataabstractFuzzy clusters, where ambiguous samples belong to multiple clusters, are common in real-world applications. Analyzing such ambiguous samples in large-scale datasets is crucial for practical applications, such as diagnosing machine learning models. A promising method to support such analysis is through hierarchical cluster-aware grid visualizations, which offer high space efficiency and clear cluster perception. However, existing cluster-aware grid layout methods cannot clarify ambiguity among fuzzy clusters, which limits their effectiveness in fuzzy cluster analysis. To tackle this issue, we introduce a hierarchical fuzzy-cluster-aware grid layout method that supports hierarchical exploration of large-scale datasets. Throughout the hierarchical exploration, it is crucial to facilitate fuzzy cluster analysis while maintaining visual continuity for users. To achieve this, we propose a two-step optimization strategy for enhancing cluster perception, clarifying ambiguity, and preserving stability during the exploration. The first step is to create cluster-aware partitions, where each partition corresponds to a cluster. This step focuses on enhancing cluster perception and maintaining the previous shapes and positions of clusters to preserve stability at the cluster level. The second step is to generate a grid layout for each partition. In addition to placing similar samples together, this step also places ambiguous samples near the boundaries to clarify ambiguity and reveal the root causes of their occurrences and maintains the relative positions of the samples in the same cluster to preserve stability at the sample level. Several quantitative experiments and a use case are conducted to demonstrate the effectiveness and usefulness of our method in analyzing large-scale datasets, especially in fuzzy cluster analysis. Yuxing Zhou, Changjian Chen, Zhiyang Shen, Jiangning Zhu, Jiashu Chen, Weikai Yang, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | ReorderBench: A Benchmark for Matrix ReorderingabstractMatrix reordering permutes the rows and columns of a matrix to reveal meaningful visual patterns, such as blocks that represent clusters. A comprehensive collection of matrices, along with a scoring method for measuring the quality of visual patterns in these matrices, contributes to building a benchmark. This benchmark is essential for selecting or designing suitable reordering algorithms for revealing specific patterns. In this paper, we build a matrix-reordering benchmark, ReorderBench, with the goal of evaluating and improving matrix-reordering techniques. This is achieved by generating a large set of representative and diverse matrices and scoring these matrices with a convolution- and entropy-based method. Our benchmark contains 2,835,000 binary matrices and 5,670,000 continuous matrices, each generated to exhibit one of four visual patterns: block, off-diagonal block, star, or band, along with 450 real-world matrices featuring hybrid visual patterns. We demonstrate the usefulness of ReorderBench through three main applications in matrix reordering: 1) evaluating different reordering algorithms, 2) creating a unified scoring model to measure the visual patterns in any matrix, and 3) developing a deep learning model for matrix reordering. Jiangning Zhu, Zhiyang Shen, Fengyuan Tian, Mengchen Liu, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Foundation models meet visualizations: Challenges and opportunitiesabstractRecent studies have indicated that foundation models, such as BERT and GPT, excel at adapting to various downstream tasks. This adaptability has made them a dominant force in building artificial intelligence (AI) systems. Moreover, a new research paradigm has emerged as visualization techniques are incorporated into these models. This study divides these intersections into two research areas: visualization for foundation model (VIS4FM) and foundation model for visualization (FM4VIS). In terms of VIS4FM, we explore the primary role of visualizations in understanding, refining, and evaluating these intricate foundation models. VIS4FM addresses the pressing need for transparency, explainability, fairness, and robustness. Conversely, in terms of FM4VIS, we highlight how foundation models can be used to advance the visualization field itself. The intersection of foundation models with visualizations is promising but also introduces a set of challenges. By highlighting these challenges and promising opportunities, this study aims to provide a starting point for the continued exploration of this research avenue. Weikai Yang, Mengchen Liu, Shixia Liu |
Comput. Vis. Media | 4 |
| 2024 | Enhancing Single-Frame Supervision for Better Temporal Action LocalizationabstractTemporal action localization aims to identify the boundaries and categories of actions in videos, such as scoring a goal in a football match. Single-frame supervision has emerged as a labor-efficient way to train action localizers as it requires only one annotated frame per action. However, it often suffers from poor performance due to the lack of precise boundary annotations. To address this issue, we propose a visual analysis method that aligns similar actions and then propagates a few user-provided annotations (e.g., boundaries, category labels) to similar actions via the generated alignments. Our method models the alignment between actions as a heaviest path problem and the annotation propagation as a quadratic optimization problem. As the automatically generated alignments may not accurately match the associated actions and could produce inaccurate localization results, we develop a storyline visualization to explain the localization results of actions and their alignments. This visualization facilitates users in correcting wrong localization results and misalignments. The corrections are then used to improve the localization results of other actions. The effectiveness of our method in improving localization performance is demonstrated through quantitative evaluation and a case study. Changjian Chen, Jiashu Chen, Weikai Yang, Haoze Wang, Johannes Knittel, Xibin Zhao, Steffen Koch 0001, Thomas Ertl, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2024 | A Unified Interactive Model Evaluation for Classification, Object Detection, and Instance Segmentation in Computer VisionabstractExisting model evaluation tools mainly focus on evaluating classification models, leaving a gap in evaluating more complex models, such as object detection. In this paper, we develop an open-source visual analysis tool, Uni-Evaluator, to support a unified model evaluation for classification, object detection, and instance segmentation in computer vision. The key idea behind our method is to formulate both discrete and continuous predictions in different tasks as unified probability distributions. Based on these distributions, we develop 1) a matrix-based visualization to provide an overview of model performance; 2) a table visualization to identify the problematic data subsets where the model performs poorly; 3) a grid visualization to display the samples of interest. These visualizations work together to facilitate the model evaluation from a global overview to individual samples. Two case studies demonstrate the effectiveness of Uni-Evaluator in evaluating model performance and making informed improvements. Changjian Chen, Yukai Guo, Fengyuan Tian, Shilong Liu 0004, Weikai Yang, Jing Wu 0004, Hang Su 0006, Hanspeter Pfister, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 10 |
| 2024 | Editorial Guest Editors' Introduction
Niklas Elmqvist, Shixia Liu, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Visual Analytics for Machine Learning: A Data Perspective SurveyabstractThe past decade has witnessed a plethora of works that leverage the power of visualization (VIS) to interpret machine learning (ML) models. The corresponding research topic, VIS4ML, keeps growing at a fast pace. To better organize the enormous works and shed light on the developing trend of VIS4ML, we provide a systematic review of these works through this survey. Since data quality greatly impacts the performance of ML models, our survey focuses specifically on summarizing VIS4ML works from the data perspective. First, we categorize the common data handled by ML models into five types, explain the unique features of each type, and highlight the corresponding ML models that are good at learning from them. Second, from the large number of VIS4ML works, we tease out six tasks that operate on these types of data (i.e., data-centric tasks) at different stages of the ML pipeline to understand, diagnose, and refine ML models. Lastly, by studying the distribution of 143 surveyed papers across the five data types, six data-centric tasks, and their intersections, we analyze the prospective research directions and envision future research trends. Junpeng Wang 0001, Shixia Liu, Wei Zhang 0189 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Interactive Reweighting for Mitigating Label Quality IssuesabstractLabel quality issues, such as noisy labels and imbalanced class distributions, have negative effects on model performance. Automatic reweighting methods identify problematic samples with label quality issues by recognizing their negative effects on validation samples and assigning lower weights to them. However, these methods fail to achieve satisfactory performance when the validation samples are of low quality. To tackle this, we develop Reweighter, a visual analysis tool for sample reweighting. The reweighting relationships between validation samples and training samples are modeled as a bipartite graph. Based on this graph, a validation sample improvement method is developed to improve the quality of validation samples. Since the automatic improvement may not always be perfect, a co-cluster-based bipartite graph visualization is developed to illustrate the reweighting relationships and support the interactive adjustments to validation samples and reweighting results. The adjustments are converted into the constraints of the validation sample improvement method to further improve validation samples. We demonstrate the effectiveness of Reweighter in improving reweighting results through quantitative evaluation and two case studies. Weikai Yang, Yukai Guo, Jing Wu 0004, Lan-Zhe Guo, Yufeng Li 0008, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Cluster-Aware Grid LayoutabstractGrid visualizations are widely used in many applications to visually explain a set of data and their proximity relationships. However, existing layout methods face difficulties when dealing with the inherent cluster structures within the data. To address this issue, we propose a cluster-aware grid layout method that aims to better preserve cluster structures by simultaneously considering proximity, compactness, and convexity in the optimization process. Our method utilizes a hybrid optimization strategy that consists of two phases. The global phase aims to balance proximity and compactness within each cluster, while the local phase ensures the convexity of cluster shapes. We evaluate the proposed grid layout method through a series of quantitative experiments and two use cases, demonstrating its effectiveness in preserving cluster structures and facilitating analysis tasks. Yuxing Zhou, Weikai Yang, Jiashu Chen, Changjian Chen, Zhiyang Shen, Lingyun Yu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2023 | In Defence of Visual Analytics Systems: Replies to CriticsabstractThe last decade has witnessed many visual analytics (VA) systems that make successful applications to wide-ranging domains like urban analytics and explainable AI. However, their research rigor and contributions have been extensively challenged within the visualization community. We come in defence of VA systems by contributing two interview studies for gathering critics and responses to those criticisms. First, we interview 24 researchers to collect criticisms the review comments on their VA work. Through an iterative coding and refinement process, the interview feedback is summarized into a list of 36 common criticisms. Second, we interview 17 researchers to validate our list and collect their responses, thereby discussing implications for defending and improving the scientific values and rigor of VA systems. We highlight that the presented knowledge is deep, extensive, but also imperfect, provocative, and controversial, and thus recommend reading with an inclusive and critical eye. We hope our work can provide thoughts and foundations for conducting VA research and spark discussions to promote the research field forward more rigorously and vibrantly. Aoyu Wu, Dazhen Deng, Furui Cheng, Yingcai Wu, Shixia Liu, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Visual Analysis of Neural Architecture Spaces for Summarizing Design PrinciplesabstractRecent advances in artificial intelligence largely benefit from better neural network architectures. These architectures are a product of a costly process of trial-and-error. To ease this process, we develop ArchExplorer, a visual analysis method for understanding a neural architecture space and summarizing design principles. The key idea behind our method is to make the architecture space explainable by exploiting structural distances between architectures. We formulate the pairwise distance calculation as solving an all-pairs shortest path problem. To improve efficiency, we decompose this problem into a set of single-source shortest path problems. The time complexity is reduced from O(kn2N) to O(knN). Architectures are hierarchically clustered according to the distances between them. A circle-packing-based architecture visualization has been developed to convey both the global relationships between clusters and local neighborhoods of the architectures in each cluster. Two case studies and a post-analysis are presented to demonstrate the effectiveness of ArchExplorer in summarizing design principles and selecting better-performing architectures. Jun Yuan 0003, Mengchen Liu, Fengyuan Tian, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Towards Better Caption Supervision for Object DetectionabstractAs training high-performance object detectors requires expensive bounding box annotations, recent methods resort to free-available image captions. However, detectors trained on caption supervision perform poorly because captions are usually noisy and cannot provide precise location information. To tackle this issue, we present a visual analysis method, which tightly integrates caption supervision with object detection to mutually enhance each other. In particular, object labels are first extracted from captions, which are utilized to train the detectors. Then, the objects detected from images are fed into caption supervision for further improvement. To effectively loop users into the object detection process, a node-link-based set visualization supported by a multi-type relational co-clustering algorithm is developed to explain the relationships between the extracted labels and the images with detected objects. The co-clustering algorithm clusters labels and images simultaneously by utilizing both their representations and their relationships. Quantitative evaluations and a case study are conducted to demonstrate the efficiency and effectiveness of the developed method in improving the performance of object detectors. Changjian Chen, Jing Wu 0004, Shouxing Xiang, Song-Hai Zhang, Qifeng Tang, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | Real-Time Visual Analysis of High-Volume Social Media PostsabstractBreaking news and first-hand reports often trend on social media platforms before traditional news outlets cover them. The real-time analysis of posts on such platforms can reveal valuable and timely insights for journalists, politicians, business analysts, and first responders, but the high number and diversity of new posts pose a challenge. In this work, we present an interactive system that enables the visual analysis of streaming social media data on a large scale in real-time. We propose an efficient and explainable dynamic clustering algorithm that powers a continuously updated visualization of the current thematic landscape as well as detailed visual summaries of specific topics of interest. Our parallel clustering strategy provides an adaptive stream with a digestible but diverse selection of recent posts related to relevant topics. We also integrate familiar visual metaphors that are highly interlinked for enabling both explorative and more focused monitoring tasks. Analysts can gradually increase the resolution to dive deeper into particular topics. In contrast to previous work, our system also works with non-geolocated posts and avoids extensive preprocessing such as detecting events. We evaluated our dynamic clustering algorithm and discuss several use cases that show the utility of our system. Johannes Knittel, Steffen Koch 0001, Tan Tang, Wei Chen 0001, Yingcai Wu, Shixia Liu, Thomas Ertl |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | PrefaceabstractThis February 2022 issue of theIEEE Transactions on Visualization and Computer Graphics (TVCG)contains the proceedings of IEEE VIS 2021, held online on October 24-29, 2021, with General Chairs from Tulane University and Universidade de Sao Paulo. With IEEE VIS 2021, the conference series is in its 32nd year. Bongshin Lee, Silvia Miksch, Anders Ynnerman, Anastasia Bezerianos, Jian Chen 0006, Wei Chen 0001, Christopher Collins 0001, Michael Gleicher, M. Eduard Gröller, Alexander Lex, Bernhard Preim, Jinwook Seo, Rüdiger Westermann, Jing Yang 0001, Xiaoru Yuan, Han-Wei Shen, Jean-Daniel Fekete, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 18 |
| 2022 | A Unified Understanding of Deep NLP Models for Text ClassificationabstractThe rapid development of deep natural language processing (NLP) models for text classification has led to an urgent need for a unified understanding of these models proposed individually. Existing methods cannot meet the need for understanding different models in one framework due to the lack of a unified measure for explaining both low-level (e.g., words) and high-level (e.g., phrases) features. We have developed a visual analysis tool, DeepNLPVis, to enable a unified understanding of NLP models for text classification. The key idea is a mutual information-based measure, which provides quantitative explanations on how each layer of a model maintains the information of input words in a sample. We model the intra- and inter-word information at each layer measuring the importance of a word to the final prediction as well as the relationships between words, such as the formation of phrases. A multi-level visualization, which consists of a corpus-level, a sample-level, and a word-level visualization, supports the analysis from the overall training set to individual samples. Two case studies on classification tasks and comparison between models demonstrate that DeepNLPVis can help users effectively identify potential problems caused by samples and model architectures and then make informed improvements. Zhen Li 0044, Xiting Wang, Weikai Yang, Jing Wu 0004, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001, Hui Zhang 0013, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2022 | Revisiting Dimensionality Reduction Techniques for Visual Cluster Analysis: An Empirical StudyabstractDimensionality Reduction (DR) techniques can generate 2D projections and enable visual exploration of cluster structures of high-dimensional datasets. However, different DR techniques would yield various patterns, which significantly affect the performance of visual cluster analysis tasks. We present the results of a user study that investigates the influence of different DR techniques on visual cluster analysis. Our study focuses on the most concerned property types, namely the linearity and locality, and evaluates twelve representative DR techniques that cover the concerned properties. Four controlled experiments were conducted to evaluate how the DR techniques facilitate the tasks of 1) cluster identification, 2) membership identification, 3) distance comparison, and 4) density comparison, respectively. We also evaluated users' subjective preference of the DR techniques regarding the quality of projected clusters. The results show that: 1) Non-linear and Local techniques are preferred in cluster identification and membership identification; 2) Linear techniques perform better than non-linear techniques in density comparison; 3) UMAP (Uniform Manifold Approximation and Projection) and t-SNE (t-Distributed Stochastic Neighbor Embedding) perform the best in cluster identification and membership identification; 4) NMF (Nonnegative Matrix Factorization) has competitive performance in distance comparison; 5) t-SNLE (t-Distributed Stochastic Neighbor Linear Embedding) has competitive performance in density comparison. Jiazhi Xia, Yang Chen 0048, Yunhai Wang, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Diagnosing Ensemble Few-Shot ClassifiersabstractThe base learners and labeled samples (shots) in an ensemble few-shot classifier greatly affect the model performance. When the performance is not satisfactory, it is usually difficult to understand the underlying causes and make improvements. To tackle this issue, we propose a visual analysis method, FSLDiagnotor. Given a set of base learners and a collection of samples with a few shots, we consider two problems: 1) finding a subset of base learners that well predict the sample collections; and 2) replacing the low-quality shots with more representative ones to adequately represent the sample collections. We formulate both problems as sparse subset selection and develop two selection algorithms to recommend appropriate learners and shots, respectively. A matrix visualization and a scatterplot are combined to explain the recommended learners and shots in context and facilitate users in adjusting them. Based on the adjustment, the algorithm updates the recommendation results for another round of improvement. Two case studies are conducted to demonstrate that FSLDiagnotor helps build a few-shot classifier efficiently and increases the accuracy by 12% and 21%, respectively. Weikai Yang, Xi Ye 0003, Xingxing Zhang 0001, Lanxi Xiao, Jiazhi Xia, Zhongyuan Wang 0006, Jun Zhu 0001, Hanspeter Pfister, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2021 | A survey of visual analytics techniques for machine learningabstractVisual analytics for machine learning has recently evolved as one of the most exciting areas in the field of visualization. To better identify which research topics are promising and to learn how to apply relevant techniques in visual analytics, we systematically review 259 papers published in the last ten years together with representative works before 2010. We build a taxonomy, which includes three first-level categories: techniques before model building, techniques during modeling building, and techniques after model building. Each category is further characterized by representative analysis tasks, and each task is exemplified by a set of recent influential works. We also discuss and highlight research challenges and promising potential future research opportunities useful for visual analytics researchers. Jun Yuan 0003, Changjian Chen, Weikai Yang, Mengchen Liu, Jiazhi Xia, Shixia Liu |
Comput. Vis. Media | 6 |
| 2021 | Special Issue on Interactive Visual Analytics for Making Explainable and Accountable Decisionsabstractresearch-article Share on Special Issue on Interactive Visual Analytics for Making Explainable and Accountable Decisions Authors: Cagatay Turkay University of Warwick, Coventry, UK University of Warwick, Coventry, UKView Profile , Tatiana Von Landesberger University of Cologne and University of Rostock, Cologne, Germany University of Cologne and University of Rostock, Cologne, GermanyView Profile , Daniel Archambault Swansea University, Swansea, Wales, UK Swansea University, Swansea, Wales, UKView Profile , Shixia Liu Tsinghua University, Beijing, People’s Republic of China Tsinghua University, Beijing, People’s Republic of ChinaView Profile , Remco Chang Tufts University, Medford, USA Tufts University, Medford, USAView Profile Authors Info & Claims ACM Transactions on Interactive Intelligent SystemsVolume 11Issue 3-4December 2021 Article No.: 17pp 1–4https://doi.org/10.1145/3471903Online:03 September 2021Publication History 0citation187DownloadsMetricsTotal Citations0Total Downloads187Last 12 Months187Last 6 weeks20 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Cagatay Turkay, Tatiana von Landesberger, Daniel Archambault, Shixia Liu, Remco Chang |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2021 | Analyzing the Noise Robustness of Deep Neural NetworksabstractAdversarial examples, generated by adding small but intentionally imperceptible perturbations to normal examples, can mislead deep neural networks (DNNs) to make incorrect predictions. Although much work has been done on both adversarial attack and defense, a fine-grained understanding of adversarial examples is still lacking. To address this issue, we present a visual analysis method to explain why adversarial examples are misclassified. The key is to compare and analyze the datapaths of both the adversarial and normal examples. A datapath is a group of critical neurons along with their connections. We formulate the datapath extraction as a subset selection problem and solve it by constructing and training a neural network. A multi-level visualization consisting of a network-level visualization of data flows, a layer-level visualization of feature maps, and a neuron-level visualization of learned features, has been designed to help investigate how datapaths of adversarial and normal examples diverge and merge in the prediction process. A quantitative evaluation and a case study were conducted to demonstrate the promise of our method to explain the misclassification of adversarial examples. Kelei Cao, Mengchen Liu, Hang Su 0006, Jing Wu 0004, Jun Zhu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Interactive Graph Construction for Graph-Based Semi-Supervised LearningabstractSemi-supervised learning (SSL) provides a way to improve the performance of prediction models (e.g., classifier) via the usage of unlabeled samples. An effective and widely used method is to construct a graph that describes the relationship between labeled and unlabeled samples. Practical experience indicates that graph quality significantly affects the model performance. In this paper, we present a visual analysis method that interactively constructs a high-quality graph for better model performance. In particular, we propose an interactive graph construction method based on the large margin principle. We have developed a river visualization and a hybrid visualization that combines a scatterplot, a node-link diagram, and a bar chart to convey the label propagation of graph-based SSL. Based on the understanding of the propagation, a user can select regions of interest to inspect and modify the graph. We conducted two case studies to showcase how our method facilitates the exploitation of labeled and unlabeled samples for improving model performance. Changjian Chen, Jing Wu 0004, Xiting Wang, Lan-Zhe Guo, Yufeng Li 0008, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | OoDAnalyzer: Interactive Analysis of Out-of-Distribution SamplesabstractOne major cause of performance degradation in predictive models is that the test samples are not well covered by the training data. Such not well-represented samples are called OoD samples. In this article, we propose OoDAnalyzer, a visual analysis approach for interactively identifying OoD samples and explaining them in context. Our approach integrates an ensemble OoD detection method and a grid-based visualization. The detection method is improved from deep ensembles by combining more features with algorithms in the same family. To better analyze and understand the OoD samples in context, we have developed a novelkNN-based grid layout algorithm motivated by Hall's theorem. The algorithm approximates the optimal layout and has O(kN2)O(kN2) time complexity, faster than the grid layout algorithm with overall best performance but O(N3)O(N3) time complexity. Quantitative evaluation and case studies were performed on several datasets to demonstrate the effectiveness and usefulness of OoDAnalyzer. Changjian Chen, Jun Yuan 0003, Yafeng Lu, Yang Liu 0014, Hang Su 0006, Songtao Yuan, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | Visual Analysis of Discrimination in Machine LearningabstractThe growing use of automated decision-making in critical applications, such as crime prediction and college admission, has raised questions about fairness in machine learning. How can we decide whether different treatments are reasonable or discriminatory? In this paper, we investigate discrimination in machine learning from a visual analytics perspective and propose an interactive visualization tool, DiscriLens, to support a more comprehensive analysis. To reveal detailed information on algorithmic discrimination, DiscriLens identifies a collection of potentially discriminatory itemsets based on causal modeling and classification rules mining. By combining an extended Euler diagram with a matrix-based visualization, we develop a novel set visualization to facilitate the exploration and interpretation of discriminatory itemsets. A user study shows that users can interpret the visually encoded information in DiscriLens quickly and accurately. Use cases demonstrate that DiscriLens provides informative guidance in understanding and reducing algorithmic discrimination. Qianwen Wang 0001, Zhenhua Xu 0003, Chen Zhu-Tian, Yong Wang 0021, Shixia Liu, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Interactive Steering of Hierarchical ClusteringabstractHierarchical clustering is an important technique to organize big data for exploratory data analysis. However, existing one-size-fits-all hierarchical clustering methods often fail to meet the diverse needs of different users. To address this challenge, we present an interactive steering method to visually supervise constrained hierarchical clustering by utilizing both public knowledge (e.g., Wikipedia) and private knowledge from users. The novelty of our approach includes 1) automatically constructing constraints for hierarchical clustering using knowledge (knowledge-driven) and intrinsic data distribution (data-driven), and 2) enabling the interactive steering of clustering through a visual interface (user-driven). Our method first maps each data item to the most relevant items in a knowledge base. An initial constraint tree is then extracted using the ant colony optimization algorithm. The algorithm balances the tree width and depth and covers the data items with high confidence. Given the constraint tree, the data items are hierarchically clustered using evolutionary Bayesian rose tree. To clearly convey the hierarchical clustering results, an uncertainty-aware tree visualization has been developed to enable users to quickly locate the most uncertain sub-hierarchies and interactively improve them. The quantitative evaluation and case study demonstrate that the proposed approach facilitates the building of customized clustering trees in an efficient and effective manner. Weikai Yang, Xiting Wang, Wenwen Dou, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Evaluation of Sampling Methods for ScatterplotsabstractGiven a scatterplot with tens of thousands of points or even more, a natural question is which sampling method should be used to create a small but "good" scatterplot for a better abstraction. We present the results of a user study that investigates the influence of different sampling strategies on multi-class scatterplots. The main goal of this study is to understand the capability of sampling methods in preserving the density, outliers, and overall shape of a scatterplot. To this end, we comprehensively review the literature and select seven typical sampling strategies as well as eight representative datasets. We then design four experiments to understand the performance of different strategies in maintaining: 1) region density; 2) class density; 3) outliers; and 4) overall shape in the sampling results. The results show that: 1) random sampling is preferred for preserving region density; 2) blue noise sampling and random sampling have comparable performance with the three multi-class sampling strategies in preserving class density; 3) outlier biased density based sampling, recursive subdivision based sampling, and blue noise sampling perform the best in keeping outliers; and 4) blue noise sampling outperforms the others in maintaining the overall shape of a scatterplot. Jun Yuan 0003, Shouxing Xiang, Jiazhi Xia, Lingyun Yu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Preserving Minority Structures in Graph SamplingabstractSampling is a widely used graph reduction technique to accelerate graph computations and simplify graph visualizations. By comprehensively analyzing the literature on graph sampling, we assume that existing algorithms cannot effectively preserve minority structures that are rare and small in a graph but are very important in graph analysis. In this work, we initially conduct a pilot user study to investigate representative minority structures that are most appealing to human viewers. We then perform an experimental study to evaluate the performance of existing graph sampling algorithms regarding minority structure preservation. Results confirm our assumption and suggest key points for designing a new graph sampling approach named mino-centric graph sampling (MCGS). In this approach, a triangle-based algorithm and a cut-point-based algorithm are proposed to efficiently identify minority structures. A set of importance assessment criteria are designed to guide the preservation of important minority structures. Three optimization objectives are introduced into a greedy strategy to balance the preservation between minority and majority structures and suppress the generation of new minority structures. A series of experiments and case studies are conducted to evaluate the effectiveness of the proposed MCGS. Ying Zhao 0001, Haojin Jiang, Qi'an Chen, Yaqi Qin, Huixuan Xie, Shixia Liu, Zhiguang Zhou, Jiazhi Xia |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2020 | Visual Genealogy of Deep Neural NetworksabstractA comprehensive and comprehensible summary of existing deep neural networks (DNNs) helps practitioners understand the behaviour and evolution of DNNs, offers insights for architecture optimization, and sheds light on the working mechanisms of DNNs. However, this summary is hard to obtain because of the complexity and diversity of DNN architectures. To address this issue, we develop DNN Genealogy, an interactive visualization tool, to offer a visual summary of representative DNNs and their evolutionary relationships. DNN Genealogy enables users to learn DNNs from multiple aspects, including architecture, performance, and evolutionary relationships. Central to this tool is a systematic analysis and visualization of 66 representative DNNs based on our analysis of 140 papers. A directed acyclic graph is used to illustrate the evolutionary relationships among these DNNs and highlight the representative DNNs. A focus + context visualization is developed to orient users during their exploration. A set of network glyphs is used in the graph to facilitate the understanding and comparing of DNNs in the context of the evolution. Case studies demonstrate that DNN Genealogy provides helpful guidance in understanding, applying, and optimizing DNNs. DNN Genealogy is extensible and will continue to be updated to reflect future advances in DNNs. Qianwen Wang 0001, Jun Yuan 0003, Hang Su 0006, Huamin Qu, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2019 | An Interactive Method to Improve Crowdsourced AnnotationsabstractIn order to effectively infer correct labels from noisy crowdsourced annotations, learning-from-crowds models have introduced expert validation. However, little research has been done on facilitating the validation procedure. In this paper, we propose an interactive method to assist experts in verifying uncertain instance labels and unreliable workers. Given the instance labels and worker reliability inferred from a learning-from-crowds model, candidate instances and workers are selected for expert validation. The influence of verified results is propagated to relevant instances and workers through the learning-from-crowds model. To facilitate the validation of annotations, we have developed a confusion visualization to indicate the confusing classes for further exploration, a constrained projection method to show the uncertain labels in context, and a scatter-plot-based visualization to illustrate worker reliability. The three visualizations are tightly integrated with the learning-from-crowds model to provide an iterative and progressive environment for data validation. Two case studies were conducted that demonstrate our approach offers an efficient method for validating and improving crowdsourced annotations. Shixia Liu, Changjian Chen, Yafeng Lu, Fang-Xin Ou-Yang, Bin Wang 0021 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Bridging Text Visualization and Mining: A Task-Driven SurveyabstractVisual text analytics has recently emerged as one of the most prominent topics in both academic research and the commercial world. To provide an overview of the relevant techniques and analysis tasks, as well as the relationships between them, we comprehensively analyzed 263 visualization papers and 4,346 mining papers published between 1992-2017 in two fields: visualization and text mining. From the analysis, we derived around 300 concepts (visualization techniques, mining techniques, and analysis tasks) and built a taxonomy for each type of concept. The co-occurrence relationships between the concepts were also extracted. Our research can be used as a stepping-stone for other researchers to 1) understand a common set of concepts used in this research topic; 2) facilitate the exploration of the relationships between visualization techniques, mining techniques, and analysis tasks; 3) understand the current practice in developing visual text analytics tools; 4) seek potential research opportunities by narrowing the gulf between visualization and mining techniques based on the analysis tasks; and 5) analyze other interdisciplinary research areas in a similar way. We have also contributed a web-based visualization tool for analyzing and understanding research trends and opportunities in visual text analytics. Shixia Liu, Xiting Wang, Christopher Collins 0001, Wenwen Dou, Fang-Xin Ou-Yang, Mennatallah El-Assady, Liu Jiang, Daniel A. Keim |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | Scalable Training of Hierarchical Topic ModelsabstractLarge-scale topic models serve as basic tools for feature extraction and dimensionality reduction in many practical applications. As a natural extension of flat topic models, hierarchical topic models (HTMs) are able to learn topics of different levels of abstraction, which lead to deeper understanding and better generalization than their flat counterparts. However, existing scalable systems for flat topic models cannot handle HTMs, due to their complicated data structures such as trees and concurrent dynamically growing matrices, as well as their susceptibility to local optima. In this paper, we study the hierarchical latent Dirichlet allocation (hLDA) model which is a powerful nonparametric Bayesian HTM. We propose an efficient partially collapsed Gibbs sampling algorithm for hLDA, as well as an initialization strategy to deal with local optima introduced by tree-structured models. We also identify new system challenges in building scalable systems for HTMs, and propose efficient data layout for vectorizing HTM as well as distributed data structures including dynamic matrices and trees. Empirical studies show that our system is 87 times more efficient than the previous open-source implementation for hLDA, and can scale to thousands of CPU cores. We demonstrate our scalability on a 131-million-document corpus with 28 billion tokens, which is 4--5 orders of magnitude larger than previously used corpus. Our distributed implementation can extract 1,722 topics from the corpus with 50 machines in just 7 hours. Jianfei Chen 0001, Jun Zhu 0001, Shixia Liu |
Proc. VLDB Endow. | 4 |
| 2018 | PrefaceabstractEditorsThis January 2018 issue of the IEEE Transactions on Visualization and Computer Graphics contains the proceedings of IEEE VIS 2017, held during 1-6 October 2017. In 2017, IEEE VIS returns to the city of Phoenix, AZ, USA, for the conference's 28th year. The conference will be held at the Hyatt Regency Phoenix hotel. VIS consists of three conferences, held concurrently: the IEEE Visual Analytics Science and Technology Conference (VAST 2017), the IEEE Information Visualization Conference (InfoVis 2017), and the IEEE Scientific Visualization Conference (SciVis 2017). Information on the paper review process is provided along with an overview of each conference. Tim Dwyer, Niklas Elmqvist, Brian D. Fisher, Steven Franconeri, Ingrid Hotz, Robert M. Kirby, Shixia Liu, Tobias Schreck, Xiaoru Yuan |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2018 | PhotoRecomposer: Interactive Photo Recomposition by CroppingabstractWe present a visual analysis method for interactively recomposing a large number of photos based on example photos with high-quality composition. The recomposition method is formulated as a matching problem between photos. The key to this formulation is a new metric for accurately measuring the composition distance between photos. We have also developed an earth-mover-distance-based online metric learning algorithm to support the interactive adjustment of the composition distance based on user preferences. To better convey the compositions of a large number of example photos, we have developed a multi-level, example photo layout method to balance multiple factors such as compactness, aspect ratio, composition distance, stability, and overlaps. By introducing an EulerSmooth-based straightening method, the composition of each photos is clearly displayed. The effectiveness and usefulness of the method has been demonstrated by the experimental results, user study, and case studies. Xiting Wang, Song-Hai Zhang, Shi-Min Hu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Analyzing the Training Processes of Deep Generative ModelsabstractAmong the many types of deep models, deep generative models (DGMs) provide a solution to the important problem of unsupervised and semi-supervised learning. However, training DGMs requires more skill, experience, and know-how because their training is more complex than other types of deep models such as convolutional neural networks (CNNs). We develop a visual analytics approach for better understanding and diagnosing the training process of a DGM. To help experts understand the overall training process, we first extract a large amount of time series data that represents training dynamics (e.g., activation changes over time). A blue-noise polyline sampling scheme is then introduced to select time series samples, which can both preserve outliers and reduce visual clutter. To further investigate the root cause of a failed training process, we propose a credit assignment algorithm that indicates how other neurons contribute to the output of the neuron causing the training failure. Two case studies are conducted with machine learning experts to demonstrate how our approach helps understand and diagnose the training processes of DGMs. We also show how our approach can be directly used to analyze other types of deep models, such as CNNs. Mengchen Liu, Jiaxin Shi, Kelei Cao, Jun Zhu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Visual Diagnosis of Tree Boosting MethodsabstractTree boosting, which combines weak learners (typically decision trees) to generate a strong learner, is a highly effective and widely used machine learning method. However, the development of a high performance tree boosting model is a time-consuming process that requires numerous trial-and-error experiments. To tackle this issue, we have developed a visual diagnosis tool, BOOSTVis, to help experts quickly analyze and diagnose the training process of tree boosting. In particular, we have designed a temporal confusion matrix visualization, and combined it with a t-SNE projection and a tree visualization. These visualization components work together to provide a comprehensive overview of a tree boosting model, and enable an effective diagnosis of an unsatisfactory training process. Two case studies that were conducted on the Otto Group Product Classification Challenge dataset demonstrate that BOOSTVis can provide informative feedback and guidance to improve understanding and diagnosis of tree boosting algorithms. Shixia Liu, Jiannan Xiao, Xiting Wang, Jing Wu 0004, Jun Zhu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2018 | StreamExplorer: A Multi-Stage System for Visually Exploring Events in Social StreamsabstractAnalyzing social streams is important for many applications, such as crisis management. However, the considerable diversity, increasing volume, and high dynamics of social streams of large events continue to be significant challenges that must be overcome to ensure effective exploration. We propose a novel framework by which to handle complex social streams on a budget PC. This framework features two components: 1) an online method to detect important time periods (i.e., subevents), and 2) a tailored GPU-assisted Self-Organizing Map (SOM) method, which clusters the tweets of subevents stably and efficiently. Based on the framework, we present StreamExplorer to facilitate the visual analysis, tracking, and comparison of a social stream at three levels. At a macroscopic level, StreamExplorer uses a new glyph-based timeline visualization, which presents a quick multi-faceted overview of the ebb and flow of a social stream. At a mesoscopic level, a map visualization is employed to visually summarize the social stream from either a topical or geographical aspect. At a microscopic level, users can employ interactive lenses to visually examine and explore the social stream from different perspectives. Two case studies and a task-based evaluation are used to demonstrate the effectiveness and usefulness of StreamExplorer.Analyzing social streams is important for many applications, such as crisis management. However, the considerable diversity, increasing volume, and high dynamics of social streams of large events continue to be significant challenges that must be overcome to ensure effective exploration. We propose a novel framework by which to handle complex social streams on a budget PC. This framework features two components: 1) an online method to detect important time periods (i.e., subevents), and 2) a tailored GPU-assisted Self-Organizing Map (SOM) method, which clusters the tweets of subevents stably and efficiently. Based on the framework, we present StreamExplorer to facilitate the visual analysis, tracking, and comparison of a social stream at three levels. At a macroscopic level, StreamExplorer uses a new glyph-based timeline visualization, which presents a quick multi-faceted overview of the ebb and flow of a social stream. At a mesoscopic level, a map visualization is employed to visually summarize the social stream from either a topical or geographical aspect. At a microscopic level, users can employ interactive lenses to visually examine and explore the social stream from different perspectives. Two case studies and a task-based evaluation are used to demonstrate the effectiveness and usefulness of StreamExplorer. Yingcai Wu, Chen Zhu-Tian, Guodao Sun, Xiao Xie, Nan Cao 0001, Shixia Liu, Weiwei Cui 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | Steering data quality with visual analytics: The complexity challengeabstractData quality management, especially data cleansing, has been extensively studied for many years in the areas of data management and visual analytics. In the paper, we first review and explore the relevant work from the research areas of data management, visual analytics and human-computer interaction. Then for different types of data such as multimedia data, textual data, trajectory data, and graph data, we summarize the common methods for improving data quality by leveraging data cleansing techniques at different analysis stages. Based on a thorough analysis, we propose a general visual analytics framework for interactively cleansing data. Finally, the challenges and opportunities are analyzed and discussed in the context of data and humans. Shixia Liu, Gennady L. Andrienko, Yingcai Wu, Nan Cao 0001, Liu Jiang, Conglei Shi, Yu-Shuen Wang, Seok-Hee Hong 0001 |
Vis. Informatics | 1 |
| 2017 | Improving Learning-from-Crowds through Expert ValidationabstractAlthough several effective learning-from-crowd methods have been developed to infer correct labels from noisy crowdsourced labels, a method for post-processed expert validation is still needed. This paper introduces a semi-supervised learning algorithm that is capable of selecting the most informative instances and maximizing the influence of expert labels. Specifically, we have developed a complete uncertainty assessment to facilitate the selection of the most informative instances. The expert labels are then propagated to similar instances via regularized Bayesian inference. Experiments on both real-world and simulated datasets indicate that given a specific accuracy goal (e.g., 95%) our method reduces expert effort from 39% to 60% compared with the state-of-the-art method. Mengchen Liu, Liu Jiang, Xiting Wang, Jun Zhu 0001, Shixia Liu |
IJCAI | 6 |
| 2017 | PrefaceabstractThe papers in this special issue were presented at IEEE VIS 2016, held during October 23-28, 2016 in Baltimore, MD. VIS contains three conferences, held concurrently: the IEEE Visual Analytics Science and Technology Conference (IEEE VAST 2016), the IEEE Information Visualization Conference (IEEE InfoVis 2016), and the IEEE Scientific Visualization Conference (IEEE SciVis2016). Gennady L. Andrienko, Shixia Liu, John T. Stasko, Niklas Elmqvist, Bongshin Lee, Kwan-Liu Ma, James P. Ahrens, Robert M. Kirby, Jos B. T. M. Roerdink |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Towards Better Analysis of Deep Convolutional Neural NetworksabstractDeep convolutional neural networks (CNNs) have achieved breakthrough performance in many pattern recognition tasks such as image classification. However, the development of high-quality deep models typically relies on a substantial amount of trial-and-error, as there is still no clear understanding of when and why a deep model works. In this paper, we present a visual analytics approach for better understanding, diagnosing, and refining deep CNNs. We formulate a deep CNN as a directed acyclic graph. Based on this formulation, a hybrid visualization is developed to disclose the multiple facets of each neuron and the interactions between them. In particular, we introduce a hierarchical rectangle packing algorithm and a matrix reordering algorithm to show the derived features of a neuron cluster. We also propose a biclustering-based edge bundling method to reduce visual clutter caused by a large number of connections between neurons. We evaluated our method on a set of CNNs and the results are generally favorable. Mengchen Liu, Jiaxin Shi, Zhen Li 0044, Chongxuan Li, Jun Zhu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2017 | Towards better analysis of machine learning models: A visual analytics perspectiveabstractInteractive model analysis, the process of understanding, diagnosing, and refining a machine learning model with the help of interactive visualization, is very important for users to efficiently solve real-world artificial intelligence and data mining problems. Dramatic advances in big data analytics have led to a wide variety of interactive model analysis tasks. In this paper, we present a comprehensive analysis and interpretation of this rapidly developing area. Specifically, we classify the relevant work into three categories: understanding, diagnosis, and refinement. Each category is exemplified by recent influential work. Possible future research opportunities are also explored and discussed. Shixia Liu, Xiting Wang, Mengchen Liu, Jun Zhu 0001 |
Vis. Informatics | 1 |
| 2016 | Tracking Idea Flows between Social GroupsabstractIn many applications, ideas that are described by a set of words often flow between different groups. To facilitate users in analyzing the flow, we present a method to model the flow behaviors that aims at identifying the lead-lag relationships between word clusters of different user groups. In particular, an improved Bayesian conditional cointegration based on dynamic time warping is employed to learn links between words in different groups. A tensor-based technique is developed to cluster these linked words into different clusters (ideas) and track the flow of ideas. The main feature of the tensor representation is that we introduce two additional dimensions to represent both time and lead-lag relationships. Experiments on both synthetic and real datasets show that our method is more effective than methods based on traditional clustering techniques and achieves better accuracy. A case study was conducted to demonstrate the usefulness of our method in helping users understand the flow of ideas between different user groups on social media. Yangxin Zhong, Shixia Liu, Xiting Wang, Jiannan Xiao, Yangqiu Song |
AAAI | 2 |
| 2016 | Spot-tracking lens: A zoomable user interface for animated bubble chartsabstractZoomable user interfaces are widely used in static visualizations and have many benefits. However, they are not well supported in animated visualizations due to problems such as change blindness and information overload. We propose the spot-tracking lens, a new zoomable user interface for animated bubble charts, to tackle these problems. It couples zooming with automatic panning and provides a rich set of auxiliary techniques to enhance its effectiveness. Our preliminary user studies suggested that, besides allowing users to examine detail information, it can be an engaging approach to exploratory analysis for dynamic data. Yueqi Hu, Tom Polk, Jing Yang 0001, Ye Zhao 0003, Shixia Liu |
PacificVis | 5 |
| 2016 | Scaling up Dynamic Topic ModelsabstractDynamic topic models (DTMs) are very effective in discovering topics and capturing their evolution trends in time series data. To do posterior inference of DTMs, existing methods are all batch algorithms that scan the full dataset before each update of the model and make inexact variational approximations with mean-field assumptions. Due to a lack of a more scalable inference algorithm, despite the usefulness, DTMs have not captured large topic dynamics. This paper fills this research void, and presents a fast and parallelizable inference algorithm using Gibbs Sampling with Stochastic Gradient Langevin Dynamics that does not make any unwarranted assumptions. We also present a Metropolis-Hastings based $O(1)$ sampler for topic assignments for each word token. In a distributed environment, our algorithm requires very little communication between workers during sampling (almost embarrassingly parallel) and scales up to large-scale applications. We are able to learn the largest Dynamic Topic Model to our knowledge, and learned the dynamics of 1,000 topics from 2.6 million documents in less than half an hour, and our empirical results show that our algorithm is not only orders of magnitude faster than the baselines but also achieves lower perplexity. Arnab Bhadury, Jianfei Chen 0001, Jun Zhu 0001, Shixia Liu |
WWW | 4 |
| 2016 | Towards building a high-quality microblog-specific Chinese sentiment lexicon
Fangzhao Wu, Yongfeng Huang 0001, Yangqiu Song, Shixia Liu |
Decis. Support Syst. | 4 |
| 2016 | Adaptive Multi-Compositionality for Recursive Neural Network ModelsabstractRecursive neural network models have achieved promising results in many natural language processing tasks. The main difference among these models lies in the composition function, i.e., how to obtain the vector representation for a phrase or sentence using the representations of words it contains. This paper introduces a novel Adaptive Multi-Compositionality (AdaMC) layer to recursive neural network models. The basic idea is to use more than one composition function and adaptively select them depending on input vectors. We develop a general framework to model the semantic composition as a distribution of these composition functions. The composition functions and parameters used for adaptive selection are jointly learnt from the supervision of specific tasks. We integrate AdaMC into existing recursive neural network models and conduct extensive experiments on the Stanford Sentiment Treebank and semantic relation classification task. The experimental results demonstrate that AdaMC improves the performance of recursive neural network models and outperforms the baseline methods. Li Dong 0004, Furu Wei, Ke Xu 0001, Shixia Liu, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | An Uncertainty-Aware Approach for Exploratory Microblog RetrievalabstractAlthough there has been a great deal of interest in analyzing customer opinions and breaking news in microblogs, progress has been hampered by the lack of an effective mechanism to discover and retrieve data of interest from microblogs. To address this problem, we have developed an uncertainty-aware visual analytics approach to retrieve salient posts, users, and hashtags. We extend an existing ranking technique to compute a multifaceted retrieval result: the mutual reinforcement rank of a graph node, the uncertainty of each rank, and the propagation of uncertainty among different graph nodes. To illustrate the three facets, we have also designed a composite visualization with three visual components: a graph visualization, an uncertainty glyph, and a flow map. The graph visualization with glyphs, the flow map, and the uncertainty analysis together enable analysts to effectively find the most uncertain results and interactively refine them. We have applied our approach to several Twitter datasets. Qualitative evaluation and two real-world case studies demonstrate the promise of our approach for retrieving high-quality microblog data. Mengchen Liu, Shixia Liu, Xizhou Zhu, Qinying Liao, Furu Wei, Shimei Pan |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | Guest Editors' Introduction: Special Section on the IEEE Pacific Visualization Symposium 2015abstractThe papers in this special section were presenteda at the 2015 IEEE Pacific Visualization Symposium (PacificVis’15) that was held in Hangzhou from April 14 to 17, 2015. Shixia Liu, Gerik Scheuermann, Shigeo Takahashi, Tim Dwyer, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Online Visual Analytics of Text StreamsabstractWe present an online visual analytics approach to helping users explore and understand hierarchical topic evolution in high-volume text streams. The key idea behind this approach is to identify representative topics in incoming documents and align them with the existing representative topics that they immediately follow (in time). To this end, we learn a set of streaming tree cuts from topic trees based on user-selected focus nodes. A dynamic Bayesian network model has been developed to derive the tree cuts in the incoming topic trees to balance the fitness of each tree cut and the smoothness between adjacent tree cuts. By connecting the corresponding topics at different times, we are able to provide an overview of the evolving hierarchical topics. A sedimentation-based visualization has been designed to enable the interactive analysis of streaming text data from global patterns to local details. We evaluated our method on real-world datasets and the results are generally favorable. Shixia Liu, Jialun Yin, Xiting Wang, Weiwei Cui 0001, Kelei Cao, Jian Pei 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | TopicPanorama: A Full Picture of Relevant TopicsabstractThis paper presents a visual analytics approach to analyzing a full picture of relevant topics discussed in multiple sources, such as news, blogs, or micro-blogs. The full picture consists of a number of common topics covered by multiple sources, as well as distinctive topics from each source. Our approach models each textual corpus as a topic graph. These graphs are then matched using a consistent graph matching method. Next, we develop a level-of-detail (LOD) visualization that balances both readability and stability. Accordingly, the resulting visualization enhances the ability of users to understand and analyze the matched graph from multiple perspectives. By incorporating metric learning and feature selection into the graph matching algorithm, we allow users to interactively modify the graph matching result based on their information needs. We have applied our approach to various types of data, including news articles, tweets, and blog data. Quantitative evaluation and real-world case studies demonstrate the promise of our approach, especially in support of examining a topic-graph-based full picture at different levels of detail. Xiting Wang, Shixia Liu, Jianfei Chen 0001, Jun Zhu 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Clutter-aware label layoutabstractA high-quality label layout is critical for effective information understanding and consumption. Existing labeling methods fail to help users quickly gain an overview of visualized data when the number of labels is large. Visual clutter is a major challenge preventing these methods from being applied to real-world applications. To address this, we propose a context-aware label layout that can measure and reduce visual clutter during the layout process. Our method formulates the clutter model using four factors: confusion, visual connection, distance, and intersection. Based on this clutter model, an effective clutter-aware labeling method has been developed that can generate clear and legible label layouts in different visualizations. We have applied our method to several types of visualizations and the results show promise, especially in support of an uncluttered and informative label layout. Hui Zhang 0013, Mengchen Liu, Shixia Liu |
PacificVis | 4 |
| 2015 | Exploring Topical Lead-Lag across CorporaabstractIdentifying which text corpus leads in the context of a topic presents a great challenge of considerable interest to researchers. Recent research into lead-lag analysis has mainly focused on estimating the overall leads and lags between two corpora. However, real-world applications have a dire need to understand lead-lag patterns both globally and locally. In this paper, we introduce TextPioneer, an interactive visual analytics tool for investigating lead-lag across corpora from the global level to the local level. In particular, we extend an existing lead-lag analysis approach to derive two-level results. To convey multiple perspectives of the results, we have designed two visualizations, a novel hybrid tree visualization that couples a radial space-filling tree with a node-link diagram and a twisted-ladder-like visualization. We have applied our method to several corpora and the evaluation shows promise, especially in support of text comparison at different levels of detail. Shixia Liu, Yang Chen 0048, Jing Yang 0001, Kun Zhou 0001, Steven Mark Drucker |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Evolutionary Bayesian Rose TreesabstractWe present an evolutionary multi-branch tree clustering method to model hierarchical topics and their evolutionary patterns over time. The method builds evolutionary trees in a Bayesian online filtering framework. The tree construction is formulated as an online posterior estimation problem, which well balances both the fitness of the current tree and the smoothness between trees. The state-of-the-art multi-branch clustering method, Bayesian rose trees, is employed to generate a topic tree with a high fitness value. A constraint model is also introduced to preserve the smoothness between trees. A set of comprehensive experiments on real world news data demonstrates that the proposed method better incorporates historical tree information and is more efficient and effective than the traditional evolutionary hierarchical clustering algorithm. In contrast to our previous method[31], we implement two additional baseline algorithms to compare them with our algorithm. We also evaluate the performance of the clustering algorithm based on multiple constraint trees. Furthermore, two case studies are conducted to demonstrate the effectiveness and usefulness of our algorithm in helping users understand the major hierarchical topic evolutionary patterns in text data. Shixia Liu, Xiting Wang, Yangqiu Song, Baining Guo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Automatic Taxonomy Construction from Keywords via Scalable Bayesian Rose TreesabstractIn this paper, we study a challenging problem of deriving a taxonomy from a set of keyword phrases. A solution can benefit many real-world applications because i) keywords give users the flexibility and ease to characterize a specific domain; and ii) in many applications, such as online advertisements, the domain of interest is already represented by a set of keywords. However, it is impossible to create a taxonomy out of a keyword set itself. We argue that additional knowledge and context are needed. To this end, we first use a general-purpose knowledgebase and keyword search to supply the required knowledge and context. Then, we develop a Bayesian approach to build a hierarchical taxonomy for a given set of keywords. We reduce the complexity of previous hierarchical clustering approaches from$O(n^2\; \log\; n)$to$O(n\; \log\; n)$using a nearest-neighbor-based approximation, so that we can derive a domain-specific taxonomy from one million keyword phrases in less than an hour. Finally, we conduct comprehensive large scale experiments to show the effectiveness and efficiency of our approach. A real life example of building an insurance-related web search query taxonomy illustrates the usefulness of our approach for specific domains. Yangqiu Song, Shixia Liu, Xueqing Liu 0001, Haixun Wang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Let It Flow: A Static Method for Exploring Dynamic GraphsabstractResearch into social network analysis has shown that graph metrics, such as degree and closeness, are often used to summarize structural changes in a dynamic graph. However there have been few visual analytics approaches that have been proposed to help analysts study graph evolutions in the context of graph metrics. In this paper, we present a novel approach, called GraphFlow, to visualize dynamic graphs. In contrast to previous approaches that provide users with an animated visualization, GraphFlow offers a static flow visualization that summarizes the graph metrics of the entire graph and its evolution over time. Our solution supports the discovery of high-level patterns that are difficult to identify in an animation or in individual static representations. In addition, GraphFlow provides users with a set of interactions to create filtered views. These views allow users to investigate why a particular pattern has occurred. We showcase the versatility of GraphFlow using two different datasets and describe how it can help users gain insights into complex dynamic graphs. Weiwei Cui 0001, Xiting Wang, Shixia Liu, Nathalie Henry Riche, Tara M. Madhyastha, Kwan-Liu Ma, Baining Guo |
PacificVis | 3 |
| 2014 | How Hierarchical Topics Evolve in Large Text CorporaabstractUsing a sequence of topic trees to organize documents is a popular way to represent hierarchical and evolving topics in text corpora. However, following evolving topics in the context of topic trees remains difficult for users. To address this issue, we present an interactive visual text analysis approach to allow users to progressively explore and analyze the complex evolutionary patterns of hierarchical topics. The key idea behind our approach is to exploit a tree cut to approximate each tree and allow users to interactively modify the tree cuts based on their interests. In particular, we propose an incremental evolutionary tree cut algorithm with the goal of balancing 1) the fitness of each tree cut and the smoothness between adjacent tree cuts; 2) the historical and new information related to user interests. A time-based visualization is designed to illustrate the evolving topics over time. To preserve the mental map, we develop a stable layout algorithm. As a result, our approach can quickly guide users to progressively gain profound insights into evolving hierarchical topics. We evaluate the effectiveness of the proposed method on Amazon's Mechanical Turk and real-world news data. The results show that users are able to successfully analyze evolving topics in text data. Weiwei Cui 0001, Shixia Liu, Zhuofeng Wu 0002 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | LoyalTracker: Visualizing Loyalty Dynamics in Search EnginesabstractThe huge amount of user log data collected by search engine providers creates new opportunities to understand user loyalty and defection behavior at an unprecedented scale. However, this also poses a great challenge to analyze the behavior and glean insights into the complex, large data. In this paper, we introduce LoyalTracker, a visual analytics system to track user loyalty and switching behavior towards multiple search engines from the vast amount of user log data. We propose a new interactive visualization technique (flow view) based on a flow metaphor, which conveys a proper visual summary of the dynamics of user loyalty of thousands of users over time. Two other visualization techniques, a density map and a word cloud, are integrated to enable analysts to gain further insights into the patterns identified by the flow view. Case studies and the interview with domain experts are conducted to demonstrate the usefulness of our technique in understanding user loyalty and switching behavior in search engines. Conglei Shi, Yingcai Wu, Shixia Liu, Hong Zhou 0004, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | EvoRiver: Visual Analysis of Topic Coopetition on Social MediaabstractCooperation and competition (jointly called "coopetition") are two modes of interactions among a set of concurrent topics on social media. How do topics cooperate or compete with each other to gain public attention? Which topics tend to cooperate or compete with one another? Who plays the key role in coopetition-related interactions? We answer these intricate questions by proposing a visual analytics system that facilitates the in-depth analysis of topic coopetition on social media. We model the complex interactions among topics as a combination of carry-over, coopetition recruitment, and coopetition distraction effects. This model provides a close functional approximation of the coopetition process by depicting how different groups of influential users (i.e., "topic leaders") affect coopetition. We also design EvoRiver, a time-based visualization, that allows users to explore coopetition-related interactions and to detect dynamically evolving patterns, as well as their major causes. We test our model and demonstrate the usefulness of our system based on two Twitter data sets (social topics data and business topics data). Guodao Sun, Yingcai Wu, Shixia Liu, Tai-Quan Peng, Jonathan J. H. Zhu, Ronghua Liang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | OpinionFlow: Visual Analysis of Opinion Diffusion on Social MediaabstractIt is important for many different applications such as government and business intelligence to analyze and explore the diffusion of public opinions on social media. However, the rapid propagation and great diversity of public opinions on social media pose great challenges to effective analysis of opinion diffusion. In this paper, we introduce a visual analysis system called OpinionFlow to empower analysts to detect opinion propagation patterns and glean insights. Inspired by the information diffusion model and the theory of selective exposure, we develop an opinion diffusion model to approximate opinion propagation among Twitter users. Accordingly, we design an opinion flow visualization that combines a Sankey graph with a tailored density map in one view to visually convey diffusion of opinions among many users. A stacked tree is used to allow analysts to select topics of interest at different levels. The stacked tree is synchronized with the opinion flow visualization to help users examine and compare diffusion patterns across topics. Experiments and case studies on Twitter data demonstrate the effectiveness and usability of OpinionFlow. Yingcai Wu, Shixia Liu, Kai Yan 0003, Mengchen Liu, Fangzhao Wu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | A survey on information visualization: recent advances and challenges
Shixia Liu, Weiwei Cui 0001, Yingcai Wu, Mengchen Liu |
Vis. Comput. | 1 |
| 2013 | A Hierarchical Aspect-Sentiment Model for Online ReviewsabstractTo help users quickly understand the major opinions from massive online reviews, it is important to automatically reveal the latent structure of the aspects, sentiment polarities, and the association between them. However, there is little work available to do this effectively. In this paper, we propose a hierarchical aspect sentiment model (HASM) to discover a hierarchical structure of aspect-based sentiments from unlabeled online reviews. In HASM, the whole structure is a tree. Each node itself is a two-level tree, whose root represents an aspect and the children represent the sentiment polarities associated with it. Each aspect or sentiment polarity is modeled as a distribution of words. To automatically extract both the structure and parameters of the tree, we use a Bayesian nonparametric model, recursive Chinese Restaurant Process (rCRP), as the prior and jointly infer the aspect-sentiment tree from the review texts. Experiments on two real datasets show that our model is comparable to two other hierarchical topic models in terms of quantitative measures of topic trees. It is also shown that our model achieves better sentence-level classification accuracy than previously proposed aspect-sentiment joint models. Suin Kim, Zheng Chen 0001, Alice Oh, Shixia Liu |
AAAI | 5 |
| 2013 | Lead-lag analysis via sparse co-projection in correlated text streamsabstractCorrelated topical trend detection is very useful in analyzing public and social media influence. In this paper, we propose an algorithm that can both detect the correlation and discover the corresponding keywords that trigger the correlation. To detect the correlation, we use a projection vector to project two text streams onto the same space, and then use a least square cost function to regress one text stream over the other with different time lags. To extract the corresponding keywords, we impose the non-negative sparsity constraints over the projection parameters. In addition, we present an accelerated algorithm based on Nesterov's method to efficiently solve the optimization problem. In our experiments, we use both syntehtic and real data sets to demonstrate the advantages and capabilities of the proposed algorithm over CCA on the follower link prediction problem. Fangzhao Wu, Yangqiu Song, Shixia Liu, Yongfeng Huang 0001 |
CIKM | 3 |
| 2013 | Optimizing temporal topic segmentation for intelligent text visualizationabstractWe are building a topic-based, interactive visual analytic tool that aids users in analyzing large collections of text. To help users quickly discover content evolution and significant content transitions within a topic over time, here we present a novel, constraint-based approach to temporal topic segmentation. Our solution splits a discovered topic into multiple linear, non-overlapping sub-topics along a timeline by satisfying a diverse set of semantic, temporal, and visualization constraints simultaneously. For each derived sub-topic, our solution also automatically selects a set of representative keywords to summarize the main content of the sub-topic. Our extensive evaluation, including a crowd-sourced user study, demonstrates the effectiveness of our method over an existing baseline. Shimei Pan, Michelle X. Zhou, Yangqiu Song, Weihong Qian, Fei Wang 0001, Shixia Liu |
IUI | 6 |
| 2013 | Mining evolutionary multi-branch trees from text streamsabstractUnderstanding topic hierarchies in text streams and their evolution patterns over time is very important in many applications. In this paper, we propose an evolutionary multi-branch tree clustering method for streaming text data. We build evolutionary trees in a Bayesian online filtering framework. The tree construction is formulated as an online posterior estimation problem, which considers both the likelihood of the current tree and conditional prior given the previous tree. We also introduce a constraint model to compute the conditional prior of a tree in the multi-branch setting. Experiments on real world news data demonstrate that our algorithm can better incorporate historical tree information and is more efficient and effective than the traditional evolutionary hierarchical clustering algorithm. Xiting Wang, Shixia Liu, Yangqiu Song, Baining Guo |
KDD | 2 |
| 2013 | A Survey of Visual Analytics Techniques and Applications: State-of-the-Art Research and Future Challenges
Guodao Sun, Yingcai Wu, Ronghua Liang, Shixia Liu |
J. Comput. Sci. Technol. | 4 |
| 2013 | Constrained Text Coclustering with Supervised and Unsupervised ConstraintsabstractIn this paper, we propose a novel constrained coclustering method to achieve two goals. First, we combine information-theoretic coclustering and constrained clustering to improve clustering performance. Second, we adopt both supervised and unsupervised constraints to demonstrate the effectiveness of our algorithm. The unsupervised constraints are automatically derived from existing knowledge sources, thus saving the effort and cost of using manually labeled constraints. To achieve our first goal, we develop a two-sided hidden Markov random field (HMRF) model to represent both document and word constraints. We then use an alternating expectation maximization (EM) algorithm to optimize the model. We also propose two novel methods to automatically construct and incorporate document and word constraints to support unsupervised constrained clustering: 1) automatically construct document constraints based on overlapping named entities (NE) extracted by an NE extractor; 2) automatically construct word constraints based on their semantic distance inferred from WordNet. The results of our evaluation over two benchmark data sets demonstrate the superiority of our approaches against a number of existing approaches. Yangqiu Song, Shimei Pan, Shixia Liu, Furu Wei, Michelle X. Zhou, Weihong Qian |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | StoryFlow: Tracking the Evolution of StoriesabstractStoryline visualizations, which are useful in many applications, aim to illustrate the dynamic relationships between entities in a story. However, the growing complexity and scalability of stories pose great challenges for existing approaches. In this paper, we propose an efficient optimization approach to generating an aesthetically appealing storyline visualization, which effectively handles the hierarchical relationships between entities over time. The approach formulates the storyline layout as a novel hybrid optimization approach that combines discrete and continuous optimization. The discrete method generates an initial layout through the ordering and alignment of entities, and the continuous method optimizes the initial layout to produce the optimal one. The efficient approach makes real-time interactions (e.g., bundling and straightening) possible, thus enabling users to better understand and track how the story evolves. Experiments and case studies are conducted to demonstrate the effectiveness and usefulness of the optimization approach. Shixia Liu, Yingcai Wu, Enxun Wei, Mengchen Liu, Yang Liu 0014 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2013 | ViSizer: A Visualization Resizing FrameworkabstractVisualization resizing is useful for many applications where users may use different display devices. General resizing techniques (e.g., uniform scaling) and image-resizing techniques suffer from several drawbacks, as they do not consider the content of the visualizations. This work introduces ViSizer, a perception-based framework for automatically resizing a visualization to fit any display. We formulate an energy function based on a perception model (feature congestion), which aims to determine the optimal deformation for every local region. We subsequently transform the problem into an optimization problem by the energy function. An efficient algorithm is introduced to iteratively solve the problem, allowing for automatic visualization resizing. Yingcai Wu, Shixia Liu, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2013 | Visual Analysis of Topic Competition on Social MediaabstractHow do various topics compete for public attention when they are spreading on social media? What roles do opinion leaders play in the rise and fall of competitiveness of various topics? In this study, we propose an expanded topic competition model to characterize the competition for public attention on multiple topics promoted by various opinion leaders on social media. To allow an intuitive understanding of the estimated measures, we present a timeline visualization through a metaphoric interpretation of the results. The visual design features both topical and social aspects of the information diffusion process by compositing ThemeRiver with storyline style visualization. ThemeRiver shows the increase and decrease of competitiveness of each topic. Opinion leaders are drawn as threads that converge or diverge with regard to their roles in influencing the public agenda change over time. To validate the effectiveness of the visual analysis techniques, we report the insights gained on two collections of Tweets: the 2012 United States presidential election and the Occupy Wall Street movement. Yingcai Wu, Enxun Wei, Tai-Quan Peng, Shixia Liu, Jonathan J. H. Zhu, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | PIWI: Visually Exploring Graphs Based on Their Community StructureabstractCommunity structure is an important characteristic of many real networks, which shows high concentrations of edges within special groups of vertices and low concentrations between these groups. Community related graph analysis, such as discovering relationships among communities, identifying attribute-structure relationships, and selecting a large number of vertices with desired structural features and attributes, are common tasks in knowledge discovery in such networks. The clutter and the lack of interactivity often hinder efforts to apply traditional graph visualization techniques in these tasks. In this paper, we propose PIWI, a novel graph visual analytics approach to these tasks. Instead of using Node-Link Diagrams (NLDs), PIWI provides coordinated, uncluttered visualizations, and novel interactions based on graph community structure. The novel features, applicability, and limitations of this new technique have been discussed in detail. A set of case studies and preliminary user studies have been conducted with real graphs containing thousands of vertices, which provide supportive evidence about the usefulness of PIWI in community related tasks. Jing Yang 0001, Yujie Liu 0009, Xiaoru Yuan, Ye Zhao 0003, Scott Barlowe, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2012 | Breaking news on twitterabstractAfter the news of Osama Bin Laden's death leaked through Twitter, many people wondered if Twitter would fundamentally change the way we produce, spread, and consume news. In this paper we provide an in-depth analysis of how the news broke and spread on Twitter. We confirm the claim that Twitter broke the news first, and find evidence that Twitter had convinced a large number of its audience before mainstream media confirmed the news. We also discover that attention on Twitter was highly concentrated on a small number of "opinion leaders" and identify three groups of opinion leaders who played key roles in spreading the news: individuals affiliated with media played a large part in breaking the news, mass media brought the news to a wider audience and provided eager Twitter users with content on external sites, and celebrities helped to spread the news and stimulate conversation. Our findings suggest Twitter has great potential as a news medium. Mengdie Hu, Shixia Liu, Furu Wei, Yingcai Wu, John T. Stasko, Kwan-Liu Ma |
CHI | 2 |
| 2012 | Automatic taxonomy construction from keywordsabstractTaxonomies, especially the ones in specific domains, are becoming indispensable to a growing number of applications. State-of-the-art approaches assume there exists a text corpus to accurately characterize the domain of interest, and that a taxonomy can be derived from the text corpus using information extraction techniques. In reality, neither assumption is valid, especially for highly focused or fast-changing domains. In this paper, we study a challenging problem: Deriving a taxonomy from a set of keyword phrases. A solution can benefit many real life applications because i) keywords give users the flexibility and ease to characterize a specific domain; and ii) in many applications, such as online advertisements, the domain of interest is already represented by a set of keywords. However, it is impossible to create a taxonomy out of a keyword set itself. We argue that additional knowledge and contexts are needed. To this end, we first use a general purpose knowledgebase and keyword search to supply the required knowledge and context. Then we develop a Bayesian approach to build a hierarchical taxonomy for a given set of keywords. We reduce the complexity of previous hierarchical clustering approaches from O(n2 log n) to O(n log n), so that we can derive a domain specific taxonomy from one million keyword phrases in less than an hour. Finally, we conduct comprehensive large scale experiments to show the effectiveness and efficiency of our approach. A real life example of building an insurance-related query taxonomy illustrates the usefulness of our approach for specific domains. Xueqing Liu 0001, Yangqiu Song, Shixia Liu, Haixun Wang |
KDD | 3 |
| 2012 | Introduction to the Special Section on Intelligent Visual Interfaces for Text AnalysisabstractElsevier’s Scopus, the largest abstract and citation database of peer-reviewed literature. Search and access research from the science, technology, medicine, social sciences and arts and humanities fields. Shixia Liu, Michelle X. Zhou, Giuseppe Carenini, Huamin Qu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2012 | TIARA: Interactive, Topic-Based Visual Text Summarization and AnalysisabstractWe are building an interactive visual text analysis tool that aids users in analyzing large collections of text. Unlike existing work in visual text analytics, which focuses either on developing sophisticated text analytic techniques or inventing novel text visualization metaphors, ours tightly integrates state-of-the-art text analytics with interactive visualization to maximize the value of both. In this article, we present our work from two aspects. We first introduce an enhanced, LDA-based topic analysis technique that automatically derives a set of topics to summarize a collection of documents and their content evolution over time. To help users understand the complex summarization results produced by our topic analysis technique, we then present the design and development of a time-based visualization of the results. Furthermore, we provide users with a set of rich interaction tools that help them further interpret the visualized results in context and examine the text collection from multiple perspectives. As a result, our work offers three unique contributions. First, we present an enhanced topic modeling technique to provide users with a time-sensitive and more meaningful text summary. Second, we develop an effective visual metaphor to transform abstract and often complex text summarization results into a comprehensible visual representation. Third, we offer users flexible visual interaction tools as alternatives to compensate for the deficiencies of current text summarization techniques. We have applied our work to a number of text corpora and our evaluation shows promise, especially in support of complex text analyses. Shixia Liu, Michelle X. Zhou, Shimei Pan, Yangqiu Song, Weihong Qian, Weijia Cai, Xiaoxiao Lian |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2012 | Whisper: Tracing the Spatiotemporal Process of Information Diffusion in Real TimeabstractWhen and where is an idea dispersed? Social media, like Twitter, has been increasingly used for exchanging information, opinions and emotions about events that are happening across the world. Here we propose a novel visualization design, "Whisper", for tracing the process of information diffusion in social media in real time. Our design highlights three major characteristics of diffusion processes in social media: the temporal trend, social-spatial extent, and community response of a topic of interest. Such social, spatiotemporal processes are conveyed based on a sunflower metaphor whose seeds are often dispersed far away. In Whisper, we summarize the collective responses of communities on a given topic based on how tweets were retweeted by groups of users, through representing the sentiments extracted from the tweets, and tracing the pathways of retweets on a spatial hierarchical layout. We use an efficient flux line-drawing algorithm to trace multiple pathways so the temporal and spatial patterns can be identified even for a bursty event. A focused diffusion series highlights key roles such as opinion leaders in the diffusion process. We demonstrate how our design facilitates the understanding of when and where a piece of information is dispersed and what are the social responses of the crowd, for large-scale events including political campaigns and natural disasters. Initial feedback from domain experts suggests promising use for today's information consumption and dispersion in the wild. Nan Cao 0001, Yu-Ru Lin, David Lazer, Shixia Liu, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2012 | RankExplorer: Visualization of Ranking Changes in Large Time Series DataabstractFor many applications involving time series data, people are often interested in the changes of item values over time as well as their ranking changes. For example, people search many words via search engines like Google and Bing every day. Analysts are interested in both the absolute searching number for each word as well as their relative rankings. Both sets of statistics may change over time. For very large time series data with thousands of items, how to visually present ranking changes is an interesting challenge. In this paper, we propose RankExplorer, a novel visualization method based on ThemeRiver to reveal the ranking changes. Our method consists of four major components: 1) a segmentation method which partitions a large set of time series curves into a manageable number of ranking categories; 2) an extended ThemeRiver view with embedded color bars and changing glyphs to show the evolution of aggregation values related to each ranking category over time as well as the content changes in each ranking category; 3) a trend curve to show the degree of ranking changes over time; 4) rich user interactions to support interactive exploration of ranking changes. We have applied our method to some real time series data and the case studies demonstrate that our method can reveal the underlying patterns related to ranking changes which might otherwise be obscured in traditional visualizations. Conglei Shi, Weiwei Cui 0001, Shixia Liu, Wei Chen 0001, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Tracking and Connecting Topics via Incremental Hierarchical Dirichlet ProcessesabstractMuch research has been devoted to topic detection from text, but one major challenge has not been addressed: revealing the rich relationships that exist among the detected topics. Finding such relationships is important since many applications are interested in how topics come into being, how they develop, grow, disintegrate, and finally disappear. In this paper, we present a novel method that reveals the connections between topics discovered from the text data. Specifically, our method focuses on how one topic splits into multiple topics, and how multiple topics merge into one topic. We adopt the hierarchical Dirichlet process (HDP) model, and propose an incremental Gibbs sampling algorithm to incrementally derive and refine the labels of clusters. We then characterize the splitting and merging patterns among clusters based on how labels change. We propose a global analysis process that focuses on cluster splitting and merging, and a finer granularity analysis process that helps users to better understand the content of the clusters and the evolution patterns. We also develop a visualization process to present the results. Zekai Gao, Yangqiu Song, Shixia Liu, Haixun Wang, Yang Chen 0048, Weiwei Cui 0001 |
ICDM | 3 |
| 2011 | Semantic-Preserving Word Clouds by Seam CarvingabstractAbstract Word clouds are proliferating on the Internet and have received much attention in visual analytics. Although word clouds can help users understand the major content of a document collection quickly, their ability to visually compare documents is limited. This paper introduces a new method to create semantic‐preserving word clouds by leveraging tailored seam carving, a well‐established content‐aware image resizing operator. The method can optimize a word cloud layout by removing a left‐to‐right or top‐to‐bottom seam iteratively and gracefully from the layout. Each seam is a connected path of low energy regions determined by a Gaussian‐based energy function. With seam carving, we can pack the word cloud compactly and effectively, while preserving its overall semantic structure. Furthermore, we design a set of interactive visualization techniques for the created word clouds to facilitate visual text analysis and comparison. Case studies are conducted to demonstrate the effectiveness and usefulness of our techniques. Yingcai Wu, Thomas Provan, Furu Wei, Shixia Liu, Kwan-Liu Ma |
Comput. Graph. Forum | 4 |
| 2011 | TextFlow: Towards Better Understanding of Evolving Topics in TextabstractUnderstanding how topics evolve in text data is an important and challenging task. Although much work has been devoted to topic analysis, the study of topic evolution has largely been limited to individual topics. In this paper, we introduce TextFlow, a seamless integration of visualization and topic mining techniques, for analyzing various evolution patterns that emerge from multiple topics. We first extend an existing analysis technique to extract three-level features: the topic evolution trend, the critical event, and the keyword correlation. Then a coherent visualization that consists of three new visual components is designed to convey complex relationships between them. Through interaction, the topic mining model and visualization can communicate with each other to help users refine the analysis result and gain insights into the data progressively. Finally, two case studies are conducted to demonstrate the effectiveness and usefulness of TextFlow in helping users understand the major topic evolution patterns in time-varying text data. Weiwei Cui 0001, Shixia Liu, Conglei Shi, Yangqiu Song, Zekai Gao, Huamin Qu, Xin Tong 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | Constrained Coclustering for Textual DocumentsabstractIn this paper, we present a constrained co-clustering approach for clustering textual documents. Our approach combines the benefits of information-theoretic co-clustering and constrained clustering. We use a two-sided hidden Markov random field (HMRF) to model both the document and word constraints. We also develop an alternating expectation maximization (EM) algorithm to optimize the constrained co-clustering model. We have conducted two sets of experiments on a benchmark data set: (1) using human-provided category labels to derive document and word constraints for semi-supervised document clustering, and (2) using automatically extracted named entities to derive document constraints for unsupervised document clustering. Compared to several representative constrained clustering and co-clustering approaches, our approach is shown to be more effective for high-dimensional, sparse text data. Yangqiu Song, Shimei Pan, Shixia Liu, Furu Wei, Michelle X. Zhou, Weihong Qian |
AAAI | 3 |
| 2010 | Context preserving dynamic word cloud visualizationabstractIn this paper, we introduce a visualization method that couples a trend chart with word clouds to illustrate temporal content evolutions in a set of documents. Specifically, we use a trend chart to encode the overall semantic evolution of document content over time. In our work, semantic evolution of a document collection is modeled by varied significance of document content, represented by a set of representative keywords, at different time points. At each time point, we also use a word cloud to depict the representative keywords. Since the words in a word cloud may vary one from another over time (e.g., words with increased importance), we use geometry meshes and an adaptive force-directed model to lay out word clouds to highlight the word differences between any two subsequent word clouds. Our method also ensures semantic coherence and spatial stability of word clouds over time. Our work is embodied in an interactive visual analysis system that helps users to perform text analysis and derive insights from a large collection of documents. Our preliminary evaluation demonstrates the usefulness and usability of our work. Weiwei Cui 0001, Yingcai Wu, Shixia Liu, Furu Wei, Michelle X. Zhou, Huamin Qu |
PacificVis | 3 |
| 2010 | Workshop on intelligent visual interfaces for text analysisabstractThis workshop brought together researchers and practitioners from both text analytics and interactive visualization communities to explore, define, and develop intelligent visual interfaces that help enhance the consumption and quality of complex text analysis results. Using this workshop as a starting point, we aim to foster closer, interdisciplinary relationships among researchers from text analytics and interactive visualization communities, so they can combine their expertise together to better tackle the difficult problems that face the text analytics community today. Shixia Liu, Michelle X. Zhou, Giuseppe Carenini, Huamin Qu |
IUI | 1 |
| 2010 | TIARA: a visual exploratory text analytic systemabstractIn this paper, we present a novel exploratory visual analytic system called TIARA (Text Insight via Automated Responsive Analytics), which combines text analytics and interactive visualization to help users explore and analyze large collections of text. Given a collection of documents, TIARA first uses topic analysis techniques to summarize the documents into a set of topics, each of which is represented by a set of keywords. In addition to extracting topics, TIARA derives time-sensitive keywords to depict the content evolution of each topic over time. To help users understand the topic-based summarization results, TIARA employs several interactive text visualization techniques to explain the summarization results and seamlessly link such results to the original text. We have applied TIARA to several real-world applications, including email summarization and patient record analysis. To measure the effectiveness of TIARA, we have conducted several experiments. Our experimental results and initial user feedback suggest that TIARA is effective in aiding users in their exploratory text analytic tasks. Furu Wei, Shixia Liu, Yangqiu Song, Shimei Pan, Michelle X. Zhou, Weihong Qian, Lei Shi 0002 |
KDD | 2 |
| 2010 | Evolutionary hierarchical dirichlet processes for multiple correlated time-varying corporaabstractMining cluster evolution from multiple correlated time-varying text corpora is important in exploratory text analytics. In this paper, we propose an approach called evolutionary hierarchical Dirichlet processes (EvoHDP) to discover interesting cluster evolution patterns from such text data. We formulate the EvoHDP as a series of hierarchical Dirichlet processes~(HDP) by adding time dependencies to the adjacent epochs, and propose a cascaded Gibbs sampling scheme to infer the model. This approach can discover different evolving patterns of clusters, including emergence, disappearance, evolution within a corpus and across different corpora. Experiments over synthetic and real-world multiple correlated time-varying data sets illustrate the effectiveness of EvoHDP on discovering cluster evolution patterns. Yangqiu Song, Changshui Zhang, Shixia Liu |
KDD | 4 |
| 2010 | ContexTour: Contextual Contour Analysis on Dynamic Multi-relational ClusteringabstractHuge amounts of rich context social network data are generated everyday from various applications such as FaceBook and Twitter. These data involve multiple social relations which are community-driven and dynamic in nature. The complex interplay of these characteristics poses tremendous challenges on the users who try to understand the underlying patterns in the social media. We introduce an exploratory analytical framework, ContexTour, which generates visual representations for exploring multiple dimensions of community activities, including relevant topics, representative users and the community-generated content, as well as their evolutions. ContexTour consists of two novel and complementary components: (1) Dynamic Relational Clustering (DRC) that efficiently tracks the community evolution and smoothly adapts to the community changes, and (2) Dynamic Network Contour-map (DNC) that visualizes the community activities and evolutions in various dimensions. In our experiments, we demonstrate ContexTour through case studies on the DBLP dataset. The visual results capture interesting and reasonable evolution in Computer Science research communities. Quantitatively, we show 85-165X performance gain of our DRC algorithm over the baseline method. Yu-Ru Lin, Jimeng Sun 0001, Nan Cao 0001, Shixia Liu |
SDM | 4 |
| 2010 | Midas: integrating public financial dataabstractThe primary goal of the Midas project is to build a system that enables easy and scalable integration of unstructured and semi-structured information present across multiple data sources. As a first step in this direction, we have built a system that extracts and integrates information from regulatory filings submitted to the U.S. Securities and Exchange Commission (SEC) and the Federal Deposit Insurance Corporation (FDIC). Midas creates a repository of entities, events, and relationships by extracting, conceptualizing, integrating, and aggregating data from unstructured and semi-structured documents. This repository enables applications to use the extracted and integrated data in a variety of ways including mashups with other public data and complex risk analysis. Sreeram Balakrishnan, Vivian Chu, Mauricio A. Hernández, C. T. Howard Ho, Rajasekar Krishnamurthy, Shixia Liu, Jan Pieper, Jeffrey S. Pierce, Lucian Popa 0001, Christine Robson, Lei Shi 0002, Ioana Stanoi, Edison L. Ting, Shivakumar Vaithyanathan, Huahai Yang |
SIGMOD Conference | 6 |
| 2010 | iRANK: A rank-learn-combine framework for unsupervised ensemble rankingabstractAbstract The authors address the problem of unsupervised ensemble ranking. Traditional approaches either combine multiple ranking criteria into a unified representation to obtain an overall ranking score or to utilize certain rank fusion or aggregation techniques to combine the ranking results. Beyond the aforementioned “combine‐then‐rank” and “rank‐then‐combine” approaches, the authors propose a novel “rank‐learn‐combine” ranking framework, called Interactive Ranking (iRANK), which allows two base rankers to “teach” each other before combination during the ranking process by providing their own ranking results as feedback to the others to boost the ranking performance. This mutual ranking refinement process continues until the two base rankers cannot learn from each other any more. The overall performance is improved by the enhancement of the base rankers through the mutual learning mechanism. The authors further design two ranking refinement strategies to efficiently and effectively use the feedback based on reasonable assumptions and rational analysis. Although iRANK is applicable to many applications, as a case study, they apply this framework to the sentence ranking problem in query‐focused summarization and evaluate its effectiveness on the DUC 2005 and 2006 data sets. The results are encouraging with consistent and promising improvements. Furu Wei, Wenjie Li 0002, Shixia Liu |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2010 | FacetAtlas: Multifaceted Visualization for Rich Text CorporaabstractDocuments in rich text corpora usually contain multiple facets of information. For example, an article about a specific disease often consists of different facets such as symptom, treatment, cause, diagnosis, prognosis, and prevention. Thus, documents may have different relations based on different facets. Powerful search tools have been developed to help users locate lists of individual documents that are most related to specific keywords. However, there is a lack of effective analysis tools that reveal the multifaceted relations of documents within or cross the document clusters. In this paper, we present FacetAtlas, a multifaceted visualization technique for visually analyzing rich text corpora. FacetAtlas combines search technology with advanced visual analytical tools to convey both global and local patterns simultaneously. We describe several unique aspects of FacetAtlas, including (1) node cliques and multifaceted edges, (2) an optimized density map, and (3) automated opacity pattern enhancement for highlighting visual patterns, (4) interactive context switch between facets. In addition, we demonstrate the power of FacetAtlas through a case study that targets patient education in the health care domain. Our evaluation shows the benefits of this work, especially in support of complex multifaceted data analysis. Nan Cao 0001, Jimeng Sun 0001, Yu-Ru Lin, David Gotz, Shixia Liu, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2010 | OpinionSeer: Interactive Visualization of Hotel Customer FeedbackabstractThe rapid development of Web technology has resulted in an increasing number of hotel customers sharing their opinions on the hotel services. Effective visual analysis of online customer opinions is needed, as it has a significant impact on building a successful business. In this paper, we present OpinionSeer, an interactive visualization system that could visually analyze a large collection of online hotel customer reviews. The system is built on a new visualization-centric opinion mining technique that considers uncertainty for faithfully modeling and analyzing customer opinions. A new visual representation is developed to convey customer opinions by augmenting well-established scatterplots and radial visualization. To provide multiple-level exploration, we introduce subjective logic to handle and organize subjective opinions with degrees of uncertainty. Several case studies illustrate the effectiveness and usefulness of OpinionSeer on analyzing relationships among multiple data dimensions and comparing opinions of different groups. Aside from data on hotel customer feedback, OpinionSeer could also be applied to visually analyze customer opinions on other products or services. Yingcai Wu, Furu Wei, Shixia Liu, Norman Au, Weiwei Cui 0001, Hong Zhou 0004, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2009 | HiMap: Adaptive visualization of large-scale online social networksabstractVisualizing large-scale online social network is a challenging yet essential task. This paper presents HiMap, a system that visualizes it by clustered graph via hierarchical grouping and summarization. HiMap employs a novel adaptive data loading technique to accurately control the visual density of each graph view, and along with the optimized layout algorithm and the two kinds of edge bundling methods, to effectively avoid the visual clutter commonly found in previous social network visualization tools. HiMap also provides an integrated suite of interactions to allow the users to easily navigate the social map with smooth and coherent view transitions to keep their momentum. Finally, we confirm the effectiveness of HiMap algorithms through graph-travesal based evaluations. Lei Shi 0002, Nan Cao 0001, Shixia Liu, Weihong Qian, Jimeng Sun 0001, Ching-Yung Lin |
PacificVis | 3 |
| 2009 | Interactive, topic-based visual text summarization and analysisabstractWe are building an interactive, visual text analysis tool that aids users in analyzing a large collection of text. Unlike existing work in text analysis, which focuses either on developing sophisticated text analytic techniques or inventing novel visualization metaphors, ours is tightly integrating state-of-the-art text analytics with interactive visualization to maximize the value of both. In this paper, we focus on describing our work from two aspects. First, we present the design and development of a time-based, visual text summary that effectively conveys complex text summarization results produced by the Latent Dirichlet Allocation (LDA) model. Second, we describe a set of rich interaction tools that allow users to work with a created visual text summary to further interpret the summarization results in context and examine the text collection from multiple perspectives. As a result, our work offers two unique contributions. First, we provide an effective visual metaphor that transforms complex and even imperfect text summarization results into a comprehensible visual summary of texts. Second, we offer users a set of flexible visual interaction tools as the alternatives to compensate for the deficiencies of current text summarization techniques. We have applied our work to a number of text corpora and our evaluation shows the promise of the work, especially in support of complex text analyses. Shixia Liu, Michelle X. Zhou, Shimei Pan, Weihong Qian, Weijia Cai, Xiaoxiao Lian |
CIKM | 1 |
| 2009 | Topic and keyword re-ranking for LDA-based topic modelingabstractTopic-based text summaries promise to help average users quickly understand a text collection and derive insights. Recent research has shown that the Latent Dirichlet Allocation (LDA) model is one of the most effective approaches to topic analysis. However, the LDA-based results may not be ideal for human understanding and consumption. In this paper, we present several topic and keyword re-ranking approaches that can help users better understand and consume the LDA-derived topics in their text analysis. Our methods process the LDA output based on a set of criteria that model a user's information needs. Our evaluation demonstrates the usefulness of the methods in summarizing several large-scale, real world data sets. Yangqiu Song, Shimei Pan, Shixia Liu, Michelle X. Zhou, Weihong Qian |
CIKM | 3 |
| 2009 | SmallBlue: Social Network Analysis for Expertise Search and Collective IntelligenceabstractSmallBlue is a social networking application that unlocks the valuable business intelligence of 'who knows what?', 'who knows whom?' and 'who knows what about whom' within an organization, without requiring explicit involvement of individuals. The aim of SmallBlue is to locate knowledgeable colleagues, communities, and knowledge networks in companies. The suite also helps users manage their personal networks, and reach out to their extended network (the friends of their friends) to find and access expertise and information. Ching-Yung Lin, Nan Cao 0001, Shixia Liu, Spiros Papadimitriou, Jimeng Sun 0001, Xifeng Yan |
ICDE | 3 |
| 2009 | MultiVis: Content-Based Social Network Exploration through Multi-way Visual AnalysisabstractWith the explosion of social media, scalability becomes a key challenge. There are two main aspects of the problems that arise: 1) data volume: how to manage and analyze huge datasets to efficiently extract patterns, 2) data understanding: how to facilitate understanding of the patterns by users? To address both aspects of the scalability challenge, we present a hybrid approach that leverages two complementary disciplines, data mining and information visualization. In particular, we propose 1) an analytic data model for content-based networks using tensors; 2) an efficient high-order clustering framework for analyzing the data; 3) a scalable context-sensitive graph visualization to present the clusters. We evaluate the proposed methods using both synthetic and real datasets. In terms of computational efficiency, the proposed methods are an order of magnitude faster compared to the baseline. In terms of effectiveness, we present several case studies of real corporate social networks. Jimeng Sun 0001, Spiros Papadimitriou, Ching-Yung Lin, Nan Cao 0001, Shixia Liu, Weihong Qian |
SDM | 5 |
| 2008 | Interactive Visual Analysis of the NSF Funding InformationabstractThis paper presents an interactive visualization toolkit for navigating and analyzing the National Science Foundation (NSF) funding information. Our design builds upon the treemap layout and the stacked graph to contribute customized techniques for visually navigating and interacting with the hierarchical data of NSF programs and proposals, supporting visual search and analysis, and allowing the user to make informed decision. In this visualization toolkit, we propose two visualization techniques to simplify the navigation of the hierarchical data: 2.5 Dimensional treemaps to make the hierarchical structure more easily to be recognized, and labeled treemap to help the user to get a clear overview of the content of the structure and make the internal area of rectangles correspond to the weights of the data set. Furthermore, an incremental layout method is adopted to handle information on a large scale. The improved treemap visualization will help to visually analyze the static funding data and the stacked graph is utilized to analyze the time-series data. Through these visual analysis techniques, research trends of NSF, popular NSF programs are quickly identified. The primary contribution is a demonstration of novel ways to effectively present and analyze NSF funding data. Shixia Liu, Nan Cao 0001 |
PacificVis | 1 |
| 2008 | FanLens: A Visual Toolkit for Dynamically Exploring the Distribution of Hierarchical AttributesabstractRadial, space-filling visualization is very useful for representing the distribution of attributes in hierarchical data; however it also suffers from its drawbacks in terms of view transition, context preservation, thin slices, flexibility and large sized data support. To address these problems, we propose FanLens, an enhancement upon existing approaches with new features like incremental layout and fisheye distortion based selecting. This visual toolkit also features dynamic hierarchy specification, dynamic visual property mapping, smooth animation, etc. We illustrate the effectiveness of our technique with two examples of case study and results from informal user experiments. Xinghua Lou, Shixia Liu |
PacificVis | 2 |
| 2007 | An integrated system for building enterprise taxonomies
Li Zhang 0007, Shixia Liu |
Inf. Retr. | 3 |
| 2006 | The visual funding navigator: analysis of the NSF funding informationabstractThis paper presents an interactive visualization toolkit for navigating and analyzing the National Science Foundation (NSF) funding information. Our design builds upon an improved 2.5D treemap layout and the stacked graph to contribute customized techniques for visually navigating and interacting with the hierarchical data of NSF programs and proposals. Furthermore, an incremental layout method is adopted to handle information on a large scale. The improved treemap visualization will help to visually analyze the static funding related data and the stacked graph is utilized to analyze the time-series data. Through these visual analysis techniques, research trends of NSF, popular NSF programs are quickly identified. Shixia Liu, Nan Cao 0001, Hui Su |
CIKM | 1 |
| 2005 | An LOD Model for Graph Visualization and Its Application in Web Navigation
Shixia Liu, Wenyin Liu |
APWeb | 1 |
| 2004 | InfoAnalyzer: a computer-aided tool for building enterprise taxonomiesabstractIn this paper we study the problem of collecting training samples for building enterprise taxonomies. We develop a computer-aided tool named InfoAnalyzer, which can effectively assist the enterprise to prepare large set of samples used for machine learning in text categorization. In our system, the enterprise category tree is initially defined by some keywords, then the Google search engine is used to construct a small set of labeled documents, and topic tracking algorithm based on document length normalization is applied to enlarge the training corpus on the bases of the seed stories. Furthermore, we design a method to check the consistency of the training corpus. Experiments show that the training corpus is good enough for statistical classification methods and meets human's requirements as well. Li Zhang 0007, Shixia Liu |
CIKM | 2 |