Han-Wei Shen

dblp:61/6829 · DBLP profile ↗
← Back
172ranked-venue papers
19as first author
52since 2021 · last 2026
0000-0002-1211-2320ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 136 · 13 first-author · 42 since 2021Human-computer interaction and ubiquitous computing · 20 · 6 first-author · 4 since 2021Systems, architecture and hardware · 10Artificial intelligence and machine learning · 8 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Beauty in the Eye of AI: Aligning LLMs and Vision Models with Human Aesthetics in Network Visualization
abstract
Abstract Network visualization has traditionally relied on heuristic metrics, such as stress, under the assumption that optimizing them leads to aesthetic and informative layouts. However, no single metric consistently produces the most effective results. A data‐driven alternative is to learn from human preferences, where labelers select their favored visualization among multiple layouts of the same graphs. These human‐preference labels can then be used to train a generative model that approximates human aesthetic preferences. However, obtaining human labels at scale is costly and time‐consuming. As a result, this generative approach has so far been tested only with machine‐labeled data [WYHS24]. In this paper, we explore the use of large language models (LLMs) and vision models (VMs) as proxies for human judgment. Through a carefully designed user study involving 27 participants, we curated a large set of human preference labels. We used this data both to better understand human preferences and to bootstrap LLM/VM labelers. We show that prompt engineering that combines few‐shot examples and diverse input formats, such as image embeddings, significantly improves LLM‐‐human alignment, and additional filtering by the confidence score of the LLM pushes the alignment to human‐‐human levels. Furthermore, we demonstrate that carefully trained VMs can achieve VM‐human alignment at a level comparable to that between human labelers. Our results suggest that AI can feasibly serve as a scalable proxy for human labelers.
Han-Wei Shen, Yifan Hu 0001
Comput. Graph. Forum4
2026 VizGenie: Toward Self-Refining, Domain-Aware Workflows for Next-Generation Scientific Visualization
abstract
We present VizGenie, a self-improving, agentic framework that advances scientific visualization through large language model (LLM) by orchestrating of a collection of domain-specific and dynamically generated modules. Users initially access core functionalities-such as threshold-based filtering, slice extraction, and statistical analysis-through pre-existing tools. For tasks beyond this baseline, VizGenie autonomously employs LLMs to generate new visualization scripts (e.g., VTK Python code), expanding its capabilities on-demand. Each generated script undergoes automated backend validation and is seamlessly integrated upon successful testing, continuously enhancing the system's adaptability and robustness. A distinctive feature of VizGenie is its intuitive natural language interface, allowing users to issue high-level feature-based queries (e.g., "visualize the skull" or "highlight tissue boundaries"). The system leverages image-based analysis and visual question answering (VQA) via fine-tuned vision models to interpret these queries precisely, bridging domain expertise and technical implementation. Additionally, users can interactively query generated visualizations through VQA, facilitating deeper exploration. Reliability and reproducibility are further strengthened by Retrieval-Augmented Generation (RAG), providing context-driven responses while maintaining comprehensive provenance records. Evaluations on complex volumetric datasets demonstrate significant reductions in cognitive overhead for iterative visualization tasks. By integrating curated domain-specific tools with LLM-driven flexibility, VizGenie not only accelerates insight generation but also establishes a sustainable, continuously evolving visualization practice. The resulting platform dynamically learns from user interactions, consistently enhancing support for feature-centric exploration and reproducible research in scientific visualization.
Ayan Biswas 0001, Terece L. Turton, Nishath Rajiv Ranasinghe, Shawn M. Jones, Bradley C. Love, William M. Jones, Aric A. Hagberg, Han-Wei Shen, Nathan DeBardeleben, Earl Lawrence
IEEE Trans. Vis. Comput. Graph.8
2026 FLUID: A Neural Operator-Based Framework for Learning Multi-Fidelity of Unstructured Data
abstract
With increasing computational power, scientists often employ high-fidelity simulations to study complex scientific phenomena. However, these simulations remain slow and costly in terms of computation and storage, prompting the use of faster low-fidelity alternatives. Yet, the distribution gap between low- and high-fidelity data, due to missing fine-scale details and simplified physics, hampers accurate scientific understanding. To overcome these challenges, we propose a neural operator-based framework for multi-fidelity prediction on unstructured data. Our framework leverages a graph neural operator to effectively map low-fidelity data to the high-fidelity counterparts. It further incorporates a spectral-based module that captures fine-scale details, enhancing the reconstruction of high-fidelity fields. Extensive experiments across diverse datasets demonstrate that our method consistently surpasses strong baselines, including state-of-the-art neural operators and learning methods for unstructured data, highlighting its robustness and effectiveness in bridging fidelity gaps.
Yi-Tang Chen, Xihaier Luo, Wei Xu 0020, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.4
2025 Completing A Systematic Review in Hours instead of Months with Interactive AI Agents
abstract
Systematic reviews (SRs) are vital for evidencebased practice in high stakes disciplines, such as healthcare, but are often impeded by laborintensive and lengthy processes that can span months.Due to the high demand for domain expertise, existing automatic summarization methods fail to accurately identify relevant studies and generate high-quality summaries.To that end, we introduce InsightAgent, a human-centered interactive AI agent powered by large language models that revolutionizes the systematic review workflow.In-sightAgent partitions a large literature corpus based on semantics and employs a multi-agent design for more focused processing of literature, leading to significant improvement in the quality of generated SRs.InsightAgent also provides intuitive visualizations of the corpus and agent trajectories, allowing users to effortlessly monitor the actions of the agent and provide real-time feedback based on their expertise.Our user studies with 9 medical professionals demonstrate that the visualization and interaction mechanisms can effectively improve the quality of synthesized SRs by 27.2%, reaching 79.7% of human-written quality.At the same time, user satisfaction is improved by 34.4%.With InsightAgent, it only takes a clinician about 1.5 hours, rather than months, to complete a high-quality systematic review.InsightAgent demonstrates great potential in facilitating more timely and informed decisionmaking in high stake application scenarios 1 .
Yu Su 0001, Po-Yin Yen, Han-Wei Shen
ACL (1)5
2025 AMGSRN++: Improved Adaptive SRN for Scientific Visualization
abstract
We present AMGSRN++, which advances previous state of the art APMGSRN along three key directions. First, we implement efficient CUDA kernels to fuse the encoding operation into a single kernel, reducing VRAM requirement by over $50 \%$ improving throughput, enabling faster training and rendering with the availability for lower-end hardware to perform neural volume rendering efficiently. Second, we introduce a compression-aware training strategy for efficient feature grid compression when saving, reducing storage costs by $80 \%$. Lastly, we extend the method to time-varying data with a 3D+time approach, allowing parallel training with no dependence between timesteps for highly efficient model fitting. We extend the previously released rendering tool to support the new model, including seamless time-varying dataset visualization. As a result, time-varying datasets over 100 GB can be rendered in real time on consumer hardware with as little as 1 GB of VRAM, and using only 88 MB of storage space. Comparisons with state of the art compressors and other SRNs are provided, displaying continued strong representation capability and higher compressive capabilities. All code is released publicly at https://github.com/skywolf829/AMGSRN.
Skylar W. Wurster, Han-Wei Shen
PacificVis2
2025 Explorable INR: An Implicit Neural Representation for Ensemble Simulation Enabling Efficient Spatial and Parameter Exploration
abstract
With the growing computational power available for high-resolution ensemble simulations in scientific fields such as cosmology and oceanology, storage and computational demands present significant challenges. Current surrogate models fall short in the flexibility of point- or region-based predictions as the entire field reconstruction is required for each parameter setting, hence hindering the efficiency of parameter space exploration. Limitations exist in capturing physical attribute distributions and pinpointing optimal parameter configurations. In this work, we propose Explorable INR, a novel implicit neural representation-based surrogate model, designed to facilitate exploration and allow point-based spatial queries without computing full-scale field data. In addition, to further address computational bottlenecks of spatial exploration, we utilize probabilistic affine forms (PAFs) for uncertainty propagation through Explorable INR to obtain statistical summaries, facilitating various ensemble analysis and visualization tasks that are expensive with existing models. Furthermore, we reformulate the parameter exploration problem as optimization tasks using gradient descent and KL divergence minimization that ensures scalability. We demonstrate that the Explorable INR with the proposed approach for spatial and parameter exploration can significantly reduce computation and memory costs while providing effective ensemble analysis.
Yi-Tang Chen, Neng Shi, Xihaier Luo, Wei Xu 0020, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.6
2025 Improving Efficiency of Iso-Surface Extraction on Implicit Neural Representations Using Uncertainty Propagation
abstract
Implicit Neural representations (INRs) are widely used for scientific data reduction and visualization by modeling the function that maps a spatial location to a data value. Without any prior knowledge about the spatial distribution of values, we are forced to sample densely from INRs to perform visualization tasks like iso-surface extraction which can be very computationally expensive. Recently, range analysis has shown promising results in improving the efficiency of geometric queries, such as ray casting and hierarchical mesh extraction, on INRs for 3D geometries by using arithmetic rules to bound the output range of the network within a spatial region. However, the analysis bounds are often too conservative for complex scientific data. In this article, we present an improved technique for range analysis by revisiting the arithmetic rules and analyzing the probability distribution of the network output within a spatial region. We model this distribution efficiently as a Gaussian distribution by applying the central limit theorem. Excluding low probability values, we are able to tighten the output bounds, resulting in a more accurate estimation of the value range, and hence more accurate identification of iso-surface cells and more efficient iso-surface extraction on INRs. Our approach demonstrates superior performance in terms of the iso-surface extraction time on four datasets compared to the original range analysis method and can also be generalized to other geometric query tasks.
Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2025 VADIS: A Visual Analytics Pipeline for Dynamic Document Representation and Information-Seeking
abstract
In the biomedical domain, visualizing the document embeddings of an extensive corpus has been widely used in information-seeking tasks. However, three key challenges with existing visualizations make it difficult for clinicians to find information efficiently. First, the document embeddings used in these visualizations are generated statically by pretrained language models, which cannot adapt to the user's evolving interest. Second, existing document visualization techniques cannot effectively display how the documents are relevant to users' interest, making it difficult for users to identify the most pertinent information. Third, existing embedding generation and visualization processes suffer from a lack of interpretability, making it difficult to understand, trust and use the result for decision-making. In this paper, we present a novel visual analytics pipeline for user-driven document representation and iterative information seeking (VADIS). VADIS introduces a prompt-based attention model (PAM) that generates dynamic document embedding and document relevance adjusted to the user's query. To effectively visualize these two pieces of information, we design a new document map that leverages a circular grid layout to display documents based on both their relevance to the query and the semantic similarity. Additionally, to improve the interpretability, we introduce a corpus-level attention visualization method to improve the user's understanding of the model focus and to enable the users to identify potential oversight. This visualization, in turn, empowers users to refine, update and introduce new queries, thereby facilitating a dynamic and iterative information-seeking experience. We evaluated VADIS quantitatively and qualitatively on a real-world dataset of biomedical research papers to demonstrate its effectiveness.
Yamei Tu, Po-Yin Yen, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.4
2025 Message from the Editor-in-Chief
abstract
Welcome to the January 2025 issue of the IEEE Transactions on Visualization and Computer Graphics (TVCG). This is my second IEEE VIS special issue as Editor-in-Chief, and I am very excited to introduce it to you all. As in the previous three years, the papers submitted to IEEE VIS were categorized into six major research subareas: Theoretical & Empirical (112 papers), Applications (154), Systems & Rendering (51), Representations & Interaction (110), Data Transformations (53) and Analytics & Decisions (77). The conference took place in St. Pete Beach, Florida, USA, during October 13-18, 2024. Included in this special issue are the top 124 papers selected by the Program Committee from a total of 557 submissions.
Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.1
2025 SurroFlow: A Flow-Based Surrogate Model for Parameter Space Exploration and Uncertainty Quantification
abstract
Existing deep learning-based surrogate models facilitate efficient data generation, but fall short in uncertainty quantification, efficient parameter space exploration, and reverse prediction. In our work, we introduce SurroFlow, a novel normalizing flow-based surrogate model, to learn the invertible transformation between simulation parameters and simulation outputs. The model not only allows accurate predictions of simulation outcomes for a given simulation parameter but also supports uncertainty quantification in the data generation process. Additionally, it enables efficient simulation parameter recommendation and exploration. We integrate SurroFlow and a genetic algorithm as the backend of a visual interface to support effective user-guided ensemble simulation exploration and visualization. Our framework significantly reduces the computational costs while enhancing the reliability and exploration capabilities of scientific surrogate models.
Jingyi Shen, Yuhan Duan, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.3
2025 IEEE VR 2025 Introducing the Special Issue
Han-Wei Shen, Kiyoshi Kiyokawa, Maud Marchal
IEEE Trans. Vis. Comput. Graph.1
2025 IEEE ISMAR 2025 Introducing the Special Issue
Han-Wei Shen, Kiyoshi Kiyokawa, Maud Marchal
IEEE Trans. Vis. Comput. Graph.1
2025 2024 VGTC Visualization Technical Achievement Award
abstract
The 2024 VGTC Visualization Technical Achievement Award goes to Han-Wei Shen for his research on extreme-scale and multivariate time-varying data visualization, uncertainty visualization, and novel approaches to inclusion of AI in scientific workflows. The 2024 VGTC Visualization Technical Achievement Award goes to Bongshin Lee for her groundbreaking contributions in advancing visualization technologies through intuitive interaction designs, data-driven storytelling, and integration of machine learning.
Han-Wei Shen, Bongshin Lee
IEEE Trans. Vis. Comput. Graph.1
2025 Regularized Multi-Decoder Ensemble for an Error-Aware Scene Representation Network
abstract
Feature grid Scene Representation Networks (SRNs) have been applied to scientific data as compact functional surrogates for analysis and visualization. As SRNs are black-box lossy data representations, assessing the prediction quality is critical for scientific visualization applications to ensure that scientists can trust the information being visualized. Currently, existing architectures do not support inference time reconstruction quality assessment, as coordinate-level errors cannot be evaluated in the absence of ground truth data. By employing the uncertain neural network architecture in feature grid SRNs, we obtain prediction variances during inference time to facilitate confidence-aware data reconstruction. Specifically, we propose a parameter-efficient multi-decoder SRN (MDSRN) architecture consisting of a shared feature grid with multiple lightweight multilayer perceptron decoders. MDSRN can generate a set of plausible predictions for a given input coordinate to compute the mean as the prediction of the multi-decoder ensemble and the variance as a confidence score. The coordinate-level variance can be rendered along with the data to inform the reconstruction quality, or be integrated into uncertainty-aware volume visualization algorithms. To prevent the misalignment between the quantified variance and the prediction quality, we propose a novel variance regularization loss for ensemble learning that promotes the Regularized multi-decoder SRN (RMDSRN) to obtain a more reliable variance that correlates closely to the true model error. We comprehensively evaluate the quality of variance quantification and data reconstruction of Monte Carlo Dropout (MCD), Mean Field Variational Inference (MFVI), Deep Ensemble (DE), and Predicting Variance (PV) in comparison with our proposed MDSRN and RMDSRN applied to state-of-the-art feature grid SRNs across diverse scalar field datasets. We demonstrate that RMDSRN attains the most accurate data reconstruction and competitive variance-error correlation among uncertain SRNs under the same neural network parameter budgets. Furthermore, we present an adaptation of uncertainty-aware volume rendering and shed light on the potential of incorporating uncertain predictions in improving the quality of volume rendering for uncertain SRNs. Through ablation studies on the regularization strength and decoder count, we show that MDSRN and RMDSRN are expected to perform sufficiently well with a default configuration without requiring customized hyperparameter settings for different datasets.
Tianyu Xiong, Skylar W. Wurster, Hanqi Guo 0001, Tom Peterka, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.5
2024 USE: Universal Segment Embeddings for Open-Vocabulary Image Segmentation
abstract
The open-vocabulary image segmentation task involves partitioning images into semantically meaningful segments and classifying them with flexible text-defined categories. The recent vision-based foundation models such as the Segment Anything Model (SAM) have shown superior performance in generating class-agnostic image segments. The main challenge in open-vocabulary image segmentation now lies in accurately classifying these segments into text-defined categories. In this paper, we introduce the Universal Segment Embedding (USE) framework to address this challenge. This framework is comprised of two key components: 1) a data pipeline designed to efficiently curate a large amount of segment-text pairs at various granularities, and 2) a universal segment embedding model that enables precise segment classification into a vast range of text-defined categories. The USE model can not only help open-vocabulary image segmentation but also facilitate other downstream tasks (e.g., querying and ranking). Through comprehensive experimental studies on semantic segmentation and part segmentation benchmarks, we demonstrate that the USE framework outperforms state-of-the-art open-vocabulary segmentation methods.
Xiwei Xuan, Clint Sebastian, Jorge Henrique Piazentin Ono, Sima Behpour, Thang Doan, Liang Gou, Han-Wei Shen, Liu Ren 0001
CVPR10
2024 GNNBoundary: Towards Explaining Graph Neural Networks through the Lens of Decision Boundaries
abstract
While Graph Neural Networks (GNNs) have achieved remarkable performance on various machine learning tasks on graph data, they also raised questions regarding their transparency and interpretability. Recently, there have been extensive research efforts to explain the decision-making process of GNNs. These efforts often focus on explaining why a certain prediction is made for a particular instance, or what discriminative features the GNNs try to detect for each class. However, to the best of our knowledge, there is no existing study on understanding the decision boundaries of GNNs, even though the decision-making process of GNNs is directly determined by the decision boundaries. To bridge this research gap, we propose a model-level explainability method called GNNBoundary, which attempts to gain deeper insights into the decision boundaries of graph classifiers. Specifically, we first develop an algorithm to identify the pairs of classes whose decision regions are adjacent. For an adjacent class pair, the near-boundary graphs between them are effectively generated by optimizing a novel objective function specifically designed for boundary graph generation. Thus, by analyzing the nearboundary graphs, the important characteristics of decision boundaries can be uncovered. To evaluate the efficacy of GNNBoundary, we conduct experiments on both synthetic and public real-world datasets. The results demonstrate that, via the analysis of faithful near-boundary graphs generated by GNNBoundary, we can thoroughly assess the robustness and generalizability of the explained GNNs. The official implementation can be found at https://github.com/yolandalalala/GNNBoundary.
Han-Wei Shen
ICLR2
2024 FedNE: Surrogate-Assisted Federated Neighbor Embedding for Dimensionality Reduction
abstract
Federated learning (FL) has rapidly evolved as a promising paradigm that enables collaborative model training across distributed participants without exchanging their local data. Despite its broad applications in fields such as computer vision, graph learning, and natural language processing, the development of a data projection model that can be effectively used to visualize data in the context of FL is crucial yet remains heavily under-explored. Neighbor embedding (NE) is an essential technique for visualizing complex high-dimensional data, but collaboratively learning a joint NE model is difficult. The key challenge lies in the objective function, as effective visualization algorithms like NE require computing loss functions among pairs of data. In this paper, we introduce \textsc{FedNE}, a novel approach that integrates the \textsc{FedAvg} framework with the contrastive NE technique, without any requirements of shareable data. To address the lack of inter-client repulsion which is crucial for the alignment in the global embedding space, we develop a surrogate loss function that each client learns and shares with each other. Additionally, we propose a data-mixing strategy to augment the local data, aiming to relax the problems of invisible neighbors and false neighbors constructed by the local $k$NN graphs. We conduct comprehensive experiments on both synthetic and real-world datasets. The results demonstrate that our \textsc{FedNE} can effectively preserve the neighborhood data structures and enhance the alignment in the global embedding space compared to several baseline methods.
Hong-You Chen, Han-Wei Shen, Wei-Lun Chao
NeurIPS4
2024 Efficient Level-Crossing Probability Calculation for Gaussian Process Modeled Data
abstract
Almost all scientific data have uncertainties originating from different sources. Gaussian process regression (GPR) models are a natural way to model data with Gaussian-distributed uncertainties. GPR also has the benefit of reducing I/O bandwidth and storage requirements for large scientific simulations. However, the reconstruction from the GPR models suffers from high computation complexity. To make the situation worse, classic approaches for visualizing the data uncertainties, like probabilistic marching cubes, are also computationally very expensive, especially for data of high resolutions. In this paper, we accelerate the level-crossing probability calculation efficiency on GPR models by subdividing the data spatially into a hierarchical data structure and only reconstructing values adaptively in the regions that have a non-zero probability. For each region, leveraging the known GPR kernel and the saved data observations, we propose a novel approach to efficiently calculate an upper bound for the level-crossing probability inside the region and use this upper bound to make the subdivision and reconstruction decisions. We demonstrate that our value occurrence probability estimation is accurate with a low computation cost by experiments that calculate the level-crossing probability fields on different datasets.
Isaac J. Michaud, Ayan Biswas 0001, Han-Wei Shen
PacificVis4
2024 KG-PRE-view: Democratizing a TVCG Knowledge Graph through Visual Explorations
abstract
IEEE Transactions on Visualization and Computer Graphics (TVCG) publishes cutting-edge research in the fields of visualization, computer graphics, and virtual and augmented realities. Within the TVCG ecosystem, different stakeholders make decisions based on available information related to TVCG almost on a daily basis. The decisions involve various tasks such as the retrieval of research ideas and trends, the invitation of peer reviewers, and the selection of editorial board members, just to name a few. To make well-informed decisions in these contexts, a data-driven approach is necessary. However, the current IEEE digital library only provides access to individual papers. Transforming this wealth of data into valuable insights is a daunting task, requiring specialized expertise and effort in tasks such as data crawling, cleaning, analysis, and visualizations. To address the needs of the community in facilitating more efficient and transparent decision-making, we construct and publicly release a TVCG knowledge graph (TVCG-KG). TVCG-KG is a structured representation of heterogeneous information, including the metadata of each publication such as author, affiliation, title, and semantic information such as method, task, data. Despite the widespread use of KGs in various downstream applications, a noticeable gap exists in the visualization literature regarding the full exploitation of the rich semantics embedded within KGs. While it might seem intuitive to just employ interactive graph-based visualization for KGs, we propose that knowledge discovery over KG is a series of visual exploratory tasks that can benefit from using multiple visualization techniques and designs. We conducted an evaluation of TVCG-KG quality and demonstrated its practical utility through several real-world cases. Our data and code are accessible via the following URL: https://github.com/yasmineTYM/TVCG-KG.git.
Yamei Tu, Han-Wei Shen
PacificVis3
2024 DocFlow: A Visual Analytics System for Question-Based Document Retrieval and Categorization
abstract
A systematic review (SR) is essential with up-to-date research evidence to support clinical decisions and practices. However, the growing literature volume makes it challenging for SR reviewers and clinicians to discover useful information efficiently. Many human-in-the-loop information retrieval approaches (HIR) have been proposed to rank documents semantically similar to users' queries and provide interactive visualizations to facilitate document retrieval. Given that the queries are mainly composed of keywords and keyphrases retrieving documents that are semantically similar to a query does not necessarily respond to the clinician's need. Clinicians still have to review many documents to find the solution. The problem motivates us to develop a visual analytics system, DocFlow, to facilitate information-seeking. One of the features of our DocFlow is accepting natural language questions. The detailed description enables retrieving documents that can answer users' questions. Additionally, clinicians often categorize documents based on their backgrounds and with different purposes (e.g., populations, treatments). Since the criteria are unknown and cannot be pre-defined in advance, existing methods can only achieve categorization by considering the entire information in documents. In contrast, by locating answers in each document, our DocFlow can intelligently categorize documents based on users' questions. The second feature of our DocFlow is a flexible interface where users can arrange a sequence of questions to customize their rules for document retrieval and categorization. The two features of this visual analytics system support a flexible information-seeking process. The case studies and the feedback from domain experts demonstrate the usefulness and effectiveness of our DocFlow.
Yamei Tu, Yu-Shuen Wang, Po-Yin Yen, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.5
2024 Message from the Editor-in-Chief
Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.1
2024 Message from the Editor-in-Chief and from the Associate Editor-in-Chief
abstract
Welcome to the13th IEEE Transactions on Visualization and Computer Graphics (TVCG)special issue on IEEE Virtual Reality and 3D User Interfaces. This volume contains a total of 80 full papers selected for and presented at the IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR 2024), held in Orlando, Florida, USA, from March 16 to 21, 2024.
Han-Wei Shen, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.1
2024 Message from the Editor-in-Chief and from the Associate Editor-in-Chief
abstract
Welcome to the 10th IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue on IEEE International Symposium on Mixed and Augmented Reality (ISMAR). This volume contains a total of 44 full papers selected for and presented at ISMAR 2024, held from October 21 to 25, 2024 in the Greater Seattle Area, USA, in a hybrid mode.
Han-Wei Shen, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.1
2024 PSRFlow: Probabilistic Super Resolution with Flow-Based Models for Scientific Data
abstract
Although many deep-learning-based super-resolution approaches have been proposed in recent years, because no ground truth is available in the inference stage, few can quantify the errors and uncertainties of the super-resolved results. For scientific visualization applications, however, conveying uncertainties of the results to scientists is crucial to avoid generating misleading or incorrect information. In this paper, we propose PSRFlow, a novel normalizing flow-based generative model for scientific data super-resolution that incorporates uncertainty quantification into the super-resolution process. PSRFlow learns the conditional distribution of the high-resolution data based on the low-resolution counterpart. By sampling from a Gaussian latent space that captures the missing information in the high-resolution data, one can generate different plausible super-resolution outputs. The efficient sampling in the Gaussian latent space allows our model to perform uncertainty quantification for the super-resolved results. During model training, we augment the training data with samples across various scales to make the model adaptable to data of different scales, achieving flexible super-resolution for a given input. Our results demonstrate superior performance and robust uncertainty quantification compared with existing methods such as interpolation and GAN-based super-resolution networks.
Jingyi Shen, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2024 PhraseMap: Attention-Based Keyphrases Recommendation for Information Seeking
abstract
Many Information Retrieval (IR) approaches have been proposed to extract relevant information from a large corpus. Among these methods, phrase-based retrieval methods have been proven to capture more concrete and concise information than word-based and paragraph-based methods. However, due to the complex relationship among phrases and a lack of proper visual guidance, achieving user-driven interactive information-seeking and retrieval remains challenging. In this study, we present a visual analytic approach for users to seek information from an extensive collection of documents efficiently. The main component of our approach is a PhraseMap, where nodes and edges represent the extracted keyphrases and their relationships, respectively, from a large corpus. To build the PhraseMap, we extract keyphrases from each document and link the phrases according to word attention determined using modern language models, i.e., BERT. As can be imagined, the graph is complex due to the extensive volume of information and the massive amount of relationships. Therefore, we develop a navigation algorithm to facilitate information seeking. It includes (1) a question-answering (QA) model to identify phrases related to users' queries and (2) updating relevant phrases based on users' feedback. To better present the PhraseMap, we introduce a resource-controlled self-organizing map (RC-SOM) to evenly and regularly display phrases on grid cells while expecting phrases with similar semantics to stay close in the visualization. To evaluate our approach, we conducted case studies with three domain experts in diverse literature. The results and feedback demonstrate its effectiveness, usability, and intelligence.
Yamei Tu, Yu-Shuen Wang, Po-Yin Yen, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.5
2024 SmartGD: A GAN-Based Graph Drawing Framework for Diverse Aesthetic Goals
abstract
While a multitude of studies have been conducted on graph drawing, many existing methods only focus on optimizing a single aesthetic aspect of graph layouts, which can lead to sub-optimal results. There are a few existing methods that have attempted to develop a flexible solution for optimizing different aesthetic aspects measured by different aesthetic criteria. Furthermore, thanks to the significant advance in deep learning techniques, several deep learning-based layout methods were proposed recently. These methods have demonstrated the advantages of deep learning approaches for graph drawing. However, none of these existing methods can be directly applied to optimizing non-differentiable criteria without special accommodation. In this work, we propose a novel Generative Adversarial Network (GAN) based deep learning framework for graph drawing, called SmartGD, which can optimize different quantitative aesthetic goals, regardless of their differentiability. To demonstrate the effectiveness and efficiency of SmartGD, we conducted experiments on minimizing stress, minimizing edge crossing, maximizing crossing angle, maximizing shape-based metrics, and a combination of multiple aesthetics. Compared with several popular graph drawing algorithms, the experimental results show that SmartGD achieves good performance both quantitatively and qualitatively.
Kevin Yen, Yifan Hu 0001, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.4
2024 Adaptively Placed Multi-Grid Scene Representation Networks for Large-Scale Data Visualization
abstract
Scene representation networks (SRNs) have been recently proposed for compression and visualization of scientific data. However, state-of-the-art SRNs do not adapt the allocation of available network parameters to the complex features found in scientific data, leading to a loss in reconstruction quality. We address this shortcoming with an adaptively placed multi-grid SRN (APMGSRN) and propose a domain decomposition training and inference technique for accelerated parallel training on multi-GPU systems. We also release an open-source neural volume rendering application that allows plug-and-play rendering with any PyTorch-based SRN. Our proposed APMGSRN architecture uses multiple spatially adaptive feature grids that learn where to be placed within the domain to dynamically allocate more neural network resources where error is high in the volume, improving state-of-the-art reconstruction accuracy of SRNs for scientific data without requiring expensive octree refining, pruning, and traversal like previous adaptive models. In our domain decomposition approach for representing large-scale data, we train an set of APMGSRNs in parallel on separate bricks of the volume to reduce training time while avoiding overhead necessary for an out-of-core solution for volumes too large to fit in GPU memory. After training, the lightweight SRNs are used for realtime neural volume rendering in our open-source renderer, where arbitrary view angles and transfer functions can be explored. A copy of this paper, all code, all models used in our experiments, and all supplemental materials and videos are available at https://github.com/skywolf829/APMGSRN.
Skylar W. Wurster, Tianyu Xiong, Han-Wei Shen, Hanqi Guo 0001, Tom Peterka
IEEE Trans. Vis. Comput. Graph.3
2023 On the Importance and Applicability of Pre-Training for Federated Learning
Hong-You Chen, Cheng-Hao Tu 0001, Han-Wei Shen, Wei-Lun Chao
ICLR4
2023 GNNInterpreter: A Probabilistic Generative Model-Level Explanation for Graph Neural Networks
Han-Wei Shen
ICLR2
2023 Neural Stream Functions
abstract
We present a neural network approach to compute stream functions, which are scalar functions with gradients orthogonal to a given vector field. As a result, isosurfaces of the stream function extract stream surfaces, which can be visualized to analyze flow features. Our approach takes a vector field as input and trains an implicit neural representation to learn a stream function for that vector field. The network learns to map input coordinates to a stream function value by minimizing the inner product of the gradient of the neural network’s output and the vector field. Since stream function solutions may not be unique, we give optional constraints for the network to learn particular stream functions of interest. Specifically, we introduce regularizing loss functions that can optionally be used to generate stream function solutions whose stream surfaces follow the flow field’s curvature, or that can learn a stream function that includes a stream surface passing through a seeding rake. We also discuss considerations for properly visualizing the trained implicit network and extracting artifact-free surfaces. We compare our results with other implicit solutions and present qualitative and quantitative results for several synthetic and simulated vector fields.
Skylar W. Wurster, Hanqi Guo 0001, Tom Peterka, Han-Wei Shen
PacificVis4
2023 Visual Analytics on Network Forgetting for Task-Incremental Learning
abstract
Abstract Task‐incremental learning (Task‐IL) aims to enable an intelligent agent to continuously accumulate knowledge from new learning tasks without catastrophically forgetting what it has learned in the past. It has drawn increasing attention in recent years, with many algorithms being proposed to mitigate neural network forgetting. However, none of the existing strategies is able to completely eliminate the issues. Moreover, explaining and fully understanding what knowledge and how it is being forgotten during the incremental learning process still remains under‐explored. In this paper, we propose KnowledgeDrift, a visual analytics framework, to interpret the network forgetting with three objectives: (1) to identify when the network fails to memorize the past knowledge, (2) to visualize what information has been forgotten, and (3) to diagnose how knowledge attained in the new model interferes with the one learned in the past. Our analytical framework first identifies the occurrence of forgetting by tracking the task performance under the incremental learning process and then provides in‐depth inspections of drifted information via various levels of data granularity. KnowledgeDrift allows analysts and model developers to enhance their understanding of network forgetting and compare the performance of different incremental learning algorithms. Three case studies are conducted in the paper to further provide insights and guidance for users to effectively diagnose catastrophic forgetting over time.
Jiayi Xu 0001, Wei-Lun Chao, Han-Wei Shen
Comput. Graph. Forum4
2023 Local Latent Representation Based on Geometric Convolution for Particle Data Feature Exploration
abstract
Feature related particle data analysis plays an important role in many scientific applications such as fluid simulations, cosmology simulations and molecular dynamics. Compared to conventional methods that use hand-crafted feature descriptors, some recent studies focus on transforming the data into a new latent space, where features are easier to be identified, compared and extracted. However, it is challenging to transform particle data into latent representations, since the convolution neural networks used in prior studies require the data presented in regular grids. In this article, we adopt Geometric Convolution, a neural network building block designed for 3D point clouds, to create latent representations for scientific particle data. These latent representations capture both the particle positions and their physical attributes in the local neighborhood so that features can be extracted by clustering in the latent space, and tracked by applying tracking algorithms such as mean-shift. We validate the extracted features and tracking results from our approach using datasets from three applications and show that they are comparable to the methods that define hand-crafted features for each specific dataset.
Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2023 Editor's Note
abstract
T HE IEEE Computer Society's policy mandates term
Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.1
2023 Editorial A Message from the New Editor-in-Chief
Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.1
2023 IEEE VR 2023 Introducing the Special Issue
abstract
Welcome to the 12th IEEE Transactions on Visualization and Computer Graphics (TVCG) special issue on IEEE Virtual Reality and 3D User Interfaces. This volume contains a total of 61 full papers selected for and presented at the IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR 2023), held in a hybrid style physically in Shanghai, China and virtually online, from March 25 to 29, 2023.
Han-Wei Shen, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.1
2023 Message from the Editor-in-Chief and from the Associate Editor-in-Chief
abstract
Welcome to the November 2023 issue of theIEEE Transactions on Visualization and Computer Graphics (TVCG). This issue contains selected papers accepted at the IEEE International Symposium on Mixed and Augmented Reality (ISMAR). The conference took place from October 16 to 20, 2023 in Sydney, Australia, in a hybrid mode.
Han-Wei Shen, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.1
2023 IDLat: An Importance-Driven Latent Generation Method for Scientific Data
abstract
Deep learning based latent representations have been widely used for numerous scientific visualization applications such as isosurface similarity analysis, volume rendering, flow field synthesis, and data reduction, just to name a few. However, existing latent representations are mostly generated from raw data in an unsupervised manner, which makes it difficult to incorporate domain interest to control the size of the latent representations and the quality of the reconstructed data. In this paper, we present a novel importance-driven latent representation to facilitate domain-interest-guided scientific data visualization and analysis. We utilize spatial importance maps to represent various scientific interests and take them as the input to a feature transformation network to guide latent generation. We further reduced the latent size by a lossless entropy encoding algorithm trained together with the autoencoder, improving the storage and memory efficiency. We qualitatively and quantitatively evaluate the effectiveness and efficiency of latent representations generated by our method with data from multiple scientific visualization applications.
Jingyi Shen, Jiayi Xu 0001, Ayan Biswas 0001, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.5
2023 VDL-Surrogate: A View-Dependent Latent-based Model for Parameter Space Exploration of Ensemble Simulations
abstract
We propose VDL-Surrogate, a view-dependent neural-network-latent-based surrogate model for parameter space exploration of ensemble simulations that allows high-resolution visualizations and user-specified visual mappings. Surrogate-enabled parameter space exploration allows domain scientists to preview simulation results without having to run a large number of computationally costly simulations. Limited by computational resources, however, existing surrogate models may not produce previews with sufficient resolution for visualization and analysis. To improve the efficient use of computational resources and support high-resolution exploration, we perform ray casting from different viewpoints to collect samples and produce compact latent representations. This latent encoding process reduces the cost of surrogate model training while maintaining the output quality. In the model training stage, we select viewpoints to cover the whole viewing sphere and train corresponding VDL-Surrogate models for the selected viewpoints. In the model inference stage, we predict the latent representations at previously selected viewpoints and decode the latent representations to data space. For any given viewpoint, we make interpolations over decoded data at selected viewpoints and generate visualizations with user-specified visual mappings. We show the effectiveness and efficiency of VDL-Surrogate in cosmological and ocean simulations with quantitative and qualitative evaluations. Source code is publicly available at https://github.com/trainsn/VDL-Surrogate.
Neng Shi, Jiayi Xu 0001, Hanqi Guo 0001, Jonathan Woodring, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.6
2023 SDRQuerier: A Visual Querying Framework for Cross-National Survey Data Recycling
abstract
Public opinion surveys constitute a widespread, powerful tool to study peoples' attitudes and behaviors from comparative perspectives. However, even global surveys can have limited geographic and temporal coverage, which can hinder the production of comprehensive knowledge. To expand the scope of comparison, social scientists turn to ex-post harmonization of variables from datasets that cover similar topics but in different populations and/or at different times. These harmonized datasets can be analyzed as a single source and accessed through various data portals. However, the Survey Data Recycling (SDR) research project has identified three challenges faced by social scientists when using data portals: the lack of capability to explore data in-depth or query data based on customized needs, the difficulty in efficiently identifying related data for studies, and the incapability to evaluate theoretical models using sliced data. To address these issues, the SDR research project has developed the SDRQuerier, which is applied to the harmonized SDR database. The SDRQuerier includes a BERT-based model that allows for customized data queries through research questions or keywords (Query-by-Question), a visual design that helps users determine the availability of harmonized data for a given research question (Query-by-Condition), and the ability to reveal the underlying relational patterns among substantive and methodological variables in the database (Query-by-Relation), aiding in the rigorous evaluation or improvement of regression models. Case studies with multiple social scientists have demonstrated the usefulness and effectiveness of the SDRQuerier in addressing daily challenges.
Yamei Tu, Olga Li, Junpeng Wang 0001, Han-Wei Shen, Przemek Powalko, Irina Tomescu-Dubrow, Kazimierz M. Slomczynski, Spyros Blanas, J. Craig Jenkins
IEEE Trans. Vis. Comput. Graph.4
2023 Deep Hierarchical Super Resolution for Scientific Data
abstract
We present a novel technique for hierarchical super resolution (SR) with neural networks (NNs), which upscales volumetric data represented with an octree data structure to a high-resolution uniform gridwith minimal seam artifacts on octree node boundaries. Our method uses existing state-of-the-art SR models and adds flexibility to upscale input data with varying levels of detail across the domain, instead of only uniform grid data that are supported in previous approaches.The key is to use a hierarchy of SR NNs, each trained to perform 2× SR between two levels of detail, with a hierarchical SR algorithm that minimizes seam artifacts by starting from the coarsest level of detail and working up.We show that our hierarchical approach outperforms baseline interpolation and hierarchical upscaling methods, and demonstrate the usefulness of our proposed approach across three use cases including data reduction using hierarchical downsampling+SR instead of uniform downsampling+SR, computation savings for hierarchical finite-time Lyapunov exponent field calculation, and super-resolving low-resolution simulation results for a high-resolution approximation visualization.
Skylar W. Wurster, Hanqi Guo 0001, Han-Wei Shen, Tom Peterka, Jiayi Xu 0001
IEEE Trans. Vis. Comput. Graph.3
2023 Reinforcement Learning for Load-Balanced Parallel Particle Tracing
abstract
We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.
Jiayi Xu 0001, Hanqi Guo 0001, Han-Wei Shen, Mukund Raj, Skylar W. Wurster, Tom Peterka
IEEE Trans. Vis. Comput. Graph.3
2022 Preface
abstract
This February 2022 issue of theIEEE Transactions on Visualization and Computer Graphics (TVCG)contains the proceedings of IEEE VIS 2021, held online on October 24-29, 2021, with General Chairs from Tulane University and Universidade de Sao Paulo. With IEEE VIS 2021, the conference series is in its 32nd year.
Bongshin Lee, Silvia Miksch, Anders Ynnerman, Anastasia Bezerianos, Jian Chen 0006, Wei Chen 0001, Christopher Collins 0001, Michael Gleicher, M. Eduard Gröller, Alexander Lex, Bernhard Preim, Jinwook Seo, Rüdiger Westermann, Jing Yang 0001, Xiaoru Yuan, Han-Wei Shen, Jean-Daniel Fekete, Shixia Liu
IEEE Trans. Vis. Comput. Graph.16
2022 GNN-Surrogate: A Hierarchical and Adaptive Graph Neural Network for Parameter Space Exploration of Unstructured-Mesh Ocean Simulations
abstract
We propose GNN-Surrogate, a graph neural network-based surrogate model to explore the parameter space of ocean climate simulations. Parameter space exploration is important for domain scientists to understand the influence of input parameters (e.g., wind stress) on the simulation output (e.g., temperature). The exploration requires scientists to exhaust the complicated parameter space by running a batch of computationally expensive simulations. Our approach improves the efficiency of parameter space exploration with a surrogate model that predicts the simulation outputs accurately and efficiently. Specifically, GNN-Surrogate predicts the output field with given simulation parameters so scientists can explore the simulation parameter space with visualizations from user-specified visual mappings. Moreover, our graph-based techniques are designed for unstructured meshes, making the exploration of simulation outputs on irregular grids efficient. For efficient training, we generate hierarchical graphs and use adaptive resolutions. We give quantitative and qualitative evaluations on the MPAS-Ocean simulation to demonstrate the effectiveness and efficiency of GNN-Surrogate. Source code is publicly available at https://github.com/trainsn/GNN-Surrogate.
Neng Shi, Jiayi Xu 0001, Skylar W. Wurster, Hanqi Guo 0001, Jonathan Woodring, Luke P. Van Roekel, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.7
2022 Geometry-Driven Detection, Tracking and Visual Analysis of Viscous and Gravitational Fingers
abstract
Viscous and gravitational flow instabilities cause a displacement front to break up into finger-like fluids. The detection and evolutionary analysis of these fingering instabilities are critical in multiple scientific disciplines such as fluid mechanics and hydrogeology. However, previous detection methods of the viscous and gravitational fingers are based on density thresholding, which provides limited geometric information of the fingers. The geometric structures of fingers and their evolution are important yet little studied in the literature. In this article, we explore the geometric detection and evolution of the fingers in detail to elucidate the dynamics of the instability. We propose a ridge voxel detection method to guide the extraction of finger cores from three-dimensional (3D) scalar fields. After skeletonizing finger cores into skeletons, we design a spanning tree based approach to capture how fingers branch spatially from the finger skeletons. Finally, we devise a novel geometric-glyph augmented tracking graph to study how the fingers and their branches grow, merge, and split over time. Feedback from earth scientists demonstrates the usefulness of our approach to performing spatio-temporal geometric analyses of fingers.
Jiayi Xu 0001, Soumya Dutta, Joachim Moortgat, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.5
2021 KeywordMap: Attention-based Visual Exploration for Keyword Analysis
abstract
With the high growth rate of text data, extracting meaningful information from a large corpus becomes increasingly difficult. Keyword extraction and analysis is a common approach to tackle the problem, but it is non-trivial to identify important words in the text and represent the multifaceted properties of those words effectively. Traditional topic modeling based keyword analysis algorithms require hyper-parameters which are often difficult to tune without enough prior knowledge. In addition, the relationships among the keywords are often difficult to obtain. In this paper, we utilize the attention scores extracted from Transformer-based language models to capture word relationships. We propose a domain-driven attention tuning method, guiding the attention to learn domain-specific word relationships. From the attention, we build a keyword network and propose a novel algorithm, Attention-based Word Influence (AWI), to compute how influential each word is in the network. An interactive visual analytics system, KeywordMap, is developed to support multi-level analysis of keywords and keyword relationships through coordinated views. We measure the quality of keywords captured by our AWI algorithm quantitatively. We also evaluate the usefulness and effectiveness of KeywordMap through case studies.
Yamei Tu, Jiayi Xu 0001, Han-Wei Shen
PacificVis3
2021 Document Domain Randomization for Deep Learning Document Layout Extraction
abstract
We present document domain randomization (DDR), the first successful transfer of convolutional neural networks (CNNs) trained only on graphically rendered pseudo-paper pages to real-world document segmentation. DDR renders pseudo-document pages by modeling randomized textual and non-textual contents of interest, with user-defined layout and font styles to support joint learning of fine-grained classes. We demonstrate competitive results using our DDR approach to extract nine document classes from the benchmark CS-150 and papers published in two domains, namely annual meetings of Association for Computational Linguistics (ACL) and IEEE Visualization (VIS). We compare DDR to conditions of style mismatch, fewer or more noisy samples that are more easily obtained in the real world. We show that high-fidelity semantic information is not necessary to label semantic classes but style mismatch between train and test can lower model accuracy. Using smaller training samples had a slightly detrimental effect. Finally, network models still achieved high test accuracy when correct labels are diluted towards confusing labels; this behavior hold across several classes.
Meng Ling, Jian Chen 0006, Torsten Möller, Petra Isenberg, Tobias Isenberg 0001, Michael Sedlmair, Robert S. Laramee, Han-Wei Shen, Jian Wu 0006, C. Lee Giles
ICDAR (1)8
2021 VIS30K: A Collection of Figures and Tables From IEEE Visualization Conference Publications
abstract
We present the VIS30K dataset, a collection of 29,689 images that represents 30 years of figures and tables from each track of the IEEE Visualization conference series (Vis, SciVis, InfoVis, VAST). VIS30K's comprehensive coverage of the scientific literature in visualization not only reflects the progress of the field but also enables researchers to study the evolution of the state-of-the-art and to find relevant work based on graphical content. We describe the dataset and our semi-automatic collection process, which couples convolutional neural networks (CNN) with curation. Extracting figures and tables semi-automatically allows us to verify that no images are overlooked or extracted erroneously. To improve quality further, we engaged in a peer-search process for high-quality figures from early IEEE Visualization papers. With the resulting data, we also contribute VISImageNavigator (VIN, visimagenavigator.github.io), a web-based tool that facilitates searching and exploring VIS30K by author names, paper keywords, title and abstract, and years.
Jian Chen 0006, Meng Ling, Rui Li 0067, Petra Isenberg, Tobias Isenberg 0001, Michael Sedlmair, Torsten Möller, Robert S. Laramee, Han-Wei Shen, Katharina Wünsche
IEEE Trans. Vis. Comput. Graph.9
2021 Preface
abstract
This February 2021 issue of the IEEE Transactions on Visualization and Computer Graphics (TVCG) contains the proceedings of IEEE VIS 2020, held online between 25-30 October 2020, hosted by General Chairs from the University of Utah. With IEEE VIS 2020, the conference series is in its 31st year. IEEE VIS consists of three conferences, held concurrently: the IEEE Visual Analytics Science and Technology Conference (VAST), the IEEE Information Visualization Conference (InfoVis), and the IEEE Scientific Visualization Conference (SciVis). These three conferences are the premier venues for the visualization community to exchange the latest ideas and developments, attracting researchers and practitioners alike.
Niklas Elmqvist, Brian D. Fisher, Peter Lindstrom 0001, Ross Maciejewski, Miriah D. Meyer, Silvia Miksch, Luis Gustavo Nonato, Nathalie Henry Riche, Han-Wei Shen, Rüdiger Westermann, Jo Wood, Jing Yang 0001
IEEE Trans. Vis. Comput. Graph.9
2021 FTK: A Simplicial Spacetime Meshing Framework for Robust and Scalable Feature Tracking
abstract
We present the Feature Tracking Kit (FTK), a framework that simplifies, scales, and delivers various feature-tracking algorithms for scientific data. The key of FTK is our simplicial spacetime meshing scheme that generalizes both regular and unstructured spatial meshes to spacetime while tessellating spacetime mesh elements into simplices. The benefits of using simplicial spacetime meshes include (1) reducing ambiguity cases for feature extraction and tracking, (2) simplifying the handling of degeneracies using symbolic perturbations, and (3) enabling scalable and parallel processing. The use of simplicial spacetime meshing simplifies and improves the implementation of several feature-tracking algorithms for critical points, quantum vortices, and isosurfaces. As a software framework, FTK provides end users with VTK/ParaView filters, Python bindings, a command line interface, and programming interfaces for feature-tracking applications. We demonstrate use cases as well as scalability studies through both synthetic data and scientific applications including tokamak, fluid dynamics, and superconductivity simulations. We also conduct end-to-end performance studies on the Summit supercomputer. FTK is open sourced under the MIT license: https://github.com/hguo/ftk.
Hanqi Guo 0001, David Lenz 0002, Jiayi Xu 0001, Xin Liang 0001, Iulian R. Grindeanu, Han-Wei Shen, Tom Peterka, Todd S. Munson, Ian T. Foster
IEEE Trans. Vis. Comput. Graph.7
2021 CNNPruner: Pruning Convolutional Neural Networks with Visual Analytics
abstract
Convolutional neural networks (CNNs) have demonstrated extraordinarily good performance in many computer vision tasks. The increasing size of CNN models, however, prevents them from being widely deployed to devices with limited computational resources, e.g., mobile/embedded devices. The emerging topic of model pruning strives to address this problem by removing less important neurons and fine-tuning the pruned networks to minimize the accuracy loss. Nevertheless, existing automated pruning solutions often rely on a numerical threshold of the pruning criteria, lacking the flexibility to optimally balance the trade-off between efficiency and accuracy. Moreover, the complicated interplay between the stages of neuron pruning and model fine-tuning makes this process opaque, and therefore becomes difficult to optimize. In this paper, we address these challenges through a visual analytics approach, named CNNPruner. It considers the importance of convolutional filters through both instability and sensitivity, and allows users to interactively create pruning plans according to a desired goal on model size or accuracy. Also, CNNPruner integrates state-of-the-art filter visualization techniques to help users understand the roles that different filters played and refine their pruning plans. Through comprehensive case studies on CNNs with real-world sizes, we validate the effectiveness of CNNPruner.
Guan Li 0002, Junpeng Wang 0001, Han-Wei Shen, Kaixin Chen 0004, Guihua Shan, Zhonghua Lu
IEEE Trans. Vis. Comput. Graph.3
2021 Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis
abstract
We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. We demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.
Jiayi Xu 0001, Hanqi Guo 0001, Han-Wei Shen, Mukund Raj, Xueyun Wang, Xueqiao Xu, Zhehui Wang, Tom Peterka
IEEE Trans. Vis. Comput. Graph.3
2021 USEVis: Visual analytics of attention-based neural embedding in information retrieval
abstract
Neural attention-based encoders, which effectively attend sentence tokens to their associated context without being restricted by long-term distance or dependency, have demonstrated outstanding performance in embedding sentences into meaningful representations (embeddings). The Universal Sentence Encoder (USE) is one of the most well-recognized deep neural network (DNN) based solutions, which is facilitated with an attention-driven transformer architecture and has been pre-trained on a large number of sentences from the Internet. Besides the fact that USE has been widely used in many downstream applications, including information retrieval (IR), interpreting its complicated internal working mechanism remains challenging. In this work, we present a visual analytics solution towards addressing this challenge. Specifically, focused on semantics and syntactics (concepts and relations) that are critical to domain clinical IR, we designed and developed a visual analytics system, i.e., USEVis. The system investigates the power of USE in effectively extracting sentences’ semantics and syntactics through exploring and interpreting how linguistic properties are captured by attentions. Furthermore, by thoroughly examining and comparing the inherent patterns of these attentions, we are able to exploit attentions to retrieve sentences/documents that have similar semantics or are closely related to a given clinical problem in IR. By collaborating with domain experts, we demonstrate use cases with inspiring findings to validate the contribution of our work and the effectiveness of our system.
Xiaonan Ji, Yamei Tu, Junpeng Wang 0001, Han-Wei Shen, Po-Yin Yen
Vis. Informatics5
2020 DynamicsExplorer: Visual Analytics for Robot Control Tasks involving Dynamics and LSTM-based Control Policies
abstract
Deep reinforcement learning (RL), where a policy represented by a deep neural network is trained, has shown some success in playing video games and chess. However, applying RL to real-world tasks like robot control is still challenging. Because generating a massive number of samples to train control policies using RL on real robots is very expensive, hence impractical, it is common to train in simulations, and then transfer to real environments. The trained policy, however, may fail in the real world because of the difference between the training and the real environments, especially the difference in dynamics. To diagnose the problems, it is crucial for experts to understand (1) how the trained policy behaves under different dynamics settings, (2) which part of the policy affects the behaviors the most when the dynamics setting changes, and (3) how to adjust the training procedure to make the policy robust.This paper presents DynamicsExplorer, a visual analytics tool to diagnose the trained policy on robot control tasks under different dynamics settings. DynamicsExplorer allows experts to overview the results of multiple tests with different dynamics-related parameter settings so experts can visually detect failures and analyze the sensitivity of different parameters. Experts can further examine the internal activations of the policy for selected tests and compare the activations between success and failure tests. Such comparisons help experts form hypotheses about the policy and allows them to verify the hypotheses via DynamicsExplorer. Multiple use cases are presented to demonstrate the utility of DynamicsExplorer.
Teng-Yok Lee, Jeroen van Baar, Kent Wittenburg, Han-Wei Shen
PacificVis5
2020 Distribution-based Particle Data Reduction for In-situ Analysis and Visualization of Large-scale N-body Cosmological Simulations
abstract
Cosmological N-body simulation is an important tool for scientists to study the evolution of the universe. With the increase of computing power, billions of particles of high space-time fidelity can be simulated by supercomputers. However, limited computer storage can only hold a small subset of the simulation output for analysis, which makes the understanding of the underlying cosmological phenomena difficult. To alleviate the problem, we design an in-situ data reduction method for large-scale unstructured particle data. During the data generation phase, we use a combined k-dimensional partitioning and Gaussian mixture model approach to reduce the data by utilizing probability distributions. We offer a model evaluation criterion to examine the quality of the probabilistic distribution models, which allows us to identify and improve low-quality models. After the in-situ processing, the particle data size is greatly reduced, which satisfies the requirements from the domain experts. By comparing the astronomical attributes and visualizations of the reconstructed data with the raw data, we demonstrate the effectiveness of our in-situ particle data reduction technique.
Guan Li 0002, Jiayi Xu 0001, Tianchi Zhang 0003, Guihua Shan, Han-Wei Shen, Ko-Chih Wang, Shihong Liao, Zhonghua Lu
PacificVis5
2020 NNVA: Neural Network Assisted Visual Analysis of Yeast Cell Polarization Simulation
abstract
Complex computational models are often designed to simulate real-world physical phenomena in many scientific disciplines. However, these simulation models tend to be computationally very expensive and involve a large number of simulation input parameters, which need to be analyzed and properly calibrated before the models can be applied for real scientific studies. We propose a visual analysis system to facilitate interactive exploratory analysis of high-dimensional input parameter space for a complex yeast cell polarization simulation. The proposed system can assist the computational biologists, who designed the simulation model, to visually calibrate the input parameters by modifying the parameter values and immediately visualizing the predicted simulation outcome without having the need to run the original expensive simulation for every instance. Our proposed visual analysis system is driven by a trained neural network-based surrogate model as the backend analysis framework. In this work, we demonstrate the advantage of using neural networks as surrogate models for visual analysis by incorporating some of the recent advances in the field of uncertainty quantification, interpretability and explainability of neural network-based models. We utilize the trained network to perform interactive parameter sensitivity analysis of the original simulation as well as recommend optimal parameter configurations using the activation maximization framework of neural networks. We also facilitate detail analysis of the trained network to extract useful insights about the simulation model, learned by the network, during the training process. We performed two case studies, and discovered multiple new parameter configurations, which can trigger high cell polarization results in the original simulation model. We evaluated our results by comparing with the original simulation model outcomes as well as the findings from previous parameter analysis performed by our experts.
Subhashis Hazarika, Ko-Chih Wang, Han-Wei Shen, Ching-Shan Chou
IEEE Trans. Vis. Comput. Graph.4
2020 eFESTA: Ensemble Feature Exploration with Surface Density Estimates
abstract
We propose surface density estimate (SDE) to model the spatial distribution of surface features-isosurfaces, ridge surfaces, and streamsurfaces-in 3D ensemble simulation data. The inputs of SDE computation are surface features represented as polygon meshes, and no field datasets are required (e.g., scalar fields or vector fields). The SDE is defined as the kernel density estimate of the infinite set of points on the input surfaces and is approximated by accumulating the surface densities of triangular patches. We also propose an algorithm to guide the selection of a proper kernel bandwidth for SDE computation. An ensemble Feature Exploration method based on Surface densiTy EstimAtes (eFESTA) is then proposed to extract and visualize the major trends of ensemble surface features. For an ensemble of surface features, each surface is first transformed into a density field based on its contribution to the SDE, and the resulting density fields are organized into a hierarchical representation based on the pairwise distances between them. The hierarchical representation is then used to guide visual exploration of the density fields as well as the underlying surface features. We demonstrate the application of our method using isosurface in ensemble scalar fields, Lagrangian coherent structures in uncertain unsteady flows, and streamsurfaces in ensemble fluid flows.
Hanqi Guo 0001, Han-Wei Shen, Tom Peterka
IEEE Trans. Vis. Comput. Graph.3
2020 InSituNet: Deep Image Synthesis for Parameter Space Exploration of Ensemble Simulations
abstract
We propose InSituNet, a deep learning based surrogate model to support parameter space exploration for ensemble simulations that are visualized in situ. In situ visualization, generating visualizations at simulation time, is becoming prevalent in handling large-scale simulations because of the I/O and storage constraints. However, in situ visualization approaches limit the flexibility of post-hoc exploration because the raw simulation data are no longer available. Although multiple image-based approaches have been proposed to mitigate this limitation, those approaches lack the ability to explore the simulation parameters. Our approach allows flexible exploration of parameter space for large-scale ensemble simulations by taking advantage of the recent advances in deep learning. Specifically, we design InSituNet as a convolutional regression model to learn the mapping from the simulation and visualization parameters to the visualization results. With the trained model, users can generate new images for different simulation parameters under various visualization settings, which enables in-depth analysis of the underlying ensemble simulations. We demonstrate the effectiveness of InSituNet in combustion, cosmology, and ocean simulations through quantitative and qualitative evaluations.
Junpeng Wang 0001, Hanqi Guo 0001, Ko-Chih Wang, Han-Wei Shen, Mukund Raj, Youssef S. G. Nashed, Tom Peterka
IEEE Trans. Vis. Comput. Graph.5
2020 Ray-Based Exploration of Large Time-Varying Volume Data Using Per-Ray Proxy Distributions
abstract
The analysis and visualization of data created from simulations on modern supercomputers is a daunting challenge because the incredible compute power of modern supercomputers allow scientists to generate datasets with very high spatial and temporal resolutions. The limited bandwidth and capacity of networking and storage devices connecting supercomputers to analysis machines become the major bottleneck for data analysis such that simply moving the whole dataset from the supercomputer to a data analysis machine is infeasible. A common approach to visualize high temporal resolution simulation datasets under constrained I/O is to reduce the sampling rate in the temporal domain while preserving the original spatial resolution at the time steps. Data interpolation between the sampled time steps alone may not be a viable option since it may suffer from large errors, especially when using a lower sampling rate. We present a novel ray-based representation storing ray based histograms and depth information that recovers the evolution of volume data between sampled time steps. Our view-dependent proxy allows for a good trade off between compactly representing the time-varying data and leveraging temporal coherence within the data by utilizing interpolation between time steps, ray histograms, depth information, and codebooks. Our approach is able to provide fast rendering in the context of transfer function exploration to support visualization of feature evolution in time-varying data.
Ko-Chih Wang, Tzu-Hsuan Wei, Naeem Shareef, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.4
2020 Token-wise sentiment decomposition for ConvNet: Visualizing a sentiment classifier
abstract
Convolutional neural networks are one of the most important and widely used constructs in natural language processing and AI in general. In many applications, they have achieved state-of-the-art performance, with training time faster than the other alternatives. However, due to their limited interpretability, they are less favored by practitioners over attention-based models, like RNNs and self-attention (Transformers), which can be visualized and interpreted more intuitively by analyzing the attention-weight heat-maps. In this work, we present a visualization technique that can be used to understand the inner workings of text-based CNN models. We also show how this method can be used to generate adversarial examples and learn the shortcomings of the training data.
Piyush Chawla, Subhashis Hazarika, Han-Wei Shen
Vis. Informatics3
2020 CECAV-DNN: Collective Ensemble Comparison and Visualization using Deep Neural Networks
abstract
We propose a deep learning approach to collectively compare two or multiple ensembles, each of which is a collection of simulation outputs. The purpose of collective comparison is to help scientists understand differences between simulation models by comparing their ensemble simulation outputs. However, the collective comparison is non-trivial because the spatiotemporal distributions of ensemble simulation outputs reside in a very high dimensional space. To this end, we choose to train a deep discriminative neural network to measure the dissimilarity between two given ensembles, and to identify when and where the two ensembles are different. We also design and develop a visualization system to help users understand the collective comparison results based on the discriminative network. We demonstrate the effectiveness of our approach with two real-world applications, including the ensemble comparison of the community atmosphere model (CAM) and the rapid radiative transfer model for general circulation models (RRTMG) for climate research, and the comparison of computational fluid dynamics (CFD) ensembles with different spatial resolutions.
Junpeng Wang 0001, Hanqi Guo 0001, Han-Wei Shen, Tom Peterka
Vis. Informatics4
2020 Foreword to the Special Issue on PacificVis 2020 Workshop on Visualization Meets AI
abstract
The task of data visualization generally involves a design step, which requires the knowledge of the data domain and visualization methods to do well.Because of the immense space for design optimization, it can take both novices and experts a tremendous effort to derive desired visualization results from data for exploration or communication.Following the resurgence of artificial intelligence technology in recent years, in the field of visualization, there is the growing interest and opportunity in applying AI to perform data transformation and to assist the generation of visualization, aiming to strike a balance between cost and quality.The use of visualization to enhance AI is the other active line of research.The PacificVis 2020 Workshop on Visualization Meets AI aims at exploring this emerging area of research and practice by fostering communication between visualization researchers and practitioners.This issue of Visual Informatics features the six papers chosen by the Workshop.Yiran Li et al. introduce a visual analytics system for comparing tree-based machine learning methods with respect to the reliability and interpretability of their predictions on patient records. Piyush Chawla et al. develop a technique to analyze a CNN based
Kwan-Liu Ma, Han-Wei Shen
Vis. Informatics2
2020 Visual exploration of latent space for traditional Chinese music
abstract
Generating compact and effective numerical representations of data is a fundamental step for many machine learning tasks. Traditionally, handcrafted features are used but as deep learning starts to show its potential, using deep learning models to extract compact representations becomes a new trend. Among them, adopting vectors from the model’s latent space is the most popular. There are several studies focused on visual analysis of latent space in NLP and computer vision . However, relatively little work has been done for music information retrieval (MIR) especially incorporating visualization. To bridge this gap, we propose a visual analysis system utilizing Autoencoders to facilitate analysis and exploration of traditional Chinese music. Due to the lack of proper traditional Chinese music data, we construct a labeled dataset from a collection of pre-recorded audios and then convert them into spectrograms . Our system takes music features learned from two deep learning models (a fully-connected Autoencoder and a Long Short-Term Memory (LSTM) Autoencoder) as input. Through interactive selection, similarity calculation, clustering and listening, we show that the latent representations of the encoded data allow our system to identify essential music elements, which lay the foundation for further analysis and retrieval of Chinese music in the future.
Jingyi Shen, Runqi Wang, Han-Wei Shen
Vis. Informatics3
2019 Object-in-Hand Feature Displacement with Physically-Based Deformation
abstract
Data deformation has been widely used in visualization to obtain an improved view that better helps the comprehension of the data. It has been a consistent pursuit to conduct interactive deformation by operations that are natural to users. In this paper, we propose a deformation system following the object-in-hand metaphor. We utilize a touchscreen to directly manipulate the shape of the data by using fingers. Users can drag data features and move them along with the fingers. Users can also press their fingers to hold other parts of the data fixed during the deformation, or perform cutting on the data using a finger. The deformation is executed using a physically-based mesh, which is constructed to incorporate data properties to make the deformation authentic as well as informative. By manipulating data features as if handling an object in hand, we can successfully achieve less occluded view of the data, or improved feature layout for better view comparison. We present case studies on various types of scientific datasets, including particle data, volumetric data, and streamlines.
Cheng Li 0062, Han-Wei Shen
PacificVis2
2019 Statistical Super Resolution for Data Analysis and Visualization of Large Scale Cosmological Simulations
abstract
Cosmologists build simulations for the evolution of the universe using different initial parameters. By exploring the datasets from different simulation runs, cosmologists can understand the evolution of our universe and approach its initial conditions. A cosmological simulation nowadays can generate datasets on the order of petabytes. Moving datasets from the supercomputers to post data analysis machines is infeasible. We propose a novel approach called statistical super-resolution to tackle the big data problem for cosmological data analysis and visualization. It uses datasets from a few simulation runs to create a prior knowledge, which captures the relation between low-and high-resolution data. We apply in situ statistical down-sampling to datasets generated from simulation runs to minimize the requirements of I/O bandwidth and storage. High-resolution datasets are reconstructed from the statistical down-sampled data by using the prior knowledge for scientists to perform advanced data analysis and render high-quality visualizations.
Ko-Chih Wang, Jiayi Xu 0001, Jonathan Woodring, Han-Wei Shen
PacificVis4
2019 Extreme-Scale Stochastic Particle Tracing for Uncertain Unsteady Flow Visualization and Analysis
abstract
We present an efficient and scalable solution to estimate uncertain transport behaviors-stochastic flow maps (SFMs)-for visualizing and analyzing uncertain unsteady flows. Computing flow maps from uncertain flow fields is extremely expensive because it requires many Monte Carlo runs to trace densely seeded particles in the flow. We reduce the computational cost by decoupling the time dependencies in SFMs so that we can process shorter sub time intervals independently and then compose them together for longer time periods. Adaptive refinement is also used to reduce the number of runs for each location. We parallelize over tasks-packets of particles in our design-to achieve high efficiency in MPI/thread hybrid programming. Such a task model also enables CPU/GPU coprocessing. We show the scalability on two supercomputers, Mira (up to 256K Blue Gene/Q cores) and Titan (up to 128K Opteron cores and 8K GPUs), that can trace billions of particles in seconds.
Hanqi Guo 0001, Han-Wei Shen, Emil M. Constantinescu, Tom Peterka
IEEE Trans. Vis. Comput. Graph.4
2019 CoDDA: A Flexible Copula-based Distribution Driven Analysis Framework for Large-Scale Multivariate Data
abstract
CoDDA (Copula-based Distribution Driven Analysis) is a flexible framework for large-scale multivariate datasets. A common strategy to deal with large-scale scientific simulation data is to partition the simulation domain and create statistical data summaries. Instead of storing the high-resolution raw data from the simulation, storing the compact statistical data summaries results in reduced storage overhead and alleviated I/O bottleneck. Such summaries, often represented in the form of statistical probability distributions, can serve various post-hoc analysis and visualization tasks. However, for multivariate simulation data using standard multivariate distributions for creating data summaries is not feasible. They are either storage inefficient or are computationally expensive to be estimated in simulation time (in situ) for large number of variables. In this work, using copula functions, we propose a flexible multivariate distribution-based data modeling and analysis framework that offers significant data reduction and can be used in an in situ environment. The framework also facilitates in storing the associated spatial information along with the multivariate distributions in an efficient representation. Using the proposed multivariate data summaries, we perform various multivariate post-hoc analyses like query-driven visualization and sampling-based visualization. We evaluate our proposed method on multiple real-world multivariate scientific datasets. To demonstrate the efficacy of our framework in an in situ environment, we apply it on a large-scale flow simulation.
Subhashis Hazarika, Soumya Dutta, Han-Wei Shen, Jen-Ping Chen
IEEE Trans. Vis. Comput. Graph.3
2019 Visual Exploration of Neural Document Embedding in Information Retrieval: Semantics and Feature Selection
abstract
Neural embeddings are widely used in language modeling and feature generation with superior computational power. Particularly, neural document embedding - converting texts of variable-length to semantic vector representations - has shown to benefit widespread downstream applications, e.g., information retrieval (IR). However, the black-box nature makes it difficult to understand how the semantics are encoded and employed. We propose visual exploration of neural document embedding to gain insights into the underlying embedding space, and promote the utilization in prevalent IR applications. In this study, we take an IR application-driven view, which is further motivated by biomedical IR in healthcare decision-making, and collaborate with domain experts to design and develop a visual analytics system. This system visualizes neural document embeddings as a configurable document map and enables guidance and reasoning; facilitates to explore the neural embedding space and identify salient neural dimensions (semantic features) per task and domain interest; and supports advisable feature selection (semantic analysis) along with instant visual feedback to promote IR performance. We demonstrate the usefulness and effectiveness of this system and present inspiring findings in use cases. This work will help designers/developers of downstream applications gain insights and confidence in neural document embedding, and exploit that to achieve more favorable performance in application domains.
Xiaonan Ji, Han-Wei Shen, Alan Ritter, Raghu Machiraju, Po-Yin Yen
IEEE Trans. Vis. Comput. Graph.2
2019 DQNViz: A Visual Analytics Approach to Understand Deep Q-Networks
abstract
Deep Q-Network (DQN), as one type of deep reinforcement learning model, targets to train an intelligent agent that acquires optimal actions while interacting with an environment. The model is well known for its ability to surpass professional human players across many Atari 2600 games. Despite the superhuman performance, in-depth understanding of the model and interpreting the sophisticated behaviors of the DQN agent remain to be challenging tasks, due to the long-time model training process and the large number of experiences dynamically generated by the agent. In this work, we propose DQNViz, a visual analytics system to expose details of the blind training process in four levels, and enable users to dive into the large experience space of the agent for comprehensive analysis. As an initial attempt in visualizing DQN models, our work focuses more on Atari games with a simple action space, most notably the Breakout game. From our visual analytics of the agent's experiences, we extract useful action/reward patterns that help to interpret the model and control the training. Through multiple case studies conducted together with deep learning experts, we demonstrate that DQNViz can effectively help domain experts to understand, diagnose, and potentially improve DQN models.
Junpeng Wang 0001, Liang Gou, Han-Wei Shen, Hao Yang 0007
IEEE Trans. Vis. Comput. Graph.3
2019 DeepVID: Deep Visual Interpretation and Diagnosis for Image Classifiers via Knowledge Distillation
abstract
Deep Neural Networks (DNNs) have been extensively used in multiple disciplines due to their superior performance. However, in most cases, DNNs are considered as black-boxes and the interpretation of their internal working mechanism is usually challenging. Given that model trust is often built on the understanding of how a model works, the interpretation of DNNs becomes more important, especially in safety-critical applications (e.g., medical diagnosis, autonomous driving). In this paper, we propose DeepVID, a Deep learning approach to Visually Interpret and Diagnose DNN models, especially image classifiers. In detail, we train a small locally-faithful model to mimic the behavior of an original cumbersome DNN around a particular data instance of interest, and the local model is sufficiently simple such that it can be visually interpreted (e.g., a linear model). Knowledge distillation is used to transfer the knowledge from the cumbersome DNN to the small model, and a deep generative model (i.e., variational auto-encoder) is used to generate neighbors around the instance of interest. Those neighbors, which come with small feature variances and semantic meanings, can effectively probe the DNN's behaviors around the interested instance and help the small model to learn those behaviors. Through comprehensive evaluations, as well as case studies conducted together with deep learning experts, we validate the effectiveness of DeepVID.
Junpeng Wang 0001, Liang Gou, Wei Zhang 0189, Hao Yang 0007, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.5
2019 Visualization and Visual Analysis of Ensemble Data: A Survey
abstract
Over the last decade, ensemble visualization has witnessed a significant development due to the wide availability of ensemble data, and the increasing visualization needs from a variety of disciplines. From the data analysis point of view, it can be observed that many ensemble visualization works focus on the same facet of ensemble data, use similar data aggregation or uncertainty modeling methods. However, the lack of reflections on those essential commonalities and a systematic overview of those works prevents visualization researchers from effectively identifying new or unsolved problems and planning for further developments. In this paper, we take a holistic perspective and provide a survey of ensemble visualization. Specifically, we study ensemble visualization works in the recent decade, and categorize them from two perspectives: (1) their proposed visualization techniques; and (2) their involved analytic tasks. For the first perspective, we focus on elaborating how conventional visualization techniques (e.g., surface, volume visualization techniques) have been adapted to ensemble data; for the second perspective, we emphasize how analytic tasks (e.g., comparison, clustering) have been performed differently for ensemble data. From the study of ensemble visualization literature, we have also identified several research trends, as well as some future research opportunities.
Junpeng Wang 0001, Subhashis Hazarika, Cheng Li 0062, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.4
2018 In Situ Prediction Driven Feature Analysis in Jet Engine Simulations
abstract
Efficient feature exploration in large-scale data sets using traditional post-hoc analysis approaches is becoming prohibitive due to the bottleneck stemming from I/O and output data sizes. This problem becomes more challenging when an ensemble of simulations are required to run for studying the influence of input parameters on the model output. As a result, scientists are inclining more towards analyzing the data in situ while it resides in the memory. In situ analysis aims at minimizing expensive data movement while maximizing the resource utilization for extraction of important information from the data. In this work, we study the evolution of rotating stall in jet engines using data generated from a large-scale flow simulation under various input conditions. Since the features of interest lack a precise descriptor, we adopt a fuzzy rule-based machine learning algorithm for efficient and robust extraction of such features. For scalable exploration, we advocate for an off-line learning and in situ prediction driven strategy that facilitates in-depth study of the stall. Task-specific information estimated in situ is visualized interactively during the post-hoc analysis revealing important details about the inception and evolution of stall. We verify and validate our method through comprehensive expert evaluation demonstrating the efficacy of our approach.
Soumya Dutta, Han-Wei Shen, Jen-Ping Chen
PacificVis2
2018 An Automatic Deformation Approach for Occlusion Free Egocentric Data Exploration
abstract
Occlusion management is an important task for three dimension data exploration. For egocentric data exploration, the occlusion problems, caused by the camera being too close to opaque data elements, have not been well addressed by previous studies. In this paper, we propose an automatic approach to resolve these problems and provide an occlusion free egocentric data exploration. Our system utilizes a state transition model to monitor both the camera and the data, and manages the initiation, duration, and termination of deformation with animation. Our method can be applied to multiple types of scientific datasets, including volumetric data, polygon mesh data, and particle data. We demonstrate our method with different exploration tasks, including camera navigation, isovalue adjustment, transfer function adjustment, and time varying exploration. We have collaborated with a domain expert and received positive feedback.
Cheng Li 0062, Joachim Moortgat, Han-Wei Shen
PacificVis3
2018 Image and Distribution Based Volume Rendering for Large Data Sets
abstract
Analyzing scientific datasets created from simulations on modern supercomputers is a daunting challenge due to the fast pace at which these datasets continue to grow. Low cost post analysis machines used by scientists to view and analyze these massive datasets are severely limited by their deficiencies in storage bandwidth, capacity, and computational power. Trying to simply move these datasets to these platforms is infeasible. Any approach to view and analyze these datasets on post analysis machines will have to effectively address the inevitable problem of data loss. Image based approaches are well suited for handling very large datasets on low cost platforms. Three challenges with these approaches are how to effectively represent the original data with minimal data loss, analyze the data in regards to transfer function exploration, which is a key analysis tool, and quantify the error from data loss during analysis. We present a novel image based approach using distributions to preserve data integrity. At each view sample, view dependent data is summarized at each pixel with distributions to define a compact proxy for the original dataset. We present this representation along with how to manipulate and render large scale datasets on post analysis machines. We show that our approach is a good trade off between rendering quality and interactive speed and provides uncertainty quantification for the information that is lost.
Ko-Chih Wang, Naeem Shareef, Han-Wei Shen
PacificVis3
2018 Information Guided Data Sampling and Recovery Using Bitmap Indexing
abstract
Creating a data representation is a common approach for efficient and effective data management and exploration. The compressed bitmap indexing is one of the emerging data representation used for large-scale data exploration. Performing sampling on the bitmapindexing based data representation allows further reduction of storage overhead and be more flexible to meet the requirements of different applications. In this paper, we propose two approaches to solve two potential limitations when exploring and visualizing the data using sampling-based bitmap indexing data representation. First, we propose an adaptive sampling approach called information guided stratified sampling (IGStS) for creating compact sampled datasets that preserves the important characteristics of the raw data. Furthermore, we propose a novel data recovery approach to reconstruct the irregular subsampled dataset into a volume dataset with regular grid structure for qualitative post-hoc data exploration and visualization. The quantitative and visual efficacy of our proposed data sampling and recovery approaches are demonstrated through multiple experiments and applications.
Tzu-Hsuan Wei, Soumya Dutta, Han-Wei Shen
PacificVis3
2018 CorrelatedMultiples: Spatially Coherent Small Multiples With Constrained Multi-Dimensional Scaling
abstract
Abstract Displaying small multiples is a popular method for visually summarizing and comparing multiple facets of a complex data set. If the correlations between the data are not considered when displaying the multiples, searching and comparing specific items become more difficult since a sequential scan of the display is often required. To address this issue, we introduce CorrelatedMultiples, a spatially coherent visualization based on small multiples, where the items are placed so that the distances reflect their dissimilarities. We propose a constrained multi‐dimensional scaling (CMDS) solver that preserves spatial proximity while forcing the items to remain within a fixed region. We evaluate the effectiveness of our approach by comparing CMDS with other competing methods through a controlled user study and a quantitative study, and demonstrate the usefulness of CorrelatedMultiples for visual search and comparison in three real‐world case studies.
Yifan Hu 0001, Stephen C. North, Han-Wei Shen
Comput. Graph. Forum4
2018 Uncertainty Visualization Using Copula-Based Analysis in Mixed Distribution Models
abstract
Distributions are often used to model uncertainty in many scientific datasets. To preserve the correlation among the spatially sampled grid locations in the dataset, various standard multivariate distribution models have been proposed in visualization literature. These models treat each grid location as a univariate random variable which models the uncertainty at that location. Standard multivariate distributions (both parametric and nonparametric) assume that all the univariate marginals are of the same type/family of distribution. But in reality, different grid locations show different statistical behavior which may not be modeled best by the same type of distribution. In this paper, we propose a new multivariate uncertainty modeling strategy to address the needs of uncertainty modeling in scientific datasets. Our proposed method is based on a statistically sound multivariate technique called Copula, which makes it possible to separate the process of estimating the univariate marginals and the process of modeling dependency, unlike the standard multivariate distributions. The modeling flexibility offered by our proposed method makes it possible to design distribution fields which can have different types of distribution (Gaussian, Histogram, KDE etc.) at the grid locations, while maintaining the correlation structure at the same time. Depending on the results of various standard statistical tests, we can choose an optimal distribution representation at each location, resulting in a more cost efficient modeling without significantly sacrificing on the analysis quality. To demonstrate the efficacy of our proposed modeling strategy, we extract and visualize uncertain features like isocontours and vortices in various real world datasets. We also study various modeling criterion to help users in the task of univariate model selection.
Subhashis Hazarika, Ayan Biswas 0001, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.3
2018 GANViz: A Visual Analytics Approach to Understand the Adversarial Game
abstract
Generative models bear promising implications to learn data representations in an unsupervised fashion with deep learning. Generative Adversarial Nets (GAN) is one of the most popular frameworks in this arena. Despite the promising results from different types of GANs, in-depth understanding on the adversarial training process of the models remains a challenge to domain experts. The complexity and the potential long-time training process of the models make it hard to evaluate, interpret, and optimize them. In this work, guided by practical needs from domain experts, we design and develop a visual analytics system, GANViz, aiming to help experts understand the adversarial process of GANs in-depth. Specifically, GANViz evaluates the model performance of two subnetworks of GANs, provides evidence and interpretations of the models' performance, and empowers comparative analysis with the evidence. Through our case studies with two real-world datasets, we demonstrate that GANViz can provide useful insight into helping domain experts understand, interpret, evaluate, and potentially improve GAN models.
Junpeng Wang 0001, Liang Gou, Hao Yang 0007, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.4
2018 Uncertainty visualization for variable associations analysis
Dezhan Qu, Quanle Liu, Qi Shang, Yafang Hou, Han-Wei Shen
Vis. Comput.6
2017 Homogeneity guided probabilistic data summaries for analysis and visualization of large-scale data sets
abstract
High-resolution simulation data sets provide plethora of information, which needs to be explored by application scientists to gain enhanced understanding about various phenomena. Visual-analytics techniques using raw data sets are often expensive due to the data sets' extreme sizes. But, interactive analysis and visualization is crucial for big data analytics, because scientists can then focus on the important data and make critical decisions quickly. To assist efficient exploration and visualization, we propose a new region-based statistical data summarization scheme. Our method is superior in quality, as compared to the existing statistical summarization techniques, with a more compact representation, reducing the overall storage cost. The quantitative and visual efficacy of our proposed method is demonstrated using several data sets along with an in situ application study for an extreme-scale flow simulation.
Soumya Dutta, Jonathan Woodring, Han-Wei Shen, Jen-Ping Chen, James P. Ahrens
PacificVis3
2017 Range likelihood tree: A compact and effective representation for visual exploration of uncertain data sets
abstract
Uncertain data visualization plays a fundamental role in many applications such as weather forecast and analysis of fluid flows. Exploring scalar uncertain data modeled as probability distribution fields is a challenging task because the underlying features are often more complex, and the data associated with each grid point are high dimensional. In this work, we present a compact and effective representation, called range likelihood tree, to summarize and explore probability distribution fields. The key idea is to decompose and summarize each complex probability distribution over a few representative subranges by cumulative probabilities, and allow users to consider the roles that different subranges play in understanding the probability distributions. In our method, the value domain is first partitioned into subranges, then the distribution at each grid point is transformed according to the cumulative probabilities of the point's distribution in those subranges. Organizing the subranges into a hierarchical structure based on how these cumulative probabilities are spatially distributed in the grid points, the new range likelihood tree representation allows effective classification and identification of features through user query and exploration. We present an exploration framework with multiple interactive views to explore probability distribution fields, and provide guidelines for visual exploration using our framework. We demonstrate the effectiveness and usefulness of our approach in exploratory analysis using several representative uncertain data sets.
Han-Wei Shen, Scott M. Collis, Jonathan J. Helmus
PacificVis3
2017 Virtual retractor: An interactive data exploration system using physically based deformation
abstract
Interactive data exploration plays a fundamental role in analyzing three dimensional scientific data. Occlusion management and context preservation are among the key factors to ensure effective identification and extraction of three-dimensional features. In this paper, we present an interactive data exploration system that utilizes a physically based deformation method to investigate hidden structures of data in three dimensional data sets. While non-physically based methods are popular for visual analytic applications due to their lower computational cost, physically based deformation methods can often better preserve features and their context. Our physically based deformation method preserves data features by setting the mesh properties according to interesting data attributes. We design effective and intuitive interfaces by using a metaphor of virtual retractor, which reflects the cutting and splitting of data that our system is simulating. We demonstrate case studies on multiple particle datasets and volume datasets, and present feedback from a domain user.
Cheng Li 0062, Xin Tong 0012, Han-Wei Shen
PacificVis3
2017 Multivariate volumetric data analysis and visualization through bottom-up subspace exploration
abstract
Multivariate volumetric datasets are often encountered in results generated by scientific simulations. Compared to univariate datasets, analysis and visualization of multivariate datasets are much more challenging due to the complex relationships among the variables. As an effective way to visualize and analyze multivariate datasets, volume rendering has been frequently used, although designing good multivariate transfer functions is still non-trivial. In this paper, we present an interactive workflow to allow users to design multivariate transfer functions. To handle large scale datasets, in the preprocessing stage we reduce the number of data points through data binning and aggregation, and then a new set of data points with a much smaller size are generated. The relationship between all pairs of variables is presented in a matrix juxtaposition view, where users can navigate through the different subspaces. An entropy based method is used to help users to choose which subspace to explore. We proposed two weights: scatter weight and size weight that are associated with each projected point in those different subspaces. Based on those two weights, data point filter and kernel density estimation operations are employed to assist users to discover interesting features. For each user-selected feature, a Gaussian function is constructed and updated incrementally. Finally, all those selected features are visualized through multivariate volume rendering to reveal the structure of data. With our system, users can interactively explore different subspaces and specify multivariate transfer functions in an effective way. We demonstrate the effectiveness of our system with several multivariate volumetric datasets.
Kewei Lu, Han-Wei Shen
PacificVis2
2017 Statistical visualization and analysis of large data using a value-based spatial distribution
abstract
The size of large-scale scientific datasets created from simulations and computed on modern supercomputers continues to grow at a fast pace. A daunting challenge is to analyze and visualize these intractable datasets on commodity hardware. A recent and promising area of research is to replace the dataset with a distribution based proxy representation that summarizes scalar information into a much reduced memory footprint. Proposed representations subdivide the dataset into local blocks, where each block holds important statistical information, such as a histogram. A key drawback is that a distribution representing the scalar values in a block lacks spatial information. This manifests itself as large errors in visualization algorithms. We present a novel statistically-based representation by augmenting the block-wise distribution based representation with location information, called a value-based spatial distribution. Information from both spatial and scalar spaces are combined using Bayes' rule to accurately estimate the data value at a given spatial location. The representation is compact using the Gaussian Mixture Model. We show that our approach is able to preserve important features in the data and alleviate uncertainty.
Ko-Chih Wang, Kewei Lu, Tzu-Hsuan Wei, Naeem Shareef, Han-Wei Shen
PacificVis5
2017 Efficient distribution-based feature search in multi-field datasets
abstract
Local distribution search is used in query-driven visualization for identifying salient features. Due to the high computational and storage costs, local distribution search in multi-field datasets is challenging. In this paper, we introduce two high performance, memory efficient algorithms for searching for local distributions that are characterized by marginal and joint features in multi-field datasets. They leverage bitmap indexing and local voting to efficiently extract regions that match a target distribution, by first approximating search results and refining to generate the final result. The first algorithm, merged-bin-comparison (MBC), reduces the computation of histogram dissimilarity measures by clustering bins. The second algorithm, sampled-active voxels (SAV), adopts stratified sampling to reduce the workload for searching local distributions with large spatial neighborhoods. The efficiency and efficacy of our algorithms are demonstrated in multiple experiments.
Tzu-Hsuan Wei, Chun-Ming Chen, Jonathan Woodring, Han-Wei Shen
PacificVis5
2017 Tutorial on information theory in visualization
abstract
course Public Access Share on Tutorial on information theory in visualization Authors: Mateu Sbert University of Girona (Spain) and Tianjin University (China) University of Girona (Spain) and Tianjin University (China)View Profile , Han-Wei Shen View Profile , Ivan Viola View Profile , Min Chen View Profile , Anton Bardera View Profile , Miquel Feixas View Profile Authors Info & Claims SA '17: SIGGRAPH Asia 2017 CoursesNovember 2017 Article No.: 17Pages 1–165https://doi.org/10.1145/3134472.3134507Published:27 November 2017Publication History 1citation495DownloadsMetricsTotal Citations1Total Downloads495Last 12 Months76Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Mateu Sbert, Han-Wei Shen, Ivan Viola, Min Chen 0001, Anton Bardera, Miquel Feixas
SIGGRAPH ASIA (Courses)2
2017 Visualization of Time-Varying Weather Ensembles across Multiple Resolutions
abstract
Uncertainty quantification in climate ensembles is an important topic for the domain scientists, especially for decision making in the real-world scenarios. With powerful computers, simulations now produce time-varying and multi-resolution ensemble data sets. It is of extreme importance to understand the model sensitivity given the input parameters such that more computation power can be allocated to the parameters with higher influence on the output. Also, when ensemble data is produced at different resolutions, understanding the accuracy of different resolutions helps the total time required to produce a desired quality solution with improved storage and computation cost. In this work, we propose to tackle these non-trivial problems on the Weather Research and Forecasting (WRF) model output. We employ a moment independent sensitivity measure to quantify and analyze parameter sensitivity across spatial regions and time domain. A comparison of clustering structures across three resolutions enables the users to investigate the sensitivity variation over the spatial regions of the five input parameters. The temporal trend in the sensitivity values is explored via an MDS view linked with a line chart for interactive brushing. The spatial and temporal views are connected to provide a full exploration system for complete spatio-temporal sensitivity analysis. To analyze the accuracy across varying resolutions, we formulate a Bayesian approach to identify which regions are better predicted at which resolutions compared to the observed precipitation. This information is aggregated over the time domain and finally encoded in an output image through a custom color map that guides the domain experts towards an adaptive grid implementation given a cost model. Users can select and further analyze the spatial and temporal error patterns for multi-resolution accuracy analysis via brushing and linking on the produced image. In this work, we collaborate with a domain expert whose feedback shows the effectiveness of our proposed exploration work-flow.
Ayan Biswas 0001, Guang Lin 0001, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.4
2017 In Situ Distribution Guided Analysis and Visualization of Transonic Jet Engine Simulations
abstract
Study of flow instability in turbine engine compressors is crucial to understand the inception and evolution of engine stall. Aerodynamics experts have been working on detecting the early signs of stall in order to devise novel stall suppression technologies. A state-of-the-art Navier-Stokes based, time-accurate computational fluid dynamics simulator, TURBO, has been developed in NASA to enhance the understanding of flow phenomena undergoing rotating stall. Despite the proven high modeling accuracy of TURBO, the excessive simulation data prohibits post-hoc analysis in both storage and I/O time. To address these issues and allow the expert to perform scalable stall analysis, we have designed an in situ distribution guided stall analysis technique. Our method summarizes statistics of important properties of the simulation data in situ using a probabilistic data modeling scheme. This data summarization enables statistical anomaly detection for flow instability in post analysis, which reveals the spatiotemporal trends of rotating stall for the expert to conceive new hypotheses. Furthermore, the verification of the hypotheses and exploratory visualization using the summarized data are realized using probabilistic visualization techniques such as uncertain isocontouring. Positive feedback from the domain scientist has indicated the efficacy of our system in exploratory stall analysis.
Soumya Dutta, Chun-Ming Chen, Gregory Heinlein, Han-Wei Shen, Jen-Ping Chen
IEEE Trans. Vis. Comput. Graph.4
2017 GlyphLens: View-Dependent Occlusion Management in the Interactive Glyph Visualization
abstract
Glyph as a powerful multivariate visualization technique is used to visualize data through its visual channels. To visualize 3D volumetric dataset, glyphs are usually placed on 2D surface, such as the slicing plane or the feature surface, to avoid occluding each other. However, the 3D spatial structure of some features may be missing. On the other hand, placing large number of glyphs over the entire 3D space results in occlusion and visual clutter that make the visualization ineffective. To avoid the occlusion, we propose a view-dependent interactive 3D lens that removes the occluding glyphs by pulling the glyphs aside through the animation. We provide two space deformation models and two lens shape models to displace the glyphs based on their spatial distributions. After the displacement, the glyphs around the user-interested region are still visible as the context information, and their spatial structures are preserved. Besides, we attenuate the brightness of the glyphs inside the lens based on their depths to provide more depth cue. Furthermore, we developed an interactive glyph visualization system to explore different glyph-based visualization applications. In the system, we provide a few lens utilities that allows users to pick a glyph or a feature and look at it from different view directions. We compare different display/interaction techniques to visualize/manipulate our lens and glyphs.
Xin Tong 0012, Cheng Li 0062, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.3
2017 Multi-Resolution Climate Ensemble Parameter Analysis with Nested Parallel Coordinates Plots
abstract
Due to the uncertain nature of weather prediction, climate simulations are usually performed multiple times with different spatial resolutions. The outputs of simulations are multi-resolution spatial temporal ensembles. Each simulation run uses a unique set of values for multiple convective parameters. Distinct parameter settings from different simulation runs in different resolutions constitute a multi-resolution high-dimensional parameter space. Understanding the correlation between the different convective parameters, and establishing a connection between the parameter settings and the ensemble outputs are crucial to domain scientists. The multi-resolution high-dimensional parameter space, however, presents a unique challenge to the existing correlation visualization techniques. We present Nested Parallel Coordinates Plot (NPCP), a new type of parallel coordinates plots that enables visualization of intra-resolution and inter-resolution parameter correlations. With flexible user control, NPCP integrates superimposition, juxtaposition and explicit encodings in a single view for comparative data visualization and analysis. We develop an integrated visual analytics system to help domain scientists understand the connection between multi-resolution convective parameters and the large spatial temporal ensembles. Our system presents intricate climate ensembles with a comprehensive overview and on-demand geographic details. We demonstrate NPCP, along with the climate ensemble visualization system, based on real-world use-cases from our collaborators in computational and predictive science.
Junpeng Wang 0001, Han-Wei Shen, Guang Lin 0001
IEEE Trans. Vis. Comput. Graph.3
2016 Visualizing the variations of ensemble of isosurfaces
abstract
Visualizing the similarities and differences among an ensemble of isosurfaces is a challenging problem mainly because the isosurfaces cannot be displayed together at the same time. For ensemble of isosurfaces, visualizing these spatial differences among the surfaces is essential to get useful insights as to how the individual ensemble simulations affect different isosurfaces. We propose a scheme to visualize the spatial variations of isosurfaces with respect to statistically significant isosurfaces within the ensemble. Understanding such variations among ensemble of isosurfaces at different spatial regions is helpful in analyzing the influence of different ensemble runs over the spatial domain. In this regard, we propose an isosurface-entropy based clustering scheme to divide the spatial domain into regions of high and low isosurface variation. We demonstrate the efficacy of our method by successfully applying it on real-world ensemble data sets from ocean simulation experiments and weather forecasts.
Subhashis Hazarika, Soumya Dutta, Han-Wei Shen
PacificVis3
2016 A Bayesian approach for probabilistic streamline computation in uncertain flows
abstract
Streamline-based techniques play an important role in visualizing and analyzing uncertain steady vector fields. It is a challenging problem to generate accurate streamlines in uncertain vector fields due to the global uncertainty transportation. In this work, we present a novel probabilistic method for streamline computation on uncertain steady vector fields using a Bayesian framework. In our framework, a streamline is modeled as a state space model which captures the spatial coherence of integration steps and uncertainty in local distributions using the conditional prior density and the likelihood function. To approximate the posterior distribution for all the possible traces originating from a given seed position, a set of weighted samples are iteratively updated from which streamlines with higher likelihood can be derived. We qualitatively and quantitatively compare our method with alternative methods on different types of flow field data sets. Our method can generate possible streamlines with higher certainty and hence more accurate flow traces.
Chun-Ming Chen, Han-Wei Shen
PacificVis4
2016 Visualization and Analysis of Rotating Stall for Transonic Jet Engine Simulation
abstract
Identification of early signs of rotating stall is essential for the study of turbine engine stability. With recent advancements of high performance computing, high-resolution unsteady flow fields allow in depth exploration of rotating stall and its possible causes. Performing stall analysis, however, involves Significant effort to process large amounts of simulation data, especially when investigating abnormalities across many time steps. In order to assist scientists during the exploration process, we present a visual analytics framework to identify suspected spatiotemporal regions through a comparative visualization so that scientists are able to focus on relevant data in more detail. To achieve this, we propose efficient stall analysis algorithms derived from domain knowledge and convey the analysis results through juxtaposed interactive plots. Using our integrated visualization system, scientists can visually investigate the detected regions for potential stall initiation and further explore these regions to enhance the understanding of this phenomenon. Positive feedback from scientists demonstrate the efficacy of our system in analyzing rotating stall.
Chun-Ming Chen, Soumya Dutta, Gregory Heinlein, Han-Wei Shen, Jen-Ping Chen
IEEE Trans. Vis. Comput. Graph.5
2016 Distribution Driven Extraction and Tracking of Features for Time-varying Data Analysis
abstract
Effective analysis of features in time-varying data is essential in numerous scientific applications. Feature extraction and tracking are two important tasks scientists rely upon to get insights about the dynamic nature of the large scale time-varying data. However, often the complexity of the scientific phenomena only allows scientists to vaguely define their feature of interest. Furthermore, such features can have varying motion patterns and dynamic evolution over time. As a result, automatic extraction and tracking of features becomes a non-trivial task. In this work, we investigate these issues and propose a distribution driven approach which allows us to construct novel algorithms for reliable feature extraction and tracking with high confidence in the absence of accurate feature definition. We exploit two key properties of an object, motion and similarity to the target feature, and fuse the information gained from them to generate a robust feature-aware classification field at every time step. Tracking of features is done using such classified fields which enhances the accuracy and robustness of the proposed algorithm. The efficacy of our method is demonstrated by successfully applying it on several scientific data sets containing a wide range of dynamic time-varying features.
Soumya Dutta, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2016 Finite-Time Lyapunov Exponents and Lagrangian Coherent Structures in Uncertain Unsteady Flows
abstract
The objective of this paper is to understand transport behavior in uncertain time-varying flow fields by redefining the finite-time Lyapunov exponent (FTLE) and Lagrangian coherent structure (LCS) as stochastic counterparts of their traditional deterministic definitions. Three new concepts are introduced: the distribution of the FTLE (D-FTLE), the FTLE of distributions (FTLE-D), and uncertain LCS (U-LCS). The D-FTLE is the probability density function of FTLE values for every spatiotemporal location, which can be visualized with different statistical measurements. The FTLE-D extends the deterministic FTLE by measuring the divergence of particle distributions. It gives a statistical overview of how transport behaviors vary in neighborhood locations. The U-LCS, the probabilities of finding LCSs over the domain, can be extracted with stochastic ridge finding and density estimation algorithms. We show that our approach produces better results than existing variance-based methods do. Our experiments also show that the combination of D-FTLE, FTLE-D, and U-LCS can help users understand transport behaviors and find separatrices in ensemble simulations of atmospheric processes.
Hanqi Guo 0001, Tom Peterka, Han-Wei Shen, Scott M. Collis, Jonathan J. Helmus
IEEE Trans. Vis. Comput. Graph.4
2016 Association Analysis for Visual Exploration of Multivariate Scientific Data Sets
abstract
The heterogeneity and complexity of multivariate characteristics poses a unique challenge to visual exploration of multivariate scientific data sets, as it requires investigating the usually hidden associations between different variables and specific scalar values to understand the data's multi-faceted properties. In this paper, we present a novel association analysis method that guides visual exploration of scalar-level associations in the multivariate context. We model the directional interactions between scalars of different variables as information flows based on association rules. We introduce the concepts of informativeness and uniqueness to describe how information flows between scalars of different variables and how they are associated with each other in the multivariate domain. Based on scalar-level associations represented by a probabilistic association graph, we propose the Multi-Scalar Informativeness-Uniqueness (MSIU) algorithm to evaluate the informativeness and uniqueness of scalars. We present an exploration framework with multiple interactive views to explore the scalars of interest with confident associations in the multivariate spatial domain, and provide guidelines for visual exploration using our framework. We demonstrate the effectiveness and usefulness of our approach through case studies using three representative multivariate scientific data sets.
Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2016 View-Dependent Streamline Deformation and Exploration
abstract
Occlusion presents a major challenge in visualizing 3D flow and tensor fields using streamlines. Displaying too many streamlines creates a dense visualization filled with occluded structures, but displaying too few streams risks losing important features. We propose a new streamline exploration approach by visually manipulating the cluttered streamlines by pulling visible layers apart and revealing the hidden structures underneath. This paper presents a customized view-dependent deformation algorithm and an interactive visualization tool to minimize visual clutter in 3D vector and tensor fields. The algorithm is able to maintain the overall integrity of the fields and expose previously hidden structures. Our system supports both mouse and direct-touch interactions to manipulate the viewing perspectives and visualize the streamlines in depth. By using a lens metaphor of different shapes to select the transition zone of the targeted area interactively, the users can move their focus and examine the vector or tensor field freely.
Xin Tong 0012, John Edwards 0002, Chun-Ming Chen, Han-Wei Shen, Chris R. Johnson 0001, Pak Chung Wong
IEEE Trans. Vis. Comput. Graph.4
2015 An uncertainty-driven approach to vortex analysis using oracle consensus and spatial proximity
abstract
Although vortex analysis and detection have been extensively investigated in the past, none of the existing techniques are able to provide fully robust and reliable identification results. Local vortex detection methods are popular as they are efficient and easy to implement, and produce binary outputs based on a user-specified, hard threshold. However, vortices are global features, which present challenges for local detectors. On the other hand, global detectors are computationally intensive and require considerable user input. In this work, we propose a consensus-based uncertainty model and introduce spatial proximity to enhance vortex detection results obtained using point-based methods. We use four existing local vortex detectors and convert their outputs into fuzzy possibility values using a sigmoid-based soft-thresholding approach. We apply a majority voting scheme that enables us to identify candidate vortex regions with a higher degree of confidence. Then, we introduce spatial proximity- based analysis to discern the final vortical regions. Thus, by using spatial proximity coupled with fuzzy inputs, we propose a novel uncertainty analysis approach for vortex detection. We use expert's input to better estimate the system parameters and results from two real-world data sets demonstrate the efficacy of our method.
Ayan Biswas 0001, David S. Thompson, Chun-Ming Chen, Han-Wei Shen, Raghu Machiraju, Anand Rangarajan 0001
PacificVis6
2015 Uncertainty modeling and error reduction for pathline computation in time-varying flow fields
abstract
When the spatial and temporal resolutions of a time-varying simulation become very high, it is not possible to process or store data from every time step due to the high computation and storage cost. Although using uniformly down-sampled data for visualization is a common practice, important information in the un-stored data can be lost. Currently, linear interpolation is a popular method used to approximate data between the stored time steps. For pathline computation, however, errors from the interpolated velocity in the time dimension can accumulate quickly and make the trajectories rather unreliable. To inform the scientist the error involved in the visualization, it is important to quantify and display the uncertainty, and more importantly, to reduce the error whenever possible. In this paper, we present an algorithm to model temporal interpolation error, and an error reduction scheme to improve the data accuracy for temporally down-sampled data. We show that it is possible to compute polynomial regression and measure the interpolation errors incrementally with one sequential scan of the time-varying flow field. We also show empirically that when the data sequence is fitted with least-squares regression, the errors can be approximated with a Gaussian distribution. With the end positions of particle traces stored, we show that our error modeling scheme can better estimate the intermediate particle trajectories between the stored time steps based on a maximum likelihood method that utilizes forward and backward particle traces.
Chun-Ming Chen, Ayan Biswas 0001, Han-Wei Shen
PacificVis3
2015 Interactive streamline exploration and manipulation using deformation
abstract
Occlusion presents a major challenge in visualizing 3D flow fields using streamlines. Displaying too many streamlines creates a dense visualization filled with occluded structures, but displaying too few streams risks losing important features. A more ideal streamline exploration approach is to visually manipulate the cluttered streamlines by pulling visible layers apart and revealing the hidden structures underneath. This paper presents a customized deformation algorithm and an interactive visualization tool to minimize visual cluttering. The algorithm is able to maintain the overall integrity of the flow field and expose the previously hidden structures. Our system supports both mouse and direct-touch interactions to manipulate the viewing perspectives and visualize the streamlines in depth. By using a lens metaphor of different shapes to select the transition zone of the targeted area interactively, the users can move their focus and examine the flow field freely.
Xin Tong 0012, Chun-Ming Chen, Han-Wei Shen, Pak Chung Wong
PacificVis3
2015 The Effects of Representation and Juxtaposition on Graphical Perception of Matrix Visualization
abstract
Analyzing multiple networks at once is a common yet difficult task in many domains. Using adjacency matrices for this purpose, however, can be effective because of its superior ability to accommodate dense networks in a small area. We evaluate various representations and juxtaposition designs for visualizing adjacency matrices through a series of controlled experiments. We investigate the effect of using square matrices and triangular matrices on the speed and accuracy of performing graphical-perception tasks. Based on human symmetric perception, we propose two alternative juxtaposition designs to the conventional side-by-side juxtaposition, and study how users perform visual search and comparison tasks regarding different juxtaposition types. Our results show that the matrix representations have similar performance, and the matrix juxtaposition types perform differently. With the design guidelines derived from our studies, we present a compact visualization termed TileMatrix for juxtaposing a large number of matrices, and demonstrate its effectiveness in analyzing multi-faceted, time-varying networks using real-world data.
Han-Wei Shen
CHI2
2014 Efficient Range Distribution Query for Visualizing Scientific Data
abstract
Visualization applications implicitly run queries on the data to retrieve distributions and statistical measures derivable from distributions. Distribution based data summaries can substitute for the raw data to answer statistical queries of different kinds. However, frequent access to the raw data is no longer practical, if possible at all, for answering large number of queries on large-scale data. Our work addresses the issue by accelerating range distribution query, which returns the distribution of an axis-aligned query region. Maintaining the interactivity of such query is a challenging task because the workload and the response time of such queries scale up with the data and the query size. In this paper, we present a framework for answering range distribution queries for any arbitrary region in near constant time, regardless of data and query size. We adapt an integral histogram based data structure to bound the workload which is a combination of computation, I/O and communication cost. We propose two novel transformations of this data structure -- a decomposition and a similarity-driven indexing -- to reduce the huge storage cost associated with it. In addition to studying the performance of range distribution query, we also demonstrate the benefits that our technique offers to visualization applications which directly or indirectly require distributions.
Abon Chaudhuri, Tzu-Hsuan Wei, Teng-Yok Lee, Han-Wei Shen, Tom Peterka
PacificVis4
2014 EmailMap: Visualizing Event Evolution and Contact Interaction within Email Archives
abstract
Email archives contain rich information about how we interact with different contacts and how events evolve throughout time. Making sense of the archived messages can be a good way to understand how things evolved and progressed in the past. Although much work has been devoted to email visualization, most work has focused on presenting one of the two aspects of email archives: discovering the evolution of emails and events, or the relationship between the email owner and his/her contacts over time. In this paper, we present Email Map, an email visualization which integrates the information of both events and contacts into a single view, enabling users to make sense of their email archives with complementary contextual information. Two visualization components are designed to portray complex information within the email archives: event flow and contact tracks. The event flow illustrates the evolution of past events, helping the users to grasp high-level pictures and patterns of their email archives. The contact tracks reveal the interaction between the email owner and his/her contacts.
Sheng-Jie Luo, Liting Huang, Bing-Yu Chen 0004, Han-Wei Shen
PacificVis4
2014 Supporting correlation analysis on scientific datasets in parallel and distributed settings
abstract
With growing computational capabilities of parallel machines, scientific simulations are being performed at finer spatial and temporal scales, leading to a data explosion. Careful analysis of this data holds much promise for future scientific discoveries. Particularly, correlation analysis, which focuses on studying the potential relationships among multiple variables, is becoming a useful method for scientific analysis. This paper focuses on the problem of correlation analysis across large-scale simulation datasets, including 1) accelerating this analysis with the use of bitmap indexing as a representative summary of the data, 2) developing efficient algorithms for parallel execution, 3) performing analysis in distributed environments, i.e., for cases where different attributes are stored in geographically distributed repositories, and 4) combining sampling with correlation analysis. These algorithms have been implemented in a system that provides a high-level API for specification of the analyses, including allowing correlation analysis on specified value-based and dimension-based subsets of the data, and supports interactive and incremental analysis. We have extensively evaluated our framework for efficiency, and have also carried out case studies with domain scientists to establish how it can aid data-driven discovery process.
Yu Su 0011, Gagan Agrawal, Jonathan Woodring, Ayan Biswas 0001, Han-Wei Shen
HPDC5
2014 Scalable Computation of Stream Surfaces on Large Scale Vector Fields
abstract
Stream surfaces and streamlines are two popular methods for visualizing three-dimensional flow fields. While several parallel streamline computation algorithms exist, relatively little research has been done to parallelize stream surface generation. This is because load-balanced parallel stream surface computation is nontrivial, due to the strong dependency in computing the positions of the particles forming the stream surface front. In this paper, we present a new algorithm that computes stream surfaces efficiently. In our algorithm, seeding curves are divided into segments, which are then assigned to the processes. Each process is responsible for integrating the segments assigned to it. To ensure a balanced computational workload, work stealing and dynamic refinement of seeding curve segments are employed to improve the overall performance. We demonstrate the effectiveness of our parallel stream surface algorithm using several large scale flow field data sets, and show the performance and scalability on HPC systems.
Kewei Lu, Han-Wei Shen, Tom Peterka
SC2
2014 Boosting Techniques for Physics-Based Vortex Detection
abstract
Abstract Robust automated vortex detection algorithms are needed to facilitate the exploration of large‐scale turbulent fluid flow simulations. Unfortunately, robust non‐local vortex detection algorithms are computationally intractable for large data sets and local algorithms, while computationally tractable, lack robustness. We argue that the deficiencies inherent to the local definitions occur because of two fundamental issues: the lack of a rigorous definition of a vortex and the fact that a vortex is an intrinsically non‐local phenomenon. As a first step towards addressing this problem, we demonstrate the use of machine learning techniques to enhance the robustness of local vortex detection algorithms. We motivate the presence of an expert‐in‐the‐loop using empirical results based on machine learning techniques. We employ adaptive boosting to combine a suite of widely used, local vortex detection algorithms, which we term weak classifiers, into a robust compound classifier. Fundamentally, the training phase of the algorithm, in which an expert manually labels small, spatially contiguous regions of the data, incorporates non‐local information into the resulting compound classifier. We demonstrate the efficacy of our approach by applying the compound classifier to two data sets obtained from computational fluid dynamical simulations. Our results demonstrate that the compound classifier has a reduced misclassification rate relative to the component classifiers.
Raghu Machiraju, Anand Rangarajan 0001, David S. Thompson, D. Keith Walters, Han-Wei Shen
Comput. Graph. Forum7
2014 Exploring Flow Fields Using Space-Filling Analysis of Streamlines
abstract
Large scale scientific simulations frequently use streamline based techniques to visualize flow fields. As the shape of a streamline is often related to some underlying property of the field, it is important to identify streamlines (or their parts) with unique geometric features. In this paper, we introduce a metric, called the box counting ratio, which measures the geometric complexity of streamlines by measuring their space-filling capacity at different scales. We propose a novel interactive visualization framework which utilizes this metric to extract, organize and visualize features of varying density and complexity hidden in large numbers of streamlines. The proposed framework extracts complex regions of varying density from the streamlines, and organizes and presents them on an interactive 2D information space, allowing user selection and visualization of streamlines. We also extend this framework to support exploration using an ensemble of measures including box counting ratio. Our framework allows the user to easily visualize and interact with features otherwise hidden in large vector field data. We strengthen our claims with case studies using combustion and climate simulation data sets.
Abon Chaudhuri, Teng-Yok Lee, Han-Wei Shen, Rephael Wenger
IEEE Trans. Vis. Comput. Graph.3
2013 Exploring vector fields with distribution-based streamline analysis
abstract
Streamline-based techniques are designed based on the idea that properties of streamlines are indicative of features in the underlying field. In this paper, we show that statistical distributions of measurements along the trajectory of a streamline can be used as a robust and effective descriptor to measure the similarity between streamlines. With the distribution-based approach, we present a framework for interactive exploration of 3D vector fields with streamline query and clustering. Streamline queries allow us to rapidly identify streamlines that share similar geometric features to the target streamline. Streamline clustering allows us to group together streamlines of similar shapes. Based on user's selection, different clusters with different features at different levels of detail can be visualized to highlight features in 3D flow fields. We demonstrate the utility of our framework with simulation data sets of varying nature and size.
Kewei Lu, Abon Chaudhuri, Teng-Yok Lee, Han-Wei Shen, Pak Chung Wong
PacificVis4
2013 Transformations for volumetric range distribution queries
abstract
Volumetric datasets continue to grow in size, and there is continued demand for interactive analysis on these datasets. Because storage device throughputs are not increasing as quickly, interactive analysis workflows are becoming working set-constrained. In an ideal workflow, the working set complexity of the interactive analysis portion of the workflow should depend primarily on the size of the analysis result being produced, rather than on the size of the data being analyzed. Past works in online analytical processing and visualization have addressed this problem within application-specific contexts, but have not generalized their solutions to a wider variety of visualization applications. We propose a general framework for reducing the working set complexity of the interactive portion of visualization workflows that can be built on top of distribution range queries, as well as a technique within this framework able to support multiple visualization applications. Transformations are applied in the preprocessing phase of the workflow to enable fast, approximate volumetric distribution range queries with low working set complexity. Interactive application algorithms are then adapted to make use of these distribution range queries, enabling efficient interactive workflows on large-scale data. We show that the proposed technique enables these applications to be scaled primarily in terms of the application result dataset size, rather than the input data size, enabling increased interactivity and scalability.
Han-Wei Shen
PacificVis2
2013 GraphCharter: Combining browsing with query to explore large semantic graphs
abstract
Large scale semantic graphs such as social networks and knowledge graphs contain rich and useful information. However, due to combined challenges in scale, density, and heterogeneity, it is impractical for users to answer many interesting questions by visual inspection alone. This is because even a semantically simple question, such as which of my extended friends are also fans of my favorite band, can in fact require information from a non-trivial number of nodes to answer. In this paper, we propose a method that combines graph browsing with query to overcome the limitation of visual inspection. By using query as the main way for information discovery in graph exploration, our “query, expand, and query again” model enables users to probe beyond the visible part of the graph and only bring in the interesting nodes, leaving the view clutter-free. We have implemented a prototype called GraphCharter and demonstrated its effectiveness and usability in a case study and a user study on Freebase knowledge graph with millions of nodes and edges.
Ying Tu, Han-Wei Shen
PacificVis2
2013 CompactMap: A mental map preserving visual interface for streaming text data
abstract
As text streams become increasingly available from social media such as Facebook and Twitter, visual analysis of streaming text data is playing an important role in most business sectors. A fundamental challenge in visualizing a large amount of streaming text data is to preserve the user's mental map to enable tracking dynamic changes in topics, while simultaneously utilizing the display space efficiently. In this paper, we present CompactMap, an online visual interface that packs text clusters efficiently, with stable updates to maintain the user's mental map. It achieves spatiotemporally coherent layouts by dynamically matching clusters across time, and removing cluster overlaps according to spatial proximity and constraints. We developed a visual search engine based on CompactMaps for exploring a large amount of text streams in details on demand. We demonstrate the effectiveness of our approach in a controlled user study compared with a competing method.
Yifan Hu 0001, Stephen C. North, Han-Wei Shen
IEEE BigData4
2013 Evaluating Isosurfaces with Level-set-based Information Maps
abstract
Abstract While isosurfaces have been widely used for scalar data visualization, it is often difficult to determine if the selected isosurfaces for visualization are sufficient to represent the entire scalar field. In this paper, we present an information‐theoretic approach to evaluate the representativeness of a given isosurface set. Our basic idea is that given two isosurfaces that enclose a subvolume, if the intermediate isosurfaces in the subvolume can be generated by smoothly morphing from one isosurface to the other, no additional isosurfaces are needed since the geometry of the true isosurfaces within the subvolume can be easily inferred. To realize this idea, given a pair of isosurfaces, to determine if such a smooth condition in the enclosed region is satisfied, we use a level‐set approach to generate the intermediate surfaces. On each intermediate surface, we sample the values from the scalar field and exam the distribution. If the entropy of the distribution is low, this intermediate surface is aligned well with a true isosurface in the scalar field. For the intermediate surfaces generated by the level‐set method from the boundary isosurfaces, the distributions of scalar values from the level‐set surfaces form a 2D distribution, called isosurface information map. This information map can be used as an indicator of the representativeness of the boundary isosurfaces for the data in the subregion, allowing a quantitative measurement of information representable by the input isosurfaces. Based on this information‐theoretic approach, this paper presents an isosurface selection algorithm that can automatically select isosurfaces for more effective visualization of scalar fields.
Tzu-Hsuan Wei, Teng-Yok Lee, Han-Wei Shen
Comput. Graph. Forum3
2013 An Information-Aware Framework for Exploring Multivariate Data Sets
abstract
Information theory provides a theoretical framework for measuring information content for an observed variable, and has attracted much attention from visualization researchers for its ability to quantify saliency and similarity among variables. In this paper, we present a new approach towards building an exploration framework based on information theory to guide the users through the multivariate data exploration process. In our framework, we compute the total entropy of the multivariate data set and identify the contribution of individual variables to the total entropy. The variables are classified into groups based on a novel graph model where a node represents a variable and the links encode the mutual information shared between the variables. The variables inside the groups are analyzed for their representativeness and an information based importance is assigned. We exploit specific information metrics to analyze the relationship between the variables and use the metrics to choose isocontours of selected variables. For a chosen group of points, parallel coordinates plots (PCP) are used to show the states of the variables and provide an interface for the user to select values of interest. Experiments with different data sets reveal the effectiveness of our proposed framework in depicting the interesting regions of the data sets taking into account the interaction among the variables.
Ayan Biswas 0001, Soumya Dutta, Han-Wei Shen, Jonathan Woodring
IEEE Trans. Vis. Comput. Graph.3
2013 Efficient Local Statistical Analysis via Integral Histograms with Discrete Wavelet Transform
abstract
Histograms computed from local regions are commonly used in many visualization applications, and allowing the user to query histograms interactively in regions of arbitrary locations and sizes plays an important role in feature identification and tracking. Computing histograms in regions with arbitrary location and size, nevertheless, can be time consuming for large data sets since it involves expensive I/O and scan of data elements. To achieve both performance- and storage-efficient query of local histograms, we present a new algorithm called WaveletSAT, which utilizes integral histograms, an extension of the summed area tables (SAT), and discrete wavelet transform (DWT). Similar to SAT, an integral histogram is the histogram computed from the area between each grid point and the grid origin, which can be be pre-computed to support fast query. Nevertheless, because one histogram contains multiple bins, it will be very expensive to store one integral histogram at each grid point. To reduce the storage cost for large integral histograms, WaveletSAT treats the integral histograms of all grid points as multiple SATs, each of which can be converted into a sparse representation via DWT, allowing the reconstruction of axis-aligned region histograms of arbitrary sizes from a limited number of wavelet coefficients. Besides, we present an efficient wavelet transform algorithm for SATs that can operate on each grid point separately in logarithmic time complexity, which can be extended to parallel GPU-based implementation. With theoretical and empirical demonstration, we show that WaveletSAT can achieve fast preprocessing and smaller storage overhead than the conventional integral histogram approach with close query performance.
Teng-Yok Lee, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2012 A flow-guided file layout for out-of-core streamline computation
abstract
We present a file layout algorithm for flow fields to improve runtime I/O efficiency for out-of-core streamline computation. Because of the increasing discrepancy between the speed of processors and storage devices, the cost of I/O becomes a major bottleneck for out-of-core computation. To reduce the I/O cost, loading data with better spatial locality has proved to be effective. It is also known that sequential file access is more efficient. To facilitate efficient streamline computation, we propose to reorganize the data blocks in a file following the data access pattern so that more efficient I/O and effective prefetching can be accomplished. To achieve the goal, we divide the domain into small spatial blocks and order the blocks into a linear layout based on the underlying flow directions. The ordering is done using a weighted directed graph model which can be formulated as a linear graph arrangement problem. Our goal is to arrange the file in a way consistent with the data access pattern during streamline computation. This allows us to prefetch a contiguous segment of data at a time from disk and minimize the memory cache miss rate. We use a recursive partitioning method to approximate the optimal layout. Our experimental results show that the resulting file layout reduces I/O cost and hence enables more efficient out-of-core streamline computation.
Chun-Ming Chen, Lijie Xu, Teng-Yok Lee, Han-Wei Shen
PacificVis4
2012 Parallel particle advection and FTLE computation for time-varying flow fields
abstract
Flow fields are an important product of scientific simulations. One popular flow visualization technique is particle advection, in which seeds are traced through the flow field. One use of these traces is to compute a powerful analysis tool called the Finite-Time Lyapunov Exponent (FTLE) field, but no existing particle tracing algorithms scale to the particle injection frequency required for high-resolution FTLE analysis. In this paper, a framework to trace the massive number of particles necessary for FTLE computation is presented. A new approach is explored, in which processes are divided into groups, and are responsible for mutually exclusive spans of time. This pipelining over time intervals reduces overall idle time of processes and decreases I/O overhead. Our parallel FTLE framework is capable of advecting hundreds of millions of particles at once, with performance scaling up to tens of thousands of processes.
Boonthanome Nouanesengsy, Teng-Yok Lee, Kewei Lu, Han-Wei Shen, Tom Peterka
SC4
2012 Coherent Time-Varying Graph Drawing with Multifocus+Context Interaction
abstract
We present a new approach for time-varying graph drawing that achieves both spatiotemporal coherence and multifocus+context visualization in a single framework. Our approach utilizes existing graph layout algorithms to produce the initial graph layout, and formulates the problem of generating coherent time-varying graph visualization with the focus+context capability as a specially tailored deformation optimization problem. We adopt the concept of the super graph to maintain spatiotemporal coherence and further balance the needs for aesthetic quality and dynamic stability when interacting with time-varying graphs through focus+context visualization. Our method is particularly useful for multifocus+context visualization of time-varying graphs where we can preserve the mental map by preventing nodes in the focus from undergoing abrupt changes in size and location in the time sequence. Experiments demonstrate that our method strikes a good balance between maintaining spatiotemporal coherence and accentuating visual foci, thus providing a more engaging viewing experience for the users.
Kun-Chuan Feng, Chaoli Wang 0001, Han-Wei Shen, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.3
2011 View point evaluation and streamline filtering for flow visualization
abstract
Visualization of flow fields with geometric primitives is often challenging due to occlusion that is inevitably introduced by 3D streamlines. In this paper, we present a novel view-dependent algorithm that can minimize occlusion and reveal important flow features for three dimensional flow fields. To analyze regions of higher importance, we utilize Shannon's entropy as a measure of vector complexity. An entropy field in the form of a three dimensional volume is extracted from the input vector field. To utilize this view-independent complexity measure for view-dependent calculations, we introduce the notion of a maximal entropy projection (MEP) framebuffer, which stores maximal entropy values as well as the corresponding depth values for a given viewpoint. With this information, we develop a view-dependent streamline selection algorithm that can evaluate and choose streamlines that will cause minimum occlusion to regions of higher importance. Based on a similar concept, we also propose a viewpoint selection algorithm that works hand-in-hand with our streamline selection algorithm to maximize the visibility of high complexity regions in the flow field.
Teng-Yok Lee, Oleg Mishchenko, Han-Wei Shen, Roger Crawfis
PacificVis3
2011 A Study of Parallel Particle Tracing for Steady-State and Time-Varying Flow Fields
abstract
Particle tracing for streamline and path line generation is a common method of visualizing vector fields in scientific data, but it is difficult to parallelize efficiently because of demanding and widely varying computational and communication loads. In this paper we scale parallel particle tracing for visualizing steady and unsteady flow fields well beyond previously published results. We configure the 4D domain decomposition into spatial and temporal blocks that combine in-core and out-of-core execution in a flexible way that favors faster run time or smaller memory. We also compare static and dynamic partitioning approaches. Strong and weak scaling curves are presented for tests conducted on an IBM Blue Gene/P machine at up to 32 K processes using a parallel flow visualization library that we are developing. Datasets are derived from computational fluid dynamics simulations of thermal hydraulics, liquid mixing, and combustion.
Tom Peterka, Robert B. Ross, Boonthanome Nouanesengsy, Teng-Yok Lee, Han-Wei Shen, Wesley Kendall, Jian Huang 0007
IPDPS5
2011 Guest Editor's Introduction: Special Section on the IEEE Pacific Visualization Symposium
Stephen C. North, Han-Wei Shen, Jarke J. van Wijk
IEEE Trans. Vis. Comput. Graph.2
2011 Load-Balanced Parallel Streamline Generation on Large Scale Vector Fields
abstract
Because of the ever increasing size of output data from scientific simulations, supercomputers are increasingly relied upon to generate visualizations. One use of supercomputers is to generate field lines from large scale flow fields. When generating field lines in parallel, the vector field is generally decomposed into blocks, which are then assigned to processors. Since various regions of the vector field can have different flow complexity, processors will require varying amounts of computation time to trace their particles, causing load imbalance, and thus limiting the performance speedup. To achieve load-balanced streamline generation, we propose a workload-aware partitioning algorithm to decompose the vector field into partitions with near equal workloads. Since actual workloads are unknown beforehand, we propose a workload estimation algorithm to predict the workload in the local vector field. A graph-based representation of the vector field is employed to generate these estimates. Once the workloads have been estimated, our partitioning algorithm is hierarchically applied to distribute the workload to all partitions. We examine the performance of our workload estimation and workload-aware partitioning algorithm in several timings studies, which demonstrates that by employing these methods, better scalability can be achieved with little overhead.
Boonthanome Nouanesengsy, Teng-Yok Lee, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.3
2010 CycleStack: Inferring periodic behavior via temporal sequence visualization in ultrasound video
abstract
A range of well-known treatment methods for destroying tumor and similar harmful growth in human body utilizes the coherence between the inherently periodic movement of the affected body part and periodic respiratory signal of the patient, with the objective of minimizing damage to surrounding normal tissues. Such methods require constant monitoring by an operator who observes the 3D body motion via its 2D projection onto an ultrasound imaging plane and studies the synchronism of this motion with the respiratory signal. Keeping an attentive eye on the respiratory signal as well as the ultrasound video for the entire treatment period is often inconvenient and burdensome. In this paper, we propose a video visualization technique called CycleStack Plot which reduces this cognitive overhead by blending the video and the signal together in a stack-like layout. This visualization reveals the inherent synchronism between the target's movement and the respiratory signal, visually highlights significant phase shifts of either of the two cyclic phenomena, with the hope of arresting the operator's attention. Our proposed visualization also provides a visual overview for the post-treatment analysis which enables educated users to quickly and effectively skim through the excessively long process. This paper demonstrates the utility of CycleStack Plot with a case study using real ultrasound videos. In addition, a user study has been performed to evaluate the merits and limitations of the proposed method with respect to the conventional way of watching a video and a signal side-by-side. Even though the motivation of the proposed visualization is improvement of medical applications that use ultrasound, the core techniques discussed here have potential to be extended to other application domains requiring analysis of cyclic patterns from videos.
Teng-Yok Lee, Abon Chaudhuri, Fatih Porikli, Han-Wei Shen
PacificVis4
2010 A message from the program chairs
abstract
Greetings from the program committee of Pacific Visualization 2010. This was the third year of the conference. The quality of this year's program conveys exciting work that is being done in many areas of visualization. In the few years since the inception of the meeting, there has been a steady increase in the number of submissions and their overall technical strength. This reflects the maturity of the field, and perhaps increasing support for a spring meeting to complement VisWeek held in the fall.
Stephen C. North, Han-Wei Shen, Jarke J. van Wijk
PacificVis2
2010 Novel Geometrical Voxelization Approach with Application to Streamlines
Hsien-Hsi Hsieh, Chin-Chen Chang 0002, Wen-Kai Tai, Han-Wei Shen
J. Comput. Sci. Technol.4
2010 An Information-Theoretic Framework for Flow Visualization
abstract
The process of visualization can be seen as a visual communication channel where the input to the channel is the raw data, and the output is the result of a visualization algorithm. From this point of view, we can evaluate the effectiveness of visualization by measuring how much information in the original data is being communicated through the visual communication channel. In this paper, we present an information-theoretic framework for flow visualization with a special focus on streamline generation. In our framework, a vector field is modeled as a distribution of directions from which Shannon's entropy is used to measure the information content in the field. The effectiveness of the streamlines displayed in visualization can be measured by first constructing a new distribution of vectors derived from the existing streamlines, and then comparing this distribution with that of the original data set using the conditional entropy. The conditional entropy between these two distributions indicates how much information in the original data remains hidden after the selected streamlines are displayed. The quality of the visualization can be improved by progressively introducing new streamlines until the conditional entropy converges to a small value. We describe the key components of our framework with detailed analysis, and show that the framework can effectively visualize 2D and 3D flow data.
Lijie Xu, Teng-Yok Lee, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.3
2009 A self-adaptive treemap-based technique for visualizing hierarchical data in 3D
abstract
In this paper, we present a novel adaptive visualization technique where the constituting polygons dynamically change their geometry and other visual attributes depending on user interaction. These changes take place with the objective of conveying required level of detail to the user through each view. Our proposed technique is successfully applied to build a treemap-based but 3D visualization of hierarchical data, a widely used information structure. This new visualization exploits its adaptive nature to address issues like cluttered display, imperceptible hierarchy, lack of smooth zoom-in and out technique which are common in tree visualization. We also present an algorithm which utilizes the flexibility of our proposed technique to deal with occlusion, a problem inherent in any 3D information visualization. On one hand, our work establishes adaptive visualization as a means of displaying tree-structured data in 3D. On the other, it promotes the technique as a potential candidate for being employed to visualize other information structures also.
Abon Chaudhuri, Han-Wei Shen
PacificVis2
2009 Out-of-core volume rendering for time-varying fields using a space-partitioning time (SPT) tree
abstract
In this paper, we propose a novel out-of-core volume rendering algorithm for large time-varying fields. Exploring temporal and spatial coherences has been an important direction for speeding up the rendering of time-varying data. Previously, there were techniques that hierarchically partition both the time and space domains into a data structure so as to re-use some results from the previous time step in multiresolution rendering; however, it has not been studied on which domain should be partitioned first to obtain a better re-use rate. We address this open question, and show both theoretically and experimentally that partitioning the time domain first is better. We call the resulting structure (a binary time tree as the primary structure and an octree as the secondary structure) the space-partitioning time (SPT) tree. Typically, our SPT-tree rendering has a higher level of details, a higher re-use rate, and runs faster. In addition, we devise a novel cut-finding algorithm to facilitate efficient out-of-core volume rendering using our SPT tree, we develop a novel out-of-core preprocessing algorithm to build our SPT tree I/O-efficiently, and we propose modified error metrics with a theoretical guarantee of a monotonicity property that is desirable for the tree search. The experiments on datasets as large as 25GB using a PC with only 2GB of RAM demonstrated the efficacy of our new approach.
Zhiyan Du, Yi-Jen Chiang, Han-Wei Shen
PacificVis3
2009 Visualizing time-varying features with TAC-based distance fields
abstract
To analyze time-varying data sets, tracking features over time is often necessary to better understand the dynamic nature of the underlying physical process. Tracking 3D time-varying features, however, is non-trivial when the boundaries of the features cannot be easily defined. In this paper, we propose a new framework to visualize time-varying features and their motion without explicit feature segmentation and tracking. In our framework, a time-varying feature is described by a time series or Time Activity Curve (TAC). To compute the distance, or similarity, between a voxel's time series and the feature, we use the Dynamic Time Warping (DTW) distance metric. The purpose of DTW is to compare the shape similarity between two time series with an optimal warping of time so that the phase shift of the feature in time can be accounted for. After applying DTW to compare each voxel's time series with the feature, a time-invariant distance field can be computed. The amount of time warping required for each voxel to match the feature provides an estimate of the time when the feature is most likely to occur. Based on the TAC-based distance field, several visualization methods can be derived to highlight the position and motion of the feature. We present several case studies to demonstrate and compare the effectiveness of our framework.
Teng-Yok Lee, Han-Wei Shen
PacificVis2
2009 A configurable algorithm for parallel image-compositing applications
abstract
Collective communication operations can dominate the cost of large-scale parallel algorithms. Image compositing in parallel scientific visualization is a reduction operation where this is the case. We present a new algorithm called Radix-k that in many cases performs better than existing compositing algorithms. It does so through a set of configurable parameters, the radices, that determine the number of communication partners in each message round. The algorithm embodies and unifies binary swap and direct-send, two of the best-known compositing methods, and enables numerous other configurations through appropriate choices of radices. While the algorithm is not tied to a particular computing architecture or network topology, the selection of radices allows Radix-k to take advantage of new supercomputer interconnect features such as multiporting. We show scalability across image size and system size, including both powers of two and nonpowers-of-two process counts.
Tom Peterka, David Goodell, Robert B. Ross, Han-Wei Shen, Rajeev Thakur
SC4
2009 Visual Analysis of Brain Activity from fMRI Data
abstract
Abstract Classically, analysis of the time‐varying data acquired during fMRI experiments is done using static activation maps obtained by testing voxels for the presence of significant activity using statistical methods. The models used in these analysis methods have a number of parameters, which profoundly impact the detection of active brain areas. Also, it is hard to study the temporal dependencies and cascading effects of brain activation from these static maps. In this paper, we propose a methodology to visually analyze the time dimension of brain function with a minimum amount of processing, allowing neurologists to verify the correctness of the analysis results, and develop a better understanding of temporal characteristics of the functional behaviour. The system allows studying time‐series data through specific volumes‐of‐interest in the brain‐cortex, the selection of which is guided by a hierarchical clustering algorithm performed in the wavelet domain. We also demonstrate the utility of this tool by presenting results on a real data‐set.
Firdaus Janoos, Boonthanome Nouanesengsy, Raghu Machiraju, Han-Wei Shen, Steffen Sammet, Michael V. Knopp, István Ákos Mórocz
Comput. Graph. Forum4
2009 Semi-Automatic Time-Series Transfer Functions via Temporal Clustering and Sequencing
abstract
Abstract When creating transfer functions for time‐varying data, it is not clear what range of values to use for classification, as data value ranges and distributions change over time. In order to generate time‐varying transfer functions, we search the data for classes that have similar behavior over time, assuming that data points that behave similarly belong to the same feature. We utilize a method we call temporal clustering and sequencing to find dynamic features in value space and create a corresponding transfer function. First, clustering finds groups of data points that have the same value space activity over time. Then, sequencing derives a progression of clusters over time, creating chains that follow value distribution changes. Finally, the cluster sequences are used to create transfer functions, as sequences describe the value range distributions over time in a data set.
Jonathan Woodring, Han-Wei Shen
Comput. Graph. Forum2
2009 Enhancing Realism of Wet Surfaces in Temporal Bone Surgical Simulation
abstract
We present techniques to improve visual realism in an interactive surgical simulation application: a mastoidectomy simulator that offers a training environment for medical residents as a complement to using a cadaver. As well as displaying the mastoid bone through volume rendering, the simulation allows users to experience haptic feedback and appropriate sound cues while controlling a virtual bone drill and suction/irrigation device. The techniques employed to improve realism consist of a fluid simulator and a shading model. The former allows for deformable boundaries based on volumetric bone data, while the latter gives a wet look to the rendered bone to emulate more closely the appearance of the bone in a surgical environment. The fluid rendering includes bleeding effects, meniscus rendering, and refraction. We incorporate a planar computational fluid dynamics simulation into our three-dimensional rendering to effect realistic blood diffusion. Maintaining real-time performance while drilling away bone in the simulation is critical for engagement with the system.
Thomas Kerwin, Han-Wei Shen, Don Stredney
IEEE Trans. Vis. Comput. Graph.2
2009 Visualization and Exploration of Temporal Trend Relationships in Multivariate Time-Varying Data
abstract
We present a new algorithm to explore and visualize multivariate time-varying data sets. We identify important trend relationships among the variables based on how the values of the variables change over time and how those changes are related to each other in different spatial regions and time intervals. The trend relationships can be used to describe the correlation and causal effects among the different variables. To identify the temporal trends from a local region, we design a new algorithm called SUBDTW to estimate when a trend appears and vanishes in a given time series. Based on the beginning and ending times of the trends, their temporal relationships can be modeled as a state machine representing the trend sequence. Since a scientific data set usually contains millions of data points, we propose an algorithm to extract important trend relationships in linear time complexity. We design novel user interfaces to explore the trend relationships, to visualize their temporal characteristics, and to display their spatial distributions. We use several scientific data sets to test our algorithm and demonstrate its utilities.
Teng-Yok Lee, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2009 Multiscale Time Activity Data Exploration via Temporal Clustering Visualization Spreadsheet
abstract
Time-varying data is usually explored by animation or arrays of static images. Neither is particularly effective for classifying data by different temporal activities. Important temporal trends can be missed due to the lack of ability to find them with current visualization methods. In this paper, we propose a method to explore data at different temporal resolutions to discover and highlight data based upon time-varying trends. Using the wavelet transform along the time axis, we transform data points into multi-scale time series curve sets. The time curves are clustered so that data of similar activity are grouped together, at different temporal resolutions. The data are displayed to the user in a global time view spreadsheet where she is able to select temporal clusters of data points, and filter and brush data across temporal scales. With our method, a user can interact with data based on time activities and create expressive visualizations.
Jonathan Woodring, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2008 Interactive Exploration of Remote Isosurfaces with Point-Based Non-Photorealistic Rendering
abstract
We present a non-photo realistic rendering technique for interactive exploration of isosurfaces generated from remote volumetric data. Instead of relying on the conventional smooth shading technique to render the isosurfaces, a point-based technique is used to represent and render the isosurfaces in a remote client-server environment. The non-photo realistic nature of the proposed rendering method enables the server to transmit only the essential surface features, which substantially reduces the network traffic. The algorithm also utilizes frame coherence and efficiently encodes the isosurface configuration inside each voxel cell to further minimize the network overhead. Finally, our algorithm can adjust the point distributions using different illumination settings to adapt to different network speeds.
Guangfeng Ji, Han-Wei Shen, Jinzhu Gao
PacificVis2
2008 Illustrative Streamline Placement and Visualization
abstract
Inspired by the abstracting, focusing and explanatory qualities of diagram drawing in art, in this paper we propose a novel seeding strategy to generate representative and illustrative streamlines in 2D vector fields to enforce visual clarity and evidence. A particular focus of our algorithm is to depict the underlying flow patterns effectively and succinctly with a minimum set of streamlines. To achieve this goal, 2D distance fields are generated to encode the distances from each grid point in the field to the nearby streamlines. A local metric is derived to measure the dissimilarity between the vectors from the original field and an approximate field computed from the distance fields. A global metric is used to measure the dissimilarity between streamlines based on the local errors to decide whether to drop a new seed at a local point. This process is iterated to generate streamlines until no more streamlines can be found that are dissimilar to the existing ones. We present examples of images generated from our algorithm and report results from qualitative analysis and user studies.
Liya Li, Hsien-Hsi Hsieh, Han-Wei Shen
PacificVis3
2008 Interactive Storyboard for Overall Time-Varying Data Visualization
abstract
Large amounts of time-varying datasets create great challenges for users to understand and explore them. This paper proposes an efficient visualization method for observing overall data contents and changes throughout an entire time-varying dataset. We develop an interactive storyboard approach by composing sample volume renderings and descriptive geometric primitives that are generated through data analysis processes. Our storyboard system integrates automatic visualization generation methods and interactive adjustment procedures to provide new tools for visualizing and exploring time-varying datasets. We also provide a flexible framework to quantify data differences and automatically select representative datasets through exploring scientific data distribution features. Since this approach reduces the visualized data amount into a more understandable size and format for users, it can be used to effectively visualize, represent, and explore a large time-varying dataset. Initial user study results show that our approach shortens the exploration time and reduces the number of datasets that users visualized individually. This visualization method is especially useful for situations that require close observance or are not capable of interactive rendering, such as documentation and demonstration.
Aidong Lu, Han-Wei Shen
PacificVis2
2008 Efficient Rendering of Extrudable Curvilinear Volumes
abstract
We present a technique for memory-efficient and time-efficient volume rendering of curvilinear adaptive mesh refinement data defined within extrudable computational spaces. One of the main challenges in the ray casting of curvilinear volumes is that a linear viewing ray in physical space will typically correspond to a curved ray in computational space. The proposed method utilizes a specialized representation of curvilinear space that provides for the compact representation of parameters for transformations between computational space and physical space, without requiring extensive preprocessing. By simplifying the representation of computational space positions using an extrusion of a profile surface, the requisite transformations can be greatly simplified. Our implementation achieves interactive rates with minimal load time and memory overhead using commodity graphics hardware with real-world data.
Han-Wei Shen, Ravi Samtaney
PacificVis2
2008 Supporting a visualization application on a self-adapting grid middleware
abstract
This paper describes how we have used a self-adapting middleware to implement a distributed and adaptive volume rendering application. The middleware we have used is GATES (grid-based adaptive execution on streams), which allows processing of streaming data in a distributed environment. A challenge in supporting such an application on streaming data is to balance the visualization quality and the speed of processing, which can be automatically done by the GATES middleware. We describe how we divide the application into a number of processing stages, and what adaptation parameters we use. Our experimental studies have focused on evaluating the self-adaptation enabled by the middleware, and measuring the overhead associated with the use of middleware.
Liang Chen 0019, Han-Wei Shen, Gagan Agrawal
IPDPS2
2008 Parallel reflective symmetry transformation for volume data
Han-Wei Shen
Comput. Graph.2
2008 Balloon Focus: a Seamless Multi-Focus+Context Method for Treemaps
abstract
The treemap is one of the most popular methods for visualizing hierarchical data. When a treemap contains a large number of items, inspecting or comparing a few selected items in a greater level of detail becomes very challenging. In this paper, we present a seamless multi-focus and context technique, called Balloon Focus, that allows the user to smoothly enlarge multiple treemap items served as the foci, while maintaining a stable treemap layout as the context. Our method has several desirable features. First, this method is quite general and hence can be used with different treemap layout algorithms. Second, as the foci are enlarged, the relative positions among all items are preserved. Third, the foci are placed in a way that the remaining space is evenly distributed back to the non-focus treemap items. When Balloon Focus enlarges the focus items to a maximum degree, the above features ensure that the treemap will maintain a consistent appearance and avoid any abrupt layout changes. In our algorithm, a DAG (Directed Acyclic Graph) is used to maintain the positional constraints, and an elastic model is employed to govern the placement of the treemap items. We demonstrate a treemap visualization system that integrates data query, manual focus selection, and our novel multi-focus+context technique, Balloon Focus, together. A user study was conducted. Results show that with Balloon Focus, users can better perform the tasks of comparing the values and the distribution of the foci.
Ying Tu, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2007 Parallel Reflective Symmetry Transformation for Volume Data
Han-Wei Shen
EGPGV2
2007 Image-Based Streamline Generation and Rendering
abstract
Seeding streamlines in 3D flow fields without considering their projections in screen space can produce visually cluttered rendering results. Streamlines will overlap or intersect with each other in the output image, which makes it difficult for the user to perceive the underlying flow structure. This paper presents a method to control the seeding and generation of streamlines in image space to avoid visual cluttering and allow a more flexible exploration of flow fields. In our algorithm, 2D images with depth maps generated by a variety of visualization techniques can be used as input from which seeds are placed and streamlines are generated. The density and rendering styles of streamlines can be flexibly controlled based on various criteria to improve visual clarity. With our image space approach, it is straightforward to implement the level of detail rendering, depth peeling, and stylized rendering of streamlines to allow for more effective visualization of 3D flow fields.
Liya Li, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2007 Visualizing Changes of Hierarchical Data using Treemaps
abstract
While the treemap is a popular method for visualizing hierarchical data, it is often difficult for users to track layout and attribute changes when the data evolve over time. When viewing the treemaps side by side or back and forth, there exist several problems that can prevent viewers from performing effective comparisons. Those problems include abrupt layout changes, a lack of prominent visual patterns to represent layouts, and a lack of direct contrast to highlight differences. In this paper, we present strategies to visualize changes of hierarchical data using treemaps. A new treemap layout algorithm is presented to reduce abrupt layout changes and produce consistent visual patterns. Techniques are proposed to effectively visualize the difference and contrast between two treemap snapshots in terms of the map items' colors, sizes, and positions. Experimental data show that our algorithm can achieve a good balance in maintaining a treemap's stability, continuity, readability, and average aspect ratio. A software tool is created to compare treemaps and generate the visualizations. User studies show that the users can better understand the changes in the hierarchy and layout, and more quickly notice the color and size differences using our method.
Ying Tu, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2007 Interactive Level-of-Detail Selection Using Image-Based Quality Metric for Large Volume Visualization
abstract
For large volume visualization, an image-based quality metric is difficult to incorporate for level-of-detail selection and rendering without sacrificing the interactivity. This is because it is usually time-consuming to update view-dependent information as well as to adjust to transfer function changes. In this paper, we introduce an image-based level-of-detail selection algorithm for interactive visualization of large volumetric data. The design of our quality metric is based on an efficient way to evaluate the contribution of multiresolution data blocks to the final image. To ensure real-time update of the quality metric and interactive level-of-detail decisions, we propose a summary table scheme in response to runtime transfer function changes and a GPU-based solution for visibility estimation. Experimental results on large scientific and medical data sets demonstrate the effectiveness and efficiency of our algorithm.
Chaoli Wang 0001, Antonio Garcia, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.3
2006 Ultra-scale visualization - Workshop on ultra-scale visualization
abstract
The output from the massively parallel scientific simulations is so voluminous and complex that advanced visualization technologies are necessary to interpret the calculated results. Even though visualization technology has progressed significantly in recent years, we are barely capable of visualizing and analyzing terascale data to its full extent, and petascale datasets are on the horizon. This workshop aims at addressing this pressing issue by fostering communication between visualization researchers and practitioners. The workshop attendees will be introduced to the latest and greatest research innovations in large data visualization and also help direct further research direction through an open discussion session.
James P. Ahrens, Hank Childs, John P. Clyne, E. Wes Bethel, Jian Huang 0007, Scott Klasky, Kwan-Liu Ma, Kenneth Moreland, Michael E. Papka, Valerio Pascucci, Han-Wei Shen, Deborah Silver
SC11
2006 Dynamic View Selection for Time-Varying Volumes
abstract
Animation is an effective way to show how time-varying phenomena evolve over time. A key issue of generating a good animation is to select ideal views through which the user can perceive the maximum amount of information from the time-varying dataset. In this paper, we first propose an improved view selection method for static data. The method measures the quality of a static view by analyzing the opacity, color and curvature distributions of the corresponding volume rendering images from the given view. Our view selection metric prefers an even opacity distribution with a larger projection area, a larger area of salient features' colors with an even distribution among the salient features, and more perceived curvatures. We use this static view selection method and a dynamic programming approach to select time-varying views. The time-varying view selection maximizes the information perceived from the time-varying dataset based on the constraints that the time-varying view should show smooth changes of direction and near-constant speed. We also introduce a method that allows the user to generate a smooth transition between any two views in a given time step, with the perceived information maximized as well. By combining the static and dynamic view selection methods, the users are able to generate a time-varying view that shows the maximum amount of information from a time-varying data set.
Guangfeng Ji, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2006 LOD Map - A Visual Interface for Navigating Multiresolution Volume Visualization
abstract
In multiresolution volume visualization, a visual representation of level-of-detail (LOD) quality is important for us to examine, compare, and validate different LOD selection algorithms. While traditional methods rely on ultimate images for quality measurement, we introduce the LOD map--an alternative representation of LOD quality and a visual interface for navigating multiresolution data exploration. Our measure for LOD quality is based on the formulation of entropy from information theory. The measure takes into account the distortion and contribution of multiresolution data blocks. A LOD map is generated through the mapping of key LOD ingredients to a treemap representation. The ordered treemap layout is used for relative stable update of the LOD map when the view or LOD changes. This visual interface not only indicates the quality of LODs in an intuitive way, but also provides immediate suggestions for possible LOD improvement through visually-striking features. It also allows us to compare different views and perform rendering budget control. A set of interactive techniques is proposed to make the LOD adjustment a simple and easy task. We demonstrate the effectiveness and efficiency of our approach on large scientific and medical data sets.
Chaoli Wang 0001, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2006 Multi-variate, Time Varying, and Comparative Visualization with Contextual Cues
abstract
Time-varying, multi-variate, and comparative data sets are not easily visualized due to the amount of data that is presented to the user at once. By combining several volumes together with different operators into one visualized volume, the user is able to compare values from different data sets in space over time, run, or field without having to mentally switch between different renderings of individual data sets. In this paper, we propose using a volume shader where the user is given the ability to easily select and operate on many data volumes to create comparison relationships. The user specifies an expression with set and numerical operations and her data to see relationships between data fields. Furthermore, we render the contextual information of the volume shader by converting it to a volume tree. We visualize the different levels and nodes of the volume tree so that the user can see the results of suboperations. This gives the user a deeper understanding of the final visualization, by seeing how the parts of the whole are operationally constructed.
Jonathan Woodring, Han-Wei Shen
IEEE Trans. Vis. Comput. Graph.2
2005 Hierarchical Navigation Interface: Leveraging Multiple Coordinated Views for Level-of-Detail Multiresolution Volume Rendering of Large Scientific Data Sets
abstract
We present a new hierarchical navigation interface for level-of-detail selection and rendering of multiresolution volumetric data. The interface consists of multiple coordinated views based on concepts from information visualization as well as scientific visualization literature. With key features such as brushing and linking, and focus and context, it gives the users full control over the level-of-detail selection when navigating through large multiresolution data hierarchies. The navigation interface can also be integrated with traditional level-of-detail selection methods for more effective visual data exploration. We test the utility and effectiveness of this hierarchical navigation interface on a couple of large-scale three-dimensional steady and time-varying data sets.
Chaoli Wang 0001, Han-Wei Shen
IV2
2005 View Selection for Volume Rendering
abstract
In a visualization of a three-dimensional dataset, the insights gained are dependent on what is occluded and what is not. Suggestion of interesting viewpoints can improve both the speed and efficiency of data understanding. This paper presents a view selection method designed for volume rendering. It can be used to find informative views for a given scene, or to find a minimal set of representative views which capture the entire scene. It becomes particularly useful when the visualization process is non-interactive - for example, when visualizing large datasets or time-varying sequences. We introduce a viewpoint "goodness" measure based on the formulation of entropy from information theory. The measure takes into account the transfer function, the data distribution and the visibility of the voxels. Combined with viewpoint properties like view-likelihood and view-stability, this technique can be used as a guide, which suggests "interesting" viewpoints for further exploration. Domain knowledge is incorporated into the algorithm via an importance transfer function or volume. This allows users to obtain view selection behaviors tailored to their specific situations. We generate a view space partitioning, and select one representative view for each partition. Together, this set of views encapsulates the "interesting" and distinct views of the data. Viewpoints in this set can be used as starting points for interactive exploration of the data, thus reducing the human effort in visualization. In non-interactive situations, such a set can be used as a representative visualization of the dataset from all directions.
Udeepta Bordoloi, Han-Wei Shen
IEEE Visualization2
2005 A parallel multiresolution volume rendering algorithm for large data visualization
Jinzhu Gao, Chaoli Wang 0001, Liya Li, Han-Wei Shen
Parallel Comput.4
2005 Parallel graphics and visualization
Bruno Raffin, Han-Wei Shen, Dirk Bartz
Parallel Comput.2
2005 GPU-based 3D wavelet reconstruction with tileboarding
Antonio Garcia, Han-Wei Shen
Vis. Comput.2
2004 Parallel Multiresolution Volume Rendering of Large Data Sets with Error-Guided Load Balancing
Chaoli Wang 0001, Jinzhu Gao, Han-Wei Shen
EGPGV3
2004 Visibility Culling for Time-Varying Volume Rendering Using Temporal Occlusion Coherence
abstract
Typically there is a high coherence in data values between neighboring time steps in an iterative scientific software simulation; this characteristic similarly contributes to a corresponding coherence in the visibility of volume blocks when these consecutive time steps are rendered. Yet traditional visibility culling algorithms were mainly designed for static data, without consideration of such potential temporal coherency. We explore the use of temporal occlusion coherence (TOC) to accelerate visibility culling for time-varying volume rendering. In our algorithm, the opacity of volume blocks is encoded by means of plenoptic opacity functions (POFs). A coherence-based block fusion technique is employed to coalesce time-coherent data blocks over a span of time steps into a single, representative block. Then POFs need only be computed for these representative blocks. To quickly determine the subvolumes that do not require updates in their visibility status for each subsequent time step, a hierarchical "TOC tree" data structure is constructed to store the spans of coherent time steps. To achieve maximal culling potential, while remaining conservative, we have extended our previous POP into an optimized POP (OPOP) encoding scheme for this specific scenario. To test our general TOC and OPOF approach, we have designed a parallel time-varying volume rendering algorithm accelerated by visibility culling. Results from experimental runs on a 32-processor cluster confirm both the effectiveness and scalability of our approach.
Jinzhu Gao, Han-Wei Shen, Jian Huang 0007, James Arthur Kohl
IEEE Visualization2
2004 Interactive Visualization of Three-Dimensional Vector Fields with Flexible Appearance Control
abstract
In this paper, we present an interactive texture-based algorithm for visualizing three-dimensional steady and unsteady vector fields. The goal of the algorithm is to provide a general volume rendering framework allowing the user to compute three-dimensional flow textures interactively and to modify the appearance of the visualization on the fly. To achieve our goal, we decouple the visualization pipeline into two disjoint stages. First, flow lines are generated from the 3D vector data. Various geometric properties of the flow paths are extracted and converted into a volumetric form using a hardware-assisted slice sweeping algorithm. In the second phase of the algorithm, the attributes stored in the volume are used as texture coordinates to look up an appearance texture to generate both informative and aesthetic representations of the vector field. Our algorithm allows the user to interactively navigate through different regions of interest in the underlying field and experiment with various appearance textures. With our algorithm, visualizations with enhanced structural perception using various visual cues can be rendered in real time. A myriad of existing geometry-based and texture-based visualization techniques can also be emulated.
Han-Wei Shen, Guo-Shi Li, Udeepta Bordoloi
IEEE Trans. Vis. Comput. Graph.1
2003 QoS-Aware Middleware for Cluster-Based Servers to support Interactive and Resource-Adaptive Applications
abstract
Advances in commodity processor and network technologies have made cluster-based servers very attractive for supporting a large number of interactive applications (such as visualization and data mining) in the domains of Grid computing and distributed computing. These applications involve accesses to huge amounts of data within the servers and heavy computations on the accessed data before sending out the results to the clients. The interactive nature of these applications requires some kind of QoS support (such as guarantees on response time) from the underlying server. Unfortunately, the current generation cluster-based servers with the popular interconnect (Gigabit Ethernet, Myrinet, or Quadrics) do not provide any kinds of QoS support. Fortunately, many of these applications are resource-adaptive, i.e., application parameters can be changed to suit user demands and available system resources. To solve these problems, a new QoS-aware middleware layer is proposed in this paper for cluster-based servers with Myrinet interconnect. The middleware is built on top of a simple NIC-based rate control scheme that provides proportional bandwidth allocation. Three major components of the middleware (profiler, QoS translator, and resource allocator), their functionalities, designs, and the associated algorithms are presented. These components work together to execute a requested job in a predictable manner with an efficient allocation of system resources while exploiting the resource-adaptive property of the application. The complete middleware is designed, developed, and implemented on a Myrinet cluster. It is evaluated for two visualization applications: polygon rendering and ray-tracing. Experimental evaluations demonstrate that the proposed QoS framework enables multiple interactive and resource-adaptive applications to be executed in a predictable manner while keeping the allocation of system resources efficient. It is shown that the QoS-aware middleware helps applications to obtain response times within 7% of the expected times, compared to increases of up to 117% in the absence of any QoS support.
S. Senapathi, B. Chandrasekaran 0001, Don Stredney, Han-Wei Shen, Dhabaleswar K. Panda 0001
HPDC4
2003 Space Efficient Fast Isosurface Extraction for Large Datasets
abstract
In this paper, we present a space efficient algorithm for speeding up isosurface extraction. Even though there exist algorithms that can achieve optimal search performance to identify isosurface cells, they prove impractical for large datasets due to a high storage overhead. With the dual goals of achieving fast isosurface extraction and simultaneously reducing the space requirement, we introduce an algorithm based on transform coding to compress the interval information of the cells in a dataset. Compression is achieved by first transforming the cell intervals (minima, maxima) into a form which allows more efficient compaction. It is followed by a dataset optimized non-uniform quantization stage. The compressed data is stored in a data structure that allows fast searches in the compression domain, thus eliminating the need to retrieve the original representation of the intervals at run-time. The space requirement of our search data structure is the mandatory cost of storing every cell ID once, plus an overhead for quantization information. The overhead is typically in the order of a few hundredths of the dataset size.
Udeepta Bordoloi, Han-Wei Shen
IEEE Visualization2
2003 Visibility Culling Using Plenoptic Opacity Functions for Large Data Visualization
abstract
Visibility culling has the potential to accelerate large data visualization in significant ways. Unfortunately, existing algorithms do not scale well when parallelized, and require full re-computation whenever the opacity transfer function is modified. To address these issues, we have designed a Plenoptic Opacity Function (POF) scheme to encode the view-dependent opacity of a volume block. POFs are computed off-line during a pre-processing stage, only once for each block. We show that using POFs is (i) an efficient, conservative and effective way to encode the opacity variations of a volume block for a range of views, (ii) flexible for re-use by a family of opacity transfer functions without the need for additional off-line processing, and (iii) highly scalable for use in massively parallel implementations. Our results confirm the efficacy of POFs for visibility culling in large-scale parallel volume rendering; we can interactively render the Visible Woman dataset using software ray-casting on 32 processors, with interactive modification of the opacity transfer function on-the-fly.
Jinzhu Gao, Jian Huang 0007, Han-Wei Shen, James Arthur Kohl
IEEE Visualization3
2003 Volume Tracking Using Higher Dimensional Isocontouring
abstract
Tracking and visualizing local features from a time-varying volumetric data allows the user to focus on selected regions of interest, both in space and time, which can lead to a better understanding of the underlying dynamics. In this paper, we present an efficient algorithm to track time-varying isosurfaces and interval volumes using isosurfacing in higher dimensions. Instead of extracting the data features such as isosurfaces or interval volumes separately from multiple time steps and computing the spatial correspondence between those features, our algorithm extracts the correspondence directly from the higher dimensional geometry and thus can more efficiently follow the user selected local features in time. In addition, by analyzing the resulting higher dimensional geometry, it becomes easier to detect important topological events and the corresponding critical time steps for the selected features. With our algorithm, the user can interact with the underlying time-varying data more easily. The computation cost for performing time-varying volume tracking is also minimized.
Guangfeng Ji, Han-Wei Shen, Rephael Wenger
IEEE Visualization2
2003 Chameleon: An interactive texture-based rendering framework for visualizing three-dimensional vector fields
abstract
In this paper we present an interactive texture-based technique for visualizing three-dimensional vector fields. The goal of the algorithm is to provide a general volume rendering framework allowing the user to compute three-dimensional flow textures interactively, and to modify the appearance of the visualization on the fly. To achieve our goal, we decouple the visualization pipeline into two disjoint stages. First, streamlines are generated from the 3D vector data. Various geometric properties of the streamlines are extracted and converted into a volumetric form using a hardware-assisted slice sweeping algorithm. In the second phase of the algorithm, the attributes stored in the volume are used as texture coordinates to look up an appearance texture to generate both informative and aesthetic representations of the underlying vector field. Users can change the input textures and instantaneously visualize the rendering results. With our algorithm, visualizations with enhanced structural perception using various visual cues can be rendered in real time. A myriad of existing geometry-based and texture-based visualization techniques can also be emulated.
Guo-Shi Li, Udeepta Bordoloi, Han-Wei Shen
IEEE Visualization3
2003 High Dimensional Direct Rendering of Time-Varying Volumetric Data
abstract
We present an alternative method for viewing time-varying volumetric data. We consider such data as a four-dimensional data field, rather than considering space and time as separate entities. If we treat the data in this manner, we can apply high dimensional slicing and projection techniques to generate an image hyperplane. The user is provided with an intuitive user interface to specify arbitrary hyperplanes in 4D, which can be displayed with standard volume rendering techniques. From the volume specification, we are able to extract arbitrary hyperslices, combine slices together into a hyperprojection volume, or apply a 4D raycasting method to generate the same results. In combination with appropriate integration operators and transfer functions, we are able to extract and present different space-time features to the user.
Jonathan Woodring, Chaoli Wang 0001, Han-Wei Shen
IEEE Visualization3
2002 Hardware Accelerated Interactive Vector Field Visualization: A level of detail approach
abstract
This paper presents an interactive global visualization technique for dense vector fields using levels of detail. We introduce a novel scheme which combines an error-controlled hierarchical approach and hardware acceleration to produce high resolution visualizations at interactive rates. Users can control the trade-off between computation time and image quality, producing visualizations amenable for situations ranging from high frame-rate previewing to accurate analysis. Use of hardware texture mapping allows the user to interactively zoom in and explore the data, and also to configure various texture parameters to change the look and feel of the visualization. We are able to achieve sub-second rates for dense LIC-like visualizations with resolutions in the order of a million pixels for data of similar dimensions. Categories and Subject Descriptors (according to ACM CCS): I.3 [Computer Graphics]: Applications
Udeepta Bordoloi, Han-Wei Shen
Comput. Graph. Forum2
1999 A Fast Volume Rendering Algorithm for Time-Varying Fields Using a Time-Space Partitioning (TSP) Tree
abstract
We present a fast volume rendering algorithm for time-varying fields. We propose a new data structure, called time-space partitioning (TSP) tree, that can effectively capture both the spatial and the temporal coherence from a time-varying field. Using the proposed data structure, the rendering speed is substantially improved. In addition, our data structure helps to maintain the memory access locality and to provide the sparse data traversal so that our algorithm becomes suitable for large-scale out-of-core applications. Finally, our algorithm allows flexible error control for both the temporal and the spatial coherence so that a trade-off between image quality and rendering speed is possible. We demonstrate the utility and speed of our algorithm with data from several time-varying CFD simulations. Our rendering algorithm can achieve substantial speedup while the storage space overhead for the TSP tree is kept at a minimum.
Han-Wei Shen, Ling-Jan Chiang, Kwan-Liu Ma
IEEE Visualization1
1998 Isosurface extraction in time-varying fields using a temporal hierarchical index tree
abstract
Many high-performance isosurface extraction algorithms have been proposed in the past several years as a result of intensive research efforts. When applying these algorithms to large-scale time-varying fields, the storage overhead incurred from storing the search index often becomes overwhelming. This paper proposes an algorithm for locating isosurface cells in time-varying fields. We devise a new data structure, called the temporal hierarchical index tree, which utilizes the temporal coherence that exists in a time-varying field and adaptively coalesces the cells' extreme values over time; the resulting extreme values are then used to create the isosurface cell search index. For a typical time-varying scalar data set, not only does this temporal hierarchical index tree require much less storage space, but also the amount of I/O required to access the indices from the disk at different time steps is substantially reduced. We illustrate the utility and speed of our algorithm with data from several large-scale time-varying CFD simulations. Our algorithm can achieve more than 80% of disk-space savings when compared with the existing techniques, while the isosurface extraction time is nearly optimal.
Han-Wei Shen
IEEE Visualization1
1998 A New Line Integral Convolution Algorithm for Visualizing Time-Varying Flow Fields
abstract
New challenges on vector field visualization emerge as time dependent numerical simulations become ubiquitous in the field of computational fluid dynamics (CFD). To visualize data generated from these simulations, traditional techniques, such as displaying particle traces, can only reveal flow phenomena in preselected local regions and thus, are unable to track the evolution of global flow features over time. The paper presents an algorithm, called UFLIC (Unsteady Flow LIC), to visualize vector data in unsteady flow fields. Our algorithm extends a texture synthesis technique, called Line Integral Convolution (LIC), by devising a new convolution algorithm that uses a time-accurate value scattering scheme to model the texture advection. In addition, our algorithm maintains the coherence of the flow animation by successively updating the convolution results over time. Furthermore, we propose a parallel UFLIC algorithm that can achieve high load balancing for multiprocessor computers with shared memory architecture. We demonstrate the effectiveness of our new algorithm by presenting image snapshots from several CFD case studies.
Han-Wei Shen, David L. Kao
IEEE Trans. Vis. Comput. Graph.1
1997 UFLIC: a line integral convolution algorithm for visualizing unsteady flows
abstract
The paper presents an algorithm, UFLIC (Unsteady Flow LIC), to visualize vector data in unsteady flow fields. Using line integral convolution (LIC) as the underlying method, a new convolution algorithm is proposed that can effectively trace the flow's global features over time. The new algorithm consists of a time-accurate value depositing scheme and a successive feedforward method. The value depositing scheme accurately models the flow advection, and the successive feedforward method maintains the coherence between animation frames. The new algorithm can produce time-accurate, highly coherent flow animations to highlight global features in unsteady flow fields. CFD scientists, for the first time, are able to visualize unsteady surface flows using the algorithm.
Han-Wei Shen, David L. Kao
IEEE Visualization1
1996 Isosurfacing in Span Space with Utmost Efficiency (ISSUE)
abstract
We present efficient sequential and parallel algorithms for isosurface extraction. Based on the Span Space data representation, new data subdivision and searching methods are described. We also present a parallel implementation with an emphasis on load balancing. The performance of our sequential algorithm to locate the cell elements intersected by isosurfaces is faster than the Kd tree searching method originally used for the Span Space algorithm. The parallel algorithm can achieve high load balancing for massively parallel machines with distributed memory architectures.
Han-Wei Shen, Charles D. Hansen, Yarden Livnat, Chris R. Johnson 0001
IEEE Visualization1
1996 A Near Optimal Isosurface Extraction Algorithm Using the Span Space
abstract
Presents the "Near Optimal IsoSurface Extraction" (NOISE) algorithm for rapidly extracting isosurfaces from structured and unstructured grids. Using the span space, a new representation of the underlying domain, we develop an isosurface extraction algorithm with a worst case complexity of o(/spl radic/n+k) for the search phase, where n is the size of the data set and k is the number of cells intersected by the isosurface. The memory requirement is kept at O(n) while the preprocessing step is O(n log n). We utilize the span space representation as a tool for comparing isosurface extraction methods on structured and unstructured grids. We also present a fast triangulation scheme for generating and displaying unstructured tetrahedral grids.
Yarden Livnat, Han-Wei Shen, Chris R. Johnson 0001
IEEE Trans. Vis. Comput. Graph.2
1996 Correction: A Near Optimal Isosurface Extraction Algorithm Using the Span Space
Yarden Livnat, Han-Wei Shen, Chris R. Johnson 0001
IEEE Trans. Vis. Comput. Graph.2
1995 Sweeping Simplices: A Fast Iso-Surface Extraction Algorithm for Unstructured Grids
abstract
Presents an algorithm that accelerates the extraction of iso-surfaces from unstructured grids by avoiding the traversal of the entire set of cells in the volume. The algorithm consists of a sweep algorithm and a data decomposition scheme. The sweep algorithm incrementally locates intersected elements, and the data decomposition scheme restricts the algorithm's worst-case performance. For data sets consisting of hundreds of thousands of elements, our algorithm can reduce the cell traversal time by more than 90% over the naive iso-surface extraction algorithm, thus facilitating interactive probing of scalar fields for large-scale problems on unstructured three-dimensional grids.
Han-Wei Shen, Chris R. Johnson 0001
IEEE Visualization1
1994 Differential Volume Rendering: A Fast Volume Visualization Technique for Flow Animation
abstract
We present a direct volume rendering algorithm to speed up volume animation for flow visualizations. Data coherency between consecutive simulation time steps is used to avoid casting rays from those pixels retaining color values assigned to the previous image. The algorithm calculates the differential information among a sequence of 3D volumetric simulation data. At each time step the differential information is used to compute the locations of pixels that need updating and a ray-casting method as utilized to produce the updated image. We illustrate the utility and speed of the differential volume rendering algorithm with simulation data from computational bioelectric and fluid dynamics applications. We can achieve considerable disk-space savings and nearly real-time rendering of 3D flows using low-cost, single processor workstations for models which contain hundreds of thousands of data points.>
Han-Wei Shen, Chris R. Johnson 0001
IEEE Visualization1