Daniel A. Keim

dblp:k/DanielAKeim · DBLP profile ↗
← Back
212ranked-venue papers
34as first author
49since 2021 · last 2026
0000-0001-7966-9740ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 123 · 12 first-author · 37 since 2021Databases, data management, data science and information retrieval · 50 · 12 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 37 · 10 first-author · 4 since 2021Artificial intelligence and machine learning · 22 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 2 since 2021Security and privacy · 5 · 1 since 2021Theory of computation · 3 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Evaluating Keyframe Layouts for Visual Known-Item Search in Homogeneous Collections
abstract
Multimodal deep-learning models power interactive video retrieval by ranking keyframes in response to textual queries. Despite these advances, users must still browse ranked candidates manually to locate a target. Keyframe arrangement within the search grid highly affects browsing effectiveness and user efficiency, yet remains underexplored. We report a study with 49 participants evaluating seven keyframe layouts for the Visual Known-Item Search task. Beyond efficiency and accuracy, we relate browsing phenomena, such as overlooks, to layout characteristics. Our results show that a video-grouped layout is the most efficient, while a four-column, rank-preserving grid achieves the highest accuracy. Sorted grids reveal potential and trade-offs, enabling rapid scanning of uninteresting regions but down-ranking relevant targets to less prominent positions, delaying first arrival times and increasing overlooks. These findings motivate hybrid designs that preserve positions of top-ranked items while sorting or grouping the remainder, and offer guidance for searching in grids beyond video retrieval.
Bastian Jäckl, Jirí Kruchina, Lucas Joos, Daniel A. Keim, Ladislav Peska, Jakub Lokoc
ICMR4
2026 What Drove Success at the 15th Video Browser Showdown? A Comprehensive Interaction-Logging Analysis
abstract
In 2026, the Multimedia Modeling conference in Prague hosted the fifteenth edition of the Video Browser Showdown (VBS) competition. Yet, for the first time, two participating systems implemented full-scale interaction logging frameworks, enabling a detailed analysis of the search process beyond traditional score-based evaluation. In this paper, both systems are introduced and described with a focus on their user interactions. To enable compact presentation and analysis of logs, all interaction types are further grouped into more abstract events, forming an interaction taxonomy hierarchy. Finally, we reveal the applied search strategies and analyze which factors drive success and failure. The results reveal a clear dominance of iterative, high-frequency text query reformulation with result set inspection, leveraging the power of modern CLIP-based models across most competition categories. In around 10% of cases, users also relied on advanced system features to achieve good performance at VBS, mostly on challenging homogeneous datasets.
Bastian Jäckl, Omar Shahbaz Khan, Benjamin Verner, Zuzana Vopálková, Udo Schlegel, Daniel A. Keim, Jakub Lokoc
ICMR6
2026 PraK V4 at the Video Browser Showdown 2026
Bastian Jäckl, Benjamin Verner, Michael Stroh, Vojtech Kloda, Ladislav Nagy, Oliver Deussen, Daniel A. Keim, Jakub Lokoc
MMM (4)7
2026 Leveraging LLMs for semi-automatic corpus filtration in systematic literature reviews
abstract
The creation of systematic literature reviews (SLR) is critical for analyzing the landscape of a research field and guiding future research directions. However, retrieving and filtering the literature corpus for an SLR is highly time-consuming and requires extensive manual effort, as keyword-based searches in digital libraries often return numerous irrelevant publications. In this work, we propose a pipeline leveraging multiple large language models (LLMs), classifying papers based on descriptive prompts and deciding jointly using a consensus scheme. The entire process is human-supervised and interactively controlled via our open-source visual analytics web interface, LLMSurver, which enables real-time inspection and modification of model outputs. We evaluate our approach using ground-truth data from a recent SLR comprising 8323 candidate papers, benchmarking both open and commercial state-of-the-art LLMs from mid-2024 and fall 2025. Results demonstrate that our pipeline significantly reduces manual effort while achieving lower error rates than single human annotators. Furthermore, modern open-source models prove sufficient for this task, making the method accessible and cost-effective. Overall, our work demonstrates how responsible human-AI collaboration can accelerate and enhance systematic literature reviews within academic workflows. • A new pipeline for semi-automated literature review. • Large language models vote using a consensus scheme. • User keeps the control through a visual analytics approach. • Evaluation with previous and modern state-of-the-art models (open and commercial). • Available online demo tool and open-source code.
Lucas Joos, Daniel A. Keim, Maximilian T. Fischer
Comput. Graph.2
2026 A user study on localized sub-region search in homogeneous marine collections
abstract
Searching in highly homogeneous video domains is challenging, especially when relying solely on human memory. This difficulty arises because homogeneous content often requires domain-specific vocabulary to describe effectively, and a single text label might fit a substantial subset of the database. This paper investigates supplementing text queries with spatial location information to better address specific search intents—a strategy applicable when users possess strong visual memory of a target object, or have external knowledge of its position. To process this location information effectively, we formally define and evaluate static grid and dynamic segmentation strategies for video frame partitioning. Furthermore, we present a large-scale cognitive user study involving 220 participants, designed to simulate realistic memory constraints of Known-Item Search (KIS) tasks commonly found in interactive retrieval benchmarks. The study utilizes standard working memory interference techniques, comprising a target exposure phase, a distraction period, and subsequent query specification from memory. These queries are then evaluated against the proposed spatial retrieval models to determine how human spatial memory decay impacts retrieval effectiveness. Our results reveal that while dynamic segmentation models achieve the highest theoretical retrieval bounds under perfect conditions, their strict geometric boundaries are severely affected by human memory decay. Consequently, static overlapping grids demonstrate high robustness, outperforming dynamic models under more realistic, memory-driven search constraints. Finally, our experiments show that the observed spatial annotation perturbations can be modeled using a four-dimensional Kernel Density Estimation (KDE) method, enabling the simulation of realistic human memory decay on ideal bounding boxes.
Vojtech Kloda, Bastian Jäckl, Daniel A. Keim, Jakub Lokoc
Inf. Syst.3
2026 Motif Simplification for BioFabric Network Visualizations: Improving Pattern Recognition and Interpretation
abstract
Detecting and interpreting common patterns in relational data is crucial for understanding complex topological structures across various domains. These patterns, or network motifs, can often be detected algorithmically. However, visual inspection remains vital for exploring and discovering patterns. This paper focuses on presenting motifs within BioFabric network visualizations-a unique technique that opens opportunities for research on scaling to larger networks, design variations, and layout algorithms to better expose motifs. Our goal is to show how highlighting motifs can assist users in identifying and interpreting patterns in BioFabric visualizations. To this end, we leverage existing motif simplification techniques. We replace edges with glyphs representing fundamental motifs such as staircases, cliques, paths, and connector nodes. The results of our controlled experiment and usage scenarios demonstrate that motif simplification for BioFabric is useful for detecting and interpreting network patterns. Our participants were faster and more confident using the simplified view without sacrificing accuracy. The efficacy of our current motif simplification approach depends on which extant layout algorithm is used. We hope our promising findings on user performance will motivate future research on layout algorithms tailored to maximizing motif presentation. Our supplemental material is available at https://osf.io/f8s3g/?view_only=7e2df9109dfd4e6c85b89ed828320843.
Johannes Fuchs 0001, Cody Dunne, Maria-Viktoria Heinle, Daniel A. Keim, Sara Di Bartolomeo
IEEE Trans. Vis. Comput. Graph.4
2025 EnMRgy: Energy Network Analysis in Mixed Reality (Poster Abstract)
abstract
The shifting and ever-growing demand for energy, for instance, driven by transformations towards new technologies such as electric vehicles, heat pumps, battery storage, or rooftop solar, requires urban infrastructure to adapt. Upgrading legacy infrastructure, such as undersized electric cables, is costly, time-consuming, and disruptive, and therefore requires a holistic perspective and thorough urban planning that considers multi energy systems and co-located utilities. We present EnMRgy, a mixed-reality decision-support system that enables experts and decision-makers to explore a city’s energy distribution networks, together with demand simulations and scenarios for infrastructure development. Within an immersive 3D city context, an energy network such as a power grid, modelled as a weighted graph, is visualised. Interactive functionalities allow users to adjust visual representations and compare scenarios across three different views. Our work enables evidence-based strategic planning for future-ready energy networks.
Lucas Joos, Maximilian T. Fischer, Alexander Frings, Daniel A. Keim
GD4
2025 Show Me Your Best Side: Characteristics of User-Preferred Perspectives for 3D Graph Drawings
Lucas Joos, Gavin J. Mooney, Maximilian T. Fischer, Daniel A. Keim, Falk Schreiber, Helen C. Purchase, Karsten Klein 0001
GD4
2025 ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning
abstract
Binary choices, as often used for reinforcement learning from human feedback (RLHF), convey only the *direction* of a preference. A person may choose apples over oranges and bananas over grapes, but *which preference is stronger*? Strength is crucial for decision-making under uncertainty and generalization of preference models, but hard to measure reliably. Metadata such as response times and inter-annotator agreement can serve as proxies for strength, but are often noisy and confounded. We propose ResponseRank to address the challenge of learning from noisy strength signals. Our method uses relative differences in proxy signals to *rank responses to pairwise comparisons by their inferred preference strength*. To control for systemic variation, we compare signals only locally within carefully constructed strata. This enables robust learning of utility differences consistent with strength-derived rankings while making minimal assumptions about the strength signal. Our contributions are threefold: (1) ResponseRank, a novel method that robustly learns preference strength by leveraging locally valid relative strength signals; (2) empirical evidence of improved sample efficiency and robustness across diverse tasks: synthetic preference learning (with simulated response times), language modeling (with annotator agreement), and RL control tasks (with simulated episode returns); and (3) the *Pearson Distance Correlation (PDC)*, a novel metric that isolates cardinal utility learning from ordinal accuracy.
Timo Kaufmann, Yannick Metz, Daniel A. Keim, Eyke Hüllermeier
NeurIPS3
2025 Dynamic Sub-region Search In Homogeneous Collections Using CLIP
Bastian Jäckl, Vojtech Kloda, Daniel A. Keim, Jakub Lokoc
SISAP3
2025 Experimental Evaluation Of Static Image Sub-Region-Based Search Models Using CLIP
Bastian Jäckl, Vojtech Kloda, Daniel A. Keim, Jakub Lokoc
SISAP3
2025 MultiInv: Inverting multidimensional scaling projections and computing decision maps by multilateration
abstract
Inverse projections enable a variety of tasks such as the exploration of classifier decision boundaries, creating counterfactual explanations, and generating synthetic data. Yet, many existing inverse projection methods are difficult to implement, challenging to predict, and sensitive to parameter settings. To address these, we propose to invert distance-preserving projections like Multidimensional Scaling (MDS) projections by using multilateration – a method used for geopositioning. Our approach finds data values for locations where no data point is projected under the key assumption that a given projection technique preserves pairwise distances among data samples in the low-dimensional space. Being based on a geometrical relationship, our technique is more interpretable than comparable machine learning-based approaches and can invert 2-dimensional projections up to D − 1 dimensional spaces if given at least D data points. We compare several strategies for multilateration point selection, show the application of our technique on three additional projection techniques apart from MDS, and use established quality metrics to evaluate its accuracy in comparison to existing inverse projections. We also show its application to computing decision maps for exploring the behavior of trained classification models. When the projection to invert captures data distances well, our inverse performs similarly to existing approaches while being interpretable and considerably simpler to compute.
Daniela Blumberg, Yu Wang 0188, Alexandru C. Telea, Daniel A. Keim, Frederik L. Dennig
Comput. Graph.4
2025 Visually Assessing 1-D Orderings of Contiguous Spatial Polygons
abstract
Abstract One‐dimensional orderings of spatial entities have been researched in many contexts, e.g. spatial indexing structures or visualizations for spatiotemporal trend analysis. While plenty of studies have been conducted to evaluate orderings of point‐based data, polygonal shapes, despite their different topological properties, have received less attention. Existing measures to quantify errors in projections or orderings suffer from generic neighborhood definitions and over‐simplification of distances when applied to polygonal data. In this work, we address these shortcomings by introducing measures that adapt to a varying neighborhood size depending on the number of contiguous neighbors and thus, address the limitations of existing measures for polygonal shapes. To guide experts in determining a suitable ordering, we propose a user‐steerable visual analytics prototype capable of locally and globally inspecting ordering errors, investigating the impact of geographic obstacles, and comparing ordering strategies using our measures. We demonstrate the effectiveness of our approach through a use case and conducted an expert study with 8 data scientists as a qualitative evaluation of our approach. Our results show that users are capable of identifying ordering errors, comparing ordering strategies on a global and local scale, as well as assessing the impact of semantically relevant geographic obstacles.
Julius Rauscher, Frederik L. Dennig, Udo Schlegel, Daniel A. Keim, Johannes Fuchs 0001
Comput. Graph. Forum4
2025 An Analysis of the Interplay and Mutual Benefits of Grounded Theory and Visualization
abstract
Grounded theory (GT) is a research methodology that entails a systematic workflow for theory generation grounded on emergent data. In this article, we juxtapose GT workflows with typical workflows in visualization and visual analytics (VIS), unveiling the characteristics shared by these workflows. We explore the research landscape of VIS to study where GT is applied to generate VIS theories, explicitly as well as implicitly. We discuss "why" GT can potentially play a significant role in VIS. We outline a "how" methodology for conducting GT research in VIS, which addresses the need for theoretical advancement in VIS while benefiting from other methods and techniques in VIS. We illustrate this "how" methodology with a use case of adopting GT approaches in studying visualization guidelines.
Alexandra Diehl, Alfie Abdul-Rahman, Benjamin Bach, Mennatallah El-Assady, Matthias Kraus 0002, Robert S. Laramee, Daniel A. Keim, Min Chen 0001
IEEE Trans. Vis. Comput. Graph.7
2025 Quality Metrics and Reordering Strategies for Revealing Patterns in BioFabric Visualizations
abstract
Visualizing relational data is crucial for understanding complex connections between entities in social networks, political affiliations, or biological interactions. Well-known representations like node-link diagrams and adjacency matrices offer valuable insights, but their effectiveness relies on the ability to identify patterns in the underlying topological structure. Reordering strategies and layout algorithms play a vital role in the visualization process since the arrangement of nodes, edges, or cells influences the visibility of these patterns. The BioFabric visualization combines elements of node-link diagrams and adjacency matrices, leveraging the strengths of both, the visual clarity of node-link diagrams and the tabular organization of adjacency matrices. A unique characteristic of BioFabric is the possibility to reorder nodes and edges separately. This raises the question of which combination of layout algorithms best reveals certain patterns. In this paper, we discuss patterns and anti-patterns in BioFabric, such as staircases or escalators, relate them to already established patterns, and propose metrics to evaluate their quality. Based on these quality metrics, we compared combinations of well-established reordering techniques applied to BioFabric with a well-known benchmark data set. Our experiments indicate that the edge order has a stronger influence on revealing patterns than the node layout. The results show that the best combination for revealing staircases is a barycentric node layout, together with an edge order based on node indices and length. Our research contributes a first building block for many promising future research directions, which we also share and discuss. A free copy of this paper and all supplemental materials are available at https://osf.io/9mt8r/?view_only=b7t0dfbe550e3404f83059afdc60184c6.
Johannes Fuchs 0001, Alexander Frings, Maria-Viktoria Heinle, Daniel A. Keim, Sara Di Bartolomeo
IEEE Trans. Vis. Comput. Graph.4
2025 TreEducation: A Visual Education Platform for Teaching Treemap Layout Algorithms
abstract
Treemaps are a powerful tool for representing hierarchical data in a space-efficient manner and are used in various domains, including network security or software development. However, interpreting the topology encoded by nested rectangles can be challenging, particularly compared to tree-structured representations like node-link diagrams or icicle plots. To address this challenge, we introduce TreEducation, a visual education platform designed to improve the visualization literacy skills required for reading treemaps among non-expert users. TreEducation is an online application that combines visualizations, interactions, and gamification elements to facilitate understanding of eight different treemap layout algorithms and enhance students' learning process. We evaluated TreEducation in a classroom setting and a controlled environment. Our results indicate a significant knowledge gain of students training exclusively with TreEducation and the usefulness of competition as a social gamification element included in our competitive quiz.
Johannes Fuchs 0001, Bastian Jäckl, Michael Jüttler, Daniel A. Keim, Rita Sevastjanova
IEEE Trans. Vis. Comput. Graph.4
2024 Known-Item Search in Video: An Eye Tracking-Based Study
abstract
Deep learning has revolutionized multimedia retrieval, yet effectively searching within large video collections remains a complex challenge. This paper focuses on the design and evaluation of known-item search systems, leveraging the strengths of CLIP-based deep neural networks for ranking. At events like the Video Browser Showdown, these models have shown promise in effectively ranking the video frames. While ranking models can be pre-selected automatically based on a benchmark collection, the selection of an optimal browsing interface, crucial for refining top-ranked items, is complex and heavily influenced by user behavior. Our study addresses this by presenting an eye tracking-based analysis of user interaction with different image grid layouts. This approach offers novel insights into search patterns and user preferences, particularly examining the trade-off between displaying fewer but larger images versus more but smaller images. Our findings reveal a preference for grids with fewer images and detail how image similarity and grid position affect user search behavior. These results not only enhance our understanding of effective video retrieval interface design but also set the stage for future advancements in the field.
Lucas Joos, Bastian Jäckl, Daniel A. Keim, Maximilian T. Fischer, Ladislav Peska, Jakub Lokoc
ICMR3
2024 Cluster-Faithful Graph Visualization: New Metrics and Algorithms
abstract
The cluster faithfulness metrics CQ measure how faithfully the ground truth clustering of a graph is represented as the geometric clustering in a drawing of the graph. Existing CQ metrics use k-means clustering, which effectively compute a geometric clustering when the cluster sizes are even, resulting in accurate CQ metrics. However, k-means clustering tends to compute clusters of even sizes and thus often fails to compute an accurate geometric clustering when the cluster sizes are uneven, leading to inaccurate CQ metrics.In this paper, we present a new cluster faithfulness metric CQ-HAC, using HAC (Hierarchical Agglomerative Clustering). HAC can compute a more accurate geometric clustering for uneven cluster sizes than k-means clustering. Consequently, CQ-HAC can more accurately measure cluster faithfulness, regardless of whether the sizes of clusters are even or uneven. Moreover, we present two algorithms, Cluster-kmeans and Cluster-HAC, for optimizing cluster faithfulness of graph drawings. Extensive experiments show that in practice, both algorithms always compute perfectly cluster-faithful drawings (i.e., CQ = 1) in our experiments using various graphs with both even and uneven cluster sizes, achieving significant improvement over existing graph layouts, including cluster-focused layouts.
Shijun Cai, Seok-Hee Hong 0001, Amyra Meidiana, Peter Eades, Daniel A. Keim
PacificVis5
2024 An Image Quality Dataset with Triplet Comparisons for Multi-dimensional Scaling
abstract
In the early days of perceptual image quality research more than 30 years ago, the multidimensionality of distortions in perceptual space was considered important. However, research focused on scalar quality as measured by mean opinion scores. With our work, we intend to revive interest in this relevant area by presenting a first pilot dataset of annotated triplet comparisons for image quality assessment. It contains one source stimulus together with distorted versions derived from 7 distortion types at 12 levels each. Our crowdsourced and curated dataset contains roughly 50,000 responses to 7,000 triplet comparisons. We show that the multidimensional embedding of the dataset poses a challenge for many established triplet embedding algorithms. Finally, we propose a new reconstruction algorithm, dubbed logistic triplet embedding (LTE) with Tikhonov regularization. It shows promising performance. This study helps researchers to create larger datasets and better embedding techniques for multidimensional image quality. The dataset includes images and ratings and can be accessed at https://github.com/jenadeleh/multidimensionalIQA-dataset/tree/main.
Mohsen Jenadeleh, Frederik L. Dennig, René Cutura, Quynh Quang Ngo, Daniel A. Keim, Michael Sedlmair, Dietmar Saupe
QoMEX5
2024 Exploring the Design Space of BioFabric Visualization for Multivariate Network Analysis
abstract
Abstract The visual analysis of multivariate network data is a common yet difficult task in many domains. The major challenge is to visualize the network's topology and additional attributes for entities and their connections. Although node‐link diagrams and adjacency matrices are widespread, they have inherent limitations. Node‐link diagrams struggle to scale effectively, while adjacency matrices can fail to represent network topologies clearly. In this paper, we delve into the design space of BioFabric, which aligns entities along rows and relationships along columns, providing a way to encapsulate multiple attributes for both. We explore how we can leverage the unique opportunities offered by BioFabric's design space to visualize multivariate network data — focusing on three main categories: juxtaposed visualizations, embedded on‐node and on‐edge encoding, and transformed node and edge encoding. We complement our exploration with a quantitative assessment comparing BioFabric to adjacency matrices. We postulate that the expansive design possibilities introduced in BioFabric network visualization have the potential for the visualization of multivariate data, and we advocate for further evaluation of the associated design space. Our supplemental material is available on osf.io.
Johannes Fuchs 0001, Frederik L. Dennig, Maria-Viktoria Heinle, Daniel A. Keim, Sara Di Bartolomeo
Comput. Graph. Forum4
2024 -generAItor: Tree-in-the-loop Text Generation for Language Model Explainability and Adaptation
abstract
Large language models (LLMs) are widely deployed in various downstream tasks, e.g., auto-completion, aided writing, or chat-based text generation. However, the considered output candidates of the underlying search algorithm are under-explored and under-explained. We tackle this shortcoming by proposing a tree-in-the-loop approach, where a visual representation of the beam search tree is the central component for analyzing, explaining, and adapting the generated outputs. To support these tasks, we present generAItor, a visual analytics technique, augmenting the central beam search tree with various task-specific widgets, providing targeted visualizations and interaction possibilities. Our approach allows interactions on multiple levels and offers an iterative pipeline that encompasses generating, exploring, and comparing output candidates, as well as fine-tuning the model based on adapted data. Our case study shows that our tool generates new insights in gender bias analysis beyond state-of-the-art template-based methods. Additionally, we demonstrate the applicability of our approach in a qualitative user study. Finally, we quantitatively evaluate the adaptability of the model to few samples, as occurring in text-generation use cases.
Thilo Spinner, Rebecca Kehlbeck, Rita Sevastjanova, Tobias Stähle, Daniel A. Keim, Oliver Deussen, Mennatallah El-Assady
ACM Trans. Interact. Intell. Syst.5
2024 FS/DS: A Theoretical Framework for the Dual Analysis of Feature Space and Data Space
abstract
With the surge of data-driven analysis techniques, there is a rising demand for enhancing the exploration of large high-dimensional data by enabling interactions for the joint analysis of features (i.e., dimensions). Such a dual analysis of the feature space and data space is characterized by three components, 1) a view visualizing feature summaries, 2) a view that visualizes the data records, and 3) a bidirectional linking of both plots triggered by human interaction in one of both visualizations, e.g., Linking & Brushing. Dual analysis approaches span many domains, e.g., medicine, crime analysis, and biology. The proposed solutions encapsulate various techniques, such as feature selection or statistical analysis. However, each approach establishes a new definition of dual analysis. To address this gap, we systematically reviewed published dual analysis methods to investigate and formalize the key elements, such as the techniques used to visualize the feature space and data space, as well as the interaction between both spaces. From the information elicited during our review, we propose a unified theoretical framework for dual analysis, encompassing all existing approaches extending the field. We apply our proposed formalization describing the interactions between each component and relate them to the addressed tasks. Additionally, we categorize the existing approaches using our framework and derive future research directions to advance dual analysis by including state-of-the-art visual analysis techniques to improve data exploration.
Frederik L. Dennig, Matthias Miller, Daniel A. Keim, Mennatallah El-Assady
IEEE Trans. Vis. Comput. Graph.3
2024 Automorphism Faithfulness Metrics for Symmetric Graph Drawings
abstract
In this article, we present new quality metrics for symmetric graph drawing based on group theory. Roughly speaking, the new metrics are faithfulness metrics, i.e., they measure how faithfully a drawing of a graph displays the ground truth (i.e., geometric automorphisms) of the graph as symmetries. More specifically, we introduce two types of automorphism faithfulness metrics for displaying: (1) a single geometric automorphism as a symmetry (axial or rotational), and (2) a group of geometric automorphisms (cyclic or dihedral). We present algorithms to compute the automorphism faithfulness metrics in O(n logn) time. Moreover, we also present efficient algorithms to detect exact symmetries in a graph drawing. We then validate our automorphism faithfulness metrics using deformation experiments. Finally, we use the metrics to evaluate existing graph drawing algorithms to compare how faithfully they display geometric automorphisms of a graph as symmetries.
Amyra Meidiana, Seok-Hee Hong 0001, Peter Eades, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.4
2024 SkiVis: Visual Exploration and Route Planning in Ski Resorts
abstract
Optimal ski route selection is a challenge based on a multitude of factors, such as the steepness, compass direction, or crowdedness. The personal preferences of every skier towards these factors require individual adaptations, which aggravate this task. Current approaches within this domain do not combine automated routing capabilities with user preferences, missing out on the possibility of integrating domain knowledge in the analysis process. We introduce SkiVis, a visual analytics application to interactively explore ski slopes and provide routing recommendations based on user preferences. In collaboration with ski guides and enthusiasts, we elicited requirements and guidelines for such an application and propose different workflows depending on the skiers' familiarity with the resort. In a case study on the resort of Ski Arlberg, we illustrate how to leverage volunteered geographic information to enable a numerical comparison between slopes. We evaluated our approach through a pair-analytics study and demonstrate how it supports skiers in discovering relevant and preference-based ski routes. Besides the tasks investigated in the study, we derive additional use cases from the interviews that showcase the further potential of SkiVis, and contribute directions for further research opportunities.
Julius Rauscher, Raphael Buchmüller, Daniel A. Keim, Matthias Miller
IEEE Trans. Vis. Comput. Graph.3
2024 Comparative Evaluation of Animated Scatter Plot Transitions
abstract
Scatter plots are popular for displaying 2D data, but in practice, many data sets have more than two dimensions. For the analysis of such multivariate data, it is often necessary to switch between scatter plots of different dimension pairs, e.g., in a scatter plot matrix (SPLOM). Alternative approaches include a "grand tour" for an overview of the entire data set or creating artificial axes from dimensionality reduction (DR). A cross-cutting concern in all techniques is the ability of viewers to find correspondence between data points in different views. Previous work proposed animations to preserve the mental map between view changes and to trace points as well as clusters between scatter plots of the same underlying data set. In this article, we evaluate a variety of spline- and rotation-based view transitions in a crowdsourced user study focusing on ecological validity. Using the study results, we assess each animation's suitability for tracing points and clusters across view changes. We evaluate whether the order of horizontal and vertical rotation is relevant for task accuracy. The results show that rotations with an orthographic camera or staged expansion of a depth axis significantly outperform all other animation techniques for the traceability of individual points. Further, we provide a ranking of the animated transition techniques for traceability of individual points. However, we could not find any significant differences for the traceability of clusters. Furthermore, we identified differences by animation direction that could guide further studies to determine potential confounds for these differences. We publish the study data for reuse and provide the animation framework as a D3.js plug-in.
Nils Rodrigues, Frederik L. Dennig, Vincent Brandt, Daniel A. Keim, Daniel Weiskopf
IEEE Trans. Vis. Comput. Graph.4
2023 Exploring Trajectory Data in Augmented Reality: A Comparative Study of Interaction Modalities
abstract
The visual exploration of trajectory data is crucial in domains such as animal behavior, molecular dynamics, and transportation. With the emergence of immersive technology, trajectory data, which is often inherently three-dimensional, can be analyzed in stereoscopic 3D, providing new opportunities for perception, engagement, and understanding. However, the interaction with the presented data remains a key challenge. While most applications depend on hand tracking, we see eye tracking as a promising yet under-explored interaction modality, while challenges such as imprecision or inadvertently triggered actions need to be addressed. In this work, we explore the potential of eye gaze interaction for the visual exploration of trajectory data within an AR environment. We integrate hand- and eye-based interaction techniques specifically designed for three common use cases and address known eye tracking challenges. We refine our techniques and setup based on a pilot user study (n=6) and find in a follow-up study (n=20) that gaze interaction can compete with hand-tracked interaction regarding effectiveness, efficiency, and task load for selection and cluster exploration tasks. However, time step analysis comes with higher answer times and task load. In general, we find the results and preferences to be user-dependent. Our work contributes to the field of immersive data exploration, underscoring the need for continued research on eye tracking interaction.
Lucas Joos, Karsten Klein 0001, Maximilian T. Fischer, Frederik L. Dennig, Daniel A. Keim, Michael Krone
ISMAR5
2023 VISITOR: Visual Interactive State Sequence Exploration for Reinforcement Learning
abstract
Abstract Understanding the behavior of deep reinforcement learning agents is a crucial requirement throughout their development. Existing work has addressed the identification of observable behavioral patterns in state sequences or analysis of isolated internal representations; however, the overall decision‐making of deep‐learning RL agents remains opaque. To tackle this, we present VISITOR, a visual analytics system enabling the analysis of entire state sequences, the diagnosis of singular predictions, and the comparison between agents. A sequence embedding view enables the multiscale analysis of state sequences, utilizing custom embedding techniques for a stable spatialization of the observations and internal states. We provide multiple layers: (1) a state space embedding, highlighting different groups of states inside the state‐action sequences, (2) a trajectory view, emphasizing decision points, (3) a network activation mapping, visualizing the relationship between observations and network activations, (4) a transition embedding, enabling the analysis of state‐to‐state transitions. The embedding view is accompanied by an interactive reward view that captures the temporal development of metrics, which can be linked directly to states in the embedding. Lastly, a model list allows for the quick comparison of models across multiple metrics. Annotations can be exported to communicate results to different audiences. Our two‐stage evaluation with eight experts confirms the effectiveness in identifying states of interest, comparing the quality of policies, and reasoning about the internal decision‐making processes.
Yannick Metz, Eugene Bykovets, Lucas Joos, Daniel A. Keim, Mennatallah El-Assady
Comput. Graph. Forum4
2023 Visual Analytics of Co-Occurrences to Discover Subspaces in Structured Data
abstract
We present an approach that shows all relevant subspaces of categorical data condensed in a single picture. We model the categorical values of the attributes as co-occurrences with data partitions generated from structured data using pattern mining. We show that these co-occurrences are a-priori , allowing us to greatly reduce the search space, effectively generating the condensed picture where conventional approaches filter out several subspaces as these are deemed insignificant. The task of identifying interesting subspaces is common but difficult due to exponential search spaces and the curse of dimensionality. One application of such a task might be identifying a cohort of patients defined by attributes such as gender, age, and diabetes type that share a common patient history, which is modeled as event sequences. Filtering the data by these attributes is common but cumbersome and often does not allow a comparison of subspaces. We contribute a powerful multi-dimensional pattern exploration approach (MDPE-approach) agnostic to the structured data type that models multiple attributes and their characteristics as co-occurrences, allowing the user to identify and compare thousands of subspaces of interest in a single picture. In our MDPE-approach, we introduce two methods to dramatically reduce the search space, outputting only the boundaries of the search space in the form of two tables. We implement the MDPE-approach in an interactive visual interface (MDPE-vis) that provides a scalable, pixel-based visualization design allowing the identification, comparison, and sense-making of subspaces in structured data. Our case studies using a gold-standard dataset and external domain experts confirm our approach’s and implementation’s applicability. A third use case sheds light on the scalability of our approach and a user study with 15 participants underlines its usefulness and power.
Wolfgang Jentner, Giuliana Lindholz, Hanna Hauptmann, Mennatallah El-Assady, Kwan-Liu Ma, Daniel A. Keim
ACM Trans. Interact. Intell. Syst.6
2023 Investigating the Sketchplan: A Novel Way of Identifying Tactical Behavior in Massive Soccer Datasets
abstract
Coaches and analysts prepare for upcoming matches by identifying common patterns in the positioning and movement of the competing teams in specific situations. Existing approaches in this domain typically rely on manual video analysis and formation discussion using whiteboards; or expert systems that rely on state-of-the-art video and trajectory visualization techniques and advanced user interaction. We bridge the gap between these approaches by contributing a light-weight, simplified interaction and visualization system, which we conceptualized in an iterative design study with the coaching team of a European first league soccer team. Our approach is walk-up usable by all domain stakeholders, and at the same time, can leverage advanced data retrieval and analysis techniques: a virtual magnetic tactic-board. Users place and move digital magnets on a virtual tactic-board, and these interactions get translated to spatio-temporal queries, used to retrieve relevant situations from massive team movement data. Despite such seemingly imprecise query input, our approach is highly usable, supports quick user exploration, and retrieval of relevant results via query relaxation. Appropriate simplified result visualization supports in-depth analyses to explore team behavior, such as formation detection, movement analysis, and what-if analysis. We evaluated our approach with several experts from European first league soccer clubs. The results show that our approach makes the complex analytical processes needed for the identification of tactical behavior directly accessible to domain experts for the first time, demonstrating our support of coaches in preparation for future encounters.
Daniel Seebacher, Tom Polk, Halldór Janetzko, Daniel A. Keim, Tobias Schreck, Manuel Stein
IEEE Trans. Vis. Comput. Graph.4
2022 Immersive Analytics with Abstract 3D Visualizations: A Survey
abstract
Abstract After a long period of scepticism, more and more publications describe basic research but also practical approaches to how abstract data can be presented in immersive environments for effective and efficient data understanding. Central aspects of this important research question in immersive analytics research are concerned with the use of 3D for visualization, the embedding in the immersive space, the combination with spatial data, suitable interaction paradigms and the evaluation of use cases. We provide a characterization that facilitates the comparison and categorization of published works and present a survey of publications that gives an overview of the state of the art, current trends, and gaps and challenges in current research.
Matthias Kraus 0002, Johannes Fuchs 0001, Björn Sommer 0001, Karsten Klein 0001, Ulrich Engelke, Daniel A. Keim, Falk Schreiber
Comput. Graph. Forum6
2022 Augmenting Digital Sheet Music through Visual Analytics
abstract
Abstract Music analysis tasks, such as structure identification and modulation detection, are tedious when performed manually due to the complexity of the common music notation (CMN). Fully automated analysis instead misses human intuition about relevance. Existing approaches use abstract data‐driven visualizations to assist music analysis but lack a suitable connection to the CMN. Therefore, music analysts often prefer to remain in their familiar context. Our approach enhances the traditional analysis workflow by complementing CMN with interactive visualization entities as minimally intrusive augmentations. Gradual step‐wise transitions empower analysts to retrace and comprehend the relationship between the CMN and abstract data representations. We leverage glyph‐based visualizations for harmony, rhythm and melody to demonstrate our technique's applicability. Design‐driven visual query filters enable analysts to investigate statistical and semantic patterns on various abstraction levels. We conducted pair analytics sessions with 16 participants of different proficiency levels to gather qualitative feedback about the intuitiveness, traceability and understandability of our approach. The results show that MusicVis supports music analysts in getting new insights about feature characteristics while increasing their engagement and willingness to explore.
Matthias Miller, Daniel Fürst, Hanna Hauptmann, Daniel A. Keim, Mennatallah El-Assady
Comput. Graph. Forum4
2022 CorpusVis: Visual Analysis of Digital Sheet Music Collections
abstract
Abstract Manually investigating sheet music collections is challenging for music analysts due to the magnitude and complexity of underlying features, structures, and contextual information. However, applying sophisticated algorithmic methods would require advanced technical expertise that analysts do not necessarily have. Bridging this gap, we contribute CorpusVis, an interactive visual workspace, enabling scalable and multi‐faceted analysis. Our proposed visual analytics dashboard provides access to computational methods, generating varying perspectives on the same data. The proposed application uses metadata including composers, type, epoch, and low‐level features, such as pitch, melody, and rhythm. To evaluate our approach, we conducted a pair‐analytics study with nine participants. The qualitative results show that CorpusVis supports users in performing exploratory and confirmatory analysis, leading them to new insights and findings. In addition, based on three exemplary workflows, we demonstrate how to apply our approach to different tasks, such as exploring musical features or comparing composers.
Matthias Miller, Julius Rauscher, Daniel A. Keim, Mennatallah El-Assady
Comput. Graph. Forum3
2022 Multiscale Visualization: A Structured Literature Analysis
abstract
Multiscale visualizations are typically used to analyze multiscale processes and data in various application domains, such as the visual exploration of hierarchical genome structures in molecular biology. However, creating such multiscale visualizations remains challenging due to the plethora of existing work and the expression ambiguity in visualization research. Up to today, there has been little work to compare and categorize multiscale visualizations to understand their design practices. In this article, we present a structured literature analysis to provide an overview of common design practices in multiscale visualization research. We systematically reviewed and categorized 122 published journal or conference articles between 1995 and 2020. We organized the reviewed articles in a taxonomy that reveals common design factors. Researchers and practitioners can use our taxonomy to explore existing work to create new multiscale navigation and visualization techniques. Based on the reviewed articles, we examine research trends and highlight open research challenges.
Eren Cakmak, Dominik Jäckle, Tobias Schreck, Daniel A. Keim, Johannes Fuchs 0001
IEEE Trans. Vis. Comput. Graph.4
2022 Visual Comparison of Networks in VR
abstract
Networks are an important means for the representation and analysis of data in a variety of research and application areas. While there are many efficient methods to create layouts for networks to support their visual analysis, approaches for the comparison of networks are still underexplored. Especially when it comes to the comparison of weighted networks, which is an important task in several areas, such as biology and biomedicine, there is a lack of efficient visualization approaches. With the availability of affordable high-quality virtual reality (VR) devices, such as head-mounted displays (HMDs), the research field of immersive analytics emerged and showed great potential for using the new technology for visual data exploration. However, the use of immersive technology for the comparison of networks is still underexplored. With this work, we explore how weighted networks can be visually compared in an immersive VR environment and investigate how visual representations can benefit from the extended 3D design space. For this purpose, we develop different encodings for 3D node-link diagrams supporting the visualization of two networks within a single representation and evaluate them in a pilot user study. We incorporate the results into a more extensive user study comparing node-link representations with matrix representations encoding two networks simultaneously. The data and tasks designed for our experiments are similar to those occurring in real-world scenarios. Our evaluation shows significantly better results for the node-link representations, which is contrary to comparable 2D experiments and indicates a high potential for using VR for the visual comparison of networks.
Lucas Joos, Sabrina Jaeger-Honz, Falk Schreiber, Daniel A. Keim, Karsten Klein 0001
IEEE Trans. Vis. Comput. Graph.4
2022 VisInReport: Complementing Visual Discourse Analytics Through Personalized Insight Reports
abstract
We present VisInReport, a visual analytics tool that supports the manual analysis of discourse transcripts and generates reports based on user interaction. As an integral part of scholarly work in the social sciences and humanities, discourse analysis involves an aggregation of characteristics identified in the text, which, in turn, involves a prior identification of regions of particular interest. Manual data evaluation requires extensive effort, which can be a barrier to effective analysis. Our system addresses this challenge by augmenting the users' analysis with a set of automatically generated visualization layers. These layers enable the detection and exploration of relevant parts of the discussion supporting several tasks, such as topic modeling or question categorization. The system summarizes the extracted events visually and verbally, generating a content-rich insight into the data and the analysis process. During each analysis session, VisInReport builds a shareable report containing a curated selection of interactions and annotations generated by the analyst. We evaluate our approach on real-world datasets through a qualitative study with domain experts from political science, computer science, and linguistics. The results highlight the benefit of integrating the analysis and reporting processes through a visual analytics system, which supports the communication of results among collaborating researchers.
Rita Sevastjanova, Mennatallah El-Assady, Adam Bradley, Christopher Collins 0001, Miriam Butt, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.6
2022 Task-Based Visual Interactive Modeling: Decision Trees and Rule-Based Classifiers
abstract
Visual analytics enables the coupling of machine learning models and humans in a tightly integrated workflow, addressing various analysis tasks. Each task poses distinct demands to analysts and decision-makers. In this survey, we focus on one canonical technique for rule-based classification, namely decision tree classifiers. We provide an overview of available visualizations for decision trees with a focus on how visualizations differ with respect to 16 tasks. Further, we investigate the types of visual designs employed, and the quality measures presented. We find that (i) interactive visual analytics systems for classifier development offer a variety of visual designs, (ii) utilization tasks are sparsely covered, (iii) beyond classifier development, node-link diagrams are omnipresent, (iv) even systems designed for machine learning experts rarely feature visual representations of quality measures other than accuracy. In conclusion, we see a potential for integrating algorithmic techniques, mathematical quality measures, and tailored interactive visualizations to enable human experts to utilize their knowledge more effectively.
Dirk Streeb, Yannick Metz, Udo Schlegel, Bruno Schneider, Mennatallah El-Assady, Hansjörg Neth, Min Chen 0001, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.8
2021 VulnEx: Exploring Open-Source Software Vulnerabilities in Large Development Organizations to Understand Risk Exposure
abstract
The prevalent usage of open-source software (OSS) has led to an increased interest in resolving potential third-party security risks by fixing common vulnerabilities and exposures (CVEs). However, even with automated code analysis tools in place, security analysts often lack the means to obtain an overview of vulnerable OSS reuse in large software organizations. In this design study, we propose VULNEX (Vulnerability Explorer), a tool to audit entire software development organizations. We introduce three complementary table based representations to identify and assess vulnerability exposures due to OSS, which we designed in collaboration with security analysts. The presented tool allows examining problematic projects and applications (repositories), third-party libraries, and vulnerabilities across a software organization. We show the applicability of our tool through a use case and preliminary expert feedback.
Frederik L. Dennig, Eren Cakmak, Henrik Plate, Daniel A. Keim
VizSec4
2021 SpatialRugs: A compact visualization of space and time for analyzing collective movement data
Juri Buchmüller, Udo Schlegel, Eren Cakmak, Daniel A. Keim, Evanthia Dimara
Comput. Graph.4
2021 Co-adaptive visual data analysis and guidance processes
Fabian Sperrle, Astrik Jeitler, Jürgen Bernard, Daniel A. Keim, Mennatallah El-Assady
Comput. Graph.4
2021 ParSetgnostics: Quality Metrics for Parallel Sets
abstract
Abstract While there are many visualization techniques for exploring numeric data, only a few work with categorical data. One prominent example is Parallel Sets, showing data frequencies instead of data points ‐ analogous to parallel coordinates for numerical data. As nominal data does not have an intrinsic order, the design of Parallel Sets is sensitive to visual clutter due to overlaps, crossings, and subdivision of ribbons hindering readability and pattern detection. In this paper, we propose a set of quality metrics, called ParSetgnostics (Parallel Sets diagnostics), which aim to improve Parallel Sets by reducing clutter. These quality metrics quantify important properties of Parallel Sets such as overlap, orthogonality, ribbon width variance, and mutual information to optimize the category and dimension ordering. By conducting a systematic correlation analysis between the individual metrics, we ensure their distinctiveness. Further, we evaluate the clutter reduction effect of ParSetgnostics by reconstructing six datasets from previous publications using Parallel Sets measuring and comparing their respective properties. Our results show that ParSetgostics facilitates multi‐dimensional analysis of categorical data by automatically providing optimized Parallel Set designs with a clutter reduction of up to 81% compared to the originally proposed Parallel Sets visualizations.
Frederik L. Dennig, Maximilian T. Fischer, Michael Blumenschein, Johannes Fuchs 0001, Daniel A. Keim, Evanthia Dimara
Comput. Graph. Forum5
2021 CommAID: Visual Analytics for Communication Analysis through Interactive Dynamics Modeling
abstract
Abstract Communication consists of both meta‐information as well as content. Currently, the automated analysis of such data often focuses either on the network aspects via social network analysis or on the content, utilizing methods from text‐mining. However, the first category of approaches does not leverage the rich content information, while the latter ignores the conversation environment and the temporal evolution, as evident in the meta‐information. In contradiction to communication research, which stresses the importance of a holistic approach, both aspects are rarely applied simultaneously, and consequently, their combination has not yet received enough attention in automated analysis systems. In this work, we aim to address this challenge by discussing the difficulties and design decisions of such a path as well as contribute CommAID, a blueprint for a holistic strategy to communication analysis. It features an integrated visual analytics design to analyze communication networks through dynamics modeling, semantic pattern retrieval, and a user‐adaptable and problem‐specific machine learning‐based retrieval system. An interactive multi‐level matrix‐based visualization facilitates a focused analysis of both network and content using inline visuals supporting cross‐checks and reducing context switches. We evaluate our approach in both a case study and through formative evaluation with eight law enforcement experts using a real‐world communication corpus. Results show that our solution surpasses existing techniques in terms of integration level and applicability. With this contribution, we aim to pave the path for a more holistic approach to communication analysis.
Maximilian T. Fischer, Daniel Seebacher, Rita Sevastjanova, Daniel A. Keim, Mennatallah El-Assady
Comput. Graph. Forum4
2021 A Survey of Human-Centered Evaluations in Human-Centered Machine Learning
abstract
Abstract Visual analytics systems integrate interactive visualizations and machine learning to enable expert users to solve complex analysis tasks. Applications combine techniques from various fields of research and are consequently not trivial to evaluate. The result is a lack of structure and comparability between evaluations. In this survey, we provide a comprehensive overview of evaluations in the field of human‐centered machine learning. We particularly focus on human‐related factors that influence trust, interpretability, and explainability. We analyze the evaluations presented in papers from top conferences and journals in information visualization and human‐computer interaction to provide a systematic review of their setup and findings. From this survey, we distill design dimensions for structured evaluations, identify evaluation gaps, and derive future research opportunities.
Fabian Sperrle, Mennatallah El-Assady, Grace Guo 0001, Rita Borgo, Polo Chau, Alex Endert, Daniel A. Keim
Comput. Graph. Forum7
2021 Learning Contextualized User Preferences for Co-Adaptive Guidance in Mixed-Initiative Topic Model Refinement
abstract
Abstract Mixed‐initiative visual analytics systems support collaborative human‐machine decision‐making processes. However, many multi‐objective optimization tasks, such as topic model refinement, are highly subjective and context‐dependent. Hence, systems need to adapt their optimization suggestions throughout the interactive refinement process to provide efficient guidance. To tackle this challenge, we present a technique for learning context‐dependent user preferences and demonstrate its applicability to topic model refinement. We deploy agents with distinct associated optimization strategies that compete for the user's acceptance of their suggestions. To decide when to provide guidance, each agent maintains an intelligible, rule‐based classifier over context vectorizations that captures the development of quality metrics between distinct analysis states. By observing implicit and explicit user feedback, agents learn in which contexts to provide their specific guidance operation. An agent in topic model refinement might, for example, learn to react to declining model coherence by suggesting to split a topic. Our results confirm that the rules learned by agents capture contextual user preferences. Further, we show that the learned rules are transferable between similar datasets, avoiding common cold‐start problems and enabling a continuous refinement of agents across corpora.
Fabian Sperrle, Hanna Hauptmann, Daniel A. Keim, Mennatallah El-Assady
Comput. Graph. Forum3
2021 Integrating Data and Model Space in Ensemble Learning by Visual Analytics
abstract
Ensembles of classifier models typically deliver superior performance and can outperform single classifier models given a dataset and classification task at hand. However, the gain in performance comes together with the lack of comprehensibility, posing a challenge to understand how each model affects the classification outputs and from where the errors come. We propose a tight visual integration of the data and the model space for exploring and combining classifier models. We introduce an interactive workflow that builds upon the visual integration and enables the effective exploration of classification outputs and models. The involvement of the user is key to our approach. Therefore, we elaborate on the role of the human and connect our approach to theoretical frameworks on human-centered machine learning. We showcase the usefulness of our approach and the integration of the user via binary and multiclass classification problems. Based on ensembles automatically selected by a standard ensemble selection algorithm, the user can manipulate models and alternative combinations.
Bruno Schneider, Dominik Jäckle, Florian Stoffel, Alexandra Diehl, Johannes Fuchs 0001, Daniel A. Keim
IEEE Trans. Big Data6
2021 Visual Analysis of Spatio-Temporal Event Predictions: Investigating the Spread Dynamics of Invasive Species
abstract
Invasive species are a major cause of ecological damage and commercial losses. A current problem spreading in North America and Europe is the vinegar fly Drosophila suzukii. Unlike other Drosophila, it infests non-rotting and healthy fruits and is therefore of concern to fruitgrowers, such as vintners. Consequently, large amounts of data about infestations have been collected in recent years. However, there is a lack of interactive methods to investigate this data. We employ ensemble-based classification to predict areas susceptible to infestation by D.suzukii and bring them into a spatio-temporal context using maps and glyph-based visualizations. Following the information-seeking mantra, we provide a visual analysis system Drosophigatorfor spatio-temporal event prediction, enabling the investigation of the spread dynamics of invasive species. We demonstrate the usefulness of this approach in two use cases.
Daniel Seebacher, Johannes Häußler, Michael Hundt, Manuel Stein, Hannes Müller, Ulrich Engelke, Daniel A. Keim
IEEE Trans. Big Data7
2021 Multiscale Snapshots: Visual Analysis of Temporal Summaries in Dynamic Graphs
abstract
The overview-driven visual analysis of large-scale dynamic graphs poses a major challenge. We propose Multiscale Snapshots, a visual analytics approach to analyze temporal summaries of dynamic graphs at multiple temporal scales. First, we recursively generate temporal summaries to abstract overlapping sequences of graphs into compact snapshots. Second, we apply graph embeddings to the snapshots to learn low-dimensional representations of each sequence of graphs to speed up specific analytical tasks (e.g., similarity search). Third, we visualize the evolving data from a coarse to fine-granular snapshots to semi-automatically analyze temporal states, trends, and outliers. The approach enables us to discover similar temporal summaries (e.g., reoccurring states), reduces the temporal data to speed up automatic analysis, and to explore both structural and temporal properties of a dynamic graph. We demonstrate the usefulness of our approach by a quantitative evaluation and the application to a real-world dataset.
Eren Cakmak, Udo Schlegel, Dominik Jäckle, Daniel A. Keim, Tobias Schreck
IEEE Trans. Vis. Comput. Graph.4
2021 Visual Analytics for Temporal Hypergraph Model Exploration
abstract
Many processes, from gene interaction in biology to computer networks to social media, can be modeled more precisely as temporal hypergraphs than by regular graphs. This is because hypergraphs generalize graphs by extending edges to connect any number of vertices, allowing complex relationships to be described more accurately and predict their behavior over time. However, the interactive exploration and seamless refinement of such hypergraph-based prediction models still pose a major challenge. We contribute Hyper-Matrix, a novel visual analytics technique that addresses this challenge through a tight coupling between machine-learning and interactive visualizations. In particular, the technique incorporates a geometric deep learning model as a blueprint for problem-specific models while integrating visualizations for graph-based and category-based data with a novel combination of interactions for an effective user-driven exploration of hypergraph models. To eliminate demanding context switches and ensure scalability, our matrix-based visualization provides drill-down capabilities across multiple levels of semantic zoom, from an overview of model predictions down to the content. We facilitate a focused analysis of relevant connections and groups based on interactive user-steering for filtering and search tasks, a dynamically modifiable partition hierarchy, various matrix reordering techniques, and interactive model feedback. We evaluate our technique in a case study and through formative evaluation with law enforcement experts using real-world internet forum communication data. The results show that our approach surpasses existing solutions in terms of scalability and applicability, enables the incorporation of domain knowledge, and allows for fast search-space traversal. With the proposed technique, we pave the way for the visual analytics of temporal hypergraphs in a wide variety of domains.
Maximilian T. Fischer, Devanshu Arya, Dirk Streeb, Daniel Seebacher, Daniel A. Keim, Marcel Worring
IEEE Trans. Vis. Comput. Graph.5
2021 MultiSegVA: Using Visual Analytics to Segment Biologging Time Series on Multiple Scales
abstract
Segmenting biologging time series of animals on multiple temporal scales is an essential step that requires complex techniques with careful parameterization and possibly cross-domain expertise. Yet, there is a lack of visual-interactive tools that strongly support such multi-scale segmentation. To close this gap, we present our MultiSegVA platform for interactively defining segmentation techniques and parameters on multiple temporal scales. MultiSegVA primarily contributes tailored, visual-interactive means and visual analytics paradigms for segmenting unlabeled time series on multiple scales. Further, to flexibly compose the multi-scale segmentation, the platform contributes a new visual query language that links a variety of segmentation techniques. To illustrate our approach, we present a domain-oriented set of segmentation techniques derived in collaboration with movement ecologists. We demonstrate the applicability and usefulness of MultiSegVA in two real-world use cases from movement ecology, related to behavior analysis after environment-aware segmentation, and after progressive clustering. Expert feedback from movement ecologists shows the effectiveness of tailored visual-interactive means and visual analytics paradigms at segmenting multi-scale data, enabling them to perform semantically meaningful analyses. A third use case demonstrates that MultiSegVA is generalizable to other domains.
Philipp Meschenmoser, Juri Buchmüller, Daniel Seebacher, Martin Wikelski, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.5
2021 Why Visualize? Untangling a Large Network of Arguments
abstract
Visualization has been deemed a useful technique by researchers and practitioners, alike, leaving a trail of arguments behind that reason why visualization works. In addition, examples of misleading usages of visualizations in information communication have occasionally been pointed out. Thus, to contribute to the fundamental understanding of our discipline, we require a comprehensive collection of arguments on "why visualize?" (or "why not?"), untangling the rationale behind positive and negative viewpoints. In this paper, we report a theoretical study to understand the underlying reasons of various arguments; their relationships (e.g., built-on, and conflict); and their respective dependencies on tasks, users, and data. We curated an argumentative network based on a collection of arguments from various fields, including information visualization, cognitive science, psychology, statistics, philosophy, and others. Our work proposes several categorizations for the arguments, and makes their relations explicit. We contribute the first comprehensive and systematic theoretical study of the arguments on visualization. Thereby, we provide a roadmap towards building a foundation for visualization theory and empirical research as well as for practical application in the critique and design of visualizations. In addition, we provide our argumentation network and argument collection online at https://whyvis.dbvis.de, supported by an interactive visualization.
Dirk Streeb, Mennatallah El-Assady, Daniel A. Keim, Min Chen 0001
IEEE Trans. Vis. Comput. Graph.3
2020 Quality Metrics for Symmetric Graph Drawings
abstract
In this paper, we present a framework for quality metrics that measure symmetry, that is, how faithfully a drawing of a graph displays the ground truth geometric automorphisms as symmetries. The quality metrics are based on group theory as well as geometry. More specifically, we introduce two types of symmetry quality metrics for displaying: (1) a single geometric automorphism as a symmetry (axial or rotational) and (2) a group of geometric automorphisms (cyclic or dihedral). We also present algorithms to compute the symmetry quality metrics in O(n log n) time. We validate our symmetry quality metrics using deformation experiments. We then use the metrics to evaluate existing graph layouts to compare how faithfully they display geometric automorphisms of a graph as symmetries.
Amyra Meidiana, Seok-Hee Hong 0001, Peter Eades, Daniel A. Keim
PacificVis4
2020 Assessing 2D and 3D Heatmaps for Comparative Analysis: An Empirical Study
abstract
Heatmaps are a popular visualization technique that encode 2D density distributions using color or brightness. Experimental studies have shown though that both of these visual variables are inaccurate when reading and comparing numeric data values. A potential remedy might be to use 3D heatmaps by introducing height as a third dimension to encode the data. Encoding abstract data in 3D, however, poses many problems, too. To better understand this tradeoff, we conducted an empirical study (N=48) to evaluate the user performance of 2D and 3D heatmaps for comparative analysis tasks. We test our conditions on a conventional 2D screen, but also in a virtual reality environment to allow for real stereoscopic vision. Our main results show that 3D heatmaps are superior in terms of error rate when reading and comparing single data items. However, for overview tasks, the well-established 2D heatmap performs better.
Matthias Kraus 0002, Katrin Angerbauer, Juri Buchmüller, Daniel Schweitzer, Daniel A. Keim, Michael Sedlmair, Johannes Fuchs 0001
CHI5
2020 A Comparative Study of Orientation Support Tools in Virtual Reality Environments with Virtual Teleportation
abstract
Movement-compensating interactions like teleportation are commonly deployed techniques in virtual reality environments. Although practical, they tend to cause disorientation while navigating. Previous studies show the effectiveness of orientation-supporting tools, such as trails, in reducing such disorientation and reveal different strengths and weaknesses of individual tools. However, to date, there is a lack of a systematic comparison of those tools when teleportation is used as a movement-compensating technique, in particular under consideration of different tasks. In this paper, we compare the effects of three orientation-supporting tools, namely minimap, trail, and heatmap. We conducted a quantitative user study with 48 participants to investigate the accuracy and efficiency when executing four exploration and search tasks. As dependent variables, task performance, completion time, space coverage, amount of revisiting, retracing time, and memorability were measured. Overall, our results indicate that orientation-supporting tools improve task completion times and revisiting behavior. The trail and heatmap tools were particularly useful for speed-focused tasks, minimal revisiting, and space coverage. The minimap increased memorability and especially supported retracing tasks. These results suggest that virtual reality systems should provide orientation aid tailored to the specific tasks of the users.
Matthias Kraus 0002, Hanna Hauptmann, Philipp Meschenmoser, Daniel Schweitzer, Daniel A. Keim, Michael Sedlmair, Johannes Fuchs 0001
ISMAR5
2020 Towards visual debugging for multi-target time series classification
abstract
Multi-target classification of multivariate time series data poses a challenge in many real-world applications (e.g., predictive maintenance). Machine learning methods, such as random forests and neural networks, support training these classifiers. However, the debugging and analysis of possible misclassifications remain challenging due to the often complex relations between targets, classes, and the multivariate time series data. We propose a model-agnostic visual debugging workflow for multi-target time series classification that enables the examination of relations between targets, partially correct predictions, potential confusions, and the classified time series data. The workflow, as well as the prototype, aims to foster an in-depth analysis of multi-target classification results to identify potential causes of mispredictions visually. We demonstrate the usefulness of the workflow in the field of predictive maintenance in a usage scenario to show how users can iteratively explore and identify critical classes, as well as, relationships between targets.
Udo Schlegel, Eren Cakmak, Hiba Arnout, Mennatallah El-Assady, Daniela Oelke, Daniel A. Keim
IUI6
2020 v-plots: Designing Hybrid Charts for the Comparative Analysis of Data Distributions
abstract
Abstract Comparing data distributions is a core focus in descriptive statistics, and part of most data analysis processes across disciplines. In particular, comparing distributions entails numerous tasks, ranging from identifying global distribution properties, comparing aggregated statistics (e.g., mean values), to the local inspection of single cases. While various specialized visualizations have been proposed (e.g., box plots, histograms, or violin plots), they are not usually designed to support more than a few tasks, unless they are combined. In this paper, we present the v‐plot designer; a technique for authoring custom hybrid charts, combining mirrored bar charts, difference encodings, and violin‐style plots. v‐plots are customizable and enable the simultaneous comparison of data distributions on global, local, and aggregation levels. Our system design is grounded in an expert survey that compares and evaluates 20 common visualization techniques to derive guidelines for the task‐driven selection of appropriate visualizations. This knowledge externalization step allowed us to develop a guiding wizard that can tailor v‐plots to individual tasks and particular distribution properties. Finally, we confirm the usefulness of our system design and the user‐guiding process by measuring the fitness for purpose and applicability in a second study with four domain and statistic experts.
Michael Blumenschein, Luka J. Debbeler, Nadine C. Lages, Britta Renner, Daniel A. Keim, Mennatallah El-Assady
Comput. Graph. Forum5
2020 Evaluating Reordering Strategies for Cluster Identification in Parallel Coordinates
abstract
Abstract The ability to perceive patterns in parallel coordinates plots (PCPs) is heavily influenced by the ordering of the dimensions. While the community has proposed over 30 automatic ordering strategies, we still lack empirical guidance for choosing an appropriate strategy for a given task. In this paper, we first propose a classification of tasks and patterns and analyze which PCP reordering strategies help in detecting them. Based on our classification, we then conduct an empirical user study with 31 participants to evaluate reordering strategies for cluster identification tasks. We particularly measure time, identification quality, and the users’ confidence for two different strategies using both synthetic and real‐world datasets. Our results show that, somewhat unexpectedly, participants tend to focus on dissimilar rather than similar dimension pairs when detecting clusters, and are more confident in their answers. This is especially true when increasing the amount of clutter in the data. As a result of these findings, we propose a new reordering strategy based on the dissimilarity of neighboring dimension pairs.
Michael Blumenschein, David Pomerenke, Daniel A. Keim, Johannes Fuchs 0001
Comput. Graph. Forum4
2020 MotionGlyphs: Visual Abstraction of Spatio-Temporal Networks in Collective Animal Behavior
abstract
Abstract Domain experts for collective animal behavior analyze relationships between single animal movers and groups of animals over time and space to detect emergent group properties. A common way to interpret this type of data is to visualize it as a spatio‐temporal network. Collective behavior data sets are often large, and may hence result in dense and highly connected node‐link diagrams, resulting in issues of node‐overlap and edge clutter. In this design study, in an iterative design process, we developed glyphs as a design for seamlessly encoding relationships and movement characteristics of a single mover or clusters of movers. Based on these glyph designs, we developed a visual exploration prototype, MotionGlyphs, that supports domain experts in interactively filtering, clustering, and animating spatio‐temporal networks for collective animal behavior analysis. By means of an expert evaluation, we show how MotionGlyphs supports important tasks and analysis goals of our domain experts, and we give evidence of the usefulness for analyzing spatio‐temporal networks of collective animal behavior.
Eren Cakmak, Hanna Hauptmann, Juri Buchmüller, Johannes Fuchs 0001, Tobias Schreck, Alex Jordan, Daniel A. Keim
Comput. Graph. Forum7
2020 Semantic Concept Spaces: Guided Topic Model Refinement using Word-Embedding Projections
abstract
We present a framework that allows users to incorporate the semantics of their domain knowledge for topic model refinement while remaining model-agnostic. Our approach enables users to (1) understand the semantic space of the model, (2) identify regions of potential conflicts and problems, and (3) readjust the semantic relation of concepts based on their understanding, directly influencing the topic modeling. These tasks are supported by an interactive visual analytics workspace that uses word-embedding projections to define concept regions which can then be refined. The user-refined concepts are independent of a particular document collection and can be transferred to related corpora. All user interactions within the concept space directly affect the semantic relations of the underlying vector space model, which, in turn, change the topic modeling. In addition to direct manipulation, our system guides the users' decision-making process through recommended interactions that point out potential improvements. This targeted refinement aims at minimizing the feedback required for an efficient human-in-the-loop process. We confirm the improvements achieved through our approach in two user studies that show topic model quality improvements through our visual knowledge externalization and learning process.
Mennatallah El-Assady, Rebecca Kehlbeck, Christopher Collins 0001, Daniel A. Keim, Oliver Deussen
IEEE Trans. Vis. Comput. Graph.4
2020 The Impact of Immersion on Cluster Identification Tasks
abstract
Recent developments in technology encourage the use of head-mounted displays (HMDs) as a medium to explore visualizations in virtual realities (VRs). VR environments (VREs) enable new, more immersive visualization design spaces compared to traditional computer screens. Previous studies in different domains, such as medicine, psychology, and geology, report a positive effect of immersion, e.g., on learning performance or phobia treatment effectiveness. Our work presented in this paper assesses the applicability of those findings to a common task from the information visualization (InfoVis) domain. We conducted a quantitative user study to investigate the impact of immersion on cluster identification tasks in scatterplot visualizations. The main experiment was carried out with 18 participants in a within-subjects setting using four different visualizations, (1) a 2D scatterplot matrix on a screen, (2) a 3D scatterplot on a screen, (3) a 3D scatterplot miniature in a VRE and (4) a fully immersive 3D scatterplot in a VRE. The four visualization design spaces vary in their level of immersion, as shown in a supplementary study. The results of our main study indicate that task performance differs between the investigated visualization design spaces in terms of accuracy, efficiency, memorability, sense of orientation, and user preference. In particular, the 2D visualization on the screen performed worse compared to the 3D visualizations with regard to the measured variables. The study shows that an increased level of immersion can be a substantial benefit in the context of 3D data and cluster detection.
Matthias Kraus 0002, Niklas Weiler, Daniela Oelke, Johannes Kehrer, Daniel A. Keim, Johannes Fuchs 0001
IEEE Trans. Vis. Comput. Graph.5
2019 A Quality Metric for Visualization of Clusters in Graphs
Amyra Meidiana, Seok-Hee Hong 0001, Peter Eades, Daniel A. Keim
GD4
2019 From Movement to Events: Improving Soccer Match Annotations
Manuel Stein, Daniel Seebacher, Tassilo Karge, Tom Polk, Michael Grossniklaus, Daniel A. Keim
MMM (1)6
2019 SurgeryCuts: Embedding Additional Information in Maps without Occluding Features
abstract
Abstract Visualizing contextual information to a map often comes at the expense of overplotting issues. Especially for use cases with relevant map features in the immediate vicinity of an information to add, occlusion of the relevant map context should be avoided. We present SurgeryCuts, a map manipulation technique for the creation of additional canvas area for contextual visualizations on maps. SurgeryCuts is occlusion‐free and does not shift, zoom or alter the map viewport. Instead, relevant parts of the map can be cut apart. The affected area is controlledly distorted using a parameterizable warping function fading out the map distortion depending on the distance to the cut. We define extended metrics for our approach and compare to related approaches. As well, we demonstrate the applicability of our approach at the example of tangible use cases and a comparative user study.
Marco Angelini, Juri Buchmüller, Daniel A. Keim, Philipp Meschenmoser, Giuseppe Santucci
Comput. Graph. Forum3
2019 Commercial Visual Analytics Systems-Advances in the Big Data Analytics Field
abstract
Five years after the first state-of-the-art report on Commercial Visual Analytics Systems we present a reevaluation of the Big Data Analytics field. We build on the success of the 2012 survey, which was influential even beyond the boundaries of the InfoVis and Visual Analytics (VA) community. While the field has matured significantly since the original survey, we find that innovation and research-driven development are increasingly sacrificed to satisfy a wide range of user groups. We evaluate new product versions on established evaluation criteria, such as available features, performance, and usability, to extend on and assure comparability with the previous survey. We also investigate previously unavailable products to paint a more complete picture of the commercial VA landscape. Furthermore, we introduce novel measures, like suitability for specific user groups and the ability to handle complex data types, and undertake a new case study to highlight innovative features. We explore the achievements in the commercial sector in addressing VA challenges and propose novel developments that should be on systems' roadmaps in the coming years.
Michael Behrisch 0001, Dirk Streeb, Florian Stoffel, Daniel Seebacher, Brian Matejek, Stefan Weber 0004, Sebastian Mittelstädt, Hanspeter Pfister, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.9
2019 MotionRugs: Visualizing Collective Trends in Space and Time
abstract
Understanding the movement patterns of collectives, such as flocks of birds or fish swarms, is an interesting open research question. The collectives are driven by mutual objectives or react to individual direction changes and external influence factors and stimuli. The challenge in visualizing collective movement data is to show space and time of hundreds of movements at the same time to enable the detection of spatiotemporal patterns. In this paper, we propose MotionRugs, a novel space efficient technique for visualizing moving groups of entities. Building upon established space-partitioning strategies, our approach reduces the spatial dimensions in each time step to a one-dimensional ordered representation of the individual entities. By design, MotionRugs provides an overlap-free, compact overview of the development of group movements over time and thus, enables analysts to visually identify and explore group-specific temporal patterns. We demonstrate the usefulness of our approach in the field of fish swarm analysis and report on initial feedback of domain experts from the field of collective behavior.
Juri Buchmüller, Dominik Jäckle, Eren Cakmak, Ulrik Brandes, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.5
2019 Visual Analytics for Topic Model Optimization based on User-Steerable Speculative Execution
abstract
To effectively assess the potential consequences of human interventions in model-driven analytics systems, we establish the concept of speculative execution as a visual analytics paradigm for creating user-steerable preview mechanisms. This paper presents an explainable, mixed-initiative topic modeling framework that integrates speculative execution into the algorithmic decisionmaking process. Our approach visualizes the model-space of our novel incremental hierarchical topic modeling algorithm, unveiling its inner-workings. We support the active incorporation of the user's domain knowledge in every step through explicit model manipulation interactions. In addition, users can initialize the model with expected topic seeds, the backbone priors. For a more targeted optimization, the modeling process automatically triggers a speculative execution of various optimization strategies, and requests feedback whenever the measured model quality deteriorates. Users compare the proposed optimizations to the current model state and preview their effect on the next model iterations, before applying one of them. This supervised human-in-the-loop process targets maximum improvement for minimum feedback and has proven to be effective in three independent studies that confirm topic model quality improvements.
Mennatallah El-Assady, Fabian Sperrle, Oliver Deussen, Daniel A. Keim, Christopher Collins 0001
IEEE Trans. Vis. Comput. Graph.4
2019 Bridging Text Visualization and Mining: A Task-Driven Survey
abstract
Visual text analytics has recently emerged as one of the most prominent topics in both academic research and the commercial world. To provide an overview of the relevant techniques and analysis tasks, as well as the relationships between them, we comprehensively analyzed 263 visualization papers and 4,346 mining papers published between 1992-2017 in two fields: visualization and text mining. From the analysis, we derived around 300 concepts (visualization techniques, mining techniques, and analysis tasks) and built a taxonomy for each type of concept. The co-occurrence relationships between the concepts were also extracted. Our research can be used as a stepping-stone for other researchers to 1) understand a common set of concepts used in this research topic; 2) facilitate the exploration of the relationships between visualization techniques, mining techniques, and analysis tasks; 3) understand the current practice in developing visual text analytics tools; 4) seek potential research opportunities by narrowing the gulf between visualization and mining techniques based on the analysis tasks; and 5) analyze other interdisciplinary research areas in a similar way. We have also contributed a web-based visualization tool for analyzing and understanding research trends and opportunities in visual text analytics.
Shixia Liu, Xiting Wang, Christopher Collins 0001, Wenwen Dou, Fang-Xin Ou-Yang, Mennatallah El-Assady, Liu Jiang, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.8
2019 VIS4ML: An Ontology for Visual Analytics Assisted Machine Learning
abstract
While many VA workflows make use of machine-learned models to support analytical tasks, VA workflows have become increasingly important in understanding and improving Machine Learning (ML) processes. In this paper, we propose an ontology (VIS4ML) for a subarea of VA, namely "VA-assisted ML". The purpose of VIS4ML is to describe and understand existing VA workflows used in ML as well as to detect gaps in ML processes and the potential of introducing advanced VA techniques to such processes. Ontologies have been widely used to map out the scope of a topic in biology, medicine, and many other disciplines. We adopt the scholarly methodologies for constructing VIS4ML, including the specification, conceptualization, formalization, implementation, and validation of ontologies. In particular, we reinterpret the traditional VA pipeline to encompass model-development workflows. We introduce necessary definitions, rules, syntaxes, and visual notations for formulating VIS4ML and make use of semantic web technologies for implementing it in the Web Ontology Language (OWL). VIS4ML captures the high-level knowledge about previous workflows where VA is used to assist in ML. It is consistent with the established VA concepts and will continue to evolve along with the future developments in VA and ML. While this ontology is an effort for building the theoretical foundation of VA, it can be used by practitioners in real-world applications to optimize model-development workflows by systematically examining the potential benefits that can be brought about by either machine or human capabilities. Meanwhile, VIS4ML is intended to be extensible and will continue to be updated to reflect future advancements in using VA for building high-quality data-analytical models or for building such models rapidly.
Dominik Sacha, Matthias Kraus 0002, Daniel A. Keim, Min Chen 0001
IEEE Trans. Vis. Comput. Graph.3
2018 Provenance-Based Visual Data Exploration with EVLIN
abstract
Tools for visual data exploration allow users to visually browse through and analyze datasets to possibly reveal interesting infor- mation hidden in the data that users are a priori unaware of. Such tools rely on both query recommendations to select data to be visualized and visualization recommendations for these data to best support users in their visual data exploration process. EVLIN ( e xploring v isually with lin eage) is a system that assists users in visually exploring relational data stored in a data ware- house. EVLIN implements novel techniques for recommending both queries and their result visualization in an integrated and interactive way [ 3 ]. Recommendations rely on provenance (aka lineage) that describes the production process of displayed data . The demonstration of EVLIN includes an introduction to its features and functionality through sample exploration sessions. Conference attendees will then have the opportunity to gain hands- on experience of provenance-based visual data exploration by performing their own exploration sessions. These sessions will explore real-world data from several domains. While exploration sessions use a Web-based visual interface, the demonstration also features a researcher console, where attendees may have a look behind the scenes to get a more in-depth understanding of the underlying recommendation algorithms.
Houssem Ben Lahmar, Melanie Herschel, Michael Blumenschein, Daniel A. Keim
EDBT4
2018 G-Rap: interactive text synthesis using recurrent neural network suggestions
Udo Schlegel, Eren Cakmak, Juri Buchmüller, Daniel A. Keim
ESANN4
2018 Viewing Visual Analytics as Model Building
abstract
Abstract To complement the currently existing definitions and conceptual frameworks of visual analytics, which focus mainly on activities performed by analysts and types of techniques they use, we attempt to define the expected results of these activities. We argue that the main goal of doing visual analytics is to build a mental and/or formal model of a certain piece of reality reflected in data. The purpose of the model may be to understand, to forecast or to control this piece of reality. Based on this model‐building perspective, we propose a detailed conceptual framework in which the visual analytics process is considered as a goal‐oriented workflow producing a model as a result. We demonstrate how this framework can be used for performing an analytical survey of the visual analytics research field and identifying the directions and areas where further research is needed.
Natalia V. Andrienko, Tim Lammarsch, Gennady L. Andrienko, Georg Fuchs, Daniel A. Keim, Silvia Miksch, Alexander Rind
Comput. Graph. Forum5
2018 Quality Metrics for Information Visualization
abstract
Abstract The visualization community has developed to date many intuitions and understandings of how to judge thequalityof views in visualizing data. The computation of a visualization's quality and usefulness ranges from measuring clutter and overlap, up to the existence and perception of specific (visual) patterns. This survey attempts to report, categorize and unify the diverse understandings and aims to establish a common vocabulary that will enable a wide audience to understand their differences and subtleties. For this purpose, we present a commonly applicable quality metric formalization that should detail and relate all constituting parts of a quality metric. We organize our corpus of reviewed research papers along the data types established in the information visualization community: multi‐ and high‐dimensional, relational, sequential, geospatial and text data. For each data type, we select the visualization subdomains in which quality metrics are an active research field and report their findings, reason on the underlying concepts, describe goals and outline the constraints and requirements. One central goal of this survey is to provide guidance on future research opportunities for the field and outline how different visualization communities could benefit from each other by applying or transferring knowledge to their respective subdomain. Additionally, we aim to motivate the visualization community to compare computed measures to the perception of humans.
Michael Behrisch 0001, Michael Blumenschein, Lin Shao 0001, Mennatallah El-Assady, Johannes Fuchs 0001, Daniel Seebacher, Alexandra Diehl, Ulrik Brandes, Hanspeter Pfister, Tobias Schreck, Daniel Weiskopf, Daniel A. Keim
Comput. Graph. Forum13
2018 ThreadReconstructor: Modeling Reply-Chains to Untangle Conversational Text through Visual Analytics
abstract
Abstract We present ThreadReconstructor, a visual analytics approach for detecting and analyzing the implicit conversational structure of discussions, e.g., in political debates and forums. Our work is motivated by the need to reveal and understand single threads in massive online conversations and verbatim text transcripts. We combine supervised and unsupervised machine learning models to generate a basic structure that is enriched by user‐defined queries and rule‐based heuristics. Depending on the data and tasks, users can modify and create various reconstruction models that are presented and compared in the visualization interface. Our tool enables the exploration of the generated threaded structures and the analysis of the untangled reply‐chains, comparing different models and their agreement. To understand the inner‐workings of the models, we visualize their decision spaces, including all considered candidate relations. In addition to a quantitative evaluation, we report qualitative feedback from an expert user study with four forum moderators and one machine learning expert, showing the effectiveness of our approach.
Mennatallah El-Assady, Rita Sevastjanova, Daniel A. Keim, Christopher Collins 0001
Comput. Graph. Forum3
2018 Urban Mobility Analysis With Mobile Network Data: A Visual Analytics Approach
abstract
Urban planning and intelligent transportation management are facing key challenges in today's ever more urbanized world. Providing the right tools to city planners is crucial to cope with these challenges. Data collected from citizens' mobile communication can be used as the foundation for such tools. These kinds of data can facilitate various analysis tasks, such as the extraction of human movement patterns or determining the urban dynamics of a city. City planners can closely monitor such patterns based on which strategic decisions can be taken to improve a city's infrastructure. In this paper, we introduce a novel visual analytics approach for pattern exploration and search in global system for mobile communications mobile networks. We define geospatial and matrix representations of data, which can be interactively navigated. The approach integrates data visualization with suitable data analysis algorithms, allowing to spatially and temporally compare mobile usage, identify regularities, as well as anomalies in daily mobility patterns across regions and user groups. As an extension to our visual analytics approach, we further introduce space-time prisms with uncertain markers to visually analyze the uncertainty of urban mobility patterns.
Hansi Senaratne, Manuel Müller, Michael Behrisch 0001, Felipe Lalanne, Javier Bustos-Jiménez, Jörn Schneidewind, Daniel A. Keim, Tobias Schreck
IEEE Trans. Intell. Transp. Syst.7
2018 Progressive Learning of Topic Modeling Parameters: A Visual Analytics Framework
abstract
Topic modeling algorithms are widely used to analyze the thematic composition of text corpora but remain difficult to interpret and adjust. Addressing these limitations, we present a modular visual analytics framework, tackling the understandability and adaptability of topic models through a user-driven reinforcement learning process which does not require a deep understanding of the underlying topic modeling algorithms. Given a document corpus, our approach initializes two algorithm configurations based on a parameter space analysis that enhances document separability. We abstract the model complexity in an interactive visual workspace for exploring the automatic matching results of two models, investigating topic summaries, analyzing parameter distributions, and reviewing documents. The main contribution of our work is an iterative decision-making technique in which users provide a document-based relevance feedback that allows the framework to converge to a user-endorsed topic distribution. We also report feedback from a two-stage study which shows that our technique results in topic model quality improvements on two independent measures.
Mennatallah El-Assady, Rita Sevastjanova, Fabian Sperrle, Daniel A. Keim, Christopher Collins 0001
IEEE Trans. Vis. Comput. Graph.4
2018 SOMFlow: Guided Exploratory Cluster Analysis with Self-Organizing Maps and Analytic Provenance
abstract
Clustering is a core building block for data analysis, aiming to extract otherwise hidden structures and relations from raw datasets, such as particular groups that can be effectively related, compared, and interpreted. A plethora of visual-interactive cluster analysis techniques has been proposed to date, however, arriving at useful clusterings often requires several rounds of user interactions to fine-tune the data preprocessing and algorithms. We present a multi-stage Visual Analytics (VA) approach for iterative cluster refinement together with an implementation (SOMFlow) that uses Self-Organizing Maps (SOM) to analyze time series data. It supports exploration by offering the analyst a visual platform to analyze intermediate results, adapt the underlying computations, iteratively partition the data, and to reflect previous analytical activities. The history of previous decisions is explicitly visualized within a flow graph, allowing to compare earlier cluster refinements and to explore relations. We further leverage quality and interestingness measures to guide the analyst in the discovery of useful patterns, relations, and data partitions. We conducted two pair analytics experiments together with a subject matter expert in speech intonation research to demonstrate that the approach is effective for interactive data analysis, supporting enhanced understanding of clustering results as well as the interactive process itself.
Dominik Sacha, Matthias Kraus 0002, Jürgen Bernard, Michael Behrisch 0001, Tobias Schreck, Yuki Asano 0003, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.7
2018 Bring It to the Pitch: Combining Video and Movement Data to Enhance Team Sport Analysis
abstract
Analysts in professional team sport regularly perform analysis to gain strategic and tactical insights into player and team behavior. Goals of team sport analysis regularly include identification of weaknesses of opposing teams, or assessing performance and improvement potential of a coached team. Current analysis workflows are typically based on the analysis of team videos. Also, analysts can rely on techniques from Information Visualization, to depict e.g., player or ball trajectories. However, video analysis is typically a time-consuming process, where the analyst needs to memorize and annotate scenes. In contrast, visualization typically relies on an abstract data model, often using abstract visual mappings, and is not directly linked to the observed movement context anymore. We propose a visual analytics system that tightly integrates team sport video recordings with abstract visualization of underlying trajectory data. We apply appropriate computer vision techniques to extract trajectory data from video input. Furthermore, we apply advanced trajectory and movement analysis techniques to derive relevant team sport analytic measures for region, event and player analysis in the case of soccer analysis. Our system seamlessly integrates video and visualization modalities, enabling analysts to draw on the advantages of both analysis forms. Several expert studies conducted with team sport analysts indicate the effectiveness of our integrated approach.
Manuel Stein, Halldór Janetzko, Andreas Lamprecht, Thorsten Breitkreutz, Philipp Zimmermann, Bastian Goldlücke, Tobias Schreck, Gennady L. Andrienko, Michael Grossniklaus, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.10
2018 Making machine intelligence less scary for criminal analysts: reflections on designing a visual comparative case analysis tool
Wolfgang Jentner, Dominik Sacha, Florian Stoffel, Geoffrey P. Ellis, Leishi Zhang, Daniel A. Keim
Vis. Comput.6
2017 Visual Analytics and Similarity Search: Concepts and Challenges for Effective Retrieval Considering Users, Tasks, and Data
Daniel Seebacher, Johannes Häußler, Manuel Stein, Halldór Janetzko, Tobias Schreck, Daniel A. Keim
SISAP6
2017 On the Impact of the Medium in the Effectiveness of 3D Software Visualizations
abstract
Many visualizations have proven to be effective in supporting various software related tasks. Although multiple media can be used to display a visualization, the standard computer screen is used the most. We hypothesize that the medium has a role in their effectiveness. We investigate our hypotheses by conducting a controlled user experiment. In the experiment we focus on the 3D city visualization technique used for software comprehension tasks. We deploy 3D city visualizations across a standard computer screen (SCS), an immersive 3D environment (I3D), and a physical 3D printed model (P3D). We asked twenty-seven participants (whom we divided in three groups for each medium) to visualize software systems of various sizes, solve a set of uniform comprehension tasks, and complete a questionnaire. We measured the effectiveness of visualizations in terms of performance, recollection, and user experience. We found that even though developers using P3D required the least time to identify outliers, they perceived the least difficulty when visualizing systems based on SCS. Moreover, developers using I3D obtained the highest recollection.
Leonel Merino, Johannes Fuchs 0001, Michael Blumenschein, Craig Anslow, Mohammad Ghafari, Oscar Nierstrasz, Michael Behrisch 0001, Daniel A. Keim
VISSOFT8
2017 NEREx: Named-Entity Relationship Exploration in Multi-Party Conversations
abstract
Abstract We present NEREx, an interactive visual analytics approach for the exploratory analysis of verbatim conversational transcripts. By revealing different perspectives on multi‐party conversations, NEREx gives an entry point for the analysis through high‐level overviews and provides mechanisms to form and verify hypotheses through linked detail‐views. Using a tailored named‐entity extraction, we abstract important entities into ten categories and extract their relations with a distance‐restricted entity‐relationship model. This model complies with the often ungrammatical structure of verbatim transcripts, relating two entities if they are present in the same sentence within a small distance window. Our tool enables the exploratory analysis of multi‐party conversations using several linked views that reveal thematic and temporal structures in the text. In addition to distant‐reading, we integrated close‐reading views for a text‐level investigation process. Beyond the exploratory and temporal analysis of conversations, NEREx helps users generate and validate hypotheses and perform comparative analyses of multiple conversations. We demonstrate the applicability of our approach on real‐world data from the 2016 U.S. Presidential Debates through a qualitative study with three domain experts from political science.
Mennatallah El-Assady, Rita Sevastjanova, Bela Gipp, Daniel A. Keim, Christopher Collins 0001
Comput. Graph. Forum4
2017 Dynamic Visual Abstraction of Soccer Movement
abstract
Abstract Trajectory‐based visualization of coordinated movement data within a bounded area, such as player and ball movement within a soccer pitch, can easily result in visual crossings, overplotting, and clutter. Trajectory abstraction can help to cope with these issues, but it is a challenging problem to select the right level of abstraction (LoA) for a given data set and analysis task. We present a novel dynamic approach that combines trajectory simplification and clustering techniques with the goal to support interpretation and understanding of movement patterns. Our technique provides smooth transitions between different abstraction types that can be computed dynamically and on‐the‐fly. This enables the analyst to effectively navigate and explore the space of possible abstractions in large trajectory data sets. Additionally, we provide a proof of concept for supporting the analyst in determining the LoA semi‐automatically with a recommender system. Our approach is illustrated and evaluated by case studies, quantitative measures, and expert feedback. We further demonstrate that it allows analysts to solve a variety of analysis tasks in the domain of soccer.
Dominik Sacha, F. Al-amoody, Manuel Stein, Tobias Schreck, Daniel A. Keim, Gennady L. Andrienko, Halldór Janetzko
Comput. Graph. Forum5
2017 Interactive Ambiguity Resolution of Named Entities in Fictional Literature
abstract
Abstract Named entity recognition (NER) denotes the task to detect entities and their corresponding classes, such as person or location, in unstructured text data. For most applications, state of the art NER software is producing reasonable results. However, as a consequence of the methodological limitations and the well‐known pitfalls when analyzing natural language data, the NER results are likely to contain ambiguities. In this paper, we present an interactive NER ambiguity resolution technique, which enables users to create (post‐processing) rules for named entity recognition data based on the content and entity context of the analyzed documents. We specifically address the problem that in use‐cases where ambiguities are problematic, such as the attribution of fictional characters with traits, it is often unfeasible to train models on custom data to improve state of the art NER software. We derive an iterative process model for improving NER results, show an interactive NER ambiguity resolution prototype, illustrate our approach with contemporary literature, and discuss our work and future research.
Florian Stoffel, Wolfgang Jentner, Michael Behrisch 0001, Johannes Fuchs 0001, Daniel A. Keim
Comput. Graph. Forum5
2017 What you see is what you can change: Human-centered machine learning by interactive visualization
Dominik Sacha, Michael Sedlmair, Leishi Zhang, John A. Lee 0001, Jaakko Peltonen, Daniel Weiskopf, Stephen C. North, Daniel A. Keim
Neurocomputing8
2017 A Systematic Review of Experimental Studies on Data Glyphs
abstract
We systematically reviewed 64 user-study papers on data glyphs to help researchers and practitioners gain an informed understanding of tradeoffs in the glyph design space. The glyphs we consider are individual representations of multi-dimensional data points, often meant to be shown in small-multiple settings. Over the past 60 years many different glyph designs were proposed and many of these designs have been subjected to perceptual or comparative evaluations. Yet, a systematic overview of the types of glyphs and design variations tested, the tasks under which they were analyzed, or even the study goals and results does not yet exist. In this paper we provide such an overview by systematically sampling and tabulating the literature on data glyph studies, listing their designs, questions, data, and tasks. In addition we present a concise overview of the types of glyphs and their design characteristics analyzed by researchers in the past, and a synthesis of the study results. Based on our meta analysis of all results we further contribute a set of design implications and a discussion on open research directions.
Johannes Fuchs 0001, Petra Isenberg, Anastasia Bezerianos, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.4
2017 Visual Interaction with Dimensionality Reduction: A Structured Literature Analysis
abstract
Dimensionality Reduction (DR) is a core building block in visualizing multidimensional data. For DR techniques to be useful in exploratory data analysis, they need to be adapted to human needs and domain-specific problems, ideally, interactively, and on-the-fly. Many visual analytics systems have already demonstrated the benefits of tightly integrating DR with interactive visualizations. Nevertheless, a general, structured understanding of this integration is missing. To address this, we systematically studied the visual analytics and visualization literature to investigate how analysts interact with automatic DR techniques. The results reveal seven common interaction scenarios that are amenable to interactive control such as specifying algorithmic constraints, selecting relevant features, or choosing among several DR algorithms. We investigate specific implementations of visual analysis systems integrating DR, and analyze ways that other machine learning methods have been combined with DR. Summarizing the results in a "human in the loop" process model provides a general lens for the evaluation of visual interactive DR systems. We apply the proposed model to study and classify several systems previously described in the literature, and to derive future research opportunities.
Dominik Sacha, Leishi Zhang, Michael Sedlmair, John A. Lee 0001, Jaakko Peltonen, Daniel Weiskopf, Stephen C. North, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.8
2016 Human-centered machine learning through interactive visualization: review and open challenges
Dominik Sacha, Michael Sedlmair, Leishi Zhang, John A. Lee 0001, Daniel Weiskopf, Stephen C. North, Daniel A. Keim
ESANN7
2016 ConToVi: Multi-Party Conversation Exploration using Topic-Space Views
abstract
Abstract We introduce a novel visual analytics approach to analyze speaker behavior patterns in multi‐party conversations. We propose Topic‐Space Views to track the movement of speakers across the thematic landscape of a conversation. Our tool is designed to assist political science scholars in exploring the dynamics of a conversation over time to generate and prove hypotheses about speaker interactions and behavior patterns. Moreover, we introduce a glyph‐based representation for each speaker turn based on linguistic and statistical cues to abstract relevant text features. We present animated views for exploring the general behavior and interactions of speakers over time and interactive steady visualizations for the detailed analysis of a selection of speakers. Using a visual sedimentation metaphor we enable the analysts to track subtle changes in the flow of a conversation over time while keeping an overview of all past speaker turns. We evaluate our approach on real‐world datasets and the results have been insightful to our domain experts.
Mennatallah El-Assady, Valentin Gold, Carmela Acevedo, Christopher Collins 0001, Daniel A. Keim
Comput. Graph. Forum5
2016 A Survey on Visual Analytics of Social Media Data
abstract
The unprecedented availability of social media data offers substantial opportunities for data owners, system operators, solution providers, and end users to explore and understand social dynamics. However, the exponential growth in the volume, velocity, and variability of social media data prevents people from fully utilizing such data. Visual analytics, which is an emerging research direction, has received considerable attention in recent years. Many visual analytics methods have been proposed across disciplines to understand large-scale structured and unstructured social media data. This objective, however, also poses significant challenges for researchers to obtain a comprehensive picture of the area, understand research challenges, and develop new techniques. In this paper, we present a comprehensive survey to characterize this fast-growing area and summarize the state-of-the-art techniques for analyzing social media data. In particular, we classify existing techniques into two categories: gathering information and understanding user behaviors. We aim to provide a clear overview of the research area through the established taxonomy. We then explore the design space and identify the research trends. Finally, we discuss challenges and open questions for future studies.
Yingcai Wu, Nan Cao 0001, David Gotz, Yap-Peng Tan, Daniel A. Keim
IEEE Trans. Multim.5
2016 Temporal MDS Plots for Analysis of Multivariate Data
abstract
Multivariate time series data can be found in many application domains. Examples include data from computer networks, healthcare, social networks, or financial markets. Often, patterns in such data evolve over time among multiple dimensions and are hard to detect. Dimensionality reduction methods such as PCA and MDS allow analysis and visualization of multivariate data, but per se do not provide means to explore multivariate patterns over time. We propose Temporal Multidimensional Scaling (TMDS), a novel visualization technique that computes temporal one-dimensional MDS plots for multivariate data which evolve over time. Using a sliding window approach, MDS is computed for each data window separately, and the results are plotted sequentially along the time axis, taking care of plot alignment. Our TMDS plots enable visual identification of patterns based on multidimensional similarity of the data evolving over time. We demonstrate the usefulness of our approach in the field of network security and show in two case studies how users can iteratively explore the data to identify previously unknown, temporally evolving patterns.
Dominik Jäckle, Fabian Fischer 0001, Tobias Schreck, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.4
2016 The Role of Uncertainty, Awareness, and Trust in Visual Analytics
abstract
Visual analytics supports humans in generating knowledge from large and often complex datasets. Evidence is collected, collated and cross-linked with our existing knowledge. In the process, a myriad of analytical and visualisation techniques are employed to generate a visual representation of the data. These often introduce their own uncertainties, in addition to the ones inherent in the data, and these propagated and compounded uncertainties can result in impaired decision making. The user's confidence or trust in the results depends on the extent of user's awareness of the underlying uncertainties generated on the system side. This paper unpacks the uncertainties that propagate through visual analytics systems, illustrates how human's perceptual and cognitive biases influence the user's awareness of such uncertainties, and how this affects the user's trust building. The knowledge generation model for visual analytics is used to provide a terminology and framework to discuss the consequences of these aspects in knowledge construction and though examples, machine uncertainty is compared to human trust measures with provenance. Furthermore, guidelines for the design of uncertainty-aware systems are presented that can aid the user in better decision making.
Dominik Sacha, Hansi Senaratne, Bum Chul Kwon, Geoffrey P. Ellis, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.5
2015 Visual Analysis of Car Fleet Trajectories to Find Representative Routes for Automotive Research
abstract
Testing is an important and wide spread practice in the development of automotive components. For the design of test methods two types of input data are often considered: (1) load data gathered from real life vehicle fleets, and (2) information of the driving routes based on road features. The development of new technologies is though complicated not only by the need to join those two data sources, but also by the too limited knowledge of the parameters and their useful combinations. As a result, information about representative driving profiles is needed. To address these problems we present a visual analytics approach for analyzing multivariate trajectories as a combination of vehicle's location and road elevation data. Our system combines trajectory clustering, interval-based user-driven trip segmentation, and frequent sequences analysis, supported by contingency table and interval-based Parallel Coordinates visualization and enables the expert user to find representative driving profiles for the definition of very compact test courses.
David Spretke, Manuel Stein, Lyubka Sharalieva, Alexander Warta, Valentin Licht, Tobias Schreck, Daniel A. Keim
IV7
2015 Subspace Nearest Neighbor Search - Problem Statement, Approaches, and Discussion - Position Paper
Michael Blumenschein, Michael Behrisch 0001, Ines Färber, Michael Sedlmair, Tobias Schreck, Thomas Seidl 0001, Daniel A. Keim
SISAP7
2015 Visual Analytics for Exploring Local Impact of Air Traffic
abstract
Abstract The environmental and noise impact of airports often causes extensive political discussion which in some cases even lead to transnational tensions. Analyzing local approach and departure patterns around an airport is difficult since it depends on a variety of complex variables like weather, local and general regulations and many more. Yet, understanding these movements and the expected amount of flights during arrival and departure is of great interest to both casual and expert users, as planes have a higher impact on the areas beneath during these phases. We present a Visual Analytics framework that enables users to develop an understanding of local flight behavior through visual exploration of historical data and interactive manipulation of prediction models with direct feedback, as well as a classification quality visualization using a random noise metaphor. We showcase our approach using real world data from the Zurich International Airport region, where aircraft noise has led to an ongoing conflict between Germany and Switzerland. The use cases, findings and expert feedback demonstrate how our approach helps in understanding the situation and to substantiate the otherwise often subjective discourse on the topic.
Juri Buchmüller, Halldór Janetzko, Gennady L. Andrienko, Natalia V. Andrienko, Georg Fuchs, Daniel A. Keim
Comput. Graph. Forum6
2015 Efficient Contrast Effect Compensation with Personalized Perception Models
abstract
Abstract Color is one of the most effective visual variables and is frequently used to encode metric quantities. Contrast effects are considered harmful in data visualizations since they significantly bias our perception of colors. For instance, a gray patch appears brighter on a black background than on a white background. Accordingly, the perception of color‐encoded data items depends on the surround in the rendered visualization. A method that compensates for contrast effects has been presented previously, which significantly improves the users’ accuracy in reading and comparing color encoded data. The method utilizes established perception models to compensate for contrast effects, assuming an average human observer. In this paper, we provide experiments that show a significant difference in the perception of users. We introduce methods to personalize contrast effect compensation and show that this outperforms the original method with a user study. We, further, overcome the major limitation of the original method, which is a runtime of several minutes. With the use of efficient optimization and surrogate models, we are able to reduce runtime to milliseconds, making the method applicable in interactive visualizations.
Sebastian Mittelstädt, Daniel A. Keim
Comput. Graph. Forum2
2015 Interactive feature space extension for multidimensional data projection
Daniel Pérez 0001, Leishi Zhang, Matthias Schäfer 0001, Tobias Schreck, Daniel A. Keim, Ignacio Díaz Blanco
Neurocomputing5
2014 Advanced visual analytics interfaces for adverse drug event detection
abstract
Adverse reactions to drugs are a major public health care issue. Currently, the Food and Drug Administration (FDA) publishes quarterly reports that typically contain on the order of 200,000 adverse incidents. In such numerous incidents, low frequency events that are clinically highly significant often remain undetected. In this paper, we introduce a visual analytics system to solve this problem using (1) high scalable interfaces for analyzing correlations between a number of complex variables (e.g., drug and reaction); (2) enhanced statistical computations and interactive relevance filters to quickly identify significant events including those with a low frequency; and (3) a tight integration of expert knowledge for detecting and validating adverse drug events. We applied these techniques to the FDA Adverse Event Reporting System and were able to identify important adverse drug events, such as the known association of the drug Avandia with myocardial infarction and Seroquel with diabetes mellitus, as well as low frequency events such as the association of Boniva with femur fracture. In our evaluation, we found over 90% of the adverse drug events that were published in the Institute for Safe Medication Practices (ISMP) reports from 2009 to 2012. In addition, our domain expert was able to identify some previously unknown adverse drug events.
Sebastian Mittelstädt, Ming C. Hao, Umeshwar Dayal, Meichun Hsu, Joseph Terdiman, Daniel A. Keim
AVI6
2014 NStreamAware: real-time visual analytics for data streams to enhance situational awareness
abstract
The analysis of data streams is important in many security-related domains to gain situational awareness. To provide monitoring and visual analysis of such data streams, we propose a system, called NStreamAware, that uses modern distributed processing technologies to analyze streams using stream slices, which are presented to analysts in a web-based visual analytics application, called NVisAware. Furthermore, we visually guide the user in the feature selection process to summarize the slices to focus on the most interesting parts of the stream based on introduced expert knowledge of the analyst. We show through case studies, how the system can be used to gain situational awareness and eventually enhance network security. Furthermore, we apply the system to a social media data stream to compete in an international challenge to evaluate the applicability of our approach to other domains.
Fabian Fischer 0001, Daniel A. Keim
VizSEC2
2014 Anomaly detection for visual analytics of power consumption data
Halldór Janetzko, Florian Stoffel, Sebastian Mittelstädt, Daniel A. Keim
Comput. Graph.4
2014 Visual Analysis of Sets of Heterogeneous Matrices Using Projection-Based Distance Functions and Semantic Zoom
abstract
Abstract Matrix visualization is an established technique in the analysis of relational data. It is applicable to large, dense networks, where node‐link representations may not be effective. Recently, domains have emerged in which the comparative analysis of sets of matrices of potentially varying size is relevant. For example, to monitor computer network traffic a dynamic set of hosts and their peer‐to‐peer connections on different ports must be analysed. A matrix visualization focused on the display of one matrix at a time cannot cope with this task. We address the research problem of the visual analysis of sets of matrices. We present a technique for comparing matrices of potentially varying size. Our approach considers the rows and/or columns of a matrix as the basic elements of the analysis. We project these vectors for pairs of matrices into a low‐dimensional space which is used as the reference to compare matrices and identify relationships among them. Bipartite graph matching is applied on the projected elements to compute a measure of distance. A key advantage of this measure is that it can be interpreted and manipulated as a visual distance function, and serves as a comprehensible basis for ranking, clustering and comparison in sets of matrices. We present an interactive system in which users may explore the matrix distances and understand potential differences in a set of matrices. A flexible semantic zoom mechanism enables users to navigate through sets of matrices and identify patterns at different levels of detail. We demonstrate the effectiveness of our approach through a case study and provide a technical evaluation to illustrate its strengths.
Michael Behrisch 0001, James Davey, Fabian Fischer 0001, Olivier Thonnard, Tobias Schreck, Daniel A. Keim, Jörn Kohlhammer
Comput. Graph. Forum6
2014 Methods for Compensating Contrast Effects in Information Visualization
abstract
Abstract Color, as one of the most effective visual variables, is used in many techniques to encode and group data points according to different features. Relations between features and groups appear as visual patterns in the visualization. However, optical illusions may bias the perception at the first level of the analysis process. For instance, in pixel‐based visualizations contrast effects make pixels appear brighter if surrounded by a darker area, which distorts the encoded metric quantity of the data points. Even if we are aware of these perceptual issues, our visual cognition system is not able to compensate these effects accurately. To overcome this limitation, we present a color optimization algorithm based on perceptual metrics and color perception models to reduce physiological contrast or color effects. We evaluate our technique with a user study and find that the technique doubles the accuracy of users comparing and estimating color encoded data values. Since the presented technique can be used in any application without adaption to the visualization itself, we are able to demonstrate its effectiveness on data visualizations in different domains.
Sebastian Mittelstädt, Andreas Stoffel, Daniel A. Keim
Comput. Graph. Forum3
2014 Visual Analysis of Time-Series Similarities for Anomaly Detection in Sensor Networks
abstract
Abstract We present a system to analyze time‐series data in sensor networks. Our approach supports exploratory tasks for the comparison of univariate, geo‐referenced sensor data, in particular for anomaly detection. We split the recordings into fixed‐length patterns and show them in order to compare them over time and space using two linked views. Apart from geo‐based comparison across sensors we also support different temporal patterns to discover seasonal effects, anomalies and periodicities. The methods we use are best practices in the information visualization domain. They cover the daily, the weekly and seasonal and patterns of the data. Daily patterns can be analyzed in a clustering‐based view, weekly patterns in a calendar‐based view and seasonal patters in a projection‐based view. The connectivity of the sensors can be analyzed through a dedicated topological network view. We assist the domain expert with interaction techniques to make the results understandable. As a result, the user can identify and analyze erroneous and suspicious measurements in the network. A case study with a domain expert verified the usefulness of our approach.
Martin Steiger, Jürgen Bernard, Sebastian Mittelstädt, Hendrik Lücke-Tieke, Daniel A. Keim, Thorsten May, Jörn Kohlhammer
Comput. Graph. Forum5
2014 Introduction to the Special Issue on Interactive Computational Visual Analytics
abstract
This editorial introduction describes the aims and scope of ACM Transactions on Interactive Intelligent Systems 's special issue on interactive computational visual analytics. It explains why visual analytics is crucial to the growing needs surrounding data analysis, and it shows how the four articles selected for this issue reflect this theme.
Remco Chang, David S. Ebert, Daniel A. Keim
ACM Trans. Interact. Intell. Syst.3
2014 Knowledge Generation Model for Visual Analytics
abstract
Visual analytics enables us to analyze huge information spaces in order to support complex decision making and data exploration. Humans play a central role in generating knowledge from the snippets of evidence emerging from visual data analysis. Although prior research provides frameworks that generalize this process, their scope is often narrowly focused so they do not encompass different perspectives at different levels. This paper proposes a knowledge generation model for visual analytics that ties together these diverse frameworks, yet retains previously developed models (e.g., KDD process) to describe individual segments of the overall visual analytic processes. To test its utility, a real world visual analytics system is compared against the model, demonstrating that the knowledge generation process model provides a useful guideline when developing and evaluating such systems. The model is used to effectively compare different data analysis systems. Furthermore, the model provides a common language and description of visual analytic processes, which can be used for communication between researchers. At the end, our model reflects areas of research that future researchers can embark on.
Dominik Sacha, Andreas Stoffel, Florian Stoffel, Bum Chul Kwon, Geoffrey P. Ellis, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.6
2013 Visualization of streaming data: Observing change and context in information visualization techniques
abstract
Visualizing data streams poses numerous challenges in the data, image and user space. In the era of big data, we need incremental visualization methods that will allow the analysts to explore data faster and help them make important decisions on time. In this paper, we have reviewed several well-known information visualization methods that are commonly used to visualize static datasets and analyzed their degrees of freedom. By observing which independent visual variables can change in each method, we described how these changes are related to the attribute and structure changes that can occur in the data stream. Most of the changes in the data stream lead to potential loss of temporal and relational context between the new data and the past data. We present potential directions for measuring the amount of change and loss of context by reviewing related work and identify open issues for future work in this domain.
Milos Krstajic, Daniel A. Keim
IEEE BigData2
2013 Finding anomalies in time-series using visual correlation for interactive root cause analysis
abstract
Monitoring computer networks often includes gathering vast amounts of time-series data from thousands of computer systems and network devices. Threshold alerting is easy to accomplish with state-of-the-art technologies. However, to find correlations and similar behaviors between the different devices is challenging. We developed a visual analytics application to tackle this challenge by integrating similarity models and analytics combined with well-known, but task-adapted, time-series visualizations. We show in a case study, how this system can be used to visually identify correlations and anomalies in large data sets and identify and investigate security-related events.
Florian Stoffel, Fabian Fischer 0001, Daniel A. Keim
VizSEC3
2013 Fingerprint Matrices: Uncovering the dynamics of social networks in prose literature
abstract
Abstract In prose literature often complex dynamics of interpersonal relationships can be observed between the different characters. Traditionally, node‐link diagrams are used to depict the social network of a novel. However, static graphs can only visualize the overall social network structure but not the development of the networks over the course of the story, while dynamic graphs have the serious problem that there are many sudden changes between different portions of the overall social network. In this paper we explore means to show the relationships between the characters of a plot and at the same time their development over the course of a novel. Based on a careful exploration of the design space, we suggest a new visualization technique called Fingerprint Matrices. A case study exemplifies the usage of Fingerprint Matrices and shows that they are an effective means to analyze prose literature with respect to the development of relationships between the different characters.
Daniela Oelke, Dimitrios Kokkinakis, Daniel A. Keim
Comput. Graph. Forum3
2013 Preface: Intelligent interactive data visualization
Barbara Hammer, Daniel A. Keim, Neil D. Lawrence, Guy Lebanon
Data Min. Knowl. Discov.2
2013 A GIS-based decision support system for hotel room rate estimation and temporal price prediction: The hotel brokers' context
Slava Kisilevich, Daniel A. Keim, Lior Rokach
Decis. Support Syst.2
2012 Solving Problems with Visual Analytics: Challenges and Applications
Daniel A. Keim
ECML/PKDD (1)1
2012 The World's Languages Explorer: Visual Analysis of Language Features in Genealogical and Areal Contexts
abstract
Abstract This paper presents a novel Visual Analytics approach that helps linguistic researchers to explore the world's languages with respect to several important tasks: (1) The comparison of manually and automatically extracted language features across languages and within the context of language genealogy, (2) the exploration of interrelations among several of such features as well as their homogeneity and heterogeneity within subtrees of the language genealogy, and (3) the exploration of genealogical and areal influences on the features. We introduce theWorld'sLanguagesExplorer, which provides the required functionalities in one single Visual Analytics environment. Contributions are made for different parts of the system: We introduce an extended Sunburst visualization whose so‐called feature‐rings allow for a cross‐comparison of a large number of features at once, within the hierarchical context of the language genealogy. We suggest a mapping of homogeneity measures to all levels of the hierarchy. In addition, we suggest an integration of information from the areal data space into the hierarchical data space. With our approach we bring Visual Analytics research to a new application field, namely Historical Comparative Linguistics, and Linguistic and Areal Typology. Finally, we provide evidence of the good performance of our system in this area through two application case studies conducted by domain experts.
Christian Rohrdantz, Michael Blumenschein, Thomas Mayer 0001, Bernhard Wälchli, Daniel A. Keim
Comput. Graph. Forum5
2012 Document Thumbnails with Variable Text Scaling
abstract
Abstract Document reader applications usually offer an overview of the layout for each page as thumbnail views. Reading the text in these becomes impossible when the font size becomes very small. We improve the readability of these thumbnails using a distortion method, which retains a readable font size of interesting text while shrinking less interesting text further. In contrast to existing approaches, our method preserves the global layout of a page and is able to show context around important terms. We evaluate our technique and show application examples.
Andreas Stoffel, Hendrik Strobelt, Oliver Deussen, Daniel A. Keim
Comput. Graph. Forum4
2012 Rolled-out Wordles: A Heuristic Method for Overlap Removal of 2D Data Representatives
abstract
Abstract When representing 2D data points with spacious objects such as labels, overlap can occur. We present a simple algorithm which modifies the (Mani‐) Wordle idea with scan‐line based techniques to allow a better placement. We give an introduction to common placement techniques from different fields and compare our method to these techniques w.r.t. euclidean displacement, changes in orthogonal ordering as well as shape and size preservation. Especially in dense scenarios our method preserves the overall shape better than known techniques and allows a good trade‐off between the other measures. Applications on real world data are given and discussed.
Hendrik Strobelt, Marc Spicker, Andreas Stoffel, Daniel A. Keim, Oliver Deussen
Comput. Graph. Forum4
2012 Improving 3D similarity search by enhancing and combining 3D descriptors
Benjamin Bustos, Tobias Schreck, Michael Walter 0001, Juan Manuel Barrios, Matthias Schäfer 0001, Daniel A. Keim
Multim. Tools Appl.6
2012 Feature-Based Visual Sentiment Analysis of Text Document Streams
abstract
This article describes automatic methods and interactive visualizations that are tightly coupled with the goal to enable users to detect interesting portions of text document streams. In this scenario the interestingness is derived from the sentiment, temporal density, and context coherence that comments about features for different targets (e.g., persons, institutions, product attributes, topics, etc.) have. Contributions are made at different stages of the visual analytics pipeline, including novel ways to visualize salient temporal accumulations for further exploration. Moreover, based on the visualization, an automatic algorithm aims to detect and preselect interesting time interval patterns for different features in order to guide analysts. The main target group for the suggested methods are business analysts who want to explore time-stamped customer feedback to detect critical issues. Finally, application case studies on two different datasets and scenarios are conducted and an extensive evaluation is provided for the presented intelligent visual interface for feature-based sentiment exploration over time.
Christian Rohrdantz, Ming C. Hao, Umeshwar Dayal, Lars-Erik Haug, Daniel A. Keim
ACM Trans. Intell. Syst. Technol.5
2012 EventRiver: Visually Exploring Text Collections with Temporal References
abstract
Many text collections with temporal references, such as news corpora and weblogs, are generated to report and discuss real life events. Thus, event-related tasks, such as detecting real life events that drive the generation of the text documents, tracking event evolutions, and investigating reports and commentaries about events of interest, are important when exploring such text collections. To incorporate and leverage human efforts in conducting such tasks, we propose a novel visual analytics approach named EventRiver. EventRiver integrates event-based automated text analysis and visualization to reveal the events motivating the text generation and the long term stories they construct. On the visualization, users can interactively conduct tasks such as event browsing, tracking, association, and investigation. A working prototype of EventRiver has been implemented for exploring news corpora. A set of case studies, experiments, and a preliminary user test have been conducted to evaluate its effectiveness and efficiency.
Dongning Luo, Jing Yang 0001, Milos Krstajic, William Ribarsky, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.5
2012 Visual Readability Analysis: How to Make Your Writings Easier to Read
abstract
We present a tool that is specifically designed to support a writer in revising a draft version of a document. In addition to showing which paragraphs and sentences are difficult to read and understand, we assist the reader in understanding why this is the case. This requires features that are expressive predictors of readability, and are also semantically understandable. In the first part of the paper, we, therefore, discuss a semiautomatic feature selection approach that is used to choose appropriate measures from a collection of 141 candidate readability features. In the second part, we present the visual analysis tool VisRA, which allows the user to analyze the feature values across the text and within single sentences. Users can choose between different visual representations accounting for differences in the size of the documents and the availability of information about the physical and logical layout of the documents. We put special emphasis on providing as much transparency as possible to ensure that the user can purposefully improve the readability of a sentence. Several case studies are presented that show the wide range of applicability of our tool. Furthermore, an in-depth evaluation assesses the quality of the measure and investigates how well users do in revising a text with the help of the tool.
Daniela Oelke, David Spretke, Andreas Stoffel, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.4
2011 Geodesic distances for web document clustering
abstract
While traditional distance measures are often capable of properly describing similarity between objects, in some application areas there is still potential to fine-tune these measures with additional information provided in the data sets. In this work we combine such traditional distance measures for document analysis with link information between documents to improve clustering results. In particular, we test the effectiveness of geodesic distances as similarity measures under the space assumption of spherical geometry in a 0-sphere. Our proposed distance measure is thus a combination of the cosine distance of the term-document matrix and some curvature values in the geodesic distance formula. To estimate these curvature values, we calculate clustering coefficient values for every document from the link graph of the data set and increase their distinctiveness by means of a heuristic as these clustering coefficient values are rough estimates of the curvatures. To evaluate our work, we perform clustering tests with the k-means algorithm on the English Wikipedia hyperlinked data set with both traditional cosine distance and our proposed geodesic distance. The effectiveness of our approach is measured by computing micro-precision values of the clusters based on the provided categorical information of each article.
Selma Tekir, Florian Mansmann, Daniel A. Keim
CIDM3
2011 A Visual Analytics Approach for Peak-Preserving Prediction of Large Seasonal Time Series
abstract
Abstract Time series prediction methods are used on a daily basis by analysts for making important decisions. Most of these methods use some variant of moving averages to reduce the number of data points before prediction. However, to reach a good prediction in certain applications (e.g., power consumption time series in data centers) it is important to preserve peaks and their patterns. In this paper, we introduce automated peak‐preserving smoothing and prediction algorithms, enabling a reliable long term prediction for seasonal data, and combine them with an advanced visual interface: (1) using high resolution cell‐based time series to explore seasonal patterns, (2) adding new visual interaction techniques (multi‐scaling, slider, and brushing & linking) to incorporate human expert knowledge, and (3) providing both new visual accuracy color indicators for validating the predicted results and certainty bands communicating the uncertainty of the prediction. We have integrated these techniques into a well‐fitted solution to support the prediction process, and applied and evaluated the approach to predict both power consumption and server utilization in data centers with 70–80% accuracy.
Ming C. Hao, Halldór Janetzko, Sebastian Mittelstädt, W. Hill, Umeshwar Dayal, Daniel A. Keim, Manish Marwah, Ratnesh K. Sharma
Comput. Graph. Forum6
2011 Visual Boosting in Pixel-based Visualizations
abstract
Abstract Pixel‐based visualizations have become popular, because they are capable of displaying large amounts of data and at the same time provide many details. However, pixel‐based visualizations are only effective if the data set is not sparse and the data distribution not random. Single pixels – no matter if they are in an empty area or in the middle of a large area of differently colored pixels – are perceptually difficult to discern and may therefore easily be missed. Furthermore, trends and interesting passages may be camouflaged in the sea of details. In this paper we compare different approaches for visual boosting in pixel‐based visualizations. Several boosting techniques such as halos, background coloring, distortion, and hatching are discussed and assessed with respect to their effectiveness in boosting single pixels, trends, and interesting passages. Application examples from three different domains (document analysis, genome analysis, and geospatial analysis) show the general applicability of the techniques and the derived guidelines.
Daniela Oelke, Halldór Janetzko, Svenja Simon, Klaus Neuhaus, Daniel A. Keim
Comput. Graph. Forum5
2011 Quality Metrics in High-Dimensional Data Visualization: An Overview and Systematization
abstract
In this paper, we present a systematization of techniques that use quality metrics to help in the visual exploration of meaningful patterns in high-dimensional data. In a number of recent papers, different quality metrics are proposed to automate the demanding search through large spaces of alternative visualizations (e.g., alternative projections or ordering), allowing the user to concentrate on the most promising visualizations suggested by the quality metrics. Over the last decade, this approach has witnessed a remarkable development but few reflections exist on how these methods are related to each other and how the approach can be developed further. For this purpose, we provide an overview of approaches that use quality metrics in high-dimensional data visualization and propose a systematization based on a thorough literature review. We carefully analyze the papers and derive a set of factors for discriminating the quality metrics, visualization techniques, and the process itself. The process is described through a reworked version of the well-known information visualization pipeline. We demonstrate the usefulness of our model by applying it to several existing approaches that use quality metrics, and we provide reflections on implications of our model for future research.
Enrico Bertini, Andrada Tatu, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.3
2011 VisWeek Keynote Address
abstract
This talk will discuss the role of visual thinking in scientific discovery and technological invention. Visual thinking uses picture-like representations as internal mental models or as external depictions such as diagrams. The first part of the talk will analyze the role of visual thinking in 100 great discoveries and 100 great inventions. The second part will discuss the contribution of visual thinking to developing new theories in the social sciences based on advances in cognitive science. Cognitive-affective mapping is a new technique for visualizing the role of emotion in social cognition. EMPATHICA is a new graphical system for resolving conflicts by increasing empathy using cognitive-affective maps.
Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.1
2011 CloudLines: Compact Display of Event Episodes in Multiple Time-Series
abstract
We propose incremental logarithmic time-series technique as a way to deal with time-based representations of large and dynamic event data sets in limited space. Modern data visualization problems in the domains of news analysis, network security and financial applications, require visual analysis of incremental data, which poses specific challenges that are normally not solved by static visualizations. The incremental nature of the data implies that visualizations have to necessarily change their content and still provide comprehensible representations. In particular, in this paper we deal with the need to keep an eye on recent events together with providing a context on the past and to make relevant patterns accessible at any scale. Our technique adapts to the incoming data by taking care of the rate at which data items occur and by using a decay function to let the items fade away according to their relevance. Since access to details is also important, we also provide a novel distortion magnifying lens technique which takes into account the distortions introduced by the logarithmic time scale to augment readability in selected areas of interest. We demonstrate the validity of our techniques by applying them on incremental data coming from online news streams in different time frames.
Milos Krstajic, Enrico Bertini, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.3
2011 Automated Analytical Methods to Support Visual Exploration of High-Dimensional Data
abstract
Visual exploration of multivariate data typically requires projection onto lower dimensional representations. The number of possible representations grows rapidly with the number of dimensions, and manual exploration quickly becomes ineffective or even unfeasible. This paper proposes automatic analysis methods to extract potentially relevant visual structures from a set of candidate visualizations. Based on features, the visualizations are ranked in accordance with a specified user task. The user is provided with a manageable number of potentially useful candidate visualizations, which can be used as a starting point for interactive data analysis. This can effectively ease the task of finding truly useful visualizations and potentially speed up the data exploration task. In this paper, we present ranking measures for class-based as well as non-class-based scatterplots and parallel coordinates visualizations. The proposed analysis methods are evaluated on different data sets.
Andrada Tatu, Georgia Albuquerque, Martin Eisemann, Peter Bak, Holger Theisel, Marcus A. Magnor, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.7
2010 Advanced visual analytics interfaces
abstract
Advanced visual interfaces, like the ones found in information visualization, intend to offer a view on abstract data spaces to enable users to make sense of them. By mapping data to visual representations and providing interactive tools to explore and navigate, it is possible to get an understanding of the data and possibly discover new knowledge. With the advent of modern data collection and analysis technologies, the direct visualization of data starts to show its limitations due to limited scalability in terms of volumes and to the complexity of required analytical reasoning. Many analytical problems we encounter today require approaches that go beyond pure analytics or pure visualization. Visual analytics provides an answer to this problems by advocating a tight integration between automatic computation and interactive visualization, proposing a more holistic approach. In this paper, we argue for Advanced Visual Analytics Interfaces (AVAIs), visual interfaces in which neither the analytics nor the visualization needs to be advanced in itself but where the synergy between automation and visualization is in fact advanced. We offer a detailed argumentation around the needs and challenges of AVAIs and provide several examples of this type of interfaces.
Daniel A. Keim, Peter Bak, Enrico Bertini, Daniela Oelke, David Spretke, Hartmut Ziegler
AVI1
2010 Visual quality metrics and human perception: an initial study on 2D projections of large multidimensional data
abstract
Visual quality metrics have been recently devised to automatically extract interesting visual projections out of a large number of available candidates in the exploration of high-dimensional databases. The metrics permit for instance to search within a large set of scatter plots (e.g., in a scatter plot matrix) and select the views that contain the best separation among clusters. The rationale behind these techniques is that automatic selection of "best" views is not only useful but also necessary when the number of potential projections exceeds the limit of human interpretation. While useful as a concept in general, such metrics received so far limited validation in terms of human perception. In this paper we present a perceptual study investigating the relationship between human interpretation of clusters in 2D scatter plots and the measures automatically extracted out of them. Specifically we compare a series of selected metrics and analyze how they predict human detection of clusters. A thorough discussion of results follows with reflections on their impact and directions for future research.
Andrada Tatu, Peter Bak, Enrico Bertini, Daniel A. Keim, Jörn Schneidewind
AVI4
2010 Event-Based Analysis of People's Activities and Behavior Using Flickr and Panoramio Geotagged Photo Collections
abstract
Photo-sharing websites such as Flickr and Panoramio contain millions of geotagged images contributed by people from all over the world. Characteristics of these data pose new challenges in the domain of spatio-temporal analysis. In this paper, we define several different tasks related to analysis of attractive places, points of interest and comparison of behavioral patterns of different user communities on geotagged photo data. We perform analysis and comparison of temporal events, rankings of sightseeing places in a city, and study mobility of people using geotagged photos. We take a systematic approach to accomplish these tasks by applying scalable computational techniques, using statistical and data mining algorithms, combined with interactive geo-visualization. We provide exploratory visual analysis environment, which allows the analyst to detect spatial and temporal patterns and extract additional knowledge from large geotagged photo collections. We demonstrate our approach by applying the methods to several regions in the world.
Slava Kisilevich, Milos Krstajic, Daniel A. Keim, Natalia V. Andrienko, Gennady L. Andrienko
IV3
2010 Applied Visual Exploration on Real-time News Feeds using Polarity and Geo-spatial Analysis
Milos Krstajic, Peter Bak, Daniela Oelke, Daniel A. Keim, Martin Atkinson, William Ribarsky
WEBIST (1)4
2010 Space-in-Time and Time-in-Space Self-Organizing Maps for Exploring Spatiotemporal Patterns
abstract
Abstract Spatiotemporal data pose serious challenges to analysts in geographic and other domains. Owing to the complexity of the geospatial and temporal components, this kind of data cannot be analyzed by fully automatic methods but require the involvement of the human analyst's expertise. For a comprehensive analysis, the data need to be considered from two complementary perspectives: (1) as spatial distributions (situations) changing over time and (2) as profiles of local temporal variation distributed over space. In order to support the visual analysis of spatiotemporal data, we suggest a framework based on the “Self‐Organizing Map” (SOM) method combined with a set of interactive visual tools supporting both analytic perspectives. SOM can be considered as a combination of clustering and dimensionality reduction. In the first perspective, SOM is applied to the spatial situations at different time moments or intervals. In the other perspective, SOM is applied to the local temporal evolution profiles. The integrated visual analytics environment includes interactive coordinated displays enabling various transformations of spatiotemporal data and post‐processing of SOM results. The SOM matrix display offers an overview of the groupings of data objects and their two‐dimensional arrangement by similarity. This view is linked to a cartographic map display, a time series graph, and a periodic pattern view. The linkage of these views supports the analysis of SOM results in both the spatial and temporal contexts. The variable SOM grid coloring serves as an instrument for linking the SOM with the corresponding items in the other displays. The framework has been validated on a large dataset with real city traffic data, where expected spatiotemporal patterns have been successfully uncovered. We also describe the use of the framework for discovery of previously unknown patterns in 41‐years time series of 7 crime rate attributes in the states of the USA.
Gennady L. Andrienko, Natalia V. Andrienko, Sebastian Bremm, Tobias Schreck, Tatiana von Landesberger, Peter Bak, Daniel A. Keim
Comput. Graph. Forum7
2010 Application of Visual Analytics for Thermal State Management in Large Data Centres
abstract
Abstract Today's large data centres are the computational hubs of the next generation of IT services. With the advent of dynamic smart cooling and rack level sensing, the need for visual data exploration is growing. If administrators know the rack level thermal state changes and catch problems in real time, energy consumption can be greatly reduced. In this paper, we apply a cell‐based spatio‐temporal overall view with high‐resolution time series to simultaneously analyze complex thermal state changes over time across hundreds of racks. We employ cell‐based visualization techniques for trouble shooting and abnormal state detection. These techniques are based on the detection of sensor temperature relations and events to help identify the root causes of problems. In order to optimize the data centre cooling system performance, we derive new non‐overlapped scatter plots to visualize the correlations between the temperatures and chiller utilization. All these techniques have been used successfully to monitor various time‐critical thermal states in real‐world large‐scale production data centres and to derive cooling policies. We are starting to embed these visualization techniques into a handheld device to add mobile monitoring capability.
Ming C. Hao, Ratnesh K. Sharma, Daniel A. Keim, Umeshwar Dayal, Chandrakant D. Patel, Ravigopal Vennelakanti
Comput. Graph. Forum3
2009 Analysis of community-contributed space-and time-referenced data (example of Panoramio photos)
abstract
Space- and time-referenced data published on the Web by general people can be viewed in a dual way: as independent spatio-temporal events and as trajectories of people in the geographical space. These two views suppose different approaches to the analysis, which can yield different kinds of valuable knowledge about places and about people. We present several analysis methods corresponding to these two views. The methods are suited to the large amounts of the data.
Gennady L. Andrienko, Natalia V. Andrienko, Peter Bak, Slava Kisilevich, Daniel A. Keim
GIS5
2009 Personalized News Video Recommendation
Hangzai Luo, Jianping Fan 0001, Daniel A. Keim, Shin'ichi Satoh 0001
MMM3
2009 Analyzing Document Collections via Context-Aware Term Extraction
Daniel A. Keim, Daniela Oelke, Christian Rohrdantz
NLDB1
2009 JustClick: Personalized Image Recommendation via Exploratory Search From Large-Scale Flickr Images
abstract
In this paper, we have developed a novel framework calledJustClickto enable personalized image recommendation via exploratory search from large-scale collections of Flickr images. First, a topic network is automatically generated to summarize large-scale collections of Flickr images at a semantic level. Hyperbolic visualization is further used to enable interactive navigation and exploration of the topic network, so that users can gain insights of large-scale image collections at the first glance, build up their mental query models interactively and specify their queries (i.e., image needs) more precisely by selecting the image topics on the topic network directly. Thus, our personalized query recommendation framework can effectively address both the problem of query formulation and the problem of vocabulary discrepancy and null returns. Second, a small set of most representative images are recommended for the given image topic according to their representativeness scores. Kernel principal component analysis and hyperbolic visualization are seamlessly integrated to organize and layout the recommended images (i.e., most representative images) according to their nonlinear visual similarity contexts, so that users can assess the relevance between the recommended images and their real query intentions interactively. An interactive interface is implemented to allow users to express their time-varying query intentions precisely and to direct ourJustClicksystem to more relevant images according to their personal preferences. Our experiments on large-scale collections of Flickr images show very positive results.
Jianping Fan 0001, Daniel A. Keim, Yuli Gao, Hangzai Luo, Zongmin Li
IEEE Trans. Circuits Syst. Video Technol.2
2009 An Interactive Approach for Filtering Out Junk Images From Keyword-Based Google Search Results
abstract
The keyword-based Google images search engine is now becoming very popular for online image search. Unfortunately, only the text terms that are explicitly or implicitly linked with the images are used for image indexing but the associated text terms may not have exact correspondence with the underlying image semantics, thus the keyword-based Google images search engine may return large amounts of junk images which are irrelevant to the given keyword-based queries. Based on this observation, we have developed an interactive approach to filter out the junk images from keyword-based Google images search results and our approach consists of the following major components. a) A kernel-based image clustering technique is developed to partition the returned images into multiple clusters and outliers. b) Hyperbolic visualization is incorporated to display large amounts of returned images according to their nonlinear visual similarity contexts, so that users can assess the relevance between the returned images and their real query intentions interactively and select one or multiple images to express their query intentions and personal preferences precisely. c) An incremental kernel learning algorithm is developed to translate the users' query intentions and personal preferences for updating the mixture-of-kernels and generating better hypotheses to achieve more accurate clustering of the returned images and filter out the junk images more effectively. Experiments on diverse keyword-based queries from Google images search engine have obtained very positive results. Our junk image filtering system is released for public evaluation at: http://www.cs.uncc.edu/~jfan/google-demo/.
Yuli Gao, Jinye Peng 0001, Hangzai Luo, Daniel A. Keim, Jianping Fan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2009 Spatiotemporal Analysis of Sensor Logs using Growth Ring Maps
abstract
Spatiotemporal analysis of sensor logs is a challenging research field due to three facts: a) traditional two-dimensional maps do not support multiple events to occur at the same spatial location, b) three-dimensional solutions introduce ambiguity and are hard to navigate, and c) map distortions to solve the overlap problem are unfamiliar to most users. This paper introduces a novel approach to represent spatial data changing over time by plotting a number of non-overlapping pixels, close to the sensor positions in a map. Thereby, we encode the amount of time that a subject spent at a particular sensor to the number of plotted pixels. Color is used in a twofold manner; while distinct colors distinguish between sensor nodes in different regions, the colors' intensity is used as an indicator to the temporal property of the subjects' activity. The resulting visualization technique, called Growth Ring Maps, enables users to find similarities and extract patterns of interest in spatiotemporal data by using humans' perceptual abilities. We demonstrate the newly introduced technique on a dataset that shows the behavior of healthy and Alzheimer transgenic, male and female mice. We motivate the new technique by showing that the temporal analysis based on hierarchical clustering and the spatial analysis based on transition matrices only reveal limited results. Results and findings are cross-validated using multidimensional scaling. While the focus of this paper is to apply our visualization for monitoring animal behavior, the technique is also applicable for analyzing data, such as packet tracing, geographic monitoring of sales development, or mobile phone capacity planning.
Peter Bak, Florian Mansmann, Halldór Janetzko, Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.4
2009 Document Cards: A Top Trumps Visualization for Documents
abstract
Finding suitable, less space consuming views for a document's main content is crucial to provide convenient access to large document collections on display devices of different size. We present a novel compact visualization which represents the document's key semantic as a mixture of images and important key terms, similar to cards in a top trumps game. The key terms are extracted using an advanced text mining approach based on a fully automatic document structure extraction. The images and their captions are extracted using a graphical heuristic and the captions are used for a semi-semantic image weighting. Furthermore, we use the image color histogram for classification and show at least one representative from each non-empty image class. The approach is demonstrated for the IEEE InfoVis publications of a complete year. The method can easily be applied to other publication collections and sets of documents which contain images.
Hendrik Strobelt, Daniela Oelke, Christian Rohrdantz, Andreas Stoffel, Daniel A. Keim, Oliver Deussen
IEEE Trans. Vis. Comput. Graph.5
2008 Visual Analytics: Combining Automated Discovery with Interactive Visualizations
Daniel A. Keim, Florian Mansmann, Daniela Oelke, Hartmut Ziegler
ALT1
2008 Visual Analytics: Combining Automated Discovery with Interactive Visualizations
Daniel A. Keim, Florian Mansmann, Daniela Oelke, Hartmut Ziegler
Discovery Science1
2008 Visual Analytics on the Financial Market: Pixel-based Analysis and Comparison of Long-Term Investments
abstract
In this paper, we describe solutions how pixel-based visualization techniques can support the decision making process for investors on the financial market. We especially focus on explorative interactive techniques where analysts try to analyze large amounts of financial data for long-term investments, and show how visualization can effectively support an investor to gain insight into large amounts of financial time series data. After presenting methods for improving the traditional performance/risk computation in order to take user-specific regions of interest into account, we present a novel visualization approach that demonstrates how changes in these regions of interest affect the ranking of assets in a long-term investment strategy.
Hartmut Ziegler, Tilo Nietzschmann, Daniel A. Keim
IV3
2008 Personalized news video recommendation
abstract
In this paper, we have developed an interactive system to enable personalized news video recommendation. First, multi-modal information channels (audio, video and closed captions) are seamlessly integrated and synchronized to achieve more reliable news topic detection, and the contextual relationships between the news topics are extracted automatically. Second, topic network and hyperbolic visualization are seamlessly integrated to achieve interactive navigation and exploration of large-scale collections of news videos at the topic level, so that users can have a good global overview of large-scale collections of news videos at the first glance. In such interactive topic network navigation and exploration process, the users' personal background knowledge can be taken into consideration for obtaining the news topics of interest interactively, building up their mental models of news needs precisely and formulating their searches easily by selecting the visible news topics on the screen directly. Our system can further recommend the relevant web news, the new search directions, and the most relevant news videos according to their importance and representativeness scores.
Hangzai Luo, Jianping Fan 0001, Daniel A. Keim
ACM Multimedia3
2008 Large-Scale Network Monitoring for Visual Analysis of Attacks
Fabian Fischer 0001, Florian Mansmann, Daniel A. Keim, Stephan Pietzko, Marcel Waldvogel
VizSEC3
2008 Density Displays for Data Stream Monitoring
abstract
Abstract In many business applications, large data workloads such as sales figures or process performance measures need to be monitored in real‐time. The data analysts want to catch problems in flight to reveal the root cause of anomalies. Immediate actions need to be taken before the problems become too expensive or consume too many resources. In the meantime, analysts need to have the “big picture” of what the information is about. In this paper, we derive and analyze two real‐time visualization techniques for managing density displays: (1) circular overlay d isplays which visualize large volumes of data without data shift movements after the display is full, thus freeing the analyst from adjusting the mental picture of the data after each data shift; and (2) variable resolution density displays which allow users to get the entire view without cluttering. We evaluate these techniques with respect to a number of evaluation measures, such as constancy of the display and usage of display space, and compare them to conventional d isplays with periodic shifts. Our real time data monitoring system also provides advanced interactions such as a local root cause analysis for further exploration. The applications using a number of real‐world data sets show the wide applicability and usefulness of our ideas.
Ming C. Hao, Daniel A. Keim, Umeshwar Dayal, Daniela Oelke, Chantal Tremblay
Comput. Graph. Forum2
2008 COPERNICUS: Context-Preserving Engine for Route Navigation with Interactive User-modifiable Scaling
abstract
Abstract In this paper, we present an automated system for generating context‐preserving route maps that depict navigation routes as a path between nodes and edges inside a topographic network. Our application identifies relevant context information to support navigation and orientation, and generates customizable route maps according to design principles that communicate all relevant context information clearly visible on one single page. Interactive scaling allows seamless transition between the original undistorted map and our new map design, and supports user‐specified scaling of regions of interest to create personalized driving directions according to the drivers needs.
Hartmut Ziegler, Daniel A. Keim
Comput. Graph. Forum2
2007 Visual Exploration and Discovery of Atypical Behavior in Financial Time Series Data using Two-Dimensional Colormaps
abstract
This work describes two pixel-based paradigms for visual financial time series data analysis that allow analyzing assets in overview and in detail, and offer improved insights into the characteristics of assets compared to traditional visualization methods. We contribute a two-dimensional color coding scheme for inter-Zintra-asset analysis that extends the two paradigms and supports the discovery of significant characteristics of assets, such as atypical " against-the-market" -behavior inform of exceptional stability in case of whole market losses, or the discovery of assets in a portfolio that - while still being profitable - notably underperform the market median. We apply our techniques on real world data sets, and focus on assets of the banks of Switzerland.
Hartmut Ziegler, Tilo Nietzschmann, Daniel A. Keim
IV3
2007 Multi-Resolution Techniques for Visual Exploration of Large Time-Series Data
abstract
Time series are a data type of utmost importance in many domains such as business management and service monitoring. We address the problem of visualizing large time-related data sets which are difficult to visualize effectively with standard techniques given the limitations of current display devices. We propose a framework for intelligent time- and data-dependent visual aggregation of data along multiple resolution levels. This idea leads to effective visualization support for long time-series data providing both focus and context. The basic idea of the technique is that either data-dependent or application-dependent, display space is allocated in proportion to the degree of interest of data subintervals, thereby (a) guiding the user in perceiving important information, and (b) freeing required display space to visualize all the data. The automatic part of the framework can accommodate any time series analysis algorithm yielding a numeric degree of interest scale. We apply our techniques on real-world data sets, compare it with the standard visualization approach, and conclude the usefulness and scalability of the approach.
Ming C. Hao, Umeshwar Dayal, Daniel A. Keim, Tobias Schreck
EuroVis3
2007 Relevance Driven Visualization of Financial Performance Measures
abstract
Visual data analysis has received a lot of research interest in recent years, and a wide variety of new visualization techniques and applications have been developed to improve insight into the various application domains. In financial data analysis, however, analysts still primarily rely on a set of statistical performance parameters in combination with traditional line charts in order to evaluate assets and to make decisions, and information visualization is only very slowly entering this important domain. In this paper, we analyze some of the standard statistical measures for technical financial data analysis and demonstrate cases where they produce insufficient and misleading results that do not reflect the real performance of an asset. We propose a technique for visualizing financial time series data that eliminates these inadequacies, offering a complete view on the real performance of an asset. The technique is enhanced by relevance and weighting functions according to the users' preferences in order to emphasize specific regions of interest. Based on these principles we redefine some of the standard performance measures. We apply our technique on real world financial data sets and combine it with higher-level financial analysis techniques such as performance/risk analysis, dominance evaluation, and efficiency curves in order to show how traditional techniques from economics can be improved by modern visual data analysis techniques.
Hartmut Ziegler, Tilo Nietzschmann, Daniel A. Keim
EuroVis3
2007 Visualization of Host Behavior for Network Security
Florian Mansmann, Daniel A. Keim
VizSEC3
2007 Introduction
Jörn Kohlhammer, Daniel A. Keim, David S. Ebert
Comput. Graph.2
2007 Geovisual analytics for spatial decision support: Setting the research agenda
abstract
This article summarizes the results of the workshop on Visualization, Analytics & Spatial Decision Support, which took place at the GIScience conference in September 2006. The discussions at the workshop and analysis of the state of the art have revealed a need in concerted cross‐disciplinary efforts to achieve substantial progress in supporting space‐related decision making. The size and complexity of real‐life problems together with their ill‐defined nature call for a true synergy between the power of computational techniques and the human capabilities to analyze, envision, reason, and deliberate. Existing methods and tools are yet far from enabling this synergy. Appropriate methods can only appear as a result of a focused research based on the achievements in the fields of geovisualization and information visualization, human‐computer interaction, geographic information science, operations research, data mining and machine learning, decision science, cognitive science, and other disciplines. The name ‘Geovisual Analytics for Spatial Decision Support’ suggested for this new research direction emphasizes the importance of visualization and interactive visual interfaces and the link with the emerging research discipline of Visual Analytics. This article, as well as the whole special issue, is meant to attract the attention of scientists with relevant expertise and interests to the major challenges requiring multidisciplinary efforts and to promote the establishment of a dedicated research community where an appropriate range of competences is combined with an appropriate breadth of thinking.
Gennady L. Andrienko, Natalia V. Andrienko, Piotr Jankowski 0001, Daniel A. Keim, Menno-Jan Kraak, Alan M. MacEachren, Stefan Wrobel
Int. J. Geogr. Inf. Sci.4
2007 Highlighting space-time patterns: Effective visual encodings for interactive decision-making
abstract
The research reported in this paper focuses on integrating analytical and visual methods in order to explore complex patterns in geo‐related multivariate data sets and to understand the changes in patterns over time. The goal is to provide techniques that are able to analyse real‐world Data Warehouses, a typical architecture to manage such geo‐related multidimensional data sets, in order to support the analyst's decision‐making process. Challenges arise because real‐world applications usually have to deal with millions of records, with dozens of dimensions, and spatio‐temporal context. Therefore, a tight integration of automated analysis and interactive visualizations is needed (as proposed in the context of Visual Analytics). Our approach uses the well‐studied capabilities provided by Data Warehouses supporting knowledge discovery and decision‐making to analyse spatio‐temporal behaviour of pattern in high‐dimensional spaces. The topic of the paper is to show possible interplays between automated analysis and geo‐spatial visualization.
Mike Sips, Jörn Schneidewind, Daniel A. Keim
Int. J. Geogr. Inf. Sci.3
2007 Value-Cell Bar Charts for Visualizing Large Transaction Data Sets
abstract
One of the common problems businesses need to solve is how to use large volumes of sales histories, Web transactions, and other data to understand the behavior of their customers and increase their revenues. Bar charts are widely used for daily analysis, but only show highly aggregated data. Users often need to visualize detailed multidimensional information reflecting the health of their businesses. In this paper, we propose an innovative visualization solution based on the use of value cells within bar charts to represent business metrics. The value of a transaction can be discretized into one or multiple cells: high-value transactions are mapped to multiple value cells, whereas many small-value transactions are combined into one cell. With value-cell bar charts, users can 1) visualize transaction value distributions and correlations, 2) identify high-value transactions and outliers at a glance, and 3) instantly display values at the transaction record level. Value-Cell Bar Charts have been applied with success to different sales and IT service usage applications, demonstrating the benefits of the technique over traditional charting techniques. A comparison with two variants of the well-known Treemap technique and our earlier work on Pixel Bar Charts is also included.
Daniel A. Keim, Ming C. Hao, Umeshwar Dayal, Martha Lyons
IEEE Trans. Vis. Comput. Graph.1
2007 Visual Analysis of Network Traffic for Resource Planning, Interactive Monitoring, and Interpretation of Security Threats
abstract
The Internet has become a wild place: malicious code is spread on personal computers across the world, deploying botnets ready to attack the network infrastructure. The vast number of security incidents and other anomalies overwhelms attempts at manual analysis, especially when monitoring service provider backbone links. We present an approach to interactive visualization with a case study indicating that interactive visualization can be applied to gain more insight into these large data sets. We superimpose a hierarchy on IP address space, and study the suitability of Treemap variants for each hierarchy level. Because viewing the whole IP hierarchy at once is not practical for most tasks, we evaluate layout stability when eliding large parts of the hierarchy, while maintaining the visibility and ordering of the data of interest.
Florian Mansmann, Daniel A. Keim, Stephen C. North, Brian Rexroad, Daniel Sheleheda
IEEE Trans. Vis. Comput. Graph.2
2006 Task-at-hand interface for change detection in stock market data
abstract
Companies trading stocks need to store information on stock prices over specific time intervals, which results in very large databases. Large quantities of numerical data (thousands of records) are virtually impossible to understand quickly and require the use of a visual model, since that is the fastest way for a human brain to absorb those enormous collections of data. However, little work has been done on verifying which visualizations are more suitable to represent these data sets. Such work is of crucial importance, since it enables us to identify those useful visual models and, in addition, opens our minds to new research possibilities. This paper presents an empirical study of different visualizations, that have been employed for stock market data, by comparing the results obtained by all studied techniques in typical exploratory data analysis tasks. This work provides several research contributions to the design of advanced visual data exploration interfaces.
Carmen Sanz Merino, Mike Sips, Daniel A. Keim, Christian Panse, Robert Spence
AVI3
2006 Finding Correlations in Functionally Equivalent Proteins by Integrating Automated and Visual Data Exploration
abstract
The analysis of alignments of functionally equivalent proteins can reveal regularities such as correlated positions or residue patterns which are important to ensure a specific fold and various cellular functions. Many approaches are found in the literature which try to identify correlated positions to predict the residues that are close to each other in the three-dimensional folded structure. However, the quality of the predictions remains disappointing. One of the problems is that the statistical correlation measures that were used cannot do justice to the underlying complex biological and physicochemical realities. In this paper we evaluate the biological requirements for a correlation measure and explain why a completely automatic approach is unlikely to succeed. We then propose a novel and flexible criteria for correlation of residue positions in protein sequences, which can be optimized for different requirements. To apply this definition we developed the tool VisAlign that combines an automatic calculation of correlations with an interactive visualization. This allows the user to visually explore alternative alignments and thereby conveniently test various hypothesis and to detect regularities in the aligned sequences
Daniel A. Keim, Daniela Oelke, Royal Truman, Klaus Neuhaus
BIBE1
2006 Visual Feature Space Analysis for Unsupervised Effectiveness Estimation and Feature Engineering
abstract
The feature vector approach is one of the most popular schemes for managing multimedia data. For many data types such as audio, images, or 3D models, an abundance of different feature vector extractors are available. The automatic (unsupervised) identification of the best suited feature extractor for a given multimedia database is a difficult and largely unsolved problem. We here address the problem of comparative unsupervised feature space analysis. We propose two interactive approaches for the visual analysis of certain feature space characteristics contributing to estimated discrimination power provided in the respective feature spaces. We apply the approaches on a database of 3D objects represented in different feature spaces, and we experimentally show the methods to be useful (a) for unsupervised comparative estimation of discrimination power and (b) for visually analyzing important properties of the components (dimensions) of the respective feature spaces. The results of the analysis are useful for feature selection and engineering
Tobias Schreck, Daniel A. Keim, Christian Panse
ICME2
2006 European Research Forum Panel Session Envisioning Research Challenges in Visual Analytics
abstract
Visual Analytics is the science of analytical reasoning supported by interactive visual interfaces. People use visual analytics tools and techniques to synthesize information; derive insight from massive, dynamic, and often conflicting data; detect the expected and discover the unexpected; provide timely, defensible, and understandable assessments; and communicate assessments effectively for action. The issues stimulating this body of research provide a grand challenge in science: turning information overload into the opportunity of the decade. Visual analytics requires interdisciplinary science beyond traditional scientific and information visualization to include statistics, data mining, knowledge and discovery technologies, cognitive science and humancomputer interaction, production and presentation, and more. An important research agenda "Illuminating the Path" provides recommendations for the next generation suite of visual analytics technologies and is available at http://nvac.pnl.gov/agenda.stm .
Mikael Jern, Ebad Banissi, Gennady L. Andrienko, Wolfgang Müller 0004, Daniel A. Keim
IV5
2006 Challenges in Visual Data Analysis
abstract
In today's applications data is produced at unprecedented rates. While the capacity to collect and store new data grows rapidly, the ability to analyze these data volumes increases at much lower pace. This gap leads to new challenges in the analysis process, since analysts, decision makers, engineers, or emergency response teams depend on information "concealed" in the data. The emerging field of visual analytics focuses on handling massive, heterogenous, and dynamic volumes of information through integration of human judgement by means of visual representations and interaction techniques in the analysis process. Furthermore, it is the combination of related research areas including visualization, data mining, and statistics that turns visual analytics into a promising field of research. This paper aims at providing an overview of visual analytics, its scope and concepts, and details the most important technical research challenges in the field
Daniel A. Keim, Florian Mansmann, Jörn Schneidewind, Hartmut Ziegler
IV1
2006 Scalable Pixel-based Visual Interfaces: Challenges and Solutions
abstract
The information revolution is creating and publishing vast data sets, such as records of business transactions, environmental statistics and census demographics. In many application domains, this data is collected and indexed by geo-spatial location. The discovery of interesting patterns in such databases through visual analytics is a key to turn this data into valuable information. Challenges arise because newly available geo-spatial data sets often have millions of records, or even far more, they are from multiple and heterogeneous data sources, and the output devices have significantly changed, e.g. high-resolution pixilated displays are increasingly available in both wall-sized and desktop units. New techniques are needed to cope with this scale. In this paper, we focus on ways to increase the scalability of pixel-based visual interfaces by adding task on hands scenarios that tightly integrate the data analyst into the exploration of geo-spatial data sets
Mike Sips, Jörn Schneidewind, Daniel A. Keim, Heidrun Schumann
IV3
2006 A Spectral Visualization System for Analyzing Financial Time Series Data
abstract
Visual data analysis of time related data sets has attracted much research interest recently, and a number of sophisticated visualization methods have been proposed in the past. In financial analysis, however, the most important and most common visualization techniques for time series data is the traditional line- or bar chart. Although these are intuitive and make it easy to spot the effect of key events on a asset’s price, and its return over a given period of time, price charts do not allow the easy perception of relative movements in terms of growth rates, which is the key feature of any price-related time series. This paper presents a novel Growth Matrix visualization technique for analyzing assets. It extends the ability of existing chart techniques by not only visualizing asset return rates over fixed time frames, but over the full spectrum of all subintervals present in a given time frame, in a single view. At the same time, the technique allows a comparison of subinterval return rates among groups of even a few hundreds of assets. This provides a powerful way for analyzing financial data, since it allows the identification of strong and weak periods of assets as compared to global market characteristics, and thus allows a more encompassing visual classification into "good" and "poor" performers than existing chart techniques. We illustrate the technique by real-world examples showing the abilities of the new approach, and its high relevance for financial analysis tasks.
Daniel A. Keim, Tilo Nietzschmann, Norman Schelwies, Jörn Schneidewind, Tobias Schreck, Hartmut Ziegler
EuroVis1
2006 Guest Editorial: Special Section on Visual Analytics
abstract
ISUAL analytics is the science of analytical reasoning supportedbyhighlyinteractivevisualinterfaces.People use visual analytics tools and techniques to synthesize information; derive insight from massive, dynamic, and often conflicting data; detect the expected and discover the unexpected; provide timely, defensible, and understandable assessments; and communicate assessments effectively for action. The issues stimulating this body of research provide a grandchallengeinscience:turninginformationoverloadinto the opportunity of the decade. Visual analytics requires interdisciplinary science beyond traditional scientific and information visualization to include statistics, mathematics, knowledge representation, management and discovery technologies, cognitive and perceptual sciences, decision sciences, and more. An important research agenda to develop the next generation suite of visual analytics technologies is described in the book Illuminating the Path: The Research and Development Agenda for Visual Analytics, which is available at http:// nvac.pnl.gov/agenda.stm. The papers in this special section address a number of the issues described in the visual analytics research agenda. They are grouped into five major areas: multidimensional data, graphs and networks, communication network analysis, space and time, and fundamentals. The first two papers address issues of visual analysis of multidimensional data. The first paper, “High-Dimensional
Daniel A. Keim, George G. Robertson, James J. Thomas, Jarke J. van Wijk
IEEE Trans. Vis. Comput. Graph.1
2006 Visualization of Geo-spatial Point Sets via Global Shape Transformation and Local Pixel Placement
abstract
In many applications, data is collected and indexed by geo-spatial location. Discovering interesting patterns through visualization is an important way of gaining insight about such data. A previously proposed approach is to apply local placement functions such as PixelMaps that transform the input data set into a solution set that preserves certain constraints while making interesting patterns more obvious and avoid data loss from overplotting. In experience, this family of spatial transformations can reveal fine structures in large point sets, but it is sometimes difficult to relate those structures to basic geographic features such as cities and regional boundaries. Recent information visualization research has addressed other types of transformation functions that make spatially-transformed maps with recognizable shapes. These types of spatial-transformation are called global shape functions. In particular, cartogram-based map distortion has been studied. On the other hand, cartogram-based distortion does not handle point sets readily. In this study, we present a framework that allows the user to specify a global shape function and a local placement function. We combine cartogram-based layout (global shape) with PixelMaps (local placement), obtaining some of the benefits of each toward improved exploration of dense geo-spatial data sets.
Christian Panse, Mike Sips, Daniel A. Keim, Stephen C. North
IEEE Trans. Vis. Comput. Graph.3
2005 VisBiz: A Business Process Visualization Case Study
abstract
Business process management involves many parameters and relationships and is modeled as complex business process workflows. A common way to analyze the process data is by using flowcharts. Visual analysis of a largescale chart, however, is too complex. In this case study, we employ a novel visualization technique, called VisBiz. VisBiz reduces data complexity by automatically analyzing operational data and abstracting the most critical parameters that influence business process. The basic idea is to select the most relevant parameters and layout them on a “triple-attributes” circular graph based on their relationships and user domain knowledge. VizBiz transforms the attributes to nodes and the process flows to lines. VisBiz derives a new process flow matrix to link the process of multiple circular graphs as the analyst introduces more parameters for further analysis. The results of the real-world credit card fraud study show the significant advantages of this technique in finding fraud distribution patterns and root causes of frauds.
Ming C. Hao, Daniel A. Keim, Umeshwar Dayal, Jörn Schneidewind
EuroVis2
2005 Mail Explorer - Spatial and Temporal Exploration of Electronic Mail
abstract
In today’s world, e-mail has become one of the most important means of communication in business and private lives due to its efficiency. However, the problems start as soon as mail volumes go beyond the scope of human information processing capabilities. Firstly, time does not allow for leaving certain messages unanswered for a long time, and in certain cases, for reading all messages. Secondly, the dilemma of electronic filters leaves a choice of too many junk mails getting through versus a risk of solicited mails being dumped. In this paper we present a new interactive visual data mining approach for analyzing individual e-mail communication. It combines classical visual analytics (help to identify pattern such as peaks and trends over time) with geo-spatial map distortions (help to understand the routes of e-mails). Experiments show that our visual e-mail explorer produces useful and interesting visualizations of large collections of e-mail and is practical for exploring temporal and geo-spatial patterns hidden in the e-mail data.
Daniel A. Keim, Florian Mansmann, Christian Panse, Jörn Schneidewind, Mike Sips
EuroVis1
2004 CircleView: a new approach for visualizing time-related multidimensional data sets
abstract
This paper introduces a new approach for visualizing multidimensional time-referenced data sets, called Circle View. The Circle View technique is a combination of hierarchical visualization techniques, such as treemaps [6], and circular layout techniques such as Pie Charts and Circle Segments [2]. The main goal is to compare continuous data changing their characteristics over time in order to identify patterns, exceptions and similarities in the data.To achieve this goal Circle View is a intuitive and easy to understand visualization interface to enable the user very fast to acquire the information needed. This is an important feature for fast changing visualization caused by time related data streams. Circle View supports the visualization of the changing characteristics over time, to allow the user the observation of changes in the data. Additionally it provides user interaction and drill down mechanism depending on user demands for a effective exploratory data analysis. There is also the capability of exploring correlations and exceptions in the data by using similarity and ordering algorithms.
Daniel A. Keim, Jörn Schneidewind, Mike Sips
AVI1
2004 Similarity Search in Multimedia Databases
abstract
The research on multimedia databases involves different areas in Computer Science, such as computer graphics, databases, and information retrieval. There are many practical applications that benefit from this research, e.g., molecular biology, medicine, CAD/CAM, and geography. An important characteristic of these applications is the variety of data that should be supported, e.g., text, images (both still and moving), and audio. This implies that the development of a multimedia information system is considerably more complex than a traditional information system. An important research issue in the field of multimedia databases is the content-based retrieval of similar objects. Given a multimedia query object, the search for an exact match in a database is not meaningful in most applications, because the probability that two multimedia objects are identical is negligible (unless they are digital copies from the same source). For this reason, the development of efficient and effective similarity search techniques has become an important topic in the multimedia database research community. The goal of this advanced technology seminar is to provide an overview of the similarity search problem and to present the state-of-art techniques for performing efficient and effective similarity queries in multimedia databases. The seminar begins with an introduction and a motivation of multimedia databases. The two main approaches for describing multimedia objects (as elements in a metric space or in a vector space) are introduced, as well as a description of the ”Multimedia Content Description Interface” (MPEG)-7 standard. The efficiency issue is addressed for both metric and vector space approaches, describing the data structures and algorithms used to answer similarity queries. For the effectiveness issue, the seminar introduces some widely used retrieval performance measures. Several examples of techniques for particular multimedia applications (text, image, CAD, 3D objects, audio and video) are presented. The seminar outline is as follows:
Daniel A. Keim, Benjamin Bustos
ICDE1
2004 Using entropy impurity for improved 3D object similarity search
abstract
Similarity search in 3D object databases is becoming an important problem in multimedia retrieval, with many practical applications. We investigate methods for improving the effectiveness in a retrieval system that implements multiple feature extraction algorithms to choose from. Our techniques are based on the entropy impurity measure, widely used in the context of decision trees. We propose a method for the a priori estimation of individual feature vector performance, given a query. We then define two approaches that use this estimator to improve the retrieval effectiveness. Our experimental results show that significant improvements are achievable using these methods.
Benjamin Bustos, Daniel A. Keim, Dietmar Saupe, Tobias Schreck, Dejan V. Vranic
ICME2
2004 2D Maps for Visual Analysis and Retrieval in Large Multi-Feature 3D Model Databases
abstract
Multimedia objects are often described by high-dimensional feature vectors which can be used for retrieval and clustering tasks. We have built an interactive retrieval system for 3D model databases that implements a variety of different feature transforms. Recently, we have enhanced the functionality of our system by integrating a SOM-based visualization module. In this poster demo, we show how 2D maps can be used to improve the effectiveness of retrieval, clustering, and over-viewing tasks in a 3D multimedia system.
Benjamin Bustos, Daniel A. Keim, Christian Panse, Tobias Schreck
IEEE Visualization2
2004 VisBiz: A Simplified Visualization of Business Operation
abstract
In this poster, we present a new technique, VisBiz, for interactively visualizing business operations. The basic idea of this technique is to visually mining relationships between important operation parameters (attributes) and to map the parameters into visualizations. VisBiz simplifies the complexity by partitioning the operation into multiple attribute circular graphs. VisBiz allows the analysis of business data as follows:
Ming C. Hao, Daniel A. Keim, Umeshwar Dayal
IEEE Visualization2
2004 Pixel based visual data mining of geo-spatial data
Daniel A. Keim, Christian Panse, Mike Sips, Stephen C. North
Comput. Graph.1
2004 Introduction to special issue with best papers from KDD 2002
Daniel A. Keim, Nick Koudas
Inf. Syst.1
2004 Multiresolution similarity search in image databases
Martin Heczko, Alexander Hinneburg, Daniel A. Keim, Markus Wawryniuk
Multim. Syst.3
2004 Guest Editor's Introduction: Special Section on InfoVis
abstract
Three papers in this issue of TVCG are expanded versions of ones presented at the IEEE Symposium on Information Visualization (InfoVis) 2003. These examples convey the breadth of work in the InfoVis community, attacking practically motivated problems from sophisticated mathematical and scientific perspectives. They demonstrate an interaction technique for navigation, a mathematically well-founded approach to the difficult problem of dimensionality reduction, and an evaluation of a particular channel of visual perception. These are briefly summarized.
Daniel A. Keim, Tamara Munzner, Stephen C. North
IEEE Trans. Vis. Comput. Graph.1
2004 CartoDraw: A Fast Algorithm for Generating Contiguous Cartograms
abstract
Cartograms are a well-known technique for showing geography-related statistical information, such as population demographics and epidemiological data. The basic idea is to distort a map by resizing its regions according to a statistical parameter, but in a way that keeps the map recognizable. In this study, we formally define a family of cartogram drawing problems. We show that even simple variants are unsolvable in the general case. Because the feasible variants are NP-complete, heuristics are needed to solve the problem. Previously proposed solutions suffer from problems with the quality of the generated drawings. For a cartogram to be recognizable, it is important to preserve the global shape or outline of the input map, a requirement that has been overlooked in the past. To address this, our objective function for cartogram drawing includes both global and local shape preservation. To measure the degree of shape preservation, we propose a shape similarity function, which is based on a Fourier transformation of the polygons' curvatures. Also, our application is visualization of dynamic data, for which we need an algorithm that recalculates a cartogram in a few seconds. None of the previous algorithms provides adequate performance with an acceptable level of quality for this application. In this paper, we therefore propose an efficient iterative scanline algorithm to reposition edges while preserving local and global shapes. Scanlines may be generated automatically or entered interactively to guide the optimization process more closely. We apply our algorithm to several example data sets and provide a detailed comparison of the two variants of our algorithm and previous approaches.
Daniel A. Keim, Stephen C. North, Christian Panse
IEEE Trans. Vis. Comput. Graph.1
2003 HD-Eye - Visual Clustering of High dimensional Data
abstract
Clustering of large databases is an important research area with a large variety of applications in the data base context. Missing in most of the research efforts are means for guiding the clustering process and understand the results, which is especially important if the data under consideration is high dimensional and has not been collected for the purpose of being analyzed. Visualization technology may help to solve this problem since it allows an effective support of different clustering paradigms and provides means for a visual inspection of the results. Our HD-Eye (high-dimensional eye) system (A. Hinneburg et al., 1999) shows that a tight integration of advanced clustering algorithms and state-of-the-art visualization techniques is powerful for a better understanding and effective guidance of the clustering process, and therefore can help to significantly improve the clustering results. The demonstration shows how the user can visually explore the data by focusing on interesting projections and guide the important steps of the clustering process. Due to its interactive nature, the HD-Eye system allows a combination of multiple clustering paradigms, leading to clustering models, which fit, well to the intended tasks and the users interests. In addition, the integrated data visualization capabilities of the HD-Eye system lead to a better understanding of the clustering results. The applications to be demonstrated include clustering of large image as well as molecular biology databases.
Alexander Hinneburg, Daniel A. Keim, Markus Wawryniuk
ICDE2
2003 Analyzing High-Dimensional Data by Subspace Validity
abstract
We are proposing a novel method that makes it possible to analyze high-dimensional data with arbitrary shaped projected clusters and high noise levels. At the core of our method lies the idea of subspace validity. We map the data in a way that allows us to test the quality of subspaces using statistical tests. Experimental results, both on synthetic and real data sets, demonstrate the potential of our method.
Amihood Amir, Reuven Kashi, Nathan S. Netanyahu, Daniel A. Keim, Markus Wawryniuk
ICDM4
2003 PixelMaps: A New Visual Data Mining Approach for Analyzing Large Spatial Data Sets
abstract
PixelMaps are a new pixel-oriented visual data mining technique for large spatial datasets. They combine kernel-density-based clustering with pixel-oriented displays to emphasize clusters while avoiding overlap in locally dense point sets on maps. Because a full evaluation of density functions is prohibitively expensive, we also propose an efficient approximation, Fast-PixelMap, based on a synthesis of the quadtree and gridfile data structures.
Daniel A. Keim, Christian Panse, Mike Sips, Stephen C. North
ICDM1
2003 A Database Striptease or How to Manage Your Personal Databases
Martin L. Kersten, Gerhard Weikum, Michael J. Franklin, Daniel A. Keim, Alejandro P. Buchmann, Surajit Chaudhuri
VLDB4
2003 Guest Editorial
David J. Hand, Daniel A. Keim, Raymond T. Ng
Data Min. Knowl. Discov.2
2003 Guest Editorial
David J. Hand, Daniel A. Keim, Raymond T. Ng
Data Min. Knowl. Discov.2
2003 A General Approach to Clustering in Large Databases with Noise
Alexander Hinneburg, Daniel A. Keim
Knowl. Inf. Syst.2
2002 HD-Eye: visual clustering of high dimensional data
abstract
Clustering of large data bases is an important research area with a large variety of applications in the data base context.Missing in most of the research efforts are means for guiding the clustering process and understanding the results, which is especially important for high dimensional data.Visualization technology may help to solve this problem since it provides effective support of different clustering paradigms and allows a visual inspection of the results.The HD-Eye (high-dim.eye) system shows that a tight integration of advanced clustering algorithms and state-of-the-art visualization techniques is powerful for a better understanding and effective guidance of the clustering process, and therefore can help to significantly improve the clustering results.
Alexander Hinneburg, Daniel A. Keim, Markus Wawryniuk
SIGMOD Conference2
2002 Information Visualization and Visual Data Mining
abstract
Never before in history has data been generated at such high volumes as it is today. Exploring and analyzing the vast volumes of data is becoming increasingly difficult. Information visualization and visual data mining can help to deal with the flood of information. The advantage of visual data exploration is that the user is directly involved in the data mining process. There are a large number of information visualization techniques which have been developed over the last decade to support the exploration of large data sets. In this paper, we propose a classification of information visualization and visual data mining techniques which is based on the data type to be visualized, the visualization technique, and the interaction and distortion technique. We exemplify the classification using a few examples, most of them referring to techniques and systems presented in this special section.
Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.1
2002 Hierarchical Pixel Bar Charts
abstract
Simple presentation graphics are intuitive and easy-to-use, but only show highly aggregated data. Bar charts, for example, only show a rather small number of data values and x-y-plots often have a high degree of overlap. Presentation techniques are often chosen depending on the considered data type, bar charts, for example, are used for categorical data and x-y plots are used for numerical data. We propose a combination of traditional bar charts and x-y-plots, which allows the visualization of large amounts of data with categorical and numerical data. The categorical data dimensions are used for the partitioning into the bars and the numerical data dimensions are used for the ordering arrangement within the bars. The basic idea is to use the pixels within the bars to present the detailed information of the data records. Our so-called pixel bar charts retain the intuitiveness of traditional bar charts while applying the principle of x-y charts within the bars. In many applications, a natural hierarchy is defined on the categorical data dimensions such as time, region, or product type. In hierarchical pixel bar charts, the hierarchy is exploited to split the bars for selected portions of the hierarchy. Our application to a number of real-world e-business and Web services data sets shows the wide applicability and usefulness of our new idea.
Daniel A. Keim, Ming C. Hao, Umeshwar Dayal
IEEE Trans. Vis. Comput. Graph.1
2001 On the Surprising Behavior of Distance Metrics in High Dimensional Spaces
Charu C. Aggarwal, Alexander Hinneburg, Daniel A. Keim
ICDT3
2001 On Optimizing Nearest Neighbor Queries in High-Dimensional Data Spaces
Stefan Berchtold, Christian Böhm 0001, Daniel A. Keim, Florian Krebs, Hans-Peter Kriegel
ICDT3
2000 Clustering 3D-structures of Small Aminoacid-chains for Detecting Dependence from Their Sequential Context in Proteins
abstract
In the past, a good number of rotamer libraries have been published, which allow a deeper understanding of the conformational behavior of amino acid residues in proteins. Since the number of available high-resolution X-ray protein structures has grown significantly over the last years, a more comprehensive analysis of the conformational behavior is possible today. In this paper, we present a method to compile a new class of rotamer libraries for detecting interesting relationships between residue conformations and their sequential context in proteins. The method is based on a new algorithm for clustering residue conformations. To demonstrate the effectiveness of our method, we apply our algorithm to a library consisting of all 8000 tripeptide fragments formed by the 20 native amino acids. The analysis shows some very interesting new results, namely that some specific tripeptide fragments show some unexpected conformation of residues instead of the highly preferred conformation. In the neighborhood of two asparagine residues, for example, threonine avoids the conformation which is most likely to occur otherwise. The new insights obtained by the analysis are important in understanding the formation and prediction of secondary structure elements and will consequently be crucial for improving the state-of-the-art of protein folding.
Alexander Hinneburg, Daniel A. Keim, Wolfgang Brandt
BIBE2
2000 Optimal Multidimensional Query Processing Using Tree Striping
Stefan Berchtold, Christian Böhm 0001, Daniel A. Keim, Hans-Peter Kriegel, Xiaowei Xu 0001
DaWaK3
2000 Indexing High-Dimensional Spaces: Database Support for Next Decade's Applications
abstract
Summary form only given. The tutorial is structured as follows: In the first section, we describe two examples of new database applications, which demonstrate the need for efficient query processing techniques in highdimensional spaces. In the second section, we discuss the effects occurring in high-dimensional spaces - first from a pure mathematical point of view and then from a database perspective. Next, we describe the different approaches for modeling the costs of processing queries on high-dimensional data. The description of the different approaches demonstrates nicely what happens if we ignore the special properties of high-dimensional spaces. In the fourth section, we then provide a structured overview of the proposed querying and indexing techniques, discussing their advantages and drawbacks. In this section, we also cover a number of additional techniques dealing with optimization and parallelization. In concluding the tutorial, we try to stir further research activities by presenting a number of interesting research problems.
Stefan Berchtold, Daniel A. Keim
ICDE2
2000 What Is the Nearest Neighbor in High Dimensional Spaces?
Alexander Hinneburg, Charu C. Aggarwal, Daniel A. Keim
VLDB3
2000 Indexing the Solution Space: A New Technique for Nearest Neighbor Search in High-Dimensional Space
abstract
Similarity search in multimedia databases requires an efficient support of nearest-neighbor search on a large set of high-dimensional points as a basic operation for query processing. As recent theoretical results show, state of the art approaches to nearest-neighbor search are not efficient in higher dimensions. In our new approach, we therefore precompute the result of any nearest-neighbor search which corresponds to a computation of the Voronoi cell of each data point. In a second step, we store conservative approximations of the Voronoi cells in an index structure efficient for high-dimensional data spaces. As a result, nearest neighbor search corresponds to a simple point query on the index structure. Although our technique is based on a precomputation of the solution space, it is dynamic, i.e., it supports insertions of new data points. An extensive experimental evaluation of our technique demonstrates the high efficiency for uniformly distributed as well as real data. We obtained a significant reduction of the search time compared to nearest neighbor search in other index structures such as the X-tree.
Stefan Berchtold, Daniel A. Keim, Hans-Peter Kriegel, Thomas Seidl 0001
IEEE Trans. Knowl. Data Eng.2
2000 Designing Pixel-Oriented Visualization Techniques: Theory and Applications
abstract
Visualization techniques are of increasing importance in exploring and analyzing large amounts of multidimensional information. One important class of visualization techniques which is particularly interesting for visualizing very large multidimensional data sets is the class of pixel-oriented techniques. The basic idea of pixel-oriented visualization techniques is to represent as many data objects as possible on the screen at the same time by mapping each data value to a pixel of the screen and arranging the pixels adequately. A number of different pixel-oriented visualization techniques have been proposed in recent years and it has been shown that the techniques are useful for visual data exploration in a number of different application contexts. In this paper, we discuss a number of issues which are important in developing pixel-oriented visualization techniques. The major goal of this article is to provide a formal basis of pixel-oriented visualization techniques and show that the design decisions in developing them can be seen as solutions of well-defined optimization problems. This is true for the mapping of the data values to colors, the arrangement of pixels inside the subwindows, the shape of the subwindows, and the ordering of the dimension subwindows. The paper also discusses the design issues of special variants of pixel-oriented techniques for visualizing large spatial data sets.
Daniel A. Keim
IEEE Trans. Vis. Comput. Graph.1
1999 Clustering Methods for Large Databases: From the Past to the Future
abstract
No abstract available.
Alexander Hinneburg, Daniel A. Keim
SIGMOD Conference2
1999 Efficient Geometry-based Similarity Search of 3D Spatial Databases
abstract
Searching a database of 3D-volume objects for objects which are similar to a given 3D search object is an important problem which arises in number of database applications — for example, in Medicine and CAD. In this paper, we present a new geometry-based solution to the problem of searching for similar 3D-volume objects. The problem is motivated from a real application in the medical domain where volume similarity is used as a basis for surgery decisions. Our solution for an efficient similarity search on large databases of 3D volume objects is based on a new geometric index structure. The basic idea of our new approach is to use the concept of hierarchical approximations of the 3D objects to speed up the search process. We formally show the correctness of our new approach and introduce two instantiations of our general idea, which are based on cuboid and octree approximations. We finally provide a performance evaluation of our new index structure revealing significant performance improvements over existing approaches.
Daniel A. Keim
SIGMOD Conference1
1999 Visualizing Large-Scale Telecommunication Networks and Services
abstract
Visual exploration of massive datasets arising from telecommunication networks and services is a challenge. This paper describes SWIFT-3D, an integrated data visualization and exploration system created at AT&T Labs for large scale network analysis. SWIFT-3D integrates a collection of interactive tools that includes pixel-oriented 2D maps, interactive 3D maps, statistical displays, network topology diagrams and an interactive drill-down query interface. Example applications are described, demonstrating a successful application to analyze unexpected network events (high volumes of unanswered calls), and comparison of usage of an Internet service with voice network traffic and local access coverage.
Eleftherios Koutsofios, Stephen C. North, Russell Truscott, Daniel A. Keim
IEEE Visualization4
1999 Optimal Grid-Clustering: Towards Breaking the Curse of Dimensionality in High-Dimensional Clustering
Alexander Hinneburg, Daniel A. Keim
VLDB2
1998 Fast Nearest Neighbor Search in High-Dimensional Space
abstract
Similarity search in multimedia databases requires an efficient support of nearest neighbor search on a large set of high dimensional points as a basic operation for query processing. As recent theoretical results show, state of the art approaches to nearest neighbor search are not efficient in higher dimensions. In our new approach, we therefore precompute the result of any nearest neighbor search which corresponds to a computation of the voronoi cell of each data point. In a second step, we store the voronoi cells in an index structure efficient for high dimensional data spaces. As a result, nearest neighbor search corresponds to a simple point query on the index structure. Although our technique is based on a precomputation of the solution space, it is dynamic, i.e. it supports insertions of new data points. An extensive experimental evaluation of our technique demonstrates the high efficiency for uniformly distributed as well as real data. We obtained a significant reduction of the search time compared to nearest neighbor search in the X tree (up to a factor of 4).
Stefan Berchtold, Bernhard Ertl, Daniel A. Keim, Hans-Peter Kriegel, Thomas Seidl 0001
ICDE3
1998 An Efficient Approach to Clustering in Large Multimedia Databases with Noise
Alexander Hinneburg, Daniel A. Keim
KDD2
1998 High-Dimensional Index Structures, Database Support for Next Decade's Applications (Tutorial)
abstract
No abstract available.
Stefan Berchtold, Daniel A. Keim
SIGMOD Conference2
1998 The Gridfit algorithm: an efficient and effective approach to visualizing large amounts of spatial data
abstract
In a large number of applications, data is collected and referenced by their spatial locations. Visualizing large amounts of spatially referenced data on a limited-size screen display often results in poor visualizations due to the high degree of overplotting of neighboring datapoints. We introduce a new approach to visualizing large amounts of spatially referenced data. The basic idea is to intelligently use the unoccupied pixels of the display instead of overplotting data points. After formally describing the problem, we present two solutions which are based on: placing overlapping data points on the nearest unoccupied pixel; and shifting data points along a screen-filling curve (e.g., Hilbert-curve). We then develop a more sophisticated approach called Gridfit, which is based on a hierarchical partitioning of the data space. We evaluate all three approaches with respect to their efficiency and effectiveness and show the superiority of the Gridfit approach. For measuring the effectiveness, we not only present the resulting visualizations but also introduce mathematical effectiveness criteria measuring properties of the generated visualizations with respect to the original data such as distance- and position-preservation.
Daniel A. Keim, Annemarie Herrmann
IEEE Visualization1
1997 A Cost Model For Nearest Neighbor Search in High-Dimensional Data Space
abstract
In this paper, we present a new cost model for nearest neighbor search in high-dimensional data space. We first analyze different nearest neighbor algorithms, present a generalization of an algorithm which has been originally proposed for Quadtrees [13], and show that this algorithm is optimal. Then, we develop a cost model which- in contrast to previous models- takes boundary effects into account and therefore also works in high dimensions. The advantages of our model are in particular: Our model works for data sets with an arbitrary number of dimensions and an arbitrary number of data points, is applicable to different data distributions and index structures, and provides accurate estimates of the expected query execution time. To show the practical relevance and accuracy of our model, we perform a detailed analysis using synthetic and real data. The results of applying our model to
Stefan Berchtold, Christian Böhm 0001, Daniel A. Keim, Hans-Peter Kriegel
PODS3
1997 Fast Parallel Similarity Search in Multimedia Databases
abstract
Most similarity search techniques map the data objects into some high-dimensional feature space. The similarity search then corresponds to a nearest-neighbor search in the feature space which is computationally very intensive. In this paper, we present a new parallel method for fast nearest-neighbor search in high-dimensional feature spaces. The core problem of designing a parallel nearest-neighbor algorithm is to find an adequate distribution of the data onto the disks. Unfortunately, the known declustering methods to not perform well for high-dimensional nearest-neighbor search. In contrast, our method has been optimized based on the special properties of high-dimensional spaces and therefore provides a near-optimal distribution of the data items among the disks. The basic idea of our data declustering technique is to assign the buckets corresponding to different quadrants of the data space to different disks. We show that our technique - in contrast to other declustering methods - guarantees that all buckets corresponding to neighboring quadrants are assigned to different disks. We evaluate our method using large amounts of real data (up to 40 MBytes) and compare it with the best known data declustering method, the Hilbert curve. Our experiments show that our method provides an almost linear speed-up and a constant scale-up. Additionally, it outperforms the Hilbert approach by a factor of up to 5.
Stefan Berchtold, Christian Böhm 0001, Bernhard Braunmüller, Daniel A. Keim, Hans-Peter Kriegel
SIGMOD Conference4
1997 Using Extended Feature Objects for Partial Similarity Retrieval
Stefan Berchtold, Daniel A. Keim, Hans-Peter Kriegel
VLDB J.2
1996 Databases and Visualization
abstract
No abstract available.
Daniel A. Keim
SIGMOD Conference1
1996 The X-tree : An Index Structure for High-Dimensional Data
Stefan Berchtold, Daniel A. Keim, Hans-Peter Kriegel
VLDB2
1996 Visualization Techniques for Mining Large Databases: A Comparison
abstract
Visual data mining techniques have proven to be of high value in exploratory data analysis, and they also have a high potential for mining large databases. In this article, we describe and evaluate a new visualization-based approach to mining large databases. The basic idea of our visual data mining techniques is to represent as many data items as possible on the screen at the same time by mapping each data value to a pixel of the screen and arranging the pixels adequately. The major goal of this article is to evaluate our visual data mining techniques and to compare them to other well-known visualization techniques for multidimensional data: the parallel coordinate and stick-figure visualization techniques. For the evaluation of visual data mining techniques, the perception of data properties counts most, while the CPU time and the number of secondary storage accesses are only of secondary importance. In addition to testing the visualization techniques using real data, we developed a testing environment for database visualizations similar to the benchmark approach used for comparing the performance of database systems. The testing environment allows the generation of test data sets with predefined data characteristics which are important for comparing the perceptual abilities of visual data mining techniques.
Daniel A. Keim, Hans-Peter Kriegel
IEEE Trans. Knowl. Data Eng.1
1995 VisDB: A System for Visualizing Large Databases
abstract
The VisDB system developed at the University of Munich is a sophisticated tool for visualizing and analyzing large databases. The key idea of the VisDB system is to support the exploration of large databases by using the phenomenal abilities of the human vision system which is able to analyze visualizations of mid-size to large amounts of data very efficiently. The goal of the VisDB system is to provide visualizations of large portions of the database, allowing properties of the data and structure in the data to become perceptually apparent.
Daniel A. Keim, Hans-Peter Kriegel
SIGMOD Conference1
1995 Recursive Pattern: A Technique for Visualizing Very Large Amounts of Data
abstract
An important goal of visualization technology is to support the exploration and analysis of very large amounts of data. In this paper, we propose a new visualization technique called a 'recursive pattern', which has been developed for visualizing large amounts of multidimensional data. The technique is based on a generic recursive scheme which generalizes a wide range of pixel-oriented arrangements for displaying large data sets. By instantiating the technique with adequate data- and application-dependent parameters, the user may greatly influence the structure of the resulting visualizations. Since the technique uses one pixel for presenting each data value, the amount of data which can be displayed is only limited by the resolution of current display technology and by the limitations of human perceptibility. Beside describing the basic idea of the 'recursive pattern' technique, we provide several examples of useful parameter settings for the various recursion levels. We further show that our 'recursive pattern' technique is particularly advantageous for the large class of data sets which have a natural order according to one dimension (e.g. time series data). We demonstrate the usefulness of our technique by using a stock market application.
Daniel A. Keim, Mihael Ankerst, Hans-Peter Kriegel
IEEE Visualization1
1994 Query Translation Supporting the Migration of Legacy Databases into Cooperative Information Systems
Daniel A. Keim, Hans-Peter Kriegel, Andreas Miethsam
CoopIS1
1994 Supporting Data Mining of Large Databases by Visual Feedback Queries
abstract
Describes a query system that provides visual relevance feedback in querying large databases. The goal is to support the process of data mining by representing as many data items as possible on the display. By arranging and coloring the data items as pixels according to their relevance for the query, the user gets a visual impression of the resulting data set. Using an interactive query interface, the user may change the query dynamically and receives immediate feedback by the visual representation of the resulting data set. Furthermore, by using multiple windows for different parts of a complex query, the user gets visual feedback for each part of the query and, therefore, may easier understand the overall result. The system allows one to represent the largest amount of data that can be visualized on current display technology, provides valuable feedback in querying the database, and allows the user to find results which would otherwise remain hidden in the database.>
Daniel A. Keim, Hans-Peter Kriegel, Thomas Seidl 0001
ICDE1
1993 Object-Oriented Querying of Existing Relations Databases
Daniel A. Keim, Hans-Peter Kriegel, Andreas Miethsam
DEXA1
1993 Visual Feedback in Querying Large Databases
abstract
In this paper, we describe a database query system that provides visual relevance feedback in querying large databases. The goal of our system is to support the query specification process by using each pixel of the display to represent one data item of the database. By arranging and coloring the pixels according to their relevance for the query, the user gets a visual impression of the resulting data set. Using sliders for each condition of the query, the user may change the query dynamically and receives immediate feedback by the visual representation of the resulting data set. By using multiple windows for different parts of a complex query, the user gets visual feedback for each part of the query and, therefore, will easier understand the overall result. The system may be used to query any database that contains tens of thousands to millions of data items, but it is especially helpful to explore large data sets with an unknown distribution of values and to find the interesting hot spots in huge amounts of data. The direct feedback allows to visually display the influence of incremental query refinements and, therefore, allows a better, easier and faster query specification.>
Daniel A. Keim, Hans-Peter Kriegel, Thomas Seidl 0001
IEEE Visualization1
1992 Visual Query Specification in a Multimedia Database System
abstract
A visual interface for a multimedia database management system (MDBMS) is described. DBMS query languages are linear in syntax. Although natural language interfaces have been found to be useful, natural language is ambiguous and difficult to process. For queries on standard (relational) data, these difficulties can be avoided with the use of a visual, graphical interface to guide the user in specifying the query. For image and other media data which are ambiguous in nature, natural language processing, combined with direct graphical access to the domain knowledge, is used to interpret and evaluate the natural language query. The system fully supports graphical and image input/output in different formats. The combination of visual effect and natural language specification, the support of media data, and the allowance of incremental query specification simplify the process of query specification not only for image or multimedia databases but also for all databases.>
Daniel A. Keim, Vincent Y. Lum
IEEE Visualization1
1991 A Friendly and Intelligent Approach to Data Retrieval in a Multimedia DBMS
Daniel A. Keim, Kyung-Chang Kim, Vincent Y. Lum
DEXA1