VLDB 2026 Research / reviewers in the wild / expert
Mennatallah El-Assady
dblp:183/8957
· DBLP profile ↗
59ranked-venue papers
7as first author
45since 2021 · last 2026
0000-0001-8526-2613ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 12 · 11 since 2021Human-computer interaction and ubiquitous computing · 10 · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PleaSQLarify: Visual Pragmatic Repair for Natural Language Database QueryingabstractNatural language database interfaces broaden data access, yet they remain brittle under input ambiguity. Standard approaches often collapse uncertainty into a single query, offering little support for mismatches between user intent and system interpretation. We reframe this challenge through pragmatic inference: while users economize expressions, systems operate on priors over the action space that may not align with the users’. In this view, pragmatic repair—incremental clarification through minimal interaction—is a natural strategy for resolving underspecification. We present PleaSQLarify, which operationalizes pragmatic repair by structuring interaction around interpretable decision variables that enable efficient clarification1. A visual interface2 complements this by surfacing the action space for exploration, requesting user disambiguation, and making belief updates traceable across turns. In a study with twelve participants, PleaSQLarify helped users recognize alternative interpretations and efficiently resolve ambiguity. Our findings highlight pragmatic repair as a design principle that fosters effective user control in natural language interfaces. Robin Shing Moon Chan, Rita Sevastjanova, Mennatallah El-Assady |
CHI | 3 |
| 2026 | StepMIND: A Visual Framework for Stepwise, Multimodal, and Bidirectional Explanations of AI-Generated Data Analysis PipelineabstractArtificial intelligence (AI) enables users to generate data visualizations from natural language descriptions, lowering the barrier to data exploration. However, AI-generated visualizations often present only the final output, lacking transparency and limiting users’ ability to verify, interpret, or refine the results. To address this, we introduce StepMIND, a generalizable visual framework that enhances explainability and interactivity in AI-generated data analysis pipelines. StepMIND integrates four dimensions: (1) Stepwise Refinement, allowing users to engage in the AI decision process; (2) Multimodal Explanations, combining natural language, structured notation, direct manipulation, and content visualization for accessible interpretation; (3) Bidirectional Editing, enabling seamless updates across modalities; and (4) Familiar Interaction Models, such as code editor and spreadsheet-based manipulations, to support both technical and non-technical users. To demonstrate its utility, we apply StepMIND in STAGE, a case study system for AI-assisted data visualization. A within-subject user study (N=20) shows that STAGE significantly improves user confidence and trust, reduces cognitive load, and facilitates both exploratory and corrective refinements. Our findings further suggest that StepMIND can generalize to broader AI-assisted workflows, offering a visible and interactive approach to explainable AI. Yang Wu 0010, Yao Wan 0001, Mennatallah El-Assady, April Yi Wang |
IUI | 3 |
| 2026 | A unified intelligence-augmented framework for building audits via physical and virtual inspectionabstract• Unified design space for intelligence-augmented building audits workflow. • 5A design principles enable scalable, modular, and automated audit workflow. • R²PIVS pipeline integrates retrieval, capture, prediction, and visualization. • Building inventory tasks are automated through human-in-the-loop prediction. • Validated framework and application schema via expert evaluation. Building audits are essential for informed decision-making in maintenance, renovation, and end-of-life planning. However, current practices remain predominantly manual and time-consuming due to fragmented data, limiting resource-efficient management of existing building stock. This paper presents a unified, intelligence-augmented framework designed to enhance the efficiency and reliability of both physical and virtual building inspection workflows. Five core design principles of adaptability, accessibility, affordability, acceleration, and alignment are derived from a multi-phase formal analysis to guide the development of the R²PIVS pipeline, which transforms the existing audit process into six modules: retrieval, reality capture, prediction, interaction, visualization, and summarization. The framework leverages human-AI collaboration in key building inventory tasks, including geometry measurement, visual assessment, and hazard estimation, through interactive annotation, model refinement, and output validation. Findings from expert elicitation studies indicate that the proposed application schema is promising for improving the efficiency and scalability of existing workflows. By aligning machine learning capabilities with domain-specific requirements, this research lays the foundation for a human-in-the-loop building audit system that enables standardized inspection and inventory information management to support circular construction practices throughout the building life cycle. Pei-Yu Wu, Mennatallah El-Assady, Catherine De Wolf |
Expert Syst. Appl. | 2 |
| 2026 | Understanding Large Language Model Behaviors Through Interactive Counterfactual Generation and AnalysisabstractUnderstanding the behavior of large language models (LLMs) is crucial for ensuring their safe and reliable use. However, existing explainable AI (XAI) methods for LLMs primarily rely on word-level explanations, which are often computationally inefficient and misaligned with human reasoning processes. Moreover, these methods often treat explanation as a one-time output, overlooking its inherently interactive and iterative nature. In this paper, we present LLM Analyzer, an interactive visualization system that addresses these limitations by enabling intuitive and efficient exploration of LLM behaviors through counterfactual analysis. Our system features a novel algorithm that generates fluent and semantically meaningful counterfactuals via targeted removal and replacement operations at user-defined levels of granularity. These counterfactuals are used to compute feature attribution scores, which are then integrated with concrete examples in a table-based visualization, supporting dynamic analysis of model behavior. A user study with LLM practitioners and interviews with experts demonstrate the system's usability and effectiveness, emphasizing the importance of involving humans in the explanation process as active participants rather than passive recipients. Furui Cheng, Vilém Zouhar, Robin Shing Moon Chan, Daniel Fürst, Hendrik Strobelt, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Finding Needles in Document Haystacks: Augmenting Serendipitous Claim Retrieval Workflows
Moritz Dück, Steffen Holter, Robin Shing Moon Chan, Rita Sevastjanova, Mennatallah El-Assady |
CHI | 5 |
| 2025 | Reward Learning from Multiple Feedback TypesabstractLearning rewards from preference feedback has become an important tool in the alignment of agentic models. Preference-based feedback, often implemented as a binary comparison between multiple completions, is an established method to acquire large-scale human feedback. However, human feedback in other contexts is often much more diverse. Such diverse feedback can better support the goals of a human annotator, and the simultaneous use of multiple sources might be mutually informative for the learning process or carry type-dependent biases for the reward learning process.
Despite these potential benefits, learning from different feedback types has yet to be explored extensively.
In this paper, we bridge this gap by enabling experimentation and evaluating multi-type feedback in a wide set of environments. We present a process to generate high-quality simulated feedback of six different types. Then, we implement reward models and downstream RL training for all six feedback types.
Based on the simulated feedback, we investigate the use of types of feedback across ten RL environments and compare them to pure preference-based baselines. We show empirically that diverse types of feedback can be utilized and lead to strong reward modeling performance. This work is the first strong indicator of the potential of multi-type feedback for RLHF. Yannick Metz, András Geiszl, Raphaël Baur, Mennatallah El-Assady |
ICLR | 4 |
| 2025 | A Design Space for Intelligent Dialogue Augmentation
Robin Shing Moon Chan, Anne Marx, Alison Kim, Mennatallah El-Assady |
IUI | 4 |
| 2025 | DxHF: Providing High-Quality Human Feedback for LLM Alignment with Interactive DecompositionabstractHuman preferences are widely used to align large language models (LLMs) through methods such as reinforcement learning from human feedback (RLHF).However, the current user interfaces require annotators to compare text paragraphs, which is cognitively challenging when the texts are long or unfamiliar.This paper contributes by studying the decomposition principle as an approach to improving the quality of human feedback for LLM alignment.This approach breaks down the text into individual claims instead of directly comparing two long-form text responses.Based on the principle, we build a novel user interface DxHF.It enhances the comparison process by showing decomposed claims, visually encoding the relevance of claims to the conversation and linking similar claims.This allows users to skim through key information and identify differences for better and quicker judgment.Our technical evaluation shows evidence that decomposition generally improves feedback accuracy regarding the ground truth, particularly for users with uncertainty.A crowdsourcing study with 160 participants indicates that using DxHF improves feedback accuracy by an average of 5%, although it increases the average feedback time by 18 seconds.Notably, accuracy is significantly higher in situations where users have less certainty.The finding of the study highlights the potential of HCI as an effective method for improving human-AI alignment. Danqing Shi, Furui Cheng, Tino Weinkauf, Antti Oulasvirta, Mennatallah El-Assady |
UIST | 5 |
| 2025 | LayerFlow: Layer-wise Exploration of LLM Embeddings using Uncertainty-aware Interlinked ProjectionsabstractAbstract Large language models (LLMs) represent words through contextual word embeddings encoding different language properties like semantics and syntax. Understanding these properties is crucial, especially for researchers investigating language model capabilities, employing embeddings for tasks related to text similarity, or evaluating the reasons behind token importance as measured through attribution methods. Applications for embedding exploration frequently involve dimensionality reduction techniques, which reduce high‐dimensional vectors to two dimensions used as coordinates in a scatterplot. This data transformation step introduces uncertainty that can be propagated to the visual representation and influence users' interpretation of the data. To communicate such uncertainties, we present LayerFlow – a visual analytics workspace that displays embeddings in an interlinked projection design and communicates the transformation, representation, and interpretation uncertainty. In particular, to hint at potential data distortions and uncertainties, the workspace includes several visual components, such as convex hulls showing 2D and HD clusters, data point pairwise distances, cluster summaries, and projection quality metrics. We show the usability of the presented workspace through replication and expert case studies that highlight the need to communicate uncertainty through multiple visual components and different data perspectives. Rita Sevastjanova, Robin Gerling, Thilo Spinner, Mennatallah El-Assady |
Comput. Graph. Forum | 4 |
| 2025 | f-RecX: A framework for designing effective textual explanations in recommender systems' user interfacesabstractRecommender systems (RecSys) have become ubiquitous in users’ daily digital interactions, significantly influencing decision-making processes. As these systems grow in algorithmic complexity, effective explanations for non-expert users become essential to fostering understanding and trust. While academic research explores diverse explanation methods, commercial applications predominantly employ textual explanations due to their implementation efficiency and user familiarity. However, the effectiveness of these textual explanations is often compromised by suboptimal presentation within RecSys user interfaces (UIs), leading to reduced user engagement and comprehension. This issue is particularly relevant given the recent emergence of large language models (LLMs) for generating RecSys explanations. We introduce f-RecX, a conceptual framework for characterizing and designing effective textual explanations in RecSys UIs. Based on a two-phase methodology combining qualitative user studies and quantitative evaluations, f-RecX maps four input dimensions (Explanation Style, Goals, Domain Dynamics, and Recommender Systems Technique) to an output dimension focused on visual representation. The framework aims to enhance the’consumability’ of textual explanations by making them easier to locate and comprehend, and more valuable for non-expert users. We demonstrate f-RecX’s applicability through a usage scenario and analysis of existing RecSys UIs, offering valuable insights for enhancing explainability and user experience. • f-RecX uniquely bridges algorithmic and human-centered design by integrating four input dimensions (Explanation Style, Goals, Domain Dynamics, and Recommender Techniques) with visual presentation parameters, addressing the gap between explanation content generation and effective UI implementation. • The framework introduces the concept of explanation ”consumability” - how easily users can locate, understand, and derive value from explanations - providing empirically validated visual design guidelines including optimal font sizes (24-32px), positioning (top-left), and explanation length ( 20 words). • Through a comprehensive two-phase methodology combining empathy workshops and quantitative surveys, f-RecX provides design guidance that demonstrates how visual characteristics significantly impact explanation effectiveness, particularly relevant for LLM-generated explanations in commercial applications. Ibrahim Al Hazwani, Gabriela Morgenshtern, Mennatallah El-Assady, Jürgen Bernard |
Int. J. Hum. Comput. Stud. | 3 |
| 2025 | An Analysis of the Interplay and Mutual Benefits of Grounded Theory and VisualizationabstractGrounded theory (GT) is a research methodology that entails a systematic workflow for theory generation grounded on emergent data. In this article, we juxtapose GT workflows with typical workflows in visualization and visual analytics (VIS), unveiling the characteristics shared by these workflows. We explore the research landscape of VIS to study where GT is applied to generate VIS theories, explicitly as well as implicitly. We discuss "why" GT can potentially play a significant role in VIS. We outline a "how" methodology for conducting GT research in VIS, which addresses the need for theoretical advancement in VIS while benefiting from other methods and techniques in VIS. We illustrate this "how" methodology with a use case of adopting GT approaches in studying visualization guidelines. Alexandra Diehl, Alfie Abdul-Rahman, Benjamin Bach, Mennatallah El-Assady, Matthias Kraus 0002, Robert S. Laramee, Daniel A. Keim, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | 2024 VGTC Visualization Significant New Researcher AwardabstractThe 2024 VGTC Visualization Significant New Researcher Award goes to Dominik Moritz for impactful work in the development of software infrastructure that has influenced both research and practice. The 2024 VGTC Visualization Significant New Researcher Award goes to Mennatallah El-Assady in recognition of her ground-breaking work at the intersection of visualization and machine learning. Dominik Moritz, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | ProvenanceWidgets: A Library of UI Control Elements to Track and Dynamically Overlay Analytic ProvenanceabstractWe present ProvenanceWidgets, a Javascript library of UI control elements such as radio buttons, checkboxes, and dropdowns to track and dynamically overlay a user's analytic provenance. These in situ overlays not only save screen space but also minimize the amount of time and effort needed to access the same information from elsewhere in the UI. In this paper, we discuss how we design modular UI control elements to track how often and how recently a user interacts with them and design visual overlays showing an aggregated summary as well as a detailed temporal history. We demonstrate the capability of ProvenanceWidgets by recreating three prior widget libraries: (1) Scented Widgets, (2) Phosphor objects, and (3) Dynamic Query Widgets. We also evaluated its expressiveness and conducted case studies with visualization developers to evaluate its effectiveness. We find that ProvenanceWidgets enables developers to implement custom provenance-tracking applications effectively. ProvenanceWidgets is available as open-source software at https://github.com/ProvenanceWidgets to help application developers build custom provenance-based systems. Arpit Narechania, Kaustubh Odak, Mennatallah El-Assady, Alex Endert |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | A Wizard of Oz Study of Guidance Strategies and DynamicsabstractCo-adaptive guidance in visual analytics is a mixed-initiative process in which both the user and the system work together to support each other in solving a given analysis task. While previous studies show the effectiveness of guidance, the impact of guidance design decisions (e.g., the suggestions' timing, contextualization, or adaptation) and misguidance often remain under-investigated. To investigate these aspects and examine a variety of guidance interaction patterns in a realistic analysis scenario, we present a Wizard of Oz (WOz) study setup in which pairs of participants take the user's and the system's roles, respectively. As users perform their analysis tasks, they are observed by wizards who provide just-in-time guidance as they see fit. Moreover, we designed the study so that wizards would occasionally and unknowingly provide misguidance during the analysis to investigate the users' confidence in guidance systems. We recruited two groups of participants (12 wizards and 12 users) and paired each participant with two from the other group, obtaining 48 observations. We report insights on interactions between users and wizards. By analyzing these interaction dynamics and the guidance strategies the wizards apply, we derive recommendations for implementing and evaluating future co-adaptive guidance systems. Fabian Sperrle, Mennatallah El-Assady, Alessio Arleo, Davide Ceneda |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | RELIC: Investigating Large Language Model Responses using Self-ConsistencyabstractLarge Language Models (LLMs) are notorious for blending fact with fiction and generating non-factual content, known as hallucinations. To address this challenge, we propose an interactive system that helps users gain insight into the reliability of the generated text. Our approach is based on the idea that the self-consistency of multiple samples generated by the same LLM relates to its confidence in individual claims in the generated texts. Using this idea, we design RELIC, an interactive system that enables users to investigate and verify semantic-level variations in multiple long-form responses. This allows users to recognize potentially inaccurate information in the generated text and make necessary corrections. From a user study with ten participants, we demonstrate that our approach helps users better verify the reliability of the generated text. We further summarize the design implications and lessons learned from this research for future studies of reliable human-LLM interactions. Furui Cheng, Vilém Zouhar, Simran Arora, Mrinmaya Sachan, Hendrik Strobelt, Mennatallah El-Assady |
CHI | 6 |
| 2024 | On Affine Homotopy between Language EncodersabstractPre-trained language encoders---functions that represent text as vectors---are an integral component of many NLP tasks.
We tackle a natural question in language encoder analysis: What does it mean for two encoders to be similar?
We contend that a faithful measure of similarity needs to be \emph{intrinsic}, that is, task-independent, yet still be informative of \emph{extrinsic} similarity---the performance on downstream tasks.
It is common to consider two encoders similar if they are \emph{homotopic}, i.e., if they can be aligned through some transformation.
In this spirit, we study the properties of \emph{affine} alignment of language encoders and its implications on extrinsic similarity.
We find that while affine alignment is fundamentally an asymmetric notion of similarity, it is still informative of extrinsic similarity.
We confirm this on datasets of natural language representations.
Beyond providing useful bounds on extrinsic similarity, affine intrinsic similarity also allows us to begin uncovering the structure of the space of pre-trained encoders by defining an order over them. Robin Shing Moon Chan, Reda Boumasmoud, Anej Svete, Qipeng Guo, Zhijing Jin 0001, Shauli Ravfogel, Mrinmaya Sachan, Bernhard Schölkopf, Mennatallah El-Assady, Ryan Cotterell |
NeurIPS | 10 |
| 2024 | Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and MetricsabstractExplainable AI (XAI) is a rapidly growing domain with a myriad of proposed methods as well as metrics aiming to evaluate their efficacy. However, current studies are often of limited scope, examining only a handful of XAI methods and ignoring underlying design parameters for performance, such as the model architecture or the nature of input data. Moreover, they often rely on one or a few metrics and neglect thorough validation, increasing the risk of selection bias and ignoring discrepancies among metrics. These shortcomings leave practitioners confused about which method to choose for their problem. In response, we introduce LATEC, a large-scale benchmark that critically evaluates 17 prominent XAI methods using 20 distinct metrics. We systematically incorporate vital design parameters like varied architectures and diverse input modalities, resulting in 7,560 examined combinations. Through LATEC, we showcase the high risk of conflicting metrics leading to unreliable rankings and consequently propose a more robust evaluation scheme. Further, we comprehensively evaluate various XAI methods to assist practitioners in selecting appropriate methods aligning with their needs. Curiously, the emerging top-performing method, Expected Gradients, is not examined in any relevant related study. LATEC reinforces its role in future XAI research by publicly releasing all 326k saliency maps and 378k metric scores as a (meta-)evaluation dataset. The benchmark is hosted at: https://github.com/IML-DKFZ/latec. Lukas Klein, Carsten T. Lüth, Udo Schlegel, Till J. Bungert, Mennatallah El-Assady, Paul F. Jaeger |
NeurIPS | 5 |
| 2024 | PowerGraph: A power grid benchmark dataset for graph neural networksabstractPower grids are critical infrastructures of paramount importance to modern society and, therefore, engineered to operate under diverse conditions and failures. The ongoing energy transition poses new challenges for the decision-makers and system operators. Therefore, we must develop grid analysis algorithms to ensure reliable operations. These key tools include power flow analysis and system security analysis, both needed for effective operational and strategic planning. The literature review shows a growing trend of machine learning (ML) models that perform these analyses effectively. In particular, Graph Neural Networks (GNNs) stand out in such applications because of the graph-based structure of power grids. However, there is a lack of publicly available graph datasets for training and benchmarking ML models in electrical power grid applications. First, we present PowerGraph, which comprises GNN-tailored datasets for i) power flows, ii) optimal power flows, and iii) cascading failure analyses of power grids. Second, we provide ground-truth explanations for the cascading failure analysis. Finally, we perform a complete benchmarking of GNN methods for node-level and graph-level tasks and explainability. Overall, PowerGraph is a multifaceted GNN dataset for diverse tasks that includes power flow and fault scenarios with real-world explanations, providing a valuable resource for developing improved GNN models for node-level, graph-level tasks and explainability methods in power system modeling. The dataset is available at https://figshare.com/articles/dataset/PowerGraph/22820534 and the code at https://github.com/PowerGraph-Datasets. Anna Varbella, Kenza Amara, Blazhe Gjorgiev, Mennatallah El-Assady, Giovanni Sansavini |
NeurIPS | 4 |
| 2024 | Trust Junk and Evil Knobs: Calibrating Trust in AI VisualizationabstractMany papers make claims about specific visualization techniques that are said to enhance or calibrate trust in AI systems. But a design choice that enhances trust in some cases appears to damage it in others. In this paper, we explore this inherent duality through an analogy with "knobs". Turning a knob too far in one direction may result in under-trust, too far in the other, over-trust or, turned up further still, in a confusing distortion. While the designs or so-called "knobs" are not inherently evil, they can be misused or used in an adversarial context and thereby manipulated to mislead users or promote unwarranted levels of trust in AI systems. When a visualization that has no meaningful connection with the underlying model or data is employed to enhance trust, we refer to the result as "trust junk." From a review of 65 papers, we identify nine commonly made claims about trust calibration. We synthesize them into a framework of knobs that can be used for good or "evil," and distill our findings into observed pitfalls for the responsible design of human-AI systems. Emily Wall 0001, Laura E. Matzen, Mennatallah El-Assady, Peta Masters, Helia Hosseinpour, Alex Endert, Rita Borgo, Polo Chau, Adam Perer, Harald T. Schupp, Hendrik Strobelt, Lace M. K. Padilla |
PacificVis | 3 |
| 2024 | Deconstructing Human-AI Collaboration: Agency, Interaction, and AdaptationabstractAbstract As full AI‐based automation remains out of reach in most real‐world applications, the focus has instead shifted to leveraging the strengths of both human and AI agents, creating effective collaborative systems. The rapid advances in this area have yielded increasingly more complex systems and frameworks, while the nuance of their characterization has gotten more vague. Similarly, the existing conceptual models no longer capture the elaborate processes of these systems nor describe the entire scope of their collaboration paradigms. In this paper, we propose a new unified set of dimensions through which to analyze and describe human‐AI systems. Our conceptual model is centered around three high‐level aspects ‐ agency, interaction, and adaptation ‐ and is developed through a multi‐step process. Firstly, an initial design space is proposed by surveying the literature and consolidating existing definitions and conceptual frameworks. Secondly, this model is iteratively refined and validated by conducting semi‐structured interviews with nine researchers in this field. Lastly, to illustrate the applicability of our design space, we utilize it to provide a structured description of selected human‐AI systems. Steffen Holter, Mennatallah El-Assady |
Comput. Graph. Forum | 2 |
| 2024 | -generAItor: Tree-in-the-loop Text Generation for Language Model Explainability and AdaptationabstractLarge language models (LLMs) are widely deployed in various downstream tasks, e.g., auto-completion, aided writing, or chat-based text generation. However, the considered output candidates of the underlying search algorithm are under-explored and under-explained. We tackle this shortcoming by proposing a tree-in-the-loop approach, where a visual representation of the beam search tree is the central component for analyzing, explaining, and adapting the generated outputs. To support these tasks, we present generAItor, a visual analytics technique, augmenting the central beam search tree with various task-specific widgets, providing targeted visualizations and interaction possibilities. Our approach allows interactions on multiple levels and offers an iterative pipeline that encompasses generating, exploring, and comparing output candidates, as well as fine-tuning the model based on adapted data. Our case study shows that our tool generates new insights in gender bias analysis beyond state-of-the-art template-based methods. Additionally, we demonstrate the applicability of our approach in a qualitative user study. Finally, we quantitatively evaluate the adaptability of the model to few samples, as occurring in text-generation use cases. Thilo Spinner, Rebecca Kehlbeck, Rita Sevastjanova, Tobias Stähle, Daniel A. Keim, Oliver Deussen, Mennatallah El-Assady |
ACM Trans. Interact. Intell. Syst. | 7 |
| 2024 | A Heuristic Approach for Dual Expert/End-User Evaluation of Guidance in Visual AnalyticsabstractGuidance can support users during the exploration and analysis of complex data. Previous research focused on characterizing the theoretical aspects of guidance in visual analytics and implementing guidance in different scenarios. However, the evaluation of guidance-enhanced visual analytics solutions remains an open research question. We tackle this question by introducing and validating a practical evaluation methodology for guidance in visual analytics. We identify eight quality criteria to be fulfilled and collect expert feedback on their validity. To facilitate actual evaluation studies, we derive two sets of heuristics. The first set targets heuristic evaluations conducted by expert evaluators. The second set facilitates end-user studies where participants actually use a guidance-enhanced system. By following such a dual approach, the different quality criteria of guidance can be examined from two different perspectives, enhancing the overall value of evaluation studies. To test the practical utility of our methodology, we employ it in two studies to gain insight into the quality of two guidance-enhanced visual analytics solutions, one being a work-in-progress research prototype, and the other being a publicly available visualization recommender system. Based on these two evaluations, we derive good practices for conducting evaluations of guidance in visual analytics and identify pitfalls to be avoided during such studies. Davide Ceneda, Christopher Collins 0001, Mennatallah El-Assady, Silvia Miksch, Christian Tominski, Alessio Arleo |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | FS/DS: A Theoretical Framework for the Dual Analysis of Feature Space and Data SpaceabstractWith the surge of data-driven analysis techniques, there is a rising demand for enhancing the exploration of large high-dimensional data by enabling interactions for the joint analysis of features (i.e., dimensions). Such a dual analysis of the feature space and data space is characterized by three components, 1) a view visualizing feature summaries, 2) a view that visualizes the data records, and 3) a bidirectional linking of both plots triggered by human interaction in one of both visualizations, e.g., Linking & Brushing. Dual analysis approaches span many domains, e.g., medicine, crime analysis, and biology. The proposed solutions encapsulate various techniques, such as feature selection or statistical analysis. However, each approach establishes a new definition of dual analysis. To address this gap, we systematically reviewed published dual analysis methods to investigate and formalize the key elements, such as the techniques used to visualize the feature space and data space, as well as the interaction between both spaces. From the information elicited during our review, we propose a unified theoretical framework for dual analysis, encompassing all existing approaches extending the field. We apply our proposed formalization describing the interactions between each component and relate them to the addressed tasks. Additionally, we categorize the existing approaches using our framework and derive future research directions to advance dual analysis by including state-of-the-art visual analysis techniques to improve data exploration. Frederik L. Dennig, Matthias Miller, Daniel A. Keim, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Personalized Language Model Selection Through Gamified Elicitation of Contrastive Concept PreferencesabstractLanguage models are widely used for different Natural Language Processing tasks while suffering from a lack of personalization. Personalization can be achieved by, e.g., fine-tuning the model on training data that is created by the user (e.g., social media posts). Previous work shows that the acquisition of such data can be challenging. Instead of adapting the model's parameters, we thus suggest selecting a model that matches the user's mental model of different thematic concepts in language. In this article, we attempt to capture such individual language understanding of users. In this process, two challenges have to be considered. First, we need to counteract disengagement since the task of communicating one's language understanding typically encompasses repetitive and time-consuming actions. Second, we need to enable users to externalize their mental models in different contexts, considering that language use changes depending on the environment. In this article, we integrate methods of gamification into a visual analytics (VA) workflow to engage users in sharing their knowledge within various contexts. In particular, we contribute the design of a gameful VA playground called Concept Universe. During the four-phased game, the users build personalized concept descriptions by explaining given concept names through representative keywords. Based on their performance, the system reacts with constant visual, verbal, and auditory feedback. We evaluate the system in a user study with six participants, showing that users are engaged and provide more specific input when facing a virtual opponent. We use the generated concepts to make personalized language model suggestions. Rita Sevastjanova, Hanna Hauptmann, Sebastian Deterding, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | A Diachronic Perspective on User Trust in AI under UncertaintyabstractIn a human-AI collaboration, users build a mental model of the AI system based on its reliability and how it presents its decision, e.g. its presentation of system confidence and an explanation of the output.Modern NLP systems are often uncalibrated, resulting in confidently incorrect predictions that undermine user trust.In order to build trustworthy AI, we must understand how user trust is developed and how it can be regained after potential trust-eroding events.We study the evolution of user trust in response to these trust-eroding events using a betting game.We find that even a few incorrect instances with inaccurate confidence estimates damage user trust and performance, with very slow recovery.We also show that this degradation in trust reduces the success of human-AI collaboration and that different types of miscalibration-unconfidently correct and confidently incorrect-have different negative effects on user trust.Our findings highlight the importance of calibration in user-facing AI applications and shed light on what aspects help users decide whether to trust the AI system. Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, Mrinmaya Sachan |
EMNLP | 3 |
| 2023 | Explore, Compare, and Predict Investment Opportunities through What-If Analysis: US Housing Market InvestigationabstractA key challenge in data analysis tools for domain-specific applications with high-dimensional time series data is to provide an intuitive way for users to explore their datasets, analyze trends and understand the models developed for these applications through human-computer interaction. To address this challenge, we propose a three-stage workflow that allows domain experts to explore their data, compare the different entities’ features, and predict the variable’s long-term trend using what-if analyses. Based on this workflow, we created a data visualization workspace for real estate investment using data from the US housing market at state and city level. The underlying machine learning model ARIMAX uses house price data together with socio-economic data from 2000 to 2021 to learn the dependencies of the house prices on the socio-economic factors and make informative and robust predictions for future years. Hongruyu Chen, Fernando Gonzalez, Oto Mraz, Sophia Kuhn, Cristina Guzman, Mennatallah El-Assady |
VINCI | 6 |
| 2023 | VISITOR: Visual Interactive State Sequence Exploration for Reinforcement LearningabstractAbstract Understanding the behavior of deep reinforcement learning agents is a crucial requirement throughout their development. Existing work has addressed the identification of observable behavioral patterns in state sequences or analysis of isolated internal representations; however, the overall decision‐making of deep‐learning RL agents remains opaque. To tackle this, we present VISITOR, a visual analytics system enabling the analysis of entire state sequences, the diagnosis of singular predictions, and the comparison between agents. A sequence embedding view enables the multiscale analysis of state sequences, utilizing custom embedding techniques for a stable spatialization of the observations and internal states. We provide multiple layers: (1) a state space embedding, highlighting different groups of states inside the state‐action sequences, (2) a trajectory view, emphasizing decision points, (3) a network activation mapping, visualizing the relationship between observations and network activations, (4) a transition embedding, enabling the analysis of state‐to‐state transitions. The embedding view is accompanied by an interactive reward view that captures the temporal development of metrics, which can be linked directly to states in the embedding. Lastly, a model list allows for the quick comparison of models across multiple metrics. Annotations can be exported to communicate results to different audiences. Our two‐stage evaluation with eight experts confirms the effectiveness in identifying states of interest, comparing the quality of policies, and reasoning about the internal decision‐making processes. Yannick Metz, Eugene Bykovets, Lucas Joos, Daniel A. Keim, Mennatallah El-Assady |
Comput. Graph. Forum | 5 |
| 2023 | Doom or Deliciousness: Challenges and Opportunities for Visualization in the Age of Generative ModelsabstractGenerative text-to-image models (as exemplified by DALL-E, MidJourney, and Stable Diffusion) have recently made enormous technological leaps, demonstrating impressive results in many graphical domains-from logo design to digital painting to photographic composition. However, the quality of these results has led to existential crises in some fields of art, leading to questions about the role of human agency in the production of meaning in a graphical context. Such issues are central to visualization, and while these generative models have yet to be widely applied in visualization, it seems only a matter of time until their integration is manifest. Seeking to circumvent similar ponderous dilemmas, we attempt to understand the roles that generative models might play across visualization. We do so by constructing a framework that characterizes what these technologies offer at various stages of the visualization workflow, augmented and analyzed through semi-structured interviews with 21 experts from related domains. Through this work, we map the space of opportunities and risks that might arise in this intersection, identifying doomsday prophecies and delicious low-hanging fruits that are ripe for research. Victor Schetinger, Sara Di Bartolomeo, Mennatallah El-Assady, Andrew M. McNutt, Matthias Miller, J. P. A. Passos, Jane Lydia Adams |
Comput. Graph. Forum | 3 |
| 2023 | Visual Analytics of Co-Occurrences to Discover Subspaces in Structured DataabstractWe present an approach that shows all relevant subspaces of categorical data condensed in a single picture. We model the categorical values of the attributes as co-occurrences with data partitions generated from structured data using pattern mining. We show that these co-occurrences are a-priori , allowing us to greatly reduce the search space, effectively generating the condensed picture where conventional approaches filter out several subspaces as these are deemed insignificant. The task of identifying interesting subspaces is common but difficult due to exponential search spaces and the curse of dimensionality. One application of such a task might be identifying a cohort of patients defined by attributes such as gender, age, and diabetes type that share a common patient history, which is modeled as event sequences. Filtering the data by these attributes is common but cumbersome and often does not allow a comparison of subspaces. We contribute a powerful multi-dimensional pattern exploration approach (MDPE-approach) agnostic to the structured data type that models multiple attributes and their characteristics as co-occurrences, allowing the user to identify and compare thousands of subspaces of interest in a single picture. In our MDPE-approach, we introduce two methods to dramatically reduce the search space, outputting only the boundaries of the search space in the form of two tables. We implement the MDPE-approach in an interactive visual interface (MDPE-vis) that provides a scalable, pixel-based visualization design allowing the identification, comparison, and sense-making of subspaces in structured data. Our case studies using a gold-standard dataset and external domain experts confirm our approach’s and implementation’s applicability. A third use case sheds light on the scalability of our approach and a user study with 15 participants underlines its usefulness and power. Wolfgang Jentner, Giuliana Lindholz, Hanna Hauptmann, Mennatallah El-Assady, Kwan-Liu Ma, Daniel A. Keim |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2023 | Visual Comparison of Language Model AdaptationabstractNeural language models are widely used; however, their model parameters often need to be adapted to the specific domains and tasks of an application, which is time- and resource-consuming. Thus, adapters have recently been introduced as a lightweight alternative for model adaptation. They consist of a small set of task-specific parameters with a reduced training time and simple parameter composition. The simplicity of adapter training and composition comes along with new challenges, such as maintaining an overview of adapter properties and effectively comparing their produced embedding spaces. To help developers overcome these challenges, we provide a twofold contribution. First, in close collaboration with NLP researchers, we conducted a requirement analysis for an approach supporting adapter evaluation and detected, among others, the need for both intrinsic (i.e., embedding similarity-based) and extrinsic (i.e., prediction-based) explanation methods. Second, motivated by the gathered requirements, we designed a flexible visual analytics workspace that enables the comparison of adapter properties. In this paper, we discuss several design iterations and alternatives for interactive, comparative visual explanation methods. Our comparative visualizations show the differences in the adapted embedding vectors and prediction outcomes for diverse human-interpretable concepts (e.g., person names, human qualities). We evaluate our workspace through case studies and show that, for instance, an adapter trained on the language debiasing task according to context-0 (decontextualized) embeddings introduces a new type of bias where words (even gender-independent words such as countries) become more similar to female- than male pronouns. We demonstrate that these are artifacts of context-0 embeddings, and the adapter effectively eliminates the gender information from the contextualized word representations. Rita Sevastjanova, Eren Cakmak, Shauli Ravfogel, Ryan Cotterell, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Lotse: A Practical Framework for Guidance in Visual AnalyticsabstractCo-adaptive guidance aims to enable efficient human-machine collaboration in visual analytics, as proposed by multiple theoretical frameworks. This paper bridges the gap between such conceptual frameworks and practical implementation by introducing an accessible model of guidance and an accompanying guidance library, mapping theory into practice. We contribute a model of system-provided guidance based on design templates and derived strategies. We instantiate the model in a library called Lotse that allows specifying guidance strategies in definition files and generates running code from them. Lotse is the first guidance library using such an approach. It supports the creation of reusable guidance strategies to retrofit existing applications with guidance and fosters the creation of general guidance strategy patterns. We demonstrate its effectiveness through first-use case studies with VA researchers of varying guidance design expertise and find that they are able to effectively and quickly implement guidance with Lotse. Further, we analyze our framework's cognitive dimensions to evaluate its expressiveness and outline a summary of open research questions for aligning guidance practice with its intricate theory. Fabian Sperrle, Davide Ceneda, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Automatic Generation of Socratic Subquestions for Teaching Math Word ProblemsabstractSocratic questioning is an educational method that allows students to discover answers to complex problems by asking them a series of thoughtful questions.Generation of didactically sound questions is challenging, requiring understanding of the reasoning process involved in the problem.We hypothesize that such questioning strategy can not only enhance the human performance, but also assist the math word problem (MWP) solvers.In this work, we explore the ability of large language models (LMs) in generating sequential questions for guiding math word problem-solving.We propose various guided question generation schemes based on input conditioning and reinforcement learning.On both automatic and human quality evaluations, we find that LMs constrained with desirable question properties generate superior questions and improve the overall performance of a math word problem solver.We conduct a preliminary user study to examine the potential value of such question generation models in the education domain.Results suggest that the difficulty level of problems plays an important role in determining whether questioning improves or hinders human performance.We discuss the future of using such questioning strategies in education.https://github.com/eth-nlped/ scaffolding-generation Kumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha, Manu Kapur, Mrinmaya Sachan |
EMNLP | 3 |
| 2022 | Augmenting Digital Sheet Music through Visual AnalyticsabstractAbstract Music analysis tasks, such as structure identification and modulation detection, are tedious when performed manually due to the complexity of the common music notation (CMN). Fully automated analysis instead misses human intuition about relevance. Existing approaches use abstract data‐driven visualizations to assist music analysis but lack a suitable connection to the CMN. Therefore, music analysts often prefer to remain in their familiar context. Our approach enhances the traditional analysis workflow by complementing CMN with interactive visualization entities as minimally intrusive augmentations. Gradual step‐wise transitions empower analysts to retrace and comprehend the relationship between the CMN and abstract data representations. We leverage glyph‐based visualizations for harmony, rhythm and melody to demonstrate our technique's applicability. Design‐driven visual query filters enable analysts to investigate statistical and semantic patterns on various abstraction levels. We conducted pair analytics sessions with 16 participants of different proficiency levels to gather qualitative feedback about the intuitiveness, traceability and understandability of our approach. The results show that MusicVis supports music analysts in getting new insights about feature characteristics while increasing their engagement and willingness to explore. Matthias Miller, Daniel Fürst, Hanna Hauptmann, Daniel A. Keim, Mennatallah El-Assady |
Comput. Graph. Forum | 5 |
| 2022 | CorpusVis: Visual Analysis of Digital Sheet Music CollectionsabstractAbstract Manually investigating sheet music collections is challenging for music analysts due to the magnitude and complexity of underlying features, structures, and contextual information. However, applying sophisticated algorithmic methods would require advanced technical expertise that analysts do not necessarily have. Bridging this gap, we contribute CorpusVis, an interactive visual workspace, enabling scalable and multi‐faceted analysis. Our proposed visual analytics dashboard provides access to computational methods, generating varying perspectives on the same data. The proposed application uses metadata including composers, type, epoch, and low‐level features, such as pitch, melody, and rhythm. To evaluate our approach, we conducted a pair‐analytics study with nine participants. The qualitative results show that CorpusVis supports users in performing exploratory and confirmatory analysis, leading them to new insights and findings. In addition, based on three exemplary workflows, we demonstrate how to apply our approach to different tasks, such as exploring musical features or comparing composers. Matthias Miller, Julius Rauscher, Daniel A. Keim, Mennatallah El-Assady |
Comput. Graph. Forum | 4 |
| 2022 | A Typology of Guidance Tasks in Mixed-Initiative Visual Analytics EnvironmentsabstractAbstract Guidance has been proposed as a conceptual framework to understand how mixed‐initiative visual analytics approaches can actively support users as they solve analytical tasks. While user tasks received a fair share of attention, it is still not completely clear how they could be supported with guidance and how such support could influence the progress of the task itself. Our observation is that there is a research gap in understanding the effect of guidance on the analytical discourse, in particular, for the knowledge generation in mixed‐initiative approaches. As a consequence, guidance in a visual analytics environment is usually indistinguishable from common visualization features, making user responses challenging to predict and measure. To address these issues, we take a system perspective to propose the notion of guidance tasks and we present it as a typology closely aligned to established user task typologies. We derived the proposed typology directly from a model of guidance in the knowledge generation process and illustrate its implications for guidance design. By discussing three case studies, we show how our typology can be applied to analyze existing guidance systems. We argue that without a clear consideration of the system perspective, the analysis of tasks in mixed‐initiative approaches is incomplete. Finally, by analyzing matchings of user and guidance tasks, we describe how guidance tasks could either help the user conclude the analysis or change its course. Ignacio Pérez-Messina, Davide Ceneda, Mennatallah El-Assady, Silvia Miksch, Fabian Sperrle |
Comput. Graph. Forum | 3 |
| 2022 | LMFingerprints: Visual Explanations of Language Model Embedding Spaces through Layerwise Contextualization ScoresabstractAbstract Language models, such as BERT, construct multiple, contextualized embeddings for each word occurrence in a corpus. Understanding how the contextualization propagates through the model's layers is crucial for deciding which layers to use for a specific analysis task. Currently, most embedding spaces are explained by probing classifiers; however, some findings remain inconclusive. In this paper, we present LMFingerprints, a novel scoring‐based technique for the explanation of contextualized word embeddings. We introduce two categories of scoring functions, which measure (1) the degree of contextualization, i.e., the layerwise changes in the embedding vectors, and (2) the type of contextualization, i.e., the captured context information. We integrate these scores into an interactive explanation workspace. By combining visual and verbal elements, we provide an overview of contextualization in six popular transformer‐based language models. We evaluate hypotheses from the domain of computational linguistics, and our results not only confirm findings from related work but also reveal new aspects about the information captured in the embedding spaces. For instance, we show that while numbers are poorly contextualized, stopwords have an unexpected high contextualization in the models' upper layers, where their neighborhoods shift from similar functionality tokens to tokens that contribute to the meaning of the surrounding sentences. Rita Sevastjanova, Aikaterini-Lida Kalouli, Christin Beck, Hanna Hauptmann, Mennatallah El-Assady |
Comput. Graph. Forum | 5 |
| 2022 | VisInReport: Complementing Visual Discourse Analytics Through Personalized Insight ReportsabstractWe present VisInReport, a visual analytics tool that supports the manual analysis of discourse transcripts and generates reports based on user interaction. As an integral part of scholarly work in the social sciences and humanities, discourse analysis involves an aggregation of characteristics identified in the text, which, in turn, involves a prior identification of regions of particular interest. Manual data evaluation requires extensive effort, which can be a barrier to effective analysis. Our system addresses this challenge by augmenting the users' analysis with a set of automatically generated visualization layers. These layers enable the detection and exploration of relevant parts of the discussion supporting several tasks, such as topic modeling or question categorization. The system summarizes the extracted events visually and verbally, generating a content-rich insight into the data and the analysis process. During each analysis session, VisInReport builds a shareable report containing a curated selection of interactions and annotations generated by the analyst. We evaluate our approach on real-world datasets through a qualitative study with domain experts from political science, computer science, and linguistics. The results highlight the benefit of integrating the analysis and reporting processes through a visual analytics system, which supports the communication of results among collaborating researchers. Rita Sevastjanova, Mennatallah El-Assady, Adam Bradley, Christopher Collins 0001, Miriam Butt, Daniel A. Keim |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Task-Based Visual Interactive Modeling: Decision Trees and Rule-Based ClassifiersabstractVisual analytics enables the coupling of machine learning models and humans in a tightly integrated workflow, addressing various analysis tasks. Each task poses distinct demands to analysts and decision-makers. In this survey, we focus on one canonical technique for rule-based classification, namely decision tree classifiers. We provide an overview of available visualizations for decision trees with a focus on how visualizations differ with respect to 16 tasks. Further, we investigate the types of visual designs employed, and the quality measures presented. We find that (i) interactive visual analytics systems for classifier development offer a variety of visual designs, (ii) utilization tasks are sparsely covered, (iii) beyond classifier development, node-link diagrams are omnipresent, (iv) even systems designed for machine learning experts rarely feature visual representations of quality measures other than accuracy. In conclusion, we see a potential for integrating algorithmic techniques, mathematical quality measures, and tailored interactive visualizations to enable human experts to utilize their knowledge more effectively. Dirk Streeb, Yannick Metz, Udo Schlegel, Bruno Schneider, Mennatallah El-Assady, Hansjörg Neth, Min Chen 0001, Daniel A. Keim |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Explaining Contextualization in Language Models using Visual AnalyticsabstractRita Sevastjanova, Aikaterini-Lida Kalouli, Christin Beck, Hanna Schäfer, Mennatallah El-Assady. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Rita Sevastjanova, Aikaterini-Lida Kalouli, Christin Beck, Hanna Hauptmann, Mennatallah El-Assady |
ACL/IJCNLP (1) | 5 |
| 2021 | Co-adaptive visual data analysis and guidance processes
Fabian Sperrle, Astrik Jeitler, Jürgen Bernard, Daniel A. Keim, Mennatallah El-Assady |
Comput. Graph. | 5 |
| 2021 | CommAID: Visual Analytics for Communication Analysis through Interactive Dynamics ModelingabstractAbstract Communication consists of both meta‐information as well as content. Currently, the automated analysis of such data often focuses either on the network aspects via social network analysis or on the content, utilizing methods from text‐mining. However, the first category of approaches does not leverage the rich content information, while the latter ignores the conversation environment and the temporal evolution, as evident in the meta‐information. In contradiction to communication research, which stresses the importance of a holistic approach, both aspects are rarely applied simultaneously, and consequently, their combination has not yet received enough attention in automated analysis systems. In this work, we aim to address this challenge by discussing the difficulties and design decisions of such a path as well as contribute CommAID, a blueprint for a holistic strategy to communication analysis. It features an integrated visual analytics design to analyze communication networks through dynamics modeling, semantic pattern retrieval, and a user‐adaptable and problem‐specific machine learning‐based retrieval system. An interactive multi‐level matrix‐based visualization facilitates a focused analysis of both network and content using inline visuals supporting cross‐checks and reducing context switches. We evaluate our approach in both a case study and through formative evaluation with eight law enforcement experts using a real‐world communication corpus. Results show that our solution surpasses existing techniques in terms of integration level and applicability. With this contribution, we aim to pave the path for a more holistic approach to communication analysis. Maximilian T. Fischer, Daniel Seebacher, Rita Sevastjanova, Daniel A. Keim, Mennatallah El-Assady |
Comput. Graph. Forum | 5 |
| 2021 | A Survey of Human-Centered Evaluations in Human-Centered Machine LearningabstractAbstract Visual analytics systems integrate interactive visualizations and machine learning to enable expert users to solve complex analysis tasks. Applications combine techniques from various fields of research and are consequently not trivial to evaluate. The result is a lack of structure and comparability between evaluations. In this survey, we provide a comprehensive overview of evaluations in the field of human‐centered machine learning. We particularly focus on human‐related factors that influence trust, interpretability, and explainability. We analyze the evaluations presented in papers from top conferences and journals in information visualization and human‐computer interaction to provide a systematic review of their setup and findings. From this survey, we distill design dimensions for structured evaluations, identify evaluation gaps, and derive future research opportunities. Fabian Sperrle, Mennatallah El-Assady, Grace Guo 0001, Rita Borgo, Polo Chau, Alex Endert, Daniel A. Keim |
Comput. Graph. Forum | 2 |
| 2021 | Learning Contextualized User Preferences for Co-Adaptive Guidance in Mixed-Initiative Topic Model RefinementabstractAbstract Mixed‐initiative visual analytics systems support collaborative human‐machine decision‐making processes. However, many multi‐objective optimization tasks, such as topic model refinement, are highly subjective and context‐dependent. Hence, systems need to adapt their optimization suggestions throughout the interactive refinement process to provide efficient guidance. To tackle this challenge, we present a technique for learning context‐dependent user preferences and demonstrate its applicability to topic model refinement. We deploy agents with distinct associated optimization strategies that compete for the user's acceptance of their suggestions. To decide when to provide guidance, each agent maintains an intelligible, rule‐based classifier over context vectorizations that captures the development of quality metrics between distinct analysis states. By observing implicit and explicit user feedback, agents learn in which contexts to provide their specific guidance operation. An agent in topic model refinement might, for example, learn to react to declining model coherence by suggesting to split a topic. Our results confirm that the rules learned by agents capture contextual user preferences. Further, we show that the learned rules are transferable between similar datasets, avoiding common cold‐start problems and enabling a continuous refinement of agents across corpora. Fabian Sperrle, Hanna Hauptmann, Daniel A. Keim, Mennatallah El-Assady |
Comput. Graph. Forum | 4 |
| 2021 | QuestionComb: A Gamification Approach for the Visual Explanation of Linguistic Phenomena through Interactive LabelingabstractLinguistic insight in the form of high-level relationships and rules in text builds the basis of our understanding of language. However, the data-driven generation of such structures often lacks labeled resources that can be used as training data for supervised machine learning. The creation of such ground-truth data is a time-consuming process that often requires domain expertise to resolve text ambiguities and characterize linguistic phenomena. Furthermore, the creation and refinement of machine learning models is often challenging for linguists as the models are often complex, in-transparent, and difficult to understand. To tackle these challenges, we present a visual analytics technique for interactive data labeling that applies concepts from gamification and explainable Artificial Intelligence (XAI) to support complex classification tasks. The visual-interactive labeling interface promotes the creation of effective training data. Visual explanations of learned rules unveil the decisions of the machine learning model and support iterative and interactive optimization. The gamification-inspired design guides the user through the labeling process and provides feedback on the model performance. As an instance of the proposed technique, we present QuestionComb , a workspace tailored to the task of question classification (i.e., in information-seeking vs. non-information-seeking questions). Our evaluation studies confirm that gamification concepts are beneficial to engage users through continuous feedback, offering an effective visual analytics technique when combined with active learning and XAI. Rita Sevastjanova, Wolfgang Jentner, Fabian Sperrle, Rebecca Kehlbeck, Jürgen Bernard, Mennatallah El-Assady |
ACM Trans. Interact. Intell. Syst. | 6 |
| 2021 | Why Visualize? Untangling a Large Network of ArgumentsabstractVisualization has been deemed a useful technique by researchers and practitioners, alike, leaving a trail of arguments behind that reason why visualization works. In addition, examples of misleading usages of visualizations in information communication have occasionally been pointed out. Thus, to contribute to the fundamental understanding of our discipline, we require a comprehensive collection of arguments on "why visualize?" (or "why not?"), untangling the rationale behind positive and negative viewpoints. In this paper, we report a theoretical study to understand the underlying reasons of various arguments; their relationships (e.g., built-on, and conflict); and their respective dependencies on tasks, users, and data. We curated an argumentative network based on a collection of arguments from various fields, including information visualization, cognitive science, psychology, statistics, philosophy, and others. Our work proposes several categorizations for the arguments, and makes their relations explicit. We contribute the first comprehensive and systematic theoretical study of the arguments on visualization. Thereby, we provide a roadmap towards building a foundation for visualization theory and empirical research as well as for practical application in the critique and design of visualizations. In addition, we provide our argumentation network and argument collection online at https://whyvis.dbvis.de, supported by an interactive visualization. Dirk Streeb, Mennatallah El-Assady, Daniel A. Keim, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Towards visual debugging for multi-target time series classificationabstractMulti-target classification of multivariate time series data poses a challenge in many real-world applications (e.g., predictive maintenance). Machine learning methods, such as random forests and neural networks, support training these classifiers. However, the debugging and analysis of possible misclassifications remain challenging due to the often complex relations between targets, classes, and the multivariate time series data. We propose a model-agnostic visual debugging workflow for multi-target time series classification that enables the examination of relations between targets, partially correct predictions, potential confusions, and the classified time series data. The workflow, as well as the prototype, aims to foster an in-depth analysis of multi-target classification results to identify potential causes of mispredictions visually. We demonstrate the usefulness of the workflow in the field of predictive maintenance in a usage scenario to show how users can iteratively explore and identify critical classes, as well as, relationships between targets. Udo Schlegel, Eren Cakmak, Hiba Arnout, Mennatallah El-Assady, Daniela Oelke, Daniel A. Keim |
IUI | 4 |
| 2020 | v-plots: Designing Hybrid Charts for the Comparative Analysis of Data DistributionsabstractAbstract Comparing data distributions is a core focus in descriptive statistics, and part of most data analysis processes across disciplines. In particular, comparing distributions entails numerous tasks, ranging from identifying global distribution properties, comparing aggregated statistics (e.g., mean values), to the local inspection of single cases. While various specialized visualizations have been proposed (e.g., box plots, histograms, or violin plots), they are not usually designed to support more than a few tasks, unless they are combined. In this paper, we present the v‐plot designer; a technique for authoring custom hybrid charts, combining mirrored bar charts, difference encodings, and violin‐style plots. v‐plots are customizable and enable the simultaneous comparison of data distributions on global, local, and aggregation levels. Our system design is grounded in an expert survey that compares and evaluates 20 common visualization techniques to derive guidelines for the task‐driven selection of appropriate visualizations. This knowledge externalization step allowed us to develop a guiding wizard that can tailor v‐plots to individual tasks and particular distribution properties. Finally, we confirm the usefulness of our system design and the user‐guiding process by measuring the fitness for purpose and applicability in a second study with four domain and statistic experts. Michael Blumenschein, Luka J. Debbeler, Nadine C. Lages, Britta Renner, Daniel A. Keim, Mennatallah El-Assady |
Comput. Graph. Forum | 6 |
| 2020 | Semantic Concept Spaces: Guided Topic Model Refinement using Word-Embedding ProjectionsabstractWe present a framework that allows users to incorporate the semantics of their domain knowledge for topic model refinement while remaining model-agnostic. Our approach enables users to (1) understand the semantic space of the model, (2) identify regions of potential conflicts and problems, and (3) readjust the semantic relation of concepts based on their understanding, directly influencing the topic modeling. These tasks are supported by an interactive visual analytics workspace that uses word-embedding projections to define concept regions which can then be refined. The user-refined concepts are independent of a particular document collection and can be transferred to related corpora. All user interactions within the concept space directly affect the semantic relations of the underlying vector space model, which, in turn, change the topic modeling. In addition to direct manipulation, our system guides the users' decision-making process through recommended interactions that point out potential improvements. This targeted refinement aims at minimizing the feedback required for an efficient human-in-the-loop process. We confirm the improvements achieved through our approach in two user studies that show topic model quality improvements through our visual knowledge externalization and learning process. Mennatallah El-Assady, Rebecca Kehlbeck, Christopher Collins 0001, Daniel A. Keim, Oliver Deussen |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | explAIner: A Visual Analytics Framework for Interactive and Explainable Machine LearningabstractWe propose a framework for interactive and explainable machine learning that enables users to (1) understand machine learning models; (2) diagnose model limitations using different explainable AI methods; as well as (3) refine and optimize the models. Our framework combines an iterative XAI pipeline with eight global monitoring and steering mechanisms, including quality monitoring, provenance tracking, model comparison, and trust building. To operationalize the framework, we present explAIner, a visual analytics system for interactive and explainable machine learning that instantiates all phases of the suggested pipeline within the commonly used TensorBoard environment. We performed a user-study with nine participants across different expertise levels to examine their perception of our workflow and to collect suggestions to fill the gap between our system and framework. The evaluation confirms that our tightly integrated system leads to an informed machine learning process while disclosing opportunities for further extensions. Thilo Spinner, Udo Schlegel, Hanna Hauptmann, Mennatallah El-Assady |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Augmenting Music Sheets with Harmonic FingerprintsabstractCommon Music Notation (CMN) is the well-established foundation for the written communication of musical information, such as rhythm or harmony. CMN suffers from the complexity of its visual encoding and the need for extensive training to acquire proficiency and legibility. While alternative notations using additional visual variables (e.g., color to improve pitch identification) have been proposed, the community does not readily accept notation systems that vary widely from the CMN. Therefore, to support student musicians in understanding harmonic relationships, instead of replacing the CMN, we present a visualization technique that augments digital sheet music with a harmonic fingerprint glyph. Our design exploits the circle of fifths, a fundamental concept in music theory, as visual metaphor. By attaching such glyphs to each bar of a composition we provide additional information about the salient harmonic features available in a musical piece. We conducted a user study to analyze the performance of experts and non-experts in an identification and comparison task of recurring patterns. The evaluation shows that the harmonic fingerprint supports these tasks without the need for close-reading, as when compared to a not-annotated music sheet. Matthias Miller, Alexandra Bonnici, Mennatallah El-Assady |
DocEng | 3 |
| 2019 | Visual Analytics for Topic Model Optimization based on User-Steerable Speculative ExecutionabstractTo effectively assess the potential consequences of human interventions in model-driven analytics systems, we establish the concept of speculative execution as a visual analytics paradigm for creating user-steerable preview mechanisms. This paper presents an explainable, mixed-initiative topic modeling framework that integrates speculative execution into the algorithmic decisionmaking process. Our approach visualizes the model-space of our novel incremental hierarchical topic modeling algorithm, unveiling its inner-workings. We support the active incorporation of the user's domain knowledge in every step through explicit model manipulation interactions. In addition, users can initialize the model with expected topic seeds, the backbone priors. For a more targeted optimization, the modeling process automatically triggers a speculative execution of various optimization strategies, and requests feedback whenever the measured model quality deteriorates. Users compare the proposed optimizations to the current model state and preview their effect on the next model iterations, before applying one of them. This supervised human-in-the-loop process targets maximum improvement for minimum feedback and has proven to be effective in three independent studies that confirm topic model quality improvements. Mennatallah El-Assady, Fabian Sperrle, Oliver Deussen, Daniel A. Keim, Christopher Collins 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Bridging Text Visualization and Mining: A Task-Driven SurveyabstractVisual text analytics has recently emerged as one of the most prominent topics in both academic research and the commercial world. To provide an overview of the relevant techniques and analysis tasks, as well as the relationships between them, we comprehensively analyzed 263 visualization papers and 4,346 mining papers published between 1992-2017 in two fields: visualization and text mining. From the analysis, we derived around 300 concepts (visualization techniques, mining techniques, and analysis tasks) and built a taxonomy for each type of concept. The co-occurrence relationships between the concepts were also extracted. Our research can be used as a stepping-stone for other researchers to 1) understand a common set of concepts used in this research topic; 2) facilitate the exploration of the relationships between visualization techniques, mining techniques, and analysis tasks; 3) understand the current practice in developing visual text analytics tools; 4) seek potential research opportunities by narrowing the gulf between visualization and mining techniques based on the analysis tasks; and 5) analyze other interdisciplinary research areas in a similar way. We have also contributed a web-based visualization tool for analyzing and understanding research trends and opportunities in visual text analytics. Shixia Liu, Xiting Wang, Christopher Collins 0001, Wenwen Dou, Fang-Xin Ou-Yang, Mennatallah El-Assady, Liu Jiang, Daniel A. Keim |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | ADD-up: Visual Analytics for Augmented Deliberative DemocracyabstractWe demonstrate the first prototype of the ADD-up visual analytics system. The Augmented Deliberative Democracy (ADD-up) project aims to enhance public deliberations by providing argument analytics in real time. The system will ultimately take a stenographic feed of a public deliberation meeting, automatically extract the arguments therein and project visual analytics intended to improve the deliberative quality of the event. Brian Plüss, Mennatallah El-Assady, Fabian Sperrle, Valentin Gold, Katarzyna Budzynska, Annette Hautli-Janisz, Chris Reed 0001 |
COMMA | 2 |
| 2018 | Visual Text Analytics: Techniques for Linguistic Information VisualizationabstractVisual Text Analytics has been an active area of interdisciplinary research (http://textvis.lnu.se/). This interactive tutorial is designed to give attendees an introduction to the area of information visualization, with a focus on linguistic visualization. After an introduction to the basic principles of information visualization and visual analytics, this tutorial will give an overview of the broad spectrum of linguistic and text visualization techniques, as well as their application areas [3]. This will be followed by a hands-on session that will allow participants to design their own visualizations using tools (e.g., Tableau), libraries (e.g., d3.js), or applying sketching techniques [4]. Some sample datasets will be provided by the instructor. Besides general techniques, special access will be provided to use the VisArgue framework [1] for the analysis of selected datasets. Mennatallah El-Assady |
DocEng | 1 |
| 2018 | Quality Metrics for Information VisualizationabstractAbstract The visualization community has developed to date many intuitions and understandings of how to judge thequalityof views in visualizing data. The computation of a visualization's quality and usefulness ranges from measuring clutter and overlap, up to the existence and perception of specific (visual) patterns. This survey attempts to report, categorize and unify the diverse understandings and aims to establish a common vocabulary that will enable a wide audience to understand their differences and subtleties. For this purpose, we present a commonly applicable quality metric formalization that should detail and relate all constituting parts of a quality metric. We organize our corpus of reviewed research papers along the data types established in the information visualization community: multi‐ and high‐dimensional, relational, sequential, geospatial and text data. For each data type, we select the visualization subdomains in which quality metrics are an active research field and report their findings, reason on the underlying concepts, describe goals and outline the constraints and requirements. One central goal of this survey is to provide guidance on future research opportunities for the field and outline how different visualization communities could benefit from each other by applying or transferring knowledge to their respective subdomain. Additionally, we aim to motivate the visualization community to compare computed measures to the perception of humans. Michael Behrisch 0001, Michael Blumenschein, Lin Shao 0001, Mennatallah El-Assady, Johannes Fuchs 0001, Daniel Seebacher, Alexandra Diehl, Ulrik Brandes, Hanspeter Pfister, Tobias Schreck, Daniel Weiskopf, Daniel A. Keim |
Comput. Graph. Forum | 5 |
| 2018 | ThreadReconstructor: Modeling Reply-Chains to Untangle Conversational Text through Visual AnalyticsabstractAbstract We present ThreadReconstructor, a visual analytics approach for detecting and analyzing the implicit conversational structure of discussions, e.g., in political debates and forums. Our work is motivated by the need to reveal and understand single threads in massive online conversations and verbatim text transcripts. We combine supervised and unsupervised machine learning models to generate a basic structure that is enriched by user‐defined queries and rule‐based heuristics. Depending on the data and tasks, users can modify and create various reconstruction models that are presented and compared in the visualization interface. Our tool enables the exploration of the generated threaded structures and the analysis of the untangled reply‐chains, comparing different models and their agreement. To understand the inner‐workings of the models, we visualize their decision spaces, including all considered candidate relations. In addition to a quantitative evaluation, we report qualitative feedback from an expert user study with four forum moderators and one machine learning expert, showing the effectiveness of our approach. Mennatallah El-Assady, Rita Sevastjanova, Daniel A. Keim, Christopher Collins 0001 |
Comput. Graph. Forum | 1 |
| 2018 | Progressive Learning of Topic Modeling Parameters: A Visual Analytics FrameworkabstractTopic modeling algorithms are widely used to analyze the thematic composition of text corpora but remain difficult to interpret and adjust. Addressing these limitations, we present a modular visual analytics framework, tackling the understandability and adaptability of topic models through a user-driven reinforcement learning process which does not require a deep understanding of the underlying topic modeling algorithms. Given a document corpus, our approach initializes two algorithm configurations based on a parameter space analysis that enhances document separability. We abstract the model complexity in an interactive visual workspace for exploring the automatic matching results of two models, investigating topic summaries, analyzing parameter distributions, and reviewing documents. The main contribution of our work is an iterative decision-making technique in which users provide a document-based relevance feedback that allows the framework to converge to a user-endorsed topic distribution. We also report feedback from a two-stage study which shows that our technique results in topic model quality improvements on two independent measures. Mennatallah El-Assady, Rita Sevastjanova, Fabian Sperrle, Daniel A. Keim, Christopher Collins 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2017 | NEREx: Named-Entity Relationship Exploration in Multi-Party ConversationsabstractAbstract We present NEREx, an interactive visual analytics approach for the exploratory analysis of verbatim conversational transcripts. By revealing different perspectives on multi‐party conversations, NEREx gives an entry point for the analysis through high‐level overviews and provides mechanisms to form and verify hypotheses through linked detail‐views. Using a tailored named‐entity extraction, we abstract important entities into ten categories and extract their relations with a distance‐restricted entity‐relationship model. This model complies with the often ungrammatical structure of verbatim transcripts, relating two entities if they are present in the same sentence within a small distance window. Our tool enables the exploratory analysis of multi‐party conversations using several linked views that reveal thematic and temporal structures in the text. In addition to distant‐reading, we integrated close‐reading views for a text‐level investigation process. Beyond the exploratory and temporal analysis of conversations, NEREx helps users generate and validate hypotheses and perform comparative analyses of multiple conversations. We demonstrate the applicability of our approach on real‐world data from the 2016 U.S. Presidential Debates through a qualitative study with three domain experts from political science. Mennatallah El-Assady, Rita Sevastjanova, Bela Gipp, Daniel A. Keim, Christopher Collins 0001 |
Comput. Graph. Forum | 1 |
| 2016 | ConToVi: Multi-Party Conversation Exploration using Topic-Space ViewsabstractAbstract We introduce a novel visual analytics approach to analyze speaker behavior patterns in multi‐party conversations. We propose Topic‐Space Views to track the movement of speakers across the thematic landscape of a conversation. Our tool is designed to assist political science scholars in exploring the dynamics of a conversation over time to generate and prove hypotheses about speaker interactions and behavior patterns. Moreover, we introduce a glyph‐based representation for each speaker turn based on linguistic and statistical cues to abstract relevant text features. We present animated views for exploring the general behavior and interactions of speakers over time and interactive steady visualizations for the detailed analysis of a selection of speakers. Using a visual sedimentation metaphor we enable the analysts to track subtle changes in the flow of a conversation over time while keeping an overview of all past speaker turns. We evaluate our approach on real‐world datasets and the results have been insightful to our domain experts. Mennatallah El-Assady, Valentin Gold, Carmela Acevedo, Christopher Collins 0001, Daniel A. Keim |
Comput. Graph. Forum | 1 |