VLDB 2026 Research / reviewers in the wild / expert
Cagatay Turkay
dblp:45/7528
· DBLP profile ↗
42ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0001-6788-251XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KAVA-PM: Knowledge-assisted visual process miningabstractThis article aims to foster a collaborative environment between the visual analytics and process mining communities by bringing together analysis methods, techniques, and tools from the process mining and visual analytics domains to devise a new knowledge-assisted, human-in-the-loop approach to process mining. Building on recent advances in methods emphasizing the role of human knowledge in analysis, we introduce knowledge-assisted interactive visual process mining (KAVA-PM) as a framework where analysts’ tacit knowledge and the externalizations of this knowledge play a key role. To achieve this, we extend an established conceptual model of KAVA that combines interactive visualizations and automated methods to support a richer process mining analysis practice that has human experts and their knowledge at its core. The paper outlines the key components of KAVA-PM as a conceptual model and its relations, proposes key analytical patterns, and demonstrates the use and validity of the patterns through usage scenarios. We then present challenges and open problems which we validate through a survey with experts. We anticipate that along with the conceptual model, these challenges will bring the VA and PM communities together along a shared research agenda where the role of humans and their knowledge is better established. • Conceptual model adapted to process mining for distinguishing between tacit and explicit knowledge. • Key research challenges validated by the visual analytics and process mining communities through an international survey. Daniel Schuster 0001, Wolfgang Aigner, Chiara Di Francescomarino, Cagatay Turkay, Francesca Zerbato |
Inf. Syst. | 4 |
| 2025 | Recoverable Facial Identity Protection via Adaptive Makeup Transfer Adversarial AttacksabstractUnauthorised face recognition (FR) systems have posed significant threats to digital identity and privacy protection. To alleviate the risk of compromised identities, recent makeup transfer-based attack methods embed adversarial signals in order to confuse unauthorised FR systems. However, their major weakness is that they set up a fixed image unrelated to both the protected and the makeup reference images as the confusion identity, which in turn has a negative impact on both attack success rate and visual quality of transferred photos. In addition, the generated images cannot be recognised by authorised FR systems once attacks are triggered. To address these challenges, in this paper, we propose a Recoverable Makeup Transferred Generative Adversarial Network (RMT-GAN) which has the distinctive feature of improving its image-transfer quality by selecting a suitable transfer reference photo as the target identity. Moreover, our method offers a solution to recover the protected photos to their original counterparts that can be recognised by authorised systems. Experimental results demonstrate that our method provides significantly improved attack success rates while maintaining higher visual quality compared to state-of-the-art makeup transfer-based adversarial attack methods. Our code and supplementary materials are available on Github. Xiyao Liu 0001, Junxing Ma, Xinda Wang 0006, Qianyu Lin, Jian Zhang 0048, Gerald Schaefer, Cagatay Turkay, Hui Fang 0003 |
AAAI | 7 |
| 2025 | Domain Experience and Expertise in Explainable AI Applications: A Bearing Fault Diagnosis Case StudyabstractThe importance of human-centred explainable artificial intelligence (XAI) has been widely recognised, leading to a growing focus on users and practitioners during the explanation design and deployment processes. Previous studies have identified that users with different domain expertise may have diverse needs for explainability. However, current XAI research often conflates a user's practical experience with domain expertise, ignoring the distinctions between the two; generally, experience relates to acquiring skill and insight through active participation or observation, while domain expertise denotes a high level of (often highly local and/or specific) knowledge. This paper investigates the impact of users' practical experience and domain expertise on how AI recommendations are considered in a high-risk decision-making context, using the example of ball bearing fault diagnosis in the manufacturing sector. As an interdisciplinary team of human-computer interaction (HCI) researchers and mechanical engineers, we co-design an XAI-based simulated ball bearing fault diagnostic task. We conduct task-led interviews with several professionals, structured around three distinct decision processes, and use an innovative sketch-based exercise to gather data to demonstrate how their decision-making behaviours change under ML recommendations and AI explanations. Our results show that highly experienced and knowledgeable practitioners understand but rely less on the explanations, while those with high experience but low expertise are more easily misled. Practitioners with high expertise but low experience trust XAI but struggle to use the explanations effectively. Based on these observations, we reflect on our methods and argue for considering both domain expertise and practical experience when designing and deploying AI explanations. Zibin Zhao, Michael Castelle, Cagatay Turkay |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2025 | LightVA: Lightweight Visual Analytics With LLM Agent-Based Task Planning and ExecutionabstractVisual analytics (VA) requires analysts to iteratively propose analysis tasks based on observations and execute tasks by creating visualizations and interactive exploration to gain insights. This process demands skills in programming, data processing, and visualization tools, highlighting the need for a more intelligent, streamlined VA approach. Large language models (LLMs) have recently been developed as agents to handle various tasks with dynamic planning and tool-using capabilities, offering the potential to enhance the efficiency and versatility of VA. We propose LightVA, a lightweight VA framework that supports task decomposition, data analysis, and interactive exploration through human-agent collaboration. Our method is designed to help users progressively translate high-level analytical goals into low-level tasks, producing visualizations and deriving insights. Specifically, we introduce an LLM agent-based task planning and execution strategy, employing a recursive process involving a planner, executor, and controller. The planner is responsible for recommending and decomposing tasks, the executor handles task execution, including data analysis, visualization generation and multi-view composition, and the controller coordinates the interaction between the planner and executor. Building on the framework, we develop a system with a hybrid user interface that includes a task flow diagram for monitoring and managing the task planning process, a visualization panel for interactive data exploration, and a chat view for guiding the model through natural language instructions. We examine the effectiveness of our method through a usage scenario and an expert study. Yuheng Zhao, Linbing Xiang, Zifei Guo, Cagatay Turkay, Yu Zhang 0043, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | LEVA: Using Large Language Models to Enhance Visual AnalyticsabstractVisual analytics supports data analysis tasks within complex domain problems. However, due to the richness of data types, visual designs, and interaction designs, users need to recall and process a significant amount of information when they visually analyze data. These challenges emphasize the need for more intelligent visual analytics methods. Large language models have demonstrated the ability to interpret various forms of textual data, offering the potential to facilitate intelligent support for visual analytics. We propose LEVA, a framework that uses large language models to enhance users' VA workflows at multiple stages: onboarding, exploration, and summarization. To support onboarding, we use large language models to interpret visualization designs and view relationships based on system specifications. For exploration, we use large language models to recommend insights based on the analysis of system status and data to facilitate mixed-initiative exploration. For summarization, we present a selective reporting strategy to retrace analysis history through a stream visualization and generate insight reports with the help of large language models. We demonstrate how LEVA can be integrated into existing visual analytics systems. Two usage scenarios and a user study suggest that LEVA effectively aids users in conducting visual analytics. Yuheng Zhao, Yu Zhang 0043, Zekai Shao 0001, Cagatay Turkay, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | TransforLearn: Interactive Visual Tutorial for the Transformer ModelabstractThe widespread adoption of Transformers in deep learning, serving as the core framework for numerous large-scale language models, has sparked significant interest in understanding their underlying mechanisms. However, beginners face difficulties in comprehending and learning Transformers due to its complex structure and abstract data representation. We present TransforLearn, the first interactive visual tutorial designed for deep learning beginners and non-experts to comprehensively learn about Transformers. TransforLearn supports interactions for architecture-driven exploration and task-driven exploration, providing insight into different levels of model details and their working processes. It accommodates interactive views of each layer's operation and mathematical formula, helping users to understand the data flow of long text sequences. By altering the current decoder-based recursive prediction results and combining the downstream task abstractions, users can deeply explore model processes. Our user study revealed that the interactions of TransforLearn are positively received. We observe that TransforLearn facilitates users' accomplishment of study tasks and a grasp of key concepts in Transformer effectively. Zekai Shao 0001, Ziqin Luo, Haibo Hu 0002, Cagatay Turkay, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Visual Explanation for Open-Domain Question Answering With BERTabstractOpen-domain question answering (OpenQA) is an essential but challenging task in natural language processing that aims to answer questions in natural language formats on the basis of large-scale unstructured passages. Recent research has taken the performance of benchmark datasets to new heights, especially when these datasets are combined with techniques for machine reading comprehension based on Transformer models. However, as identified through our ongoing collaboration with domain experts and our review of literature, three key challenges limit their further improvement: (i) complex data with multiple long texts, (ii) complex model architecture with multiple modules, and (iii) semantically complex decision process. In this paper, we present VEQA, a visual analytics system that helps experts understand the decision reasons of OpenQA and provides insights into model improvement. The system summarizes the data flow within and between modules in the OpenQA model as the decision process takes place at the summary, instance and candidate levels. Specifically, it guides users through a summary visualization of dataset and module response to explore individual instances with a ranking visualization that incorporates context. Furthermore, VEQA supports fine-grained exploration of the decision flow within a single module through a comparative tree visualization. We demonstrate the effectiveness of VEQA in promoting interpretability and providing insights into model enhancement through a case study and expert evaluation. Zekai Shao 0001, Shuran Sun, Yuheng Zhao, Siyuan Wang 0025, Zhongyu Wei, Tao Gui, Cagatay Turkay, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | From Asymptomatics to Zombies: Visualization-Based Education of Disease Modeling for ChildrenabstractThroughout the COVID-19 pandemic, visualizations became commonplace in public communications to help people make sense of the world and the reasons behind government-imposed restrictions. Though the adult population were the main target of these messages, children were affected by restrictions through not being able to see friends and virtual schooling. However, through these daily models and visualizations, the pandemic response provided a way for children to understand what data scientists really do and provided new routes for engagement with STEM subjects. In this paper, we describe the development of an interactive and accessible visualization tool to be used in workshops for children to explain computational modeling of diseases, in particular COVID-19. We detail our design decisions based on approaches evidenced to be effective and engaging such as unplugged activities and interactivity. We share reflections and learnings from delivering these workshops to 140 children and assess their effectiveness. Graham Mcneill, Max Sondag, Stewart Powell, Phoebe Asplin, Cagatay Turkay, Faron Moller, Daniel Archambault |
CHI | 5 |
| 2023 | Visual Analytics for Phishing Scam Identification in Blockchain Transactions with Multiple Model ComparisonabstractThe phishing scam is a major kind of fraudulence in blockchain. And it has become an urgent issue to discern and prevent the fraudulent behaviors. However, the large-scale and dynamic nature of transaction network imposes great challenges on the identification and analysis. While there have been many sophisticated machine learning approaches providing predictive capability in terms of detecting such cases, they usually offer little insight into the essence of those behaviors and the occasion when phishing scam activities happen. Motivated by these shortcomings and bottlenecks, this paper proposes a suite of visual analytical methods for interpretable and explorable fraudulence identification in large-scale blockchain transaction networks, incorporating an anomaly detection model based on multiple feature extraction manners. In this paper, we adopt two types of graph embedding methods and variable derivation to generate features from transaction data. Then we use machine learning classification approaches to fit the three sets of features. Evaluations show that all kinds of features perform well in classification. Besides, we design an interactive visualization system displaying the transaction networks and classification models, which allows users better explore the data and understand the models. Furthermore, we demonstrate two cases through the visualization system to unearth fraudulent patterns and interpret classification results. Finally, we close with discussions for further improvements of our models and system. Zishu Qin, Zengfeng Huang, Haoyun Guo, Richen Liu, Cagatay Turkay, Siming Chen 0001 |
VINCI | 8 |
| 2023 | Dashboard Design PatternsabstractThis paper introduces design patterns for dashboards to inform dashboard design processes. Despite a growing number of public examples, case studies, and general guidelines there is surprisingly little design guidance for dashboards. Such guidance is necessary to inspire designs and discuss tradeoffs in, e.g., screenspace, interaction, or information shown. Based on a systematic review of 144 dashboards, we report on eight groups of design patterns that provide common solutions in dashboard design. We discuss combinations of these patterns in "dashboard genres" such as narrative, analytical, or embedded dashboard. We ran a 2-week dashboard design workshop with 23 participants of varying expertise working on their own data and dashboards. We discuss the application of patterns for the dashboard design processes, as well as general design tradeoffs and common challenges. Our work complements previous surveys and aims to support dashboard designers and researchers in co-creation, structured design decisions, as well as future user evaluations about dashboard design guidelines. Detailed pattern descriptions and workshop material can be found online: https://dashboarddesignpatterns.github.io. Benjamin Bach, Euan Freeman, Alfie Abdul-Rahman, Cagatay Turkay, Saiful Khan, Yulei Fan, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Development and Evaluation of Two Approaches of Visual Sensitivity Analysis to Support Epidemiological ModelingabstractComputational modeling is a commonly used technology in many scientific disciplines and has played a noticeable role in combating the COVID-19 pandemic. Modeling scientists conduct sensitivity analysis frequently to observe and monitor the behavior of a model during its development and deployment. The traditional algorithmic ranking of sensitivity of different parameters usually does not provide modeling scientists with sufficient information to understand the interactions between different parameters and model outputs, while modeling scientists need to observe a large number of model runs in order to gain actionable information for parameter optimization. To address the above challenge, we developed and compared two visual analytics approaches, namely: algorithm-centric and visualization-assisted, and visualization-centric and algorithm-assisted. We evaluated the two approaches based on a structured analysis of different tasks in visual sensitivity analysis as well as the feedback of domain experts. While the work was carried out in the context of epidemiological modeling, the two approaches developed in this work are directly applicable to a variety of modeling processes featuring time series outputs, and can be extended to work with models with other types of outputs. Erik Rydow, Rita Borgo, Hui Fang 0003, Thomas Torsney-Weir, Ben Swallow, Thibaud Porphyre, Cagatay Turkay, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | Foreword to the Special Section on Visual Analytics
Katerina Vrotsou, Cagatay Turkay |
Comput. Graph. | 2 |
| 2022 | Visual Analytics of Contact Tracing Policy Simulations During an Emergency ResponseabstractAbstract Epidemiologists use individual‐based models to (a) simulate disease spread over dynamic contact networks and (b) to investigate strategies to control the outbreak. These model simulations generate complex ‘infection maps’ of time‐varying transmission trees and patterns of spread. Conventional statistical analysis of outputs offers only limited interpretation. This paper presents a novel visual analytics approach for the inspection of infection maps along with their associated metadata, developed collaboratively over 16 months in an evolving emergency response situation. We introduce the concept of representative trees that summarize the many components of a time‐varying infection map while preserving the epidemiological characteristics of each individual transmission tree. We also present interactive visualization techniques for the quick assessment of different control policies. Through a series of case studies and a qualitative evaluation by epidemiologists, we demonstrate how our visualizations can help improve the development of epidemiological models and help interpret complex transmission patterns. Max Sondag, Cagatay Turkay, Kai Xu 0003, Louise Matthews, Sibylle Mohr, Daniel Archambault |
Comput. Graph. Forum | 2 |
| 2022 | Rapid Development of a Data Visualization Service in an Emergency ResponseabstractWe present the design and development of a data visualization service (RAMPVIS) in response to the urgent need to support epidemiological modeling workflows during the COVID-19 pandemic. Facing a set of demanding requirements and several practical challenges, our small team of volunteers had to rely on existing knowledge and components of services computing, while thinking on our feet in configuring services composition and adopting suitable approaches to services engineering. Through developing the RAMPVIS service, we have gained useful experience of ensuring conformation to services computing standards, enabling rapid development and early deployment, and facilitating effective and efficient maintenance and operation with limited resources. This experience can be valuable to the ongoing effort for combating the COVID-19 pandemic, and provides a blueprint for visualization service development when future needs for visual analytics arise during emergency response. Saiful Khan, Phong Hai Nguyen, Alfie Abdul-Rahman, Euan Freeman, Cagatay Turkay, Min Chen 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2022 | Words of Estimative Correlation: Studying Verbalizations of ScatterplotsabstractNatural language and visualization are being increasingly deployed together for supporting data analysis in different ways, from multimodal interaction to enriched data summaries and insights. Yet, researchers still lack systematic knowledge on how viewers verbalize their interpretations of visualizations, and how they interpret verbalizations of visualizations in such contexts. We describe two studies aimed at identifying characteristics of data and charts that are relevant in such tasks. The first study asks participants to verbalize what they see in scatterplots that depict various levels of correlations. The second study then asks participants to choose visualizations that match a given verbal description of correlation. We extract key concepts from responses, organize them in a taxonomy and analyze the categorized responses. We observe that participants use a wide range of vocabulary across all scatterplots, but particular concepts are preferred for higher levels of correlation. A comparison between the studies reveals the ambiguity of some of the concepts. We discuss how the results could inform the design of multimodal representations aligned with the data and analytical tasks, and present a research roadmap to deepen the understanding about visualizations and natural language. Rafael Henkin, Cagatay Turkay |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Propagating Visual Designs to Numerous Plots and DashboardsabstractIn the process of developing an infrastructure for providing visualization and visual analytics (VIS) tools to epidemiologists and modeling scientists, we encountered a technical challenge for applying a number of visual designs to numerous datasets rapidly and reliably with limited development resources. In this paper, we present a technical solution to address this challenge. Operationally, we separate the tasks of data management, visual designs, and plots and dashboard deployment in order to streamline the development workflow. Technically, we utilize: an ontology to bring datasets, visual designs, and deployable plots and dashboards under the same management framework; multi-criteria search and ranking algorithms for discovering potential datasets that match a visual design; and a purposely-design user interface for propagating each visual design to appropriate datasets (often in tens and hundreds) and quality-assuring the propagation before the deployment. This technical solution has been used in the development of the RAMPVIS infrastructure for supporting a consortium of epidemiologists and modeling scientists through visualization. Saiful Khan, Phong Hai Nguyen, Alfie Abdul-Rahman, Benjamin Bach, Min Chen 0001, Euan Freeman, Cagatay Turkay |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | Special Issue on Interactive Visual Analytics for Making Explainable and Accountable Decisionsabstractresearch-article Share on Special Issue on Interactive Visual Analytics for Making Explainable and Accountable Decisions Authors: Cagatay Turkay University of Warwick, Coventry, UK University of Warwick, Coventry, UKView Profile , Tatiana Von Landesberger University of Cologne and University of Rostock, Cologne, Germany University of Cologne and University of Rostock, Cologne, GermanyView Profile , Daniel Archambault Swansea University, Swansea, Wales, UK Swansea University, Swansea, Wales, UKView Profile , Shixia Liu Tsinghua University, Beijing, People’s Republic of China Tsinghua University, Beijing, People’s Republic of ChinaView Profile , Remco Chang Tufts University, Medford, USA Tufts University, Medford, USAView Profile Authors Info & Claims ACM Transactions on Interactive Intelligent SystemsVolume 11Issue 3-4December 2021 Article No.: 17pp 1–4https://doi.org/10.1145/3471903Online:03 September 2021Publication History 0citation187DownloadsMetricsTotal Citations0Total Downloads187Last 12 Months187Last 6 weeks20 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Cagatay Turkay, Tatiana von Landesberger, Daniel Archambault, Shixia Liu, Remco Chang |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2021 | Revisiting the Modifiable Areal Unit Problem in Deep Traffic Prediction with Visual AnalyticsabstractDeep learning methods are being increasingly used for urban traffic prediction where spatiotemporal traffic data is aggregated into sequentially organized matrices that are then fed into convolution-based residual neural networks. However, the widely known modifiable areal unit problem within such aggregation processes can lead to perturbations in the network inputs. This issue can significantly destabilize the feature embeddings and the predictions - rendering deep networks much less useful for the experts. This paper approaches this challenge by leveraging unit visualization techniques that enable the investigation of many-to-many relationships between dynamically varied multi-scalar aggregations of urban traffic data and neural network predictions. Through regular exchanges with a domain expert, we design and develop a visual analytics solution that integrates 1) a Bivariate Map equipped with an advanced bivariate colormap to simultaneously depict input traffic and prediction errors across space, 2) a Moran's I Scatterplot that provides local indicators of spatial association analysis, and 3) a Multi-scale Attribution View that arranges non-linear dot plots in a tree layout to promote model analysis and comparison across scales. We evaluate our approach through a series of case studies involving a real-world dataset of Shenzhen taxi trips, and through interviews with domain experts. We observe that geographical scale variations have important impact on prediction performances, and interactive visual exploration of dynamically varying inputs and outputs benefit experts in the development of deep traffic prediction models. Wei Zeng 0004, Chengqiao Lin, Juncong Lin, Jincheng Jiang, Jiazhi Xia, Cagatay Turkay, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | LDA Ensembles for Interactive Exploration and Categorization of BehaviorsabstractWe define behavior as a set of actions performed by some actor during a period of time. We consider the problem of analyzing a large collection of behaviors by multiple actors, more specifically, identifying typical behaviors and spotting anomalous behaviors. We propose an approach leveraging topic modeling techniques - LDA (Latent Dirichlet Allocation) Ensembles - to represent categories of typical behaviors by topics that are obtained through topic modeling a behavior collection. When such methods are applied to text in natural languages, the quality of the extracted topics are usually judged based on the semantic relatedness of the terms pertinent to the topics. This criterion, however, is not necessarily applicable to topics extracted from non-textual data, such as action sets, since relationships between actions may not be obvious. We have developed a suite of visual and interactive techniques supporting the construction of an appropriate combination of topics based on other criteria, such as distinctiveness and coverage of the behavior set. Two case studies on analyzing operation behaviors in the security management system and visiting behaviors in an amusement park, and the expert evaluation of the first case study demonstrate the effectiveness of our approach. Siming Chen 0001, Natalia V. Andrienko, Gennady L. Andrienko, Linara Adilova, Jérémie Barlet, Jörg Kindermann, Phong H. Nguyen, Olivier Thonnard, Cagatay Turkay |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2020 | Supporting Story Synthesis: Bridging the Gap between Visual Analytics and StorytellingabstractVisual analytics usually deals with complex data and uses sophisticated algorithmic, visual, and interactive techniques supporting the analysis. Findings and results of the analysis often need to be communicated to an audience that lacks visual analytics expertise. This requires analysis outcomes to be presented in simpler ways than that are typically used in visual analytics systems. However, not only analytical visualizations may be too complex for target audiences but also the information that needs to be presented. Analysis results may consist of multiple components, which may involve multiple heterogeneous facets. Hence, there exists a gap on the path from obtaining analysis findings to communicating them, within which two main challenges lie: information complexity and display complexity. We address this problem by proposing a general framework where data analysis and result presentation are linked by story synthesis, in which the analyst creates and organises story contents. Unlike previous research, where analytic findings are represented by stored display states, we treat findings as data constructs. We focus on selecting, assembling and organizing findings for further presentation rather than on tracking analysis history and enabling dual (i.e., explorative and communicative) use of data displays. In story synthesis, findings are selected, assembled, and arranged in meaningful layouts that take into account the structure of information and inherent properties of its components. We propose a workflow for applying the proposed conceptual framework in designing visual analytics systems and demonstrate the generality of the approach by applying it to two diverse domains, social media and movement analysis. Siming Chen 0001, Jie Li 0006, Gennady L. Andrienko, Natalia V. Andrienko, Yun Wang 0012, Phong H. Nguyen, Cagatay Turkay |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2020 | VASABI: Hierarchical User Profiles for Interactive Visual User Behaviour AnalyticsabstractUser behaviour analytics (UBA) systems offer sophisticated models that capture users' behaviour over time with an aim to identify fraudulent activities that do not match their profiles. Motivated by the challenges in the interpretation of UBA models, this paper presents a visual analytics approach to help analysts gain a comprehensive understanding of user behaviour at multiple levels, namely individual and group level. We take a user-centred approach to design a visual analytics framework supporting the analysis of collections of users and the numerous sessions of activities they conduct within digital applications. The framework is centred around the concept of hierarchical user profiles that are built based on features derived from sessions, as well as on user tasks extracted using a topic modelling approach to summarise and stratify user behaviour. We externalise a series of analysis goals and tasks, and evaluate our methods through use cases conducted with experts. We observe that with the aid of interactive visual hierarchical user profiles, analysts are able to conduct exploratory and investigative analysis effectively, and able to understand the characteristics of user behaviour to make informed decisions whilst evaluating suspicious users and activities. Phong H. Nguyen, Rafael Henkin, Siming Chen 0001, Natalia V. Andrienko, Gennady L. Andrienko, Olivier Thonnard, Cagatay Turkay |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2019 | Understanding User Behaviour through Action Sequences: From the Usual to the UnusualabstractAction sequences, where atomic user actions are represented in a labelled, timestamped form, are becoming a fundamental data asset in the inspection and monitoring of user behaviour in digital systems. Although the analysis of such sequences is highly critical to the investigation of activities in cyber security applications, existing solutions fail to provide a comprehensive understanding due to the complex semantic and temporal characteristics of these data. This paper presents a visual analytics approach that aims to facilitate a user-involved, multi-faceted decision making process during the identification and the investigation of "unusual" action sequences. We first report the results of the task analysis and domain characterisation process. Then we describe the components of our multi-level analysis approach that comprises of constraint-based sequential pattern mining and semantic distance based clustering, and multi-scalar visualisations of users and their sequences. Finally, we demonstrate the applicability of our approach through a case study that involves tasks requiring effective decision-making by a group of domain experts. Although our solution here is tightly informed by a user-centred, domain-focused design process, we present findings and techniques that are transferable to other applications where the analysis of such sequences is of interest. Phong H. Nguyen, Cagatay Turkay, Gennady L. Andrienko, Natalia V. Andrienko, Olivier Thonnard, Jihane Zouaoui |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | User Behavior Map: Visual Exploration for Cyber Security Session DataabstractUser behavior analysis is complex and especially crucial in the cyber security domain. Understanding dynamic and multi-variate user behavior are challenging. Traditional sequential and timeline based method cannot easily address the complexity of temporal and relational features of user behaviors. We propose a map-based visual metaphor and create an interactive map for encoding user behaviors. It enables analysts to explore and identify user behavior patterns and helps them to understand why some behaviors are regarded as anomalous. We experiment with a real dataset containing multiple user sessions, consisting of sequences of diverse types of actions. In the behavior map, we encode an action as a city and user sessions as trajectories going through the cities. The position of the cities is determined by the sequential and temporal relationship of actions. Spatial and temporal patterns on the map reflect behavior patterns in the action space. In the case study, we illustrate how we explore relationships between actions, identify patterns of the typical session and detect anomaly behaviors. Siming Chen 0001, Shuai Chen 0001, Natalia V. Andrienko, Gennady L. Andrienko, Phong H. Nguyen, Cagatay Turkay, Olivier Thonnard, Xiaoru Yuan |
VizSEC | 6 |
| 2018 | Hunting High and Low: Visualising Shifting Correlations in Financial MarketsabstractAbstract The analysis of financial assets’ correlations is fundamental to many aspects of finance theory and practice, especially modern portfolio theory and the study of risk. In order to manage investment risk, in‐depth analysis of changing correlations is needed, with both high and low correlations between financial assets (and groups thereof) important to identify. In this paper, we propose a visual analytics framework for the interactive analysis of relations and structures in dynamic, high‐dimensional correlation data. We conduct a series of interviews and review the financial correlation analysis literature to guide our design. Our solution combines concepts from multi‐dimensional scaling, weighted complete graphs and threshold networks to present interactive, animated displays which use proximity as a visual metaphor for correlation and animation stability to encode correlation stability. We devise interaction techniques coupled with context‐sensitive auxiliary views to support the analysis of subsets of correlation networks. As part of our contribution, we also present behaviour profiles to help guide future users of our approach. We evaluate our approach by checking the validity of the layouts produced, presenting a number of analysis stories, and through a user study. We observe that our solutions help unravel complex behaviours and resonate well with study participants in addressing their needs in the context of correlation analysis in finance. P. M. Simon, Cagatay Turkay |
Comput. Graph. Forum | 2 |
| 2017 | On the Challenges and Opportunities in Visualization for Machine Learning and Knowledge Extraction: A Research Agenda
Cagatay Turkay, Robert S. Laramee, Andreas Holzinger |
CD-MAKE | 1 |
| 2017 | The State of the Art in Integrating Machine Learning into Visual AnalyticsabstractAbstract Visual analytics systems combine machine learning or other analytic techniques with interactive data visualization to promote sensemaking and analytical reasoning. It is through such techniques that people can make sense of large, complex data. While progress has been made, the tactful combination of machine learning and data visualization is still under‐explored. This state‐of‐the‐art report presents a summary of the progress that has been made by highlighting and synthesizing select research advances. Further, it presents opportunities and challenges to enhance the synergy between machine learning and visual analytics for impactful future research directions. Alex Endert, William Ribarsky, Cagatay Turkay, B. L. William Wong, Ian T. Nabney, Ignacio Díaz Blanco, Fabrice Rossi |
Comput. Graph. Forum | 3 |
| 2017 | Supporting theoretically-grounded model building in the social sciences through interactive visualisation
Cagatay Turkay, Aidan Slingsby, Kaisa Lahtinen, Sarah Butt, Jason Dykes |
Neurocomputing | 1 |
| 2017 | Map LineUps: Effects of spatial structure on graphical inferenceabstractFundamental to the effective use of visualization as an analytic and descriptive tool is the assurance that presenting data visually provides the capability of making inferences from what we see. This paper explores two related approaches to quantifying the confidence we may have in making visual inferences from mapped geospatial data. We adapt Wickham et al.'s 'Visual Line-up' method as a direct analogy with Null Hypothesis Significance Testing (NHST) and propose a new approach for generating more credible spatial null hypotheses. Rather than using as a spatial null hypothesis the unrealistic assumption of complete spatial randomness, we propose spatially autocorrelated simulations as alternative nulls. We conduct a set of crowdsourced experiments (n=361) to determine the just noticeable difference (JND) between pairs of choropleth maps of geographic units controlling for spatial autocorrelation (Moran's I statistic) and geometric configuration (variance in spatial unit area). Results indicate that people's abilities to perceive differences in spatial autocorrelation vary with baseline autocorrelation structure and the geometric configuration of geographic units. These results allow us, for the first time, to construct a visual equivalent of statistical power for geospatial data. Our JND results add to those provided in recent years by Klippel et al. (2011), Harrison et al. (2014) and Kay & Heer (2015) for correlation visualization. Importantly, they provide an empirical basis for an improved construction of visual line-ups for maps and the development of theory to inform geospatial tests of graphical inference. Roger Beecham, Jason Dykes, Wouter Meulemans, Aidan Slingsby, Cagatay Turkay, Jo Wood |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2017 | Small Multiples with GapsabstractSmall multiples enable comparison by providing different views of a single data set in a dense and aligned manner. A common frame defines each view, which varies based upon values of a conditioning variable. An increasingly popular use of this technique is to project two-dimensional locations into a gridded space (e.g. grid maps), using the underlying distribution both as the conditioning variable and to determine the grid layout. Using whitespace in this layout has the potential to carry information, especially in a geographic context. Yet, the effects of doing so on the spatial properties of the original units are not understood. We explore the design space offered by such small multiples with gaps. We do so by constructing a comprehensive suite of metrics that capture properties of the layout used to arrange the small multiples for comparison (e.g. compactness and alignment) and the preservation of the original data (e.g. distance, topology and shape). We study these metrics in geographic data sets with varying properties and numbers of gaps. We use simulated annealing to optimize for each metric and measure the effects on the others. To explore these effects systematically, we take a new approach, developing a system to visualize this design space using a set of interactive matrices. We find that adding small amounts of whitespace to small multiple arrays improves some of the characteristics of 2D layouts, such as shape, distance and direction. This comes at the cost of other metrics, such as the retention of topology. Effects vary according to the input maps, with degree of variation in size of input regions found to be a factor. Optima exist for particular metrics in many cases, but at different amounts of whitespace for different maps. We suggest multiple metrics be used in optimized layouts, finding topology to be a primary factor in existing manually-crafted solutions, followed by a trade-off between shape and displacement. But the rich range of possible optimized layouts leads us to challenge single-solution thinking; we suggest to consider alternative optimized layouts for small multiples with gaps. Key to our work is the systematic, quantified and visual approach to exploring design spaces when facing a trade-off between many competing criteria-an approach likely to be of value to the analysis of other design spaces. Wouter Meulemans, Jason Dykes, Aidan Slingsby, Cagatay Turkay, Jo Wood |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Designing Progressive and Interactive Analytics Processes for High-Dimensional Data AnalysisabstractIn interactive data analysis processes, the dialogue between the human and the computer is the enabling mechanism that can lead to actionable observations about the phenomena being investigated. It is of paramount importance that this dialogue is not interrupted by slow computational mechanisms that do not consider any known temporal human-computer interaction characteristics that prioritize the perceptual and cognitive capabilities of the users. In cases where the analysis involves an integrated computational method, for instance to reduce the dimensionality of the data or to perform clustering, such non-optimal processes are often likely. To remedy this, progressive computations, where results are iteratively improved, are getting increasing interest in visual analytics. In this paper, we present techniques and design considerations to incorporate progressive methods within interactive analysis processes that involve high-dimensional data. We define methodologies to facilitate processes that adhere to the perceptual characteristics of users and describe how online algorithms can be incorporated within these. A set of design recommendations and according methods to support analysts in accomplishing high-dimensional data analysis tasks are then presented. Our arguments and decisions here are informed by observations gathered over a series of analysis sessions with analysts from finance. We document observations and recommendations from this study and present evidence on how our approach contribute to the efficiency and productivity of interactive visual analysis sessions involving high-dimensional data. Cagatay Turkay, Erdem Kaya, Selim Balcisoy, Helwig Hauser |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Enhancing a social science model-building workflow with interactive visualisation
Cagatay Turkay, Aidan Slingsby, Kaisa Lahtinen, Sarah Butt, Jason Dykes |
ESANN | 1 |
| 2016 | Faceted Views of Varying Emphasis (FaVVEs): a framework for visualising multi-perspective small multiplesabstractAbstract Many datasets have multiple perspectives – for example space, time and description – and often analysts are required to study these multiple perspectives concurrently. This concurrent analysis becomes difficult when data are grouped and split into small multiples for comparison. A design challenge is thus to provide representations that enable multiple perspectives, split into small multiples, to be viewed simultaneously in ways that neither clutter nor overload. We present a design framework that allows us to do this. We claim that multi‐perspective comparison across small multiples may be possible by superimposing perspectives on one another rather than juxtaposing those perspectives side‐by‐side. This approach defies conventional wisdom and likely results in visual and informational clutter. For this reason we propose designs at three levels of abstraction for each perspective. By flexibly varying the abstraction level, certain perspectives can be brought into, or out of, focus. We evaluate our framework through laboratory‐style user tests. We find that superimposing, rather than juxtaposing, perspective views has little effect on performance of a low‐level comparison task. We reflect on the user study and its design to further identify analysis situations for which our framework may be desirable. Although the user study findings were insufficiently discriminating, we believe our framework opens up a new design space for multi‐perspective visual analysis. Roger Beecham, Chris Rooney, S. Meier, Jason Dykes, Aidan Slingsby, Cagatay Turkay, Jo Wood, B. L. William Wong |
Comput. Graph. Forum | 6 |
| 2016 | Visualizing Multiple Variables Across Scale and GeographyabstractComparing multiple variables to select those that effectively characterize complex entities is important in a wide variety of domains - geodemographics for example. Identifying variables that correlate is a common practice to remove redundancy, but correlation varies across space, with scale and over time, and the frequently used global statistics hide potentially important differentiating local variation. For more comprehensive and robust insights into multivariate relations, these local correlations need to be assessed through various means of defining locality. We explore the geography of this issue, and use novel interactive visualization to identify interdependencies in multivariate data sets to support geographically informed multivariate analysis. We offer terminology for considering scale and locality, visual techniques for establishing the effects of scale on correlation and a theoretical framework through which variation in geographic correlation with scale and locality are addressed explicitly. Prototype software demonstrates how these contributions act together. These techniques enable multiple variables and their geographic characteristics to be considered concurrently as we extend visual parameter space analysis (vPSA) to the spatial domain. We find variable correlations to be sensitive to scale and geography to varying degrees in the context of energy-based geodemographics. This sensitivity depends upon the calculation of locality as well as the geographical and statistical structure of the variable. Sarah Goodwin, Jason Dykes, Aidan Slingsby, Cagatay Turkay |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2014 | Perceptually Uniform Motion SpaceabstractFlow data is often visualized by animated particles inserted into a flow field. The velocity of a particle on the screen is typically linearly scaled by the velocities in the data. However, the perception of velocity magnitude in animated particles is not necessarily linear. We present a study on how different parameters affect relative motion perception. We have investigated the impact of four parameters. The parameters consist of speed multiplier, direction, contrast type and the global velocity scale. In addition, we investigated if multiple motion cues, and point distribution, affect the speed estimation. Several studies were executed to investigate the impact of each parameter. In the initial results, we noticed trends in scale and multiplier. Using the trends for the significant parameters, we designed a compensation model, which adjusts the particle speed to compensate for the effect of the parameters. We then performed a second study to investigate the performance of the compensation model. From the second study we detected a constant estimation error, which we adjusted for in the last study. In addition, we connect our work to established theories in psychophysics by comparing our model to a model based on Stevens' Power Law. Åsmund Birkeland, Cagatay Turkay, Ivan Viola |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | Attribute Signatures: Dynamic Visual Summaries for Analyzing Multivariate Geographical DataabstractThe visual analysis of geographically referenced datasets with a large number of attributes is challenging due to the fact that the characteristics of the attributes are highly dependent upon the locations at which they are focussed, and the scale and time at which they are measured. Specialized interactive visual methods are required to help analysts in understanding the characteristics of the attributes when these multiple aspects are considered concurrently. Here, we develop attribute signatures-interactively crafted graphics that show the geographic variability of statistics of attributes through which the extent of dependency between the attributes and geography can be visually explored. We compute a number of statistical measures, which can also account for variations in time and scale, and use them as a basis for our visualizations. We then employ different graphical configurations to show and compare both continuous and discrete variation of location and scale. Our methods allow variation in multiple statistical summaries of multiple attributes to be considered concurrently and geographically, as evidenced by examples in which the census geography of London and the wider UK are explored. Cagatay Turkay, Aidan Slingsby, Helwig Hauser, Jo Wood, Jason Dykes |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2013 | Visual cavity analysis in molecular simulationsabstractMolecular surfaces provide a useful mean for analyzing interactions between biomolecules; such as identification and characterization of ligand binding sites to a host macromolecule. We present a novel technique, which extracts potential binding sites, represented by cavities, and characterize them by 3D graphs and by amino acids. The binding sites are extracted using an implicit function sampling and graph algorithms. We propose an advanced cavity exploration technique based on the graph parameters and associated amino acids. Additionally, we interactively visualize the graphs in the context of the molecular surface. We apply our method to the analysis of MD simulations of Proteinase 3, where we verify the previously described cavities and suggest a new potential cavity to be studied. Július Parulek, Cagatay Turkay, Nathalie Reuter, Ivan Viola |
BMC Bioinform. | 2 |
| 2012 | A Perceptual-Statistics Shading ModelabstractThe process of surface perception is complex and based on several influencing factors, e.g., shading, silhouettes, occluding contours, and top down cognition. The accuracy of surface perception can be measured and the influencing factors can be modified in order to decrease the error in perception. This paper presents a novel concept of how a perceptual evaluation of a visualization technique can contribute to its redesign with the aim of improving the match between the distal and the proximal stimulus. During analysis of data from previous perceptual studies, we observed that the slant of 3D surfaces visualized on 2D screens is systematically underestimated. The visible trends in the error allowed us to create a statistical model of the perceived surface slant. Based on this statistical model we obtained from user experiments, we derived a new shading model that uses adjusted surface normals and aims to reduce the error in slant perception. The result is a shape-enhancement of visualization which is driven by an experimentally-founded statistical model. To assess the efficiency of the statistical shading model, we repeated the evaluation experiment and confirmed that the error in perception was decreased. Results of both user experiments are publicly-available datasets. Veronika Soltészová, Cagatay Turkay, Mark C. Price, Ivan Viola |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | Representative Factor Generation for the Interactive Visual Analysis of High-Dimensional DataabstractDatasets with a large number of dimensions per data item (hundreds or more) are challenging both for computational and visual analysis. Moreover, these dimensions have different characteristics and relations that result in sub-groups and/or hierarchies over the set of dimensions. Such structures lead to heterogeneity within the dimensions. Although the consideration of these structures is crucial for the analysis, most of the available analysis methods discard the heterogeneous relations among the dimensions. In this paper, we introduce the construction and utilization of representative factors for the interactive visual analysis of structures in high-dimensional datasets. First, we present a selection of methods to investigate the sub-groups in the dimension set and associate representative factors with those groups of dimensions. Second, we introduce how these factors are included in the interactive visual analysis cycle together with the original dimensions. We then provide the steps of an analytical procedure that iteratively analyzes the datasets through the use of representative factors. We discuss how our methods improve the reliability and interpretability of the analysis process by enabling more informed selections of computational tools. Finally, we demonstrate our techniques on the analysis of brain imaging study results that are performed over a large group of subjects. Cagatay Turkay, Arvid Lundervold, Astri J. Lundervold, Helwig Hauser |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | Interactive Visual Analysis of Temporal Cluster StructuresabstractAbstract Cluster analysis is a useful method which reveals underlying structures and relations of items after grouping them into clusters. In the case of temporal data, clusters are defined over time intervals where they usually exhibit structural changes. Conventional cluster analysis does not provide sufficient methods to analyze these structural changes, which are, however, crucial in the interpretation and evaluation of temporal clusters. In this paper, we present two novel and interactive visualization techniques that enable users to explore and interpret the structural changes of temporal clusters. We introduce the temporal cluster view, which visualizes the structural quality of a number of temporal clusters, and temporal signatures, which represents the structure of clusters over time. We discuss how these views are utilized to understand the temporal evolution of clusters. We evaluate the proposed techniques in the cluster analysis of mixed lipid bilayers. Cagatay Turkay, Július Parulek, Nathalie Reuter, Helwig Hauser |
Comput. Graph. Forum | 1 |
| 2011 | Integrating Information Theory in Agent-Based Crowd Simulation Behavior ModelsabstractCrowds must be simulated believable in terms of their appearance and behavior to improve a virtual environment's realism. Due to the complex nature of human behavior, realistic behavior of agents in crowd simulations is still a challenging problem. In this paper, we propose a novel behavioral model which builds analytical maps to control agents’ behavior adaptively with agent–crowd interaction formulations. We introduce information theoretical concepts to construct analytical maps automatically. Our model can be integrated into crowd simulators and enhance their behavioral complexity. We made comparative analyses of the presented behavior model with measured crowd data and two agent-based crowd simulators. Cagatay Turkay, Emre Koc, Selim Balcisoy |
Comput. J. | 1 |
| 2011 | Brushing Dimensions - A Dual Visual Analysis Model for High-Dimensional DataabstractIn many application fields, data analysts have to deal with datasets that contain many expressions per item. The effective analysis of such multivariate datasets is dependent on the user's ability to understand both the intrinsic dimensionality of the dataset as well as the distribution of the dependent values with respect to the dimensions. In this paper, we propose a visualization model that enables the joint interactive visual analysis of multivariate datasets with respect to their dimensions as well as with respect to the actual data values. We describe a dual setting of visualization and interaction in items space and in dimensions space. The visualization of items is linked to the visualization of dimensions with brushing and focus+context visualization. With this approach, the user is able to jointly study the structure of the dimensions space as well as the distribution of data items with respect to the dimensions. Even though the proposed visualization model is general, we demonstrate its application in the context of a DNA microarray data analysis. Cagatay Turkay, Peter Filzmoser, Helwig Hauser |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2009 | An information theoretic approach to camera control for crowded scenes
Cagatay Turkay, Emre Koc, Selim Balcisoy |
Vis. Comput. | 1 |